
Dataflow kit
Web Scraper API. SERP extractor. Visual point-&-click tool.
We are so excited to introduce new completely re-implemented Dataflow Kit. In particular, legacy custom web scraper has been supplemented by more focused and more understandable web services for our users.
Read more info in our blog post at https://blog.dataflowkit.com/reloaded/
Finally, #dataflowkit #kubernetes cluster has been deployed. Thanks to @digitalocean.
1 Like
4 Comments
4 Comments
Check out our new article on Medium at https://hackernoon.com/hacker-news-scraping-challenge-e0655479f85b
Like
Comment
it is not practical to keep a multi-gigabyte as a single JSON array. Taking into consideration that Dataflow kit users would require to store and parse big volumes of data we’ve implemented export to JSONL format.
Read more about JSON Lines in our medium blog post at https://medium.com/@slotix/json-lines-format-76353b4e588d
Like
Comment
Load and save existing payloads for data export with one click!
Like
Comment
Medium is possibly the best blogging platform out there right now. Read more info at https://medium.com/dataflow-kit
1 Like
Comment
MongoDB along with diskv and cassandra may be used to store intermediate scraped results.
Like
Comment
We used Splash from Scrapinghub as a Java Script Rendering service at. We’ve switched to Headless Chrome recently as is a really game changer in the scraping field.
Like
Comment
We were forced to refactor our services to meet requirement of processing of big volumes of data.
Dataflow kit engine is stable enough to process several Millions of pages from a specified website and generate result successfully.
Like
Comment
Started spreading information about Dataflow Kit among Github Open source community members https://github.com/slotix/dataflowkit
Like
Comment
About
We help people to extract data from web sites with a simple point and click UI and turn them to API.


Comment