4
5 Comments

Web Scraping Discussion

Curious to see if anyone is the community is doing advanced web scraping as part of their projects.

We current have solutions in place to scrape networks like linkedin, angellist, yelp, indeed, crunchbase, etc.

We have historically used 1 of 3 techniques to scrape website

  1. simple cURL based
  2. Chrome Extensions methods
  3. Headless Browsers

I'm curious if anyone has extensive experience in the Headless browser space. We are playing around with https://github.com/puppeteer/puppeteer recently to try and step our game up.

The most difficult networks we have come up against are Linkedin/Angellist so curious if anyone has experience here and is willing to discuss techniques.

on December 24, 2019
  1. 2

    Linkedin scraping is a battle you don't want to fight 😞(from experience scraping those sites you mentioned and several others using puppeteer and other similar solutions).
    It just takes so much effort and it's a constant fight against edge cases and intentional anti-scraping measures that many months will go by and you'll still be paddling uphill with limited RoI.

    Just my personal experience. Might not apply to you so take it with a grain of salt.

  2. 1

    Just thought I'd drop this here, we've just launched the ScrapeDiary community for anyone looking to learn or collaborate on web scraping. I've already been blown away by our members so far, all welcome, just sign up using this link;

    https://blog.scrapediary.com/community

  3. 1

    I have messed around with puppeteer on some property websites to send me some info to keep on eye on property prices. Even though the sites already offer this, I wanted something to get info from all of them at once and give me a single, no bullshit or adverts email.

    But again, this isn't "special sauce" in any way, and would be hard to monetize which seems to be the case for most web scraping solutions. It would cost more to run these scraping jobs then people would be willing to pay for them.

  4. 1

    I have been using Puppeteer is do large scale web scraping and I love it. Any specific questions you have?

  5. 1

    I did use a combination of Scrapy and Splash in the past to allow for rendering JavaScript and CSS, it worked quite well and I was fairly happy with the combo — Splash work as a micro service which also fit into my architecture at the time.