1
5 Comments

What is my best option? Realtime end user web (image) scraping in Python/Django

Hi everyone. Just found this site whilst looking for a solution to my problem and it seems great, will definitely stick around.

Essentially I am building a web app in Django that allows users to enter a URL, it should grab the image URLs (ideally based on size constraints) from that page and send the back to the site. I have been dipping in and out of this for months trying different solutions ranging from Selenium, BS4, PhantomJS and Chrome headless browsers - all of which were just too slow and too many sites were blocked - and also tried a couple of DaaS providers, but it's still not working consistently enough. Scraperapi.com was the closest I got to it working but it's being blocked by some pretty big sites. I also need to do all requests based on them being JS rendered sites as I have no way of knowing in advance, and its adding 10s onto the 5-8s it already takes to give the results.

Kit.com is a perfect example of how I would like this feature to work (when you try and add an item to your 'kit'). It is somehow consistently getting product/main images from any URL, and very quickly. The site is in Angular but I suspect they are using some high end scraping service.

I feel like a DaaS solution is where I need to look rather than coding it all myself, I do not have a lot of money to spend on a subscription at the start but as the site (hopefully) gets used the revenue could cover increased subscription fees.

Any pointers at all would be much appreciated, this is driving me crazy.

Cheers

  1. 1

    Have you tried proxycrawl.com ? I think someone posted about it in a similar thread and he was doing some millions of crawls per month successfully with them, so I guess it will work better than scraperapi. They also have javascript crawling so you should be able to do what you want to accomplish. I'll try to find the post.

    1. 1

      Found it! Here you have the post with some interesting info on how to crawl massive amounts of data: https://www.indiehackers.com/forum/online-businesses-based-on-scraping-e28025aeb8?commentId=-LEdBW0TUIY6gmNV2uMI

  2. 1

    This should certainly be possible with BeautifulSoup - maybe I misunderstood what you’re trying to do? What problems did you come across with BS4? I use it for https://string.click and it works a treat, although a much simpler use case.

    1. 1

      Hey - thanks for that link, that is great actually but still hits some of the same issues. For instance if you run this URL through it: http://www.gap.eu/browse/product.do?cid=1040054&pcid=57372&vid=1&pid=000296210001 you dont actually get the product images. The GAP site is rendered in JS and I assume string.click doesnt do JS rendering, so the response is just sort of default layout code. If you try this URL https://www.zara.com/us/en/bejeweled-blouse-with-gathered-details-p04387236.html?v1=5322962&v2=732008 it returns the images in "data:image/png;base64..." format, but if you view the source of the page in your browser the image paths are all .jpg - so something is up there as well.

      However this string.click site is at least not being blocked (both these sites were returning "forbidden" when I tried other services) so its a step in the right direction for me. Thanks!

      Also I guess i misspoke - I was still using Beautiful Soup to parse and filter the data once I have received it back (whether using the API of some DaaS, Selenium or the Requests library).

      1. 1

        Ah cool - yeah there are challenges around JS sites. Maybe something that involves headless Chrome? I imagine it could be quite involved. Good luck - sorry I can't be more helpful!