1
4 Comments

I built a $217K MRR Bot, and now I am building a Bot Protection Test Service

10 years ago I built Price2Spy, a competitor price monitoring tool. Since it basically works as a very complex bot we gained a lot of experience in this area and I can say that we truly understand how bots actually work.

This led to an idea.
We created a bot that can monitor almost any website and is capable of avoiding many obstacles such as captchas, logins, location-aware websites, etc. Now, it’s time to allow everyone to test their websites and see how protected they are from bots. This is how BotMeNot was born.

How does it work?

In essence, BotmeNot will test your website's bot protection, by executing a Test.
Depending on the type of the Test BotMeNot will send a number of different bot requests to your website. While doing so, we will try to catch your content. After you get the results, and these results are bad, you should start working on the necessary improvements.

Want to learn more?

We just launched a landing page with some basic info here, and you can leave your email there so we can notify you very soon when a much more detailed website will be available.

What are your thoughts? Do you have any problems with bots on your website?

on July 13, 2021
  1. 2

    I think that scraping is almost impossible to prevent, if someone really want to do it. It can be slowed down limiting the number of page loads per x time, or number of connections per IP, or setting hidden ban-trigger links that shouldn't be "followed" etc, but it can't be prevented fully.

    You can "control" the bots that behave nicely and present themselves as such. Add user-agent or hostname or IP range to be redirected to 404 or blocked and you're golden.

    But bots can decide to not present them as bots and then average user can't do much in terms of preventing them to visit, crawl or scrape whatever they want.

    Even less advanced bots can set random visit times, load random pages and stay on them random number of seconds, click random elements and in general mimic the actions of regular visitors, all together with user-agent randomization, proxies, browser footprint and everything else in between.

    Web design is predictable and html elements are mostly "hard-coded" so it's hard to prevent scraping of specific areas of interest. It's dead easy to set scraping with xpath or simpledom or php / python scripts for any purpose. Even if you change class name or selector order it's a matter of seconds to adjust the script and scrape whatever.

    1. 1

      Yes, bots are getting smarter, but so are anti-bot defenses.

      If you cannot stop bots, you can make it harder for them (that is – your competitors will have to pay more for such bots)

      Btw, I do not agree that Web design is predictable – the more websites we monitor, the more unexpected cases we find.

      Gone are the days of plain old HTML – nowadays the content can be in all kinds of unexpected elements: JS, JSON requiring whole different scraping techniques.

      1. 1

        Yes, every site is different, different styling and classes and manual setup have to be done for each site, but content has to be seen in browser somehow, meaning - html. JSON, js, sqlite, databases can keep the content but html is still here to show it on website, as soon as text is on the screen it can be scraped. Sometimes even can't be copied manually (stupid js and css tricks) but xpath is finding them easily.
        I am not a pro like you so I don't know usage cases and statistics, but my understanding is that most of the people want to scrape something that is visual, on screen - headlines, anchor texts + links, prices, specific paragraph below the article, stock tickers, etc. and for that purpose even cheap commercial tools are more than enough to do the job - WebHarvy, Screaming Frog, etc.
        What use cases you have (examples) where bots have to interact with inner elements or on-screen data where html front-end is not used?

        1. 1

          Price2Spy uses both methods
          A) Web Driver, parsing what is visible on-screen. Here we use a solution that automatically detects the most prominent numerical piece of content, and extracts it as price.
          B) HTTP response parsing (HTML, JS, JSON etc), where we have several different methods for extracting prices (some methods take 2-3 minutes to configure, others may take up to 1-2 hours, depending on site complexity)

          Whenever possible, we use B). Why?

          •      A) is much more likely to capture the wrong price
            
          •      A) offers no good method for capturing product availability (we need both price + availability)
            
          •      B) is superior to complex cases like
            

          o Product with variations (especially if you want to capture non-default variation)
          o Multiple products are shown on the same page
          o Products where the price in non-default currency needs to be captured
          o Price is not shown as a continual piece of text – but rather displayed in several different elements (integer price shown in large font, and decimal part in much smaller font)

          Six months ago we have conducted an experiment – we took 100 URLs from random sites in our database (we monitor prices on more than 100K sites) and compared results from A) and B)
          The results were: A) Captured correct price in 38% of cases, B) captured all 100% right

          So, the short conclusion is – if the site is of simple structure, A) should work,. However, whenever there is a more complex situation, B) is the way to go.