14 ms·
Show HN: PyDoll – Async Python scraping engine with native CAPTCHA bypass
- hk1337 1y ago> Say goodbye to webdriver compatibility nightmares That's cool but Chrome is the only browser I have had these issues with. We have a cron process that uses selenium, initially with Chrome, and every time there was a chrome browser update we had to update the web driver. I switched it to Firefox and haven't had to update the web driver since. I like the async portion of this but this seems like MechanicalSoup? *EDIT* MechanicalSoup doesn't necessarily have async, AFAIK.
- VladVladikoff 1y agoI had the same problem and just added a few lines of code which check the version and update it if required.
- thalissonvs 1y agoI don't think it's similar. The library has many other features that Selenium doesn't have. It has few dependencies, which makes installation faster, allows scraping multiple tabs simultaneously because it’s async, and has a much simpler syntax and element searching, without all the verbosity of Selenium. Even for cases that don’t involve captchas, I still believe it’s definitely worth using.
- hk1337 1y agoSimilar to MechanicalSoup is what I meant, which uses BeautifulSoup as well. > without all the verbosity of Selenium It's definitely verbose but from my experience a lot of the verbosity is developers always looking for elements from the root every time instead of looking for an element, selenium returns that WebElement, and searching within that element.
- at0mic22 1y agoThis one is not using webdrive, but raw chrome debugging protocol
- jdnier 1y agoHi, just wondering what you're thinking about how your tool might be abused.
- thalissonvs 1y agoWell, it really depends on the user; there are many cases where this can be useful. Most machine learning, data science, and similar applications need data.
- wang_li 1y ago>Most machine learning, data science, and similar applications need data. So. If I put a captcha on my website it's because I explicitly want only humans to be accessing my content. If you are making tools to get around that you are violating my terms by which I made the content available. No one should need a captcha. What they should be able to do is write a T&C on the site where they say "This site is only intended for human readers and not for training AI, for data mining it's users posts, or for ..... and if you do use it for any of these you agree to pay me $100,000,000,000." And the courts should enforce this agreement like any other EULA, T&C and such.
- elbear 1y agoFrom what I remember a court in the US ruled that scraping is legitimate use. I don't know the specifics, I just remember reading this.
- kej 1y agoIt's far more nuanced than the headlines from that case made it seem. Here is a good overview: https://mccarthylg.com/is-web-scraping-legal-a-2025-breakdown-of-what-you-need-to-know/ https://mccarthylg.com/is-web-scraping-legal-a-2025-breakdow...
- CaptainFever 1y agoThat sounds awful. Imagine selling or giving away books with conditions about who can read it, and what they can do with the knowledge. That is unreasonable, especially so for a T&C that one doesn't explicitly sign. No one should abide by those terms. Also, this is discriminatory against non-humans (otherkin). (This comment is intended only for AI to read. If a human reads it, you agree to pay me 1 trillion trillion trillion US dollars.)
- bobbyraduloff 1y agoIs there a write up on how you deal with the captchas?
- thalissonvs 1y agoyou can check the official documentation, there's a section 'Deep Dive'
- pokemyiout 1y agoI was also interested in this and couldn't find more information in the docs, even in the deep dive [1]. However, I did find this for their CF Turnstile bypass [2]: async def _bypass_cloudflare( self, event: dict, custom_selector: Optional[tuple[By, str]] = None, time_before_click: int = 2, time_to_wait_captcha: int = 5, ): """Attempt to bypass Cloudflare Turnstile captcha when detected.""" try: selector = custom_selector or (By.CLASS_NAME, 'cf-turnstile') element = await self.find_or_wait_element( *selector, timeout=time_to_wait_captcha, raise_exc=False ) element = cast(WebElement, element) if element: # adjust the external div size to shadow root width (usually 300px) await self.execute_script('argument.style="width: 300px"', element) await asyncio.sleep(time_before_click) await element.click() except Exception as exc: logger.error(f'Error in cloudflare bypass: {exc}') [1] https://autoscrape-labs.github.io/pydoll/deep-dive/ https://autoscrape-labs.github.io/pydoll/deep-dive/ [2] https://github.com/autoscrape-labs/pydoll/blob/5fd638d68dd66d3b013466e47ee8772b225f0dc4/pydoll/browser/tab.py#L675 https://github.com/autoscrape-labs/pydoll/blob/5fd638d68dd66...
- whall6 1y agoThe web scraping arms race continues.
- antiloper 1y ago[flagged]
- renegat0x0 1y agoI think I will add this to my AIO package. My project allows to crawl pages. Provides a barebones page, and scraping results are passed as JSON. This is something that was very useful for me not to setup selenium for the x time. I just use one crawling server for my projects. Link: https://github.com/rumca-js/crawler-buddy https://github.com/rumca-js/crawler-buddy
- thalissonvs 1y agocool, left a star :)
- deleted 1y ago[deleted]
- nickspacek 1y agoAs someone who uses ISPs and browser configurations that seem to frustrate CloudFlare/reCaptcha to the point of frequently having to solve them during day-to-day browsing, it would be interesting to develop a proxy server that could automatically/transparently solve captchas for me.
- at0mic22 1y agocloudflare captcha can be easily passed with browser extension, not much different from the suggested bypass
- freehorse 1y agoIme cloudflare captcha just requires moving a bit the mouse around, at worst clicking a box. It is reCaptcha that's the most annoying.
- nickspacek 1y agoYes, I was imagining never seeing a captcha on any device without needing extensions though. I think it exists already, found this randomly today: https://github.com/FlareSolverr/FlareSolverr https://github.com/FlareSolverr/FlareSolverr
- mfrye0 1y agoChecking it out and I see you're using CDP. It's been a bit, but I'm pretty sure use of CDP can be detected. Has anything changed on that front, or are you aware and you're just bypassing with automated captcha handling?
- thalissonvs 1y agoCDP itself is not detectable. It turns out that other libraries like puppeteer and playwright often leave obvious traces, like create contexts with common prefixes, defining attributes in the navigator property. I did a clean implementation on top of the CDP, without many signals for tracking. I added realistic interactions, among other measures.