9 ms·
Closer to the Metal: Leaving Playwright for CDP
- arm32 1y agoAh, yes, the classic "Playwright isn't fast enough so we're reinventing Puppeteer" trope. I'd be lying if I haven't seen this done a few times already. Now that I got my snarky remark out of the way: Puppeteer uses CDP under the hood. Just use Puppeteer.
- haolez 1y agoI've seen a team implement Go workers that would download the HTML from a target, then download some of the referenced JavaScript files, then run these JavaScript files in an embedded JavaScript engine so that they could consume less resources to get the specific things that they needed without using a full browser. It's like a browser homunculus! Of course, each new site would require custom code. This was for quant stuff. Quite cool!
- odo1242 1y agoThis exact homunculus is actually supported in Node.JS by the `jsdom` library: https://www.npmjs.com/package/jsdom https://www.npmjs.com/package/jsdom I don't know how well it would work for that use-case, but I've used it before, for example, to write a web-crawler that could handle client-side rendering.
- nikisweeting 1y agoour use primary use-case with the AI stuff is not really scraping, we're mostly going after RPA
- boredtofears 1y agoIs the case for playwright over puppeteer just in it's crossbrowser support? We're currently using Cypress for some automated testing on a recent project and its extremely brittle. Considering moving to playwright or puppeteer but not sure if that will fix the brittleness.
- rising-sky 1y agoIn my experience Playwright provided a much more stable or reliable experience with multiple browser support and asynchronous operations (which is the entire point) over Puppeteer. ymmv
- arm32 1y agoPlaywright also offers nice sugar like HTML test reports and trace viewing.
- gregpr07 1y agoFrom my experience with Playwright RR-Web recordings are MUCH better than Playwright’s replay traces, so we usually just use those.
- epolanski 1y agoWhat's RR web?
- nikisweeting 1y agohttps://github.com/rrweb-io/rrweb https://github.com/rrweb-io/rrweb
- jonkoops 1y agoThat can be integrated with Playwright, or did you mean to say it is already used under the hood for their reports?
- nikisweeting 1y agoGregor was saying it works without needing playwright, and provides more detailed trace recordings than playwright does. we plan to use rr-web and maybe browsertrix for our website archival / replay system for deterministic evals.
- nikisweeting 1y agosir we are a python library, puppeteer-python was abandoned, how exactly do you propose we use puppeteer?
- epolanski 1y agoPlaywright has Python bindings .
- nikisweeting 1y agoyes I know, I wrote the post
- arm32 1y agoHave you considered just using Playwright? ;)
- hugs 1y agoyeah, i continue to be amazed at how google dropped the ball on this one.
- wredcoll 1y agoWait, does playwright not use cdp? What does it do?!
- steveklabnik 1y agoDescribing "2011–2017" as "the dark ages" makes me feel so old. There was a ton of this stuff before Chrome or WebKit even existed! Back in my day, we used Selenium and hated it. (I was lucky enough to start after Mercury...)
- hugs 1y agoselenium creator here. hi!
- arm32 1y agoUh, ahem, <clears throat>, we meant the _other_ Selenium.
- hugs 1y agothat's what i thought. :) personal life accomplishment was seeing wikipedia add a disambiguation link on the element's page. you know, because it's right up there in importance as the periodic table, obviously.
- gregpr07 1y agoHi, the first version of Browser Use was actually built on Selenium but we quite quickly switched to Playwright
- hugs 1y agoyeah, i noticed that. apologies if i missed a post about it... what do you wish didn't suck about selenium?
- moi2388 1y agoScrolling to an element doesn’t always work because somehow the element might not be ready. You need to add ids to the element and select by that to ensure it works properly.
- dataviz1000 1y agoI made this comment yesterday but really applies to this conversation. > In the past 3 weeks I ported Playwright to run completely inside a Chrome extension without Chrome DevTools Protocol (CDP) using purely DOM APIs and Chrome extension APIs, I ported a TypeScript port of Browser Use to run in a Chrome extension side panel using my port of Playwright, in 2 days I ported Selenium ChromeDriver to run inside a Chrome Extension using chrome.debugger APIs which I call ChromeExtensionDriver, and today I'm porting Stagehand to also run in a Chrome extension using the Playwright port. This is following using VSCode's core libraries in a Chrome extension and having them drive a Chrome extension instead of an electron app. The most difficult part is managing the lifecycle of Windows, Pages, and Frames and handling race conditions, in the case of automating a user's browser, where, for example, the user switches to another tab or closes the tab.
- Tsarp 1y agoWouldnt having chrome.debugger=true also flag your requests?
- wonger_ 1y agoWhat is the benefit of porting all those tools to extensions? Have you ran into any other extension-based challenges besides lifecycles and race conditions?
- diggan 1y ago> What is the benefit of porting all those tools to extensions? Personally, I have a browser extension running in my user/personal browser instance that my agent use (with rate-limits) in order to avoid all the captchas and blocks basically. Everything else I've tried ultimately ends up getting blocked. But then I'm also doing some heavy caching so most agent "browse" calls end up not even reaching out to the internet as it's finding and using stuff already stored locally.
- dataviz1000 1y agoSome benefits (without using Chrome.debugger or Chrome DevTools Protocol): 1. There are 3,500,000,000 instances of Chrome desktop being used. [0] 2. A Chrome Extension can be installed with a click from the Chrome Web Store. 3. It is closer to the metal so runs extremely fast. 4. Can run completely contained on the users machine 5. It's just one user automating their web based workflows making it harder for bot protections to stop and with a human-in-the-loop any hang ups and snags can be solved by the human 6. Chrome extensions now have a side panel that is stationary in the window during navigation and tab switching. It is exactly like using the Cursor or VSCode side panel copilots Some limitations: 1. Can't automate ChatGPT console because they check for user agent events by testing if the `isTrusted` property on event objects is true. (The bypass is using Chrome.debugger and the ChromeExtensionDriver I created.) 2. Can't take full page screen captions however it is possible to very quickly take visible scree captions of the viewport. Currently I scroll and stitch the images together if a full page screen is required. There are other APIs which allow this in a Chrome Extension and can capture video and audio but they require the user to click on some button so it isn't useful for computer vision automation. (The bypass is once again using the Chrome.debugger and ChromeExtensionDriver I created.) 3. Chrome DevTool Protocol allows intercepting and rewriting scripts and web pages before they are evaluated. With manifest v2 this was possible but they removed this ability in manifest v3 which we still hear about today with the adblock extensions. I feel like with the limitations having a popup dialog that directs the user to do an action will work as long as it automates 98% of the user's workflows. Moreover, a lot of this automation should require explicit user acknowledgments before preceding. [0] https://www.demandsage.com/chrome-statistics/ https://www.demandsage.com/chrome-statistics/
- deleted 1y ago[deleted]
- Robdel12 1y agoAll of the approaches of driving the browser outside of the browser is going to be slow (webdriver, playwright, puppeteer, etc). Karma like approaches are where I’m at (execute in the browser)
- appcustodian2 1y ago> All of the approaches of driving the browser outside of the browser is going to be slow Why? I would think any cross-process communication through the CDP websocket would have imperceptible overhead compared to what already takes long in the browser: a ton of HTTP I/O What is Karma? What are you executing in the browser?
- nikisweeting 1y agoCDP rountrip time on a local machine is 100µs (0.1ms), it's not slow haha
- johnsmith1840 1y ago"Thousands of cdp calls" from the link. Cdp does add a good chunk of latency. Depends on what your threshold is. An image grab is around 60ms and a snapshot can range from 40ms -> 500ms The latency is pure data movement. It's like the difference of using ram vs ssd vs data from the internet.
- nikisweeting 1y agoyeah but good luck getting rid of that with a browser extension, you're just moving the latency around / moving it to chrome.runtime message passing.
- benmmurphy 1y agodirect CDP has been used by the scraping community for a long time in order to have a cleaner browser environment that is harder to fingerprint. for example nodriver (https://github.com/ultrafunkamsterdam/nodriver https://github.com/ultrafunkamsterdam/nodriver) was started in Feb 2024 and I suspect this technique was popular before that project started.
- gregpr07 1y agoI really like both nodriver and pydoll. I am definitely keeping the option of switching to them open, but we just wanted to have full control for now and see how painful CDP-use is to maintain first and then reconsider.
- spullara 1y agothis is exactly what I did when I wrote my first agent with scraping. later we switched to taking control of the users browser through a browser extension.
- saberience 1y agoTalk about "not built here" mentality. This is a project doomed to failure. Using VC money to re-write better built software which has been around for years. Good luck guys!
- Tostino 1y agoExactly what I was thinking. Instead of attempting to contribute back to Playwright to fix those hangups, or even creating a private patch to do so as a POC, they went right to building their own framework from scratch. That isn't how you launch a product.
- gregpr07 1y agoI mean... Playwright was built and is maintained by Microsoft, so I don't think VC money argument really makes sense here. By the very nature of how Playwright is built we can't contribute to it - it runs inside a JS subprocess and does not expose a bunch of CDP apis that we NEED (for example to make cross origin iframes work).
- nikisweeting 1y agoI've been trying to contribute to playwright for years! All of my issues have been closed / rejected without much consideration because they're not part of the core "QA testing" use-case that playwright is built for. Personally have not found their team to be the easiest to work with on Github. I would've loved to use puppeteer instead, their team is quite reasonable but they abandoned their python bindings and we want to stay in python.
- hugs 1y agore: ms -- thank you for calling that out. i've been thinking we had been collectively sleepwalking into ms owning everything (again). they've owned everything once before -- it wasn't great! related side-note: have you had to interact with the core chrome / cdp devs?
- 1y ago
- patrickhogan1 1y agoSelenium was very usable before 2011. This post is like saying Grafana and not mentioning Nagios
- nikisweeting 1y agoIt was, but I feel like the advent of headless browsers marked a step function explosion in browser automation. Also any earlier than 2010 is when I was like 13yo, so it's more like "the dark ages in my own memory" than "objectively dark ages in automation history".
- patrickhogan1 1y agoI get that drawing historical boundaries is arbitrary, but Selenium is a really good prior. Selenium offered headless mode and integrated with 3rd party providers like BrowserStack, which ran acceptance tests in parallel in the cloud. It seems like what browser-use.com is doing is a modern day version with many more features & adaptability.
- hugs 1y agospeaking of priors... sauce labs existed for three whole years before browserstack (selenium and sauce founder here. :-) i like that there are new startups in the space, though. things were getting pretty stale and uninspired.
- patrickhogan1 1y agoSauce Labs is excellent. I've actually used it extensively myself (not sure why BrowserStack came to mind first). I remember Sauce Labs was super active in the SF Selenium community and the Selenium meetups. Just checked my emails. Good memories. Thank you for building Selenium.
- nikisweeting 1y agoYeah I agree, I changed the history section a bit: https://github.com/browser-use/browser-use/blob/2a0f4bd93a43fd43086a47cd2465ad4528399e5d/browser_use/browser/crash_watchdog.py#L318 https://github.com/browser-use/browser-use/blob/2a0f4bd93a43...
- johnsmith1840 1y agoWhy not cdp snapshot?
- nikisweeting 1y agoWhat do you mean? We use CDP page snapshots extensively to get full html across frames but it's not nearly enough on its own, there are lots of checks still needed for individual OOPIFs or elements.
- johnsmith1840 1y agoYou can get all of that pure snapshot. There are no extra checks needed it's by a significant margin the most reliable method to see current state. I run snapshot at 10-20fps though plus the same for parallel image capture. I've been wondering if I should release just this part of my system open source seems like I'm not alone in how complex this all is. I could launch yet another automation framework!
- nikisweeting 1y agoyou still need separate calls to get the AXTree with full computed aria properties, and a bunch of Runtime.evaluate calls to scan all the dynamically-added event listeners.
- johnsmith1840 1y agoSnapshot is all you need. I get all of the info you describe in pure snapshot. Automation works fine.
- nikisweeting 1y agoyou cannot get the list of bound onclick handlers with a snapshot
- 1y ago
- ipsum2 1y agoNice thorough write up, I've had my share of annoyances with playwright for automating some menial tasks due to being blocked by captcha or other waf (I'm just logging into my own accounts and scraping my account balance, nothing nefarious), I'll try out pydoll or your library next time.
- aitchnyu 1y agoUmm, will this run on Firefox too? They deprecated CDP and favors Webdriver Bidi.
- hugs 1y agoi like that the post uses the phrase "time is a flat circle". it is indeed. once upon a time, most devs only cared about one browser -- internet explorer. then for a good chunk of time, cross-browser compatibility was highly valued. now, most devs only care about one browser -- google chrome. it's a bummer, but also a market reality... the best way to get more devs to care about non-chrome browsers is to get more people to use non-chrome browsers. easier said than done, though.
- nikisweeting 1y agoNo we are not planning to support Firefox. We do support Brave, Edge, and ungoogled-chromium though if you have a problem with Google.
- deleted 1y ago[deleted]
- 1vuio0pswjnm7 1y agohttps://web.archive.org/web/20250820155850if_/https://browser-use.com/posts/playwright-to-cdp https://web.archive.org/web/20250820155850if_/https://browse...
- pjmlp 1y agoNote how "browsers" is all about Chrome.
- OldfieldFund 1y agoChromium is the only browser that has extensive undetectable/automation support. Look at patchright: https://github.com/Kaliiiiiiiiii-Vinyzu/patchright?tab=readme-ov-file https://github.com/Kaliiiiiiiiii-Vinyzu/patchright?tab=readm...
- pjmlp 1y agoThen people complain that Google is taking over the Web, well don't help them in the process. Guess how IE became what it was before the lawsuit, it was the cool browser when all nice developer features came first. Dynamic HTML, HTML Applications, CSS shaders (backed by DirectX), VS debugging integration (via Frontpage)... Apparently a lesson gone after one generation.
- OldfieldFund 1y agoI think blaming people who are trying to make a buck using the fastest route is not the way to achieve non-monopoly. A practical point: Mozilla made design choices in the past that made it harder to hide the automation footprint. For some time it was more difficult to disable the navigator.webdriver flag in Firefox compared to Chromium.
- nikisweeting 1y agoThat's why we support Brave, Edge, Ungoogled-Chromium, and our own custom Chromium fork that we're working on. Just because we only support Chromium doesn't mean we're pro-Google-dominance. There are enough Chrome forks at this point that Google no longer has the power to unilaterally remove features from Chromium. Manifest v2 extensions still work great in Brave for example.
- nikisweeting 1y agoThat's not really true, https://github.com/daijro/camoufox https://github.com/daijro/camoufox is at parity with patchright on stealth, it's just that Firefox has way less market share so it's not worth the >2x maintanance effort for us to support multiple browsers.