4 ms·
Author here. The browser is made from scratch (not based on Chromium/Webkit), in Zig, using v8 as a JS engine. Our idea is to build a lightweight browser optim
by fbouvier 2y ago
Author here. The browser is made from scratch (not based on Chromium/Webkit), in Zig, using v8 as a JS engine.
Our idea is to build a lightweight browser optimized for AI use cases like LLM training and agent workflows. And more generally any type of web automation.
It's a work in progress, there are hundreds of Web APIs, and for now we just support some of them (DOM, XHR, Fetch). So expect most websites to fail or crash. The plan is to increase coverage over time.
Happy to answer any questions.
- toobulkeh 2y agoI’d love to see better optimized web socket support and “save” features that cache LLM queries to optimize fallback
- JoelEinbinder 2y agoWhen I've talked to people running this kind of ai scraping/agent workflow, the costs of the AI parts dwarf that of the web browser parts. This causes computational cost of the browser to become irrelevant. I'm curious what situation you got yourself in where optimizing the browser results in meaningful savings. I'd also like to be in that place! I think your ram usage benchmark is deceptive. I'd expect a minimal browser to have much lower peak memory usage than chrome on a minimal website. But it should even out or get worse as the websites get richer. The nature of web scraping is that the worst sites take up the vast majority of your cpu cycles. I don't think lowering the ram usage of the browser process will have much real world impact.
- refulgentis 2y agoGenerally, for consumer use cases, it's best to A) do it locally, preserving some of the original web contract B) run JS to get actual content C) post-process to reduce inference cost D) get latency as low as possible Then, as the article points out, the Big Guns making the LLMs are a big use case for this because they get a 10x speedup and can begin contemplating running JS. It sounds like the people you've talked to are in a messy middle: no incentive to improve efficiency of loading pages, simply because there's something else in the system that has a fixed cost to it. I'm not sure why that would rule out improving anything else, it doesn't seem they should be stuck doing nothing other than flailing around for cheaper LLM inference. > I think your ram usage benchmark is deceptive. I'd expect a minimal browser to have much lower peak memory usage than chrome on a minimal website. I'm a bit lost, the ram usage benchmark says its ~10x less, and you feel its deceptive because you'd expect ram usage to be less? Steelmanning: 10% of Chrome's usage is still too high?
- JoelEinbinder 2y agoThe benchmark shows lower ram usage on a very simple demo website. I expect that if the benchmark ran on a random set of real websites, ram usage would not be meaningfully lower than Chrome. Happy to be impressed and wrong if it remains lower.
- fbouvier 2y agoI believe it will be still significantly lower as we skip the graphical rendering. But to validate that we need to increase our Web APIs coverage.
- fbouvier 2y agoThe cost of the browser part is still a problem. In our previous startup, we were scraping >20 millions of webpages per day, with thousands of instances of Chrome headless in parallel. Regarding the RAM usage, it's still ~10x better than Chrome :) It seems to be coming mostly from v8, I guess that we could do better with a lightweight JS engine alternative.
- Tostino 2y agoYou may reduce ram, but also performance. A good JIT costs ram.
- fbouvier 2y agoYes, that's true. It's a balance to find between RAM and speed. I was thinking more on use cases that require to disable JIT anyway (WASM, iOS integration, security).
- Tostino 2y agoYeah, could be nice to allow the user to select the type of ECMAScript engine that fits their use-case / performance requirements (balancing the resources available).
- cxr 2y agoIf your target is consistent enough (perhaps even stationary), then at some point "JIT" means wasting CPU cycles.
- cush 2y ago> there are hundreds of Web APIs, and for now we just support some of them (DOM, XHR, Fetch) > it's still ~10x better than Chrome Do you expect it to stay that way once you've reached parity?
- fbouvier 2y agoI don't expect it to change a lot. All the main components are there, it's mainly a question of coverage now.
- szundi 2y agoThen came deepseek
- dtj1123 2y agoVery nice. Does this / will this support the puppeteer-extra stealth plugin?
- katiehallett 2y agoThanks! Right now no, but since we use the CDP (playwright, puppeteer), I guess it would be possible to support it
- sesm 2y agoGreat job! And good luck on your journey! One question: which JS engines did you consider and why you chose V8 in the end?
- fbouvier 2y agoWe have also considered JavaScriptCore (used by Bun) and QuickJS. We did choose v8 because it's state of the art, quite well documented and easy to embed. The code is made to support others JS engine in the future. We do want to add a lightweight alternative like QuickJS or Kiesel https://kiesel.dev/ https://kiesel.dev/
- ksec 2y agoThank You I was thinking of JSC and Bun as well. Was half expecting JSC since that combination seems to work well.
- bityard 2y agoPlease put a priority on making it hard to abuse the web with your tool. At a _bare_ minimum, that means obeying robot.txt and NOT crawling a site that doesn't want to be crawled. And there should not be an option to override that. It goes without saying that you should not allow users to make hundreds or thousands of "blind" parallel requests as these tend to have the effect of DoSing sites that are being hosted on modest hardware. You should also be measuring response times and throttling your requests accordingly. If a website issues a response code or other signal that you are hitting it too fast or too often, slow down. I say this because since around the start of the new year, AI bots have been ravaging what's left of the open web and causing REAL stress and problems for admins of small and mid-sized websites and their human visitors: https://www.heise.de/en/news/AI-bots-paralyze-Linux-news-site-and-others-10252162.html https://www.heise.de/en/news/AI-bots-paralyze-Linux-news-sit...
- gkbrk 2y agoPlease don't. Software I installed on my computer needs to the what I want as the user. I don't want every random thing I install to come with DRM. The project looks useful, and if it ends up getting popular I imagine someone would make a DRM-free version anyway.
- tossandthrow 2y agoWhere do you read DRM? Parent commenter merely and humbly asks the author of the library to make sure that it has sane defaults and support for ethical crawling. I find it disturbing that you would recommend against that.
- gkbrk 2y agoHere's what the parent comment wrote. > And there should not be an option to override that. This is not just a sane default. This is software telling you what you are allowed to do based on what the rights owner wants, literally DRM. This is exactly like Android not allowing screenshots to be taken in certain apps because the rights owner didn't allow it.
- afk1914 2y agoI am curious how Lightpanda compares to chrome-headless-shell ({headless: 'shell'} in Puppeteer) in benchmarks.
- fbouvier 2y agoWe did not run benchmarks with chrome-headless-shell (aka the old headless mode) but I guess that performance wise it's on the same scale as the new headless mode.
- xena 2y agoHow do I make sure that people can't use lightpanda to bypass bot protection tools?
- dolmen 2y agoOne of Lightpanda's goals is to ease building bots.
- danielsht 2y agoVery impressive! At Airtop.ai we looked into lightweight browsers like this one since we run a huge fleet of cloud browsers but found that anything other than a non-headless Chromium based browser would trigger bot detection pretty quickly. Even spoofing user agents triggers bot detection because fingerprinting tools like FingerprintJS will use things like JS features, canvas fingerprinting, WebGL fingerprinting, font enumeration, etc. Can you share if you've looked into how your browser fares against bot detection tools like these?
- fbouvier 2y agoThanks! No we haven't worked on bot detection.
- 867-5309 2y agodoes this work with selenium/chromedriver?
- fbouvier 2y agoFor now we just support CDP. But Selenium is definitely in our roadmap.
- keepamovin 2y agoIf you support Page.startScreencast or even just capture screenshot we could experiment with using this as a backend for BrowserBox, when lightpanda matures. Cool stuff! https://github.com/BrowserBox/BrowserBox/ https://github.com/BrowserBox/BrowserBox/
- returnofzyx 2y agoHi. Can I embed this as library? Is there C API exposed? I can't seem to find any documentation. I'd prefer this to a CDP server.
- fbouvier 2y agoNot now but we might do it in the future. It's easy to export a Zig project as a C ABI library.
- returnofzyx 2y agoOh please do. I'm sure there are many people like me who want this.
- niutech 2y agoCongratulations! But does it support Google Account login? And ReCAPTCHA?