7 ms·
Show HN: I wrote an open-source browser alternative for Computer Use for any LLM
Hey HN,
I made Browser-Use, an open-source tool that lets (all Langchain supported) LLMs execute tasks directly in the browser just with function calling.
It allows you to build agents that interact with web elements using natural language prompts. We created a layer that simplifies website interaction for LLMs by extracting xPaths and interactive elements like buttons and input fields (and other fancy things). This enables you to design custom web automation and scraping functions without manual inspection through DevTools.
Hasn't this been done a lot of times?
Good question, as a general SaaS tool yes, but I think a lot of people are going to try to make their own web automation agents from scratch, so the idea is to provide groundwork/library for the hard part so that not everyone has to repeat these steps:
- parse html in a LLM friendly way (clickable items + screenshots)
- provide a nice function calls for everything inside the browser
- create reusable agent classes
What this is NOT? An all knowing AI agent that can solve all your problems.
The vision: create repeatable tasks on the web just by prompting your agent and not care about the hows.
To better showcase the power of text extraction we made a few demos such as:
- Applying for multiple software engineering jobs in San Francisco
- Opening new tabs to search for images of Albert Einstein, Oprah Winfrey, and Steve Jobs
- Finding the cheapest one-way flight from London to Kyrgyzstan for December 25th
I’d be interested in feedback on how this tool fits into your automation workflows. Try it out and let me know how it performs on your end.
We are Gregor & Magnus and we built this in 5 days.
- coreyp_1 2y agoThis looks really interesting. The first hurdle, though, that prevents me from experimenting with this on my job is the lack of a license. I see in the readme that it claims that it is MIT licensed, but there is no actual license file or information in any of the source files that I could find.
- gregpr07 2y agoThanks you, love the feedback! Will add the license. Let me know how it goes if you try it.
- WillAdams 2y agoDoes it work with COM objects/Java applications? I'd give my interest in Hell for a way to have a script plug in data into a Java app.
- bravura 2y agoIt would be amazing if you: a) There were a test / eval suite to determine which model works best for what. It could be divided into a training suite and test suite. (Training tasks can be used for training, test tasks only for evaluation.) Possibly a combination of unit tests against known xpaths, and integration tests that are multi-step and end in a measurable result. I know the web is constantly changing, so I'm not 100% sure how this should work. b) There were some sort of wiki, or perhaps another repo or discussion board, of community-generated prompt recipes for particular actions.
- gregpr07 2y agoA) we plan on thoroughly testing that with Mind2Web dataset. They have a very robust set of (persistant) selectors B) so, shadcn for prompts for web agents haha :) but I agree, that would be SICK! Just go to browseruse and get the prompt for your specific use case
- maggreenWAI 2y agoA) For Mind2Web: because there are multiple ways to reach a goal state - any thoughts how to evaluate if a task was successful? Should we let the LLM/ other LLM evaluate it?
- firejake308 2y agoIs it decided then that screenshots are better input for LLMs than HTML, or is that still an active area of investigation? I see that y'all elected for a mostly screenshot-based approach here, wondering if that was based on evidence or just a working theory.
- maggreenWAI 2y agoNext step is to represent also the structure of the HTML tree in the extracted elements for better understanding, maybe images are then less needed.
- gregpr07 2y agoNot sure, I think there is a lot of research being done here. Actually, browser use works quite well with vision turned off, it just sometimes gets stuck at some trivial vision tasks. The interesting thing is that screenshot approach is often cheaper than cleaned up html, because some websites have HUGE action spaces. We looked at some papers (like ferret ui) but i think we can do much better on html tasks. Also, there is a lot of space to improve the current pipeline.
- arjvik 2y agoWould be really cool if you could tie this into Claude's computer use APIs!
- gregpr07 2y agoDo you think they do any super fancy magic other than for example how ferret ui does their classification of ui elements? It could be very interesting to test head to head hope much better you can make computer use by adding html (it’s much better from our quick testing, just don’t know the numbers).
- stuckkeys 2y agoMy wallet just ran way. Come back you!
- rahimnathwani 2y agoIn case anyone else was looking for the functions available to the LLM: https://github.com/gregpr07/browser-use/blob/68a3227c8bc97fe424f90f744c295c7d330ae5fd/src/controller/views.py#L59 https://github.com/gregpr07/browser-use/blob/68a3227c8bc97fe...
- gregpr07 2y agoYes, this plus reasoning and ask user for additional info. More here https://github.com/gregpr07/browser-use/blob/main/src/agent/views.py https://github.com/gregpr07/browser-use/blob/main/src/agent/...
- maggreenWAI 2y agoYou can just extend this e.g. with adding data to database, sending notifications, extracting specific data format ect... make sure to also accept your added function when its called in act() https://github.com/gregpr07/browser-use/blob/main/src/controller/service.py https://github.com/gregpr07/browser-use/blob/main/src/contro...
- daft_pink 2y agoI was really excited about the original claude computer use until I watched the youtube videos and saw it was only running in a docker container. I wish I could run something like this on a real machine.
- deleted 2y ago[deleted]
- gregpr07 2y agoYou can run browser use in your terminal, no need for Docker containers. Just clone it and run it
- craftkiller 2y agoWhat makes a docker container not a "real machine"? Docker programs are running natively, just like any other program, without emulation/virtualization. Its not like a virtual machine (unless you're on one of the lesser operating systems like Windows or OSX), its just configuring some settings in the Linux kernel to isolate the process from other processes. Its basically just an enhanced chroot.
- alexeichhorn 2y agoRunning natively doesn't make it a real machine. If I run iOS Simulator, it also runs natively, but I'm pretty sure it's not a real iPhone ;)
- daft_pink 2y agoBecause I would like to automate programs I use in Windows with files that I actually use, instead of some random linux docker container. I should have used my actual machine instead of a real machine.
- craftkiller 2y ago> Windows Ah, well there's your problem. Your problem isn't docker, nor is it claude, its that you're running Windows.
- Oras 2y agoThis looks interesting. I am really impressed with MultiOn [0], and I tried to make something similar, but it's quite challenging doing it with a Chrome extension. I also saw one doing Captcha solving with Selenium [1]. I will keep an eye on your development, good luck! [0] https://www.multion.ai/ https://www.multion.ai/ [1] https://github.com/VRSEN/agency-swarm https://github.com/VRSEN/agency-swarm
- gregpr07 2y agoThanks! Have you tried captcha solving with [1]? It's very tricky sometimes, especially with non standard "verify human" - maybe you could solve it by writing Selenium/Javascript code directly and then execute it.
- aethelingas 2y agowhat are the challenges with the Chrome extension path?
- Oras 2y agoYou need to call an API to screenshot the page, then figure out the JavaScript code to execute it. It’s not as easy as it might sound. Playwright and selenium automate the browser itself, but with the chrome extension you need to use the context of the current browser. I’m not an expert in browser automation so found it challenging moving from playwright to make it completely browser based.
- soham123 2y agoI have built something similar at https://github.com/ComposioHQ/composio/tree/master/python/composio/tools/local/browsertool/actions https://github.com/ComposioHQ/composio/tree/master/python/co... Compatible with any LLMs and agentic framework
- gregpr07 2y agoLooks nice. I find the cleaning HTML step in our cleaning pipeline extremely important, otherwise there is no real benefit from just using a general vision model and clicking coordinates (and whole HTML is just way too many tokens). How do you guys handle that?
- theredsix 2y agoAwesome project, starred! Here are some other projects for agentic browser interactions: * Cerebellum (Typescript): https://github.com/theredsix/cerebellum https://github.com/theredsix/cerebellum * Skyvern: https://github.com/Skyvern-AI/skyvern https://github.com/Skyvern-AI/skyvern Disclaimer: I am the author of Cerebellum
- gregpr07 2y agoThanks man, starred yours too, it's super cool to see all these projects getting spun up! I see Cerebellum is vision only. Did you try adding HTML + screenshot? I think that improves the performance like crazy and you don't have to use Claude only. Just saw Skyvern today on previous Show HNs haha :)
- theredsix 2y agoI had an older version that used simplified HTML, and it got to decent performance with GPT-4o and Gemini but at the cost of 10x token usage. You are right, identifying the interactable elements and pulling out their values into a prompt structure to explicitly allow the next actions can boost performance, especially if done with grammar like structured outputs or guidance-llm. However, I saw that Claude had similar levels of performance with pure vision, and I felt that vision + more training would beat a specialized DOM algorithm due to "the bitter lesson". BTW I really like your handling of browser tabs, I think it's really clever.
- gregpr07 2y agoFair, also Claude probably only gets better on this since they kinda want people to use Computer use. We are gonna try to do best of both worlds. Thanks man, Magnus came up with it this morning haha!
- dbacar 2y agoI starred both of you
- fragmede 2y agowants to have cron, so I can ask it to check with my local parking agency, every day or every 12 hours, do I have a parking ticket, and to raise a warning if I do. Or to check with county jail and see if someone is still there/not there. Or check the price of a product on Amazon every hour and warn when it's changed (aka camelcamelcamel but local). Search craigslist/zillow/Facebook marketplace for items until one shows up. etc.
- maggreenWAI 2y agoLet's say in 1 year, more agents than humans interact with the web. Do you think: 1. Websites release more API functions for agents to interact with them or 2. We will transform with tools like this the UI into functions callable by agents and maybe even cache all inferred functions for websites in a third party service?
- stuartjohnson12 2y agoCame across some folks a little while ago who just raised a few million $ for the latter (in stealth).
- gregpr07 2y agoWho (or still in stealth haha)?
- stuartjohnson12 2y agoI'll leave it to them to announce once they're ready, but it's certainly a pretty future-looking play.
- KaoruAoiShiho 2y agoMaybe can build a database for which sites / pages work best with HTML vs Screenshots, and then can choose to use HTML to save on token cost / improve latency if possible.
- G_o_D 2y agoIt is called screen scraping, where text rendered on screen/monitors are being scraped either in browser or even in windows os even on android screen , thats how softwares like autohotkey and all do automation windows or android screen can be dumped into heirarchical xml along with x y coordinates of its ui elements along with text they contain which can be uses o click scroll scrape text
- gitgud 2y agoIt's impressive, but to me it seems like the saddest development experience... agent = Agent( task='Go to hackernews on show hn and give me top 10 post titels, their points and hours. Calculate for each the ratio of points per hour.', llm=ChatOpenAI(model='gpt-4o'), ) await agent.run() Passing prompts to a LLM agent... waiting for the black box to run and do something...
- atrus 2y agoI mean, is that really much different than an API? I pass a query, and get data back, and rarely do I get to inspect the mechanisms behind what's returning that data.
- gitgud 2y agoOkay that’s a pretty good point actually. I guess I wrongly assume regular API’s are more reliable, but you’re right they’re basically black boxes too…
- Tepix 2y agoWell, i guess that with an LLM you kind of hope it will understand you whereas with a regular program you don't have to hope for that, just that it will work without throwing errors (usually with a magnitudes lower error rate).
- aabbcc1241 2y agoIt may give more transparency if it output the code to do scrapping with playwright. Then we can review the code before actually running it. And the reviewed result can be saved for future running.
- gregpr07 2y agoFor that we wanted to give you more control as well with a for loop and you take actions step by step. I think all of these crewai like agent swarms are also very much black boxes. How would you imagine the perfect scenario? What would make LLM outputs less of a black box?
- DeathArrow 2y ago>This enables you to design custom web automation and scraping functions without manual inspection through DevTools Can it use a headless browser?
- gregpr07 2y agoTechnically you can yeah, I am not sure if the performance is the same - would have to test it!
- ReD_CoDE 2y agoMany web developers use Playwright and Puppeteer, so why Selenium?