Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gregpr07
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
gregpr07
2y ago
Nice!! We are working on a higher level library over at https://github.com/gregpr07/browser-use .
62.
▲
by
gregpr07
2y ago
Haha with browser use or Computer use??
63.
▲
by
gregpr07
2y ago
Technically you can yeah, I am not sure if the performance is the same - would have to test it!
64.
▲
by
gregpr07
2y ago
Who (or still in stealth haha)?
65.
▲
by
gregpr07
2y ago
For that we wanted to give you more control as well with a for loop and you take actions step by step. I think all of these crewai like agent swarms are also very much black boxes. How would you imagine the perfect scenario? What would make
66.
▲
by
gregpr07
2y ago
Do you think they do any super fancy magic other than for example how ferret ui does their classification of ui elements? It could be very interesting to test head to head hope much better you can make computer use by adding html (it’s much
67.
▲
by
gregpr07
2y ago
Nope: 1280x1024 low resolution with gpt-4o are 85 tokens so approx $0.0002 (so 100x cheaper). For high resolution its apporx $0.002 https://openai.com/api/pricing/
68.
▲
by
gregpr07
2y ago
Hmm, but this how we handle it? We just have a CLI that outputs exactly, goal, state, and asks user for more clarity if needed, no GUI. The original idea was to make it completely headless.
69.
▲
by
gregpr07
2y ago
Could you elaborate on the CLI idea? I am intrigued but not exactly sure what you mean.
70.
▲
by
gregpr07
2y ago
I don’t know a lot about this but do you have full power of Selenium or not? That would be also very interesting aproach especially when “local” browser models get very good
71.
▲
by
gregpr07
2y ago
Will def try it.
72.
▲
by
gregpr07
2y ago
Fair, also Claude probably only gets better on this since they kinda want people to use Computer use. We are gonna try to do best of both worlds. Thanks man, Magnus came up with it this morning haha!
73.
▲
by
gregpr07
2y ago
Thanks man, starred yours too, it's super cool to see all these projects getting spun up! I see Cerebellum is vision only. Did you try adding HTML + screenshot? I think that improves the performance like crazy and you don't have t
74.
▲
by
gregpr07
2y ago
Looks nice. I find the cleaning HTML step in our cleaning pipeline extremely important, otherwise there is no real benefit from just using a general vision model and clicking coordinates (and whole HTML is just way too many tokens). How do
75.
▲
by
gregpr07
2y ago
Thanks! Have you tried captcha solving with [1]? It's very tricky sometimes, especially with non standard "verify human" - maybe you could solve it by writing Selenium/Javascript code directly and then execute it.
76.
▲
by
gregpr07
2y ago
You can run browser use in your terminal, no need for Docker containers. Just clone it and run it
77.
▲
by
gregpr07
2y ago
A) we plan on thoroughly testing that with Mind2Web dataset. They have a very robust set of (persistant) selectors B) so, shadcn for prompts for web agents haha :) but I agree, that would be SICK! Just go to browseruse and get the prompt fo
78.
▲
by
gregpr07
2y ago
Not sure, I think there is a lot of research being done here. Actually, browser use works quite well with vision turned off, it just sometimes gets stuck at some trivial vision tasks. The interesting thing is that screenshot approach is oft
79.
▲
by
gregpr07
2y ago
No I think that’s a completely different beast, this is only for html/websites. But would be interesting to see what happenes with our pipeline with pure vision model. Did you mean something else?
80.
▲
by
gregpr07
2y ago
Yes, this plus reasoning and ask user for additional info. More here https://github.com/gregpr07/browser-use/blob/main/src/agent/...
81.
▲
by
gregpr07
2y ago
Thanks you, love the feedback! Will add the license. Let me know how it goes if you try it.
82.
▲
Show HN: I wrote an open-source browser alternative for Computer Use for any LLM
(github.com)
180 points
by
gregpr07
2y ago
|
72 comments