4 ms·
Is it decided then that screenshots are better input for LLMs than HTML, or is that still an active area of investigation? I see that y'all elected for a mostly
by firejake308 2y ago
Is it decided then that screenshots are better input for LLMs than HTML, or is that still an active area of investigation? I see that y'all elected for a mostly screenshot-based approach here, wondering if that was based on evidence or just a working theory.
- maggreenWAI 2y agoNext step is to represent also the structure of the HTML tree in the extracted elements for better understanding, maybe images are then less needed.
- gregpr07 2y agoNot sure, I think there is a lot of research being done here. Actually, browser use works quite well with vision turned off, it just sometimes gets stuck at some trivial vision tasks. The interesting thing is that screenshot approach is often cheaper than cleaned up html, because some websites have HUGE action spaces. We looked at some papers (like ferret ui) but i think we can do much better on html tasks. Also, there is a lot of space to improve the current pipeline.
- arjvik 2y agoWould be really cool if you could tie this into Claude's computer use APIs!
- gregpr07 2y agoDo you think they do any super fancy magic other than for example how ferret ui does their classification of ui elements? It could be very interesting to test head to head hope much better you can make computer use by adding html (it’s much better from our quick testing, just don’t know the numbers).
- stuckkeys 2y agoMy wallet just ran way. Come back you!
- gregpr07 2y agoHaha with browser use or Computer use??
- refulgentis 2y agoPer research across companies, both help, screenshots are worse, but marginally. The computer use stuff gets me fired up enough that I end up always sharing this, even though when delivered concisely without breaking NDAs, it can sound like a hot take: The whole thing is a dead end. I saw internal work at a FAANG on this for years, and even in the case where the demo is cooked up to "get everything right", intentionally, to figure out the value of investing in chasing this further...its undesirable, for design reasons. It's easy to imagine being wow'd by the computer doing something itself, but when its us, its a boring and slow way to get things done thats scary to watch. Even with the stilted 100% success rate, our meatbrains cheerily emulated knowing its < 100%, the fear is akin to watching a toddler a month into walking, except if the toddler had your credit card and a web browser and instructions to buy a ticket. I humbly and strongly suggest to anyone interested in this space to work towards CLI versions of this concept. Now, you're nonblocking, are in a more "native" environment for the LLM, and are much cheaper. If that sounds regressive and hardheaded, Microsoft, in particular, has plenty of research on this subject, and there's a good amount from diverse sources. Note the 20%-40% success rates they report, then, note that completing a full task successfully represents a product series of 20%-40%. To get an intuition for how this affects the design experience, think how annoying it is to have to repeat a question because Siri/Assistant/whatever voice assistant don't understand it, and they have roughly ~5 errors per 100 words.
- gregpr07 2y agoCould you elaborate on the CLI idea? I am intrigued but not exactly sure what you mean.
- refulgentis 2y ago(handwaving) I'd rather be in a loop of "here's our goal. here's latest output from the CLI. what do we type into the CLI" than the GUI version of that loop. I hope that's clearer, I'm a bit over-caffeinated
- gregpr07 2y agoHmm, but this how we handle it? We just have a CLI that outputs exactly, goal, state, and asks user for more clarity if needed, no GUI. The original idea was to make it completely headless.
- its_down_again 2y agoScreenshots aren't as accurate or context-rich as HTML, but they let you bypass the hassle of building logic for permissions and authentication across different apps to pull in text content for the LLM.
- EGreg 2y agoCan’t you just make a browser extension to haveaccess to the HTML and CSS, and use LLMs from that?
- maggreenWAI 2y agoContext length + API cost is right now main bottleneck for huge HTML + CSS files. The extraction here is already quite efficient but still: with past messages + system prompt + sometimes extracted text + extracted interactive elements you are quickly already around 2500 tokens (for gpt-4o 0.01$). If you extract entire HTML and CSS your cost + inference time are quickly 10x.
- EGreg 2y agoAren't screenshots far larger than this?
- gregpr07 2y agoNope: 1280x1024 low resolution with gpt-4o are 85 tokens so approx $0.0002 (so 100x cheaper). For high resolution its apporx $0.002 https://openai.com/api/pricing/ https://openai.com/api/pricing/
- stuckkeys 2y agoYeah. I noticed a very low cost when I run it via vm, predefined resolution. Good tip.
- e-clinton 2y agoI do this for my extension [0] but the HTML is often too large for context window sizes . I end up doing scraping of the relevant pieces before sending to LLM. [0] https://chromewebstore.google.com/detail/namebrand-check-for-amazo/jacmhjjebjgliobjggngkmkmckakphel https://chromewebstore.google.com/detail/namebrand-check-for...
- EGreg 2y agoI doubt screenshots would be better input considering that eg <select> box options and other markup are hisden visually until a user interacts with something
- teaxio 2y ago[dead]
- elpalek 2y agoTried Agent-E, like the DOM distillation to reduce token method. I've started using AgentQL for a new project, which works well combining playwright.