8 ms·
Claude Computer Use – Is Vision the Ultimate API?
- simonw 2y agoIf you want to try out Computer Use (awful name) in a relatively safe environment the Docker container Anthropic provide here is very easy to start running (provided you have Docker setup, I used it with Docker Deaktop for Mac): https://github.com/anthropics/anthropic-quickstarts/tree/main/computer-use-demo https://github.com/anthropics/anthropic-quickstarts/tree/mai...
- trq_ 2y agoYes that's a good point! To be honest, I felt that I wanted to try it on the machine I used every day, but it's definitely a bit risky. Let me link that in the article.
- danielbln 2y agoI for one appreciate the name Computer Use, no flashy marketing name, just describes what is she's. LLM using a computer.
- sharpshadow 2y agoIn this context Windows Recall makes total sense now from a AI learning perspective for them. It’s actually a super cool development and I’m very exiting already to let my computer use any software like a pro infront of me. Paint me canvas of a savanna sunset with animals silhouette, produce me a track of uk garage house, etc. everything with all the layers and elements in the software not just an finished output.
- croes 2y agoLots of energy consumption just to create a remix of something that already exists.
- sharpshadow 2y agoAbsolutely we need much much more energy and many many more powerful chips. Energy is a resource and we need to harvest more of it. I don’t understand why people make a point about energy consumption as it would be something bad.
- viraptor 2y agoCome on, the trolling is too obvious.
- sharpshadow 2y agoAbsolutely not I’m serious and that is exactly what is going on, it would be trolling to pretend the opposite or accepted the status quo as final. Obviously we need to use a way to not harm and destroy our environment more but we are on a good way on that. But technically we need much much more energy.
- croes 2y agoWe are not on a good way, we are far away from our goals and at the moment AI‘s is near the same as Bitcoin. We do things fast and expensive that could be done slow but cheap. The problem is we are running out of time. If you want more energy you first build clean energy sources then you can pump up consumption not the other way around.
- CharlieDigital 2y agoVision is the ultimate API. The historical progression from text to still images to audio to moving images will hold true for AI as well. Just look at OpenAI's progression as well from LLM to multi-modal to the realtime API. A co-worker almost 20 years ago said something interesting to me as we were discussing Al Gore's CurrentTV project: the history of information is constrained by "bandwidth". He mentioned how broadcast television went from 72 hours of "bandwidth" (3 channels x 24h) per day to now having so much bandwidth that we could have a channel with citizen journalists. Of course, this was also the same time that YouTube was taking off. The pattern holds true for AI. AI is going to create "infinite bandwidth".
- ToDougie 2y agoSo long as the spectrum is open for infinity, yes. Was listening to a Seth Godin interview where he pointed out that there was a time when you had to purchase a slice of spectrum to share your voice on the radio. Nowadays you can put your thoughts on a platform, but that platform is owned by corporations who can and will put their thumb on thoughtcrime or challenges. I really do love your comment. Cheers.
- CharlieDigital 2y agoThanks! There's a related concept as well which is that as "bandwidth" increases, the ratio of producers to consumers pushes upwards towards 1. My take is that generative AI will accelerate this I write a bit more in depth about it here: https://chrlschn.dev/blog/2024/10/im-a-gen-ai-maximalist-and-why-you-should-be-too/ https://chrlschn.dev/blog/2024/10/im-a-gen-ai-maximalist-and...
- ricardo81 2y agoYou could call it bandwidth, or call it entropy. I'd lean towards the more physical definition. I think of how the USA had cable TV and hundreds of channels projecting all kinds of whatever in the 80s while here in the UK we were limited to our finite channels. To be fair those finite channels gave people something to talk about the next day, because millions of people saw the same thing. Surely a lot of what mankind has done is to tame entropy, like steam engines etc. With AI and everyone having a prompt, it's surely a game changer. How it works out, we'll see.
- viraptor 2y agoSome time ago I made a prediction that accessibility is the ultimate API for the UI agents, but unfortunately multimodal capabilities went the other way. But we can still change the course: This is a great place for people to start caring about accessibility annotations. All serious UI toolkits allow you to tell the computer what's on the screen. This allows things like Windows Automation https://learn.microsoft.com/en-us/windows/win32/winauto/entry-uiauto-win32 https://learn.microsoft.com/en-us/windows/win32/winauto/entr... to see a tree of controls with labels and descriptions without any vision/OCR. It can be inspected by apps like FlauiInspect https://github.com/FlaUI/FlaUInspect?tab=readme-ov-file#main-screen https://github.com/FlaUI/FlaUInspect?tab=readme-ov-file#main... But see how the example shows a statusbar with (Text "UIA3" "")? It could've been (Text "UIA3" "Current automation interface") instead for both a good tooltip and an accessibility label. Now we can kill two birds with one stone - actually improve the accessibility of everything and make sure custom controls adhere to the framework as well, and provide the same data to the coming automation agents. The text description will be much cheaper than a screenshot to process. Also it will help my work with manually coded app automation, so that's a win-win-win. As a side effect, it would also solve issues with UI weirdness. Have you ever had windows open something on a screen which is not connected anymore? Or under another window? Or minimised? Screenshots won't give enough information here to progress.
- tomatohs 2y ago> It is very helpful to give it things like: - A list of applications that are open - Which application has active focus - What is focused inside the application - Function calls to specifically navigate those applications, as many as possible We’ve found the same thing while building the client for testdriver.ai. This info is in every request.
- pabe 2y agoI don't think vision is the ultimate API. It wasn't with "traditional" RPA and it won't with more advanced AI-RPA. It's inefficient. If you want something to be used by a bot, write an interface for a bot. I'd make an exception for end2end testing.
- Veen 2y agoYou're looking at it from a developer's perspective. For non-developers, vision opens up all sorts of new capabilities. And they won't have to rely on the software creator's view of what should be automated and what should not.
- skydhash 2y agoMost non-developers won't bother. You have shortcut on iOS and macOS which is like Scratch for automation and still only power users use it. Others just download the shortcut they want.
- croes 2y agoIf a GUI is confusing for humans AI will be have problems too. So you still rely on developers to make reasonable GUIs
- cheevly 2y agoNo, language is the ultimate API.
- ukuina 2y agoOn the instruction-provision end, sure.
- throwup238 2y agoVision plus accessibility metadata is the ultimate API. I see little reason that poorly designed flat UIs are going to confuse LLMs any less than humans, especially when they’re missing from the training data like most internal apps or the documentation on the web is out of date. Even a basic dump of ARIA attributes or the hierarchy from OS accessibility APIs can help a lot.
- dbish 2y agoThe problem is accessibility data and apis are very bad across the board.
- unglaublich 2y agoVision here means "2d pixel space". The ultimate API is "all the raw data you can acquire from your environment".
- layer8 2y agoFor a typical GUI, the “mental model” actually needs to be 2.5D, due to stacked windows, popups, menus, modals, and so on. The article mentions that the model has difficulties with those.
- PreInternet01 2y agoCounterpoint: no, it's just more hype. Doing real-time OCR on 1280x1024 bitmaps has been possible for... the last decade or so? Sure, you can now do it on 4K or 8K bitmaps, but that's just an incremental improvement. Fact is, full-screen OCR coupled with innovations like "Google" has not lead to "ultimate" productivity improvements, and as impressive as OpenAI et al may appear right now, the impact of these technologies will end up roughly similar. (Which is to say: the landscape will change, but not in a truly fundamental way. What you're seeing demonstrated right now is, roughly speaking, the next Clippy, which, believe it or not, was hyped to a similar extent around the time it was introduced...)
- acchow 2y ago"OCR : Computer Use" is as "voice-to-text : ChatGPT Voice"
- simonw 2y agoThe way these new LLM vision models work is very different from OCR. I saw a demo this morning of someone getting Claude to play FreeCiv (admittedly extremely badly): https://twitter.com/greggyb/status/1849198544445432229 https://twitter.com/greggyb/status/1849198544445432229 Try doing that with Tesseract.
- croes 2y agoI bet Tesseract plays pretty badly too.
- KoolKat23 2y agoExisting OCR is extremely limited and requires custom narrow development.
- throwaway19972 2y agoI'd imagine you'd get higher quality leveraging accessibility integrations.
- echoangle 2y agoAm I the only one thinking this is an awful way for AI to do useful stuff for you? Why would I train an AI to use a GUI? Wouldn’t it be better to just have the AI learn API docs and use that? I don’t want the AI to open my browser, open google maps and search for Shawarma, I want the AI to call a google api and give me the result.
- voiper1 2y agoSure, it's more effecient to have it use an API. And people have been integrating those for the last while. But there's tons of applications that are locked behind a website or deckstop GUI with no API that are accessible via vision.
- famouswaffles 2y agoThe vast majority of Applications cannot be used by anything other than a GUI. We built computers to be used by humans and humans overwhelmingly operate computers with GUIs. So if you want a machine that can potentially operate computers as well as humans then you're going to have to stick to GUIs. It's the same reason we're trying to build general purpose robots in a human form factor. The fact that a car is about as wide as a two horse drawn carriage is also no coincidence. You can't ignore existing infrastructure.
- echoangle 2y agoBut I don’t want an AI to „operate a computer“… maybe I’m missing the point of this but I just can’t imagine a usecase where this is a good solution. For everything browser based, the burden of making an API is probably relatively small and if the page is simple enough, you could maybe even get away with training the AI on the page HTML and generating a response to send. And for everything that’s not browser-based, I would either want the AI embedded in the software (image editors, IDEs…) or not there are all.
- famouswaffles 2y ago>But I don’t want an AI to „operate a computer“… You don't and that's fine but certainly many people are interested in such a thing. >maybe I’m missing the point of this but I just can’t imagine a usecase where this is a good solution. If it could operate computers robustly and reliabily then why wouldn't you ? Not everything someone does on a computer is a task they wouldn't like to automate away but can't with current technology. >For everything browser based, the burden of making an API is probably relatively small It's definitely not less effort than stricking to a GUI >and if the page is simple enough, you could maybe even get away with training the AI on the page HTML and generating a response to send. Sure in special circumstances, it may be a good idea to use something else. >And for everything that’s not browser-based, I would either want the AI embedded in the software (image editors, IDEs…) or not there are all. AI embedded in software and AI operating the computer itself are entirely different things. The former is not necessarily a substitute for the latter. Having access to SORA is not at all the same thing as AI that can expertly operate Blender. And right now at least, studios would actually much prefer the latter. Even if they were equivalent (they're not), then you wouldn't be able to operate most applications without developers explicitly supporting and maintaining it first. That's infeasible.
- downWidOutaFite 2y agoVision is a crappy interface for computers but I think it could be a useful weapon against all the extremely "secure" platforms that refuse to give you access to your own data and refuse to interoperate with anything outside their militarized walled gardens.
- bev-erage 2y ago[dead]
- m3kw9 2y agoNo, Vision in this case is a brute force way for the AI to interact with our current world because we designed the interface for human vision. In the future, AI creates the UI and their control will be low level most likely at the model level as even business logic+UI will be generated live.
- freediver 2y agoAnd text is the ultimate API to human brain! ;) https://www.youtube.com/watch?v=Zctp972y_Eg https://www.youtube.com/watch?v=Zctp972y_Eg