21 ms·
Computer Use is 45x more expensive than structured APIs
- egorfine 5mo agoStructured APIs require actual thinking and in these days thinking is frowned upon.
- deleted 5mo ago[deleted]
- taormina 5mo agoThe interface designed for humans is poor for AI needs? And the interface designed for programmatic use is easier for the AI to use? In other news, the sky is blue and water is wet.
- palashawas 5mo agoYep, everyone knows computer use is more expensive. This is about quantifying the gap
- sudb 5mo agoI'm pretty unsurprised that the vision agent did worse. I'd be interested in a comparison between the different tools that now exist to let LLMs drive browsers (e.g. vercel's agent-browser, the relatively new dev-browser[1], etc.) There are usecases where the vision agent is the more obvious, or only choice though, e.g. prorprietary/locked-down desktop apps that lack an automation layer. 1. https://github.com/SawyerHood/dev-browser https://github.com/SawyerHood/dev-browser
- palashawas 5mo agoInteresting! I'll play around with agent-browser and update this article if anything comes up
- cjbarber 5mo agoI think of computer use as like last mile delivery. APIs and bash and such are the efficient logistics networks. Both have different benefits. Obviously, use the efficient methods when you can.
- svnt 5mo ago> This is not a model problem. The vision agent was reasoning about a rendered page and had no signal that the page wasn't showing everything. > To make the comparison apples-to-apples, we rewrote the vision prompt as an explicit UI walkthrough, naming the sidebar items, tabs, and form fields the agent should interact with at each step. Fourteen numbered instructions covering the navigation the agent had failed to figure out on its own. This is a model problem, though. Because the model failed to understand it could scroll, you forced it to consume multiples of the tokens. Could you come up with an alternative here? Do you know what the vision model was trained on? Because often people see “vision model” and think “human-level GUI navigator” when afaik the latter has yet to be built.
- palashawas 5mo agoThis is a fair point. The models frequently failed for many reasons on earlier runs, and the browser-use prompt ended up being pretty granular. I'll add a couple of runs that include a scroll instruction to the repo today and see how that compares Pretty hard to guess what Anthropic trained sonnet on, but general multimodals are what people are using to drive similar tools today, whether GUI-trained or not, so the comparison still holds, for now
- aurareturn 5mo agoIn an agentic world, the OS needs to be completely rethought. For example, every single app functionality should be exposable via an API while remaining human friendly. I think OpenAI designing their own phone is the next logical step. I hope they succeed which should bring major competition to Apple and Android.
- mtoner23 5mo agoOpenai should not design a phone... They should try making money first
- sophacles 5mo agoNonsense. Don't you know how bubbles work? Everyone does massive rushes for all the low hanging and medium hanging fruit. The the bubble pops and the randomized carnage of companies big and small being destroyed is sifted through by the next wave of companies actually intended to make money. The good ideas and the bad ideas don't signal success in a bubble, nor does making money or not. Its random and any notion of "this was a good business model and that was bad" is post-hoc rationalization. The number of people who make fun of pets.com but order from chewy.com is a prime example of this.
- joshstrange 5mo ago> I think OpenAI designing their own phone is the next logical step. I hope they succeed which should bring major competition to Apple and Android. This is not going to happen, or if it does it will just be Android (like Samsung reskins/modifies it) and it will certainly use Google Play Services.
- reorder9695 5mo agoPresumably on Linux at least apps could just expose a DBus API? The machinery for this is already in place as far as I can tell.
- lazide 5mo agoThis is like insisting - after the problem turns out to be harder than thought - that the worlds roads need to be completely redone to make them self driving friendly, so self driving can work. Isn’t the whole ‘promise’ of AI that it doesn’t need any of those things?
- moralestapia 5mo agoThis is obvious. The problem is that not everything has an API, while everything has a human-oriented UI.
- palashawas 5mo agoRight - we did this benchmark because we launched a plugin that makes APIs programmatically from an app's human-oriented UI (from the event handlers, to be specific). So any app that has a human-oriented UI now has an API. The benchmark is a more generally interesting part of the launch materials, so I figured it had its own separate home here.
- moralestapia 5mo agoThat is actually great, I'll definitely check it out. Thanks!
- Havoc 5mo agoIsn't it possible to somehow wire this into the window manager? Wayland or whatever. Have it speak the native window lang rather than crunch the pixels? At least for the majority. I can see the appeal in pixel route given universality but wow that seems ugly on efficiency
- QuercusMax 5mo agoimagine, if you will, that we had a windowing system that's built on Postscript... lots of folks thought it was a super awesome idea, and built NeXTSTEP around it. https://en.wikipedia.org/wiki/Display_PostScript https://en.wikipedia.org/wiki/Display_PostScript or even one based on PDF like OSX: https://en.wikipedia.org/wiki/Quartz_2D https://en.wikipedia.org/wiki/Quartz_2D
- donaldjbiden 5mo agoWayland only has pixels. It was designed to get rid of all the X11 cruft.
- lelanthran 5mo ago> Isn't it possible to somehow wire this into the window manager? Wayland or whatever. Have it speak the native window lang rather than crunch the pixels? At least for the majority. Not possible on wayland, maybe on X11 protocol?
- _boffin_ 5mo agoWhat i don't understand about "computer use" is why they're not just grabbing the window handles and storing them to determine what should be clicked after the first few iterations of using that a specific application. if a new case / path / whatever is found, drop back to screen grabbing and bounding boxes and then figure the handles that are there and store after. idk.. not really thought out too much, but has to be better
- misja111 5mo agoexactly! This should have been part of the prompt. If you choose not to do this then you shouldn't be surprised about the high token usage.
- faangguyindia 5mo agoI saw Codex was screenshotting, then clicking around. I just stopped it and never used that again. Using CLI tools is much faster and token-efficient. I developed ten apps in the last two months. One reached 10,000+ monthly active users. I ask Codex to generate SVG line by line and backtrack edit, ask it to use Inkscape to generate icons, etc... I developed all this on $20 codex sub.
- ceejayoz 5mo agoClaude does this too, with the Chrome extension. It breaks like 80% of the time for me, and it's incredibly slow. Having it use Playwright (bonus: can test in FF/Saf too) was a big improvement.
- embedding-shape 5mo agoI think it's the third or forth time I see you bragging about HN how many apps you're able to develop with AI now. Care to link any of them, especially where we can see the actual code that you've produced here? Without being able to see actual results, I'm not sure what you want people to take away from your repeated comments.
- faangguyindia 5mo agoI only write here because people are spreading doomerism here with AI and I am excited about future. Well I am competing with geoip provider like maxmind. I developed custom traceroute and ping service to geolocate IPs with very high accuracy beating products like digital element, maxmind, ipinfo These companies have huge teams. But my 3 people company already beat them. Code doesn't matter much, it's not an opensource project. My free app is http://macrocodex.app http://macrocodex.app which I've developed along with a fitness coach. I am currently beating companies with 20-30 developers and closing more deals while having 1/10th of the staff. I am simply very excited about all this. Nobody cares show you solve the problem, or if your code is ugly. As long as it's reliable and without downtime, you aren't breaking things and causing your customer headache, you are winning. Even before AI, bad code existed. Not every company had 10x developer writing beautiful idiomatic rust code. AI is just a tool, people who are trying to generate whole codebase with it are doing something very wrong. You can write code faster with AI provided you understand its strength and weakness
- dist-epoch 5mo agoIt doesn't matter. Electron uses 10x more RAM than regular apps. But it's so convenient. Python is 100x slower than C. It's in the top 3 of languages now. Worse but more convenient always wins.
- password4321 5mo agoThis is probably why MCP "code mode" (generating code once to call the MCP going forward) hasn't caught on yet... no need until the financial costs reflect reality.
- gowld 5mo agoConfusing title? "Computer Use" is actually "Browser vision"?
- antves 5mo agoI think one main point is that not all "computer use" is the same, the harness and agentic experience matters a lot. A poorly designed API experience can actually be _less_ efficient than a well designed browser or computer use experience In particular, the vision-based approach used in the evaluation has clear limitations with regard to efficiency due to its nature (small observation window, heterogeneous modality) At Smooth we use an hybrid DOM/vision approach and we index very strongly on small models. An interesting fact is that UIs are generally designed to minimize ambiguity and supply all and only the necessary context as token-efficient as possible, and the UX is cabled up to abstract the APIs in well-understood interface patterns e.g. dropdowns or autocompletes. This makes navigation easier and that's why small models can do it, which is another dimension that must be considered We typically recommend using APIs/MCP where available and well designed, but it's genuinely surprising how token-efficient agentic browser navigation can actually be
- overgard 5mo agoI've been thinking of things I'd want an agent for recently. The problem is, everything I think of is something that requires using quite a few different websites, saving a lot of data securely, and working with a lot of sensitive accounts (my email, etc.) The problem is, all the tasks are essentially: a) things agents probably just can't do, and b) things that absolutely cannot afford to be hallucinated or otherwise fucked up. So far the tasks I've thought of: - Taxes. So it needs a lot of sensitive information to get W2's. Since I have to look up a lot of this stuff in the physical world anyway, it's not like I can just let it run wild. - Background check for a new job. It took me 3 hrs to fill out one of them (mostly because the website was THAT bad). Being myself, I already was making mistakes just forgetting things like move in dates from 10 years ago, and having to do a lot of searching in my email for random documents. No way I'm trusting an agent with this. - Setting up an LLC. Nope nope nope. There's a lot of annoying work involved with this, but I'm not trusting an LLM to do this. Anyway, I guess my point is that even if an LLM was good at using my computer (so far, it seems like it wouldn't be), the kind of things I'd want an agent for are things that an LLM can't be trusted with.
- junofan 5mo agoIt’s great at 1. things you wouldn’t otherwise bother doing 2. things where it otherwise would get stuck iterating on hacky workarounds doomed to fail “Reverse engineer this app/site so we can do $common_task in one click”, “by the way, I’m logged in to $developer_portal, so try @Browser Use if you’re stuck”, etc. I just had Codex pull user flows out of a site I’m working on and organize them on a single page. It found 116. I went in and annotated where I wanted changes, and now it’s crunching away fixing them all. Then it’ll give me an updated contact sheet and I can do a second pass. I’d never do this sort of quality pass manually and instead would’ve just fixed issues as they came up, but this just runs in the background and requires 15 minutes of my time for a lot of polish.
- overgard 5mo agoI guess the problem I see here is that if the use case is "things I otherwise wouldn't bother doing", that's fine, but it's pretty niche. I dunno, if you're talking about a human "Agent" (like say in sports or entertainment), they'd be a trusted person to handle business matters outside of your competency (contract negotiations, etc.). I don't see AI "agents" being at all like that, they're more like an intern you need to supervise constantly.
- rootcage 5mo agoThe best use cases I've seen for computer/browser use is for legacy SaaS/Software. For example, hotels use archaic Property Management Systems (PMS) and they're required by corporate to use it and pay for it. These companies can barely keep the product alive, they definitely aren't incentivized to maintain an API. In such a case browser use agent seems to be the best (only) way.
- noprocrasted 5mo agoWouldn't using a coding agent to build a screenscraper be better?
- merlindru 5mo agoI'm building something that fixes this exact problem[1]. The landing page doesn't advertise it yet, but essentially, I give agents a small set of tools to explore apps' surfaces, and then an API over common macOS functions, especially those related to accessibility. The agent explores the app, then writes a repeatable workflow for it. Then it can run that workflow through CLI: `invoke chrome pinTab` Why accessibility? Well, turns out that it's just a good DOM in general. It's structure for apps. Not all apps implement it perfectly, but enough do to make it wildly useful. [1] https://getinvoke.com https://getinvoke.com - note that the landing page is targeted towards creatives right now and doesn't talk about this use case yet
- gbriel 5mo agoThis is a good solution, instead of everyone blowing tokens on repeating the same computer use task, come up with a way to share the workflows. I think you'd need to make sure there aren't workflows shared that extract user information (passwords).
- merlindru 5mo agothis is protected against at the OS level, provided the applications declare the input correctly as a SecureTextField. i so far haven't found any application that doesn't. all you're able to get out, as far as i can tell, is the length of the entered password.
- jasomill 5mo agoFrom applications that capture the screen or use accessibility APIs, perhaps, but what about, e.g., Windows applications that capture window messages, e.g., https://devblogs.microsoft.com/cppblog/spy-internals/ https://devblogs.microsoft.com/cppblog/spy-internals/ Obviously, if you can inject code into a process that receives sensitive data, you're already running in a context where all security bets are off. But with processes you yourself create, you probably can, even without elevated privileges, unless the application takes measures to prevent injection (akin to game anticheat mechanisms), so it seems worth pointing out that there are simple mechanisms to subvert such "protected" fields that don't require application-specific reverse engineering.
- janalsncm 5mo agoWall clock time tells me everything I need to know. The vision model took almost 20 minutes to do the thing that Sonnet did in 20 seconds. The only reason you wouldn’t choose an API is if it wasn’t viable.
- volume_tech 5mo ago[flagged]
- jacktu 5mo agoTotally agree. I’ve been building an AI visual tool recently and experimented with both approaches. The latency and c ost of generic "agentic" browser use are absolute dealbreakers for real-time consumer apps right now. Structured APIs (even just chained LLM calls with strict JSON schemas) are not only 40x cheaper, but more importantly, they are deterministic enough to actually build a stable product on top of. Computer use is an amazing demo, but structured APIs are what pay the server bills.
- ai_fry_ur_brain 5mo ago"Agentic engineering" were always just FADs to bring in more revenue for token providers. If I think an LLM is good for something I create well defined, very deterministic "middleware" for that purpose on top of Openrouter.
- wahnfrieden 5mo agoIt’s not a fad or without value.
- ai_fry_ur_brain 5mo agoIts very much valuable to lazy people who dont care about quality or doing hard things. I totally see the appeal for those people.
- wahnfrieden 5mo agoSounds like you are more interested in performativity / aesthetics of production if you think writing software in a harder way is an indisputable virtue just because it requires more effort. On top of that you are an elitist about it Agent use can be used to improve quality and maintainability
- k__ 5mo agoAgentic engineers can build well defined, very deterministic middleware on top of OpenRouter. Anthropic even says, that an agent based solution should only be your last resort and that most problems are well served with a one-shot. https://www.anthropic.com/engineering/building-effective-agents https://www.anthropic.com/engineering/building-effective-age...
- zephen 5mo agoI find this extremely surprising. When you think of everything it takes for an AI to use what the article calls a "vision agent" then it seems as if using a purpose-made API ought to be MANY orders of magnitude faster.
- 2001zhaozhao 5mo agoI have only found Computer Use useful for GUI app local debugging. Presumably it will also be useful for getting around protections for external apps that don't want AI to interact with them, or for interfacing with legacy apps or those built without AI in mind. I don't think any new app should ever be specifically designed for AI to interact with them through computer use
- ai_fry_ur_brain 5mo agoIts funny watching the slow mean reversion back to more deterministic tooling.
- sanderjd 5mo agoOnly 45x?
- sheepscreek 5mo agoThis tracks - has been my experience exactly. Not to mention there isn’t particularly a significant lift in inaccuracy or speed. As things stand, to me it is the worst of both worlds. Expensive and inaccurate.
- rgilliotte 5mo ago[dead]
- jasomill 5mo agoThe nicest thing about this rush to find and build "agentic" endpoints for controlling everything is that there's no reason these same endpoints can't be consumed by deterministic, non-LLM software as well. It feels like 1994 called, and it's giving me my AppleScript back.
- orliesaurus 5mo agoComputer Use? Or Browser Use? IMHO big diff The problem is that not everything from the 'past' can be accessed via APIs. It would be a fun time - remember Prism [1] - I would just run that and get all the API calls in a nice format and then replay them over and over to do things in succession. In the new world, we have access to OpenAPI.json and whatnot, but in the world where things were built in the days pre-OpenAPI and pre-specs and best practices...I am not so sure! (and a lot of world lives then) Alas, this works for a good chunk of things but not everything. Which is why the other technnology exists. [1] https://stoplight.io/open-source/prism https://stoplight.io/open-source/prism
- Worf 5mo agoIs it possible to ask the vision agent to "map" the UI and expose it to another agent as a set of interfaces that resemble an API better? From what I understand the vision agent now should both know that "next page" shows more results and that they need to get more results in the first place. If one agent just explores the UI, maybe in a test environment, and outputs a somewhat-structured description of the various UI elements and their behavior, then another agent was given that description, would the other agent perform better that an agent that both explores the UI and tries to accomplish the given task at the same time? With an example UI I made up, the description (API-like interface definition) could be something like: Get all reviews: To get all the reviews you need to go to each page and click "show full review" for every review summary in that page. Go to each page: Start at page 1 (the default when in the Reviews tab). Continue by clicking the "next" button until the "next" button is no longer available (as you've reached the last page). So the second agent can skip some thinking about how to navigate because it already has that skill. The first agent can explore the UI on its own, once, without worrying about messing up if there's a test environment. Or am I misunderstanding the article completely? Probably. But it's interesting nonetheless. Sorry if it makes no sense.
- angry_octet 5mo agoI think you're right, you can get agents to do what we do -- learn how a website works. Then expose that model as a simple API. There will still be some vision tasks for navigation but they will be just vision tasks, no thinking required.
- nijave 5mo agoThat was my first thought as well. A lot of current web development relies heavily on code generation then has obsfuscation and compression slapped on top leading to complicated structures. Then on top of that, more code (client side/JavaScript) reconfigures everything again. You end up with fairly complicated html/css/JavaScript to wade through. For better and worse, 5-10Mi isn't uncommon for a web app. Instead of trying to go "bottom up" and, effectively, do what a browser engine is doing in reverse, it seems easier to go "top down" like a human does and go off the visual representation.
- 5mo ago
- RobRivera 5mo agoUX feedback Me: hmm, this title confuses and infuriates Rob. [Clicks link] Me: Sees same title, repeat feelings of confusion and infuration [Scrolls article down on my smartphone] Me: Sees jpg with the same title, repeat feelings of co fusion and infuriation. [Closes tab] [Continues living rest of my life] I hope this feedback is well received and understood.
- deleted 5mo ago[deleted]
- deafpolygon 5mo agoThis is missing the point that AI training probably costed boatloads more to achieve to get here.
- arjunchint 5mo agoThe hard part about the web is that API's aren't just available even if the website owner wants them exposed (big if). I embedded a Google Calendar widget on my Book a demo page, I don't know the API and Google doesn't expose/maintain one either. What we are doing at Retriever AI is to instead reverse engineer the website APIs on the fly and call them directly from within the webpage so that auth/session tokens propoagate for free: https://www.rtrvr.ai/blog/ai-subroutines-zero-token-deterministic-automation https://www.rtrvr.ai/blog/ai-subroutines-zero-token-determin...
- rahulyc 5mo agoAll the websites currently blocking Claude Code or other AI agents are fighting a losing battle. Computer-use is in the early stages, and the thing preventing mass-adoption seems to be the number of tokens it takes. Agents can fumble around trying 10 CLI commands that don't work before finding the right one and we barely notice. But other visual agents (browser use / computer use etc) end up eventually fumbling on to the right thing, but we don't have the patience to wait 20 mins. to click a button. As tokens get cheaper + faster, we probably get the models that can use a UI interface just as natively as a CLI.
- boringg 5mo agoTokens cheaper? I don't think that seems to be the case ... VC funded tokens were there to build user base and token price will go up as they eventually switch from growth to profitability.
- bheadmaster 5mo agoIt will take a few years until scheduled data center construction finishes, and together with software optimizations that may come up in the meantime, it may cause a significant decrease in token price.
- Aurornis 5mo agoI wish I could place a lot of money on the opposite side of this bet. I don't think many realize how could the cheap, alternative models are becoming. I prefer SOTA models for key work, but I can also spend 10X as many tokens on an open model hosted by a non-VC subsidized provider (who is selling at a profit) for tasks that can tolerate slightly less quality. The situation is only getting better as models improve and data centers get built out.
- boringg 5mo agoFair - there are bets both ways though I wouldn't consider it to be a certainty. That revenue drive on this AI build out is going to be real and multifold.
- johnsmith1840 5mo agoText based web browsing? Would love the comparison there. Tons of systems have a dom translation layer. I'm building around this with the concept of turn a webpage into text for an agent to use directly. I actually had to move away from haiku not because of accuracy problems but because it operated the browser too fast for a human to follow what it was doing. The real loss here are bespoke webapps like a figma or google docs which are near impossible to see what they are doing via the dom. To me the browser is a translation layer. Working on the browser directly while hard enables big advantages on compatibility. The only thing I miss as of now which is on the todo is ocr of the images in the browser into text out. But an api would need to do that anyways to work. The main loss in my view of pure API based is, where do you get the data? We won't replicate human work without seeing that done. Humans work in the UI that's it. Computer use to me is the promise of being able to replicate end to end actions a human does. API can do that in theory but the data to do that is also near impossible to collect properly.
- etothet 5mo agoVision has a long way to go. I remember trying an early version of AWS's Nova Act and laughed at how slow it was. And a few months later it hadn't really seemed to improve that much. Recently, I asked Claude to log into my local grocery store chain's website and add all of the items from my shopping list to a cart. It was hilariously slow, but it did get the job done. Unless I missed it, the article doesn't explictly mention speed in the copy, but the results do show a 17 minute (!!!) total time for the vision agent vs. 0.5s - 2.8s for the API approach. A big part of the challenge with vision is that to manipulate the DOM, you first have to be sure the entire (current) DOM is loaded. In my experience this ends up in adding a lot of artificial waits for certain elements to exist on the page.
- nijave 5mo agoWould a lightweight motion detection algorithm work there? Thinking of Frigate NVR that does motion > object detection > scene description Where you build up to progressively slower and more expensive algorithms i.e. there's motion > it's a person > here's what the person is doing
- angry_octet 5mo agoGreat guidance hidden in here for making it expensive for agents to navigate your website. Move elements on screen as the mouse moves, force natural mouse movement to make the UI work, change the button labels in the JS to be randomly named every visit, force scrolling to the bottom of the screen to check for hidden extra tasks... Hang on, that sounds like common corporate SaaS apps.
- notjustanymike 5mo agoAh damn it, we invented Jira
- fooker 5mo agoJira from first principles Almost sounds like an Orielly book
- QuantumNomad_ 5mo agoThe O’Reilly animal for Jira is apparently some kind of duck or goose. Matthew B. Doar (2011). Practical JIRA Plugins. O’Reilly. https://www.oreilly.com/library/view/practical-jira-plugins/9781449311322/ https://www.oreilly.com/library/view/practical-jira-plugins/... In case anyone was wondering. Which they probably weren’t :p
- fooker 5mo agoI'm more interested in the next volume: impractical Jira plugins
- SAI_Peregrinus 5mo agoI'm sure the doohickeycorporation folks on Reddit can come up with some.
- phatskat 5mo agoAccording to O’Reilly’s Animal Menagerie [1] it’s a King Duck [1] https://www.oreilly.com/animals.csp?x-search=duck&x-sort=oldest https://www.oreilly.com/animals.csp?x-search=duck&x-sort=old...
- theabhinavdas 5mo agoFor now.
- ipunchghosts 5mo agoI have a similar finding for a website I made that collates college town bar specials and live music. Using agents with vision models works but it's not as straightforward as one would initially think. U can check out the results here. https://www.nittanynights.com https://www.nittanynights.com
- creatonez 5mo agoBrowser agents / vision agents are a menace and ISPs should outright ban subscribers who run them on the public internet.
- lacymorrow 5mo ago[dead]
- WhoffAgents 5mo ago[flagged]
- bottlepalm 5mo agoThere's no way this is true. I would argue in some cases computer use is less expensive. First for APIs that don't even exist, it's a non starter. Second most APIs are not designed for agents and are verbose as hell - returning the entire DTO and tons of unnecessary properties burns tokens. Second computer use is not as token hungry as you think it is - a single screenshot may be just 1000 tokens, it's actually competitive and beats API workflows in many cases.
- brikym 5mo agoIt would be great if institutions like banks provided proper APIs.
- mrcwinn 5mo agoWe need a superset of HTML that is designed for agents. I'm not sure it's quite as simple as "just make everything an API."
- eggplantiny 5mo ago[dead]
- 0xWTF 5mo agoSo, to make this concrete, Akasa uses computer vision to read medical records to replace medical coders because there aren't enough medical coders to get all the billing right and medical systems leave like $1T a year on the table. The EHRs could give companies like Akasa API access so Akasa could then just run NLP, but the EHR vendors don't grant various third parties API access for various reasons, so instead Akasa gets a seat license for each medical system they service and uses computer vision to read the screen (a cadre of Akasa medical coders review errors to stay up to date with unannounced changes from the EHR vendors) and then runs the NLP to figure out which CPT codes to assign to actually put in a bill and send the payer so the hospitals can stay afloat. So this 45x delta is how much more the medical systems pay Akasa because Epic won't work with Akasa. This is but one example of why US medical bills are outrageously high.
- sarmike31 5mo agoJust wondering: RPA companies like UiPath ard dead in the water, right?
- bnyhil31-afk 5mo agoI certainly would be curious how their Agentic AI compares. On another note, if RPA has taught me anything, it's 'don't rely on the UI'.
- zhxiaoliang 5mo agoI'm always skeptical of the whole "computer use" concept. It's like hiring someone and inviting him to your house and telling him to go ahead, feel free to sleep on the bed, use the toilet, eat whatever is in the fridge, watch the TV, and oh here are the combinations for the safe... and that someone you hire is a monkey.
- eddythompson80 5mo agoBut think of how comfortable and productive the monkey will feel. It might not be that hard to just build temp housing for it while you have monkey business to do.
- andrekandre 5mo ago> build temp housing for it everyone knows the real trouble starts when the monkey asks for the vote
- nijave 5mo agoIn fairness, you're hoping the monkey does all the monkey tasks you'd rather not do yourself
- titzer 5mo agoI feel like I am taking crazy pills. Are we really having an AI fart around with a mouse and clicking on things to accomplish stuff because we're not capable of making one kind of software query and command another piece of software? It kind of boggles my mind.
- hnav 5mo agoThe writing was on the wall with the MCP->CLI jump. The promise to investors is that you're replacing people. People don't make API calls.
- ex-aws-dude 5mo agoYou are because that requires you to expose that API for every single piece of software ever
- dfee 5mo agoby design: https://en.wikipedia.org/wiki/Desire_path https://en.wikipedia.org/wiki/Desire_path IMO, this is the argument for doing work in the first place.
- game_the0ry 5mo agoMy "best practice" is to use as little "visual" (computer use) tooling and as much api + cli tooling as possible specifically to save on tokens. Tokens a resource and should be managed as such.
- zmmmmm 5mo agoAnd structured APIs are about 1e9x more expensive than not invoking an LLM in the first place compared to using deterministic code to do something ... it's not like any of this is rational based on compute.
- hnav 5mo agoIt simply doesn't fit in the token/time budget to be useful. I don't think the purveyors of these technologies care about how expensive it is as long as it's "cheap enough"
- theptip 5mo agoI’m missing the premise. For internal apps why would you ever reach for Computer Use vs just having your agent whip up a cli or MCP? _of course_ computer use is worse. It is your last resort. Do not use it on state that lives in a DB that you own. If anything I am impressed that it’s only 50x worse.
- phh 5mo agoI agree with you that it is the very obvious conclusion, but it isn't obvious to everyone. But it could still be relevant to you if you find yourself discussing with someone saying "why do we even spend money making an API, the AI can just control my computer?"
- euphetar 5mo agoI wouldn't call it a benchmark since it's just one sample. They do highlight a real problem, though. Computer use is immature right now and far behind language agents Try playing fruit ninja via text and llm toolcalls though
- danpalmer 5mo agoMetadata and structure beats AI every time.
- j45 5mo agoSounds like some efficiency gains will still arrive.
- _heimdall 5mo agoWe gave up on structured APIs 20+ years ago when JSON RPC largely replaced XML REST. You can do REST in many different formats, it mainly just needs to be structured data and self-discoverable. Had we not made that wrong turn, LLMs and humans would have a much easier time reasoning about APIs they don't directly control.
- ex-aws-dude 5mo agoYeah but why would we keep that around for 20 years with no good use case
- _heimdall 5mo agoWhy do you assume theres no good use case? trpc, grpc, etc are all attempts to add schemas back into JSON. Swagger, OpenAPI, etc are attempts to add discover ability back into JSON-based RPC APIs. MCPs fall in here as well, which attempt to add schemas and discover ability back in where our APIs aren't actually RESTful.
- zepolen 5mo ago> You can do REST in many different formats, it mainly just needs to be structured data and self-discoverable. REST has nothing to do with structured data or discoverability.
- _heimdall 5mo agoSchemaed data may be more accurate than structured. Regardless I'd be curious how you would consider something RESTful when it isn't schemaed data or discoverable via the data describing actions that can be taken.
- chrismarlow9 5mo agoBlackhat SEO spamming knew this 20 years ago
- mbgerring 5mo agoHello from the distant past, when being able to easily consume a website via API was an exciting and fresh idea for humans, before robots could effectively use the computer https://en.wikipedia.org/wiki/HATEOAS https://en.wikipedia.org/wiki/HATEOAS
- mbgerring 5mo agoDoes anyone remember the conference talk in the early days of React that was titled something like “best practices considered harmful,” or something? Or maybe that was a joke someone made about it. Anyway, the Semantic Web people have been right this whole time, and it’s very funny that we can now quantify the cost of building websites upside down and backwards for more than a decade.
- momo26 5mo ago[flagged]
- hamasho 5mo agoI'm trying to use computer use and browser use (via playwright MCP) in my work. Computer use is a hit and miss (mostly miss), but playwright MCP often works very well. The downside is it takes a lot of time to complete even easy tasks. For example, to automate processing emails, it needs to 1. go to Gmail 2. log in to Google if necessary (This often requires two step verification so it's hard to completely automating, but possible) 3. read the latest mail 4. check the content and choose the action - if needed, reply the email - if it mentions tasks, add them to the todo list - if it mentions schedules, add them to the calendar 5. repeat for all emails based on specified conditions. And each step requires dozens of DOM (a11y tree) analyzes and actions (fill username/password input, check keep logging in, click submit button, etc). Based on the model used, each step can take ~100s. So easy tasks can easily add up to tens of minutes or even hours. For frequently used tasks, I write skills like /logging-in, /read-latest-emails, using playwright scripts and let the agent choose them And based on the email content, the agent chooses other tools like /write-reply, /add-todo, /add-event, etc, so that the model can only focus on the core tasks requiring thinking. It reduces the execution time drastically. But it can buries important business logic in the playwright scripts, not the agent's instructions. For examples, simplified steps to add TODO items are like; 1. read the email 2. check if it's about todos, then decide to add them to Asana 3. extract and summarize the title, content, priority, due date, tags, etc. 3. access to Asana (log in if necessary) 4. check if there are similar tasks 5. if not, add the tasks This can take tens of minutes, and each step can have important business logic, like; - how to decide the priority and due date - how to choose tags based on the content - how to decide if two tasks are similar This information should be read and updated by not only developers, but managers and other teams. And if I write those steps in skills with playwright scripts, it improves the speed, but all those business logic are buried in the code, so not accessible by non-technical people. It's also error-prone because web sites often tweak the UI and scripts can stop working. So it's very convenient if the agent processes these step once, then decides it's worth writing the playwright script so that the next time those mundate processs can be executed instantly. With automatic skill generation, the agent decides by itself if there are workflows worth writing skills with playwright scripts, like /log-in, /extract-information, /check-similar-tasks, /add-tasks. Like Just-In-Time compiler, the skills are a byproduct of the agent instruction, all business logic are written in the agent's instruction, and doesn't need to be updated manually nor tracked in a version control system. This can reduce a lot of execution time and API cost, and be applied other than browser automation, like computer use or any other agentic tasks if it's possible to write automation scripts for tasks not requiring thinking.
- Amber-chen 5mo ago[dead]
- overlord1109 5mo ago[dead]
- jasomill 5mo agoIn what world would a vision agent be the default, when whatever HTTP-based mechanism a site uses to communicate with the server can usually be reverse-engineered and easily emulated with widely available HTTP request libraries, HTML parsers, and JavaScript engines, and at worst you can use something like Puppeteer to navigate and control applications at a significantly higher level than image scraping and simulating user input? It seems like you'd need a deliberately hostile app before a vision agent would even be considered as an option.
- doctorpcgum 5mo agoBh
- doctorpcgum 5mo ago[flagged]
- morpheos137 5mo agoWho would have thunk? You know what is a great LLM agent api? bash. vast corpus, text based, already traindd in the model.
- Frannky 5mo agoI want to just talk to the Mac and have it do things. I tried computer use and other alternatives, but the latency made it unusable. I want to be able to control both Mac, apps and the browser. I also need it to figure out things by itself given a goal. Claude Code with the --chrome flag is kind of good, but it's too slow. I wanted to try faster APIs, like the one hosted on Cerebras, but it's too expensive. Any solution I might be missing?
- jasomill 5mo agoDo you want to do something that can't be done through AppleScript, macOS accessibility APIs, and something like Puppeteer to control the browser? Or something you don't understand how to do manually? Because I guess I don't understand the attraction of using an LLM for system automation where existing interfaces exist, other than as a form of documentation, or to write code using these interfaces.
- Frannky 5mo agoI don't want to think about that stuff. I just want to ask and get stuff done.
- m3kw9 5mo agoI did a simple computer use to search something, and used up 50% of my 5h plan limit from codex.
- jacktu 5mo ago[dead]
- oleg2025 5mo agoCouple of months ago I was inspired by kubectl, and built desktopctl CLI to control GUI apps. It uses combination of OCR and Accessibility API on Mac, represents UI as markdown, and exposes actions for mouse and keyboard. My core idea was that "fast" perception loop is fully local, GPU optimised for UI tokenisation and change detection. "Slow" control loop requires LLM roundtrip, and uses token-efficient markdown interface in CLI output. It uses relatively stable identifiers for controls, so agents can script common actions, eg `desktopctl pointer click --id btn_save` doesn't require UI tokenisation loop. https://github.com/yaroshevych/desktopctl/tree/main https://github.com/yaroshevych/desktopctl/tree/main
- oleg2025 5mo agoI've learned that compared to APIs, human interfaces are slow and messy, but there is actually a lot of science behind them. The good apps expose information well, and are optimised for clicks, typing, etc. The best GUIs make great use of muscle memory, which makes them perfect candidates for scripting via CLI. eg a simple sequence "open Notes app, hit Cmd+F, enter search term, read list of results" can be one Bash command invoked by AI agent.
- sneefle 5mo ago[flagged]
- BionicAI 5mo ago[flagged]
- azyc 5mo ago[flagged]
- RadiozRadioz 5mo ago> The alternative, writing an MCP or REST surface per app, is its own engineering project Well, if your backend was sufficiently decoupled from your frontend, and the server-side operations were designed thoughtfully and generically, it need not be an engineering project.
- jgalt212 5mo agoComputer Use and Agentic are primarily about (token) number go up.
- theuniverseson 5mo ago[flagged]
- netics01 5mo agoLike MCP, but it seems most AI tools these days are at the stage where they're being built so that AI can somehow operate or control them. Over time, I expect that not only Computer Use, but pretty much all other tools will eventually just become APIs.
- deleted 5mo ago[deleted]
- ggsa 5mo ago[flagged]
- amai 5mo agoWhy using Claude Sonnet for navigating the gui? You could also use selenium, playwright or some other headless browser instead. Would be much cheaper and more reproducible.
- danborn26 5mo agoThis aligns with what I've seen in production. Forcing an LLM to navigate a DOM is an incredible waste of compute when an API is available.
- westurner 5mo agoOther examples of where non-LLM methods are and/or will always be more efficient than LLM methods?
- steffs 5mo ago[flagged]