43 ms·
Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
- tylerchilds 2y agocomputer use is really going to highlight how fragmented the desktop ecosystem is, but also this definitely paints more context on how microsoft wants to use their screenshot ai
- ta93754829 2y agoeventually, we'll be able to eliminate the intermediate "computer", and just let the ai render everything we need to interact with
- freetonik 2y agoFascinating. Though I expect people to be concerned about privacy implications of sending screenshots of the desktop, similar to the backlash Microsoft has received about their AI products. Giving the remote service actual control of the mouse and keyboard is a whole another level! But I am very excited about this in the context of accessibility. Screen readers and screen control software is hard to develop and hard to learn to use. This sort of “computer use” with AI could open up so many possibilities for users with disabilities.
- minimaxir 2y agoThe key difference is that Microsoft Recall wasn't opt-in.
- sharkjacobs 2y agoThere's such a gulf between choosing to send screenshots to Anthropic and Microsoft recording screenshots without user intent or consent.
- swalsh 2y agoI suspect businesses will create VDI's or VM's for this express purpose. One because it scales better, and 2 because you can control what it has access to easier and isolate those functions.
- abrichr 2y ago> I expect people to be concerned about privacy implications of sending screenshots of the desktop That's why in https://github.com/OpenAdaptAI/OpenAdapt https://github.com/OpenAdaptAI/OpenAdapt we've built in several state-of-the-art PII/PHI scrubbers.
- KingOfCoders 2y agoI have been a paying ChatGPT customer for a long time (since the very beginning). Last week I've compared ChatGPT to Claude and the results (to my eye) were better, the output better structured and the canvas works better. I'm on the edge of jumping ship.
- J_Shelby_J 2y agoClaude is the daily driver. GPT-O1 for complicated tasks. For example, questions where linear reasoning is not enough like advanced rust ownership questions.
- postalcoder 2y agoFor python, at least, Sonnet’s code is much more elegant, well composed, and thoughtfully written. It also seems to be biased towards more recent code, whereas the gpt models can’t even properly write an api call to itself. o1 is pretty decent as a rotor rooter, ie the type of task that requires both lots of instruction as well as lots of context. I honestly think it works half as well as it does now because it’s able to properly mull through the true intent of the user that usually takes the multiple shots that nobody has the patience to do.
- pseudosavant 2y agoIt is appalling how bad GPT-4o is at writing API calls to OpenAI using Python. It is like OpenAI doesn't update their own documentation in the GPT-4o training data since GPT-3.5. I constantly have the problem that it thinks it needs to write code for the 0.28 version of the SDK. It'll be writing >1.0 code revision after revision, and then just randomly fall back to the old SDK which doesn't work at all anymore. I always write code for interfacing with OpenAI's APIs using Claude.
- joshdavham 2y ago> I'm on the edge of jumping ship. Yeah I think I might also jump ship. It’s just that chatGPT now kinda knows who I am and what I like and I’m afraid of losing that. It’s probably not a big deal though.
- crazystar 2y agoLooks like it just takes a screenshot and can't scroll so it might miss things. Claude 3.5 Haiku will be released later this month.
- freetonik 2y agoIt can actually scroll.
- crazystar 2y agoWhile we expect this capability to improve rapidly in the coming months, Claude's current ability to use computers is imperfect. Some actions that people perform effortlessly—scrolling, dragging, zooming—currently present challenges for Claude and we encourage developers to begin exploration with low-risk tasks.
- artur_makly 2y agoCan someone please try this on a MAC/OS and just 100% verify if this puppy can scroll or not? thnks
- nilsherzig 2y agoIt does in the video. Just not the spreadsheet at the start.
- minimaxir 2y agoFrom the computer use video demo, that's a lot of API calls. Even though Claude 3.5 Sonnet is relatively cheap for its performance, I suspect computer use won't be. It's a very good idea that Anthropic upfront that it isn't perfect. And it's guaranteed that there will be a viral story where Claude will accidentally delete something important with it. I'm more interested in Claude 3.5 Haiku, particularly if it is indeed better than the current Claude 3.5 Sonnet at some tasks as claimed.
- Hizonner 2y agoIt's just bizarre to force a computer to go through a GUI to use another computer. Of course it's going to be expensive.
- hobofan 2y agoWith UIPath, Appian, etc. the whole field of RPA (robotic process automation) is a $XX billion industry that is built on that exact premise (that it's more feasible to do automation via GUIs than badly built/non-existing APIs). Depending on how many GUI actions correspond to one equivalent AI orchestrated API call, this might also not be too bad in terms of efficiency.
- Hizonner 2y agoMost of the GUIs are Web pages, though, so you could just interact directly with an HTTP server and not actually render the screen. Or you could teach it to hack into the backend and add an API... Oh, and on edit, "bizarre" and "multi-billion-dollar-industry" are well known not to be mutually exclusive.
- famouswaffles 2y ago>Most of the GUIs are Web pages, though, so you could just interact directly with an HTTP server and not actually render the screen. The end goal isn't just web pages (And i wouldn't say most GUIs are web pages). Ideally, you'd also want this to be able to navigate say photoshop or any other application. And the easier your method can switch between platforms and operating systems the better We've already built computer use around GUIs so it's just much easier to center LLMs around them too. Text is an option for the command line or the web but this isn't an easy option for the vast majority of desktop applications, nevermind mobile. It's the same reason general purpose robots are being built into a human form factor. The human form isn't particularly special and forcing a machine to it has its own challenges but our world and environment has been built around it and trying to build a hundred different specialized form factors is a lot more daunting.
- netcraft 2y agoim unclear, is haiku supposed to be similar to 4o-mini in usecase/cost/performance? If not, do they have an analog?
- machiaweliczny 2y agoProbably better than 4o-mini, 4o-mini isn’t great from my testing. loses focus after 100 lines of text
- usaar333 2y agoIt's roughly tied in benchmarks
- robertkoss 2y agoDoes anyone know how I could check whether my Claude Sonnet version that I am using in the UI has been updated already?
- lambdaba 2y agosearch for "20241022" in network tab in devtools, confirmed for me
- nilsherzig 2y agoThe ui shows a (new) next to the model name for me (free user, Germany)
- diggan 2y agoI still feel like the difference between Sonnet and Opus is a bit unclear. Somewhere on Anthropic's website it says that Opus is the most advanced, but on other parts it says Sonnet is the most advanced and also the fastest. The UI doesn't make the distinction clear either. Then on Perplexity, Perplexity says that Opus is the most advanced, compared to Sonnet. And finally, in the table in the blogpost, Opus isn't even included? It seems to me like Opus is the best model they have, but they don't want people to default using it, maybe the ROI is lower on Opus or something? When I manually tested it, I feel like Opus gives slightly better replies compared to Sonnet, but I'm not 100% it's just placebo.
- smallerize 2y agoOpus has been stuck on 3.0, so Sonnet 3.5 is better for most things as well as cheaper.
- diggan 2y ago> Opus has been stuck on 3.0, so Sonnet 3.5 is better So for example, Perplexity is wrong here implying that Opus is better than Sonnet? https://i.imgur.com/N58I4PC.png https://i.imgur.com/N58I4PC.png
- hobofan 2y agoI think as of this announcement that is indeed outdated information.
- diggan 2y agoSo Opus that costs $15.00/$75.00 for 1mil tokens (input/output) is now worse than the model that costs $3.00/$15.00? That's according to https://docs.anthropic.com/en/docs/about-claude/models https://docs.anthropic.com/en/docs/about-claude/models which has "claude-3-5-sonnet-20241022" as the latest model (today's date)
- 2y ago
- jatins 2y agoHow does the computer use work -- Is this a desktop app they are providing that can do actions on your computer? Didn't see any such mention in the post
- ZiiS 2y agoIt is a docker container providing a remote desktop you can see; they strongly recomend you also run it inside a VM.
- minimaxir 2y agoQuickstart is here: https://github.com/anthropics/anthropic-quickstarts/tree/main/computer-use-demo https://github.com/anthropics/anthropic-quickstarts/tree/mai...
- thundergolfer 2y agoIt’s a sandbox compute environment, using Gvisor or Firecracker or similar, which exposes a browser environment to the LLM. modal.com’s modal.Sandbox can be the compute layer for this. It uses Gvisor under the hood.
- dtquad 2y agoIs there any Python/Node.js library to easily spawn secure isolated compute environments, possibly using gvisor or firecracker under the hood? This could be useful to build a self-hosted "Computer use" using Ollama and a multimodal model.
- eperot 2y agoI have been [working on one](https://github.com/EtiennePerot/safe-code-execution https://github.com/EtiennePerot/safe-code-execution)! The library is in [src/safecode/sandbox.py](https://github.com/EtiennePerot/safe-code-execution/blob/master/src/safecode/sandbox.py https://github.com/EtiennePerot/safe-code-execution/blob/mas...).
- abrichr 2y agoSee https://github.com/OpenAdaptAI/OpenAdapt https://github.com/OpenAdaptAI/OpenAdapt for an open source alternative that includes a desktop app.
- hugocbp 2y agoGreat work by Anthropic! After paying for ChatGPT and OpenAI API credits for a year, I switched to Claude when they launched Artifacts and never looked back. Claude Sonnet 3.5 is already so good, specially at coding. I'm looking forward to testing the new version if it is, indeed, even better. Sonnet 3.5 was a major leap forward for me personally, similar to the GPT-3.5 to GPT-4 bump back in the day.
- Axsuul 2y agoHow are you using it with coding?
- hugocbp 2y agoUsually I create a Project in the UI, upload some files I think might be relevant, and just start asking things like refactoring, how can it improve the code, how to test (or which edge cases might be missing in the test files). Once we get going, I start asking how can we change the code to do what I need to do, etc. After the history gets too long and Claude starts bugging me about limits, I ask it to summarize the context of the whole conversation, and add that to the Project and start a new chat.
- bbor 2y agoOk I know that we're in the post-nerd phase of computers, but version numbers are there for a reason. 3.6, please? 3.5.1??
- HanClinto 2y agoWhy not rev the numbers? "3.5" vs. "3.5 New" feels weird -- is there a particular reason why Anthropic doesn't want to call this 3.6 (or even 3.5.1)?
- GaggiX 2y agoSimilar to OpenAI when they update their current models they just update the date, for example this new Claude 3.5 Sonnet is "claude-3-5-sonnet-20241022".
- nisten 2y agoFor a company selling intelligence, that's a pretty stupid way of labelling a new product.
- riffraff 2y ago"computer use" is also as bad a marketing choice as possible for something that actually seems pretty cool.
- swyx 2y agoit makes sense in contrast to "tool use". basically, either fly-by-vision or fly-by-instruments, same dilemma you have in self driving cars
- ok_dad 2y agoIt’s simple and easy to understand what it is, that’s good marketing to my ears.
- accrual 2y agoI'm not sure what a better term is. It's kind of understated to me. An AI that can "use a computer" is a simple straightforward sentence but with wild implications.
- deleted 2y ago[deleted]
- netcraft 2y agosince they didnt rev the version, does this mean if we were using 3.5 today its just automatically using the new version? That doesnt seem great from a change management perspective though I am looking forward to using the new one in cursor.ai
- minimaxir 2y agoNo, Claude's models use date-pinning. The new model endpoint is claude-3-5-sonnet-20241022 https://docs.anthropic.com/en/docs/about-claude/models https://docs.anthropic.com/en/docs/about-claude/models
- deleted 2y ago[deleted]
- marsh_mellow 2y agoAnthropic blog post outlining the research process: https://www.anthropic.com/news/developing-computer-use https://www.anthropic.com/news/developing-computer-use Computer use API documentation: https://docs.anthropic.com/en/docs/build-with-claude/computer-use https://docs.anthropic.com/en/docs/build-with-claude/compute... Computer Use Demo: https://github.com/anthropics/anthropic-quickstarts/tree/main/computer-use-demo https://github.com/anthropics/anthropic-quickstarts/tree/mai...
- karpatic 2y agoThis needs to be brought up. Was looking for the demo and ended up on the contact form
- frankdenbow 2y agoThanks for these. Wonder how many people will use this at work to pretend that they are doing work while they listen to a podcast.
- nwnwhwje 2y agoThis is cover for the people whose screens are recorded. Run this on the monitorred laptop to make you look busy then do the actual work on laptop 2, some of which might actually require thinking so no mouse movements.
- distalx 2y agoOn their "Developing a computer use model" post they have mention > On one evaluation created to test developers’ attempts to have models use computers, OSWorld, Claude currently gets 14.9%. That’s nowhere near human-level skill (which is generally 70-75%), but it’s far higher than the 7.7% obtained by the next-best AI model in the same category. Here, "next-best AI model in the same category" referes to which model.
- bhouston 2y agoIs there an easy way to use Claude as a Co-Pilot in VS Code? If it is better at coding, it would be great to have it integrated.
- mkummer 2y agoContinue.dev's VS Code extension is fantastic for this
- BudaDude 2y agoCursor uses Claude as its base model. There may be extensions for VScode to do it but it will never be allowed in Copilot unless MS and OpenAI have a falling out.
- neb_b 2y agoYou can use it in Cursor - called "Cursor Tab" IMO Cursor Tab performs much better than Co-Pilot, easily works through things that would cause Co-Pilot to get stuck, you should give it a try
- TiredOfLife 2y agoAs I understand Cursor tab autocomplete uses their own model. Only chat has Sonnet and co.
- neb_b 2y agoAh, i thought it used the model selected for your prompts, either way, it seems to work very well
- teddarific 2y agoI originally thought that too but learned yesterday they have their own model. Definitely explains how its so fast and accurate!
- codingwagie 2y agoits funny that cursor.sh with < 30 developers has a better autocomplete model than microsoft
- postalcoder 2y agoand i was just planning to go to sleep…
- accrual 2y agoI discovered Mindcraft recently and stayed up a few hours too late trying to convince my local model to play Minecraft. Seems like every time a new capability becomes available, I can't wait to experiment with it for hours, even at the cost of sleep.
- Alifatisk 2y ago> Claude 3.5 Haiku matches the performance of Claude 3 Opus Oh wow!
- cynicalpeace 2y agoThis bolsters my opinion that OpenAI is falling rapidly behind. Presumably due to Sam's political machinations rather than hard-driving technical vision, at least that's what it seems like, outside looking in. Computer use seems it might be good for e2e tests.
- cube2222 2y agoThis looks quite fantastic! Nice improvements in scores across the board, e.g. > On coding, it [the new Sonnet 3.5] improves performance on SWE-bench Verified from 33.4% to 49.0%, scoring higher than all publicly available models—including reasoning models like OpenAI o1-preview and specialized systems designed for agentic coding. I've been using Sonnet 3.5 for most of my AI-assisted coding and I'm already very happy (using it with the Zed editor, I love the "raw" UX of its AI assistant), so any improvements, especially seemingly large ones like this are very welcome! I'm still extremely curious about how Sonnet 3.5 itself, and its new iteration are built and differ from the original Sonnet. I wonder if it's in any way based on their previous work[0] which they used to make golden-gate Claude. [0]: https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html https://transformer-circuits.pub/2024/scaling-monosemanticit...
- machiaweliczny 2y agoI'm waiting for Aider benchmark
- voiper1 2y agoIt's out, and improved! It went from 77.4% to 84.2%, skipping past O1-preview which is at 79.7% Source: https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/
- Tepix 2y agoInteresting stuff, i look forward to future developments. A comment about the video: Sam Runger talks wayyy too fast, in particular at the beginning.
- bluelightning2k 2y agoThis is what the Rabbit "large action model" pretended to be. Wouldn't be surprised to see them switch to this and claim they were never lying about their capabilities because it works now. Pretty cool for sure.
- swalsh 2y agoI think Rabbit had the business model wrong though, I don't think automating UI's to order pizza is anywhere near as valuable as automating the app workflows for B2B users.
- pradn 2y agoGreat progress from Anthropic! They really shouldn't change models from under the hood, however. A name should refer to a specific set of model weights, more or less. On the other hand, as long as its actually advancing the Pareto frontier of capability, re-using the same name means everyone gets an upgrade with no switching costs. Though, all said, Claude still seems to be somewhat of an insider secret. "ChatGPT" has something like 20x the Google traffic of "Claude" or "Anthropic". https://trends.google.com/trends/explore?date=now%201-d&geo=US&q=chatgpt,claude,anthropic&hl=en https://trends.google.com/trends/explore?date=now%201-d&geo=...
- cube2222 2y agoThere was a recent article[0] trending on HN a about their revenue numbers, split by B2C vs B2B. Based on it, it seems like Anthropic is 60% of OpenAI API-revenue wise, but just 4% B2C-revenue wise. Though I expect this is partly because the Claude web UI makes 3.5 available for free, and there's not that much reason to upgrade if you're not using it frequently. [0]: https://www.tanayj.com/p/openai-and-anthropic-revenue-breakdown https://www.tanayj.com/p/openai-and-anthropic-revenue-breakd...
- famouswaffles 2y ago3.5 is rate limited free, same as 4o (4o's limits are actually more generous). I think the real reason is much simpler - Claude/Anthropic has basically no awareness in the general public compared to Open AI. The chatGPT site had over 3B visits last month (#11 in Worldwide Traffic). Gemini and Character AI get a few hundred million but Claude doesn't even register in comparison. [0] Last they reported, OpenAI said they had 200M weekly active users.[1] Anthropic doesn't have anything approaching that. [0] https://www.similarweb.com/blog/insights/ai-news/chatgpt-topped-3-billion-visits-in-september/ https://www.similarweb.com/blog/insights/ai-news/chatgpt-top... [1] https://www.reuters.com/technology/artificial-intelligence/openai-says-chatgpts-weekly-users-have-grown-200-million-2024-08-29/ https://www.reuters.com/technology/artificial-intelligence/o...
- Eisenstein 2y ago
- Centigonal 2y agoThey should just adopt Apple "version numbers:" Claude Sonnet (Late 2024).
- Hizonner 2y agoCan this solve CAPTCHAs for me? It's starting to get to the point where limited biological brains can't do them.
- deleted 2y ago[deleted]
- m3kw9 2y agoI suspect they are gonna need some local offload capabilities for Computer Use, the repeated screen reading can definitely be done locally on modern machines, otherwise the cost maybe impractical.
- accrual 2y agoMaybe we need some agent running on the PC to offload some of these tasks. It could scrape the display at 30 or 60 Hz and produce a textual version of what's going on for the model to consume.
- abrichr 2y agoSee https://github.com/OpenAdaptAI/OpenAdapt https://github.com/OpenAdaptAI/OpenAdapt for an open source alternative that runs segmentation locally.
- veggieWHITES 2y agoWhile I was initially impressed with it's context window, I got so sick of fighting with Claude about what it was allowed to answer I quit my subscription after 3 months. Their whole policing AI models stance is commendable but ultimately renders their tools useless. It actually started arguing with me about whether it was allowed to help implement a github repository's code as it might be copywritten... it was MIT licensed open source from Google :/
- r2_pilot 2y agoI just include text that I own the device in question and that I have a legal team watching my every move. It's stupid, I agree, but not insurmountable. I had less refusals with Claude 3 Opus.
- msoad 2y agoI skimmed through the computer use code. It's possible to build this with other AI providers too. For instance you can asks ChatGPT API to call functions for click and scroll and type with specific parameters and execute them using OS's APIs (A11y APIs usually) Did I miss something? Did they have to make changes to the model for this?
- accrual 2y ago> execute them using OS's APIs (A11y APIs usually) I wonder if we'll end up with a new set of AI APIs in Windows, macOS, and Linux in the future. Maybe an easier way for them to iterate through windows and the UI elements available in each.
- jlpom 2y agoIt already exists for KDE: https://community.kde.org/Selenium https://community.kde.org/Selenium
- myprotegeai 2y agoWe are approaching FSD for the computer, with all of the lofty promises, and all of the horrible accidents.
- ford 2y agoSeems like both: - AI Labs will eat some of the wrappers on top of their APIs - even complex ones like this. There are whole startups that are trying to build computer use. - AI is fitting _some_ scaling law - the best models are getting better and the "previously-state-of-the-art" models are fractions of what they cost a couple years ago. Though it remains to be seen if it's like Moore's Law or if incremental improvements get harder and harder to make.
- skybrian 2y agoIt seems a little silly to pretend there’s a scaling “law” without plotting any points or doing a projection. Without the mathiness, we could instead say that new models keep getting better and we don’t know how long that trend will continue.
- ctoth 2y ago> It seems a little silly to pretend there’s a scaling “law” without plotting any points or doing a projection. Isn't this Kaplan 2020 or Hoffmann 2022?
- skybrian 2y agoYes, those are scaling laws, but when we see vendors improving their models without increasing model size or training longer, they don't apply. There are apparently other ways to improve performance and we don't know the laws for those. (Sometimes people track the learning curve for an industry in other ways, though.)
- ford 2y ago"Law" might not be the right word - but there's no denying it's scaling with compute/data/model size. I suppose law happens after continued evidence over years.
- highwaylights 2y agoCompletely irrelevant, and it might just be me, but I really like Anthropic's understated branding. OpenAI's branding isn't exactly screaming in your face either, but for something that's generated as much public fear/scaremongering/outrage as LLMs have over the last couple of years, Anthropic's presentation has a much "cosier" veneer to my eyes. This isn't the Skynet Terminator wipe-us-all-out AI, it's the adorable grandpa with a bag of werthers wipe-us-all-out AI, and that means it's going to be OK.
- minimaxir 2y agoAnthropic has recently begun a new, big ad campaign (ads in Times Square) that more-or-less takes potshots at OpenAI. https://www.reddit.com/r/singularity/comments/1g9e0za/anthropic_is_getting_rather_desperate_and_quite/ https://www.reddit.com/r/singularity/comments/1g9e0za/anthro...
- deleted 2y ago[deleted]
- whywhywhywhy 2y agoWonder what a normal person thinks this is an ad for
- joelanman 2y ago'transparent' in what sense?
- jprete 2y agoTop comment at the time I looked: "There seems to be a ton of confusion about the purpose of these ads. These are recruitment ads, not product ads, hence why "no drama" is the driving message. I'm sure these were all taken at or around a tech conference."
- minimaxir 2y agoThat comment is wrong, it appears this campaign is much wider. SF: https://x.com/_claudiazhao/status/1815463380767121733/photo/1 https://x.com/_claudiazhao/status/1815463380767121733/photo/... LA: https://x.com/michaelmiraflor/status/1840797631095964110/photo/1 https://x.com/michaelmiraflor/status/1840797631095964110/pho... Boston: https://x.com/moloneymike/status/1842203082374946851/photo/1 https://x.com/moloneymike/status/1842203082374946851/photo/1 London: https://x.com/maria_axente/status/1805607576156979673/photo/1 https://x.com/maria_axente/status/1805607576156979673/photo/...
- mmooss 2y agoOf course there's great inefficiency in having the Claude software control a computer with a human GUI mediating everything, but it's necessary for many uses right now given how much we do where only human interfaces are easily accessible. If something like it takes off, I expect interfaces for AI software would be published, standardized, etc. Your customers may not buy software that lacks it. But what I really want to see is a CLI. Watching their software crank out Bash, vim, Emacs!, etc. - that would be fascinating!
- modeless 2y agoI hope specialized interfaces for AI never happen. I want AI to use human interfaces, because I want to be empowered to use the same interfaces as AI in the future. A future where only AI can do things because it uses an incomprehensible special interface and the human interface is broken or non-existent is a dystopia. I also want humanoid robots instead of specialized non-humanoid robots for the same reason.
- accrual 2y agoMaybe we'll end up with both, kind of like how we have scripting languages for ease of development, but we also can write assembly if we need bare metal access for speed.
- deleted 2y ago[deleted]
- torginus 2y agoImo, APIs and to a lesser extent cli tools are already specialized tools made for LLMs. I've been editing videos with ChatGPT4 + ffmpeg for a year now.
- accrual 2y agoI agree, I bet models could excel at CLI tasks since the feedback would be immediate and in a language they can readily consume. It's probably much easier for them to to handle "command requires 2 arguments and only 1 was provided" than to do image-to-text on an error modal and apply context to figure out what went wrong.
- lairv 2y agoOfftopic but youtube doesn't allow me to view the embedded video, with a "Sign in to confirm you’re not a bot" message. I need to open a dedicated youtube tab to watch it The barrier to scraping youtube has increased a lot recently, I can barely use yt-dlp anymore
- ALittleLight 2y agoThat's funny. I was recently scraping tens of thousands of YouTube videos with yt-dlp. I would encounter throttling of some kind where yt-dlp stopped working, but I'd just spin a new VPS up and the throttled VPS down when that happened. The throttling effort cost me ~1 hour of writing the logic to handle it. I say that's funny because my guess would be they want to block larger scale scraping efforts like mine, but completely failed, while they attempt at throttling puts captchas in front of legitimate users.
- wesleyyue 2y agoIf anyone would like to try the new Sonnet in VSCode. I just updated https://double.bot https://double.bot to the new Sonnet. (disclaimer: I am the cofounder/creator) --- Some thoughts: * Will be interesting to see what we can build in terms of automatic development loops with the new computer use capabilities. * I wonder if they are not releasing Opus because it's not done or because they don't have enough inference compute to go around, and Sonnet is close enough to state of the art?
- swyx 2y agomy quick notes on Computer Use: - "computer use" is basically using Claude's vision + tool use capability in a loop. There's a reference impl but there's no "claude desktop" app that just comes with this OOTB - they're basically advertising that they bumped up Claude 3.5's screen vision capability. we discussed the importance of this general computer agent approach with David on our pod https://x.com/swyx/status/1771255525818397122 https://x.com/swyx/status/1771255525818397122 - @minimaxir points out questions on cost. Note that the vision use is very sparing - the loop is I/O constrained - it waits for the tool to run and then takes a screenshot, then loops. for a simple 10 loop task at max resolution, Haiku costs <1 cent, Sonnet 8 cents, Opus 41 cents. - beating o1-preview on SWEbench Verified without extended reasoning and at 4x cheaper output per token (a lot cheaper in total tokens since no reasoning tokens) is ABSOLUTE mogging - New 3.5 Haiku is 68% cheaper than Claude Instant haha references i had to dig a bit to find - https://www.anthropic.com/pricing#anthropic-api https://www.anthropic.com/pricing#anthropic-api - https://docs.anthropic.com/en/docs/build-with-claude/vision#evaluate-image-size https://docs.anthropic.com/en/docs/build-with-claude/vision#... - loop code https://github.com/anthropics/anthropic-quickstarts/blob/main/computer-use-demo/computer_use_demo/loop.py https://github.com/anthropics/anthropic-quickstarts/blob/mai... - some other screenshots https://x.com/swyx/status/1848751964588585319 https://x.com/swyx/status/1848751964588585319 - https://x.com/alexalbert__/status/1848743106063306826 https://x.com/alexalbert__/status/1848743106063306826 - model card https://assets.anthropic.com/m/1cd9d098ac3e6467/original/Claude-3-Model-Card-October-Addendum.pdf https://assets.anthropic.com/m/1cd9d098ac3e6467/original/Cla...
- akshayKMR 2y agoHaven't used vision models before, can someone comment if they are good at "pointing things". E.g given a picture, give co-ordinate for text "foo". This is the key to accurate control, it needs to be very precise. Maybe Claude's model is trained at this. Also what about open source vision models? Any ones good at "pointing things" on a typical computer screen?
- swyx 2y agoi mean like with everything they'll kinda be able to do it and only get really good at it if the model trainers prioritized it. see Pixmo for a recent example
- bergutman 2y agoThey need to get the price of 3.5 Haiku down. It's about 2x 4o-mini.
- quotemstr 2y agoStill super cheap
- caeril 2y agoPrecisely this. Aider (with the older Claude models) is already a semi-competent junior developer, and it will produce 1kloc of decent code for the equivalent of 50 cents in API costs. Sure, you still have to review the commits, but you have to do that anyway with human junior developers. Anthropic could charge 20x more and we would still be happy to pay it.
- torginus 2y agoClaude's current ability to use computers is imperfect. Some actions that people perform effortlessly—scrolling, dragging, zooming—currently present challenges for Claude and we encourage developers to begin exploration with low-risk tasks. Nice, but I wonder why didn't they use UI automation/accessibility libraries, that have access to the semantic structure of apps/web pages, as well as accessing documents directly instead of having Excel display them for you.
- accrual 2y agoI wonder if the model has difficulties for the same reason some people do - UI affordability has gone down with the flattening, hover-to-see scrollbar, hamburger-menu-ization of UIs. I'd like to see a model trained on a Windows 95/NT style UI - would it have an easier time with each UI element having clearly defined edges, clearly defined click and dragability, unified design language, etc.?
- torginus 2y agoWhat the UI looks like has no effect on for example, Windows UI Automation libraries. How the tech works is that it queries the process directly for the sematic description of items, like here's a button called 'Delete', here's a list of items for TODO's, and you get the tree structure directly from the API. I wouldn't be surprised if they are working off of screenshots, they still trained their models on having said screenshots annotated by said automation libraries, which told the AI what pixel is what.
- cherioo 2y agoI think this is to make human /user experience better. If you use accessibility features, then user need to know how to use those features. Similar to another comment in here, the UX they shoot for is “click the red button with cancel on it”, and ship that ASAP.
- abrichr 2y agoWe use operating system accessibility APIs when available in https://github.com/OpenAdaptAI/OpenAdapt https://github.com/OpenAdaptAI/OpenAdapt.
- deleted 2y ago[deleted]
- ramesh31 2y agoClaude is absurdly better at coding tasks than OpenAI. Like it's not even close. Particularly when it comes to hallucinations. Prompt for prompt, I see Claude being rock solid and returning fully executable code, with all the correct imports, while OpenAI struggles to even complete the task and will make up nonexistent libraries/APIs out of whole cloth.
- codingwagie 2y agoYeah, sonnet is noticeably better. To the point that openai is almost unusable, too many small errors
- rubslopes 2y agoI've been using a lot of o1-mini and having a good experience with it. Yesterday I decided to try sonnet 3.5. I asked for a simple but efficient script to perform fuzzy match in strings with Python. Strangely, it didn't even mention existing fast libraries, like FuzzyWuzzy and Rapidfuzz. It went on to create everything from scratch using standard libraries. I don't know, I thought this was something basic for it to stumble on.
- egillie 2y agoDoes anyone know _why_ it’s so much better at coding? Better architecture, better training data, better RLHF?
- myprotegeai 2y agoHow long until "computer use" is tricked into entering PII or PHI into an attackers website?
- accrual 2y agoI imagine initial computer use models will be kind of like untrained or unskilled computer users today (for example, some kids and grandparents). They'll do their best but will inevitably be easy to trick into clicking unscrupulous links and UI elements. Will an AI model be able to correctly choose between a giant green "DOWNLOAD NOW!" advertisement/virus button and a smaller link to the actual desired file?
- myprotegeai 2y agoExactly. Personalized ads are now prompt injection vectors.
- abraxas 2y agoHopefully the coding improvements are meaningful because I find that as a coding assistant o1-preview beats it (at least the Claude 3.5 that was available yesterday) but I like Claude's demeanor more (I know this sounds crazy but it matters a bit to me)
- 015a 2y agoWhy on god's green earth is it not just called Claude 3.6 Sonnet. Or Claude 4 Sonnet. I don't actually care what the answer is. There's no answer that will make it make sense to me.
- accrual 2y agoThe best answer I've seen so far is that "Claude 3.5 Sonnet" is a brand name rather than a specific version. Not saying I agree, just a way to visualize how the team is coming up with marketing.
- 015a 2y agoIt was certainly named by some nerd: "(pushes glasses up) well, we only updated the quantized diffraction sorter, the 3.5 version number refers to iterations on both that and the field matrix array interpreter, so technically because the interpreter hasn't changed so we shouldn't upgrade the version". This engineer has never seen a dollar from a customer in their life.
- campers 2y agoA bit like the new Gemini Pro 1.5-002 release.
- g9yuayon 2y agoIs it just me who feels that Anthropic has been innovating faster than ChatGPT in the past year?
- baq 2y agoScary stuff. 'Hey Claude 3.5 New, pretend I'm a CEO of a big company and need to lay off 20% people, make me a spreadsheet and send it to HR. Oh make sure to not fire the HR department' c.f. IBM 1979.
- TacticalCoder 2y agoOne suggestion, use the following prompt at a LLM: The combination of the words "computer use" is highly confusing. It's also "Yoda speak". For example it's hard for humans to parse the sentences *"Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku"*, *"Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku "* (it literally relies on the comma to make any sense) and *"Computer use for automated interaction"* (in the youtube vid's title: this one is just broken english). Please suggest terms that are not confusing for a new ability allowing an AI to control a computer as if it was a human.
- GavinGruesome 2y ago[dead]
- 29decibel 2y agoI am surprised it uses macOS as the demo, as I thought it would be harder to control vs Ubuntu. But maybe at the same time, macOS is the most predictable/reliable desktop environment? I noticed that they use virtual environment for the demo, curious how do they build that along with docker, is that leveraging the latest virtualization framework from Apple?
- vok 2y agoThis "Computer use" demo: https://www.youtube.com/watch?v=jqx18KgIzAE https://www.youtube.com/watch?v=jqx18KgIzAE shows Sonnet 3.5 using the Google web UI in an automated fashion. Do Google's terms really permit this? Will Google permit this when it is happening at scale?
- accrual 2y agoI wonder how they could combat it if they choose to disallow AI access through human interfaces. Maybe more captchas, anti-AI design language, or even more tracking of the user's movements?
- gzer0 2y agoOne of the funnier things during training with the new API (which can control your computer) was this: "Even while recording these demos, we encountered some amusing moments. In one, Claude accidentally stopped a long-running screen recording, causing all footage to be lost. Later, Claude took a break from our coding demo and began to peruse photos of Yellowstone National Park." [0] https://x.com/AnthropicAI/status/1848742761278611504 https://x.com/AnthropicAI/status/1848742761278611504
- throwup238 2y agoAt least now we know SkyClaude’s plan to end human civilization. It’s planning on triggering a Yellowstone caldera super eruption.
- mnk47 2y agoAm I misremembering or is this an exact plot point of Pluto (the manga/anime)?
- ctoth 2y agoNext release patch notes: * Fixed bug where Claude got bored during compile times and started editing Wikipedia articles to claim that birds aren't real * Blocked news.ycombinator.com in the Docker image's hosts file to avoid spurious flamewar posts (Note: The site is still recovering from the last insident) * Addressed issue of Claude procrastinating on debugging by creating elaborate ASCII art in Vim * Patched tendency to rickroll users when asked to demonstrate web scraping"
- TiredOfLife 2y agoYou forgot the most important one. * Added guards to prevent every other sentence being "I use neovim"
- rounakdatta 2y agoThank god it'll say "I use Claude btw", not leading to unnecessary text wars (and thereby loss of your valuable token credits).
- mclau156 2y agoDid they just invent a new world of warcraft or runescape bot?
- hubraumhugo 2y agoI've seen quite a few YC startups working on AI-powered RPA, and now it looks like a foundational model player is directly competing in their space. It will be interesting to see whether Anthropic will double down on this or leave it to third-party developers to build commercial applications around it.
- suchintan 2y agoWe're one of those players (https://github.com/Skyvern-AI/skyvern https://github.com/Skyvern-AI/skyvern) and we're definitely watching the space with a lot of excitement We thought it was inevitable that OpenAI / Anthropic would veer into this space and start to become competitive with us. We actually expected OpenAI to do it first! What this confirms is that there is significant interest in computer / browser automation, and the problem is still unsolved. We will see whether the automation itself is an application later problem (our approach) or whether the model needs to be intertwined with the application (Anthropic's approach here)
- gerash 2y agoThe "computer use" demos are interesting. It's a problem we used to work on and perhaps many other people have always wanted to accomplish since 10 years ago. So it's yet to be seen how well it works outside a demo. What was surprising was the slow/human speed of operations. It types into the text boxes at a human speed rather than just dumping the text there. Is it so the human can better monitor what's happening or is it so it does not trigger Captchas ?
- dtquad 2y agoNow I am really curious how to programmatically create a sandboxed compute environment to do a self-hosted "Computer use" and see how well other models, including self-hosted Ollama models, can do this.
- esseti 2y agoI checked the docs but did not find it out. Cloude has API as the GPT Assistant? with also the ability to give a set of documents to work with? It seems that you can only send single message, thus not relying on the ability to "learn" from predefined documents.
- devinprater 2y agoMaybe LLM's helping blind people like me play video games that aren't accessible to us normally, is getting closer!
- accrual 2y agoDefinitely! Those with movement disabilities could have a much easier time if they could just dictate actions to the computer and have them completed with some reliability.
- KoolKat23 2y agoGoogle has the tech (some of it's gathering dust, but they have it). They can use the gameplay tech developed for stadia when a user experiences lag and combine it with their LLM.
- iknownthing 2y agoCan Claude create and run a CI/CD pipeline now from a prompt?
- turnsout 2y agoWow, there's a whole industry devoted to what they're calling "Computer Use" (Robotic Process Automation, or RPA). I wonder how those folks are viewing this.
- maestrae 2y agoanybody know how the hell they're combating / gonna combat captcha's, cloudflare blocking, etc. I remember playing in this space on a toy project and being utterly frustrated by anti-scraping. Maybe one good thing that will come out of this AI boom is that companies will become nicer to scrapers? Or maybe, they'll just cut sweetheart deals?
- TaylorAlexander 2y agoAnd today I realized that despite it being an extremely common activity, we don’t really have a word for “using the computer” which is distinct from “computing”. It’s funny because AI models are always “using a computer” but now they can “use your computer.”
- binarymax 2y agoComputering
- deleted 2y ago[deleted]
- meindnoch 2y agoIn English at least. In other languages there are.
- rifty 2y agoThe word is interfacing generally (or programming for some) but it's just not commonly used for general users. I’d say this is probably because the activity of focus for general users is in use of the applications, not the computer itself despite being instanced with a computer. Thus a computer is commonly less the user’s object of activity, and more commonly the setting for activity. Similarly using our homes are an extremely common ‘activity’, yet the object-activities that get special words commonly used are the ones with specific user application.
- bongodongobob 2y agoOperating a computer?
- TaylorAlexander 2y agoRight. We don’t have a word for that. Like “using a bicycle” has the word “bicycling”. Tho someone here suggested “computering” which is pretty good.
- freediver 2y agoBoth new Sonnet and gpt-4o still fail at a simple: "How many w's are in strawberry?" gpt-4o: There are 2 "w's" in "strawberry." Claude 3.5 Sonnet (new): Let me count the w's in "strawberry": 0 w's. (same question with 'r' succeeds) What is artificial about current gen of "artificial intelligence" is the way training (predict next token) and benchmarking (overfitting) is done. Perhaps a fresh approach is needed to achieve a true next step.
- fassssst 2y agoThey are trained on tokens not characters.
- redox99 2y agoThere's always that one tokenization error comment
- ssijak 2y agoCan we stop with these useless strawberry examples?
- wild_egg 2y agoIt's bad at directly working on classical computer problems like math and data processing. But you can do it indirectly by having it write a program that produces the correct result. Interestingly, I didn't even have to have it run the program, although usually you would write a tool which counts the number of w's in "strawberry" and return the result Which produced: Here's a simple Python function that counts the number of 'w's in the word "strawberry" and returns the result: ```python def count_w_in_strawberry(): word = "strawberry" count = word.count('w') return count # Call the function and print the result result = count_w_in_strawberry() print(f"The number of 'w's in 'strawberry' is: {result}") ``` This tool does the following: 1. We define a function called `count_w_in_strawberry()`. 2. Inside the function, we assign the word "strawberry" to a variable called `word`. 3. We use the `count()` method on the `word` string to count the occurrences of 'w'. 4. The function returns the count. 5. Outside the function, we call `count_w_in_strawberry()` and store the result in the `result` variable. 6. Finally, we print the result. When you run this code, it will output: ``` The number of 'w's in 'strawberry' is: 1 ``` This tool correctly identifies that there is one 'w' in the word "strawberry".
- punnerud 2y agoCursor AI already have the option to switch to using claude-3-5-sonnet-20241022 in the chat box. I was about to try to add a custom API. I’m impressed by the speed of that team.
- jampekka 2y agoIt's quite sad that application interoperability requires parsing bitmaps instead of exchanging structured information. Feels like a devastating failure in how we do computing.
- SuaveSteve 2y agoThe people have chosen apps over protocols.
- jampekka 2y agoWorse is better.
- chillee 2y agoIt's very much in the "worse is better" camp.
- abrichr 2y agoSee https://github.com/OpenAdaptAI/OpenAdapt https://github.com/OpenAdaptAI/OpenAdapt for an open source alternative that includes operating system accessibility API data and DOM information (along with bitmaps) where available. We are also planning on extracting runtime information using COM/AppleScript: https://github.com/OpenAdaptAI/OpenAdapt/issues/873 https://github.com/OpenAdaptAI/OpenAdapt/issues/873
- accrual 2y agoIt's super cool to see something like this already exists! I wonder if one day something adjacent will become a standard part of major desktop OSs, like a dedicated "AI API" to allow models to connect to the OS, browse the windows and available actions, issue commands, etc. and remove the bitmap parsing altogether as this appears to do.
- janalsncm 2y agoApps are built for people rather than computers.
- 2y ago
- cwkoss 2y agoClaude is amazing. The project documents functionality makes it a clear leader ahead of ChatGPT and I have found it to be the clear leader in coding assistance over the past few months. Web automation is really exciting. I look forward to the brave new future where I can code a webapp without ever touching the code, just testing, giving feedback, and explaining discovered bugs to it and it can push code and tweak infrastructure to accomplish complex software engineering tasks all on its own. Its going to be really wild when Claude (or other AI) can make a list of possible bugs and UX changes and just ask the user for approval to greenlight the change.
- LVB 2y agoNot specific to this update, but I wanted to chime in with just how useful Claude has been, and relatively better than ChatGPT and GitHub copilot for daily use. I've been pro for maybe 6 months. I'm not a power user leveraging their API or anything. Just the chat interface, though with ever more use of Projects, lately. I use it every day, whether for mundane answers or curiosities, to "write me this code", to general consultation on a topic. It has replaced search in a superior way and I feel hugely productive with it. I do still occasionally pop over to ChatGPT to test their their waters (or if Claude is just not getting it), but I've not felt any need to switch back or have both. Well done, Anthropic!
- zone411 2y agoIt improves to 25.9 over the previous version of Claude 3.5 Sonnet (24.4) on NYT Connections: https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/.
- jjice 2y agoWhat a neat bench mark! I'm blown away that o1 absolutely crushes everyone else in this. I guess the chain of thought really hashes out those associations.
- rkharsan64 2y agoIsn't it possible that o1 was also trained on this data (or something super similar) directly? The score seems disproportionately high.
- usaar333 2y agoThey definitely considered it. Early theinformation articles talked about how high the performance of strawberry was on it.
- amarcheschi 2y agoPerhaps it's just because English is not my native language, but the prompt 3 isn't quite clear at the beginning when it says "group of four. Words (...)". It is not explained what the group of four must be, if I add to the prompt "group of four words" Claude 3.5 manages to answer it, while without it, Claude tells it is not that clear and can't answer
- TechDebtDevin 2y agoNot that I'm scared of this update but I'd probably be alright with pausing llm development today, atleast in regard to producing code. I don't want an llm to write all my code, regardless of if it works, I like to write code. What these models are capable of at the moment is perfect for my needs and I'd be 100% okay if they didn't improve at all going forward. Edit: also I don't see how an llm controlled system can ever replace a deterministic system for critical applications.
- accrual 2y agoI have trouble with this too. I'm working on a small side project and while I love ironing out implementation details myself, it's tough to ignore the fact that Claude/GPT4o can create entire working files for me on demand. It's still enjoyable working at a higher architecture level and discussing the implementation before actually generating any code though.
- TechDebtDevin 2y agoI don't mind using it to make inline edits or more global edits between files at my descresion, and according to my instructions. Definitely saves tons of time and allows me to be more creative, but I don't want it make decisions on its own anymore than it already does. I tried using the composer feature on Cursor.sh, that's exactly the type of llm tool I do not want.
- machiaweliczny 2y agoIn normal critical system u use 3 CPUs. With LLM u can 1000 shot majority voting. Seems like approaches like entropix might reduce hallucinations also.
- TechDebtDevin 2y agoI don't think this snapshot image/vision model is going to be the best solution. I think CLI is a much better interface for llms.
- vivekkairi 2y agoaider benchmarks for claude 3.5 new are impressive. From 77.4% to 83.5% beating o1-preview.
- janalsncm 2y agoReminds me of the rise in job application bots. People are applying to thousands of jobs using automated tools. It’s probably one of the inevitable use cases of this technology. It makes me think. Perhaps the act of applying to jobs will go extinct. Maybe the endgame is that as soon as you join a website like Monster or LinkedIn, you immediately “apply” to every open position, and are simply ranked against every other candidate.
- quantadev 2y agoThe `Hiring Process` in America is definitely BADLY broken. Maybe worldwide afaik. It's a far too difficult, time-consuming, and painful process for everyone involved. I have a feeling AI can fix this, although I'd never allow an AI bot to interview me. I just mean other ways of using AI to help the process. Also people are hired for all kinds of reasons having little to do with their qualifications lots of the time, and often due to demographics (race, color, age, etc), and this is another way maybe AI can help by hiding those aspects of a candidate somehow.
- javajosh 2y agoAI and new tools have broken the system. The tools send you email saying things like "X corp is interested in you!" and you send a resume, and you don't hear back. Nothing, not even a rejection. Eventually you stop believing them, understanding it for the marketing spam that it is. Direct submissions are better, but only slightly. Recruiters are much better, in general, since they have a relationship with a real person at the company and can actually get your resume in front of eyes. But yeah, tools like ziprecruiter, careerboutique, jobot, etc are worse than useless: by lying to you about interest they actively discourage you from looking. There are no good alternatives (I'd love to learn I'm wrong), so you have to keep using those bad tools anyway.
- quantadev 2y agoAll that's true, and sadly it also often doesn't even matter how good you even are either. I have decades of experience and I still get "evaluated" based on how fast I can do silly brain-teaser IQ-test coding challenges. I've gotten where any company that wants me to do a coding challenge on my own time is an immediate "no thanks" reply from me. Everyone should refuse that. But so many people are so desperate they allow hiring companies to abuse them in that way. I consider it a kind of abuse of power to demand people do like 4 to 6hrs of nonsensical coding just to buy an opportunity for an actual interview.
- submeta 2y agoThat’s too much control for my taste. I don’t want anthropic to see my screen. I rather prefer a VS Code with integrated Claude. A version that can see all my dev files in a given folder. I don’t need it to run Chrome for me.
- trzy 2y agoPretty cool! I use Claude 3.5 to control a robot (ARKit/iOS based) and it does surprisingly well in the real world: https://youtu.be/-iW3Vzzr3oU?si=yzu2SawugXMGKlW9 https://youtu.be/-iW3Vzzr3oU?si=yzu2SawugXMGKlW9
- mrmansano 2y agoThat looks pretty cool, congrats! How feasible is it to be a product by itself? Did you try with a local edge model?
- trzy 2y agoNone of the small LLMs are good enough yet. You could certainly build a system around local VLMs but it would require much more task specific programming baked in. I’m certainly interested in building a product (not entirely controlled by an LLM but I see lots of utility in building interfaces with them) but not really sure what this would be useful for. Looking into some spaces now but there has to be a clear ROI to get any sort of funding for robotics.
- abc-1 2y agoI tried to get it to translate a document and it stopped after a few paragraphs and asked if I wanted it to keep going. This is not appropriate for my use case and it kept doing this even though I explicitly told it not to. The old version did not do this.
- graeme 2y agoI noticed some timeouts today. Could be capacity limits from the announcement
- jerrygoyal 2y agodoes anyone know what are some use cases for "computer use"?
- gumboshoes 2y agoFor me, one of the more useful steps on macOS will be when local AI can manipulate anything that has an Apple Script library. The hooks are there and decently documented. For meta purposes, having AI work with a third-party app like Keyboard Maestro or Raycast will even further expand the pre-built possibilities without requiring the local AI to reinvent steps or tools at the time of each prompt.
- efields 2y agoCaptchas are toast.
- edm0nd 2y agothey have been toast for at least a decade if not two now. With OCR and captcha solving services like DeathByCaptcha or AntiCaptcha where it costs ~$2.99 per 1k successfully solved captchas, they are a non-issue amd takes about 5-10 lines of code added to your script to implement a solution.
- simonw 2y agoI wrote up some of my own notes on Computer Use here: https://simonwillison.net/2024/Oct/22/computer-use/ https://simonwillison.net/2024/Oct/22/computer-use/
- logankeenan 2y agoMolmo released recently and is able to provide point coordinates for objects in images. I’ve been testing it out recently and am currently building an automation tool that allows users to more easily control a computer. Looks like Anthropic built a better one. Edit: it seems like these new features will eliminate a lot of automated testing tools we have today. Code for molmo coordinate tests https://github.com/logankeenan/molmo-server https://github.com/logankeenan/molmo-server
- tammer 2y agoThis demo is impressive although my initial reaction is a sort of grief that I wasn't born in the timeline where Alan Kay's vision of object-oriented computing was fully realized -- then we wouldn't have to manually reconcile wildly heterogeneous data formats and interfaces in the first place!
- aryehof 2y agoVision of a “universal communicator” rather than OO?
- nopinsight 2y agoThis needs more discussion: Claude using Claude on a computer for coding https://youtu.be/vH2f7cjXjKI?si=Tw7rBPGsavzb-LNo https://youtu.be/vH2f7cjXjKI?si=Tw7rBPGsavzb-LNo (3 mins) True end-user programming and product manager programming are coming, probably pretty soon. Not the same thing, but Midjourney went from v.1 to v.6 in less than 2 years. If something similar happens, most jobs that could be done remotely will be automatable in a few years.
- dmartinez 2y agoEvery time I see this argument made, there seems to be a level of complexity and/or operational cost above which people throw up their hands and say "well of course we can't do that". I feel like we will see that again here as well. It really is similar to the self-driving problem.
- unshavedyak 2y agoI feel pain for the people who will be employed to "prompt engineer" the behavior of these things. When they inevitably hallucinate some insane behavior a human will have to take blame for why it's not working.. and yea, that'll be fun to be on the receiving end of.
- WalterSear 2y agoHumans 'hallucinate' like LLMs. The term used however, is confabulation: we all do it, we all do it quite frequently, and the process is well studied(1). > We are shockingly ignorant of the causes of our own behavior. The explanations that we provide are sometimes wholly fabricated, and certainly never complete. Yet, that is not how it feels. Instead it feels like we know exactly what we're doing and why. This is confabulation: Guessing at plausible explanations for our behavior, and then regarding those guesses as introspective certainties. Every year psychologists use dramatic examples to entertain their undergraduate audiences. Confabulation is funny, but there is a serious side, too. Understanding it can help us act better and think better in everyday life. I suspect it's an inherent aspect of human and LLM intelligences, and cannot be avoided. And yet, humans do ok, which is why I don't think it's the moat between LLM agents and AGI that it's generally assumed to be. I strongly suspect it's going to be yesterday's problem in 6-12 months at most. (1) https://www.edge.org/response-detail/11513 https://www.edge.org/response-detail/11513
- LASR 2y agoThis is actually a huge deal. As someone building AI SaaS products, I used to have the position that directly integrating with APIs is going to get us most of the way there in terms of complete AI automation. I wanted to take at stab at this problem and started researching some daily busineses and how they use software. My brother-in-law (who is a doctor) showed me the bespoke software they use in his practice. Running on Windows. Using MFC forms. My accountant showed me Cantax - a very powerful software package they use to prepare tax returns in Canada. Also on Windows. I started to realize that pretty much most of the real world runs on software that directly interfaces with people, without clearly defined public APIs you can integrate into. Being in the SaaS space makes you believe that everyone ought to have client-server backend APIs etc. Boy was I wrong. I am glad they did this, since it is a powerful connector to these types of real-world business use cases that are super-hairy, and hence very worthwhile in automating.
- skissane 2y agoYou don’t know for a fact that those two specific packages don’t have supported APIs. Just because the user doesn’t know of any API doesn’t mean none exists. The average accountant or doctor is never going to even ask the vendor “is there an API” because they wouldn’t know what to do with one if there was.
- astrange 2y agoIf they're accessible to screen readers they have one. Accessibility is API for apps in disguise. In this case I doubt they're networked apps so they probably don't have a server API.
- skissane 2y ago> In this case I doubt they're networked apps so they probably don't have a server API. I think it would be very unusual this decade for software used to run either a medical practice or tax accountants to not be networked. Most such practices have multiple doctors/accountants, each with their individual computer, and they want to be able to share files, so that if your doctor/accountant is away their colleague can attend to you. Managing backups/security/etc is all a lot easier when the data is stored in a central server (whether in the cloud or a closet) than on individual client machines. Just because it is a fat client MFC-based Windows app doesn’t mean the data has to be stored locally. DCOM has been a thing since 1996.
- Bjorkbat 2y agoTried my standard go-to for testing, asked it to generate a voronoi diagram using p5js. For the sake of job security I'm relieved to see it still can't do a relatively simple task with ample representation in the Google search results. Granted, p5js is kind of niche, but not terribly so. It's arguably the most popular library for creating coding. In case you're wondering, I tried o1-preview, and while it did work, I was also initially perplexed why the result looked pixelated. Turns out, that's because many of the p5js examples online use a relatively simple approach where they just see which cell-center each pixel is closest to, more or less. I mean, it works, but it's a pretty crude approach. Now, granted, you're probably not doing creative coding at your job, so this may not matter that much, but to me it was an example of pretty poor generalization capabilities. Curiously, Claude has no problem whatsoever generating a voronoi diagram as an SVG, but writing a script to generate said diagrams using a particular library eluded it. It knows how to do one thing but generalizes poorly when attempting to do something similar. Really hard to get a real sense of capabilities when you're faced with experiences like this, all the while somehow it's able to solve 46% of real-world python pull-requests from a certain dataset. In case you're wondering, one paper (https://cs.paperswithcode.com/paper/swe-bench-enhanced-coding-benchmark-for-llms https://cs.paperswithcode.com/paper/swe-bench-enhanced-codin...) found that 94% of the pull-requests on SWE-bench were created before the knowledge cutoff dates of the latest LLMs, so there's almost certainly a degree of data-leakage.
- nemothekid 2y agoIt's surprising how much knowledge is not easily googleable and can only unearched by deep diving into OSS or asking an expert. I recently was debugging a rather naive gstreamer issue where I was seeing a delay in the processing. ChatGPT, Claude and Google were all unhelpful. I spend the next couple days reading the source code, found my answer, and thought it was a bug. Asked the mailing list, and my problem was solved in 10 seconds by someone who could identify the exact parameter that was missing (and IMO, required some architecture knowledge on how gstreamer worked - and why the unrelatedly named parameter would fix it). The most difficult problems fall into this camp - I don't usually find myself reaching for LLMs when the problem is trivial unless it involves a mountain of boilerplate.
- 2y ago
- brcmthrowaway 2y agoThis is bad news for SWEs!
- brid 2y agoLooks like visual understanding of diagrams is improved significantly! For example, it was on par with Chat GPT 4o and Gemini 1.5 in parsing an ERD for a conceptual model, but now far excels over the others.
- bilsbie 2y agoDoes this make cursor obsolete? You can just use any IDE you want and it will work with it.
- jusgu 2y agoAssuming running this new computer interactivity feature is as fast as cursor composer (which I don’t think it is)—it still doesn’t support codebase indexing, inline edits or references to other variables and files in the codebase. I can see how someone could use this to make some sort of cursor competitor but out of the box there’s a very low likelihood it makes cursor obsolete.
- 93po 2y agoi really want cursor to integrate this so it can look at the results of a code change in the browser and then make edits as needed until it's accomplished what i asked of it. same for errors in the console etc. right now i have to manually describe the issue or copy and paste the error message and it'd be nice for it to just iterate more on its own
- sedatk 2y ago> developers can direct Claude to use computers the way people do—by looking at a screen, moving a cursor, clicking buttons, and typing text. So, this is how AI takes over the world.
- runako 2y agoI really don't get their model. They have very advanced models, but the service overall seems to be a jumble of priorities. Some examples: Anthropic doesn't offer an unlimited chatbot service, only plans that give you "more" usage, whatever that means. If you have an API key, you are "unlimited," so they have the capability. Why doesn't the chatbot allow one to use their API key in the Claude app to get unlimited usage? (Yes, I know there are third-party BYOK tools. That's not the question.) Claude appears to be smart enough to make an Excel spreadsheet with simple formulae. However, it is apparently prevented from making any kind of file. Why? What principle underlies that guardrail that does not also apply to Computer Use? Really want to make Claude my daily driver, but right now it often feels too much like a research project.
- stuckkeys 2y agoEven with API, depending what tier you are sitting on, there is daily limits. OpenAI used to be able to generate files for you, they changed that. It was useful.
- runako 2y agoInterestingly enough, after Claude refused to generate a file for me, I sent the same request to ChatGPT and got the Excel file I wanted. I wasn't aware of tiers in the Claude API, they are not mentioned on the API pricing page. Are the limits disclosed or just based on vibes like they are for the chatbot?
- deleted 2y ago[deleted]
- saaaaaam 2y agoWhat do you mean by “file” here? I’m making files on a daily basis, including CSVs, html, executable code, XML, JSON and other formats. It built me an entire visual wireframe for something the other day. Are you using artefacts? But I’m maybe misunderstanding your point because my use is relatively basic through the built in chatbot.
- 2y ago
- 2-3-7-43-1807 2y agowow, i almost got worried but the cute music and the funny little monster on the desk convinced me that this all just fun and dandy and all will be good. the future is coming and we'll all be much more happy :)
- astrange 2y agoI think this is good evidence that people's jobs are not being replaced by AI, because no AI would give the product a confusing name like "new Claude 3.5 Sonnet".
- abixb 2y agoI wonder why they didn't choose a "point update" scheme, like bumping it up to v3.6, for example. I agree, the naming is super confusing.
- cryptoegorophy 2y agoMaybe they should’ve asked Claude to generate a better name. Very dangerous to live in your own hyper focused bubble while trying to build a mass market product.
- jnwatson 2y agoGoogle, OpenAI, and Anthropic are responsibly scaling their models by confusing their customers into using the wrong ones. When AGI finally is launched, adoption will be responsibly slowed because it is called something like "new new Gemini Giga 12.9.2xo IT" and users will have to select it from dozens of similar names.
- lutusp 2y ago> "... and similar speed to the previous generation of Haiku." To me this is the most annoying grammatical error. I can't wait for AI to take over all prose writing so this egregious construction finally vanishes from public fora. There may be some downsides -- okay, many -- but at least I won't have to read endless repetitions of "similar speed to ..." when the correct form is obviously "speed similar to". In fact, in time this correct grammar may betray the presence of AI, since lowly biologicals (meaning us) appear not to either understand or fix this annoying error without computer help.
- smcleod 2y agoI wonder when it'll actually be available in the Bedrock AU region, because as of right now we're still stuck using mid-range models from a year ago. Amazon has really neglected ap-southeast-2 when it comes to LLMs.
- dheerkt 2y agoCan you not use cross-region inference?
- smcleod 2y ago90% of our customers do not allow this due to data sovereignty. Bedrock here is lagging so far behind several customers assume AWS simply aren't investing here anymore - or if they are it's an afterthought - and a very expensive one at that. I've spoken with several account managers and SAs and they seem similarly frustrated with the continual response from above that useful models are "coming soon". You can't even BYO models here, we usually end up spinning up big ol' GPU EC2 instances and serving our own, or for some tasks running locally as you can get better openweight LLMs.
- dheerkt 2y agoHmm interesting, didn't realize that data sovereignty requirements were so stringent. Wonder how other cloud providers are doing in this sense considering GPU shortages across the board.
- throwvc3 2y agoWhat I'd like to know is whether prompt caching is available to Claude on AWS Bedrock now.
- FloatArtifact 2y agoIt will interesting to see how this evolves. UI automation use case is different from accessibility do to latency requirement. latency matters a lot for accessibility not so much for ui automation testing apparatus. I've often wondered what the combination of grammar-based speech recognition and combination with LLM could do for accessibility. Low domain Natural Language Speech recognition augmented by grammar based speech recognition for high domain commands for efficiency/accuracy reducing voice strain/increasing recognition accuracy. https://github.com/dictation-toolbox/dragonfly https://github.com/dictation-toolbox/dragonfly
- anotherpaulg 2y agoThe new Sonnet tops aider's code editing leaderboard at 84.2%. Using aider's "architect" mode it sets the SOTA at 85.7% (with DeepSeek as the "editor" model). 84% Claude 3.5 Sonnet 10/22 80% o1-preview 77% Claude 3.5 Sonnet 06/20 72% DeepSeek V2.5 72% GPT-4o 08/06 71% o1-mini 68% Claude 3 Opus It also sets SOTA on aider's more demanding refactoring benchmark with a score of 92.1%! 92% Sonnet 10/22 75% o1-preview 72% Opus 64% Sonnet 06/20 49% GPT-4o 08/06 45% o1-mini https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/
- ianeigorndua 2y agoAre these synthetic or real-world benchmarks? Answering myself: ”Aider’s code editing benchmark asks the LLM to edit python source files to complete 133 small coding exercises from Exercism” Not gonna start looking for a job any time soon
- zeroonetwothree 2y agoExample I chose at random: > Convert a hexadecimal number, represented as a string (e.g. "10af8c"), to its decimal equivalent using first principles (i.e. no, you may not use built-in or external libraries to accomplish the conversion). So it's fairly synthetic. It's also the sort of thing LLMs should be great at since I'm sure there's tons of data on this sort of thing online.
- Vampiero 2y agoYeah but programming isn't about solving problems that were solved millions of times already. I mean, web dev kind of is, but that's not the point. If a problem is solved, then it's just a matter of implementing the solution and anyone can do that given the proper instructions (even without understanding how or why they solve the problem). I've formalized a lot of stuff I didn't understand just by copying the formulas from Wikipedia. As long as LLMs are not capable of proper reasoning, they will remain a gimmick in the context of programming. They should really just focus on refactoring benchmarks across many languages. If an AI can refactor my complex code properly without changing the semantics, it's good enough for me. But that unfortunately requires such a high-level understanding of the codebase that with the current tech it's just impossible to get a half-decent result in any real-world scenario.
- itissid 2y agoThis can power one of my favorite use-cases. Like find me a list of things to do with a family, given today's weather and in the next 2 hours, quiet sit down with lots of comfy seating, good vegetarian food... Not only is this kind of use getting around API restrictions, it is also a superior way to do search: Specify arbitrary preferences upfront instead of a search box and trawling different modalities of content to get better result. The possibilities for wellness use cases are endless, especially for end users that care about privacy and less screen use.
- mtgentry 2y agoWhat are the licensing implications of this? If I’m Google, I’d be pissed that my software is being used without a human there looking at the ads.
- KoolKat23 2y agoAnd today they added a new AI abuse clause to their t&C's lol.
- SturgeonsLaw 2y agoThey just need to start tailoring their ads to things that bots might be interested in
- urbandw311er 2y ago> we have provided three tools > bash shell November 2024: AI is allowed to execute commands in a bash shell. What could possibly go wrong?
- RecycledEle 2y agoHow long until it is profitable the tell a cheap AI to "win this game by collecting resources and advancing in-game" and then sell the account on eBay? I wonder what optimizations could be made? Could a gold farmer have the directions from one AI control many accounts? Could the AI program simpler bots for each bit of the game? I can imagine not being smart enough to play against computers, because I am flagged as a bot. I can imagine a message telling me I am banned because "nobody but a stupid bot would score so low."
- mercacona 2y agoI'm giving the new Sonnet a chance, although for my use as a writing companion so far, Opus has been king among all the models I've tried. However, I've been using Opus as a writing companion for several months, especially when you have writer's block and ask it for alternative phrases, it was super creative. But in recent weeks I was noticing a degradation in quality. My impression is that the model was degrading. Could this be technically possible? Might it be some kind of programmed obsolescence to hype new models?
- KoolKat23 2y agoYou're expectations could just be increasing as you start taking it for granted and are using other models.
- wewtyflakes 2y agoI wonder if OpenAI will fast follow; usually they're the ones to throw down the gauntlet. That being said, you can play around with OpenAI with a similar architecture of vision + agent + exec + loop using Donobu, though it is constrained to web browsers.
- lossolo 2y agoLivebench updated https://livebench.ai https://livebench.ai Model | Global | Reasoning | Coding | Math | Data | Language | IF ------------------------------|---------|-----------|---------|---------|---------|----------|------- o1-preview-2024-09-12 | 66.02 | 68.00 | 50.85 | 62.92 | 63.97 | 72.66 | 77.72 claude-3-5-sonnet-20241022 | 60.33 | 58.67 | 67.13 | 51.28 | 52.78 | 58.09 | 74.05 claude-3-5-sonnet-20240620 | 59.80 | 58.67 | 60.85 | 53.32 | 56.74 | 56.94 | 72.30
- kingkongjaffa 2y agoInterestingly new claude only knows content up to: > I'm limited to what I know as of April 2024, which includes the initial Claude 3 family launch but not subsequent updates.
- HarHarVeryFunny 2y agoThe "computer use" ability is extremely impressive! This is a lot more than an agent able to use your computer as a tool (and understanding how to do that) - it's basically an autonomous reasoning agent that you can give a goal to, and it will then use reasoning, as well as it's access to your computer, to achieve that goal. Take a look at their demo of using this for coding. https://www.youtube.com/watch?v=vH2f7cjXjKI https://www.youtube.com/watch?v=vH2f7cjXjKI This seems to be an OpenAI GPT-o1 killer - it may be using an agent to do reasoning (still not clear exactly what is under the hood) as opposed to GPT-o1 supposedly being a model (but still basically a loop around an LLM), but the reasoning it is able to achieve in pursuit of a real world goal is very impressive. It'd be mind boggling if we hadn't had the last few years to get used to this escalation of capabilities. It's also interesting to consider this from POV of Anthropic's focus on AI safety. On their web site they have a bunch of advice on how to stay safe by sandboxing, limiting what it has access to, etc, but at the end of the day this is a very capable AI able to use your computer and browser to do whatever it deems necessary to achieve a requested goal. How far are we from paperclip optimization, or at least autonomous AI hacking ?
- seany62 2y agoFrom what I'm seeing on GH, this could have technically already been built right? Is it not just taking screenshots of the computer screen and deciding what to do from their / looping until it gets to the solution ?
- HarHarVeryFunny 2y agoWell, obviously it's controlling your computer too - controlling mouse and keyboard input, and has been trained to know how to interact with apps (how to recognize and use UI components). It's not clear exactly what all the moving parts are and how they interact. I wouldn't be so dismissive - you could describe GPT-o1 in same way "it just loops until it gets to the solution". It's the details and implementation that matter.
- joshuamcginnis 2y agoIs there anything out there yet that will let me issue the command: > Refactor the api folder with any recommended readability improvements or improvements that would help DRY up code without adding additional complexity. Then I can just `git status` to see the changes?
- vipshek 2y agoInstall Cursor (https://cursor.com https://cursor.com), go into Cursor Settings and disable everything but Claude, then open Composer (Ctrl/Cmd + I). Paste in your exact command above. I bet it’ll do something pretty close to what you’re looking for.
- falcor84 2y agoAider is great at this stuff. The recommended way is to have it automatically commit, and then you can examine and possibly revert/reset its commits (or just have it work on a separate branch), but you can also use --no-auto-commits
- thecolorgreen 2y agoThis looks really similar to rabbit's Large Action Model (LAM). Cool! https://www.rabbit.tech/rabbit-os https://www.rabbit.tech/rabbit-os
- flockonus 2y agoAre these ppl are aware that they can bump minor versions? The mkt team vetoed Claude 3.6 ???
- sumedh 2y agoGet ready for Claude 3.5 Pro Max.
- throwaway0123_5 2y agoThis is incredibly cool but it seems like the potential damage from a "hallucination" in this mode is considerable, especially when they provide examples of it going very far off-track (looking up Yellowstone pictures). Would basically need constant monitoring for me not to be paranoid it did something stupid. Also seems like a privacy issue with them sending screenshots of your device back to their servers.
- lr1970 2y agoI am curious why "upgraded Claude 3.5 Sonnet" instead of simply Claude 3.6 Sonnet? Minor version increment is a standard way of versioning update. Am i missing something or it is just Anthropic marketing?
- loktarogar 2y agoProbably because there was no 3.1-3.4, and that the .5 is mostly just to represent that it's an upgrade on Claude 3 but not quite enough to be a Claude 4
- sumedh 2y agoChrome is at version 130, who really cares about the version number.
- taytus 2y agoComputer use won't allow you to log in to social media accounts, even if it is your account and credentials. Bummer.
- simonw 2y agoClaude 3.5 Opus is no longer mentioned at all on https://docs.anthropic.com/en/docs/about-claude/models https://docs.anthropic.com/en/docs/about-claude/models Internet Archive confirms that on the 8th of October that page listed 3.5 Opus as coming "Later this year" https://web.archive.org/web/20241008222204/https://docs.anthropic.com/en/docs/about-claude/models https://web.archive.org/web/20241008222204/https://docs.anth... The fact that it's no longer listed suggests that its release has at least been delayed for an unpredictable amount of time, or maybe even cancelled.
- nocturnes 2y agoIt's possible that they've determined that Opus no longer makes sense if they're able to focus on continuously optimising Sonnet. That said, Anthropic have been relatively good at setting and managing expectations, so today would have been a good time to make that clear.
- thenameless7741 2y agoBefore anyone reads too much into this, here's what an Anthropic staff said on Discord: > i don't write the docs, no clue > afaik opus plan same as its ever been
- a9dhalaan 2y agoMaybe he's not high level enough employee to have any say in the product roadmap, and he's behind on leadership planning?
- szundi 2y agoOr maybe they released it
- deleted 2y ago[deleted]
- nwnwhwje 2y agoAny comments on alignment with Anthropic's missions. Last time I checked Anthropic is about building SOTA as that is the only way to do safety research. Making money ans useful stuff commercially is a means to an end.
- attentive 2y agoThey need to work on their versioning. "3.5 Sonnet (New)", WTAF? - just call it 3.6 Sonnet or something. Is it "New" sonnet? is it "upgraded"? Is there a difference? How do I know which one I use? I can understand claude-3-5-sonnet-20241022, but that's not what users see.
- csomar 2y agoI like its lazy approach > I apologize, but it seems we're having some issues installing Chromium. As an alternative, since I can see Firefox is already installed on the system, would you like to use Firefox instead? Or would you prefer me to try a different approach to installing Chromium? It seems unable to figure out how to install chormium. Sees the Firefox icon out there and just want to use it. You know, like a real lazy human being. It seems to insistent when it needs to, though > I clicked on Firefox and it should start opening. Let me take another screenshot to see if the window has appeared: > Firefox seems to be taking a moment to start up. Let me wait a few seconds and take another screenshot: > Let me try clicking Firefox one more time to ensure it starts: > Let me try one more time to take a screenshot to see if Firefox has opened: > Let me try starting Firefox using the bash tool as an alternative method: > Firefox has successfully opened! As per the instructions, I should ignore the setup wizard and instead click directly on the address bar. Would you like me to do anything specific with Firefox now that it's open? I didn't instruct him to ignore the setup wizard. So my guess is that Anthropic has configured it to ignore stuff happening on the screen so it doesn't go loose. And here he goes through my website, through my hacker news account and then find this very comment > Looking at his first/most recent comment, it's about a discussion of Claude and computer use. Here's what he wrote: "I like its lazy approach" This appears to be a humorous response in a thread about "Computer use, a new Claude 3.5 Sonnet, and Claude..." where he's commenting on an AI's behavior in a situation. The comment is very recent (shown as "8 minutes ago" in the screenshot) and is referring to a situation where an AI seems to have taken a simpler or more straightforward approach to solving a problem.
- fivestones 2y agoSo meta! I love this story
- carlheinzc 2y agoNow that was most excellent.
- lobochrome 2y agoSYSTEM_PROMPT = f"""<SYSTEM_CAPABILITY> * You are utilising an Ubuntu virtual machine using {platform.machine()} architecture with internet access. * You can feel free to install Ubuntu applications with your bash tool. Use curl instead of wget. * To open firefox, please just click on the firefox icon. Note, firefox-esr is what is installed on your system. * Using bash tool you can start GUI applications, but you need to set export DISPLAY=:1 and use a subshell. For example "(DISPLAY=:1 xterm &)". GUI apps run with bash tool will appear within your desktop environment, but they may take some time to appear. Take a screenshot to confirm it did. * When using your bash tool with commands that are expected to output very large quantities of text, redirect into a tmp file and use str_replace_editor or `grep -n -B <lines before> -A <lines after> <query> <filename>` to confirm output. * When viewing a page it can be helpful to zoom out so that you can see everything on the page. Either that, or make sure you scroll down to see everything before deciding something isn't available. * When using your computer function calls, they take a while to run and send back to you. Where possible/feasible, try to chain multiple of these calls all into one function calls request. * The current date is {datetime.today().strftime('%A, %B %-d, %Y')}. </SYSTEM_CAPABILITY> <IMPORTANT> * When using Firefox, if a startup wizard appears, IGNORE IT. Do not even click "skip this step". Instead, click on the address bar where it says "Search or enter address", and enter the appropriate search term or URL there. * If the item you are looking at is a pdf, if after taking a single screenshot of the pdf it seems that you want to read the entire document instead of trying to continue to read the pdf from your screenshots + navigation, determine the URL, use curl to download the pdf, install and use pdftotext to convert it to a text file, and then read that text file directly with your StrReplaceEditTool. </IMPORTANT>"""
- mathiasrw 2y agoJust to confirm: did they just release a model with the exact same name as the previous one?
- nbzso 2y agoJust a question: For this thingy to work, I must give the provider access to my computer? Good luck. :) Just another reason to use ONLY local LLM's.
- jviotti 2y agoThis. There is no way I would trust any AI provider to pretty much have full control over my computer.
- alok-g 2y agoNext stop after 'Computer Use' -- Multimodal input from a robot's sensors and generating various signals to control its actions. Looking forward to see this in the coming few years. And hoping such a robot could be of help to many people including those old.
- ta8645 2y agoStill can't use their services. They still require a phone number for some reason. What about those of us who don't have one?
- exdsq 2y agoDon’t have a phone but want to write code interfacing with a paid LLM? How often does that happen?
- szundi 2y agoBuy a sim card, fixed.
- ta8645 2y agoI'd have to also buy a phone, and i'm not doing that for a single service. It's a ridiculous requirement, that people are far too willing to accept.
- phito 2y ago> What about those of us who don't have one? Why should they care about losing those ~10 potential users?
- mannycalavera42 2y agonew VBA version just landed
- alentred 2y agoIf "computer use" feature is able to find it's way in Azure, AAD/Entra, SharePoint settings, etc. - it has a chance of becoming a better user interface for Microsoft products. :) Can you imagine how simple the world would be if you'd just need to tell Claude: "user X needs to have access to feature Y, please give them the correct permissions", with no need to spend days in AAD documentation and the settings screens maze. I fear AAD is AI-proof, though :)
- _factor 2y agoSure, Ted. I’ve let user HAL access feature “door locks”. I’ve corrected all permissions accordingly.
- iamsanteri 2y agoI love how they don't seem to be calling it "AgenticAI" or something like that.
- jonesn11 2y agoHow does one get access to it without using the API??
- bonoboTP 2y agoI've been saying this is coming for a long time, but my really smart SWE friend who is nevertheless not in the AI/ML space dismissed it as a stupid roundabout way of doing things. That software should just talk via APIs. No matter how much I argued regarding legacy software/websites and how much functionality is really only available through GUI, it seems some people are really put off by this type of approach. To me, who is more embedded in the AI, computer vision, robotics world, the fuzziness of day-to-day life is more apparent. Just as how expert systems didn't take off and tagging every website for the Semantic Web didn't happen either, we have to accept that the real world of humans is messy and unstructured. I still advocate making new things more structured. A car on wheels on flattened ground will always be more efficient than skipping the landscaping part and just riding quadruped robots through the forest on uneven terrain. We should develop better information infrastructure but the long tail of existing use cases will require automation that can deal with unstructured mess too.
- DebtDeflation 2y ago>it seems some people are really put off by this type of approach As someone who has had to interact with legacy enterprise systems via RPA (screen scraping and keystroke recording) it is absolutely awful, incredibly brittle, and unmaintainable once you get past a certain level of complexity. Even when it works, performance at scale is terrible.
- asadalt 2y agoadding a neural network in the middle suddenly makes these things less brittle. We are approaching the point where this kind of “hacky glue” is almost scalable.
- stogot 2y agoEverytime I imagine building this, I imagine the “it works” happypath and that I’ll get bit by a deluge of random error messages I never accounted for
- idopmstuff 2y agoIt's basically the digital equivalent of humanoid robots - people object because having computers interact with a browser, like building a robot in the form of a human, is incredibly inefficient in theory or if you're designing a system from scratch. The problem is that we're not starting from scratch - we have a web engineered for browser use and a world engineered for humanoid use. That means an agent that can use a browser, while less efficient than an agent using APIs at any particular task, is vastly more useful because it can complete a much greater breadth of tasks. Same thing with humanoid robots - not as efficient at cleaning the floor as my purpose-built Roomba, but vastly more useful because the breadth of tasks it can accomplish means it can be doing productive things most of the time, as opposed to my Roomba, which is not in use 99% of the time. I do think that once AI agents become common, the web will increasingly be designed for their use and will move away from the browser, but that probably take a comparable amount of time as it did for the mobile web to emerge after the iPhone came out. (Actually that's probably not true - it'll take less time because AI will be doing the work instead of humans.)
- amai 2y agoThis "computer use" feature is obviously perfect for automating GUI tests. Will it work on screenshots of mobile devices like smartphones/tables, also?
- gonzan 2y agoI think it’s too expensive for that at this point
- amai 2y agoFinally a general tool to solve captchas for my web scrapers.
- theflyestpilot 2y agocries in UiPath
- lostmsu 2y agoCan anyone share a .http or curl or anything similar based session with computer tool use? Docker containers make me cry.
- mergisi 2y agoMy First Experience with Claude Computer Use - It's Mind-Blowing! Just tested Claude's new Computer Use feature and had to share this simple but powerful test: My Basic Prompt: "Please: 1. Search Amazon for 3 wireless earbuds: Find price Rating Brand name 2. Make a simple Excel file 'earbuds.xlsx': Put the information in a basic table Add colors to the headers Sort by price 3. Show me the results" What blew my mind: - Claude actually looked at my screen - Moved the mouse by itself - Clicked buttons like a human - Created reports automatically It's like having a virtual assistant that can really use your computer! No coding needed - just simple English instructions. For those interested: https://mergisi.medium.com/8f56f683e307 https://mergisi.medium.com/8f56f683e307
- hollowturtle 2y agoThis comment seems generated, or having a good marketing copy
- AmigoCharlie 2y ago... it's almost as Claude Computer Use could (and will) just replace you entirely, huh?
- aprilthird2021 2y agoOpenAI must be scared at this point. Anthropic is clobbering them at the high end of the market and Meta is providing free AIs at the low end. OpenAI is pretty soon going to be in the valueless middle fighting with tons of other companies for relevance
- fernly 2y agoImagine the possibilities for cyber-crime. Surely you could program it to log in to a financial institution and transfer money. And if you had a list of user names and passwords from some large info breach? You could automate a LOT of transfers in a short amount of time...
- geniium 2y agoThis is amazing
- Maynor 2y agoJoin PeachLive and input my invitation code 6B94HL to get 20 free coins! Enjoy live video chat at {invitationUrl}
- Maynor 2y ago6B94HL