11 ms·
Great idea Kyle! I read through the source code as an experienced desktop automation/Electron developer and felt good about trying it for some basic tasks. The
by taroth 2y ago
Great idea Kyle! I read through the source code as an experienced desktop automation/Electron developer and felt good about trying it for some basic tasks.
The implementation is a thin wrapper over the Anthropic API and the step-based approach made me confident I could kill the process before it did anything weird. Closed anything I didn't want Anthropic seeing in a screenshot. Installed smoothly on my M1 and was running in minutes.
The default task is "find flights from seattle to sf for next tuesday to thursday". I let it run with my Anthropic API key and it used chrome. Takes a few seconds per action step. It correctly opened up google flights, but booked the wrong dates!
It had aimed for november 2nd, but that option was visually blocked by the Agent.exe window itself, so it chose november 20th instead. I was curious to see if it would try to correct itself as Claude could see the wrong secondary date, but it kept the wrong date and declared itself successful thinking that it had found me a 1 week trip, not a 4 week trip as it had actually done.
The exercise cost $0.38 in credits and about 20 seconds. Will continue to experiment
- computeruseYES 2y agoThanks so much, valuable information, sounds much faster than we heard about, maybe cost could be brought down by sending some of the prompts to a cheaper model or updating how the screenshots are tokenized
- taroth 2y agoThe safety rails are indeed enforced. I asked it to send a message on Discord to a friend and got this error: > I apologize, but I cannot directly message or send communications on behalf of users. This includes sending messages to friends or contacts. While I can see that there appears to be a Discord interface open, I should not send messages on your behalf. You would need to compose and send the message yourself. error({"message":"I cannot send messages or communications on behalf of users."})
- taroth 2y agoGave it a new challenge of > add new mens socks to my amazon shopping cart Which it did! It chose the option with the best reviews. However again the Agent.exe window was covering something important (in this case, the shopping cart counter) so it couldn't verify and began browsing more socks until I killed it. Will submit a PR to autohide the window before screenshot actions.
- rossjudson 2y agoHow many sockets got delivered? Did it use a referral link?
- stefan_ 2y agoWhy on earth would that be a "safety rail"?
- ceejayoz 2y agoSending spam?
- kcorbitt 2y ago(author here) yes it often confidently declares success when it clearly hasn't performed the task, and should have enough information from the screenshots to know that. I'm somewhat surprised by this failure mode; 3.5 Sonnet is pretty good about not hallucinating for normal text API responses, at least compared to other models.
- InsideOutSanta 2y agoI asked it to send a message in WhatsApp saying that "a robot sent this message," and it refused, because it didn't want to impersonate somebody else (which it wouldn't have). Next, I asked it to find a specific group in WhatsApp. It did identify the WhatsApp window correctly, despite there being no text on screen that labelled it "WhatsApp." But then it confused the message field with the search field, sent a message with the group name to a different recipient, and declared itself successful. It's definitely interesting, and the potential is clearly there, but it's not quite smart enough to do even basic tasks reliably yet.
- arijo 2y agoWe could maybe chose the target window as the screenshot capture source instead of the full screen to prevent it to be hidden buy the Agent: ``` const getScreenshot = async (windowTitle: string) => { const { width, height } = getScreenDimensions(); const aiDimensions = getAiScaledScreenDimensions(); const sources = await desktopCapturer.getSources({ types: ['window'], thumbnailSize: { width, height }, }); const targetWindow = sources.find(source => source.name === windowTitle); if (targetWindow) { const screenshot = targetWindow.thumbnail; // Resize the screenshot to AI dimensions const resizedScreenshot = screenshot.resize(aiDimensions); // Convert the resized screenshot to a base64-encoded PNG const base64Image = resizedScreenshot.toPNG().toString('base64'); return base64Image; } throw new Error(`Window with title "${windowTitle}" not found`); }; ```
- taroth 2y agoYup that could help, although if the key content is behind the window, clicks would bug out. I'm writing a PR to hide the window for now as a simple solution. More graceful solutions would intelligently hide the window based on the mouse position and/or move it away from the action.
- arijo 2y agoI think you can use nut-js desktop automation tool to send commands straight to the target window ``` import { mouse, Window, Point, Region } from '@nut-tree-fork/nut-js'; async function clickLinkInWindow(windowTitle: string, linkCoordinates: { x: number, y: number }) { try { // Find window by title (using regex) const windows = await Window.getWindows(new RegExp(windowTitle)); if (windows.length === 0) { throw new Error(`No window found matching title: ${windowTitle}`); } const targetWindow = windows[0]; // Get window position and dimensions const windowRegion = await targetWindow.getRegion(); console.log('Window region:', windowRegion); // Focus the window await targetWindow.focus(); // Calculate absolute coordinates relative to window position const clickPoint = new Point( windowRegion.left + linkCoordinates.x, windowRegion.top + linkCoordinates.y ); // Move mouse to target and click await mouse.setPosition(clickPoint); await mouse.leftClick(); return true; } catch (error) { console.error('Error clicking link:', error); throw error; } } ```
- TechDebtDevin 2y agoSo the assistant I could pay to book me incorrect flights would cost $68.00 and hour. This makes me feel a little better about the state of things.
- malfist 2y agoYeah, but that assistant won't book the wrong flights.
- delusional 2y agoI'd say correctness would be worth another 40 bucks an hour.
- pants2 2y agoPresumably every step has to also read the tokens from the previous steps, so it gets more expensive over time. If you run it on a single task for an hour I would not be surprised if it consumed hundreds of dollars of tokens.
- vineyardmike 2y agoI’m curious how many tokens this used, and what the actual effective maximum duration it has due to the context window.
- MacsHeadroom 2y agoGenAI costs go down 95% per year. So next year it will be $3.40/hr and more reliable.
- TechDebtDevin 2y agowanna bet?
- IanCal 2y agoPer hour of computer execution is a poor measure. Imagine it did this twice as fast, and cost the same. Is that worse? A per hour figure would suggest so. What if it was far slower, would that be better?
- jrflowers 2y ago> The exercise cost $0.38 in credits and about 20 seconds I am intrigued by a future where I can burn seventy dollars per hour watching my cursor click buttons on the computer that I own
- bastawhiz 2y agoAmazingly my employer continues to pay me hundreds of dollars an hour to search Kagi and type on a computer they paid for and own!
- jrflowers 2y agoAnd to think they could be paying you to supervise the buttons clicking themselves instead! The past where the lack of a human meant a lack of input is over, all hail the future where a lack of a human could mean wasteful and counterproductive input instead
- bastawhiz 2y agoWhat I'm hearing is that now they can fire my manager
- tylerchilds 2y agoi think you’d get fired and your boss will be demoted to your position.
- ionwake 2y agoa smart take
- urbandw311er 2y agoYou wouldn’t sit there watching your paid human assistant work would you? So why would you sit watching your paid AI assistant? I think the general idea is that you’re off doing something more productive, more relaxing or more profitable!