3 ms·
Show HN: AI agent that runs real browser workflows
I’ve been experimenting with letting an AI agent execute full workflows in a browser.
In this demo I gave it my CV and asked it to find matching jobs. It scans my inbox, opens the listings, extracts the details and builds a Google Sheet automatically.
- abraxas 7mo agoI was looking for a similar produc/project the other day. Alas my need is a Linux native version. You may want to consider it as Mac seems to be overserved by the agent harness supply while Linux is the opposite
- heavymemory 7mo agolinux and windows support is on the way, i’ve designed it in a decoupled way, so should be straight forward. Just need to see if people find this version useful
- hkonte 7mo ago[dead]
- heavymemory 7mo agoYeah, instruction drift is a real problem in long agent chains. In this case the workflow gets decomposed into steps up front and each step is executed by a separate sub-agent. So the model isn’t carrying the whole instruction chain across multiple steps, it’s just solving the current task. Similar pattern to what tools like Codex CLI or Claude Code do.
- fidorka 7mo agoCool demo. The tricky bit with browser workflow agents is figuring out which workflows to automate in the first place. Most people don't even realize they're doing the same thing over and over - they just do it. I've been building MemoryLane (https://github.com/deusXmachina-dev/memorylane https://github.com/deusXmachina-dev/memorylane) which comes at this from the other side - it records screen activity, spots repeated patterns with AI, and then tells you "hey you keep doing this, want to automate it?" Works as an MCP plugin for Claude/Cursor. Feels like pattern detection (finding what to automate) + browser agents like yours (actually doing the automation) is the right combo. Are you thinking about the discovery side at all, or mostly focused on execution?
- heavymemory 7mo agoInteresting. Part of why I built this was to avoid screen capture as the control layer. Once you’re taking screenshots, guessing what to click, moving the mouse, and repeating, it gets slow and brittle fast. Here the workflow is just described in text, executed in the browser, and saved for reuse.
- june-jule 7mo agoInteresting demo, how are you thinking about prompt injection and security with web agents? Ive been facing this as well.
- heavymemory 7mo agoPrompt injection is the same problem all agents face, ChatGpt Atlas, claude cowork, openclaw, all of them. It's a known unsolved problem across the industry. I mitigate it by giving the agent a fixed action set (no scripts, no direct API calls), and breaking tasks into focused subtasks so no single agent has broad scope. The LLM prioritises its own instructions over page content, but if someone managed to hijack it, the agent can interact with authenticated sessions. Everything's visible in real time though, and all actions are logged, so you can see exactly what it's doing and kill it. Practically speaking, I use it similar to how people use Zapier or n8n, you set up specific workflows and make sure you're only pointing it at sites you trust. If you're sending it to random unknown websites then yeah, there's more risk. But even then, an attacker would need to know what apps you're authenticated with and what data the agent has access to. The chances of something actually happening are pretty low, but the risk is there. No one's fully solved this yet.
- deleted 7mo ago[deleted]
- clawbridge 7mo agoNice demo. The prompt injection concern is real for any browser agent in production. We've found that the hybrid approach (deterministic steps + AI for the unpredictable parts) combined with step-level observability is what gives teams confidence to run these unattended. The 'fixed action set' is a good start.
- vanessa49 6mo agoOne thing that surprised me while experimenting with agents is how quickly automation runs into the problem of context and memory. Executing workflows is actually the easy part. The harder problem is deciding what the agent should remember about previous interactions and how that memory should influence future behavior. Without some form of long-term memory or learning loop, agents often end up behaving more like stateless automation scripts. It makes me wonder whether the next interesting step for agents isn't just more tools, but systems that can gradually evolve from their interactions with a single user.