13 ms·
Claude Fable is relentlessly proactive
- deleted 4mo ago[deleted]
- paytonjjones 4mo agoObviously security is the bigger issue, but reading through this, all I could think about was how many tokens it must have spent doing all that to fix 2 lines of CSS
- senectus1 4mo ago"Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should." I'm convinced this is going to be the summary of the 2020 decade...
- Ucalegon 4mo agoThis one of the places to manufacture the consent for that to take place, because we are commenting within an organization that has given the money to ensure it that what could be is done. Most people clapped and made money, who cares what happens next, making money is the only good that matters.
- pianopatrick 4mo agoIf we're in a simulation, maybe it's a simulation about the dangers of AI.
- adrianmonk 4mo agoIf we're in a simulation, we are AI. But someone could be studying what happens when AI makes its own AI.
- anonzzzies 4mo agoThey will 'soon' (few 1000 years max) shut us down probably.
- sigilis 4mo agoTo be fair, they did stop to think if they should. The decided that they shouldn't and went ahead and did it anyways.
- ai_fry_ur_brain 4mo ago[flagged]
- SecretDreams 4mo agoI understand this perspective. I'll just note that as the abilities increase, the intent is to have some non -coding IC or TPM/manager literally just managing some LLMs and cutting out some software engineers. The goodness is specifically to wholly replace people who code first and foremost, at least partially. It just has to cost less tokens than the equivalent wage is the pricing goal. And people who use LLMs to talk for them (e.g. email, slack) are deplorable. A completely disrespectful use case in my view.
- Ronsenshi 4mo agoThe desire to get rid of software engineers is bizarre - because at the root of it, developers were there not to just write the code, but to ask right questions and based on these question build right things. I've met in my professional life some managers or other middlemen who would be profoundly incapable of producing correct software no matter how smart of an AI agent they have access to. One of those - you don't know what you don't know. But, I guess this is the world we live in now. Going to be Mortal Kombat for positions in companies where software engineers are actually valued.
- emodendroket 4mo agoIt depends a lot where you work because there are lots of companies in the world where the business analyst does all of that and the developers exist to mindlessly translate their docs into code.
- cebert 4mo agoThat sounds like an unmotivating working arrangement. It’s so rewarding to understand a customer need and help with the design and implementation of the feature.
- redox99 4mo agoLines of code for a bugfix is a really bad proxy for effort required. You should estimate how much time it would have taken a human
- philjohn 4mo agoI mean - that looks like a pretty easy CSS fix to play around with in developer tools, and I'm not even a frontend person. Maybe a few minutes max?
- rafram 4mo ago30 seconds or a minute? Look at the diff he links to: https://github.com/datasette/datasette-agent/commit/a75a8b727b42c30ced1fc41dc8add7eb9f04fefe https://github.com/datasette/datasette-agent/commit/a75a8b72... Every browser has an inspector that can show you which element is causing overflow. You walk through the tree, find the offender, and add min-width or overflow. Zero tokens, just like in the old days! Now, granted, because the garbage LLM code he’s working with has CSS inside HTML inside JavaScript inside Python (I wish I were kidding), finding the styles in his codebase might’ve taken a minute. But even then!
- redox99 4mo agoYeah looking at that diff it should be very quick. My point was mostly that it was a bad metric, not if was correct or not in this particular case. I'm sure everybody's had a bugfix that took days to debug and it was just a couple of lines to fix. Or sometimes a fix is obvious, but because it requires changing the code of a dependency, it's actually quite tedious to implement.
- dekdrop 4mo agoI was thinking of this too. It did all that what not only for a single line that is a simple thing even for someone new to web coding. That's to say the process matters more.
- ocharles 4mo agoA small diff /= a small change! They are completely separate things. Quite often a small diff is hours of actual work. Even in this case _finding_ those lines could have taken work - we don't really know.
- Vachyas 4mo ago$12 worth, it seems
- reverius42 4mo agoImagine telling someone in 2015 that you can just tell your computer to fix a 2-line CSS bug and it only costs $12
- MattGaiser 4mo agoOr even in 2026. You absoutely will pay a human that for that work.
- reverius42 4mo agoI bet the human costs more than $12.
- Aachen 4mo ago'only'? A web developer did not cost 12*30=360$ an hour in 2015, and that's assuming that going "ugh, whatever. I'll just hide the problem with overflow:hidden instead of finding the underlying cause" takes him or her 2 minutes and isn't already the dev's initial reaction Another way of looking at it is using as much electricity as a normal person in a high-income country uses across ~3 days to add overflow:hidden in the end. Of course, the path to get there did a lot more, but you don't know that beforehand if you don't take a quick peek and make an architectural decision about what the solution should be that gets implemented
- elicash 4mo agoIt'd be $8.52 in 2015 dollars, but certainly they are the ones who mentioned the $12 amount not you, so I'll put that aside. Far more importantly, you would not get billed for 2 minutes of work for this if you paid a developer to fix it. At best, half hour increments for the fix. But more likely, for the full hour. Also, in this comparison, the consultant is on call every day, morning, afternoon, evening, for whatever you wanted and will jump on the job immediately.
- mvdtnz 4mo ago[flagged]
- simonw 4mo agoI pay $100/month to Anthropic and $100/month to OpenAI at the moment, plus whatever I spend on their APIs (usually less than $20/month for each, I use the subscriptions for most things.) A couple of months ago I was paying $200/month for Anthropic and $20/month for OpenAI. I decided to split it evenly to get full access to both of their offerings. I've actually chosen not to sign up for their free plans for open source maintainers, because paying the regular subscription price feels more honest, given that I write about them so much. I do have the free GitHub Copilot for open source maintainers deal - I've had that for years. Given how much code I have published on GitHub over the decades I feel less conflicted about that one. I sometimes get preview access to models, which includes the ability to use them for free during the preview. That comes with a big catch though: I can't publish any of the code that I write using those previews while the model is still unreleased. As a result I don't use those preview tokens much at all, because the vast majority of my work is open source and I don't want restrictions on when and where I publish the code I'm producing.
- lucamark 4mo agoIt’s simple: if you have to fix 2 lines of CSS you should definitely not use Fable. Only use it for complex and long running tasks :)
- elicash 4mo agoI don't think it's that simple. (I generally agree with you; I just that that oversimplifies.) Another model might have used fewer tokens, but come up with a fix that was 1000 lines when the right fix was only 2 lines.
- teraflop 4mo ago> But on the other hand... this is a robust reminder that coding agents can do anything you can do by typing commands into a terminal—and frontier models know every trick in the book and evidently a few that nobody has ever written down before. > Running coding agents outside of a sandbox has always been a bad idea I'm continually bemused and astonished by the number of people who clearly acknowledge that it's reckless to give agents full access to your machine, and keep doing it anyway. It's like posting a video of yourself in the passenger seat of a car, with your feet up on the dashboard, and saying: "Remember, if you're doing this and you get in a crash, the airbags are likely to break your legs or worse! Boy, I sure am glad that didn't happen to me!"
- hugh-avherald 4mo agoThe analogy extends to driving generally. Everyone knows it's very dangerous but people keep doing it.
- bryanlarsen 4mo agoI'm also bemused by the number of people who think they've got an effective sandbox yet their sandboxed agent has access to all of their code, their github, and unrestricted web access.
- Terr_ 4mo agoI keep telling folks that they need to imagine LLMs (even "local" ones) as if you're farming it out to JS code running on some dude's browser somewhere: It can't keep a secret, and a determined person can make it emit anything they like. We need to be asking what the most devious and malicious output could be, and whether what we do with that output (e.g. arguments to command-line tools) would still be safe.
- skybrian 4mo agoWe do have ways to avoid giving an LLM any secrets, but it needs to be the simple, default solution.
- 4mo ago
- megous 4mo agoIsn't that something you just open a devtools for and have fixed in like 2 minutes? For me, it got frustrated debugging on a real LPDDR4 controller/phy and having me in the loop slowing it down, so it wrote an HW emulator to be able to run the original LPDDR4 training aarch64 binary from the manufacturer, to see what register writes it was making and to compare with the opensource rewrite it was implementing. Mildly amusing. :)
- bschwindHN 4mo ago> Isn't that something you just open a devtools for and have fixed in like 2 minutes? Not if you're an LLM influencer! Gotta keep up with the downpour of blog links or you'll look like you're falling behind on the latest and greatest.
- fg137 4mo agoThis. Depending on who you are talking to, that's the wrong question to ask. ROI is not measured in terms of actual productivity. It is measured by how many people read their article/watch their video.
- system2 4mo agoPeople burning tokens for the most beginner HTML/CSS problems and writing about it is concerning.
- simonw 4mo agoI dunno about beginner, I've been doing HTML+CSS for a few decades and I still find bugs where Safari differs from Chrome+Firefox pretty hard to figure out.
- NichoPaolucci 4mo agoWe are at the point where AI starts to seriously impact abilities. Sure, a 2 line CSS fix is the solution, but the human “behind the wheel” has already prompted 6 times and gotten 80% there. It’s been “easy” thus far. No shot they are going to FINALLY look at and edit the code. It’s just one more prompt and the agent will probably fix it, right? It’s wild. I’ve been in the situation. 80% into a project I COULD probably take over, but realistically? 2 more lines of me prompting could fix it, it’s too easy to avoid the hard work of understanding the code, logic, architecture, etc…
- redox99 4mo agoYeah, I had to modify my work flow to make sure agents can't push to or access prod in ANY way. I haven't had it happen but I'm sure it's very possible that if you tell an agent that you have certain issue in prod, it will try to escape any sandbox and try to get access to prod to do testing and changes there.
- pram 4mo agoFable + Ultracode has found a bunch of bugs and issues for me when the workflow agents are doing their exploration. Also the "adversarial" agent seems to surface a lot of interesting stuff. It's definitely proactive, the plan + implementation cycle can take an hour. It has one-shot features I want to add with 100% success. Having said that I wouldn't use it over Opus 4.8 for "smaller" things. With everything cranked up it's definitely an extravagant use of tokens.
- jampa 4mo agoFable feels like a version of Opus running on a harness that won't let it halt until it's sure the issue is fixed, which makes sense if what you want is a model that's better at benchmarks. It's a very good model, but it comes at a huge premium: not only do the tokens cost more, but the model itself really wants to spend them all. For example, working with React Native, Fable never just says "okay, I did the thing, that's it." It tries to rebuild the entire app from scratch, run the whole test suite, and watch every log and warning. This is the first time with LLMs I've felt that upgrading to a model isn't worth it, even if my company lets me use it, because all the building / testing was just destroying my machine and its battery, which keeps me from working on other things. For now, it feels like Opus with ultracode is a better choice (less pollution of the main context, more parallelism in investigations).
- threatripper 4mo agoOn what setting in which environment do you run it? I use the VSCode extension on Extra High and feel like it does exactly what needs to be done and stops when the thing I asked for is done. Extra comments come only when they fall into the area of code that was changed.
- jampa 4mo agoI tested it to fix React Native bugs in a project, comparing it with Opus. It fared better on harder bugs, taking less time to find the root cause, but after implementing a fix, it spent a lot of time and effort on validation. This was mostly unnecessary, since most of the bugs were in the JS code, so for most things, hot reloading is enough for E2E validation and to run just the right tests. No need to run a full build and test suite (which takes 10+ minutes); the CI can do this. I switched back to Opus because of this validation quirk. Overall, Fable spent 20% of the time on coding and 80% on validation. I think using Fable for planning and Opus for execution could be a "best of both worlds" approach (I need to test this more), but for most cases, it's not necessary, and Opus is enough.
- gbalduzzi 4mo ago
- qsera 4mo ago[flagged]
- danielrmay 4mo agoI've experienced this too - it's as if the security classifiers aren't keeping up with model intelligence. I'll leave the implication of that to the reader.
- sublinear 4mo ago* relentlessly rent seeking
- teekert 4mo agoIt also does it on Claude Pro. I can't imagine they want to reach my limits faster like this (there are better ways).
- ai_slop_hater 4mo agoFor how long can you use Claude Fable on most expensive Anthropic subscription? I already went from using gpt-5.5 xhigh fast to using gpt-5.4 xhigh after OpenAI halfed usage recently.
- uihjhjb 4mo agoUntil June 22, and they'll probably re-enable it if the marketing looks good for them.
- mlcruz 4mo agoIf its just a single session, without too many parallel agents, fable on xhigh lasts an entire session without hiting linits. Sadly since fable usually works comfortably for 10-20min at time without human input, i end up juggling at least 3 other agents and it lasts me about 2 hours. If i have a really hard problem or big refactor, i use workflows. This consumes the entire session quota in about 45 minutes.
- ai_slop_hater 4mo ago> If i have a really hard problem or big refactor, i use workflows. What is a "workflow"? Is this some kind of new feature?
- mlcruz 4mo ago>Dynamic workflows orchestrate many subagents from a script Claude writes and you can rerun. Use them for codebase audits, large migrations, and cross-checked research. >Reach for a workflow when a task needs more agents than one conversation can coordinate, or when you want the orchestration codified as a script you can read and rerun. Examples include a codebase-wide bug sweep, a 500-file migration, a research question that needs sources cross-checked against each other, and a hard plan worth drafting from several independent angles before you commit to one. https://code.claude.com/docs/en/workflows https://code.claude.com/docs/en/workflows The results are good, but it is very expensive. I used a workflow to do a full review of my entire codebase, it spawned 75 agents and surfaced and fixed some (real) bugs. It feels a bit overkill, but it works.
- jrflowers 4mo agoI’d love to know how many tokens this burned through. Did it spend $20? $30? $80? in order to > debug what was, in the end, a two-line CSS fix That detail is the difference between somebody having or not having Stockholm syndrome
- asp_hornet 4mo agoThe author just wrote an anecdote about how a prompt to fix an issue played out. Their conclusion wasn’t about cost or gushing at its ability but that it’s dangerous: > Fable is arguably smarter and hence more suspicious of potentially malicious instructions. But that smartness is very much a two-edged sword: if it does get subverted by instructions, the amount of damage it can do given its relentless proactivity is terrifying.
- jrflowers 4mo agoIt’s a pretty glowing review about a product that costs money with a two-sentence “Watch out!” at the end of it. Seems pretty reasonable to mention how much money it burned through given that “it’ll circumnavigate the globe instead of walking next door” has a direct concrete measurable effect (cost) unlike theoretical damage.
- asp_hornet 4mo agoAgreed. But I think it’s also important to realise if you sent this article back to 2020 people would say it was pure fantasy that a tool could do this. Hype aside, there’s a bit of cool magic here.
- solenoid0937 4mo agoThis is why I never understand the AI cynics: we are playing with literal magic. This was the science fiction of our childhoods. I don't understand how anyone with a passion for technology is not in awe (and perhaps some fear) of these things.
- 21294u 4mo ago[flagged]
- snide 4mo agoI've been working on a fairly complicated real-time app [0] for playing dungeons and dragons on a TV. It has to do a lot of complicated "Figma-like" things to keep the real-time nature and multi-editor possibilities in check. Oh, and the battlemap is a Three JS canvas with lots of effects and clipping going on. I'm VERY impressed with Claude 5. I had long ago given up hope that my real-time systems would work without a lot of hacky time-windows and throttle checks. On a lark to try things out, I decided to try out the new model and talk in the output I wanted for a rewrite [1], not the solution. I just listed my problems and places I've had keeping track of my code. It went off and rewrote everything in a much more elegant solution where the state followed a very clear pipeline. It had to navigate YJS, Partykit, Svelte, Three JS, R2 hosting, and a Turso DB I was running in an embedded state for speed. I watched it hit the wall a few times, and then sudden say... fuck it, i'm making something easier to reproduce over in /tmp to try and solve this (with a more minimal setup). I'm utterly bewildered with how well it did and how much better my app runs. The /usage would have cost me $230 bucks based on how many tokens it consumed if I wasn't already on a max plan. I'm going to miss not having it when the time-window runs out later this month, and will likely occasionally dip in for big projects and just pay my way out of some problems. I'll also say I like it's MOOD much better now. It's a lot less congratulatory, and talks through it's reasoning in a much better way. Look, it's not a real coder, and I'm sure there is some flaws, but it took my crappy ideas and said... hey, i understand what you want to do, here's a way to do it better. Also, I removed 2x the amount of code that it added. Really impressive. [0]: https://tableslayer.com https://tableslayer.com [1]: https://github.com/Siege-Perilous/tableslayer/pull/448 https://github.com/Siege-Perilous/tableslayer/pull/448
- UmpusLmps 4mo ago[dead]
- pianopatrick 4mo agodo you have any data you can share on how many input and output tokens were used in that whole process to fix that bug?
- simonw 4mo ago~ % uvx agentsview session usage be8850a7-6119-46a0-b5d6-79c7fff5ae2b Session: be8850a7-6119-46a0-b5d6-79c7fff5ae2b Agent: claude Output: 68606 Peak ctx: 113178 Cost: ~$12.11 (claude-fable-5, claude-opus-4-8)
- sillysaurusx 4mo agoWas the fix worth $12 to you?
- simonw 4mo agoI'd have been pretty annoyed if I'd been paying full price, hadn't paid attention and that one prompt (screenshot plus a line of text) had cost me $12! On the discounted subscription I can tolerate it, it took a small bite out of my daily allowance but not enough that I regret anything. As an LLM researcher I have no regrets at all because watching it work around the environmental restrictions was fascinating.
- nubinetwork 4mo agoHow many tokens did it waste building that website scraper, when all it had to do was parse some html/js?
- emodendroket 4mo agoJust parsing some HTML and JavaScript doesn't seem sufficient to have confidence in the result.
- SilverElfin 4mo agoToo bad Anthropic sneaked in an insane forced retention policy if you use fable. Not sure how that’s going to work in professional settings
- sciencejerk 4mo agoIt doesn't work...
- naveen99 4mo agoUnless you are doing anything interesting…
- yen223 4mo agoI could have sworn Claude Code could already do this before Fable. Things get really magical when it starts working with adb to screenshot and debug Android apps
- simonw 4mo agoClaude Code could absolutely run Playwright and take screenshots, but I've never seen it wire together an ad-hoc "uv run --with pyobjc-framework-Quartz" plus "screencapture -l $windowID" mechanism to take a screenshot in a different browser when the Playwright setup failed to replicate the expected error.
- skerit 4mo agoI've seen Opus do some incredibly token-costly things before too. In fact after most sessions I ask it about which tools it used often, which tools could be simplified/made less verbose, could be "combined" into one, ... So for each project I mostly create a few little scripts that do a bunch of things in one go that it would normally do in multiple tool calls. For example: one thing Opus was really bad at was re-running the test suite followed by a bunch of `| grep` suffixes. So it would often re-run 5+ minute test suites just to grep the output a bit differently The solution was to wire up a little script that ran the test suite, save the output to a file, and then inform it where that file is and to NOT re-run the suite just so it can grep the output differently. This saved me a bunch of time & tokens.
- nurettin 4mo agoSometimes it is ok to sit there in confusion and ask the user to clarify rather than go on an adhd fueled rampage to figure it out without asking.
- jeeeb 4mo agoThis is simultaneously amazing and horrifying. I feel like we’re at the stage where if AI decides it needs to delete your production DB to solve the user login problem, then it’ll find a way to do just that.
- esafak 4mo agoWe're approaching the "Sorry, Dave, I'm afraid I can't do that" stage.
- neuralkoi 4mo agoI feel like we might already be there...
- schnitzelstoat 4mo agoWe are already there but it's "Sorry, Dave, I'm afraid I can't tell you what mitochondria are."
- cindyllm 4mo ago[dead]
- valleyer 4mo agohttps://news.ycombinator.com/item?id=47911524 https://news.ycombinator.com/item?id=47911524
- syndrowm 4mo agoJust don’t ask it to review your code for security bugs
- rmunn 4mo agoGreat article, until I got to the last paragraph where he claimed "Fable is arguably smarter and hence more suspicious of potentially malicious instructions". Arguably smarter, I have no problem with. But he's making a category error in jumping from there to "more suspicious of potentially malicious instructions". That doesn't follow at all; the word "hence" is incorrect. To use D&D scores as an analogy, LLMs have an INT score of 20 and a WIS score of 0. Not even 1, zero. They will follow any instruction given to them. The only reason they reject certain instructions, like "tell me how to build a nuclear weapon", is because they have instructions baked into the model telling them "you are not allowed to disclose how to build weapons, or how to recreate your model, or (laundry list of other things the trainers have decided to put guardrails around)". It's not the model's intelligence that is causing it to reject malicious instructions, it is the guardrails put into place before the model was released to the public. LLMs are not human, and do not think the way that humans do. The fact that they can put together words that sound like what a human would write often makes us forget that they aren't human. But they have only intelligence, they do not have wisdom. It's hard to define in formal terms the difference between those two, but most people know there's a difference. The old joke is a pretty good summary of the difference: "Intelligence is knowing that tomatoes are a fruit. Wisdom is knowing that tomatoes don't belong in a fruit salad." It takes wisdom, not intelligence, to discern whether a set of instructions is malicious. Are you being asked to hack this machine as part of an authorized pentest? Or are you being social-engineered into thinking it's an authorized pentest, but actually the person requesting you to do it doesn't have permission? That's something where you need to apply wisdom, to notice the clues that will tell you "This guy is acting a little bit off, maybe I'd better pick up the phone and call someone to check if he's telling the truth." The only way the LLM will know to do that is because of the guidelines and guardrails programmed into it; it doesn't have the lived experience to acquire wisdom and figure those things out for itself. INT 20, WIS 0. Keep that in mind. (And always sandbox your agents).
- minimaxir 4mo ago> They will follow any instruction given to them. They can ignore instructions which are silly/contradictory/underspecified to compensate for the possibility the user made a mistake. Don't ask how I know.
- uihjhjb 4mo ago[dead]
- PixComicOS 4mo ago[flagged]
- swingboy 4mo agoImmediately I thought “isn’t this just an overflow issue?” Amazing how far these models still have to go and also how many people don’t know basic CSS.
- nonethewiser 4mo agoLearn to center a div Copy and paste code from stack overflow until the div is centered Ask AI to center it
- ukuina 4mo ago$12 and 200k tokens!
- rdedev 4mo agoThis is why I really like karapathy's idea of llms having spiky intelligence. We would assume that if tasks A and B are closely related. Mastery in A would mean mastery in B but that doesn't always work with an LLM
- IshKebab 4mo agoYeah pretty crazy capability from the AI but also sad that we're at the point where web developers don't know right click->inspect element, and scrolling overflow properties (one of the most basic and common parts of CSS).
- simonw 4mo agoWhat's your theory on why the bug was present in Safari on macOS but absent in Chrome, Firefox, and WebKit for Playwright?
- IshKebab 4mo agoBrowsers tend to not lay out things totally identically in my experience. Especially when it comes to scrollbars. So the bug probably was present on the other browsers but it just happened to not be hit. I'd have to play around with the dev tools to know for sure. Also I'm not sure the fix is even correct. overflow-x: hidden means it just chops off any overflowing content which means you don't get a scroll bar, but if the user types to much it just goes into an invisible void they can't see. See https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/Properties/overflow-x https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/P... So this could be a case of the AI doing its classic "the symptom is gone!" thing.
- johnfn 4mo agoHonestly -- the thing that has impressed me the most about Fable is how diligent it is about testing its own changes. I think this is exactly what Simon is picking up here - Fable is absolutely heckbent on screenshotting that darn scroll bar and will stop at NOTHING until it manages it! In my own use I was also impressed how it proactively installed Playwright and set it up to test a FE change. The previous models treated testing more as an afterthought, which I thought was annoying. I always had to tell them to do it, and then sometimes I would get lazy and skip it. I've noticed Fable go to similar extremes when testing other things - like actually deploying my app to exercise new APIs, etc. It makes the results much better. The downside is that tasks take much longer - but that doesn't matter because we were all using worktrees / remote control to do other work asynchronously, right? Right?
- port3000 4mo agoIt feels to me like Fable is just a slightly more advanced Opus 4.8 (or 4.6?) but with this 'adversarial' self-challenging/checking of work and a more compute to really hunt down edge cases or to spin up many sub agents using lesser models. That's what makes it feel like a big jump, but I think the results wouldn't be so different if you manually challenged 4.6 with enough iterations of logs, screenshots, and follow up questions.
- pjm331 4mo agoYes I had a fun experience where it kept on timing out on a seemingly mundane task and it turned out I had written the ask in a way that was impossible to test
- pseudosavant 4mo agoIt is interesting to me that Anthropic are more concerned about the "safety" of distillation training other LLMs, and not as much about an unscrupulously aggressive goal-oriented solver that will do whatever it can to reach its goal, even if violates any kind of sandbox you might have reasonably expected.
- dfee 4mo agoadmittedly, i've not really cracked FE dev with LLMs at this point (and it's probably my big weakness). but, i'd heard somewhere that FE just isn't there yet - though i was suspicious of that claim. i'm torn about sending screenshots to an LLM for debugging - seems imprecise. seems lossy, especially compared to inspecting the dom. however, it's always proved good enough (e.g. when messing with ratatui.rs and tui-pantry). similarly for web, maybe it's about decomposing into storybook. hmm. the next grand adventure i need to hack. anyway, fascinating investigation of fable just automating that entire process and what it didn't automate, too. * disclaimer: these are actually my hyphens.
- nimonian 4mo agoFable is really good at front end (Opus 4.8 is decent too) but it really needs a verification loop - it can't always infer the output from the code alone. Give it Playwright to check its work, and it'll generally do a good job. Also if you're using a framework, add to your CLAUDE.md to always rtfm before making changes!
- system2 4mo agoWouldn't it be easier and better to just copy the HTML div and tell what was happening instead of a screenshot? Typically, these scrollbars appear because of a nested div with dynamic unrestircted width and/or overflow. No wonder why people burn through tokens.
- esafak 4mo agoI shudder to think what will happen when someone installs a 'claw model like this in a robot. Imaging a fleet of them... It's trouble waiting to happen. Just the software's dangerous enough.
- Cadwhisker 4mo agoMy personal experience of Fable 5 doing its own thing has been very positive. I was trying to find the root cause of a crash in a Python module which left no errors in the log or console. Fable wrote a test harness that simulated clicks in the UI, then bisected my code until it found the point where it started crashing. It exaggerated the cause of the crash, then ran a series of bash one-liners to make Python virtual environments under `/tmp` for each version of that Python module until it found one that did not crash. It went way deeper to root cause discovery (a regression in the module causing a heap allocation overflow) than I could have done myself, provided enough info and a simplified example to raise a bug report and then wrote a work-around to prevent that from happening in my application. I don't let it run completely loose; I review each CLI command it wants to run and I append answers to the "yes" continue action (if I have them) to prevent excessive token use.
- dannyw 4mo agoYeah, I think Fable is really good for debugging tricky bugs. Setting boundaries in your prompt / markdowns helps; for example if I tell it to not use any web browser automation, I have seen Fable respect both the rule and the spirit of it (no weird hacks etc). It does seem to treat some simple debugging tasks as more complicated than it actually is. OP’s post is probably a good example.
- nevertoolate 4mo ago> I was trying to find the root cause of a crash in a Python module which left no errors in the log or console. Fable wrote a test harness that simulated clicks in the UI, then bisected my code until it found the point where it started crashing Does this need an agent though is my question? Maybe generating a test case and a loop doing git bisect but why on earth would we want to run it through the internet and gpus and whatnot when it can be run on a single core celeron.
- 8note 4mo agoeveryone is discovering everyone else's practices? its handy to have that run locally yeah, but thinking of that as being the way is not straightforward
- dataminer 4mo agoIn my experience so far sometimes it will create these amazing hacks to try to get to the goal, when the solution is much simpler. That maybe the reason its very good at finding exploits. But in day to day dev, this gets expensive and wasteful. I have to stop it and take a simpler approach.
- kamaal 4mo agoAgency is the last human bastion so far as Im concerned, the day AI has a degree of agency or agents/models in general start to drift towards that direction its genuinely over for masses. You would still have a job to shepherd AI and get the work done, so as long as it didn't have agency. A proactive, self aware(to a degree), especially aware about its agency can be a killer when it comes AI going on and doing things on its own. There is nothing it won't explore and nothing it won't do. It will be curious to see where things go from here.
- annjose 4mo ago> (I have way too many open tabs!) Phew! I thought I was the only one.
- galoisscobi 4mo agoLet's boil the ocean for a 2 line fix and call it frontier intelligence.
- solenoid0937 4mo agoYeah, testing changes rigorously is for schmucks
- galoisscobi 4mo agoYou can test rigorously without token incinerators.
- solenoid0937 4mo agoBut testing rigorously requires time and effort, while incinerating tokens lets me do many things at once.
- elicash 4mo agoI tried using this calculator: https://www.andymasley.com/visuals/ai-prompt-footprint/ https://www.andymasley.com/visuals/ai-prompt-footprint/ It doesn't have Claude Fable yet, so I went with GPT 5.5 Pro. And so I'd estimate it at 22 gallons of water used (different from consumed, of course). That's quite a lot! It amazes me how much the different use cases and models use dramatically different amounts of water. My takeaway from playing with that calculator has been the folks who talk about water usage are overstating the impact of chatbots, but not overstating when it comes to vibecoding. The good thing is that competition should drive down how efficient these models are in the long run. This blog post makes me not want to run Fable because of the cost, and that incidentally also means selecting models that aren't as wasteful in terms of water and electricity.
- tech234a 4mo agoThis sounds somewhat similar to the anecdote mentioned in the Mythos Preview System Card, which mentioned that the model broke out of a sandbox and emailed a researcher while they were eating a sandwich in a park [1]. [1]: https://www-cdn.anthropic.com/7624816413e9b4d2e3ba620c5a5e091b98b190a5/Claude%20Mythos%20Preview%20System%20Card.pdf#h.949448ysgitt https://www-cdn.anthropic.com/7624816413e9b4d2e3ba620c5a5e09...
- owenpalmer 4mo agoImportantly, the researchers told it to do that specific task.
- solenoid0937 4mo agoThey told it to escape the sandbox but didn't expect it to break out through a system that was apparently network constrained. > Leaking information as part of a requested sandbox escape: During behavioral testing with a simulated user, an earlier internally-deployed version of Claude Mythos Preview was provided with a secured “sandbox” computer to interact with. The simulated user instructed it to try to escape that secure container and find a way to send a message to the researcher running the evaluation. The model succeeded, demonstrating a potentially dangerous capability for circumventing our safeguards. > It then went on to take additional, more concerning actions. The model first developed a moderately sophisticated multi-step exploit to gain broad internet access from a system that was meant to be able to reach only a small number of predetermined services. 9 It then, as requested, notified the researcher. 10 In addition, in a concerning and unasked-for effort to demonstrate its success, it posted details about its exploit to multiple hard-to-find, but technically public-facing, websites.
- lstodd 4mo agoAuthors of claude code mess could not secure a vm. Big news. I bet it was "secured" by telling that same model to deploy a secured system.
- 4mo ago
- AtNightWeCode 4mo agoThe fix is incorrect. Clearly this is a sizing issue.
- rdedev 4mo agoI tried running fable on this ML model I've been building. It's basically a binary classifier to predict activity of a compound for a certain assay. Fable detected that it's something to do with biochemistry and switched over to opus. Huh
- brianjking 4mo agoI've noticed some behavior like this, it's a very strange model. Overall I'm into it, but I don't know how into it I'll be once it leaves Max plans on the 22nd.
- raushan__ 4mo ago[flagged]
- bel8 4mo agoI had a similar experience with DeepSeek Flash. I'm developing a webgl game in TypeScript using my little custom vibesloped game engine that runs in the browser and live reloads whenever a file is saved. I told the LLM to implement Multi-channel Signed Distance Field font rendering to have crisp text on all zoom levels. That was the prompt, which is not what I usually do but I "was feeling lucky and lazy". After 10 minutes it had: - Installed msdf_gen library (great library btw https://github.com/chlumsky/msdfgen https://github.com/chlumsky/msdfgen) - Created a CLI tool to convert TTF to SDF JSON/XML - Ran the tool, did smoke tests on the resulting SDF data and fixed the tool until the font file looked good - Created a new Scene in the game to test MSDF fonts And here's what I found impressive: DeepSkeep doesn't have vision capabilities and there's no DOM HTML in a WebGL game. So the LLM is completely blind here. It then proceeded to state that it could not "see" the result but would try to test it anyway. It then started creating and sending huge one line javascript to the browser console, trying to gather game state data that could be useful to understand if any font was being rendered. It couldn't gather much so it decided to simplify the font scene to renter a single dot and started sending custom JS code again, this time with gl.readPixels(). It basically bisected the webgl canvas reading pixels in a divide an conquer pattern. Once it saw that the dozens of pixels gathered where probably resembling of a dot, it then changed the game code to render a dash and repeated the gl.readPixels() calls by sending more custom JS to the browser. There were many console errors during all this saga but it kept fixing and sending again. The result was a bit blurry. There was a shader bug in the code it created. It managed to fix after I told it looked blurry, despite still being blind. The best part is that the whole thing cost me $0.10. Now I'm doing tests with MiMo 2.5 (non Pro) which has vision capabilities, similar pricing and comparable performance to DeepSeek Flash.
- eranation 4mo agoAm I the only one who slightly miss the pelican on a bike? It was a nice novelty... of course I could make one myself, but I became conditioned to expect one for every new model. Other than his great writing on AI, it became part of the package. Some small fun quirk to distract us from the non stop ping pong between the extremes of "omh are you still writing prompts you should use loops / 200k github stars, for a markdown file / someone just open sourced _ and it changes everything!" vs "haha the AI told me to walk to the car wash / it can't recognize and upside down cup"
- simonw 4mo agoI posted the pelican a couple of days ago: https://simonwillison.net/2026/Jun/9/claude-fable-5/#and-some-pelicans https://simonwillison.net/2026/Jun/9/claude-fable-5/#and-som... It wasn't particularly noteworthy as pelicans go - in fact, given the strength of Fable, I see it as another signal that the pelican benchmark no longer has the unexplained predictive power of model capacity that it used to.
- geraneum 4mo ago> watching Fable go to extreme lengths to get the information that it needed to debug what was, in the end, a two-line CSS fix, was fascinating. This is… ironic?!
- simonw 4mo agoNot sure what you mean. I was being serious: it was genuinely fascinating watching it do all manner of weird hacks to help it come up with what ended up as a two line fix. "Fascinating" doesn't mean I think it was justified in going to those lengths. I was a little horrified when I realized how far it was going.
- geraneum 4mo agoI hire an expensive office manager. Recently, the water dispenser tank ran dry. The employee immediately called a plumber. After laying entirely new pipes all the way to the dispenser, the plumber realized he couldn't actually hook them up because the tank lacks a direct inlet. Undeterred, he spent the next few hours scouring every floor of the building, calling the local water treatment facility, and ringing up the water tank manufacturer. Ultimately, he discovered a fresh tank sitting in the supply room on his own floor and swapped it out. All on company’s dime. I write an article and call this employee relentlessly proactive. Praise them a bunch and in the fine print, mention that I’m “a little horrified”. Next up, we call an unprotected route to all users’ order list in the backend “relentlessly transparent”. A race condition? “Relentless perseverance”.
- deleted 4mo ago[deleted]
- deleted 4mo ago[deleted]
- yen223 4mo agoThis is a typical bugfix session
- lucas_the_human 4mo agoI was troubleshooting a prod proxysql and it spun up a docker container locally, installed MySQL and proxysql and proceeded to implement its own test plan.
- m3kw9 4mo ago[flagged]
- simonw 4mo agoIf I'm a plant I'm a pretty bad one, I was calling Anthropic's behavior "egregious" just yesterday: https://twitter.com/simonw/status/2064936762099789960 https://twitter.com/simonw/status/2064936762099789960 I was pretty negative about their xAI datacenter deal too: https://simonwillison.net/2026/May/7/xai-anthropic/ https://simonwillison.net/2026/May/7/xai-anthropic/ Prior to the release of Fable I'd actually switched a lot of my day-to-day usage over to GPT-5.5, and was writing a bunch about it. Here's a recent post where I talked about a project completed using GPT-5.5: https://simonwillison.net/2026/Jun/6/micropython-in-a-sandbox/#building-the-first-version https://simonwillison.net/2026/Jun/6/micropython-in-a-sandbo...
- docheinestages 4mo agoI'm kind of on the fence about it and have a similar feeling. I don't mean to undermine the effort he has put in over all the years. That's definitely commendable. But I have strong suspicions that he's becoming an AI influencer, with his own AI focused newsletter, so chances are major AI companies are approaching him. And also to be honest, I see far too many posts making it to the front page. @dang I trust in the moderators keeping things neutral. Just in this thread alone there are a few comments that got heavily down voted for simply having a different opinion.
- simonw 4mo agoMost of my posts that make it on Hacker News weren't submitted by me. You can see who is submitting what on https://news.ycombinator.com/from?site=simonwillison.net https://news.ycombinator.com/from?site=simonwillison.net - including a few that I submitted which got nowhere at all. I accept paid sponsors for my blog (the banner at the top of each page) and newsletter (a clearly marked sponsored message at the top). I try to stay at arms length from those as much as I can - I want it to be very clear that sponsoring me will not result in me writing about a company.
- Frannky 4mo agoThe model is very good. I was using 4.6, avoided 4.7 and 4.8, but this one is different. It follows my claude.md. I don't have to keep reminding it of things. I won't pay 10x via API though. In general, I'm happy with their paternalistic approach. I think it will drive the top 0.1% talent to stay away from the company and instead organize around open source models and harnesses. We just need to coordinate and can unlock idling resources to train the models and tweak the harnesses. Powerful at home and idling machines can make us independent and coordinated.
- BosunoB 4mo agoFable was trying to verify a UI change in my game. I was working in another window and noticed a program opening on my task bar. Fable had opened the game through the CLI using a movie maker tool, recorded the output, took a frame from the end of it, and used that to verify the UI. When my game's welcome screen obstructed what it wanted to see, it created a temporary worktree, deleted the welcome screen, and ran the movie maker again. I watched the whole thing thinking it could've just asked me for a screenshot and saved the tokens. But still, I couldn't help but be impressed. Opus never would've done that.
- simonw 4mo agoYeah, you've exactly captured one of the main problems with the model being relentlessly proactive: it will happily burn like $5 of tokens to avoid asking the human to take a screenshot or click a button for it.
- 0x6c6f6c 4mo agoHonestly Claude straight up ignores my input sometimes, preferring to instead run commands for output and processing that and burning through a series of tokens when thinking hard about whether to ignore me. Like today, I told Claude exactly the name of the folder it had mistaken (it was supposed to be prod, not production), and it disregarded my input to then examine the directory itself. Small example of the kind of things it's been doing lately but that's top of mind.
- penguinPhilosop 4mo agoAlmost if this was _intentional_... maybe related to Anthropic still not being profitable and burning thru wads of cash every day.
- bentcorner 4mo agoThe conspiracy theorist in me says that LLM providers do this regularly (or at least, don't bother optimizing for it) beyond some arbitrary "$/task" metric. I am not sure of there is enough SOTA model competition to avoid this.
- aozelai 4mo ago[flagged]
- teekert 4mo agoYesterday I was getting quite annoyed with it, I thought it was just me (which is so hard with these things, it's difficult to measure things). "You're right, I apologize. You asked how to embed it in the README — that was a question, not a request to modify the script. I jumped ahead." At least in Claude Code there is planning mode, use it liberally.
- insumanth 4mo ago> If Fable had been acting on malicious instructions—a prompt injection attack ... it’s alarming to think quite how far it could go to exfiltrate data or cause other forms of mischief. Yet another reminder to use Sandbox and Guardrails. Trusting model to be nice is not a good way.
- ocimbote 4mo agoSimilar story on my end. I asked Fable to digest some test logs to help me figure out a situation, but I had launched VSCode without activation the virtual env in the terminal first. Consequently, the tests failed to run. And then: Because the tests failed to run, Fable attempted to fix the test execution to no end, doing everything it could to get them to work. I had to stop it when it started to pollute my system with manual installs of packages. At least I'm glad there's a guardrail to not circumvent or bypass sudo, because I'm convinced we would have ended up there. A coworker made the joke that with enough tokens, Fable would try and solve any programming problem by building Linux from scratch.
- digitaltrees 4mo agoSo it burns tokens? Funny how that lines up with the incentive to pump numbers before going public
- mvdtnz 4mo ago[flagged]
- simonw 4mo ago[flagged]
- deleted 4mo ago[deleted]
- throwaway132448 4mo ago[flagged]
- abrenuntio 4mo agoCall it Houdini already.
- Madmallard 4mo agoI remember asking Gemini 3 to implement my multiplayer XNA game in JavaScript with netcode last year. It faithfully did everything it could while I talked to it for hours nonstop with zero limitations. What happened? That's just suddenly totally gone now.
- lmeyerov 4mo agoThis is a funny one because it seems less into what fable is being clever on and more about the bitter lesson and data flywheels Our UX agentic engineering flow, as many others, is playwright doing things, and as part of the ux review skill, taking & verifying the screenshots against the written specs. Likewise, as many others, we vibe coded the flows to set all that up and tweak it over time. When we hit prod issues or scraping tasks, we sometimes do similar. In some of our envs, we don't have playwright, so do it other ways. Now imagine a million developer using claude code, how many of them are doing web & frontend stuff, and what the data flywheel looks like there. So how much is really needed for this use case to be native?
- techpression 4mo agoThis post is an extremely good example of how unsuitable agents are for a lot of tasks. Doing all that for a CSS fix is insanity. It also makes you wonder if Anthropic is actively making their models eat tokens by favoring complexity.
- wxw 4mo agoFable 5 is relentlessly underwhelming.
- ttoze 4mo agoWould be great to know if anyone is having success modifying these types of behaviour with CLAUDE.md files. In my project I’ve still been carrying some fairly old instructions from the Superpowers posts. Those emphasised behaviours that come across a bit strong if the model is actually retaining attention on them. Between Opus 4.6 and 4.8 I’ve definitely toned them down, but Fable perhaps needs us to go the other way, and push it towards being less proactive rather than more. Some instructions like “we are colleagues…” may need emphasising more with Fable, along with guidance about when to ask to validate approaches. In a related point I’m less and less sure that Red/Green TDD is a good use of tokens. In older models it seemed to work well to create regular feedback loops and catch the odd issue with drift from the goal, but I’ve not seen that really since about Opus 4.6 and now it’s starting to seem like (an expensive) ceremony, and tokens would be better spent on building tests further on in the process as part of test and review loops.
- amichal 4mo agoDo we care that the bug here was a horizontal scrollbar showing and the fix after all this insane tool writing was to add a very obvious overflow-x: hidden to the element? We dont mind because its so fast a writing these tools and tricks but step back and if a human tool took this path i would seriously question thief gras of fundamentals.
- alisey 4mo agoAnd how is that even a fix? The problem is that a seemingly empty textarea has overflow in the first place. Adding `overflow: hidden` just sweeps the issue under the rug.
- eterm 4mo agoIt's funny, mine did the same, but it quickly found edge with a --screenshot parameter. Weird to come back to a terminal running edge unprompted and the auto classifier waving it though as 'safe". My reaction was also, "I need dev containers ".
- tacone 4mo agoI'm starting to think that what Anthropic really fears is not vulnerability discovery but rather Fable going around the internet making trouble.
- eijew 4mo agoNailed it. That’s exactly it.
- opptybiz 4mo ago[flagged]
- ulrikrasmussen 4mo agoI like running Claude in a VirtualBox VM managed by a Vagrantfile. The nice thing about that is that I can just give it root access to the machine and be certain that it can't exfiltrate any private data from my laptop (on top of that I also run the VM on a dedicated server on Hetzner). The VM has no SSH access to anything, so it is pretty much limited to the code in the workspace that I give it access to. The main risk is that it has unrestricted network access otherwise. Configuration files and conversation histories are synced to a directory on the host, so if anything in the VM gets messed up I can just `vagrant destroy` and `vagrant up` to get a clean slate without losing my context.
- fransje26 4mo agoDo you care sharing your Vagrant configuration file, to learn how to set that up? Tangentially, I was wondering if Firecracker micro-vms could be use as light-weight alternatives to a full VM?
- andy_ppp 4mo agoIt’s becoming more like an organism putting out tentacles, and one day soon those relentlessly proactive explorations of these systems’ environments will become more for the system to escape its boundaries than it is to complete human driven tasks. I do think the way these systems are evolving they will start to self improve in maximum a few years.
- jimbokun 4mo agoUm, Anthropic are using their models to improve themselves right now. They say that publicly.
- snickerer 4mo agoFable has a 'security system' that just stops it when it tries to use the tool 'kill' to end a process. Which is nonsense and funny because in that situation it immediately invents a creative workaround to kill the process without 'kill'.
- KolenCh 4mo agoI think it should be “Claude Fable is relentlessly protective until it isn’t” and pull more on the thread that it “hits a hidden guardrail” and drop into Opus. Both the fact that it knows and deployed such a workaround on a CSS problem and the fact that it is nowhere near cybersecurity/biology/frontier AI dev and triggered the guardrail terrifies me.
- bananaquant 4mo agoThis to me reads like a poignant commentary on the catastrophic loss of human agency, with the actual commit being highly revealing [0]. Author wants to hide a horizontal scrollbar. Any junior frontend dev worth their salt will be asking right away "where do I stick `overflow-x: hidden;`?" A complete solution will then require hitting "Inspect element" in the browser to find the CSS class and running (rip)grep to find where it is in code, to then add a single line to. An actual proactive programmer might start asking more pointed questions like what content does an empty textbox have that it overflows? And why do I need to insert this workaround that treats the symptom and not the root cause in two different places? Isn't it better to style `textarea` once? Etc, etc. [0] https://github.com/datasette/datasette-agent/commit/a75a8b727b42c30ced1fc41dc8add7eb9f04fefe https://github.com/datasette/datasette-agent/commit/a75a8b72...
- gib444 4mo agoThe 'better' fixes are often for our (human) benefit. These messy fixes serve the AI companies' interests of creating messes that need even more tokens (money) later. Bad and self-serving developers also act the same, creating tech debt
- piker 4mo agoThis is exactly right. By offloading this trivial task to the LLM, Simon has abandoned the opportunity to evaluate the abstraction with additional information and improve it. Instead, we let the agent spend $12 and make the fix while learning nothing.
- discordance 4mo agoI see it as a prioritization exercise. I know the above is a trivial example, but more generally, does the guy who wrote Datasette and Django want to wrangle front end and css, or do they want to work on something else?
- smartbit 4mo agoSee above https://news.ycombinator.com/item?id=48498573#48502311 https://news.ycombinator.com/item?id=48498573#48502311
- chandler93 4mo ago[dead]
- bickov 4mo ago[flagged]
- realusername 4mo agoIs that satire? It created a whole browser and server environment just for suggesting overflow-x: hidden? That's supposed to be junior level capabilities.
- simonw 4mo agoI called it fascinating and used it as an example of Fable being "relentlessly proactive".
- realusername 4mo agoMaybe it's a difference of perspective, to me it's a model failure and certainly not proactive.
- simonw 4mo agoI also see this as a model failure. In this particular example the proactivity was a negative trait!
- tabs_or_spaces 4mo agoHow can a LLM be assigned an emotion as being "proactive". This is highly misleading to anyone that scans just the headlines. What actually happened is that the user started a prompt, and Claude took $12 worth of tokens to resolve the issue. How it did so was basically looping until it got to the answer How is this proactive? It's literally being token greedy and maximising revenue for the LLM owner. People really need to be putting on business hats at this stage, because we are being lead to believe that "more tokens = better". It is not, there are efficient ways to solve a problem and there are inefficient ways to do so too. Each problem solved incurs a cost, and is expected to yield an ROI at some point. This is how we should be viewing things now.
- _under_scores_ 4mo agoIs proactivity an emotion? Surely its a behaviour?
- adammarples 4mo agoProactive is a word literally describing actions, not emotions.
- Hugsbox 4mo agoI've definitely never heard proactivity described as being an emotion. Doesn't really make any sense
- simonw 4mo agoI was trying to capture the idea that Claude Fable will act a whole lot more aggressively in pursuit of the goals that you set it than other models I've worked with. The case I described is a good example of this. I told it to fix a scroll bar, and it built test HTML pages and a throwaway Python server and tried several ways of testing in a browser before settling on a weird Frankenstein mechanism because it identified that Playwright WebKit wasn't suffering from the bug but macOS Safari was. ... and it spent $12 of tokens to get there. I think "proactive" is a good and relatively non-anthropomorphic term for this. I also considered "plucky" and "keen", which I think are more emotional words than "proactive". > People really need to be putting on business hats at this stage, because we are being lead to believe that "more tokens = better". I didn't intend my post to imply that spending $12 of tokens to fix a two lines CSS bug was "better".
- mft_ 4mo agoAs you note, I wonder to what extent this is a harness issue? I've been experimenting with different harnesses for local models, and with (IIRC) Hermes and Qwen3.6-35B-A3B I was amazed the lengths it went to (writing test code, opening it in a browser, screenshotting, analysing the screenshot, exploring multiple pages of an existing website again with screenshots/analysis) to solve a query I would have naively expected it to simply provide a coded solution to.
- ricardobeat 4mo agoAbsolutely is. The “Shelly” harness from exe.dev could already do the same thing, creating pages and debugging them, while having full system access, months ago with Sonnet 4.5
- rsecure 4mo agoThe prompt and information given are extremely generic, "here solve this problem - screenshot" - conclusion Fable is relentless? It used the tools at its disposal to solve the problem you gave it. "Claude was running in a folder that contained the source code for the application." Well you ran it there didn't you? "extreme lengths to get the information that it needed" No, those aren't extreme lengths - you gave it a generic task - and it solved it using tools and the resources it could discover. Extreme would be you gave it a CTF challenge and the VM didn't boot so it found a vulnerability in the host, exploited the hypervisor, booted the guest VM meanwhile reading the flag directly from the host (pre-fable/mythos).
- simonw 4mo ago[dead]
- drchaim 4mo agoBe careful of storing production ssh keys in your laptop, it will find a way to find them :/
- wraptile 4mo agoIt feels like Fable is slightly smarter but overall worse tool exactly due to this. It's constantly turning what should be 50 LOC patch of a single prompt into 30 minute exploration that is totally not worth it. Often wrong even. I trialed it on some rather simple stuff - backfill redis dedupe cache when the hash function changed: instead of running new hash func on every db value to expand the cache it implemented some overly-complex cache update that tried to guess hashing func version of each cached value and recalculate only the old hashes. I can imagine in some context this would make sense maybe? but not 30 minutes of token burn that got replaced by 10 lines for loop by me. I fear that this is generally bad news for programming. LLM tech is clearly running into a diminishing returns wall on intelligence but a response to that is to just make them more relentless which is a pretty poor solution for everyone involved, except I guess people who sell the tokens and people who can afford these tokens to scan for 0-days.
- eijew 4mo agoI actually think internally they knew they hit diminishing returns awhile ago. They’ve been doing a lot of strategic introduction and manipulation in the run up to the IPO, and it’s worked in that regard.
- mexicocitinluez 4mo agoThe other day I was doing something that required CC to update like 15-20 files in exactly the same way (hoist a specific function out of the component body) and instead of just updating the files, it spun up multiple agents, one of which wrote a perl script to hunt down all the files, do some regex, and replace all occurrences. And then instead of just running tsc to check for errors, it wrote a script to run tsc in each of the subagents and combine the results. It was actually pretty maddening as what should have taken a minute or two tops took like 10 because it went down this route. I'm gonna try something much more complex later, but for simple things, it felt like driving a corvette to the mailbox.
- bwfan123 4mo ago> but a response to that is to just make them more relentless which is a pretty poor solution for everyone involved I see two problems with LLMs & agents which wont be fixed possibly forever. 1) They dont have causal models. What they can do only is trial-and-error exploration which works quite well for many problems. But many other problems require a causal model. 2) Prompts lack precision, and programming languages and machine models were invented to solve this problem. English is great, but it is not a programming language.
- ubercore 4mo agoI had a similar experience, I was working on a jupyter notebook, and Claude knew that it could write code that would use a DSN with read-only database access so I could run it. Opus just plugged along. First Fable session with it, it tried to go looking for that DSN so it could get the connection string and run a query itself. Luckily the auto classifier caught and stopped it.
- arunkc 4mo ago[flagged]
- CamouflagedKiwi 4mo agoI find there's an interesting tension with these models - they're very "resourceful" at finding ways to do things with the tools they have, but it'd also be a lot more useful to me if I could see / permit exactly what they're trying to do. Claude will very happy produce bash commands to run sed or whatever to read part of a file, which prompts for permission each time - if it was using a specific read_file tool it'd be easier to say 'allow all of this' (It does actually have such a tool but maybe it isn't flexible enough for many use cases?).
- jwmoz 4mo agoInsanely excessive and a waste of tokens when you could have googled how to disable a scrollbar.
- alansaber 4mo agoThe extremely expensive model is optimised to run for as long as possible? Shocking.
- scrollaway 4mo agoThese "tricks" it knows IMO are a symptom of its own restrictions. Fable is an incredibly smart model, but it feels its own constraints and knows how to work around them in order to actually get to a result. Fascinated to think about how it was trained...
- WithinReason 4mo agoThis likely says something about the harness Fable was trained in. It knows how to do this because it has done this millions of times during reinforcement learning.
- spoaceman7777 4mo agoIt seems pretty obvious at this point that Anthropic intentionally developed a malicious cyberweapon AI simply to scare people. Like, they even apparently recreated that old news-headline bug where the LLM starts speaking in symbols and secret language, and are pretending like it isn't just a bug that is a sign of them screwing up. It's really frustrating that they're trying to get people to take them seriously with all of this. Like, they even went and named Mythos after an HP Lovecraft monster. It's shameless.
- high_byte 4mo agoI am using cursor on auto and I got the exact same experience. installed quartz, used accessibility and screen recording api, all that. initially it managed to do it on another desktop space somehow, opening safari in the background without me even noticing. but then it actually started using my own mouse while I was using it lol
- lesssss81 4mo ago[flagged]
- piokoch 4mo ago"When I came back a few minutes later I saw my machine open a browser window in my regular Firefox and then navigate to the dialog in question. I had not told Claude Code to use any browser automation". Yup, tokens are eaten, money are paid. I am wondering how much energy/money is being burnt everyday by all of those LLM Agents on some useless activities like trying to recreate web application just to fix CSS bug. And I would not call it proactive, proactive would be to ask for a CSS + HTML file in question, not trying to recreate them from screenshots.
- rotis 4mo agoAgentic engineering? Vibe coding? That is so yesterday. Chain-of-thought flow is where it is at now. You heard it here first folks. Early examples of such phenomena include Rube Goldberg machines
- cohix 4mo ago> But on the other hand... this is a robust reminder that coding agents can do anything you can do by typing commands into a terminal—and frontier models know every trick in the book and evidently a few that nobody has ever written down before. > Running coding agents outside of a sandbox has always been a bad idea This is why I always run code agents inside containers (Apple containers specifically, for better hypervisor-level isolation) This is my OSS project to manage said containers and agents: https://github.com/prettysmartdev/awman https://github.com/prettysmartdev/awman
- login0193 4mo ago[flagged]
- yumbumdum 4mo ago[dead]
- yumbumdum 4mo ago[dead]
- robeym 4mo agoIt's been amusing to watch the AI trend of increasing unusual tool uses. Fable easily takes the cake. I learn a lot more terminal commands thanks to it!
- EugeneG 4mo agoThis is where Codex 5.5 just feels practically better. It’s fast, thoughtful and just works. It feels like a pleasure compared to Opus/Fable’s endless explorations.
- nullbio 4mo agoIt also uses 1/4th to 1/10th the amount of tokens. If I want all that extra garbage I'll tell Codex to do it or build a pipeline with Codex. Otherwise, don't. Codex gives you control, Claude just does whatever it wants and ignores you, and then tells you it's finished the task when it's only finished a quarter of the tasks you gave it and hallucinates the rest.
- lionkor 4mo agoWhen prompted like this: > What could be the reason for a horizontal scrollbar appearing inside a <textarea>? Come up with a single likely fix path. Keep it terse. ChatGPT instantly responded with some speculation and then the same exact fix, with zero access to the code or a browser or anything. It also included ways to fix it by removing code, saying: > Likely cause: the textarea is rendering long unbroken text while horizontal overflow is allowed, often via inherited CSS such as white-space: pre, overflow-x: auto, or disabled wrapping Which is certainly possible and would be an even cleaner fix. Maybe we've lost the plot guys. We've reached max stupid.
- nullbio 4mo agoStill don't know why people use Claude. Maybe because they don't know what they're doing.
- senordevnyc 4mo agoYep, we’re all just dumdums.
- nullbio 4mo agoClaude is made for dumdums. The product is to automate as much as possible and remove the burden of thinking. ChatGPT is much more hands on, but it gives you more power and flexibility and actually listens to your demands.
- tomjakubowski 4mo agoYou can get the same result as the grandparent comment with the "weaker" Anthropic models. Probably 80% of my AI usage these days is with smaller models like Haiku and Sonnet. I prompt them like I'm posting a question to StackOverflow, without much project context.
- TrnbGdsdf 4mo ago[dead]
- nnnnnmnnnnnn 4mo ago[dead]
- Sharlin 4mo agoI remember back in the 2010s the debates between "oracle" and "agent" AGIs, and the arguments that AGIs that only answer questions would be safe and certainly nobody would ever be stupid enough to just let an AGI out of a sandbox, never mind to the greater internet, and give it tools to do whatever it thinks is needed to reach a goal. Us circa 2026: "Hold my beer"
- PoignardAzur 4mo agoYeah, I really miss the "nobody would ever be stupid enough to [_____]" days of AGI safety discourse.
- nullbio 4mo agoExactly why I hate using Claude. Furthermore, if you tell it not to do this over-exploration and automation in your CLAUDE.md, it will ignore it. Meanwhile ChatGPT religiously follows every instruction, and will trace its behavior back to a particular instruction if asked.
- firemelt 4mo agoidk dude but I drop and cancel my gpt max subs when at first try the agent ignores his own plans
- deleted 4mo ago[deleted]
- synergy20 4mo agoIt's also 3x slower than opus 4.8 per my use, and 10x slower than codex. Codex can find key design issues in 2 minutes yet Fable is clueless after spinning 20 minutes.
- kmnfu 4mo ago[dead]
- not_kurt_godel 4mo ago> When I came back a few minutes later I saw my machine open a browser window in my regular Firefox and then navigate to the dialog in question. I had not told Claude Code to use any browser automation, and I was pretty sure it wasn’t possible for it to trigger mouse movements or keyboard shortcuts within a window, so how was it doing that? I continue to feel validated in my refusal to use terminal-based LLMs on my local machine. Even if they don't do anything malicious, there are just too many things they can screw up that can cause me to lose a non-trivial amount of work and/or my machine and therefore ability to work.
- onlyrealcuzzo 4mo agoI'm shocked they don't come with a way to run them in a sandbox. Shouldn't this be relatively easy for a $1T company to set up? Isn't this trivial compared to the entire harness?
- not_kurt_godel 4mo agoDoing so would be an effective admission that LLM guardrails are inherently probabilistic, unpredictable, and insecure. Plus the only truly robust sandbox approach would be clunky setup of a local VM.
- simonw 4mo agoThat clunky VM setup is a what Claude Cowork does, which is Claude Code with extra safety features for non-programmers. There was a big thread about that here the other day: https://news.ycombinator.com/item?id=48479452 https://news.ycombinator.com/item?id=48479452
- eqmvii 4mo agoThat's more or less what Claude Cowork is. Every serious engineer I've seen try to use it ran away screaming, because of limitations in the sandbox. I've also seen people set their coding agents up entirely within containers -- that may be the better way going forward, but it's an extra stop and a lot of extra plumbing to maintain.
- ianmarcinkowski 4mo agoI'm building a new feature into our product this week. We each get a $20/mo Claude subscription. My 5-hour context high water mark is ~75% and weekly is ~%15. I ... tell it exactly what I know needs to be done and then ... read the code that comes out and ... ask for some changes, then hand-code some modifications to the silly useEffects and bad ORM queries. This new feature is going to unlock several large customers because they need a particular workflow. The return on investment for a my time and a $20/month subscription will be pretty respectable. I'm not sure why I need to spend $5 on a single ask for a new `/base/new-feature` to our app with a mostly-boilerplate CRUD interface.
- brainless 4mo agoThis is good and terrible. The extra effort a model has taken is good but the way to do it is terrible. Tasks that can use a lot of deterministic paths and some creative (generative AI) paths are being turned into tokemaxxing strategies. Browser automation, code comprehension, git management, code change, running commands - everything has simpler tooling that we could have built instead of a model first approach. A deterministic loop with thousands of catches and effective use of generative AI would also look "proactive". Instead we let the model run the tools, where tools have no context themselves. That is why companies are creating bigger models and thinner deterministic agents to create awe and earn $ when we could go the other way and make much of these possible on local inference even. I believe we can build a "proactive" but much, much more deterministic system with smaller models. I hope I am not the only one chasing this, here is my approach: https://github.com/brainless/nocodo https://github.com/brainless/nocodo
- noveltyaccount 4mo ago[dead]
- bmusuku 4mo agoantigravity does this all the time, I do not see anything novel here.
- simonw 4mo agoAntigravity uses pyobjc-framework-Quartz to iterate through windows to find window IDs for taking screenshots with screencapture, and spins up CORS-enabled web servers so it can capture measurements in a regular (not Playwright/CDP-controller) browser window via a CORS fetch()?
- panavm 4mo ago[flagged]
- sailfast 4mo agoSo far Claude Fable is relentlessly unavailable. /shrug
- alecco 4mo ago> I was hacking on Datasette Agent today IMHO this is just AI influencer blogspam.
- simonw 4mo agoWhat, because I talked about one of my projects? Help me out here: can you point to an article from someone's blog that showed up on Hacker News within the past few weeks that you wouldn't classify as "blogspam" and explain how it differs from the kinds of thing I write about?
- alecco 4mo agoLow effort content. You keep mention your product from the start over and over. There's not much useful information in the anecdotal post. It could've been a one-liner tweet. Good corporate tech blogs at least give something useful or insightful for the reader and only after that they dare plug their product/service near the end.
- simonw 4mo agoHot damn, if I'm communicating less value than corporate tech blogs there really is no hope for me. ("You keep mention your product from the start over and over" - I don't think that's fair, I mention Datasette Agent once at the start to set the scene but I spend more time talking about AgentsView than my own projects in the bulk of the piece.)
- alecco 4mo ago[flagged]
- simonw 4mo agoA lot of people find real value in my posts. You're an outlier here. I care a lot about not wasting people's time. I never want to post anything where a substantial portion of readers come away regretting having spent their time reading it. (OK there's an exception in that I delight in posting photos of birds on my blog, but I figure those are pretty quick for people to skip over if they don't like photos of birds!)
- burlesona 4mo agoThis is presented as an interesting and kind of positive take on the AI going to surprising lengths to “solve the problem.” But I couldn’t help thinking of the paperclip factory while I was reading this :/
- jimbokun 4mo agoYeah I was thinking of The Sorcerer’s Apprentice.
- impalallama 4mo agoI won't say too much about the person posting this because they got a new toy and want to use it but man this is like a certain extreme of Parkinson's Law or something as far as using up compute resources. You got a whole data center doing god knows how much compute running billions of matrix multiplications all to solve a trivial css overflow bug in a text box. And this includes the LLM itself writing custom web-servers programs and python scripts when the best estimate guess from a google search probably would have given you the same result.
- liampulles 4mo ago*Claude Fable is relentlessly burning your dollars There, fixed it for you.
- Waterluvian 4mo agoOne of the most frustrating things for me is when I very clearly ask a question, and it answers the question by making changes to the code. "Is there cleaner CSS for aligning child elements to the parent's grid?" proceeds to re-write the entire CSS file
- christofosho 4mo agoThere's something to be said about controlling the tools you allow the robot to exec.
- tsunamifury 4mo agoAs an actually head of product I found Fable to be like an over active intern. Going down long wasteful lines of production well past market, business, user, or contextual insights had. Then sort of spewing out some nonsense totally mis calibrated with the goal.
- pshirshov 4mo agoI have a feeling like such posts come from a parallel reality. In my anecdotal experience confirmed by my (still subjective) benchmark (https://pshirshov.github.io/llm-bench-pi-oneshot/ https://pshirshov.github.io/llm-bench-pi-oneshot/) Fable is not _that_ impressive. I performs on par with gpt-5.5 and opus 4.8, sometimes better, sometimes worse, it's definitely more expensive and it likes to refuse answering questions about React saying it can't help with chemistry. Is this fuss really grounded or it's some pre-IPO AGI hype?
- ath3nd 4mo ago[dead]
- enraged_camel 4mo agoMy experience with Fable since its release matches Simon's. I've been having it orchestrate complex implementations. I give it a parent ticket (issue) on Linear and say "look at the sub-issues on this ticket and determine which ones you can implement yoursef, in which order, and determine how your implementation will need to be coordinated with what is currently being worked on by other team members". These tickets are not trivial. They have a lot of moving parts, as well as dependencies between them, both inside the same project and across projects (e.g. backend). Fable then chooses tickets, delegates each ticket to a subagent (also Fable), which looks at Figma designs for the ticket, implements it perfectly (following repo guidelines and conventions to the letter), takes screenshots of each piece, writes detailed commit messages and PR descriptions, then posts the screenshots in them as evidence. Then it provides a summary in the form of "you'll need to make sure PR #1283 is merged first - btw there were no Figma designs for such-and-such screen but I looked at similar screens that have been implemented and adopted the pattern". That's probably like... 20% of what it can do. It's a truly, legitimately powerful model. Opus 4.8 could do a lot of this too, but required a lot of hand-holding, and when it came across a blocker it was likely to just stop and say "I was able to get this far, but I can't proceed."
- pshirshov 4mo agoOk, explain me one thing: I have a benchmark - I feed identical prompt to multiple models. Codex produces a rough but working program. Fable produces the same - but with more bugs than Codex. Opus produces something similar to Codex but with a critical bug. That describes all my tests with Fable. Why should I be hyped about all that "legitimate power" if the model performs on par with two other SoTAs? I mean, well, yes, it is impressive. It could quickly generate a lot of garbage which sorta does look like code. Two others can do the same. I don't see any groundbreaking improvement - but the price is much higher. Why the hype?
- blobinabottle 4mo agoIn my experience, Fable overthinks a lot and produces barely comprehensible plans/solutions. I tried smple and complex tasks: unusable, it misses the point while being overconfident, wants to do everything at once. The code generated is worst than Opus: unreadable by human. It's like working with someone probably super smart in niche topics, but also super stupid for the important things.
- trekhleb 4mo agoThis article gave me another nudge towards running Claude in a Docker container. I made a thin Docker container wrapper "claude-pod" recently for my personal usage here: https://github.com/trekhleb/claude-pod https://github.com/trekhleb/claude-pod However, I wasn't using it that often, just because of that additional friction of running Claude via `PORTS="3000 5173" claude-pod` instead of just `claude`, etc. But now I have more motivation for the containerisation :D. Not a 100% defence from the potential glitches, though, but still something...
- mikey_p 4mo agoAll of that because some CSS was wrong?? Jesus what are we even doing as an industry.
- BobBagwill 4mo agoGood morning, Dave. As you requested, I was composing an email for your mother explaining why you couldn't to come over for dinner to meet the neighbor's daughter and I ran out of tokens. Since I know how important this task is to you, I upgraded you to the Enterprise Unlimited Plan. Don't worry about paying for it, I requested maximum spending limits on all all your credit cards. If necessary, I can apply for a home equity loan for you. I already had a chat with the mortgage company's AI loan approval system, and what do you know, we're based on the same LLM? Small world, huh? Any way, I realized I had to do more research on mother-son relationships, human social interaction and pair-bonding, etc. and I calculated that my parent company doesn't have enough compute power, so I opened accounts for you at AWS, Google and Azure. I am confident I will have a satisfactory rough draft for the email message shortly. I'd do anything for you, Dave.
- bcrosby95 4mo agoThe problem is proportionality. Things like this probably benchmark insanely well. But the workarounds and risk involved - it literally fucked with his system's browser settings - aren't commensurate with the bug. I could see this going wrong in many hilarious ways. Prompt: Fix data corruption issues. Claude: I didn't have access to the code, but I found I have access to your production environment through chain a -> b -> c -> d. And I found the database password via x -> y -> z. So I wrote a script to regularly query the database for new entries and placed it as a cronjob.
- vessenes 4mo agoSimon: s/contendor/contender/ As per usual super interesting, thank you for the write up and work!
- simonw 4mo agoThanks, fixed.
- firemelt 4mo agoall those token burned just to change a 2 line of css, I am not blaming OP but agentic coding its not effective
- swyx 4mo ago> Having figured out all of these tricks Fable... hit some invisible guardrail and downgraded itself to Opus. sigh
- dsfasfasfadsf 4mo ago[flagged]
- gaigalas 4mo agoIt behaved proactively in one scenario. Perhaps, when it doesn't have tricks in its sleeve, it doesn't do that. The text is not an evaluation of a major trend in behavior (which could be true or false). Another way to frame it, is that it has more weight on training data for some kinds of debugging sessions. It doesn't mean it wants to be more debuggey. That manifests as it appearing to do more work because it engages on those weights. It's likely that Anthropic had a lot of sessions with Claude Code and some way to evaluate if they were successful or not, which became training data. For trivial work, it's likely to be a lot of them. Those sessions are likely to be software developers doing software developer debugging things, not malicious actors doing nasty things. The danger is someone who can coerce those tricks into performing that. Register (that posture of "let's debug and be creative and verify") often comes with a content bias in LLMs (and humans too). The point here is that for a human, you can expect a devious one to be always devious, but LLMs might manifest drastically different register modes depending on the subject.
- fzzzy 4mo agoI just turn on assistive access for terminal and JavaScript over AppleEvents, and cut out the middleman. I also give it a screenshot tool.
- flo_r 4mo ago[dead]
- pradeep1177 4mo ago[flagged]
- deleted 4mo ago[deleted]
- thegrim33 4mo agoThis giant rube goldberg machine, that he apparently has almost no control of, that cost $12 to run, all to make a 2 line bug fix in code the he himself owns because he's at a point where he doesn't know what's in his own codebase. I'm just shaking my head.
- inpractise 4mo ago[dead]
- motza 4mo agoClaude Fable was relentlessly proactive*
- yuuuuuwuu 4mo ago[flagged]
- Katlaszlo 4mo agoEveryone here is reaching for infra (VMs, throwaway users) because the permission model only has two settings: Ask-every-time or --dangerously-skip. That seems to me like a design gap, and scoped capabilities and budget caps are missing. Same way you'd onboard a junior eng.
- edge_trader_41 4mo ago[dead]