6 ms·
Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
- blfr 3mo ago> Ploy’s agent builds and edits real marketing websites. It plans a page, reads the codebase, writes components, generates imagery, screenshots its own work, and decides when it’s done. That job description sets a very high bar for a model, and we test every frontier release against it. For the four months Opus held the default slot (first Opus 4.7, then 4.8), nothing we tested beat it. Well, unlike OP I haven't run a rigorous test, but I still would expect Fable to be significantly better at building marketing websites than Opus. It sure is way better at building decks.
- greenavocado 3mo ago4.7 is very autistic in terms of following directions so I find OPs claims plausible
- arikrahman 3mo agoVery descriptive there heh
- aeonfox 3mo agoGame recognises game
- make_it_sure 3mo agogpt 5.6 is so much better ar design than fable
- cute_boi 3mo agoCan you please provide evidence. It shouldn't be hard to give side by side comparison. I have not found any task where gpt is better than fable.
- hankbond 3mo ago[flagged]
- deleted 3mo ago[deleted]
- kristianp 3mo ago> Numbers like that buy a model a real migration effort. Such a silly choice of words. I wish the human directing the LLM writing the article put some effort into rewriting the worst examples of LLM style. > But it did extremely well, and the promise was immediate and specific: builds finishing in less than half the wall-clock time, at 27% lower cost, scoring at or above our incumbent on completed work. The way the LLMs write (Claude perhaps?) With short phrases separated by colons, commas or full stops, is so poor and frustrating. There some good insights behind this article, so it's worth reading, for example below, but it isn't easy to read. > Earlier GPT models cached implicitly on partial prefix matches, which gave decent hit rates for free. GPT-5.6 dropped partial-prefix matching:
- w4yai 3mo agoCan we get over the detective work about if the text was written by LLM or not in 2026 already ? This is a lost cause, and we could instead focus on substance over syntax.
- conjectures 3mo agoWhat substance? That they consume a newer model from the same vendor?
- spicyusername 3mo agoThe problem is that the second you suspect something is written by AI, its a pretty good signal that 50-80% of the text is empty of meaning. Maybe that will change, but LLMs are terrible and inefficient writers. Only so much time in the day, its a quick signal to not waste anymore of it.
- lewistaariq 3mo agoCorrect. AI == Credibility hit and it's increasing as more humans get used to feeling they are AI slop consumers, not worth the time for genuine human engagement. Human engagement costs are increasing. Amazing to read/watch.
- 3mo ago
- estebarb 3mo agoBut what users prefer? Given this is for marketing, which results produce more conversions? From the examples shown, personally I strongly preferred Claude Opus in all cases.
- arikrahman 3mo agoMigrating my workflow to Reasonix with cache hits on Deepseek make requests practically free, and that's on unsubsidized American providers.
- gunapologist99 3mo agoSorry, what did that have to do with the article?
- bel8 3mo agoThey also migrated and that also made the workflow cheaper. It has everything to do with the article.
- gunapologist99 3mo agoReasonix, the harness, only works with DeepSeek. Neither are mentioned in the article at all, which was about a migration from Claude Opus to GPT 5.6. Maybe a bit of DeepSeek astroturfing going on?
- arikrahman 3mo agoI wish people were astroturfing it
- desterothx 3mo agoWhat's your config, how does it compare to pi
- arikrahman 3mo agoHere's my config for agents: https://codeberg.org/arik/agents https://codeberg.org/arik/agents It's a lot better since it's optimized for the model and cache hits, unlike other frontends that try to be more general.
- przemarzec 3mo ago[flagged]
- bob1029 3mo ago> we’ve made GPT 5.6 Sol the default model powering every Ploy workspace I would consider Luna for parts of the workload that touch actual tools. It is surprisingly capable and it runs fast. Sol is great at talking to the human and orchestration of agent calls, but it's just too expensive to use everywhere. You can get 5 Luna runs for the cost of 1 Sol run. Statistically speaking, going from one to five samples is a pretty big deal.
- Tadpole9181 3mo agoThe problem I always run into with subagents is that they are isolated. This is a double-edged sword, as it keeps context down and lets them "focus", but it often means they must do their own research to continue to do work given to them, which eats uncached tokens. So depending on how heavily agents are used on what tasks, it's entirely possible that you get worse work for more cost.
- bob1029 3mo ago> they are isolated This is a feature if your goal is to obtain many samples. Independence is critical. This makes it easier to accurately model the uncertainty of a decision.
- taspeotis 3mo agoI feel Claude Code has added (and removed?) a feature that forks a subagent from the parent context, so it’s still isolated but it’s more of a continuation of what you were doing in one narrow direction and then it dies. Rather than a blank slate with a prompt of what to do.
- auspiv 3mo agoStatistically speaking if each part of the Luna run has a 90% chance of being correct, 5 of those is 0.9^5 = 0.59 = 59%. Or one Sol run being maybe 95% correct? Exact numbers vary of course. But then again having sol verify at end may be cheaper.
- bob1029 3mo ago
- luciana1u 3mo ago[flagged]
- CurbStomper 3mo ago[dead]
- thiagoperes 3mo agoWe run a lot of varied, tiny, simple workflows that were previously running on 5.4-nano and mini. We transitioned them to 5.6 and noticed exactly this range of improvement across the board. In a few cases, we had improvements in classification. I think a lot of people miss that for many companies, a model upgrade like this is basically a one liner. Even if you have an amazing model router architecture (which we do for our golden flows), it’s just not worth it. Not to mention reliability and so on
- febed 3mo agoWhat SDK are you using? Or is it custom?
- desterothx 3mo agothe article is literally about the model upgrade not being a one liner
- madaxe_again 3mo agoThe first thing I used Sol for was to assess 5.6 on our workflows - previously, it was 5.5 for everything, as the quality on simpler models was just not good enough. We’re doing a mix of text and image analysis to extract explicit and implicit structured data from a steaming pile. They do work pretty much as advertised. The bulk of our workload is now going via terra, which has cut our cost in half by itself, as well as improved response times by 50% - luna I am using as a backstop for opencv hits, and it is good enough, and so cheap as to almost be free - but very limited - and very fast. Sol only gave marginal improvements over terra for our workload. I’ve also gotta say I’m impressed as to how well Sol ultra carried out the assessment itself - it made sound recommendations, and gave me a nice big dossier of “you should look at these outputs yourself and compare and consider” along with raw and digested data, and cpm for queries. Anyway. Spent nothing beyond my pro sub, let Sol gnaw on it for a few hours, and my cost basis just dropped 50% and throughput improved by 100%. Win.
- implexa_founder 3mo ago[flagged]
- redfather918 3mo agoThe cost reduction is impressive, but I think consistency matters even more for production agents. I'd be interested to know whether prompt engineering or tool-calling workflows had to change significantly.
- narmiouh 3mo agoHe actually talks about that in the article. Including how they had to rearchitect tool calling using nulls as well as limitations of prompt caching
- jing09928 3mo ago[flagged]
- rjnz199 3mo ago[flagged]
- langs 3mo ago[dead]
- modgate 3mo ago[flagged]
- desktopentree 3mo agoI catch a lot of issues on the technical writing side.
- lcampbell 3mo ago> The fix that worked is a schema transform at the provider boundary. For OpenAI-family models only, we rewrite every optional property to be required but nullable, using anyOf: [T, null], which gives the model an explicit way to say “not using this.” I admit, I've only used a bastardized form of MCP, but this smells... wrong? It's not clear to me why the Typescript type definitions would have any influence on (what I presume is) JSONSchema being sent from the agent to the inference backend as part of the completion request. The MCP specification (which the OpenAI backend might not use, I don't know) has an explicit field to signify "optional" parameters in the JSONSchema; my read on this is there's a bug somewhere between the Typescript layer(??) and the generated tool description which is actually sent to the inference backend. It's possible the inference backend has changed from "generate valid tool responses" to "generate valid tool responses according to the JSON schema [where no parameters are optional]" but it's impossible to tell without seeing the actual requests sent to the inference backend (which I didn't see in TFA).
- dannyw 3mo agoModern frontier models, including Fable/Opus and 5.6, are often very loose with tool calls, and often don’t follow your schema precisely. For example, see this post for Claude models hallucinating properties for an edit/replace tool call in Pi: https://lucumr.pocoo.org/about/ https://lucumr.pocoo.org/about/ I suspect some part of this comes from the noticed intelligence degradation when you do constrained decoding. Yes, you’re guaranteed schema validation, but you lose a lot of intelligence. It’s fine if you just want a classifier, a summary, a prompt enhancement, etc; but I’d be careful in agentic loops. Harnesses like Claude Code do a lot of preprocessing, repairing, cleaning, etc; as the blog post shows. You usually don’t see it. In practice it’s easier and better to just make your harness “looser” and work better with the model (they’re coming out every month or two anyway, each with their own idiosyncrasies) than to assume and force perfect correctness. Welcome to vibe applied AI ;)
- JLO64 3mo agoI believe the post you were trying to link was this one: https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/ https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/
- ianberdin 3mo agoWe at Playcode.io - a company similar to Ploy are still using Opus 4.6. "Why?" you might ask. Because GPT 5.6 Sol, while fast and pleasant to use, is essentially the same model as 5.5 wrapped in new marketing packaging, just to avoid losing ground to Anthropic. In practice, it's the same quality: it generates the same garbage, tons of code, and can never solve even a single complex task. We simply don't trust it to write code for clients that they'll end up throwing away anyway. "Then why not Opus 4.8?" you might ask. Well, because Opus 4.8 and 4.7 are just another lie, a price hike with no actual quality improvement. That's why at Playcode, we give our clients the best possible quality/price - which is Opus 4.6. Regardless of what people write in articles like this.
- ianberdin 3mo agoEveryone probably has the same question: what about Fable? Fable 5 is sick. It is simply the best model in the world out of everything we have ever tried. It's absolutely fantastic. It solves almost any task from start to finish, the way it should be done — no errors, perfect code. It's a miracle. If there's any way to make it a little more affordable, that would be incredible. As for GPT-5.6 Sol — it doesn't even come close. I honestly don't understand why people even try to compare them. It feels like Sam's attempt to hold onto his audience with those endless daily limit resets. A clever trick, nothing more.
- porker 3mo ago> Fable 5 is sick. [It] solves almost any task from start to finish, the way it should be done — no errors, perfect code. It's a miracle. > As for GPT-5.6 Sol — it doesn't even come close. I honestly don't understand why people even try to compare them. What kind of problems are you working on? I like Fable but when planning work on a complex C codebase it's making more mistakes than 5.6 Sol xhigh for me. In what scenarios is Fable giving you "no errors, perfect code"?
- ianberdin 3mo agoI have a large monorepo that includes about 15 TypeScript services and many Rust services. Everything is well-documented and organized, with standardized and structured custom code. When an issue arises, I often test the systems by providing a minimal prompt, like: "this user, this is their email, this isn't working, figure it out in production." I send this to both Opus and ChatGPT, but it doesn't help. I've set up Agents.md and Quote.md identically, with the same access and linkers, so the Harness is consistent. ChatGPT rarely succeeds. If the task is complex and requires a multi-step process to identify the true cause, ChatGPT usually stops after a few initial ideas and wrongly claims it has found the solution. - For simple tasks, like identifying a missing item in a to-do list, ChatGPT performs well. - However, for issues like memory leaks or file system corruption, it struggles. On the other hand, Opus 4.8 always finds the solution, albeit slowly. I can rely on it without worrying about whether it will succeed. It just gets the job done. Recently, Fable 5 has emerged, which resolves issues without needing any prompts. It operates even faster than Opus. When I ask ChatGPT or Opus to create a new feature: - ChatGPT often produces superficial results, ignoring existing code and building unnecessary independent code. - Interestingly, the outcome from ChatGPT appears functional, but it's usually incorrect, focusing on a superficial "aha!" moment. Opus, however, plans thoroughly, executes, and cleans up, ensuring everything works correctly. If needed, I can provide more realistic examples, though it's challenging due to the monorepo's size and complexity, with hundreds of thousands of lines of code.
- htlemur_bobby 3mo agoI found Claude to be better for the first prototype. It was more likely to come up with something fast. But it kept lying and claiming it did world class work and it was just hardcoding response by the end. I found GPT never lied to me.
- znnajdla 3mo agoMy experience mirrors this: services like OpenRouter that promise “failover” are pretty much useless except for sandbox testing because models in production are not really interchangeable. Any production harness doing serious agentic work in production is dependent on more model-specific quirks than you would expect. And even if another model works without errors, performance and efficiency is a whole different story. Even the system prompt can and should be tuned to a model’s preferred speaking style, for example <xml tags> for Claude-like models because they were trained on it, while other models do better with other delimiters. Think of the whole harness, prompt, and model as one system, not really with modular parts that can be swapped out if you care about optimal performance.
- marcyb5st 3mo agoI believe part of the LLMOps (I don't like the term, but it is what it is) should be building a failover plan with proper testing that check tools trajectories and such. If you have these then you can sort the good enough models from cheaper to more expensive and have the failover you mentioned. I saw people bulding a mapping of model->{{prompts}, {tools descriptions}, ...}, but that, to me, it feels extreme. I believe it is the model that needs to adapt to your prompts after a certain point. Models that fail to do so won't get our api requests as they will be out of the failoever roster.
- hamandcheese 3mo agoOpenRouter doesn't fail over to a different model, it fails over a different provider of the same model.
- jdw64 3mo agoPersonally, could you share the code sometime later? The GPT code looks decent to me. Or even just the prompt would be fine.
- SwtCyber 3mo agoIts ironic that under an article with a ton of deep infrastructure insights half the comments are crying about the "forced writing style". What does it matter if claude helped the author clean up the text when inside is a ready-to-use blueprint on how to save 30% of the api budget and fix empty file reads?
- jtrn 3mo agoOne of many reasons I would assume is that people just hate anything related to AI, so they latch on to anything negative they can say. Another reason is that they mean what they say... That they really, really hate the style of writing, enough to fixate on that. Personally, I think the people whining about the style are silly. Maybe because I'm terrible at grammar and spelling, but I always just focus on the message, not the delivery. I just care about the concept, facts, the argument, and so forth. The actual grammar and spelling are just trees, while the forest is the point. Edit: just an infobit: The reason my text isn't full of errors is due to the awesomeness of the dictation and a custom hotkey I have created on my computer, which uses a local LLM to spellcheck any text I have selected and replaces it with the corrected one. Nothing has improved my quality of life and writing more than these two tools!
- oakst 3mo agoOn the infobit, I've been trying to build in a similar completely local LLM cleanup step into Epilude. Willing to share anything you've found especially useful in producing good spellchecks/cleanups in your local setup?
- throwa356262 3mo agoAs of today, Ploy’s agent runs on GPT-5.6 Sol, the flagship tier of the model family OpenAI released this morning. Wait a moment, did they make the switch based on half a days of playing with Sol? Are these companies ran by teenagers?
- fxwin 3mo agoI would expect they have production based datasets they evaluate new models against.
- sekai 3mo agoAI psychosis is still at all time high
- avianlyric 3mo agoThere’s every possibility they got some amount of early access to evaluate GPT5.6, precisely so they could write an article like this.
- brryant 3mo agohah - we actually skew staff, senior staff. We have been testing GPT 5.6 for about a week as a preview model through a YC relationship, providing them feedback on the model. Our evals run in github CI and we can run them all in about 15 minutes against our eval bench of 115+ web design and marketing related jobs that ploy.ai specializes in. then after we toggled it on (through a posthog feature flag) we actively monitored for failures. I came from running Webflow, which powers > 1% of the internet so trying my best to relay all of that knowledge to ploy to power more % of the internet!
- agumonkey 3mo agois it common to measure LLMs with landing pages like this ? it seems a bit too simple
- yusufnb 3mo agoWhat does a site build cost, in real $$ with Opus vs Sol? Could you share a ballpark?
- jon_steadioai 3mo ago[flagged]
- brcmthrowaway 3mo agoAI wrapper companies still exist?
- bluelightning2k 3mo agoMust admit, for this particular case I don't see the appeal in using a wrapper. Why would I not just use Codex directly? The we wrote a bunch of prompts argument is kind of meh. That sort of thing has not only diminishing value with subsequent model releases but I actually believe will turn negative. The model will know better by default. For example, initially giving the model some advice on code best practice was helpful. But now it's unhelpful because the model already knows best.
- brcmthrowaway 3mo agoIdk man, they raised money lol
- tosh 3mo agoMost agent harnesses that have not been designed from scratch for current models are over-engineered. Chances are whatever was needed to make earlier models perform well now either is no longer helping much or actively hurts performance (worse results, slower, uses more tokens …).
- bluelightning2k 3mo agoAgreed! Example: for large Eloqua/Marketo/HubSpot emails we would previously make a planner which delegated the sections to their own call. GPT5.6 can do the whole thing. The planner is unhelpful. My suggestion: feature flag your complex implementations so you can rapidly contrast with and without it. (Or a formal eval suite if you have one). If you prefer the simpler path, delete the old path. Note: the challenge is making things compatible with these tools. Obviously generating html directly has been simple for ages. (Source: mopsy.ai)
- cws_ai_buddy 3mo ago[flagged]