5 ms·
For pro mode the agents worked independently and only when they all finished did a new agent take a look at everything to merge the work into a single response.
by changoplatanero 3mo ago
For pro mode the agents worked independently and only when they all finished did a new agent take a look at everything to merge the work into a single response. The new thing involves subagents that have been trained to cooperatively pursue a task and are allowed to communicate with each other along the way.
- dools 3mo agoI tried a pro model out the other day and thought there must have been a bug in Pi’s cost calculations. But no, it’s absolutely fucking insane. Wasn’t even any better at the task.
- bombcar 3mo agoI really suspect that the models are basically the same below, it’s all in the prompt. The way I use them, surgically, they seem to perform about the same. Fable certainly hasn’t blow my socks off.
- giancarlostoro 3mo ago> Fable certainly hasn’t blow my socks off. Same. I suspect they'll get better at taking in terrible prompts over time though... Maybe that's what Fable does better, reminds me of Sora 2, it would take my crappy prompt and expound upon it. I told it once to generate a video of someone working at some company that changed its name, but the old name had historic relevance, it referred to the new company name without me telling it to, by virtue of me wanting a video of TODAY with a 90s icon.
- X-Istence 3mo agoWhere fable has blown me away is converting entire code bases and or refactoring across many different segments. It’s far more careful than opus and puts far more effort into testing and validating by default. Switching back to opus at work was a downgrade. Similar requests felt more clunky and needed far more hand holding.
- bombcar 3mo agoSome of it feels boiled down to "opus works better when told not to be dumb, fable's prompt tells it not to be dumb." If they know much of what the tool is used for, they can customize prompts to "do that usage right" even if the user doesn't know exactly how to ask for it.
- anjel 3mo ago> Fable certainly hasn’t blow my socks off. Same. Its not so much perf increase as cost increase justified by ambiguous perf increase.
- bdcravens 3mo agoThis is where I think you see the distinction between two classes of LLM users: 1. Managers: those who generally know what needs to be done, and want it done faster, so they provide a lot of instructions and context (where many developers fall) 2. Executives: those who vaguely know the end goal, but are clueless about the process, and are willing to burn resources and cycles on a black box to get the result
- andai 3mo agoYeah, the bigger models shine when it comes to complexity (making the right decisions regarding choices with second-order effects), ambiguity (esp. common sense) and time horizon (agentic steps and context size). If your tasks are well defined and don't require a very large number of steps -- e.g. you're asking for small, clearly defined changes to the code -- you're fine with grok-4-fast. (Well, you would be fine if they hadn't killed it.) I work in both of these modes, and I find that the latter actually benefits from dumber models, because smaller models are faster. The work shifts from async to realtime/interactive. So you can stay alert, keep track of what they're doing and iterate, instead of alt-tabbing, getting a coffee, and then spending extra time resynchronizing your mental model later.
- thomasahle 3mo agoDo you have a source for this, or just rumors? The responses I get from pro don't feel like ensembles. They are often very one directional.
- changoplatanero 3mo agoThis can be because the summary model just picked the output from one of the sub agents.
- thomasahle 3mo agoYes, but then it should be easy to check of the result you see match the reasoning you see
- wahnfrieden 3mo agooops
- nl 3mo agoThe source is the GPT 5.5 System Card: > We generally treat GPT-5.5’s safety results as strong proxies for GPT-5.5 Pro, which is the same underlying model using a setting that makes use of parallel test time compute. As noted below, we separately evaluate GPT-5.5 Pro in certain cases because we judge that the setting could materially impact the relevant risks or appropriate safeguards posture. https://deploymentsafety.openai.com/gpt-5-5/model-data-and-training https://deploymentsafety.openai.com/gpt-5-5/model-data-and-t... There have been multiple podcasts with people from OpenAI which have confirmed this.
- cubefox 3mo ago> makes use of parallel test time compute Any idea what that means exactly? I vaguely remember that ChatGPT Pro was originally called "deep thought", just like Geminis "deep thought" feature (or "deep think"?), so it seems likely they are using the same approach.