4 ms·
Do you have a source for this, or just rumors? The responses I get from pro don't feel like ensembles. They are often very one directional.
by thomasahle 3mo ago
Do you have a source for this, or just rumors?
The responses I get from pro don't feel like ensembles. They are often very one directional.
- changoplatanero 3mo agoThis can be because the summary model just picked the output from one of the sub agents.
- thomasahle 3mo agoYes, but then it should be easy to check of the result you see match the reasoning you see
- wahnfrieden 3mo agooops
- nl 3mo agoThe source is the GPT 5.5 System Card: > We generally treat GPT-5.5’s safety results as strong proxies for GPT-5.5 Pro, which is the same underlying model using a setting that makes use of parallel test time compute. As noted below, we separately evaluate GPT-5.5 Pro in certain cases because we judge that the setting could materially impact the relevant risks or appropriate safeguards posture. https://deploymentsafety.openai.com/gpt-5-5/model-data-and-training https://deploymentsafety.openai.com/gpt-5-5/model-data-and-t... There have been multiple podcasts with people from OpenAI which have confirmed this.
- cubefox 3mo ago> makes use of parallel test time compute Any idea what that means exactly? I vaguely remember that ChatGPT Pro was originally called "deep thought", just like Geminis "deep thought" feature (or "deep think"?), so it seems likely they are using the same approach.
- nl 3mo agoTheir methodology isn't published. Its widely accepted[1] that it runs the same query through the model in parallel and then has a model that either selects the best answer or synthesizes an answer from the multiple ones generated. I believe most people think it runs 6 sub-models, but I think that is based on the pricing. It's a pity that OpenAI doesn't publish details like this. [1]eg https://news.ycombinator.com/item?id=48799977 https://news.ycombinator.com/item?id=48799977
- dannyw 3mo agoBasically like passes@6 or passes@5 if you’re doing a benchmark, except for your real tasks. Pro is quite limited on the web UI I reckon. This approach can be highly effective for reasonably verifiable task, for example, write comprehensive unit tests pointing out a tricky bug, get multiple agents to swarm at it.
- nl 3mo agoIt's been very successful at frontier math tasks - a bunch of the Erdos questions have been solved by it - more than any other model. https://www.erdosproblems.com/ https://www.erdosproblems.com/
- cubefox 3mo ago> Basically like passes@6 or passes@5 if you’re doing a benchmark, except for your real tasks. It's unclear how they would do this when there is no signal that provides an objective ground truth.