5 ms·
Interesting... Most benchmarks show this model as being worse than o3-mini-high and sonnet3.7. What difference are you seeing from these models that makes it b
by dinobones 2y ago
Interesting... Most benchmarks show this model as being worse than o3-mini-high and sonnet3.7.
What difference are you seeing from these models that makes it better?
I say this as someone considering shelling out $200 for ChatGPT pro for this.
- ldjkfkdsjnv 2y agoI regularly push 100k+ tokens into it. So most of my code base/large portions. I use the Repo Prompt product to construct the code prompts. It finds bugs and solutions at a rate that is far better than others. I also speak into the prompt to describe my problem, and find spoken language is interpreted very well. I also frequently download all the source code of libraries I am debugging, and when running into issues, pass that code in along with my own broken code. Its very good
- jbellis 2y agoIf you're in the habit of breaking down problems to Sonnet-sized pieces you won't see a benefit. The win is that o1pro lets you stop breaking down one level up from what you're used to. It may also have a larger usable context window, not totally sure about that.
- logankeenan 2y ago> lets you stop breaking down one level up from what you're used to. Can you provide an example of what you mean by this? I provide very verbose prompts where I know what needs to be done and just let AI “do” the work. I’m curious how this is different?
- jbellis 2y agoPartly it means you can tell it to do X and it will figure out that implies Y and Z without you having to spell it out And partly it can actually execute more at the same time without starting to make mistakes
- raylad 2y agoSonnet 3.7 and O1 Pro both have 200K context windows. But O1 Pro has a 100K output window, and Sonnet 3.7 has a 128K output window. Point for Sonnet. I routinely put about 100K + of context into Sonnet 3.7 in the form of source code, and in the Extended mode, given the right prompt, it will output perhaps 20 large source files before having to make a "continue" request (for example if it's asked to convert a web app from templates to React). I'm curious whether O1 Pro actually exceeds Sonnet 3.7 in Extended mode for coding or not. Looking forward to seeing some benchmarks.
- consumer451 2y agoI am very curious how 3.7 and o1 pro perform in this regard: > We evaluate 12 popular LLMs that claim to support contexts of at least 128K tokens. While they perform well in short contexts (<1K), performance degrades significantly as context length increases. At 32K, for instance, 10 models drop below 50% of their strong short-length baselines. Even GPT-4o, one of the top-performing exceptions, experiences a reduction from an almost-perfect baseline of 99.3% to 69.7%. https://arxiv.org/abs/2502.05167 https://arxiv.org/abs/2502.05167
- futopy 2y agoAnyone ever tries to restructure a 10K text? For example, structure a 45min - 1hr interview transcript in an organized way without losing any detailed numbers / facts / supporting evidence. I find that none of OpenAI's model is capable of this task. Models are trying to summarize and omitting details. I think such task does not require much intelligence, but surprisingly OpenAI's "large" context model cannot make it.
- qeternity 2y ago"Usable" is the key word here. Not all context is created equal. Have a look at the RULER benchmark for a bit more detail.
- Tiberium 2y agoThere actually were almost no benchmarks for o1 pro before because it wasn't on the API. o1 pro is a different model from o1 (yes, even o1 with high reasoning).