5 ms·
What wall? Not a week has gone by in recent years without an LLM breaking new benchmarks. There is little evidence to suggest it will all come to a halt in 2025
by briga 2y ago
What wall? Not a week has gone by in recent years without an LLM breaking new benchmarks. There is little evidence to suggest it will all come to a halt in 2025.
- jrm4 2y agoSure, but "benchmarks" here seems roughly as useful as "benchmarks" for GPUs or CPUs, which don't much translate to what the makers of GPT need, which is 'money making use cases.'
- deleted 2y ago[deleted]
- peepeepoopoo98 2y agoO3 has demonstrated that OpenAI needs 1,000,000% more inference time compute to score 50% higher on benchmarks. If O3-High costs about $350k an hour to operate, that would mean making O4 score 50% higher would cost $3.5B (!!!) an hour. That scaling wall.
- Kuinox 2y agoWait a few month and they will have a distilled model with the same performance and 1% of the run cost.
- peepeepoopoo98 2y ago100X efficiency improvement (doubtful) still means that costs grow 200X faster than benchmark performance.
- deleted 2y ago[deleted]
- achierius 2y agoEven assuming that past rates of inference cost scaling hold up, we would only expect a 2 OoM decrease after about a year or so. And 1% of 3.5b is still a very large number.
- popcorncowboy 2y agoAnd to your point "past performance is not indicative of future results". The extrapolate to infinity approach is the mindfever of this field.
- norir 2y agoI used to run a lot of monte carlo simulations where the error is proportional to the inverse square root. There was a huge advantage of running for an hour vs a few minutes, but you hit the diminishing returns depressingly quickly. It would not surprise me at all if llms end up having similar scaling properties.
- riku_iki 2y agoAnd I suspect o3 is something like monte carlo: generates tons of CoTs, with most of them are junk, but some hit the answer.
- exhaze 2y agoSounds plausible given I’ve recently observed a ton of research papers in the space that in some way or another incorporate MCTS
- LegionMammal978 2y agoYeah, any situation you need O(n^2) runtime to obtain n bits of output (or bits of accuracy, in the Monre Carlo case) is pure pain. At every point, it's still within your means to double the amount of output (by running it 3x longer than you have so far), but it gradually becomes more and more painful, instead of there being a single point where you can call it off.
- oceanplexian 2y agoI’m convinced they’re getting good at gaming the benchmarks since 4 has deteriorated via ChatGPT, in fact I’ve used 4-0125 and 4-1106 via the API and find them far superior to o1 and o1-mini at coding problems. GPT4 is an amazing tool but the true capabilities are being hidden from the public and/or intentionally neutered.
- CSMastermind 2y ago> I’ve used 4-0125 and 4-1106 via the API and find them far superior to o1 and o1-mini at coding problems Just chiming in to say you're not alone. This has been my experience as well. The o# line of models just don't do well at coding, regardless of what the benchmarks say.
- didibus 2y agoAll the benchmarks provide substantial scaffolding and specification details, and that's if they are zero-shot at all, which they often are not. In reality, nobody wants to spend as much time providing so much details or examples just to get the AI to write the correct function, when that same time and effort you'd have used to write it yourself. Also, those benchmarks often run the model K times on the same question, and if any one of them is correct, they say it passed. That could mean if you re-ran the model 8 times, it might come up with the right answer only once. But now you have to waste your time checking if it is right or not. I want to ask: "Write a function to count unique numbers in a list" and get the correct answer the first time. What you need to ask: """ Write a Python function that takes a list of integers as input and returns the count of numbers that appear exactly once in the list. The function should: - Accept a single parameter: a list of integers - Count elements that appear exactly once - Return an integer representing the count - Handle empty lists and return 0 - Handle lists with duplicates correctly Please provide a complete implementation. """ And run it 8 times and if you're lucky it'll get it correct zero-shot. Edit: I'm not even aware of a Pass@1, zero-shot, and without detailed prompting (natural prompting) benchmark. If anyone knows one let me know.
- famouswaffles 2y agoNot really. o3-low compute still stomps the benchmarks and isn't anywhere that expensive and o3-mini seems better than o1 while being cheaper. Combine that with the fact that LLM inference has reduced orders of magnitudes in cost the last few years and hampering over the inference costs of a new release seems a bit silly.
- riku_iki 2y agoIf you are talking about ARC benchmark, then o3-low doesn't look that special if you take into account there are plenty of finetuned models with much smaller resources achieved 40-50% results on private set (not semi-private like o3-low).
- famouswaffles 2y ago- I'm not just talking about ARC. On frontier Math, we have 2 scores, one with pass@1 and another with consensus vote with 64 samples. Both scores are much better than previous Sota. - Also apparently, ARC wasn't a special fine-tune but rather some of the training set in the corpus for pre-training.
- riku_iki 2y ago> On frontier Math that result is not verifiable, not reproducable, unknown if it was leaked and how it was measured. Its kinda hype science. > ARC wasn't a special fine-tune but rather some of the training set in the corpus for pre-training. post says: Note on "tuned": OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details. So, I guess we don't know.
- famouswaffles 2y ago>that result is not verifiable, not reproducable, unknown if it was leaked and how it was measured. Its kinda hype science. It will be verifiable when the model is released. Open ai haven't released any benchmark scores that were shown falsified later so unless you have an actual reason to believe they're outright lying then it's not something to take seriously. Frontier Math is a private benchmark with its highest tier of difficulty Terrence Tao says: “These are extremely challenging. I think that in the near term basically the only way to solve them, short of having a real domain expert in the area, is by a combination of a semi-expert like a graduate student in a related field, maybe paired with some combination of a modern AI and lots of other algebra packages…” Unless you have a reason to believe answers were leaked then again, not interested in baseless speculation.