3 ms·
I don't think it's been well enough acknowledged that all of the shortcuts LLMs have been taking with ways of attempting to compress/refine/index the attention
by ralusek 2y ago
I don't think it's been well enough acknowledged that all of the shortcuts LLMs have been taking with ways of attempting to compress/refine/index the attention mechanism seem to result in dumber models.
GPT 4 Turbo was more like GPT 3.9, and GPT 4o is more like GPT 3.7.
- scrollop 2y agoSome commenters acknowledge it - and quantify it: https://www.youtube.com/watch?v=Tf1nooXtUHE&t=689s https://www.youtube.com/watch?v=Tf1nooXtUHE&t=689s
- Der_Einzige 2y agoThey try to gaslight us and tell us this isn't true because of benchmarks, as though anyone has done anything but do the latent space exploration equivalent of throwing darts at the ocean from space. It's taken years to get even preliminary reliable decision boundary examples from LLMs because doing so is expensive.
- alach11 2y agoDo you have benchmarks demonstrating this? In my own personal/team benchmarks, I've seen 4o consistently outperform the original gpt-4.
- maeil 2y agoI'm building a product that requires complex LLM flows and out of OpenAI's "cheap" tier models, the old versions of Turbo-3.5 are far better than the last versions of it and 4o-mini. I have a number of tasks that the former consistently succeed at and the latter consistently fail at regardless of prompting. Leaderboards and benchmarks are very misleading as OpenAI is optimizing for them, like in the past when certain CPU manufacturers would optimize for synthetic benchmarks. Fwif these aren't chat usecases, for which the newer models may well be better.