3 ms·
There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then
by legucy 2mo ago
There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab.
But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could multiply higher numbers. The difference is that with calculators we consciously theorized about what universal computation would require, then we built one as a step change. Despite it having low memory and slow speeds, the first one built was as theoretically universal as any computer we have today, in terms of the surface of computations it can perform.
With intelligence, it’s turned out to be less discontinuous, which I believe has convinced people that intelligence is a never ending exponential rather than an S curve approaching a horizontal asymptote. I suspect the LLMs we have today are the same kind of thing we will have in 5-10 years, but in 5-10 years we’ll consider them to be fully universal. At that point we’ll still have improvements in tokens per second and volume of context window, but not in capability per token.
- prideout 2mo agoBut aren't today's frontier models already "fully universal"? To use your Turing machine analogy, I think we're past the calculator stage.
- LarsDu88 2mo agoYou're take basically lines up with Francois Chollet: https://arxiv.org/abs/1911.01547 https://arxiv.org/abs/1911.01547 intelligence is more like polishing a ball smooth than growing the ball to infinity. For many tasks, it will be smooth enough.
- HPsquared 2mo agoAt a certain point the roughness of the ball reaches a size threshold where the imperfections are smaller than the wavelength of light, and the surface takes on a glassy smoothness. Intelligence has similar milestones, almost like phase changes, I think, where capabilities are reached. Maybe it's like a superposition of many small step functions.
- Manfrednotfunny 2mo agoBut in theory you can make an LLM A LOT faster than a human. You can also run massive amount of LLMs in parallel. There might be a limit to a normal LLM but not to theo everall system.
- tavavex 2mo agoHumans, however, are highly variable, which may produce really varied and interesting results if they work together. One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.
- boorang 2mo agoI think using different AGENTS.md can give the same model different perspectives on the same problem. For example a model with a well-tuned AGENTS.md by an expert mathematician approaching the same problem as the same model with a well-tuned AGENTS.md by an expert biologist can grind on the same problem from different perpectives. It's worth a shot at least, as a microservices architect I have a bias that we aren't networking these enough, a single main agent session orchestrating multiple subagents is different from multiple main agent sessions with their own subagents coordinating with each other.
- tavavex 2mo agoCrucially, does it make capabilities infinitely scalable? My comment just said that models may have a hard cap, and maybe doing specific setups like yours can make reaching it easier, but making the 'team' 10x larger after that optimal point may bring few to no improvements. Although I'm also not sure about just how much better models can really get with this technique. Ultimately you're still getting the same model with the same training data, which are the important parts. Asking it to pretend to be something feels like it would just put a color filter in front of the conclusion the model has already predicted, or maybe alter the path to the conclusion slightly or pick a less likely answer that it still could've provided normally.