2 ms·
In this context "thinking" was meant as an analogy to the supposed reflective "slow" mode of the cognitive process, vs the more "reflexive"/"fast" mode, not in
by sottol 2y ago
In this context "thinking" was meant as an analogy to the supposed reflective "slow" mode of the cognitive process, vs the more "reflexive"/"fast" mode, not in the sense of "thinking soul" or "thinking self-aware entity".
Concretely, what do you gain by giving a current-generation LLM more runtime? It's not trained/designed to do anything with it, more time = more tokens = more nonsense once past the "end" token/end of context. You could build an agent on top of the LLM that calls the LLM iteratively, but afaict this approach isn't strictly an improvement over the base LLM. Current architectures seem to be limited with how much improvement they can eke out of more runtime without a broader redesign or retraining.
Now with new architectures you might be able to do more with more runtime, but I'm not sold that allowing a current-gen LLM to execute code or do web-searches and re-feeding it the quetion + its output + new data is going to strictly return better results. Sometimes maybe yes, sometimes not.
So imo the current limitation is not the runtime allotted to LLMs but their fundamental (current-gen) design/training.