3 ms·
is this working because agents are just another way of chunking context across models to keep them from drifting off topic?
by beardedwizard 2y ago
is this working because agents are just another way of chunking context across models to keep them from drifting off topic?
- spdustin 2y agoAgents "work" because (in a well-designed agentic system) their context isn't polluted by tokens that aren't meaningful to the work they're doing. Their attention mechanisms are less likely to drift around with each new token if the subject matter largely stays the same.
- beardedwizard 2y agook so exactly what I said then. im sure its just me (and not directed at your reply), but I find 'agent' to be a very obnoxious term that obfuscates what is really going on. It's not a magical being working on your behalf with agency, it's that there are inherent limitations to attention, and overcoming them usually involves having multiple models cooperate on the inputs and outputs. the number of agents is not necessarily related to the number of distinct jobs you need done, after all a single job could still perform best with multiple "agents" due to limitations in attention.
- mewpmewp2 2y agoMultiple models cooperating doesn't sound much better though. In the end it could be a same model or even a same single instance where it just gets triggered with different prompts in sequence. Response to first prompt decides 10 new tasks to be done in sequence with different prompts, running them then doing a final prompt with the results of those.
- beardedwizard 2y agoThat's a fair point, multiple models is a potential implementation detail for parallelization - I do think it's fair to assume you are "restarting" the model or clearing context by some means between runs otherwise I don't think it achieves the goal of groomed attention.
- mewpmewp2 2y agoThat gets really philosophical here. What is a single run actually? Because each token is generated one by one, with all the previous tokens as input. If we remove one token from the context in the past does it make it a new "run"? What if we have a summarization memory system, where in order to keep the context size small after certain length it will start to summarize/compress the beginning until the context size is good. Then the whole input and context is constantly changing/evolving. You could have 100s of different randomly selected instances generating a new token one by one to the (last input + token generated by another instance) as input.
- beardedwizard 2y agoI think it comes down to the goal you are trying to achieve: maximize use of available attention while minimizing drift. The experiment seems clear, and the outcome being tested is the resulting generation.