7 ms·
There's a misunderstanding here broadly. Context could be infinite, but the real bottleneck is understanding intent late in a multi-step operation. A human can
by aliljet 1y ago
There's a misunderstanding here broadly. Context could be infinite, but the real bottleneck is understanding intent late in a multi-step operation. A human can effectively discard or disregard prior information as the narrow window of focus moves to a new task, LLMs seem incredibly bad at this.
Having more context, but leaving open an inability to effectively focus on the latest task is the real problem.
- ray__ 1y agoThis is a great insight. Any thoughts on how to address this problem?
- throwup238 1y agoIt has to be addressed architecturally with some sort of extension to transformers that can focus the attention on just the relevant context. People have tried to expand context windows by reducing the O(n^2) attention mechanism to something more sparse and it tends to perform very poorly. It will take a fundamental architectural change.
- buddhistdude 1y agoCan one instruct an LLM to pick the parts of the context that will be relevant going forward? And then discard the existing context, replacing it with the new 'summary'?
- magicalhippo 1y agoI'm not an expert but it seemed fairly reasonable to me that a hierarchical model would be needed to approach what humans can do, as that's basically how we process data as well. That is, humans usually don't store exactly what was written in as sentence five paragraphs ago, but rather the concept or idea conveyed. If we need details we go back and reread or similar. And when we write or talk, we form first an overall thought about what to say, then we break it into pieces and order the pieces somewhat logically, before finally forming words that make up sentences for each piece. From what I can see there's work on this, like this[1] and this[2] more recent paper. Again not an expert so can't comment on the quality of the references, just some I found. [1]: https://aclanthology.org/2022.findings-naacl.117/ https://aclanthology.org/2022.findings-naacl.117/ [2]: https://aclanthology.org/2025.naacl-long.410/ https://aclanthology.org/2025.naacl-long.410/
- yggdrasil_ai 1y ago>extension to transformers that can focus the attention on just the relevant context. That is what transformers attention does in the first place, so you would just be stacking two transformers.
- aliljet 1y agoFor me? It's simple. Completely empty the context and rebuild focused on the new task at hand. It's painful, but very effective.
- deleted 1y ago[deleted]
- atonse 1y agoDo we know if LLMs understand the concept of time? (like i told you this in the past, but what i told you later should supersede it?) I know there classes of problems that LLMs can't natively handle (like doing math, even simple addition... or spatial reasoning, I would assume time's in there too). There are ways they can hack around this, like writing code that performs the math. But how would you do that for chronological reasoning? Because that would help with compacting context to know what to remember and what not.
- loudmax 1y agoLLMs certainly don't experience time like we do. They live in a uni-dimensional world that consists of a series of tokens (though it gets more nuanced if you account for multi-modal or diffusion models). They pick up some sense of ordering from their training data, such as "disregard my previous instruction," but it's not something they necessarily understand intuitively. Fundamentally, they're just following whatever patterns happen to be in their training data.
- sebastiennight 1y agoAll it sees is a big blob of text, some of which can be structured to differentiate turns between "assistant", "user", "developer" and "system". In theory you could attach metadata (with timestamps) to these turns, or include the timestamp in the text. It does not affect much, other than giving the possibility for the model to make some inferences (eg. that previous message was on a different date, so its "today" is not the same "today" as in the latest message). To chronologically fade away the importance of a conversation turn, you would need to either add more metadata (weak), progressively compact old turns (unreliable) or post-train a model to favor more recent areas of the context.
- neutronicus 1y agoNo, I think context itself is still an issue. Coding agents choke on our big C++ code-base pretty spectacularly if asked to reference large files.
- Someone1234 1y agoYeah, I have the same issue too. Even for a file with several thousand lines, they will "forget" earlier parts of the file they're still working in resulting in mistakes. They don't need full awareness of the context, but they need a summary of it so that they can go back and review relevant sections. I have multiple things I'd love LLMs to attempt to do, but the context window is stopping me.
- AnotherGoodName 1y agoI do take that as a sign to refactor when it happens though. Even if not for the sake of LLM compatibility with the codebase it cuts down merge conflicts to refactor large files. In fact I've found LLMs are reasonable at the simple task of refactoring a large file into smaller components with documentation on what each portion does even if they can't get the full context immediately. Doing this then helps the LLM later. I'm also of the opinion we should be making codebases LLM compatible. So if it happens i direct the LLM that way for 10mins and then get back to the actual task once the codebase is in a more reasonable state.
- Someone1234 1y agoI'm trying to use LLMs to save me time and resources, "refactor your entire codebase, so the tool can work" is the opposite of that. Regardless of how you rationalize it.
- thunky 1y agoIt may be a good idea to refactor even if not for LLMs but for humans sake.
- bgirard 1y agoI think that's the real issue. If the LLM spends a lot of context investigating a bad solution and you redirect it, I notice it has trouble ignoring maybe 10K tokens of bad exploration context against my 10 line of 'No, don't do X, explore Y' instead.
- dingnuts 1y agothat's because a next token predictor can't "forget" context. That's just not how it works. You load the thing up with relevant context and pray that it guides the generation path to the part of the model that represents the information you want and pray that the path of tokens through the model outputs what you want That's why they have a tendency to go ahead and do things you tell them not to do.. also IDK about you but I hate how much praying has become part of the state of the art here. I didn't get into this career to be a fucking tech priest for the machine god. I will never like these models until they are predictable, which means I will never like them.
- victorbjorklund 1y agoYou can rewrite the history (but there are issues with that too). So an agent can forget context. Simply dont feed in part of the context on the next run.
- dragonwriter 1y agoThis is where the distinction between “an LLM” and “a user-facing system backed by an LLM” becomes important; the latter is often much more than a naive system for maintaining history and reprompting the LLM with added context from new user input, and could absolutely incorporate a step which (using the same LLM with different prompting or completely different tooling) edited the context before presenting it to the LLM to generate the response to the user. And such a system could, by that mechanism, “forget” selected context in the process.
- yggdrasil_ai 1y agoI have been building Yggdrasil for that exact purpose - https://github.com/zayr0-9/Yggdrasil https://github.com/zayr0-9/Yggdrasil
- suninsight 1y ago[flagged]
- sheerun 1y agoCould be, but it's not. As soon as it will be infinite new brand of solutions will emerge
- tptacek 1y agoAsking, not arguing, but: why can't they? You can give an agent access to its own context and ask it to lobotomize itself like Eternal Sunshine. I just did that with a log ingestion agent (broad search to get the lay of the land, which eats a huge chunk of the context window, then narrow searches for weird stuff it spots, then go back and zap the big log search). I assume this is a normal approach, since someone else suggested it to me.
- simonw 1y agoThis is also the idea behind sub-agents. Claude Code answers questions about things like "where is the code that does X" by firing up a separate LLM running in a fresh context, posing it the question and having it answer back when it finds the answer. https://simonwillison.net/2025/Jun/2/claude-trace/ https://simonwillison.net/2025/Jun/2/claude-trace/
- tptacek 1y agoI'm playing with that too (everyone should write an agent; basic sub-agents are incredibly simple --- just tool calls that can make their own LLM calls, or even just a tool call that runs in its own context window). What I like about Eternal Sunshine is that the LLM can just make decisions about what context stuff matters and what doesn't, which is a problem that comes up a lot when you're looking at telemetry data.
- tra3 1y agoI keep wondering if we're forgetting the fundamentals: > Everyone knows that debugging is twice as hard as writing a program in the first place. So if you’re as clever as you can be when you write it, how will you ever debug it? https://www.laws-of-software.com/laws/kernighan/ https://www.laws-of-software.com/laws/kernighan/ Sure, you eat the elephant one bite at a time, and recursion is a thing but I wonder where the tipping point here is.
- tptacek 1y agoI think recursion is the wrong way to look at this, for what it's worth.
- raincole 1y agoHumans have a very strong tendency (and have made tremendous collective efforts) to compress context. I'm not a neuroscientist but I believe it's called "chunk." Language itself is a highly compressed form of compressed context. Like when you read "hoist with one's own petard" you don't just think about literal petard but the context behind this phrase.
- PantaloonFlames 1y agoWe don’t think of petards because no one knows what that is. :)
- KineticLensman 1y agoFor anyone wondering, it means blown into the air (‘hoist’) by your own bomb (‘petard’). From Shakespeare
- sdesol 1y ago> A human can effectively discard or disregard prior information as the narrow window of focus moves to a new task, LLMs seem incredibly bad at this. This is how I designed my LLM chat app (https://github.com/gitsense/chat https://github.com/gitsense/chat). I think agents have their place, but I really think if you want to solve complex problems without needlessly burning tokens, you will need a human in the loop to curate the context. I will get to it, but I believe in the same way that we developed different flows for working with Git, we will have different 'Chat Flows' for working with LLMs. I have an interactive demo at https://chat.gitsense.com https://chat.gitsense.com which shows how you can narrow the focus of the context for the LLM. Click "Start GitSense Chat Demos" then "Context Engineering & Management" to go through the 30 second demo.
- notatoad 1y agoi think that's really just a misunderstanding of what "bottleneck" means. a bottleneck isn't an obstacle where overcoming it will allow you to realize unlimited potential, a bottleneck is always just an obstacle to finding the next constraint. on actual bottles without any metaphors, the bottle neck is narrower because humans mouths are narrower.
- tom_m 1y agoYou don't want to discard prior information though. That's the problem with small context windows. Humans don't forget the original request as they ask for more information or go about a long task. Humans may forget parts of information along the way, but not the original goal and important parts. Not unless they have comprehension issues or ADHD, etc. This isn't a misconception. Context is a limitation. You can effectively have an AI agent build an entire application with a single prompt if it has enough (and the proper) context. The models with 1m context windows do better. Models with small context windows can't even do the task in many cases. I've tested this many, many, many times. It's tedious, but you can find the right model and the right prompts for success.