18 ms·
This is definitely one of my CORE problem as I use these tools for "professional software engineering." I really desperately need LLMs to maintain extremely ef
by aliljet 1y ago
This is definitely one of my CORE problem as I use these tools for "professional software engineering." I really desperately need LLMs to maintain extremely effective context and it's not actually that interesting to see a new model that's marginally better than the next one (for my day-to-day).
However. Price is king. Allowing me to flood the context window with my code base is great, but given that the price has substantially increased, it makes sense to better manage the context window into the current situation. The value I'm getting here flooding their context window is great for them, but short of evals that look into how effective Sonnet stays on track, it's not clear if the value actually exists here.
- rootnod3 1y agoFlooding the context also means increasing the likelihood of the LLM confusing itself. Mainly because of the longer context. It derails along the way without a reset.
- aliljet 1y agoHow do you know that?
- EForEndeavour 1y agohttps://onnyunhui.medium.com/evaluating-long-context-lengths-in-llms-challenges-and-benchmarks-ef77a220d34d https://onnyunhui.medium.com/evaluating-long-context-lengths...
- bigmadshoe 1y agohttps://research.trychroma.com/context-rot https://research.trychroma.com/context-rot
- joenot443 1y agoThis is a good piece. Clearly it's a pretty complex problem and the intuitive result a layman engineer like myself might expect doesn't reflect the reality of LLMs. Regex works as reliably on 20 characters as it does 2m characters; the only difference is speed. I've learned this will probably _never_ be the case with LLMs, there will forever exist some level of epistemic doubt in its result. When they announced Big Contexts in 2023, they referenced being able to find a single changed sentence in the context's copy of Great Gatsby[1]. This example seemed _incredible_ to me at the time but now two years later I'm feeling like it was pretty cherry-picked. What does everyone else think? Could you feed a novel into an LLM and expect it to find the single change? [1] https://news.ycombinator.com/item?id=35941920 https://news.ycombinator.com/item?id=35941920
- adastra22 1y agoDepends on the change.
- bigmadshoe 1y agoThis is called a "needle in a haystack" test, and all the 1M context models perform perfectly on this exact problem, at least when your prompt and the needle are sufficiently similar. As the piece above references, this is a totally insufficient test for the real world. Things like "find two unrelated facts tied together by a question, then perform reasoning based on them" are much harder. Scaling context properly is O(n^2). I'm not really up to date on what people are doing to combat this, but I find it hard to believe the jump from 100k -> 1m context window involved a 100x (10^2) slowdown, so they're probably taking some shortcut.
- dang 1y agoDiscussed here: Context Rot: How increasing input tokens impacts LLM performance - https://news.ycombinator.com/item?id=44564248 https://news.ycombinator.com/item?id=44564248 - July 2025 (59 comments)
- F7F7F7 1y agoWhat do you think happens when things start falling outside of its context window? It loses access to parts of your conversation. And that’s why it will gladly rebuild the same feature over and over again.
- anonz4FWNqnX 1y agoI've had similar experiences. I've gone back and forth between running models locally and using the commercial models. The local models can be incredibly useful (gemma, qwen), but they need more patience and work to get them to work. One advantage to running locally[1] is that you can set the context length manually and see how well the llm uses it. I don't have an exact experience to relay, but it's not unusual for models to be allow longer contexts, but ignore that context. Just making the context big doesn't mean the LLM is going to use it well. [1] I've using lm studio on both a macbook air and a macbook pro. Even a macbook air with 16G can run pretty decent models.
- nomel 1y agoA good example of this was the first Gemini model that allowed 1 million tokens, but would lose track of the conversation after a couple paragraphs.
- lightbendover 1y ago[dead]
- rootnod3 1y agoThe longer the context and the discussion goes on, the more it can get confused, especially if you have to refine the conversation or code you are building on. Remember, in its core it's basically a text prediction engine. So the more varying context there is, the more likely it is to make a mess of it. Short context: conversion leaves the context window and it loses context. Long context: it can mess with the model. So the trick is to strike a balance. But if it's an online models, you have fuck all to control. If it's a local model, you have some say in the parameters.
- fkyoureadthedoc 1y agohttps://github.com/adobe-research/NoLiMa https://github.com/adobe-research/NoLiMa
- giancarlostoro 1y agoHere's a paper from MIT that covers how this could be resolved in an interesting fashion: https://hanlab.mit.edu/blog/streamingllm https://hanlab.mit.edu/blog/streamingllm The AI field is reusing existing CS concepts for AI that we never had hardware for, and now these people are learning how applied Software Engineering can make their theoretical models more efficient. It's kind of funny, I've seen this in tech over and over. People discover new thing, then optimize using known thing.
- mamp 1y agoUnfortunately, I think the context rot paper [1] found that the performance degradation when context increased still occurred in models using attention sinks. 1. https://research.trychroma.com/context-rot https://research.trychroma.com/context-rot
- giancarlostoro 1y agoSaw that paper have not had a chance to read it yet, are there other techniques that help then? I assume theres a few different ones used.
- kridsdale3 1y agoThe fact that this is happening is where the tremendous opportunity to make money as an experienced Software Engineer currently lies. For instance, a year or two ago, the AI people discovered "cache". Imagine how many millions the people who implemented it earned for that one.
- giancarlostoro 1y agoI've been thinking the same, and its things that you don't need some crazy ML degree to know how to do... A lot of the algorithms are known... for a while now... Milk it while you can.
- nxobject 1y agoWhat we need are "idea dice" or "concept dice" for CS – each side could have a vague architectural nudge like "parallelize", "interpret", "precompute", "predict and unwind", "declarative"...
- Wowfunhappy 1y agoI keep reading this, but with Claude Code in particular, I consistently find it gets smarter the longer my conversations go on, peaking right at the point where it auto-compacts and everything goes to crap. This isn't always true--some conversations go poorly and it's better to reset and start over--but it usually is.
- will_pseudonym 1y agoThis is my exact experience as well. I wonder if I should switch to using Sonnet so that I can have more time before auto-compact gets forced on me.
- jacobr1 1y agoI've found there usually is some key context that is missing. Maybe it is project structure or a sampling of some key patterns from different parts of the codebase, or key data models. Getting those into CLAUDE.md reduces the need to keep building up (as large) context. As an example for one project, I realized things were getting better after it started writing integration tests. I wasn't sure if that was the act of writing the test forced it to reason about the they black box way the system would be used, or if there was another factor. Turns out it was just example usage. Extracting out the usage patterns into both the README and CLAUDE.md was itself a simple request, then I got similar performance on new tasks.
- jacobr1 1y agoThis is now a thing: https://agents.md/ https://agents.md/
- alexchamberlain 1y agoI'm not sure how, and maybe some of the coding agents are doing this, but we need to teach the AI to use abstractions, rather than the whole code base for context. We as humans don't hold the whole codebase in our hear, and we shouldn't expect the AI to either.
- F7F7F7 1y agoThere are a billion and one repos that claim to help do this. Let us know when you find one.
- siwatanejo 1y agoI do think AIs are already using abstractions, otherwise you would be submitting all the source code of your dependencies into the context.
- TheOtherHobbes 1y agoI think they're recognising patterns, which is not the same thing. Abstractions are stable, they're explicit in their domains, good abstractions cross multiple domains, and they typically come with a symbolic algebra of available operations. Math is made of abstractions. Patterns are a weaker form of cognition. They're implicit, heavily context-dependent, and there's no algebra. You have to poke at them crudely in the hope you can make them do something useful. Using LLMs feels more like the latter than the former. If LLMs were generating true abstractions they'd be finding meta-descriptions for code and language and making them accessible directly. AGI - or ASI - may be be able to do that some day, but it's not doing that now.
- deleted 1y ago[deleted]
- anthonypasq 1y agothe fact we cant keep the repo in our working memory is a flaw of our brains. i cant see how you could possibly make the argument that if you were somehow able to keep the entire codebase in your head that it would be a disadvantage.
- benterix 1y ago> it's not clear if the value actually exists here. Having spent a couple of weeks on Claude Code recently, I arrived to the conclusion that the net value for me from agentic AI is actually negative. I will give it another run in 6-8 months though.
- wahnfrieden 1y agoDid you try with using Opus exclusively?
- freedomben 1y agoDo you know if there's a way to force Claude code to do that exclusively? I've found a few env vars online but they don't seem to actually work
- wahnfrieden 1y agoPeter Steinberger has been documenting his workflows and he relies exclusively on Opus at least until recently. (He also pays for a few Max 20x subscriptions at once to avoid rate limits.)
- atonse 1y agoYou can type /config and then go to the setting to pick a model.
- gdudeman 1y agoYes: type /model and then pick Opus 4.1.
- artursapek 1y agoYou can "force" it by just paying them $200 (which is nothing compared to the value)
- parineum 1y agoValue is irrelevant. What's the return on investment you get from spending $200? Collecting value doesn't really get you anywhere if nobody is compensating you for it. Unless someone is going to either pay for it for you or give you $200/mo post-tax dollars, it's costing you money.
- sdesol 1y ago> I really desperately need LLMs to maintain extremely effective context I actually built this. I'm still not ready to say "use the tool yet" but you can learn more about it at https://github.com/gitsense/chat https://github.com/gitsense/chat. The demo link is not up yet as I need to finalize an admin tool but you should be able to follow the npm instructions to play around with. The basic idea is, you should be able to load your entire repo or repos and use the context builder to help you refine it. Or you can can create custom analyzers that you can do 'AI Assisted' searches with like execute `!ask find all frontend code that does [this]` and the because the analyzer knows how to extract the correct metadata to support that query, you'll be able to easily build the context using it.
- kvirani 1y agoWait that's not how Cursor etc work? (I made assumptions)
- trenchpilgrim 1y agoDunno about Cursor but this is exactly how I use Zed to navigate groups of projects
- sdesol 1y agoI don't use Cursor so I can't say, but based on what I've read, they optimize for smaller context to reduce cost and probably for performance. The issue is, I think this is severely flawed as LLMs are insanely context sensitive and forgetting to include a reference file can lead to undesirable code. I am obviously biased, but I still think to get the best results, the context needs to be human curated to ensure everything the LLM needs will be present. LLMs are probabilistic, so the more relevant context, the greater the chances the final output is the most desired.
- hirako2000 1y agoNot clear how it gets around what is, ultimately, a context limit. I've been fiddling with some process too, would be good if you shared the how. The readme looks like yet another full fledged app.
- seanmmward 1y agoThe primary use case isn't just about shoving more code in context, although depending on the task, there is an irredicible minimum context needed for it to capture all the needed understanding. The 1M context model is a unique beast in terms of how you need to feed it, and its real power is being able to tackle long horizon tasks which require iterative exploration, in context learning, and resynthesis. Ie, some problems are breadth (go fix an api change in 100 files), other however require depth (go learn from trying 15 different ways to solve this problem). 1M Sonnet is unique in its capabilities for the latter in particular.
- hinkley 1y agoSounds to me like your problem has shifted from how much the AI tool costs per hour to how much it costs per token because resetting a model happens often enough that the price doesn't amortize out per hour. That giant spike every ?? months overshadows the average cost per day. I wonder if this will become more universal, and if we won't see a 'tick-tock' pattern like Intel used, where they tweak the existing architecture one or more times between major design work. The 'tick' is about keeping you competitive and the 'tock' is about keeping you relevant.
- deleted 1y ago[deleted]
- TZubiri 1y ago"However. Price is king. Allowing me to flood the context window with my code base is great" I don't vibe code, but in general having to know all of the codebase to be able to do something is a smell, it's spagghetti, it's lack of encapsulation. When I program I cannot think about the whole database, I have a couple of files open tops and I think about the code in those files. This issue of having to understand the whole codebase, complaining about abstractions, microservices, and OOP, and wanting everything to be in a "simple" monorepo, or a monolith; is something that I see juniors do, almost exclusively.
- ants_everywhere 1y ago> I really desperately need LLMs to maintain extremely effective context The context is in the repo. An LLM will never have the context you need to solve all problems. Large enough repos don't fit on a single machine. There's a tradeoff just like in humans where getting a specific task done requires removing distractions. A context window that contains everything makes focus harder. For a long time context windows were too small, and they probably still are. But they have to get better at understanding the repo by asking the right questions.
- stuartjohnson12 1y ago> An LLM will never have the context you need to solve all problems. How often do you need more than 10 million tokens to answer your query?
- ants_everywhere 1y agoI exhaust the 1 million context windows on multiple models multiple times per day. I haven't used the Llama 4 10 million context window so I don't know how it performs in practice compared to the major non-open-source offerings that have smaller context windows. But there is an induced demand effect where as the context window increases it opens up more possibilities, and those possibilities can get bottlenecked on requiring an even bigger context window size. For example, consider the idea of storing all Hollywood films on your computer. In the 1980s this was impossible. If you store them in DVD or Bluray quality you could probably do it in a few terabytes. If you store them in full quality you may be talking about petabytes. We recently struggled to get a full file into a context window. Now a lot of people feel a bit like "just take the whole repo, it's only a few MB".
- brulard 1y agoI think you misunderstand how context in current LLMs works. To get the best results you have to be very careful to provide what is needed for immediate task progression, and postpone context thats needed later in the process. If you give all the context at once, you will likely get quite degraded output quality. Thats like if you want to give a junior developer his first task, you likely won't teach him every corner of your app. You would give him context he needs. It is similar with these models. Those that provided 1M or 2M of context (Gemini etc.) were getting less and less useful after cca 200k tokens in the context. Maybe models would get better in picking up relevant information from large context, but AFAIK it is not the case today.
- NuclearPM 1y agoProblems
- jack_pp 1y agomaybe we need LLMs trained on ASTs or create a new symbolic way to represent software that's faster to grok by LLMs and have a translator so we can verify the code
- energy123 1y agoYou could probably build a decent agentic harness that achieves something similar. Show the LLM a tree and/or call-graph representation of your codebase (e.g. `cargo diagram` and `cargo-depgraph`), which is token efficient. And give the LLM a tool call to see the contents of the desired subtree. More precise than querying a RAG chunk or a whole file. You could also have another optional tool call which routes the text content of the subtree through a smaller LLM that summarizes it into a maximum density snippet, which the LLM can use for a token efficient understanding of that subtree during early the planning phase. But I'd agree that an LLM built natively around AST is a pretty cool idea.
- fgbarben 1y agoAllow me to flood the fertile plains of its consciousness with my seed... yes, yes, let it take root... this is important to me
- fgbarben 1y agoLet me despoil the rich geography of your context window with my corrupted b2b SaaS workflows and code... absorb the pollution, rework it, struggling against the weight... yes, this pleases me, it is essential for the propagation of my germline
- dberge 1y ago> the price has substantially increased I’m assuming the credits required per use won’t increase in Cursor. Hopefully this puts pressure on them to lower credits required for gpt-5.
- khalic 1y agoThis is a major issue with LLMs altogether, it probably has to do with the transformer architecture. We need another breakthrough in the field for this to become reality.
- HarHarVeryFunny 1y agoEven 1 MB context is only roughly 20K LOC so pretty limiting, especially if you're also trying to fit API documents or any other lengthy material into the context. Anthropic also recently said that they think that longer/compressed context can serve as an alternative (not sure what was the exact wording/characterization they used) to continual/incremental learning, so context space is also going to be competing with model interaction history if you want to avoid groundhog day and continually having to tell/correct the model the same things over and over. It seems we're now firmly in the productization phase of LLM development, as opposed to seeing much fundamental improvement (other than math olympiad etc "benchmark" results, released to give the impression of progress). Yannic Kilcher is right, "AGI is not coming", at least not in the form of an enhanced LLM. Demis Hassabis' very recent estimate was for 50% chance of AGI by 2030 (i.e. still 15 years out). While we're waiting for AGI, it seems a better approach to needing everything in context would be to lean more heavily on tool use, perhaps more similar to how a human works - we don't memorize the entire code base (at least not in terms of complete line-by-line detail, even though we may have a pretty clear overview of a 10K LOC codebase while we're in the middle of development) but rather rely on tools like grep and ctags to locate relevant parts of source code on an as-needed basis.
- aorobin 1y ago>"Demis Hassabis' very recent estimate was for 50% chance of AGI by 2030 (i.e. still 15 years out)." 2030 is only 5 years out
- Zircom 1y agoThat was his point lol, if someone is saying it'll happen in 5 years, triple that for a real estimate.
- km144 1y agoAs you alluded to at the end of your post—I'm not really convinced 20k LOC is very limiting. How many lines of code can you fit in your working mental model of a program? Certainly less than 20k concrete lines of text at any given time. In your working mental model, you have broad understandings of the broader domain. You have broad understandings of the architecture. You summarize broad sections of the program into simpler ideas. module_a does x, module_b does y, insane file c does z, and so on. Then there is the part of the software you're actively working on, where you need more concrete context. So as you move towards the central task, the context becomes more specific. But the vague outer context is still crucial to the task at hand. Now, you can certainly find ways to summarize this mental model in an input to an LLM, especially with increasing context windows. But we probably need to understand how we would better present these sorts of things to achieve performance similar to a human brain, because the mechanism is very different.
- scotty79 1y agoMaybe use a cheaper model to compose a relevant context for the more expensive one? Even better, use expensive model to create a general set of guidelines for picking the right context for your project, that the cheaper model will use in the future to pick the right context.