4 ms·
If a model is not making use of the whole context window - shouldn't that be very noticeable when the prompt is code? For example when querying a model to refa
by mg 11mo ago
If a model is not making use of the whole context window - shouldn't that be very noticeable when the prompt is code?
For example when querying a model to refactor a piece of code - would that really work if it forgets about one part of the code while it refactors another part?
I concatenate a lot of code files into a single prompt multiple times a day and ask LLMs to refactor them, implement features or review the code.
So far, I never had the impression that filling the context window with a lot of code causes problems.
I also use very long lists of instructions on code style on top of my prompts. And the LLMs seem to be able to follow all of them just fine.
- MallocVoidstar 11mo agoI don't think there are any up-to-date leaderboards, but models absolutely degrade in performance the more context they're dealing with. https://wandb.ai/byyoung3/ruler_eval/reports/How-to-evaluate-the-true-context-length-of-your-LLM-using-RULER---VmlldzoxNDE0OTA0OQ https://wandb.ai/byyoung3/ruler_eval/reports/How-to-evaluate... >Gpt-5-mini records 0.87 overall judge accuracy at 4k [context] and falls to 0.59 at 128k. And Llama 4 Scout claimed a 10 million token context window but in practice its performance on query tasks drops below 20% accuracy by 32k tokens.
- mg 11mo agoThat makes me wonder if we could simply test this by letting the LLM add or multiply a long list of numbers? Here is an experiment: https://www.gnod.com/search/#q=%23%20Calcuate%20the%20below%20number.%20Do%20not%20use%20a%20calculator.%20Do%20it%20in%20your%20head.%0A%0A(1.1%20%2B%202.5)%20*%20(7.1%20%2B%202.5)%20*%20(3.3%20%2B%204.4)%20*%20(12.3%20-%201.4)%20*%20(2.3%20-%201.1)%20*%20(5%20%2B%201)%20*%20(7.1%20%2B%202.5)%20*%20(3.3%20%2B%204.4)%20*%20(12.3%20-%201.4)%20*%20(2.3%20-%201.1) https://www.gnod.com/search/#q=%23%20Calcuate%20the%20below%... The correct answer: Correct: 20,192,642.460942328 Here is what I got from different models on the first try: ChatGPT: 20,384,918.24 Perplexity: 20,000,000 Google: 25,167,098.4 Mistral: 200,000,000 Grok: Timed out after 300s of thinking
- jarek83 11mo agoIsn't that LLMs are not designed to do calculations?
- mg 11mo agoNeither are humans.
- cuu508 11mo agoBut humans can still do it.
- cluckindan 11mo agoThey are not LMMs, after all…
- gcanyon 11mo ago> Do not use a calculator. Do it in your head. You wouldn't ask a human to do that, why would you ask an LLM to? I guess it's a way to test them, but it feels like the world record for backwards running: interesting, maybe, but not a good way to measure, like, anything about the individual involved.
- throwuxiytayq 11mo agoI’m starting to find it unreasonably funny how people always want language models to multiply numbers for some reason. Every god damn time. In every single HN thread. I think my sanity might be giving out.
- solatic 11mo agoA model, no, but an agent with a calculator tool? Then there's the question of why not just build the calculator tool into the model?
- KristoAI 11mo ago[flagged]