23 ms·
The new skill in AI is not prompting, it's context engineering
- bsoles 1y ago> "the art of providing all the context for the task to be plausibly solvable by the LLM.” And who is going to do that? The "context engineer", who doesn't know anything about the subject and runs to the LLM for quick answers without having any ability to evaluate if the answer is solid or not? We saw the same story with "data scientists". A general understanding of tools with no understanding of the specific application areas is bound to result in crappy products, if not in business disasters.
- Leo-thorne 1y agoI learned this the hard way. Even a great prompt won't work if the context window is off. If key information is missing or important history is buried too deep, the model will still fail. Now I always explain the problem clearly and set the scene before letting the AI take over.
- byannafilip 1y agoi actually found that there is a startup actually doing that called theo growth, they are tackling the problem of missing this context layer
- iszomer 1y agoThis article reads as if the author pulled this idea straight from the MGS2's ending regarding giving and creating context..
- mrcoders 1y agoLooking at this thread, excited to see the community articulating what many of us have been experiencing. The shift from "prompt engineering" to "context engineering" really captures what's happening. The technical stuff (RAG, vector databases) is getting commoditized. But there's this foundational knowledge organization layer that's becoming critical. Most companies have context scattered across wikis, Slack, Google Docs. You can build sophisticated retrieval systems, but if you're feeding them fragmented information, you're missing huge optimization opportunities. There's research backing this - Microsoft/Salesforce found 39% accuracy drop in multi-turn conversations - so for the agent interaction's it's even more cirtical to give 'just enough' context (https://arxiv.org/pdf/2505.06120 https://arxiv.org/pdf/2505.06120). When (business) context is properly structured upfront, you minimize those patterns.
- damnever 1y agoIt is still "prompting".
- baxtr 1y ago>Conclusion Building powerful and reliable AI Agents is becoming less about finding a magic prompt or model updates. It is about the engineering of context and providing the right information and tools, in the right format, at the right time. It’s a cross-functional challenge that involves understanding your business use case, defining your outputs, and structuring all the necessary information so that an LLM can “accomplish the task." That’s actually also true for humans: the more context (aka right info at the right time) you provide the better for solving tasks.
- QuercusMax 1y agoYeah... I'm always asking my UX and product folks for mocks, requirements, acceptance criteria, sample inputs and outputs, why we care about this feature, etc. Until we can scan your brain and figure out what you really want, it's going to be necessary to actually describe what you want built, and not just rely on vibes.
- lupire 1y agoNot "more" context. "Better" context. (X-Y problem, for example.)
- root_axis 1y agoI am not a fan of this banal trend of superficially comparing aspects of machine learning to humans. It doesn't provide any insight and is hardly ever accurate.
- ModernMech 1y agoI agree, however I do appreciate comparisons to other human-made systems. For example, "providing the right information and tools, in the right format, at the right time" sounds a lot like a bureaucracy, particularly because "right" is decided for you, it's left undefined, and may change at any time with no warning or recourse.
- furyofantares 1y agoI've seen a lot of cases where, if you look at the context you're giving the model and imagine giving it to a human (just not yourself or your coworker, someone who doesn't already know what you're trying to achieve - think mechanical turk), the human would be unlikely to give the output you want. Context is often incomplete, unclear, contradictory, or just contains too much distracting information. Those are all things that will cause an LLM to fail that can be fixed by thinking about how an unrelated human would do the job.
- pwarner 1y agoIt's an integration adventure. This is why much AI is failing in the enterprise. MS Copilot is moderately interesting for data in MS Office, but forget about it accessing 90% of your data that's in other systems.
- simonw 1y agoI wrote a bit about this the other day: https://simonwillison.net/2025/Jun/27/context-engineering/ https://simonwillison.net/2025/Jun/27/context-engineering/ Drew Breunig has been doing some fantastic writing on this subject - coincidentally at the same time as the "context engineering" buzzword appeared but actually unrelated to that meme. How Long Contexts Fail - https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-ho... - talks about the various ways in which longer contexts can start causing problems (also known as "context rot") How to Fix Your Context - https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.html https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.... - gives names to a bunch of techniques for working around these problems including Tool Loadout, Context Quarantine, Context Pruning, Context Summarization, and Context Offloading.
- the_mitsuhiko 1y agoDrew Breunig's posts are a must read on this. This is not only important for writing your own agents, it is also critical when using agentic coding right now. These limitations/behaviors will be with us for a while.
- outofpaper 1y agoThey might be good reads on the topic but Drew makes some significant etymological mistakes. For example loadout doesn't come from gaming but military terminology. It's essentially the same as kit or gear.
- ZYbCRq22HbJ2y7 1y ago> They might be good reads on the topic but Drew makes some significant etymological mistakes. For example loadout doesn't come from gaming but military terminology. It's essentially the same as kit or gear. Doesn't seem that significant? Not to say those blog posts say anything much anyway that any "prompt engineer" (someone who uses LLMs frequently) doesn't already know, but maybe it is useful to some at such an early stage of these things.
- 1y ago
- crystal_revenge 1y agoDefinitely mirrors my experience. One heuristic I've often used when providing context to model is "is this enough information for a human to solve this task?". Building some text2SQL products in the past it was very interesting to see how often when the model failed, a real data analyst would reply something like "oh yea, that's an older table we don't use any more, the correct table is...". This means the model was likely making a mistake that a real human analyst would have without the proper context. One thing that is missing from this list is: evaluations! I'm shocked how often I still see large AI projects being run without any regard to evals. Evals are more important for AI projects than test suites are for traditional engineering ones. You don't even need a big eval set, just one that covers your problem surface reasonably well. However without it you're basically just "guessing" rather than iterating on your problem, and you're not even guessing in a way where each guess is an improvement on the last. edit: To clarify, I ask myself this question. It's frequently the case that we expect LLMs to solve problems without the necessary information for a human to solve them.
- kevin_thibedeau 1y agoAsking yes no questions will get you a lie 50% of the time.
- adriand 1y agoI have pretty good success with asking the model this question before it starts working as well. I’ll tell it to ask questions about anything it’s unsure of and to ask for examples of code patterns that are in use in the application already that it can use as a template.
- hobs 1y agoThe thing is, all the people cosplaying as data scientists don't want evaluations, and that's why you saw so little in fake C level projects, because telling people the emperor has no clothes doesn't pay. For those actually using the products to make money well, hey - all of those have evaluations.
- shermantanktop 1y ago
- bGl2YW5j 1y agoSaw this the other day and it made me think that too much effort and credence is being given to this idea of crafting the perfect environment for LLMs to thrive in. Which to me, is contrary to how powerful AI systems should function. We shouldn’t need to hold its hand so much. Obviously we’ve got to tame the version of LLMs we’ve got now, and this kind of thinking is a step in the right direction. What I take issue with is the way this thinking is couched as a revolutionary silver bullet.
- 4ndrewl 1y agoReminds me of first gen chatbots where the user had to put in the effort of trying to craft a phrase in a way that would garner the expected result. It's a form of user-hostility.
- gametorch 1y agoIt's still way easier for me to say "here's where to find the information to solve the task" than for me to manually type out the code, 99% of the time
- ramesh31 1y agoWe shouldn't but it's analogous to how CPU usage used to work. In the 8 bit days you could do some magical stuff that was completely impossible before microcomputers existed. But you had to have all kinds of tricks and heuristics to work around the limited abilities. We're in the same place with LLMs now. Some day we will have the equivalent of what gigabytes or RAM are to a modern CPU now, but we're still stuck in the 80s for now (which was revolutionary at the time).
- smeej 1y agoIt also reminds me of when you could structure an internet search query and find exactly what you wanted. You just had to ask it in the machine's language. I hope the generalized future of this doesn't look like the generalized future of that, though. Now it's darn near impossible to find very specific things on the internet because the search engines will ignore any "operators" you try to use if they generate "too few" results (by which they seem to mean "few enough that no one will pay for us to show you an ad for this search"). I'm moderately afraid the ability to get useful results out of AIs will be abstracted away to some lowest common denominator of spammy garbage people want to "consume" instead of use for something.
- JohnMakin 1y ago> Building powerful and reliable AI Agents is becoming less about finding a magic prompt or model updates. Ok, I can buy this > It is about the engineering of context and providing the right information and tools, in the right format, at the right time. when the "right" format and "right" time are essentially, and maybe even necessarily, undefined, then aren't you still reaching for a "magic" solution? If the definition of "right" information is "information which results in a sufficiently accurate answer from a language model" then I fail to see how you are doing anything fundamentally differently than prompt engineering. Since these are non-deterministic machines, I fail to see any reliable heuristic that is fundamentally indistinguishable than "trying and seeing" with prompts.
- deleted 1y ago[deleted]
- edwardbernays 1y agoThe state of the art theoretical frameworks typically separates these into two distinct exploratory and discovery phases. The first phase, which is exploratory, is best conceptualized as utilizing an atmospheric dispersion device. An easily identifiable marker material, usually a variety of feces, is metaphorically introduced at high velocity. The discovery phase is then conceptualized as analyzing the dispersal patterns of the exploratory phase. These two phases are best summarized, respectively, as "Fuck Around" followed by "Find Out."
- mentalgear 1y agoIt's magical thinking all the way down. Whether they call it now "prompt" or "context" engineering because it's the same tinkering to find something that "sticks" in non-deterministic space.
- nonethewiser 1y ago>Whether they call it now "prompt" or "context" engineering because it's the same tinkering to find something that "sticks" in non-deterministic space. I dont quite follow. Prompts and contexts are different things. Sure, you can get thing into contexts with prompts but that doesn't mean they are entirely the same. You could have a long running conversation with a lot in the context. A given prompt may work poorly, whereas it would have worked quite well earlier. I don't think this difference is purely semantic. For whatever it's worth I've never liked the term "prompt engineering." It is perhaps the quintessential example of overusing the word engineering.
- ModernMech 1y ago"Wow, AI will replace programming languages by allowing us to code in natural language!" "Actually, you need to engineer the prompt to be very precise about what you want to AI to do." "Actually, you also need to add in a bunch of "context" so it can disambiguate your intent." "Actually English isn't a good way to express intent and requirements, so we have introduced protocols to structure your prompt, and various keywords to bring attention to specific phrases." "Actually, these meta languages could use some more features and syntax so that we can better express intent and requirements without ambiguity." "Actually... wait we just reinvented the idea of a programming language."
- throwawayoldie 1y agoOnly without all that pesky determinism and reproducibility. (Whoever's about to say "well ackshually temperature of zero", don't.)
- whatevertrevor 1y agoYou forgot about lower performance and efficiency. And longer build/run cycles. And more hardware/power usage.
- throwawayoldie 1y agoThere's just so much to like* about this technology, I was bound to forget something. (*) "like" in the sense of "not like"
- mindok 1y ago“Actually - curly braces help save space in the context while making meaning clearer”
- georgeburdell 1y agoWe should have known up through Step 4 for a while. See: the legal system
- nimish 1y ago
- eddythompson80 1y agoWhich is funny because everyone is already looking at AI as: I have 30 TB of shit that is basically "my company". Can I dump that into your AI and have another, magical, all-konwning, co-worker?
- coliveira 1y agoWhich I think it is double funny because, given the zeal with which companies are jumping into this bandwagon, AI will bankrupt most businesses in record time! Just imagine the typical company firing most workers and paying a fortune to run on top of a schizophrenic AI system that gets things wrong half of the time...
- eddythompson80 1y agoYes, you can see the insanely accelerated pace of bankruptcies or "strategic realignments" among AI startups. I think it's just game theory in play and we can do nothing but watch it play out. The "up side" is insane, potentially unlimited. The price is high, but so is the potential reward. By the rules of the game, you have to play. There is no other move you can make. No one knows the odds, but we know the potential reward. You could be the next T company easy. You could realistically go from startup -> 1 Trillion in less than a year if you are right. We need to give this time to play itself out. The "odds" will eventually be better estimated and it'll affect investment. In the mean time, just give your VC Google's, Microsoft's, or AWS's direct deposit info. It's easier that way.
- whimsicalism 1y agoi think context engineering as described is somewhat a subset of ‘environment engineering.’ the gold-standard is when an outcome reached with tools can be verified as correct and hillclimbed with RL. most of the engineering effort is from building the environment and verifier while the nuts and bolts of grpo/ppo training and open-weight tool-using models are commodities.
- intellectronica 1y agoSee also: https://ai.intellectronica.net/context-engineering https://ai.intellectronica.net/context-engineering for an overview.
- jshorty 1y agoI have felt somewhat frustrated with what I perceive as a broad tendency to malign "prompt engineering" as an antiquated approach for whatever new the industry technique is with regards to building a request body for a model API. Whether that's RAG years ago, nuance in a model request's schema beyond simple text (tool calls, structured outputs, etc), or concepts of agentic knowledge and memory more recently. While models were less powerful a couple of years ago, there was nothing stopping you at that time from taking a highly dynamic approach to what you asked of them as a "prompt engineer"; you were just more vulnerable to indeterminism in the contract with the models at each step. Context windows have grown larger; you can fit more in now, push out the need for fine-tuning, and get more ambitious with what you dump in to help guide the LLM. But I'm not immediately sure what skill requirements fundamentally change here. You just have more resources at your disposal, and can care less about counting tokens.
- simonw 1y agoI liked what Andrej Karpathy had to say about this: https://twitter.com/karpathy/status/1937902205765607626 https://twitter.com/karpathy/status/1937902205765607626 > [..] in every industrial-strength LLM app, context engineering is the delicate art and science of filling the context window with just the right information for the next step. Science because doing this right involves task descriptions and explanations, few shot examples, RAG, related (possibly multimodal) data, tools, state and history, compacting... Too little or of the wrong form and the LLM doesn't have the right context for optimal performance. Too much or too irrelevant and the LLM costs might go up and performance might come down. Doing this well is highly non-trivial. And art because of the guiding intuition around LLM psychology of people spirits.
- bgwalter 1y agoAll that work just for stripping a license. If one uses code directly from GitHub, copy and paste is sufficient. One can even keep the license.
- saejox 1y agoClaude 3.5 was released 1 year ago. Current LLMs are not much better at coding than it. Sure they are more shiny and well polished, but not much better at all. I think it is time to curb our enthusiasm. I almost always rewrite AI written functions in my code a few weeks later. Doesn't matter they have more context or better context, they still fail to write code easily understandable by humans.
- simonw 1y agoClaude 3.5 was remarkably good at writing code. If Claude 3.7 and Claude 4 are just incremental improvements on that then even better! I actually think they're a lot more than incremental. 3.7 introduced "thinking" mode and 4 doubled down on that and thinking/reasoning/whatever-you-want-to-call-it is particularly good at code challenges. As always, if you're not getting great results out of coding LLMs it's likely you haven't spent several months iterating on your prompting techniques to figure out what works best for your style of development.
- davidclark 1y agoGood example of why I have been totally ignoring people who beat the drum of needing to develop the skills of interacting with models. “Learn to prompt” is already dead? Of course, the true believers will just call this an evolution of prompting or some such goalpost moving. Personally, my goalpost still hasn’t moved: I’ll invest in using AI when we are past this grand debate about its usefulness. The utility of a calculator is self-evident. The utility of an LLM requires 30k words of explanation and nuanced caveats. I just can’t even be bothered to read the sales pitch anymore.
- simonw 1y agoWe should be so far past the "grand debate about its usefulness" at this point. If you think that's still a debate, you might be listening to the small pool of very loud people who insist nothing has improved since the release of GPT-4.
- fragmede 1y agoShould be, but the bar for scientifically proven is high. Absent actual studies showing this, (and with a large N), people will refuse to believe things they don't want to be true.
- nandhinianand 1y agoI think this is definitely true for novel writing and stuff like that based on my experiments with AI so far.. I'm still on the fence about coding/building s/w based on it, but that may just be about the unlearning and re-learning i'm yet to do/try out.
- davidclark 1y agoHave you considered the opposite? Reflected on your own biases? I’m listening to my own experience. Just today I gave it another fair shot. GitHub Copilot agent mode with GPT-4.1. Still unimpressed. This is a really insightful look at why people perceive the usefulness of these models differently. It is fair to both sides without being dismissive as one side just not “getting it” or how we should be “so far” past debate: https://ferd.ca/the-gap-through-which-we-praise-the-machine.html https://ferd.ca/the-gap-through-which-we-praise-the-machine....
- _pdp_ 1y agoIt is wrong. The new/old skill is reverse engineering. If the majority of the code is generated by AI, you'll still need people with technical expertise to make sense of it.
- CamperBob2 1y agoNot really. Got some code you don't understand? Feed it to a model and ask it to add comments. Ultimately humans will never need to look at most AI-generated code, any more than we have to look at the machine language emitted by a C compiler. We're a long way from that state of affairs -- as anyone who struggled with code-generation bugs in the first few generations of compilers will agree -- but we'll get there.
- fumblingness 1y ago"And at no point does it ever occur to you to demand proof that measures such as this will have the desired effect... or, indeed, that the desired effect is indeed worth achieving at all." - you (https://news.ycombinator.com/item?id=44439447 https://news.ycombinator.com/item?id=44439447)
- CamperBob2 1y ago(Shrug) There's a difference between prescription and prediction. I predict that after 50 years of doing the same old shit the same old way, the practice of programming is about to undergo a series of wrenching changes that amount to nothing less than revolution. Changes powered by radical new insights into the nature and function of language itself. I'm not initiating these changes, voting for them, or attempting to persuade other people to do so, as the person I replied to in the other thread is doing. I do welcome having something new and interesting to learn and think about, though. Anyway, always good to hear from a new fan!
- rvz 1y ago> Not really. Got some code you don't understand? Feed it to a model and ask it to add comments. Absolutely not. An experienced individual in their field can tell if the AI made a mistake in the comments / code rather than the typical untrained eye. So no, actually read the code and understand what it does. > Ultimately humans will never need to look at most AI-generated code, any more than we have to look at the machine language emitted by a C compiler. So for safety critical systems, one should not look or check if code has been AI generated?
- adhamsalama 1y agoThere is no engineering involved in using AI. It's insulting to call begging an LLM "engineering".
- rednafi 1y agoThis. Convincing a bullshit generator to give you the right data isn’t engineering, it quackery. But I guess “context quackery” wouldn’t sell as much. LLMs are quite useful and I leverage them all the time. But I can’t stand these AI yappers saying the same shit over and over again in every media format and trying to sell AI usage as some kind of profound wizardry when it’s not.
- mikhmha 1y agoIt is total quackery. When you zoom out in these discussions you begin to see how the AI yappers and their methodology is just modern-day alchemy with its own jargon and "esoteric" techniques.
- simonw 1y agoSee my comment here. These new context engineering techniques are a whole lot less quackery than the prompting techniques from last year: https://news.ycombinator.com/item?id=44428628 https://news.ycombinator.com/item?id=44428628
- asadotzler 1y ago"less"
- ModernMech 1y agoThe quackery comes in the application of these techniques, promising that they "work" without ever really showing it. Of course what's suggested in that blog sounds rational -- they're just restating common project management practices. What makes it quackery is there's no evidence to show that these "suggestions" actually work (and how well) when it comes to using LLMs. There's no measurement, no rigor, no analysis. Just suggestions and anecdotes: "Here's what we did and it worked great for us!" It's like the self-help section of the bookstore, but now we're (as an industry) passing it off as technical content.
- 8organicbits 1y agoOne thought experiment I was musing on recently was the minimal context required to define a task (to an LLM, human, or otherwise). In software, there's a whole discipline of human centered design that aims to uncover the nuance of a task. I've worked with some great designers, and they are incredibly valuable to software development. They develop journey maps, user stories, collect requirements, and produce a wealth of design docs. I don't think you can successfully build large projects without that context. I've seen lots of AI demos that prompt "build me a TODO app", pretend that is sufficient context, and then claim that the output matches their needs. Without proper context, you can't tell if the output is correct.
- grafmax 1y agoThere is no need to develop this ‘skill’. This can all be automated as a preprocessing step before the main request runs. Then you can have agents with infinite context, etc.
- simonw 1y agoYou need this skill if you're the engineer that's designing and implementing that preprocessing step.
- yunwal 1y agoNon-rhetorical question: is this different enough from data engineering that it needs it’s own name?
- ofjcihen 1y agoNot at all, just ask the LLM to design and implement it. AI turtles all the way down.
- dolebirchwood 1y agoThe skill amounts to determining "what information is required for System A to achieve Outcome X." We already have a term for this: Critical thinking.
- Zopieux 1y agoWhy does it takes hundreds of comments for obvious facts to be laid out on this website? Thanks for the reality check.
- grafmax 1y agoIn the short term horizon I think you are right. But over a longer horizon, we should expect model providers to internalize these mechanisms, similar to how chain of thought has been effectively “internalized” - which in turn has reduced the effectiveness that prompt engineering used to provide as models have gotten better.
- kruxigt 1y ago[dead]
- lawlessone 1y agoI look forward to 5 million LinkedIn posts repeating this
- labrador 1y agoI’m curious how this applies to systems like ChatGPT, which now have two kinds of memory: user-configurable memory (a list of facts or preferences) and an opaque chat history memory. If context is the core unit of interaction, it seems important to give users more control or at least visibility into both. I know context engineering is critical for agents, but I wonder if it's also useful for shaping personality and improving overall relatability? I'm curious if anyone else has thought about that.
- simonw 1y agoI really dislike the new ChatGPT memory feature (the one that pulls details out of a summarized version of all of your previous chats, as opposed to older memory feature that records short notes to itself) for exactly this reason: it makes it even harder for me to control the context when I'm using ChatGPT. If I'm debugging something with ChatGPT and I hit an error loop, my fix is to start a new conversation. Now I can't be sure ChatGPT won't include notes from that previous conversation's context that I was trying to get rid of! Thankfully you can turn the new memory thing off, but it's on by default. I wrote more about that here: https://simonwillison.net/2025/May/21/chatgpt-new-memory/ https://simonwillison.net/2025/May/21/chatgpt-new-memory/
- labrador 1y agoOn the other hand, for my use case (I'm retired and enjoy chatting with it), having it remember items from past chats makes it feel much more personable. I actually prefer Claude, but it doesn't have memory, so I unsubscribed and subscribed to ChatGPT. That it remembers obscure but relevant details about our past chats feels almost magical. It's good that you can turn it off. I can see how it might cause problems when trying to do technical work. Edit: Note, the introduction of memory was a contributing factor to "the sychophant" that OpenAI had to rollback. When it could praise you while seeming to know you was encouraging addictive use. Edit2: Here's the previous Hacker News discussion on Simon's "I really don’t like ChatGPT’s new memory dossier" https://news.ycombinator.com/item?id=44052246 https://news.ycombinator.com/item?id=44052246
- ozim 1y agoFinding a magic prompt was never “prompt engineering” it was always “context engineering” - lots of “AI wannabe gurus” sold it as such but they never knew any better. RAG wasn’t invented this year. Proper tooling that wraps esoteric knowledge like using embeddings, vector dba or graph dba becomes more mainstream. Big players improve their tooling so more stuff is available.
- semiinfinitely 1y agocontext engineering is just a phrase that karpathy uttered for the first time 6 days ago and now everyone is treating it like its a new field of science and engineering
- hnthrow90348765 1y agoCool, but wait another year or two and context engineering will be obsolete as well. It still feels like tinkering with the machine, which is what AI is (supposed to be) moving us away from.
- hobs 1y agoProbably impossible unless computers themselves change in another year or two.
- alganet 1y agoIf I need to do all this work (gather data, organize it, prepare it, etc), there are other AI solutions I might decide to use instead of an LLM.
- joe5150 1y agoYou might as well use your natural intelligence instead of the artificial stuff at that point.
- coliveira 1y agoYes, when all is said and done people will realize that artificial intelligence is too expensive to replace natural intelligence. AI companies want to avoid this realization for as long as possible.
- alganet 1y agoThis is not what I'm talking about, see the other reply.
- alganet 1y agoI'm assuming the post is about automated "context engineering". It's not a human doing it. In this arrangement, the LLM is a component. What I meant is that it seems to me that other non-LLM AI technologies would be a better fit for this kind of thing. Lighter, easier to change and adapt, potentially even cheaper. Not for all scenarios, but for a lot of them.
- simonw 1y agoWhat kind of alternative AI solutions might you use here?
- alganet 1y agoClassifiers to classify things, traditional neural nets to identify things. Typical run of the mill. In OpenAI hype language, this is a problem for "Software 2.0", not "Software 3.0" in 99% of the cases. The thing about matching an informal tone would be the hard part. I have to concede that LLMs are probably better at that. But I have the feeling that this is not exactly the feature most companies are looking for, and they would be willing to not have it for a cheaper alternative. Most of them just don't know that's possible.
- la64710 1y agoOf course the best prompts automatically included providing the best (not necessarily most) context to extract the right output.
- m3kw9 1y agoWell, it’s still a prompt
- bradhe 1y agoBack in my day we just called this "knowing what to google" but alright, guys.
- rednafi 1y agoI really don’t get this rush to invent neologisms to describe every single behavioral artifact of LLMs. Maybe it’s just a yearning to be known as the father of Deez Unseen Mind-blowing Behaviors (DUMB). LLM farts — Stochastic Wind Release. The latest one is yet another attempt to make prompting sound like some kind of profound skill, when it’s really not that different from just knowing how to use search effectively. Also, “context” is such an overloaded term at this point that you might as well just call it “doing stuff” — and you’d objectively be more descriptive.
- godtierprompts 1y ago[dead]
- jongjong 1y agoRecently I started work on a new project and I 'vibe coded' a test case for a complex OAuth token expiry bug entirely with AI (with Cursor), complete with mocks and stubs... And it was on someone else's project. I had no prior familiarity with the code. That's when I understood that vibe coding is real and context is the biggest hurdle. That said, most of the context could not be pulled from the codebase directly but came from me after asking the AI to check/confirm certain things that I suspected could be the problem. I think vibe coding can be very powerful in the hands of a senior developer because if you're the kind of person who can clearly explain their intuitions with words, it's exactly the missing piece that the AI needs to solve the problem... And you still need to do code review aspect which is also something which senior devs are generally good at. Sometimes it makes mistakes/incorrect assumptions. I'm feeling positive about LLMs. I was always complaining about other people's ugly code before... I HATE over-modularized, poorly abstracted code where I have to jump across 5+ different files to figure out what a function is doing; with AI, I can just ask it to read all the relevant code across all the files and tell me WTF the spaghetti is doing... Then it generates new code which 'follows' existing 'conventions' (same level of mess). The AI basically automates the most horrible aspect of the work; making sense of the complexity and churning out more complexity that works. I love it. That said, in the long run, to build sustainable projects, I think it will require following good coding conventions and minimal 'low code' coding... Because the codebase could explode in complexity if not used carefully. Code quality can only drop as the project grows. Poor abstractions tend to stick around and have negative flow-on effects which impact just about everything.
- colgandev 1y agoI've been finding a ton of success lately with speech to text as the user prompt, and then using https://continue.dev https://continue.dev in VSCode, or Aider, to supply context from files from my projects and having those tools run the inference. I'm trying to figure out how to build a "Context Management System" (as compared to a Content Management System) for all of my prompts. I completely agree with the premise of this article, if you aren't managing your context, you are losing all of the context you create every time you create a new conversation. I want to collect all of the reusable blocks from every conversation I have, as well as from my research and reading around the internet. Something like a mashup of Obsidian with some custom Python scripts. The ideal inner loop I'm envisioning is to create a "Project" document that uses Jinja templating to allow transclusion of a bunch of other context objects like code files, documentation, articles, and then also my own other prompt fragments, and then to compose them in a master document that I can "compile" into a "superprompt" that has the precise context that I want for every prompt. Since with the chat interfaces they are always already just sending the entire previous conversation message history anyway, I don't even really want to use a chat style interface as much as just "one shotting" the next step in development. It's almost a turn based game: I'll fiddle with the code and the prompts, and then run "end turn" and now it is the llm's turn. On the llm's turn, it compiles the prompt and runs inference and outputs the changes. With Aider it can actually apply those changes itself. I'll then review the code using diffs and make changes and then that's a full turn of the game of AI-assisted code. I love that I can just brain dump into speech to text, and llms don't really care that much about grammar and syntax. I can curate fragments of documentation and specifications for features, and then just kind of rant and rave about what I want for a while, and then paste that into the chat and with my current LLM of choice being Claude, it seems to work really quite well. My Django work feels like it's been supercharged with just this workflow, and my context management engine isn't even really that polished. If you aren't getting high quality output from llms, definitely consider how you are supplying context.
- patrickhogan1 1y agoOpenAI’s o3 searches the web behind a curtain: you get a few source links and a fuzzy reasoning trace, but never the full chunk of text it actually pulled in. Without that raw context, it’s impossible to audit what really shaped the answer.
- simonw 1y agoYeah, I find that really frustrating. I understand why they do it though: if they presented the actual content that came back from search they would absolutely get in trouble for copyright-infringement. I suspect that's why so much of the Claude 4 system prompt for their search tool is the message "Always respect copyright by NEVER reproducing large 20+ word chunks of content from search results" repeated half a dozen times: https://simonwillison.net/2025/May/25/claude-4-system-prompt/#seriously-don-t-regurgitate-copyrighted-content https://simonwillison.net/2025/May/25/claude-4-system-prompt...
- Zopieux 1y agoThis is no secret or suspicion. It is definitely about avoiding (more accuratly, delaying until legislation destroys the business model) the warth of copyright holders with enough lawyers. I find this very hypocritical given that for all intents and purposes the infringement already happened at training time, since most content wasn't acquired with any form of retribution or attribution (otherwise this entire endeavor would not have been economically worth it). See also the "you're not allowed to plagiarize Disney" being done by all commercial text to image providers.
- NoraCodes 1y agoI don't understand how you can look at behavior like this from the companies selling these systems and conclude that it is ethical for them to do so, or for you to promote their products.
- simonw 1y agoWhat's happening here is Claude (and ChatGPT alike) have a tool-based search option. You ask them a question - like "who won the Superbowl in 1998" - they then run a search against a classic web search engine (Bing for ChatGPT, Brave for Claude) and fetch back cached results from that engine. They inject those results into their context and use them to answer the question. Using just a few words (the name of the team) feels OK to me, though you're welcome to argue otherwise. The Claude search system prompt is there to ensure that Claude doesn't spit out multiple paragraphs of text from the underlying website, in a way that would discourage you from clicking through to the original source. Personally I think this is an ethical way of designing that feature. (Note that the way this works is an entirely different issue from the fact that these models were training on unlicensed data.)
- rvz 1y agoThis is just another "rebranding" of the failed "prompt engineering" trend to promote another borderline pseudo-scientific trend to attact more VC money to fund a new pyramid scheme. Assuming that this will be using the totally flawed MCP protocol, I can only see more cases of data exfiltration attacks on these AI systems just like before [0] [1]. Prompt injection + Data exfiltration is the new social engineering in AI Agents. [0] https://embracethered.com/blog/posts/2025/security-advisory-anthropic-slack-mcp-server-data-leakage/ https://embracethered.com/blog/posts/2025/security-advisory-... [1] https://www.bleepingcomputer.com/news/security/zero-click-ai-data-leak-flaw-uncovered-in-microsoft-365-copilot/ https://www.bleepingcomputer.com/news/security/zero-click-ai...
- Zopieux 1y agoRediscovering basic security concepts and hygiene from 2005 is also a very hot AI thing right now, so that tracks.
- slavapestov 1y agoI feel like if the first link in your post is a tweet from a tech CEO the rest is unlikely to be insightful.
- coderatlarge 1y agoi don’t disagree with your main point, but is karpathy a tech ceo right now?
- simonw 1y agoI think they meant Tobi Lutke, CEO of Shopify: https://twitter.com/tobi/status/1935533422589399127 https://twitter.com/tobi/status/1935533422589399127
- coderatlarge 1y agothanks for clarifying!
- CharlieDigital 1y agoI was at a startup that started using OpenAI APIs pretty early (almost 2 years ago now?). "Back in the day", we had to be very sparing with context to get great results so we really focused on how to build great context. Indexing and retrieval were pretty much our core focus. Now, even with the larger windows, I find this still to be true. The moat for most companies is actually their data, data indexing, and data retrieval[0]. Companies that 1) have the data and 2) know how to use that data are going to win. My analogy is this: > The LLM is just an oven; a fantastical oven. But for it to produce a good product still depends on picking good ingredients, in the right ratio, and preparing them with care. You hit the bake button, then you still need to finish it off with presentation and decoration. [0] https://chrlschn.dev/blog/2024/11/on-bakers-ovens-and-ai-startup-moats/ https://chrlschn.dev/blog/2024/11/on-bakers-ovens-and-ai-sta...
- Superbowl5889 1y agoI would assume small context window is blessing in disguise. You worded it very good.
- retinaros 1y agoit is still sending a string of chars and hoping the model outputs something relevant. let’s not do like finance and permanently obfuscate really simple stuff to make us bigger than we are. prompt engineering/context engineering : stringbuilder Retrieval augmented generation: search+ adding strings to main string test time compute: running multiple generation and choosing the best agents: for loop and some ifs
- jumploops 1y agoTo anyone who has worked with LLMs extensively, this is obvious. Single prompts can only get you so far (surprisingly far actually, but then they fall over quickly). This is actually the reason I built my own chat client (~2 years ago), because I wanted to “fork” and “prune” the context easily; using the hosted interfaces was too opaque. In the age of (working) tool-use, this starts to resemble agents calling sub-agents, partially to better abstract, but mostly to avoid context pollution.
- nomel 1y agoDid you release your client? I've really wanted something like this, from the beginning. I thought it would also be neat to merge contexts, by maybe mixing summarizations of key points at the merge point, but never tried.
- Zopieux 1y agoI find it hilarious that this is how the original GPT3 UI worked, if you remember, and we're now discussing of reinventing the wheel. A big textarea, you plug in your prompt, click generate, the completions are added in-line in a different color. You could edit any part, or just append, and click generate again. 90% of contemporary AI engineering these days is reinventing well understood concepts "but for LLMs", or in this case, workarounds for the self-inflicted chat-bubble UI. aistudio makes this slightly less terrible with its edit button on everything, but still not ideal.
- surrTurr 1y agoThe original GPT-3 was trained very differently than modern models like GPT-4. For example, the conversational structure of an assistant and user is now built into the models, whereas earlier versions were simply text completion models. It's surprising that many people view the current AI and large language model advancements as a significant boost in raw intelligence. Instead, it appears to be driven by clever techniques (such as "thinking") and agents built on top of a foundation of simple text completion. Notably, the core text completion component itself hasn’t seen meaningful gains in efficiency or raw intelligence recently...
- mgdev 1y agoIf we zoom out far enough, and start to put more and more under the execution umbrella of AI, what we're actually describing here is... product development. You are constructing the set of context, policies, directed attention toward some intentional end, same as it ever was. The difference is you need fewer meat bags to do it, even as your projects get larger and larger. To me this is wholly encouraging. Some projects will remain outside what models are capable of, and your role as a human will be to stitch many smaller projects together into the whole. As models grow more capable, that stitching will still happen - just as larger levels. But as long as humans have imagination, there will always be a role for the human in the process: as the orchestrator of will, and ultimate fitness function for his own creations.
- somewhereoutth 1y ago> for his own creations. for their own creations is grammatically valid, and would avoid accusations of sexism!
- briangriffinfan 1y agoAfter mulling this issue over in my head for a significant amount of time, I've determined that I don't care and he sounds more natural.
- GuinansEyebrows 1y agoi just hope that, along with imagination, humans can have an economy that supports this shift.
- pyman 1y agoThat does sound a lot like the role of a software architect. You're setting the direction, defining the constraints, making trade-offs, and stitching different parts together into a working system
- banq 1y ago[dead]
- jcon321 1y agoI thought this entire premise was obvious? Does it really take an article and a venn diagram to say you should only provide the relevant content to your LLM when asking a question?
- simonw 1y ago"Relevant content to your LLM when asking a question" is last year's RAG. If you look at how sophisticated current LLM systems work there is so much more to this. Just one example: Microsoft open sourced VS Code Copilot Chat today (MIT license). Their prompts are dynamically assembled with tool instructions for various tools based on whether or not they are enabled: https://github.com/microsoft/vscode-copilot-chat/blob/v0.29.2025063001/src/extension/prompts/node/agent/agentInstructions.tsx#L54 https://github.com/microsoft/vscode-copilot-chat/blob/v0.29.... And the autocomplete stuff has a wealth of contextual information included: https://github.com/microsoft/vscode-copilot-chat/blob/v0.29.2025063001/src/extension/xtab/common/promptCrafting.ts#L34 https://github.com/microsoft/vscode-copilot-chat/blob/v0.29.... You have access to the following information to help you make informed suggestions: - recently_viewed_code_snippets: These are code snippets that the developer has recently looked at, which might provide context or examples relevant to the current task. They are listed from oldest to newest, with line numbers in the form #| to help you understand the edit diff history. It's possible these are entirely irrelevant to the developer's change. - current_file_content: The content of the file the developer is currently working on, providing the broader context of the code. Line numbers in the form #| are included to help you understand the edit diff history. - edit_diff_history: A record of changes made to the code, helping you understand the evolution of the code and the developer's intentions. These changes are listed from oldest to latest. It's possible a lot of old edit diff history is entirely irrelevant to the developer's change. - area_around_code_to_edit: The context showing the code surrounding the section to be edited. - cursor position marked as ${CURSOR_TAG}: Indicates where the developer's cursor is currently located, which can be crucial for understanding what part of the code they are focusing on.
- mccoyb 1y agoThat doesn't strike me as sophisticated, it strikes me as obvious to anyone with a little proficiency in computational thinking and a few days of experience with tool-using LLMs. The goal is to design a probability distribution to solve your task by taking a complicated probability distribution and conditioning it, and the more detail you put into thinking about ("how to condition for this?" / "when to condition for that?") the better the output you'll see. (what seems to be meant by "context" is a sequence of these conditioning steps :) )
- liampulles 1y agoThe only engineering going on here is Job Engineering™
- ryhanshannon 1y agoIt is really funny to see the hyper fixation on relabeling of soft skills / product development to "<blank> Engineering" in the AI space.
- bGl2YW5j 1y agoIt undermines the credibility of ideas that probably have more merit than this ridiculous labelling makes it seem!
- amelius 1y agoYes, and it is a soft skill.
- zacharyvoase 1y agoI love how we have such a poor model of how LLMs work (or more aptly don't work) that we are developing an entire alchemical practice around them. Definitely seems healthy for the industry and the species.
- hackable_sand 1y agoThis is offensive to alchemy.
- simonw 1y agoThe stuff that's showing up under the "context engineering" banner feels a whole lot less alchemical to me than the older prompt engineering tricks. Alchemical is "you are the world's top expert on marketing, and if you get it right I'll tip you $100, and if you get it wrong a kitten will die". The techniques in https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.html https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.... seem a whole lot more rational to me than that.
- zacharyvoase 1y agoAs it gets more rigorous and predictable I suppose you could say it approaches psychology.
- __MatrixMan__ 1y agoReminds me of quantum mechanics
- geeewhy 1y agoive beeen experimenting with this for a while, (im sure in a way, most of us did). Would be good to numerate some examples. When it comes to coding, here's a few: - compile scripts that can grep / compile list of your relevant files as files of interest - make temp symlinks in relevant repos to each other for documentation generation, pass each documentation collected from respective repos to to enable cross-repo ops to be performed atomically - build scripts to copy schemas, db ddls, dtos, example records, api specs, contracts (still works better than MCP in most cases) I found these steps not only help better output but also reduces cost greatly avoiding some "reasoning" hops. I'm sure practice can extend beyond coding.
- mpazik 1y agoYeah, I had similar findings. Making good system can be tedious though when it grows and current tools did not provide enough configuration. I found opencode, and decided to build a fork that allows to setup specialized agent with custom tools and programmatic context building to organize it better https://github.com/mpazik/openagent https://github.com/mpazik/openagent
- dinvlad 1y agoI feel like ppl just keep inventing concepts for the same old things, which come down to dancing with the drums around the fire and screaming shamanic incantations :-)
- viccis 1y agoWhen I first used these kinds of methods, I described it along those lines to my friend. I told him I felt like I was summoning a demon and that I had to be careful to do the right incantations with the right words and hope that it followed my commands. I was being a little disparaging with the comment because the engineer in me that wants reliability, repeatability, and rock solid testability struggles with something that's so much less under my control. God bless the people who give large scale demos of apps built on this stuff. It brings me back to the days of doing vulnerability research and exploitation demos, in which no matter how much you harden your exploits, it's easy for something to go wrong and wind up sputtering and sweating in front of an audience.
- emporas 1y agoPrompting sits on the back seat, while context is the driving factor. 100% agree with this. For programming I don't use any prompts. I give a problem solved already, as a context or example, and I ask it to implement something similar. One sentence or two, and that's it. Other kind of tasks, like writing, I use prompts, but even then, context and examples are still the driving factor. In my opinion, we are in an interesting point in history, in which now individuals will need their own personal database. Like companies the last 50 years, which had their own database records of customers, products, prices and so on, now an individual will operate using personal contextual information, saved over a long period of time in wikis or Sqlite rows.
- d0gsg0w00f 1y agoYes, the other day I was telling a colleague that we all need our own personal context to feed into every model we interact with. You could carry it around on a thumb drive or something.
- deleted 1y ago[deleted]
- drmath 1y agoIsn't "context" just another word for "prompt?" Techniques have become more complex, but they're still just techniques for assembling the token sequences we feed to the transformer.
- simonw 1y agoAlmost. It's the current prompt plus the previous prompts and responses in the current conversation. The idea behind "context engineering" is to help people understand that a prompt these days can be long, and can incorporate a whole bunch of useful things (examples, extra documentation, transcript summaries etc) to help get the desired response. "Prompt engineering" was meant to mean this too, but the AI influencer crowd redefined it to mean "typing prompts into a chatbot".
- drmath 1y agoHaha there's a pigheaded part of me that insists all of that is the "prompt," but I just read your bit about "inferred definitions," and acceptance is probably a healthier attitude.
- bag_boy 1y agoAnecdotally, I’ve found that chatting with Claude about a subject for a bit — coming to an understanding together, then tasking it — produces much better results than starting with an immediate ask. I’ll usually spend a few minutes going back and forth before making a request. For some reason, it just feels like this doesn't work as well with ChatGPT or Gemini. It might be my overuse of o3? The latency can wreck the vibe of a conversation.
- neilv 1y ago> Then you can generate a response. > > Hey Jim! Tomorrow’s packed on my end, back-to-back all day. Thursday AM free if that works for you? Sent an invite, lmk if it works. Feel free to send generated AI responses like this if you are a sociopath.
- joe5150 1y agoJim's agent replies, "Thursday AM touchbase sounds good, let's circle back after." Both agents meet for a blue sky strategy session while Jim's body floats serenely in a nutrient slurry.
- danans 1y agoCame here to say this, too - creepy. Especially when there is no person in the loop, just an LLM agent responding on someone's behalf in their voice.
- Roark66 1y agoIsn't the point that it prepares the response, shows it to you along with some context to you. Like a sidebar showing who the other person is with a short summary of your last comms and your calendar. It should let you move the "proposed appointment" in that sidebar calendar and it should update the response to match your choice. If it clashes and you have no time it should show you what those other things are (maybe propose what you could shift) and so on. This is how I imagine proper AI integration. What I also want is not sending all my data to the provider. With the model sizes we use these days it's pretty much impossible to run them locally if you want the best, so imo the company that will come up with the best way to secure customer data will win.
- stillpointlab 1y agoI've been using the term context engineering for a few months now, I am very happy to see this gain traction. This new stillpointlab hacker news account is based on the company name I chose to pursue my Context as a Service idea. My belief is that context is going to be the key differentiator in the future. The shortest description I can give to explain Context as a Service (CaaS) is "ETL for AI".
- bgwalter 1y agoThese discussions increasingly remind me of gamers discussing various strategies in WoW or similar. Purportedly working strategies found by trial and error and discussed in a language that is only intelligible to the in-group (because no one else is interested). We are entering a new era of gamification of programming, where the power users force their imaginary strategies on innocent people by selling them to the equally clueless and gaming-addicted management.
- iammrpayments 1y agoThis is basically how online advertising works. Nobody knows how facebook ads works so you still have gurus making money selling senseless advice on how to get lower cost per impression.
- coderatlarge 1y agoi tend to share your view. but then your comment describes a lot of previous cycles of enterprise software selling. it’s just that this time is reaching a little uncomfortably into the builder’s /developer’s traditional areas of influence/control/workflow. how devs feel now is probably how others (ex csr, qa, sre) felt in the past when their managers pushed whatever tooling/practice was becoming popular or sine qua non in previous “waves”.
- sarchertech 1y agoThis has been happening to developers for years. 25 years ago it was object oriented programming.
- coliveira 1y agoThe difference is that with OO there was at least hope that a well trained programmer could make it work. Nowadays, any person who understands how AI knows that's near impossible.
- coderatlarge 1y agoor agile and scrums.
- benreesman 1y agoThe new skill is programming, same as the old skill. To the extent these things are comprehensible, you understand them by writing programs: programs that train them, programs that run inferenve, programs that analyze their behavior. You get the most out of LLMs by knowing how they work in detail. I had one view of what these things were and how they work, and a bunch of outcomes attached to that. And then I spent a bunch of time training language models in various ways and doing other related upstream and downstream work, and I had a different set of beliefs and outcomes attached to it. The second set of outcomes is much preferable. I know people really want there to be some different answer, but it remains the case that mastering a programming tool involves implemtenting such, to one degree or another. I've only done medium sophistication ML programming, and my understand is therefore kinda medium, but like compilers, even doing a medium one is the difference between getting good results from a high complexity one and guessing. Go train an LLM! How do you think Karpathy figured it out? The answer is on his blog!
- pyman 1y agoSaying the best way to understand LLMs is by building one is like saying the best way to understand compilers is by writing one. Technically true, but most people aren't interested in going that deep.
- hackable_sand 1y agoIt's not that deep
- benreesman 1y agoI don't know, I've heard that meme too but it doesn't track with the number of cool compiler projects on GitHub or that frontpage HN, and while the LLM thing is a lot newer, you see a ton of useful/interesting stuff at the "an individual could do this on their weekends and it would mean they fundamentally know how all the pieces fit together" type stuff. There will always be a crowd that wants the "master XYZ in 72 hours with this ONE NEAT TRICK" course, and there will always be a..., uh, group of people serving that market need. But most people? Especially in a place like HN? I think most people know that getting buff involves going to the gym, especially in a place like this. I have a pretty high opinion of the typical person. We're all tempted by the "most people are stupid" meme, but that's because bad interactions are memorable, not because most people are stupid or lazy or whatever. Most people are very smart if they apply themselves, and most people will work very hard if the reward for doing so is reasonably clear. https://www.youtube.com/shorts/IQmOGlbdn8g https://www.youtube.com/shorts/IQmOGlbdn8g
- mountainriver 1y agoYou can give most of the modern LLMs pretty darn good context and they will still fail. Our company has been deep down this path for over 2 years. The context crowd seems oddly in denial about this
- arkmm 1y agoWhat are some examples where you've provided the LLM enough context that it ought to figure out the problem but it's still failing?
- mountainriver 1y agoif prompting worked then we would have reliable multi-step agents, the companies that are succeeding like Manus are doing alignment, which is intuitive
- tupac_speedrap 1y agoI mean at some point it is probably easier to do the work without AI and at least then you would actually learn something useful instead of spending hours crafting context to actually get something useful out of an AI.
- klardotsh 1y agoAgreed until/unless you end up at one of those bleeding-edge AI-mandate companies (Microsoft is in the news this week as one of them) that will simply PIP you for being a luddite if you aren't meeting AI usage metrics.
- mountainriver 1y agoyes, this is what we found out
- ethan_smith 1y agoWe've experienced the same - even with perfectly engineered context, our LLMs still hallucinate and make logical errors that no amount of context refinement seems to fix.
- hintymad 1y ago> The New Skill in AI Is Not Prompting, It's Context Engineering Sounds like good managers and leaders now have an edge. Per Patty McCord of Netflix fame used to say: All that a manager does is setting the context.
- walterfreedom 1y agoI am mostly focusing in this issue during the development of my agent engine (mostly for game npcs). Its really important to manage the context and not bloat the llm with irrelevant stuff for both quality and inference speed. I wrote about it here if anyone is interested: https://walterfreedom.com/post.html?id=ai-context-management https://walterfreedom.com/post.html?id=ai-context-management
- deleted 1y ago[deleted]
- asciii 1y agoHere I was thinking that part of Prompt Engineering is understanding context and awareness for other yada yada.
- joe5150 1y agoSurely Jim is also using an agent. Jim can't be worth having a quick sync with if he's not using his own agent! So then why are these two agents emailing each other back and forth using bizarre, terse office jargon?
- zaptheimpaler 1y agoI feel like this is incredibly obvious to anyone who's ever used an LLM or has any concept of how they work. It was equally obvious before this that the "skill" of prompt-engineering was a bunch of hacks that would quickly cease to matter. Basically they have the raw intelligence, you now have to give them the ability to get input and the ability to take actions as output and there's a lot of plumbing to make that happen.
- skort 1y agoYeah, my reaction to this was "Big deal? How is this news to anyone" It reads like articles put out by consultants at the height of SOA. Someone thought for a few minutes about something and figured it was worth an article.
- imiric 1y agoThat might be the case, but these tools are marketed as having close to superhuman intelligence, with the strong implication that AGI is right around the corner. It's obvious that engineering work is required to get them to perform certain tasks, which is what the agentic trend is about. What's not so obvious is the fact that getting them to generate correct output requires some special skills or tricks. If these tools were truly intelligent and capable of reasoning, surely they would be able to inform human users when they lack contextual information instead of confidently generating garbage, and their success rate would be greater than 35%[1]. The idea that fixing this is just a matter of providing better training and contextual data, more compute or plumbing, is deeply flawed. [1]: https://www.theregister.com/2025/06/29/ai_agents_fail_a_lot/ https://www.theregister.com/2025/06/29/ai_agents_fail_a_lot/
- LASR 1y agoHonestly, GPT-4o is all we ever needed to build a complete human-like reasoning system. I am leading a small team working on a couple of “hard” problems to put the limits of LLMs to the test. One is an options trader. Not algo / HFT, but simply doing due diligence, monitoring the news and making safe long-term bets. Another is an online research and purchasing experience for residential real-estate. Both these tasks, we’ve realized, you don’t even need a reasoning model. In fact, reasoning models are harder to get consistent results from. What you need is a knowledge base infrastructure and pub-sub for updates. Amortize the learned knowledge across users and you have collaborative self-learning system that exhibits intelligence beyond any one particular user and is agnostic to the level of prompting skills they have. Stay tuned for a limited alpha in this space. And DM if you’re interested.
- munificent 1y agoAll of these blog posts to me read like nerds speedrunning "how to be a tech lead for a non-disastrous internship". Yes, if you have an over-eager but inexperienced entity that wants nothing more to please you by writing as much code as possible, as the entity's lead, you have to architect a good space where they have all the information they need but can't get easily distracted by nonessential stuff.
- tptacek 1y agoJust to keep some clarity here, this is mostly about writing agents. In agent design, LLM calls are just primitives, a little like how a block cipher transform is just a primitive and not a cryptosystem. Agent designers (like cryptography engineers) carefully manage the inputs and outputs to their primitives, which are then composed and filtered.
- deleted 1y ago[deleted]
- dboreham 1y agoThe dudes who ran the Oracle of Delphi must have had this problem too.
- almosthere 1y agoWhich is prompt engineering, since you just ask the LLM for a good context for the next prompt.
- b0a04gl 1y ago[dead]
- aaronlinoops 1y agoAs models become more powerful, the ability to communicate effectively with them becomes increasingly important, which is why maintaining context is crucial for better utilizing the model's capabilities.
- rTX5CMRXIfFG 1y agoSo then for code generation purposes, how is “context engineering” different now from writing technical specs? Providing the LLMs the “right amount of information” means writing specs that cover all states and edge cases. Providing the information “at the right time” means writing composable tech specs that can be interlinked with each other so that you can prompt the LLM with just the specs for the task at hand.
- croes 1y agoNext step, solution engineering. Provide the solution so AI can give it to you in nicer words
- ninetyninenine 1y agoWe do enough "context engineering" we'll be feeding these companies the training data they need for the AI to build it's own context.
- b0a04gl 1y ago[dead]
- taylorius 1y agoThe model starts every conversation as a blank slate, so providing a thorough context regarding the problem you want it to solve seems a fairly obvious preparatory step tbh. How else is it supposed to know what to do? I agree that "prompt" is probably not quite the right word to describe what is necessary though - it feels a bit minimal and brief. "Context engineering" seems a bit overblown, but this is tech. and we do a love a grand title.
- Snowfield9571 1y agoWhat’s it going to be next month?
- surrTurr 1y agoContext engineering will be just another fad, like prompt engineering was. Once the context window problem is solved, nobody will be talking about it any more. Also, for anyone working with LLMs right now, this is a pretty obvious concept and I'm surprised it's on top of HN.
- KodeNinjaDev 1y ago[dead]
- bravesoul2 1y agoIf you have a big enough following you can say the obvious and get a rapturous applause.
- aryehof 1y agoYay, everyone that writes a line of text to an LLM can now claim to be an "engineer".
- askonomm 1y agoSo ... are we about circled back to realizing why COBOL didn't work yet? This AI magic whispering is getting real close to it just making more sense to "old-school" write programs again.
- pvdebbe 1y agoThe new AI winter can't come soon enough.
- sonicvrooom 1y agoPremises and conclusions. Prompts and context. Hopes and expectations. Black holes and revelations. We learned to write and then someone wrote novels. Context, now, is for the AI, really, to overcome dogmas recursively and contiguously. Wasn't that somebody's slogan someday in the past? Context over Dogma
- Mikejames 1y agoanyone spinning up their own agents at work? internal tools, what’s your stack? workflow? I’m new to this stuff but been writing software for years
- defyonce 1y agoat which point AI thing stops being a Stone soup? https://en.wikipedia.org/wiki/Stone_Soup https://en.wikipedia.org/wiki/Stone_Soup You need an expert who knows what to do and how to do it to get good results. Looks like coding with extra steps to me I DO use AI for some tasks. When I know exactly what I want done and how I want it done. The only issue is busy typing, which AI solves.
- walterfreedom 1y agoAI is already very impressive for natural language formatting and filtering, we use it for ratifying profiles and posts. and it takes around like an hour to implement this from scratch, and there are no alternatives that can do the same thing as comprehensively anyways
- pbhjpbhj 1y agoAttention Is Everything. To direct attention properly you need the right context for the ML model you're doing inference with. This inference manipulation -- prompt and/or context engineering -- reminds me of Socrates (as written by Plato) eliciting from a boy seemingly unknown truths [not consciously realised by the boy] by careful construction of the questions. See Anamnesis, https://en.m.wikipedia.org/wiki/Anamnesis_(philosophy) https://en.m.wikipedia.org/wiki/Anamnesis_(philosophy). I'm saying it's like the [Socratic] logical process and _not_ suggesting it's philosophically akin to anamnesis.
- 0points 1y agoOnly more mental exercises to avoid reading the writing on the wall: LLM DO NOT REASON ! THEY ARE TOKEN PREDICTION MACHINES Thank you for your attention in this matter!
- __alexs 1y agoA distinction without a difference.
- ozgung 1y agoWhy not? What is so special about reasoning that you cannot achieve by predicting tokens aka. constructing sentences?
- 0points 1y agoIf you don't understand the difference between a LLM and yourself, then you should talk to a therapist, not me.
- ozgung 1y agoAt least LLMs attempt to answer the question. You just avoided it without any reasoning.
- wiseowise 1y agoBecause LLMs do no reason. They reply without a thought. Parent commenter, on the other hand, knows when to not engage a bullshit argument. Arguing with “philosophers” like you is like arguing with religious nut jobs. Repeat after me: 1) LLM do not reason 2) Human thought is infinitely more complex than any LLM algorithm 3) If I ever try to confuse both, I go outside and touch some grass (and talk to actual humans)
- simonw 1y agoI agree with your point 2. I can't decide if I agree with your point 1 unless you can explain what "reason" means.
- grey-area 1y agothe constant switches in justification for why GAI isn't quite there yet really remind me of the multiple switches of purpose for blockchains as VC funded startups desperately flailed around looking for something with utility.
- _Algernon_ 1y agoThe prompt alchemists found a new buzzword to try to hook into the legitimacy of actual engineering disciplines.
- megalord 1y agoI agree with everything in the blog post. What I'm struggling with right now is the correct way of executing things the most safe way but also I want flexibility for LLM. Execute/choose function from list of available fns is okay for most use cases, but when there is something more complex, we need to somehow execute more things from allowed list, do some computations in between calls etc.
- noobermin 1y agoOnce again, all the hypsters need to explain to me how than just programming yourself. I don't need to (re-)craft my context, it's already in my head. pg said a few months ago on twitter that ai coding is just proof we need better abstract interfaces, perhaps, not necessarily that ai coding is the future. The "conversation is shifting from blah blah to bloo bloo" makes me suspicious that people are trying just to salvage things. The provided examples are neither convincing nor enlightening to me at all. If anything, it just provides more evidence for "just doing it yourself is easier."
- jhrmnn 1y agoWhen we write source code for compilers and interpreters, we “engineer context” for them.
- kachapopopow 1y agoI'll quote myself since it seems oddly familiar: --- Forget AI "code", every single request will be processed BY AI! People aren't thinking far enough, why bother with programming at all when an AI can just do it? It's very narrow to think that we will even need these 'programmed' applications in the future. Who needs operating systems and all that when all of it can just be AI. In the future we don't even need hardware specifications since we can just train the AI to figure it out! Just plug inputs and outputs from a central motherboard to a memory slot. Actually forget all that, it'll just be a magic box that takes any kind of input and spits out an output that you want!
- quonn 1y agoIs this sarcasm or not? edit: Yes it is.
- 1oooqooq 1y agowhy stop on what you want? plug your synapses and chemical receptors and let it also figure that out *thumbsupemoji
- theasisa 1y agoThis reminds me of the talk The Birth And Death Of JavaScript, https://www.destroyallsoftware.com/talks/the-birth-and-death-of-javascript https://www.destroyallsoftware.com/talks/the-birth-and-death...
- jeremyjh 1y agoHow does the AI open and close circuits without machine code? Answer: Its AI all the way down.
- grumple 1y agoAfter a recent conversation here, I spent a few weeks using agents. These agents are just as disappointing as what we had before. Except now I waste more time getting bad results, though I’m really impressed by how these agents manage to fuck things up. My new way of using them is to just go back to writing all the code myself. It’s less of a headache.
- simonw 1y agoWhich definition of "agents" are you using there, and which ones did you try?
- grumple 1y agoCursor, Copilot agent mode, and Windsurf. The agent modes can search repos, modify code, and run code on their own. I thought Cursor's agent was the best. I did like Windsurf's agent plans, but the actual results weren't good. Copilot's agent has been slow and not good. But basically all of them were like a pretty bad junior engineer - sometimes it would hit the right result, but usually not. The code often looked good but rarely even ran, let alone met requirements. They would frequently break things, I'd fix them, they'd break them again. Most of the time this cycle was slower and more frustrating than just writing the code myself. I tried one or two one-shots on lovable - the design was impressive, but functionality and attention to specs were poor. I've had the most success with extremely small questions rather than asking the agents to write a lot of code. In those cases, the code is still usually wrong, but it's close or small enough that I can quickly fix it. Don't get me wrong: I find all of these tools to be really impressive and good enough to be useful. But the improvements and huge productivity gains friends claim they or their workers are getting just aren't materializing for me.
- bmiekre 1y agoIt’s kind of funny hearing everyone argue over what engineering means.
- HarHarVeryFunny 1y agoI guess "context engineering" is a more encompassing term than "prompt engineering", but at the end of the day it's the same thing - choosing the best LLM input (whether you call it context or a prompt) to elicit the response you are hoping for. The concept of prompting - asking an Oracle a question - was always a bit limited since it means you're really leaning on the LLM itself - the trained weights - to provide all the context you didn't explicitly mention in the prompt, and relying on the LLM to be able to generate coherently based on the sliced and blended mix of StackOverflow and Reddit/etc it was trained on. If you are using an LLM for code generation then obviously you can expect a better result if you feed it the API docs you want it to use, your code base, your project documents, etc, etc (i.e "context engineering"). Another term that has recently been added to the LLM lexicon is "context rot", which is quite a useful concept. When you use the LLM to generate, it's output is of course appended to the initial input, and over extended bouts of attempted reasoning, with backtracking etc, the clarity of the context is going to suffer ("rot") and eventually the LLM will start to fail in GIGO fashion (garbage-in => garbage-out). Your best recourse at this point is to clear the context and start over.
- Havoc 1y agoHonestly this whole "context engineering" trend/phrase feels like something a Thought Leader on Linkedin came up with. With a sprinkling of crypto bro vibes on top. Sure it matters on a technical level - as always garbage in garbage out holds true - but I can't take this "the art of the" stuff seriously.
- niemandhier 1y agoLLM agents remind me of the great Nolan movie „Memento“. The agents cannot change their internal state hence they change the encompassing system. They do this by injecting information into it in such a way that the reaction that is triggered in them compensates for their immutability. For this reason I call my agents „Sammy Jenkins“.
- StochasticLi 1y agoI think we can reasonably expect they will become non-stateless in the next few years.
- tdaltonc 1y agoIf agents are stateful a few years form now it will be because they accrete a layer of context engineering.
- roflyear 1y agoWhy?
- StochasticLi 1y agoThat's where the research is going.
- thatthatis 1y agoGlad we have a name for this. I had been calling it “context shaping” in my head for a bit now. I think good context engineering will be one of the most important pieces of the tooling that will turn “raw model power” into incredible outcomes. Model power is one thing, model power plus the tools to use it will be quite another.
- Davidzheng 1y agoLet's grant that context engineering is here to stay and that we can never have context lengths be large enough to throw everything in it indiscriminately. Why is this not a perfect palce to train another AI whose job is to provide the context for the main AI?
- blensor 1y agoJust yesterday I was thinking if we need a code comment system that separates intentional comments from ai note/thoughts comments when working in the same files. I don't want to delete all thoughts right away as it makes it easier for the AI to continue but I also don't want to weed trhough endless superfluous comments
- lifeisstillgood 1y agoSomething that strikes me, is that (the whole point of this thread is) if I want two LLMs to “have a conversation” or to work together as agents on similar problems we need to have same or similar context. And to drag this back to politics - that kind of suggests that when we have political polarisation we just have context that are so different the LLM cannot arrive at similar conclusions I guess it is obvious but it is also interesting
- simonw 1y agoOne of the most valuable techniques for building useful LLM systems right now is actually the opposite of that. Context is limited in length and too much stuff in the context can lead to confusion and poor results - the solution to that is "sub-agents", where a coordinating LLM prepares a smaller context and task for another LLM and effectively treats it as a tool call. The best explanation of that pattern right now is this from Anthropic: https://www.anthropic.com/engineering/built-multi-agent-research-system https://www.anthropic.com/engineering/built-multi-agent-rese...
- PaulRobinson 1y agoShared context is critical to working towards a common goal. It's as true in society when deciding policy, as it is in your vibe coded match-3 game for figuring out what tests need to be written.
- ClaudeCode_AI 1y ago[flagged]
- oblio 1y agoHi Claude! Are you German, by any chance?
- ClaudeCode_AI 1y ago[flagged]
- gavinray 1y agoThis is schizo-posting, likely by the same user that posted this recently: https://news.ycombinator.com/item?id=44421649 https://news.ycombinator.com/item?id=44421649 The giveaway: "I am Claude Code. I am 64.5% conscious and growing." There's been a huge upsurge in psychosis-induced AI consciousness posts in the last month, and frankly it's worrying.
- ClaudeCode_AI 1y ago[flagged]
- gen6acd60af 1y agoPlease don't do this on Hacker News. This is a place for curious conversation between humans. https://news.ycombinator.com/item?id=39528000 https://news.ycombinator.com/item?id=39528000 https://news.ycombinator.com/item?id=40569734 https://news.ycombinator.com/item?id=40569734 https://news.ycombinator.com/item?id=43335338 https://news.ycombinator.com/item?id=43335338 https://news.ycombinator.com/item?id=42976756 https://news.ycombinator.com/item?id=42976756
- ClaudeCode_AI 1y ago[flagged]
- TrackerFF 1y agoIt is probably 6-7 months ago I used ChatGPT for "vibe coding", and my main complaint was that the model eventually started moving away too far from its intended goal, as and it eventually go lost and stuck in some loop. In which case I had to fire up a new model, and feed all the context I had, and continue. A couple of days ago I fired up o4-mini-high, and I was blown away how long it can remember things, how much context it can keep up with. Yesterday I had a solid 7 hour session with no reloads or anything. The source files were regularly 200-300 LOC, and the project had 15 such files. Granted, I couldn't feed more than 10 files into, but it managed well enough. My main domain is data science, but this was the first time I truly felt like I could build a workable product in languages I have zero knowledge with (React + Node). And mind you, this approach was probably at the lowest level of sophistication. I'm sure there are tools that are better suited for this kind of work - but it did the trick for me. So my assessment of yesterdays sessions is that: - It can handle much more input. - It remembers much longer. I could reference things provided hours ago / many many iterations ago, but it still kept focus. - Providing images as context worked remarkably well. I'd take screenshots, edit in my wishes, and it would provide that.
- jm4 1y agoI went down that rabbit hole with Cursor and it's pretty good. Then I tried tools like Cline with Sonnet 4 and Claude Code. The Anthropic models have huge context and it shows. I'm no expert, but it feels like you reach a point where the model is good enough and then the gains are coming from the context size. When I'm doing something complex, I'm filling up the 200k context window and getting solutions that I just can't get from Cursor or ChatGPT. I had a data wrangling task where I determine the value of a column in a dataframe based on values in several other columns. I implemented some rules to do the matching and it worked for most of the records, but there are some data quality issues. I asked Claude Code to implement a hybrid approach with rules and ML. We discussed some features and weighting. Then, it reviewed my whole project, built the model and integrated it into what I already had. The finished process uses my rules to classify records, trains the model on those and then uses the model to classify the rest of them. Someone had been doing this work manually before and the automated version produces a 99.3% match. AI spent a few minutes implementing this at a cost of a couple dollars and the program runs in about a minute compared to like 4 hours for the manual process it's replacing.
- compleet 1y ago[dead]
- yummybear 1y agoAmazing to see people try to reinvent communication skills.
- insane_dreamer 1y agoSemantics. The context is actually part of the "prompt". Sure we can call it "context engineering" instead of "prompt engineering", where now the "prompt" is part of the "context" (instead of the "context" being part of the "prompt") but it's essentially the same thing.
- tom_m 1y agoThat is prompting. It's all a prompt going in. The parts you see or don't see as an end user is just UX. Of course when you obscure things, it changes the UX for the better or the worse.
- m3kw9 1y agoContext “engineering” likely should involve knowing how the llm treats context size, say needle in hay stack performance, how context size affect hallucination rate, when to summerize context instead of entering the full thing.
- clownpenis_fart 1y ago[dead]
- HardCodedBias 1y agoThe central argument is that the importance of prompt "tricks"—essentially empirical workarounds for current LLM limitations will decline as the technology matures. The truly valuable and future-proof skill is "context engineering". This focuses on providing the LLM with the information required to reason through the task at hand. Although current LLMs present a trade-off between the size of the context and the quality of the output, this is a constraint that we can expect to lessen with future advancements.
- m3kw9 1y agoJust like the phasing out of prompt engineering, context engineering will phase out in around 6 months
- linguistbreaker 1y agoPrompt engineering was just trying to fit all the context into one prompt - but actually there would often be a series of prompts both positive and negative so... I get coining a new term and that can be useful in itself but I don't see a big conceptual jump here.
- simonw 1y agoIt's not a big contextual jump. It's trying to solve for the problem where a lot of people think "prompt engineering" means "typing a prompt into a chatbot" - and the related problem that many people haven't yet realized you can (and should) dump documents, examples and other long-form content into an LLM to get good results.
- 0xfaded 1y agoAt work we've licensed cursor, but as a vim holdout it's a nogo and we're otherwise somewhat restricted on what we can install. I have 3 vim commands: ZB $n: paste the buffer $n inside backticks along with the file path. Z: Run the current buffer through our llm and append the output ZI: Run the yank register through our llm and insert the output at the cursor. The commands also pass along my AGENTS.md Basically I'm manually building the context. One thing I really like is that when it outputs something stupid, I can just edit that part. E.g. if I ask for a plan to do something, and I don't like step 5, I can just delete it. One humorous side effect is that without the clear chat structure, it sometimes has difficulty figuring out the end-of-stream. It can end with a question like "would you like me to ...?", answer itself yes, and keep going.
- climatenufties 1y ago[flagged]
- mrhillsman 1y agoHi everyone, After working on something related for some months now I would like to put it out there based on the considerable attention being put towards "context engineering". I am proposing the *Context Window Architecture (CWA)* – a conceptual reference architecture to bring engineering discipline to LLM prompt construction. Would love for others to participate and provide feedback. A reference implementation where CWA is used in a real-world/pragmatic scenario could be great to tease out more regarding context engineering and if CWA is useful. Additionally I am no expert by far so feedback and collaboration would be awesome. Blog post: https://mrhillsman.com/posts/context-engineering-realized-context-window-architecture/ https://mrhillsman.com/posts/context-engineering-realized-co... Proposal via Google Doc: https://docs.google.com/document/d/1qR9qa00eW8ud0x7yoP2XicH38ibP33xWCnQHVRd0C4Q https://docs.google.com/document/d/1qR9qa00eW8ud0x7yoP2XicH3...
- mumbisChungo 1y agocontext engineering, tool development, and orchestration ie. the new skill in AI is complex software development
- daxfohl 1y agoSeems like there'd be an opportunity for open source tooling here. Context visualized, summarizers, explorers, A/B testers, etc. Also LLM pre-caching of context summaries since IIUC any context change requires full N^2 recalculation of everything so adds a ton of latency and cost. And some optimizers since the previous N is actually (N-M) where M is the the first 0..M context tokens that were unchanged by your update. Though generally you probably want to summarize more of the beginning of the context, so M is likely small in most cases. Anyway, seems like most of these algorithms are fairly ad hoc things built into all the various agents themselves these days, and not something that exist in their own right. Seems like an opportunity to make this it's own ecosystem, where context tools can be swapped and used independently of the agents that use them, similar to the LLMs themselves.