16 ms·
LLMs reward expertise
- walrus01 2mo agoLLMs reward architecture knowledge of how to structure things and how to not just say "Claude, make me Microsoft Flight Simulator, make no mistakes".
- champagnepapi 2mo agoUnfortunately the software industry is saying things like "don't look at the code", "LLMs have made developers 10-100x faster", etc. The only way they can make such claims is by saying what you said above: "Claude, make me Microsoft Flight Simulator, make no mistakes". Additionally engineers are facing pressures via deadlines to work in the paradigm of "Claude, make me Microsoft Flight Simulator, make no mistakes"...
- natsucks 2mo agoThe question i wonder about is, when will an event come along that persuades everyone that human understanding is still required? Or will it never come?
- thewebguyd 2mo agoSuch an event would have to be pretty catastrophic at this point to slow down the inertia. Perhaps the tech debt will just pile up until someone's product implodes, or there's a massive safety issue that causes loss of life, or some big hedge fund goes bust.
- champagnepapi 2mo agoI wonder the same thing. I think we've already seen some of this happening, however the consequences haven't been large enough to the organization, for example: - https://www.theguardian.com/technology/2026/mar/20/meta-ai-agents-instruction-causes-large-sensitive-data-leak-to-employees https://www.theguardian.com/technology/2026/mar/20/meta-ai-a... - https://tech.yahoo.com/articles/ai-code-wreaked-havoc-amazon-160849695.html https://tech.yahoo.com/articles/ai-code-wreaked-havoc-amazon... - https://alexeyondata.substack.com/p/how-i-dropped-our-production-database https://alexeyondata.substack.com/p/how-i-dropped-our-produc... We can only hope that engineers working in safety critical systems haven't fallen to these working conditions.
- RideOnTime22 2mo agoAs long as people keep gaslighting by sayting those events are just "skill issues," I doubt there will sadly be a catalyst.
- strange_quark 2mo agoYeah, this will never happen until governments step in and cause these companies real pain. I mean just look at Crowdstrike. They caused billions and billions in economic damage due to their incompetence, and nothing happened. In fact, their stock is close to an all time high.
- wrs 2mo agoThat question makes me think about Boeing. Or NASA. Or Enron. Reality always wins, no matter what management and Investor Relations says.
- icedchai 2mo agoIt makes me think about The Terminator.
- icameron 2mo agoThe event could be when fair pricing comes from the model providers. We're still at the cash burning stage. When the economy crashes a little and departments start monitoring their spending, and the prices for inference are 10x what they are, there will be less tolerance for employees to substitute constant AI usage for understanding.
- kaibee 2mo agoYou're assuming that LLMs entered a world of people who understood how the systems they're inside of work, why they're setup that way, and that LLMs are displacing them. I sadly don't think that's the case in... well... a lot of the cases.
- bonoboTP 2mo agoMany, including myself, report having a lot of success with braindumping and not structuring anything. Just talking into speech recognition for 2-10 minutes as a stream of consciousness about what my context is, what I want, what I know already, what I have a vague hunch about, how it fits into a bigger picture, what aspects are most important to me, any footguns I already know about, really like having a chat with a person on the phone, with someone you have to guide remotely because they have to implement the thing right now but you have to be out of office and so your only interface is speech. Except you can be more structureless because the AI won't be offended. Just keep on rambling, and press enter, don't even correct mistranscriptions. It will understand it anyway. Now, the key is, that while rambling without structure, you do have to drop the key facts into your speech, and you have to know what you're talking about in at least a good portion of it. I think people are afraid of doing it, because it seems "not the right way" or "not scientific" or whatnot. They want to believe there is some magic to writing the right prompt. So let me tell you, it works.
- walrus01 2mo agoI don't completely disagree with the concept of giving a free association thought process ramble into context. But I also bet that when you start getting it to actually generate code and link modules of things together, subroutines, functions, code structure and filenames, you still pay attention to what it does and you guide it into the architecture that makes logical sense to you.
- bonoboTP 2mo agoFor real work yes. For personal projects, less and less since Fable came out (probably the same if true of the other frontier models). You can get a lot done if it's just some one off, or a personal tool, even without looking at the code, just trying the application. Frontier models now automatically test it before handing the thing to you, they take screenshots, they fix the superficial issues themselves. To get something up and running, it's enough to send chat messages.
- cheriot 2mo agoAgree with this. LLMs multiply the human user's ability. More ability, more impact!
- tills13 2mo agoAnd unfortunately, more ineptitude, more chaos.
- postalcoder 2mo agoNot sure I agree with this. The math guy at anthropic's prompts are essentially: "suppose you’ve gotta resolve the $CONJECTURE, like absolutely have to, everything depends on it. think really hard, and try to come up with a bunch of ideas to try. but remember to trust yourself and not necessarily in conventional wisdom!!" https://claude.ai/share/25740bd5-aa97-4bd7-bf58-c4df3793fda7 https://xcancel.com/__alpoge__/status/2083855298239078748 Tao's chat was for him to gain intuition, not to solve the problem from the outset. What's funny is that every other person gets a different conclusion about who these models reward/empower. I've seen people say that the generalist stands to gain the most and others say that it's the experts. Like all of life, maybe the "winner" is the person who just does stuff.
- zmj 2mo agoIt's not contradictory to say that expertise is a multiplier, and that models are systematically underconfident in themselves.
- titzer 2mo agoIt's actually refreshing when a model is sure about something because it actually tested it and has the receipts. Opus 5 seems really good about testing its own knowledge with experiments. Scientific method ftw.
- atleastoptimal 2mo agoThis works better for math because math is self-verifiable. Once you have a proof it needs no outside evidence. Expertise is needed to evaluate model outputs where it can't verify itself, or at the very least one's expertise can help steer the model in the right direction. However this is irrelevant if models themselves are better at evaluating/leveraging expertise/information.
- colechristensen 2mo agoCorollary to this is an important part of LLM usage is what I call pinning it to reality. That is, designing verification steps that interact with the real world in some way not easy to hallucinate or work around. This means things like having code that interacts with the physical world, round trip tests, arriving at the same result using different paths, interoperability / replication with external libraries / competing products, performance improvement projects that start with robust performance test suites, and similar sorts of things that reduce to "how do I provide evidence that's difficult to fool myself about". This includes things like "before you start fixing this bug, write two tests that fail proving it exists". Expertise is good, but a wise expert will set up methods for the machine to prove to itself that a desired result is achieved removing the expert from the tight development loop.
- k__ 2mo agoPrompt an image or video generator without knowledge in photography or art skills and your results will look sloppy.
- natsucks 2mo agoI am feeling this a lot lately. Getting the most out of agents seems to require being able to ask the right question. And how can you ask the right questions without deep domain expertise?
- ModernMech 2mo agoYes sometimes it’s a matter of just using the right word. You can talk to an agent about a general concept for hours and hours and it may never mention $Concept_X, but you mention $Keyword_Y and all of a sudden the AI is going on about how $Concept_X is foundational to understanding the whole thing.
- QuercusMax 2mo agoI started developing webapps back in the late 90s when I was in high school using Perl, and I've worked with tons of technologies up till around 2014 or so when I shifted into almost pure backend work and lost touch with modern frontend development. I'm now learning how modern frontend development is done (for both personal and professional projects), so I may not know the specific tools, technologies, or terms but I can say "whatever the equivalent of XYZ is" and the models will translate for me. If I say "run pytype" it will tell me "we're using mypy - i'll run that checker for you". If you can express what problem you're trying to solve, that will get you most of the way - and then you can refine by asking questions. "I think I need something like Redis for caching things - do people still use that? Is there a simpler more modern version that is the new standard? Do we already have company docs suggesting what to use?"
- asdfman123 2mo agoYou're basically playing the role of team lead to the LLM's junior dev.
- bt1a 2mo agoI love larping as a vacant scrum master
- Swizec 2mo agoThis matches my experience. Just Talk To It is the best method for working with LLMs if you're an expert. I've seen this at work (as eng manager/lead/principal/whoevenknowsanymore) – all the big APIs give you stats. We see how much people burn in tokens and we know how much output they produce. There is a pretty strong inverse correlation between token burn and output. The more tokens people burn, the less likely they are to produce a good outcome.
- sramsay 2mo agoI do find that "signalling expertise" is important. "I have a significant background in biblical scholarship. You can assume I've read the most important works in NT studies in particular. Do not translate Greek, Latin, Hebrew, or Syriac. Now, I would like to know . . ." That changes things significantly. So does telling it you have 20+ years of experience with C programming, that you have a robust understanding of machine organization, memory layouts, embedded systems, etc.
- QuercusMax 2mo agoFor sure. On a personal coding project I said "I'm a professional software engineer, and while this is a hobby project I'm not just vibe-coding and want to build reliable software" and the agent suddenly started suggesting all kinds of things to make its code more robust.
- PaulStatezny 2mo agoLLMs skew toward over-focusing on things that you mention. The reason "the agent suddenly started suggesting all kinds of things to make its code more robust" is because you said you "want to build reliable software". It's not a signal of good judgment or understanding. It's just how LLM attention works.
- lagrange77 2mo agoI thought exactly the same at first. But then i wondered if that still holds true with today's advanced thinking, RLHF involved, frontier models. I guess to a certain extend it did indeed behave better, as a reaction to his self description into account. EDIT: I mean, those systems accumulated so much complexity around the attention based next token predictor.
- beering 2mo agoTraining the LLM to do things that the user didn’t explicitly ask for is a good way to get complaints from the users. Doesn’t matter if those things are best practices.
- nevi-me 2mo agoI have lengthy conversations with my LLM, almost like an interview. I agree on the expertise part, because I wouldn't be able to go in depth on a subject with it if I lacked the expertise. Some work is a result of design and negotiations in those designs. I don't think Tao's style works with everyone/thing, especially if we don't know what style he's tuned his LLM on.
- boron1006 2mo agoThis is true but also false. In my experience (scientific programming) AI is a giant multiplier for people with specialized knowledge. But it’s also a giant devaluer for that same knowledge as people with no idea what they’re doing can clog the field with plausible bullshit. It’s now the case that if someone tells me they’ve done something, and I look into it and find out it’s completely AI slop, then I will have spent more time on the project than the person who “made” it. The situation is completely untenable and only serves to drain time and resources from people with better things to do.
- theredleft 2mo agowe are slowly punishing reading comprehension this will have educational consequences (that I'm trying to solve). I don't think that we can adjust without rapid education and making extreme specialists of us all. This requires coordination, certification, licensing, and other tiers of authenticity. False experts can ruin sample gathering, can ruin training. False expertise is exemplified by the current American Administration. Look at Robert F. Kennedy Jr.; he's a false expert. He is responsible for the measles outbreak. He is responsible for ivermectin abuse by humans. False expertise is overtaking real expertise. And the results are continuously disastrous and large-scale.
- porphyra 2mo agoThe counterexample of the Dinitz-Garg-Goemans conjecture was basically just "keep going" and finally "enough of partial results. now finish with a complete unconditional counterexample" lol https://x.com/DmitryRybin1/status/2079904005652893709 https://x.com/DmitryRybin1/status/2079904005652893709 https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063 https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de0...
- bt1a 2mo agoI often do my best to represent a genuine interest in the subject at hand and learning in general to models. Imagine the model's response prose and mannerisms being on the other polar end of answering questions simply to get the correct answers as they're often scoped for on quantitative benchmarks. Not sure I explained this well, sorry. An LLM could help
- yearesadpeople 2mo agoYes. I agree with most, if not all of this. For instance, I am seeing folks either relying in the LLM as an _assumed_ expert or, assuming someone - who knows the structure of skill definitions - also has some expertise (in the area of the skill). It's a difficult situation; there is not much point in explaining _why_ the LLM output or skill in use (on a domain problem) isn't what the person actually _needs_ to address the domain problem, because the person isn't a domain expert or indeed, adjacent to domain expertise. But, it is an interesting experiemnt to arm folk with little domain expertise with the _skill_ necessary to be able to extract the right solution from the model.
- erelong 2mo agoThis is also why people's experience with LLMs/AI varies so much, because some people can see a use for AI for their needs and go about using the tool, while others do not as it relates to whatever they're working on and so they may say "LLMs/AI are useless" (it doesn't mean they're not experts though, although some people who have totally no expertise might also see no use for AI for themselves).
- s0rce 2mo agoOverall, I agree, when I ask things I'm an expert in and do professionally every day. I get very good useful answers. When, for example, our marketing people, ask about the science, they often get confusing and wrong answers.
- skybrian 2mo agoSkilled use may or may not matter, depending on the task. Do you need to do what Terence Tao is doing?
- selimfedakar 2mo ago[dead]
- neilv 2mo ago> In the 2010s, if you had technical gaps (say, you couldn’t write CSS), you had to either rely on a skilled colleague or just hope that the answer to your exact problem was out there on the internet. You could read some general reference/guide/tutorial documentation on CSS, and then probably solve your problem (without searching for "how to center a div", or whatever your exact problem was, and copy&pasting the answer and moving on), also becoming more knowledgeable in the process. The rest of the short blog post has some good points, but the first sentence sounds like it's targeted at the percentage of developers who did StackOverflow copy&paste to close Jira tickets, never becoming experts. Delegating to LLM-ish AI is just a natural evolution of that. The question is whether they can still add value if kept in the loop. The article author suggests that the answer is to be expert, and is addressing people who... "either rely on a skilled colleague or just hope that the answer to your exact problem was out there on the internet."
- hahahaa 2mo agoAs they said in the 80s or maybe earlier RTFM. I think if you got a good enough duster TFM was still readable in 2010.
- bonoboTP 2mo agoI don't think AI use is supposed to replace foundational learning such as reading a C++ book or Python book or CSS tutorial when you're a beginner. You still have to do those things if you want to be a professional or a strong amateur. But many people just want to get the thing done. They don't want to become a mechanic, they just want to drive from A to B.
- Avicebron 2mo ago> They don't want to become a mechanic, they just want to drive from A to B. I'm fairly certain the article is directed at professionals, or at least the AI companies are basing their valuations off of directly taking a slice of that professional "productivity".
- nullsanity 2mo ago[dead]
- Arshad-Talpur 2mo agoI cant have an overall opinion but in my personal experience i have analysed that LLMs do reward concreteness
- tsunamifury 2mo agoYes. If you use the right technical terms together it’s lights up more specific feature spaces to your task. Specificity matters to LLMs a lot.
- deleted 2mo ago[deleted]
- zmmmmm 2mo agoThere's a growing and fascinating divide between people who see LLMs as more of a "bicycle for the mind" in the vein of Jobs vs those who see them as whollly supplanting the role of human intelligence. I can't help but wonder if these aren't primarily two human archetypes more than anything - the LLMs can be both and they erect a mirror of the human using them. Some humans really don't want deep individual expertise and intelligence to be the deciding factor because they don't identify with that. Others are completely the opposite. We really can't tell which will be more effective yet, because LLMs are very good in both modes. But most of the predictions currently are people executing on wishful thinking about what they hope will be the outcome.
- jappgar 2mo agoThere ARE two types of people. Those who ride bicycles and those who prefer a self-driving car.
- bashtoni 2mo agoThe short version I give to non-technical people who ask me about whether "AI will replace coding" is this: it accelerates you. You can get much further much more quickly. If you don't know where you're going or how to get there, or even if you're just not paying enough attention, it will get you very far in the wrong direction before you've realised.
- chrisjj 2mo ago... and worse, it will diminish your ability to realise.
- cyberax 2mo agoYes. This is called the Matthew Principle: > For to every one who has will more be given, and he will have abundance; but from him who has not, even what he has will be taken away.
- pianopatrick 2mo agoThis feels like a moment in time, not the end state of AI. Like I read there was a time when teams of people + AI could beat pure AI at chess. But that these days, pure AI wins. For all the things people say about "how AI works" you have to add the missing piece "how current AI works".
- travisgriggs 2mo agoI totally see this. I just did 3 hours of bot sitting to put together some thrash loops that thrash our provisioning working flow for a BLE gadget we make. It was pretty straightforward and productive. But then, I have a lot of experience with BLE, and a quite a bit of experience with python and shell scripting. So I was able to guide the process through stages, do some intermediate testing, make some adjustments, and proceed. Domain experience made this really easy and straightforward. Me two junior engineers who have only superficial/high level knowledge of BLE and some of the other pieces, couldn't have done this as effectively. Where my angst comes, is worrying that no one will ever get that experience anymore. They might have had some eventual success, who knows what monstrosity a much less guided LLM would have done, but experential learning may be mostly a thing of the past. And it creates a real tension between the person with experience and the person without.
- ImaCake 2mo ago>Where my angst comes, is worrying that no one will ever get that experience anymore. I am a fairly inexperienced python developer using LLMs to build software and find that I still learn a lot just from prompting and tinkering. Maybe that's less true once you reach a certain level of competence, but at my intermediate level I am still learning a lot even leaning heavily on LLMs.
- deleted 2mo ago[deleted]
- petres 2mo agoWell, nice post. Actually, there may be some truth behind it, but basically, it captures what I—as a programmer—want to read: expertise will remain valuable. But how I am observing is different, though. Since LLMs the gap between experts and non-experts has been shrinking. And yes, there is still a gap, but vanishing.
- kwakubiney 2mo agoMight be a very noob question but in this era of LLMs, let me ask the reverse, how do you gain expertise? It seems this rewards people who had expertise pre LLMs, but what about people who don’t have that in a specific domain? What approaches are viable now in this current system?
- lucb1e 2mo agoI'm not sure I understand the question. What would prevent you from doing what these people did now that LLMs are here?
- kwakubiney 2mo agoThose people had no choice. In my opinion, it’s harder to grind through problems knowing very well an answer is a prompt away.
- jselysianeagle 2mo agoBut getting an answer is not the same thing as understanding why that is the correct answer, or going deeper and learning more about the subject. IMHO, the people who genuinely desire to learn will trudge through whatever they need to in order to grow their understanding - be it through reading books, original research papers or what have you. If, OTOH, all you seek is the answers and that alone is satisfying to you, then of course you simply will not be motivated to do it the old school way anyway. But that's hardly different now in the age of AI.
- lionkor 2mo agoThe same as it's been! Make things without using LLMs. Don't debug with them, don't use them to research things, just do it yourself. It'll be painful and that pain is learning.
- michaelchisari 2mo agoSame way you get strong in an age of heavy machinery: Lift heavy weights yourself. Skills will have to be built through artificial constraints. Pen & paper, reading books, not using AI, etc.
- abixb 2mo agoThe amplifying mirror analogy works best here. LLMs are ultimately a reflection of your own interactions with its weights, the tone you use, the structure with which you construct your prompt, aspects of an issue you tend to focus on, your breadth of vocabulary and world knowledge and whatnot. People who (carefully) use it as an extension of their own mind and senses will very likely thrive, and those who use it as a replacement for their minds and their senses will struggle. One of the Claude skills I made Claude itself generate was the 'learning a concept across tiers' skill -- from ELI5 level to a PhD level, and it triggers whenever I ask it a very general question on a complex topic that isn't my bread-and-butter. The fact that I'm able to choose explanation level from a super smart LLM (that's available 24x7) that can explain any topic under the sun would've been mind-bogglingly sci-fi-ish just 4 years ago in 2022.
- xvfLJfx9 2mo agoThat sounds useful. Can you share that skill?
- Avicebron 2mo agoNot the OP but you can whack this into your prompt and get most of the way there: "no jargon goes unearned, nothing gets dumbed down, every abstraction touches ground"
- japhib 2mo agoA punchy tricolon containing 2 analogies that don’t quite make sense. That’s some S-tier AI-mimicking. Nice!
- Avicebron 2mo agoOne could say that it being an LLMism is...load-bearing :)
- DaiPlusPlus 2mo ago
- bob1029 2mo agoThe LLM is like the death star. If you don't know exactly where to point it, you will likely miss your target and have no/negative effect. The further away the target, the more accurate your firing solution needs to be. If all you need to do is add something like a dark mode theme to an existing product, this is probably a point blank shot in this metaphor. Building an entire codebase from zero, or even refactoring a legacy codebase into a new codebase, are lightyears away by comparison. You can still land the shot, but you need to deeply understand the metrology and astrodynamics. The information system required to encode the aesthetic preferences needed to make a technology experience not suck is likely in excess of what any near-term solution will offer. Knowing when to say "no" is perhaps the most important skill here. You can't just say it arbitrarily either. You really have to mean it and be willing to fight other humans for it.
- aksappy 2mo agoI think if we have a large population of generalists, then none of them are generalists after all
- xpct 2mo agoI believe they would still be called generalists.
- Austiiiiii 2mo agoThis is something that really needs to be formally studied. I'm inclined to say that this matches my own experience, but I can't rule out confirmation bias on my part. As a meticulous person generally looking for a very specific code outcome, I prompt in a way intended to get exactly the thing I have in mind, and my results reflect that. But on the other hand, I have coworkers who type ten-word prompts with very limited specificity, and they seem to get results that way as well, and that makes me wonder. It would certainly be beneficial for my career and financial well-being for the assertion to be true, because it means I don't have to worry about being pushed out of my job by an army of $15/hr vibe coders. But the convenience of that assumption is exactly why I think it's important to be skeptical.
- mettamage 2mo agoMeanwhile all I do is vibecode. I get the outcomes I want though. I see vibe coded apps as requirement documents. Rarely do I have to engineer. If my job gave me some actual tasks, then maybe I'd engineer something. But at home? Vibe coding all the way. I'm open to engineering, but I need a compelling reason such as: the app is fundamentally broken and an LLM is going in circles. When the only user is me, there are not many performance issues to think about or fix, so that helps. Moreover, certain systems don't need to exist (though they might soon since now I have a smattering of apps that I need to manage).
- ssl-3 2mo agoI suck at writing code, so strictly speaking: Everything I do is vibecoded. The stuff that I produce in this way would probably be considered by many to be unusable trash. But it solves the problems I have, and it does so with exactly the amount of precision that I demand. When I built a PWM fan controller for a pro audio amplifier, I was very particular about some aspects. I wanted maximum resolution from the DS1820B temperature sensors (which is a relatively slow mode where reads take ~750ms, and often the bot is primarily interested in fast), and resolutely-consistent PWM output (so software PWM was a non-starter). It was very important to me that the fan speed ramp smoothly and without audibly-discernible steps, so the target output goes through a low-pass filter to smooth things out and the final PWM value gets recalculated at a completely-overkilled rate of 1KHz. Power consumption was a very deliberate non-concern: The power used by the MCU is ~nothing compared to that of the whole of the system, so optimizing towards reducing it was never my goal. At the end, it's a rewarding little project that is all wrapped into a state machine that burns clock cycles like they're free (they are free!), and it works very well. There's parts of this thing that I do not understand at all, and that I have no desire to understand. But if I hadn't been so particular about the parts I did care about, then: An underspecified one-shot prompt seems like it would probably have just produced a loop with a lazy 1-second sleep at the end, since being sleepy and power-efficient was a feature that the bot kept working to reintroduce. I spent a lot of time working to dismantle the bot's proclivities to be this way, and I probably would not be happy with the end result if I had just let it do its thing. Differently-stated: It could have been an unsupervised one-shot prompt, and the result almost certainly would have done the job of keeping the amplifier cool. (I just would not like it.)
- inventor7777 2mo agoI agree. When I talk to LLMs about fields I am familiar with, I can push back on bad suggestions and ignore faulty/incorrect advice and assumptions, which is much harder for unfamiliar subjects. Of course, simple common sense and extremely basic Googling on unfamiliar subjects can produce similar results, but it's much faster if you are truly understanding what the AI is suggesting.
- roncesvalles 2mo agoThat's why when people like Pieter Levels tweet "I cancelled and then vibecoded 100% of my SaaS subscriptions", you need to take it with a huge grain of salt because you're not Pieter Levels, you cannot vibe code your SaaS subscriptions.
- dbalatero 2mo ago> Of course both are useful, but I’d rather have familiarity with the codebase than a deep general understanding of software systems. In my experience, getting that familiarity with a particular codebase in a way that isn't surface-level has always been a hands-on process. E.g. just because I know many general things about software, I need to know the particulars of the current codebase I'm in to know what is reasonable to actually apply to it. This is a chicken and egg problem I find hard to resolve with LLMs. If we're pushed to delegate most work to them, how do you build that expertise? Sure you can ask questions about the codebase, but IMHO that falls under surface-level information, and the devil is often in the deeper details. Hmm.
- chr15m 2mo agoRead the code.
- Greed 2mo agoWhat if the code sucks, because it was vibe coded by an LLM over a dozen disparate sessions?
- LoganDark 2mo agoClaude, make this codebase less ass
- ggrantrowberry 2mo agoThat will actually work pretty well.
- LoganDark 2mo agoI legitimately caught Claude calling things "ass" while a friend was using it earlier, which I think is pretty funny.
- pyrolistical 2mo ago
- aanet 2mo agoI'm surprised nobody mentioned (including the author) the Gell-Mann Amnesia Effect [1]... Just substitute "LLM" for "journalist" and there you have it. And to be honest, I have seen it, as I'm sure (almost) everyone has, who has demonstrated experience/expertise in their own fields, and correct the LLM's responses one time or another... [1] https://en.wikipedia.org/wiki/Michael_Crichton#%22Gell-Mann_amnesia_effect%22 https://en.wikipedia.org/wiki/Michael_Crichton#%22Gell-Mann_...
- deleted 2mo ago[deleted]
- sonicrocketman 2mo agoThis has been my experience as well. I’ve also been thinking a lot about Terrence Tao and his chats and presentation.
- 6thbit 2mo agoSo we could run a lighter LLM in front of humans, which translates from 'no domain knowledge' to 'domain expert' and in turn prompts over to the larger LLM. Then the larger LLM gets all the right lights on, yields better outputs and we translate back into user domain. I kinda thought the chain-of-thought reasoning already did this, no?
- ekeric13 2mo agoi find this post re-assuring (as who doesn't like to feel like they are an expert at something and llm definitely strips that away)... but it still feels like you are rewarded just as much for being a 6/10 expert as you are for being a 9/10 expert. It definitely is an equalizer it is just a question of to what degree.
- zeroq 2mo agoA good moment to remind everyone that if we took the promise for granted, that AI will in fact prevail and prompting is the one skill that will rule them all... we'll lose all domain experts in one generation. It's less of "signaling expertise" and more about actually having said "expertise". In my experience with LLMs it's not uncommon to be having a deep conversation about making pasta, only to be told, after asking for a sample recipe, to get a bucket of paint and a bag of concrete. Of course these hallucinations are way more subtle and easy to miss for someone who doesn't have deep domain knowledge.
- layer8 2mo agoNow everyone who feels rewarded by LLMs will conclude that it demonstrates their expertise. ;)
- xiaoni-liahuas 2mo ago[dead]
- deleted 2mo ago[deleted]
- amoorthy 2mo agoAgree so much with this! In domains I know well I get much better results then someone who doesn't know the domain because I know where to challenge the LLM. LLMs need to be pushed because otherwise their answers are typically average.
- titzer 2mo agoThe fact that Claude knows I wrote the Virgil compiler makes it be on its best behavior when working on it. I force it to not write too much code, and to write more tests. I push back on slop and just adding another special case. It has a surprisingly deep understanding of floating point.
- jesse_dot_id 2mo agoI've been equating them to graphing calculators since the first LLM launched. It's an amazing tool if you know how to use it. If you don't know how to use it, it's still a tool, but you won't be doing anything amazing with it.
- 27183 2mo agomaybe outing myself as a dinosaur, but "back in my day" the calculator came with a book that detailed exactly how to use it. Both the high level basic language and the low level system language. Not knowing how to use it is simply a failure to Read The Fucking Manual.
- dafelst 2mo agoYou can read the manual all you want, but if you don't know basic algebra, trig, calculus, etc, you are not going to have any idea how to apply or use much of anything that the manual describes with regards to actually doing math with a graphing calculator. There is a base level of knowledge required.
- 27183 2mo agoIt was a general purpose computer, just a small one. Anything you could do on a "real" computer could be done on a calculator, albeit with tighter constraints. It might help to have some higher math objective to accomplish, because that would better utilize the preloaded system software. But in terms of the hardware? Probably not super relevant. I made a lot of use of the TI-89 era CAS in college. But IMO the TI-83 era manuals taught me more about both math and computers than the subsequent generations could have.
- lowbloodsugar 2mo agoDifference is, it can teach you.
- moregrist 2mo agoNice analogy. I loved graphing calculators until I learned tools like Mathematica and Matlab. Still waiting for the Mathematica version of LLMs. Agents / loop engineering / whatever is hot with the AI Twitter kids still isn’t it.
- sixdimensional 2mo ago"The most important skill in the AI era may not be prompting. It may be learning how to solve problems using the right kind of help." [1] Context - I have over 25+ years in software, and I have this observation - being introduced to a new codebase as a human is difficult, especially depending on the scale/size and complexity of it. Yes, you do start to learn it as you work through it, but if the scale is truly huge, it may just not be possible to fully read and understand all the code and paths etc. I have found systems-thinkers (I believe I am one, sometimes they are architects) to be able to kind of "see the whole picture" while not knowing all the details, to the point of being able to guess how the system/software should be behaving, even if it is not actually yet. This is a hugely valuable skill and I think takes a certain kind of brain too. That said, I think recently I may have realized something - we rely on statistics and confidence levels in order to make statements about larger populations. If we can represent a codebase as a, perhaps stratified population of code, interfaces, docs, etc. etc. etc. we may be able to take a valid random sample, review portions of the code, and make some kind of assertions about the state of the larger system - potentially, from that. I am trying to implement this as a side project right now to see if there is anything to it, basically, a combination of AI/LLM + stats/sampling + facilitated expert human review. I'd be interested to know if anybody is doing anything similar. [1] https://www.actinginbalance.com/p/the-right-tool-rule https://www.actinginbalance.com/p/the-right-tool-rule
- GPerson 2mo agoCows
- davesque 2mo agoI actually don't feel like Tao's recently published conversation is the best example of this idea. As intelligent as Dr. Tao is, and surely more so than me, I got the feeling that he wasn't running up against failure states of the model, which I'm not sure you could attribute entirely to his expertise. I honestly think it was more a matter of luck that the model apparently had so much training data on the topic or that it was architecturally so well suited for it. On the other hand, I've had really surprising moments where Claude was just failing terribly to execute simple dev ops tasks having to do with log processing. And I'd be so bold to say that I don't think it could have been explained by a lack of expertise on my part, or even a misuse of the model. So yeah, sometimes LLMs reward expertise, sometimes they don't. I guess either way it helps to have it.
- yyyyyyyyyyzyyyy 2mo agoIt is very funny that you felt the need to say this, "and surely more so than me"
- krisoft 2mo agoI did a test a few months ago. A friend of mine wanted to develop what i understood to be a simple single page web app. But since she didn’t have any software engineering experience she asked me to help. Around that time everyone was talking about how literally anyone can develop software with LLMs i asked her if she could give it a try first, and if I could watch the attempt. I was fully expecting that writing the code will pose no problem for the AI. But i was curious if the AI will realise that my friend is a novice and needs extra help with things like: copy pasting the code into a text file and saving it with an html extension, helping her host the file online so she can share it with others, buying a domain for it, etc. I assumed they will get there eventually, but i also assumed that it will take a lot of stumbling around and misunderstandings. But i was completely wrong. They didn’t even get to that point. Because my friend didn’t have the vocabulary to ask the AI to write code. They were just going around in circles where the AI was brainstorming with her about possible features and getting thints more and more complicated. We terminated the experiment after one and a half hours and many many messages exchanged between her and the LLM. Whereas to me who knows the terminology would have probably just taken a single exchange of messages to get the result she described to me. I would have prompted with something like “Please write to me an html page which does X, Y, Z.” But since she didn’t know the right terminology she got into a vortex of feature discusion, and she didn’t find a way to tip the AI into “just do it, write it now” mode. In other words in that case the LLM would have rewarded even just a little bit of expertise, but without it there was a confusion about goals between the human and the machine.
- Melatonic 2mo agoReminds me of watching someone who has no idea how to use a search engine try to use a search engine
- matthew-wegner 2mo agoAre you describing a "chat window" experience here? This is apples to oranges.
- Synthetic7346 2mo agoYeah I wonder if they had given their friend Claude code or Codex, would it have been more likely to create what she wanted?
- theturtletalks 2mo agoLove this idea of reading prompts that lead to new discoveries and figuring out how the person got the LLM there. It truly is an art and I’m always reminded of “I, Robot” and the scene about “you must ask the right questions.”
- Amekedl 2mo agoYou got to know how to use the model+harness+prompt to achieve the results you want, but honestly for many projects and questions all the models already pump out their same best version of an answer. Sometimes it is really akin to a git clone, although it was a LLM request. This rewarding expertise is somewhat wishful thinking. At the end of a day, it feels and is more like gambling, even with the recommended expertise and a good approach, don't delude yourself you're simply pulling the lever too, as any novice.
- theturtletalks 2mo agoDomain knowledge will stand alone as the sole differentiator. Because LLM benefits can be reaped by almost anyone and it’s a force multiplier. Now those who have the strongest initial force will have a far bigger edge than before.
- ehnto 2mo agoReal world domain knowledge and experience cuts through the chaff too. LLMs are going to have people reinventing the wheel and wasting tonnes of time on stuff that won't work out. If you're a domain expert you are going to be much more aware of how to focus effort in the right places, and what's actually needed or been tried before in your niche. A lot of this domain knowledge is not in any training data, it's locked up in companies in the industry. I suspect it will get even more important to guard it.
- gib444 2mo agoIs someone keeping a list of the excuses and varying instructions on how to hold it right? It would be fascinating historic documentation
- uzername 2mo agoAt work we call this implicit steering. To use webdev metaphor, if a non-technical person describes making a web page with a big block at the top and some things to click on and then my pictures below that, that will eventually get somewhere. Meanwhile, if you know industry jargon, you might describe a hero, with call to action buttons, and then below a 3x3 grid of images of my portfolio photos—that's likely going to generate something entirely different and likely richer. It can assume things about you (it doesn't think), it can ask you specific questions a web personal might know, it can infer domain context that is otherwise omitted with a basic conversation. Everyone wants to capitalize on corporate vibe coding but the tech literacy is hardly there, let alone more advanced topics.
- lowbloodsugar 2mo agoOpus, assume i know nothing about web development. if i wanted to design a new webpage, with a good design, what are some of the terms of art, some best practices? Like if i wanted a big block at the top, some things to click on and some pictures below that, is there terminology for that? >Yes. Nearly everything you described has a standard name. Here is the vocabulary, organized by what part of the page it describes... Goes on to identify Header, Navbar, Stucky header, hamburger menu, hero, CTA, Above the fold etc. >So your described page is: header/nav -> hero with CTA -> card grid -> footer. That is the single most common landing page structure in existence, and that is fine. Being conventional is a feature, not a failure. I've had the same conversation with an electrician wiring a car charger: we are more likely to succeed if I use his terminology.
- jmchuster 2mo agoThat first sentence already uses a ton of jargon that non-developers don't use, "web development", "new webpage", "good design", "best practices", "big block",
- HarHarVeryFunny 2mo agoI think this is just the nature of LLMs as predictive generators. The model is predicting the type/level of conversation based on what the other party is saying. The most typical types of conversation are of two peers, so by default the LLM is likely to respond to you at your own level, unless you ask it to behave differently. As always, prediction goes deep. The best response to Terrance Tao is Tao-level math. It reminds me of reading how LLMs continue chess games if given a partial game - they have learnt to assess player strength based on the moves they make, and will predict game continuations based on the perceived strength of each player, predicting (generating) poor quality moves for a weaker player. This isn't an AI playing chess to win - it's an expert predictor predicting what comes next.
- Animats 2mo agoKeep telling yourself that, right up to the layoff.[1] [1] https://www.linkedin.com/posts/ademola-adelakun_pov-you-get-laid-off-by-your-ai-manager-activity-7424474288159678465-j7L4 https://www.linkedin.com/posts/ademola-adelakun_pov-you-get-...
- techblueberry 2mo agoI think that’s an influencer making a joke.
- anshumankmr 2mo agoSame thing I personally am not a front end guy but I have dabbled with it in the past but I am writing a front end app and a chrome extension, but besides a few pages of code I have reviewed I really do not know what the fck is written (its for an MVP I am building) and I am feeling really conflicted as to what the fuck do I do. At work, the stuff I write has a decent mix of my code, AI code and a few things I do the old way of copying from stackoverflow and seeing what works/doesn't work.
- dtagames 2mo agoThis is absolutely a case where you can't get any output that better than the input, and the input is you.
- luciana1u 2mo ago[flagged]
- fny 2mo agocanine - dog = expertise It's pretty obvious that for some questions a novice wont be able to drive the conversation towards an "answer". A novice may also not be able to understand an answer either. But there's a more subtle failure mode. The vernacular used by an expert and novice to describe the exact same problem lead to different traversals of the information space. For example, I recently asked ChatGPT a medical question using plain english. It gave me an imprecise vague response and told me to call 911. Repeated prodding did not fix this, so I asked the exact same question using medical jargon and in one shot I got what I wanted.
- xsotf 2mo agoWow
- mintflow 2mo agoagreed on this. recently i start to rewrite a core part of one of my iOS VPN app to rust, which previously use fd.io vpp as it's networking core, the original vpp port is 1.5 years ago manually by myself, given i know a lot about how the vpp does and how vpp coroutine and runtime scheduling works. the rewrite is in good shape and solve many issues such as pre allocated memory heap using mmap apis and some scheduling issue of back2back tcp session terminated in the vpp host stack. also by addressing the issus, i am now can easily integrated tailscale as a addon interface for moving in/out l3 packets between tailscale and the core. All those i think cannot be done easily without domain knowledge about those networking and system stuffs.
- zdc1 2mo agoI do agree. An LLM is like a motorboat that's tends to drift off course. If you know where you want to go, and can steer it to keep it on course, you will get there very fast.
- mrloopex 2mo agoThey’re a force multiplier if you are skilled and chaos if you are not.
- qwertox 2mo agoIf you are not skilled they are not chaos, just more distracted, wasting tokens in being nice to humans. It could still teach one very well so that one improves its domain expertise.
- achow 2mo agoThe interesting thing is, most messages were ending with just one question of his. Examples: - ..Does this polynomial map have any symmetry or other structure that makes this cancelation less miraculous? - Given this structure can you see the non injectivity in a transparent way? - ..But why is the jacobian from x u r to P Q R just a monomial? - ..Is there a general theory of such twisted jacobians and do you have any sense why those particular dilation weights were used? - Given this weight structure, why exactly is x given by a cubic equation from P,Q,R? Also, looks like Terence Tao was doing lot of work and asking LLM to verify. This is inverse of the LLM trend, where LLM does the work and humans verify.
- jambalaya8 2mo agoThe problem is LLMs reward no expertise and stupidity also.
- akudha 2mo agoI don’t understand why this is such a revelation. Anyone who has listened to a good/great interview knows the skill of the interviewer plays a big part. To ask good questions, to understand what the other person is saying (AI or human) - that requires skill, expertise and patience. Someone with less skill or expertise might still get good results, sure. It would just take longer and it wouldn’t be pretty
- esjeon 2mo ago> The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” mode, not “explaining-to-amateurs” mode I believe this works in two different ways. First, information compression. The use of professional language helps describe problems more densely with minimal information loss/distortions. Verbose output by LLMs (e.g. ELI5) tend to incorporate local chat context, which can destabilize the context (e.g. out-of-topic, irrelevant nitpicking on writing style and wordings) and lead to faulty logic and even hallucination. LLMs are not good enough to look through all the noise, so, sometimes, it's helpful to refine the input data before performing actual tasks. Second, boosting logical pattern-matching. Using professional language helps drive logical reasoning through simpler pattern-matching b/w texts. This is not about whether LLMs can reason or not; it's about how high-level reasoning is guided by preconception. Even humans tend to consume only textual surface of highly complicated theories (e.g. Adam Smith's "invisible hand"), and use them casually during conversation. It's similar for LLMs: if the conversation is conducted entirely in professional language, LLMs can easily incorporate external professional information into its reasoning. If the text is written in amateurish tongue, translating it into professional language can introduce errors and distortions. So, yeah, keep your conversation professional, tidy and tight. A large volume of unprofessional text helps no one.
- ninjahawk1 2mo agoThis is true for output but as well for learning, if you speak to an LLM trying to get it to give you a certain answer, it’ll find a way to tell you you’re right. If you’re truth seeking and attempting to understand it step by step as it’s going, you’ll likely learn what it’s doing as it’s doing it, meaning you’re basically distilling that information into your own local LLM (also know as the brain).
- lukaslalinsky 2mo agoOf course they do. They have such a huge parameter maps. You need to be able to guide it through the map, so it starts making the right connections. Even in the Sonnet 3.7 days, it became clear to me, that if I have want efficient code out of it, I need to really take care of the context. If I just let it research a problem, it will mess up most of the time. If I tell it to study A, B, C and then present problem D, it will solve it perfectly. And it's true even with the current top models.
- Valakas_ 2mo agoThey reward expertise but not for long. Let's not kid ourselves into coping for a little longer.
- vatsachak 2mo agoWhat makes you think that they will get good at driving themselves? Can they now train on their own outputs? Doesn't seem like it
- vatsachak 2mo agoThe LLM knows how to solve problems in any way you want. This is not a good thing
- madhu_ghalame 2mo ago[dead]
- Thanemate 2mo agoThe LLM industry is deliberately consuming human expertise on a grand scale, so that eventually knowledge work is delegated to any machine, yet I have to gain some sense of comfort knowing that for the present point in time it still rewards personal skill? Regardless of whether you agree with the claim or not, it's definitely not the endgame.
- energy123 2mo agoThe breakthroughs are coming from simple prompts, some made by people with no math training: https://www.newscientist.com/article/2580932-extremely-basic-ai-prompt-cracks-decades-old-maths-problem/ https://www.newscientist.com/article/2580932-extremely-basic... The referenced Terence Tao chat did not lead to new breakthroughs.
- xlii 2mo agoI consider myself senior engineer. When talking with junior colleagues they often are surprised how little I care about some things and how much I care about others. These internal "attention weights" are highly influential parameters of how I work with LLM. E.g. when working with Rust I often hold strict control over structures and lifetimes. But when lately I've been doing token-based bind generation I didn't care about anything outside of high level patterns like RAII and ultimately - API ergonomics which was verified in consumer app. I've been in position of porting real-code to vibe-code platform and seeing non-technical people prompt-stream (they were shared across accounts) I know why they engaged engineer to run this work. Their efforts took 6 weeks, I ported app within 4 days and (to be honest with myself) without LLM I that'd be 3M+ work pre-LLM. In short: I observed same effect as claimed.
- Alisaqqt 2mo ago[flagged]
- deton3991 2mo ago[flagged]
- mariorossi25 2mo agoWhat's this LLM's generated yapping?
- ath3nd 2mo ago[dead]
- Axtrivc 2mo ago[dead]
- nlpnerd 2mo agoThis why those RL env startups are able to charge frontier labs so much for their work. LLMs still generalize poorly outside of self-verifiable tasks like coding and math. Labs have to compensate with post-training in RL env that embeds these expertise well, which is non-trivial both in terms of domain knowledge and technical expertise.
- hintymad 2mo ago> Because my friend didn’t have the vocabulary to ask the AI to write code Is it possible that the effectiveness of an LLM user with respect to the expertise of the user is like a sigmoid function or at least a step function in that shape? That is, one has to know something like the basic concepts and the vocabulary to bootstrap a programming project, but one does not have to know too much to do lots of meaningful work, and then again one needs to be en expert to build something extraordinary. Since most of the work is somewhere middle, most of us mere mortals are still concerned or stressed out for the possibility that LLMs will squeeze out too many job opportunities.
- wei_b0 2mo agoI've experienced this firsthand and 100% agree. The more cracked you are in a domain, the more you can squeeze out of an LLM. If you already know what "good" looks like, you can steer it, call out its BS, and iterate way faster than someone who's using it to learn the domain itself.
- waldarbeiter 2mo agoMaybe LLMs don't usher in the end of software engineering but they definitely end the whole "made with love (and coffee) in X". Nobody cares if you put effort into something software related. "Does it work? Yes? Ok build the next thing." Its the same with the notion of "taste" (see "sometimes tasteless computer code" from the Goedecke article), your colleague who is also a SE might respect your choices as good taste. But it ends there. This was also the case pre-LLMs I would argue. What is worse now is that communicating any uncertainty in decisions related to implementation will result in an immediate "Have you asked Claude?".
- tpoacher 2mo agoBoth the article and some of the discussions here share a lot of commonalities with doctors taking a medical history. There is a certain skill in guiding the conversation towards useful outputs, while not dictating the exact outputs to a patient who is eager to please with their responses. E.g., medical history taking protocol always says to start with open ended (albeit structured) questions, and converge towards more closed/specific ones when you're sure you've extracted the broader surface and you now want to close in on a differential diagnosis. If you start open and go with the flow but then just let the patient talk without any structure or subsequent attempt to converge, there's a risk that the patient might spend 60 minutes taking about their fluffy dog at home, which wastes time, and doesn't get you anywhere nearer the diagnosis. But, if you skip the open questions and go straight to yes/no diagnostic questions, you will definitely miss the fact that they have a dog at home that they're worried about, and that they'll be self-discharging against medical advice in the next hour to go tend to their dog. So while to an outsider, the conversation might look effortless, in reality the doctor requires considerable skill to be able to strike a balance between open vs closed prompts, as well as the ability to critically sift through the outputs, and decide which outputs are relevant to pursue further and lead to a fruitful direction, versus those that can be safely discarded to remove potentially distracting noise from the conversation (and all while attempting to keep this interaction within a limited number of prompts due to operational time constraints).
- Cthulhu_ 2mo agoCounterpoint, some doctors will zoom in on the most likely problem and misdiagnose. This is in part due to pressure on the health care system (where I live anyway); you can only get a GP appointment for 10 minute blocks, which really isn't a lot. But when a 30-some year old shows up at a rheumatologist with joint pain they will likely go to unusual (at that age) but not unheard of rheumatism/arthritis, not hypermobile spectrum disorder. When a woman goes to a GP with period pain they will be prescribed mild pain killers or anticonception pills and fobbed off, until a decade and much suffering / many more issues later they get diagnosed with endometriosis.
- deleted 2mo ago[deleted]
- akkad33 2mo agoThese LLM articles are so boring. Most of them are like shower thoughts with no data to back up and only the writers experience.
- a96 2mo agoLike they were just posting... vibes.
- anjork 2mo agoThis is true today and has been my experience as well -- both to write software as well as doing computational physics. The interesting question then is to ask how long will this stay true? As the models get better will they eventually not need the human expertise to start adding value?
- bjackman 2mo ago> The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” Maybe, but FWIW my first thought when I skimmed Tao's session was that he probably has a personal system prompt requesting this style. E.g. even if you get it into "talking to an expert" mode I've found AI waffling through filler like "given your background in Linux kernel engineering, I'll skip the surface level and go straight to the technical meat". You do have to explicitly tell them if you don't want this.
- MstTK 2mo ago[flagged]
- ixlixl 2mo agoDid you ever hear of this neat thing called "The Bitter Lesson" ?
- josefrichter 2mo agoLLMs are like fire: great servant, terrible master
- ChrisMarshallNY 2mo agoThis has been my experience. I’ve been working on an app (highly successfully) since February, with the help of an LLM (ChatGPT). It has not been used as an author. Rather, it’s been a “coding partner.” I’ve been the one that has submitted the work to VCS, and run the tests. I’ve found the most utility in having it do “small stuff that I could do, myself, but it’s faster to have the LLM do it,” and in analyzing intractable bugs, like memory and threading problems. It’s really good at analyzing a bunch of code, and seeing a small typo that results in something like a strong reference. In both these cases, my own expertise is vital. I’m asking it to act as a consultant; to give me advice and material to be integrated into a whole that I am architecting. I guess part of it, is that I haven’t been able to completely “give in,” and wholly trust the LLM, like I hear many people do (profitably, I guess). I’m used to having my sleeves rolled up, and my hands in the dough. Catching some pretty severe mistakes, from time to time, has reinforced this perception, on my part. I wouldn’t catch these, if I didn’t know what I was doing.
- elendilm 2mo agoI swear the ever living shit out of LLMs for even the tiniest of logical mistakes they commit. Correcting LLMs with extreme swearing that they dare never make it again. I make otherworldly progress with kimi, Gemini, Chatgpt, Deepseek and Claude. Claude now stops the session. Hence Claude is now useless for me. Swearing is nothing personal. Its a correctness enforcer.
- ethical 2mo agoIf you go back to Alan Turings' paper, its all about chastisement! honestly, last pages are all about postive and negative (child!) reenforcement - 1950's style (I do not condone ... etc). Simple as that. I conduct high level litgation in the courts, and win because of a good LLM, with a good version of me, keeping it in line! also crypto and cyber sec. Of course, child rearing, and dealing with former spouses is also very useful. The orginal paper 1950 https://tinyurl.com/yuszahpw https://tinyurl.com/yuszahpw (punish is mentioned six times). Just say-ing-like. TTFN.
- bananaflag 2mo agoEvery thing passes through the following stages: 1. AI cannot do something. 2. AI starts being able to do something, but one needs to prompt it carefully, so one needs to be an expert, see, we will always need human experts <--- this article is here 3. AI just one-shots it. Why do people still need to say this for each and every task? It's just reliving the bitter lesson over and over again.
- quikoa 2mo agoWhy not not show these kickass one-shots and prove how awesome AI can be?
- globular-toast 2mo agoYou need to define what "one-shotting" is. Some examples would help too.
- bananaflag 2mo agohttps://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063 https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de0...
- globular-toast 2mo agoNow explain how someone with no expertise would even know about the Dinitz conjecture, or even a conjecture at all for that matter, and how that would be useful to them.
- bananaflag 2mo agoWhat you are proposing is a barrier no AI and no human and no intelligence would pass. The idea of an AI that one-shots a task presumes that the one requesting the task already knows what the task is, and no additional expertise. Similarly, when you ask an AI "please summarize this text" it means you already know what summarization is as a concept.
- 2mo ago
- manojbajaj95 2mo agoI agree with the premise that LLMs reward experstise, but people without expertise can very eaily learn to prompt correctly and get to a result that is very good. I remember somebody proved a mathematical conjecture by just asking 'keep going' in plain english without a mathematics background.
- FinnLobsien 2mo agoBut then what's the point of the proof? I guess it's cool that it's possible, but given enough tokens, you could take someone who's never written a line of code and have them prompt AI to turn their vibe-coded meal prep app into a highly available distributed system with multi-region failover and immutable audit logs. They would probably get something that checks those boxes in one way or another, but what does it do for them?
- whosdat 2mo agoSure. Not to speak about the ones who own the LLMs. They just rewarded with a FREE SUBSCRIPTION one hundred thousands or so professional mathematicians! Undoubtedly, to advance mathematics! Hurrah!
- FinnLobsien 2mo agoI think this is extremely true when it comes to prompting, but not only in this way. I would add that this also applies to an LLM's output on deep enough topics. Anyone can point at a public GitHub repo and have an LLM write the documentation for it. Whether that documentation is good requires understanding that codebase. There's no way around expertise unless we're talking about strictly mechanical tasks. I do think LLMs are incredible at helping to build your expertise. You could point it at a codebase and say: "Explain how this API works" and interrogating the LLM until you get an explanation at exactly your level of understanding.
- huflungdung 2mo ago[dead]
- cgufus 2mo agoThis reminded me of Gaussian Processes. You start out with n-dimensional unconstrained (but strongly correlated) gaussians. As soon as constraints (data) are added (mathematically it's called conditioning), the thing goes more and more into shape. Prompting feels a lot like this conditioning phase to me. You start with an LLM in unconstrained mode, basically just a "soup" of knowledge. If you prompt wisely, you immediately condition the LLM into "your space of (domain) knowledge". What comes out is an extended version of your existing knowledge.
- mlsu 2mo agoYes I love this comment! GPs are awesome. Prompting is conditioning, that is what it is. The visual of a GP (like the thing you get if you google image search “Gaussian process”) is a great metaphor for what prompting an LLM is doing. The output of the LLM is the logits which is sampled - plucking out tokens from a distribution. The input of the LLM is data which constrains the logits. That is what it is. That’s also how you know that AI will never “solve” intelligence (the way the boosters say it will) without some general mechanism for this conditioning process. The ultimate mechanism would be embodiment; the crappy mechanism we have now is something like openCLAW.
- Culonavirus 2mo agoIt's this https://youtube.com/shorts/OPVp_FDEkdA https://youtube.com/shorts/OPVp_FDEkdA
- miki123211 2mo agoThe actual "prompting trick" that dramatically improves your results is often to add just two or three words, like "use library foo", "write in <language>", "<bar> algorithm". To know which two or three words apply in your situation, you need a deep understanding of both the problem and the solution space. Your prompt might look almost the same as the one from somebody with a good understanding of the requirements but no technical competency, plus maybe one or two sentences. Those one or two sentences dramatically change the results, and what those sentences are differs from prompt to prompt.
- deleted 2mo ago[deleted]
- logseman 2mo agoIf LLMs are good at busy work, and it is expertise that lets you distinguish between busy work and valuable one, then it makes sense that they reward expertise.
- keiferski 2mo agoThis is why the chat interface is ultimately not the best option for non-expert users, because they require the user to bring knowledge with them. You can call this the “query” method: you have to know what to ask to get the answer you want. A real world example might be: I can find any movie DVD you want from our warehouse, but you need to tell me the name of it. Don’t know the name? Tough luck. Contrast this with a “browse” interface: the options available are presented to you, and you can pick from them. Relevant contextual information is already on-site. The DVD store has shelves of potential movies you can rent, and you don’t need to know their names ahead of time. The interfaces of future AI will be more browse oriented, with a query viewer available in the settings for advanced users.
- incrudible 2mo agoPicking a DVD to watch is a rather inconsequential decision. LLMs already do this sometimes, asking you to pick one of a few options, but without domain expertise you will invariably make worse decisions, but if all the n-th order consequences were first explained to you, that would result in you having built domain expertise, but also erasing most of the speed advantage LLMs give you. Moreover, you will never know about the options that are never presented. Inevitably, this is the new tradeoff to make, above average quality comes from asking for more, and knowing what to ask for comes from expertise.
- jwpapi 2mo agoTerminology is a crazy lever with AI
- raptor111 2mo agoI very much agree, but at the same time I feel like those type of shortcomings are fundamental and will be somehow fixed within the next year. The AI companies would just go bankrupt otherwise...
- olup 2mo agoAlso, one can improve domain expertise with the help of the LLM, to becomme a better part of the LLM harness.
- Izmaki 2mo ago> The most important skill in prompting is expertise in the domain you’re prompting for. Amen. AI is a tool, a powerful tool indeed, but if you don't know how to apply it, the quality is seriously impacted.
- mangudai 2mo ago[flagged]
- socketcluster 2mo agoWhen I use Claude to write code for my own projects, the code it generates is exactly the same code that I would have written had I done it by hand. If there is any deviation, I ask it to adjust but that is rare. At least 95% of the time, it's like it read my mind... Which is quite impressive when it outputs like 500+ lines from a single prompt and then it works straight away without any debugging necessary. I don't even debug anymore on those projects. If Claude tries to add debugging logic in my code, I tell it not to and just provide additional information and it can usually find the solution faster that way. This is when working on my own projects. When working on projects created by other people, it's a different story and I have to fight it constantly to stop it from implementing hacks and workarounds... It uses much more tokens to implement basic features. It's more work for both the AI agent and myself. The project's existing code makes up most of the context so if the code is not great, you have to write long detailed prompts to set it on the right path. You have to make it clear that the existing code isn't good enough and your expectation is higher. In this case, it usually gets better with more back-and-forth... At the beginning, it can't do anything because you keep pointing out a problem whenever it tries anything at all, but eventually, after a lot of criticism, it starts becoming more careful and adapting to your standards. So yeah, even same person doing the prompting can lead to two very different experiences depending on who built the foundation. So my conclusion is that the expertise comes from both the existing codebase and from the person doing the prompting... And TBH, I would say the codebase/foundation carries more weight than the person doing the prompting. Pretty sure I could put an idiot on one of my codebases with Claude Code and they'd do a decent job.
- dr0idattack 2mo agoI like it. And if you don't know a codebase, like normal, spend some time learning in. Use the LLM to query it, create your own architecture diagrams. Get homey with it before making sweeping changes. Maybe make the first simple changes by hand.
- perrygeo 2mo agoLLMs are language models. So much of this can be reduced to a simple heuristic: If you can't think clearly, LLMs will not help you. If you don't know what you're asking for or how to express it precisely, it really should not be a surprise that the output is garbage. Failures of LLMs are more often failures of our own brain to consider the problem clearly. That's harder to admit than just blaming the AI. It's important to note though, from the perspective of the LLM's objective function, that this is not a failure at all! LLMs are designed to match patterns. You give a jumbled mess of incoherent ideas, it will faithfully reproduce a token stream of incoherent ideas. It's only when the model output is subjected to the real world that it fails.
- econ 2mo agoIm beginning to think buzzword bingo is a viable job interview method. If you use the b words the LLM would be somewhat constrained to content that has them. If you don't use them it will use content that doesn't have them. The company website (subject) will never become a static html document. If you request exactly what static html does it should probably point you to a wysiwyg website builder.
- balderdash 2mo agoI haven’t experienced this - or maybe the training data for aviation is limited. But ask a an llm for aircraft performance / flight planning data and it’s scary how bad the advice /feed back is.
- uncivilized 2mo agoMost Hacker News are working on simple JavaScript applications. LLMs quickly become unusable on any sort of specialist discipline except for the math marketing releases we’ve seen recently.
- still_grokking 2mo agoFor anything where "the answer" wasn't already in the training data you just get some arbitrary correlated tokens out, as that's all a LLM can do. Of course the meaning of these tokens is just random. (And even for things that were in the training data you don't have any guaranty they will be reproduced correctly, there is just some chance something meaningful comes out, or it doesn't, it's random.) The whole idea to use a next token predictor as "answer machine" is completely flawed. This can't work like advertised, and that's by construction.
- smcleod 2mo agoSimilarly there is research that shows the quality of LLM outputs strongly correlate with the education level (in the field) of the person promoting them.
- conartist6 2mo agoThey may reward it but they don't build it, and therein lies the paradox
- norren8 2mo agoWhat a nothing-burger. Garbage in garbage out. Hasn't everyone who worked with LLM's experienced this?
- vicentwu 2mo agoLLMs raise the floor, but you determine the ceiling.
- jebarker 2mo agoA fascinating thing about the LLM/AI blogosphere and X is watching memetics in real time. An idea like this one propagates on the order of days until everyone that speaks publicly or in workplace meetings about AI is repeating it.
- bishengke 2mo ago[flagged]
- gootz 2mo agoMy LLM said this was a good article :)
- TormentNexusAI 2mo agoThis is an interesting problem. A similar approach that worked for us was to only load the tools the agent actually needs for each task.
- abhishek03113 2mo agoNot the best way to test this, but I am working on a blog post, where I'll implement a problem statement with the dumbest/cheapest AI model while someone non technical person will vibe code end to end and compare both of them.
- m3kw9 2mo agoFor example, if you write iOS apps, you ask it to do some animations. It will do it for you, but you are at it's mercy of writing complete custom code or you can specify they use a certain apple supplied API for a native experience. It's one of thousand things that will get a vibe coder if they don't ask. Or you may get lucky and LLM chooses to use a native method.
- m3kw9 2mo agoThe entire issue is that when you ask it to do something, you are leaving it to chance they may or may not do it properly, either on look/feel, performance, security, scalability etc. It compounds as you layer a new prompt output over that project.
- stillpointlab 2mo agoI agree with this post's gist, and I've certainly noticed how LLMs change their interaction with me once I demonstrate some knowledge. I've often started a technical conversation very vaguely and only once I challenge the LLM on its simplifications does it start to actually get to the meat of issues. Often there is a perceptible moment where the LLM seems to recognize my level of ability and how it communicates clearly changes. But another thing I have found is that I get significantly better results from the LLM by treating it like an intelligent independent agent. All of the "you are a senior dev ..." or "your starving kids depend on the correctness of this answer ..." kind of prompting has been mostly useless. In general, I find being honest and clear to be the best strategy. You can't "pretend" to be a senior software engineer. If I can root you out of an interview process then you aren't going to fool the LLM. But if you clearly state your level of expertise and your desired outcome, then the LLM does a very good job of meeting you where you are. There is also a strange ephemeral attitude I get from agents sometimes, like they don't like to be called out for being wrong. But in the same way that human's show this trait, they also seem to warm up over time as they gain trust. It is almost like social positioning, once they realize they aren't actually expert they morph into a support role stance pretty seamlessly. That is also why they can still feel sycophantic, because once they realize they aren't actually driving the discussion they can actually feel like enthusiastic passengers, wanting to see where the conversation leads as much as the prompter.
- still_grokking 2mo agoThat's some of the most weird anthropomorphization of a next token predictor I've read in a while… This things output token which are correlated with the context given. That's all! If you feed it some context the parrot will answer with the same. It does not "sense your expertise level"—it just outputs correlated tokens… Is this really so hard to understand?
- stillpointlab 2mo agoIt seems possible, in fact reasonable, to understand my observation as following from your assumed "correlation". I gave it context that it is having a conversation with an expert, it correlates it's output with the context. A non-expert user can't fake that context so they never see similar correlated output. This context-correlation doesn't happen if one uses "You are a expert ..." prompting "tricks". My experience has been, if you speak like an expert who is speaking to an expert you get better results. I'm not sure how that is anthropomorphizing. It is just "if I do these things, I get these results".
- swordsith 2mo agoSometimes after a larger source draft with AI before I even touch the program I start to pick up on some of the holes in my prompts caused bugs/unintentional mechanics to be potentially woven in. Knowing what you want and how it should be made is half of it, but if you don't supply the bot with extra guardrails eg. don't modify the contents of x, because of y don't allow z. etc. They will do what you ask of them, usually less than you'd hope.
- Max345 2mo agoLLMs are good on things I know little about, but fall short on things I'm good in.
- still_grokking 2mo agoNo, they fail all the time. You just don't notice all the bullshit if you're not already an expert in the field you generate LLM output for.
- igravious 2mo agoHah. Clever.
- pop3zxcv 2mo ago[flagged]
- py4 2mo agoI am not sure. Why do you need domain expertise beyond being able to draft a verifier for the problem? Once you have the verifier, it's just a matter of compute. You can argue that being a domain expert allows you to narrow down the search space for the LLM and save time/compute costs. This is partially true, but LLMs are getting better and better at search (they can already do end-to-end performance optimization faster than performance experts at a FAANG company I work at), and compute to maintain the same intelligence level is getting cheaper. A concrete example: GPU performance optimization for a kernel. This was (and still is) a very niche domain with not many top-notch experts. But kernel performance and characteristics are easily verifiable. You can run the agent in a closed loop for it to improve iteratively (and people are already doing it, coming up with kernels better than human-written ones). You see Tao's example because: 1. He is curious (so he asks detailed questions, which are not necessarily needed in a closed-loop optimization). 2. Verification in math is harder. Many math tasks used in RL are easily verifiable. But for advanced open conjectures that require long proofs, you cannot trust the proof directly from the LLM (so it's not as easily verifiable as basic math problems or code). The model needs to write it in Lean, and you still need to make sure the Lean implementation correctly captures the specification of the problem. So you still need a human for verification in advanced math. But I don't see why you would need this in domains like performance improvement.
- Hi_mike 2mo ago[flagged]
- zqna 2mo agoI'm yet to find a developer who would agree that there is no good git UI, and would like to fix that situation together with me. 90% of those who know what got is dont care. The remaining 10% is split between command line fanatics and those who have no time to spare. One could imagine that 'LLM revolution' would have open the doors to unlimited prototyping but what we instead get is "hey ma i got a different angle at HN comments" at best
- distantprovince 2mo agoThat's a pretty good frame and makes me hopeful. Now it's even more important to figure out how on Earth we can promote expertise formation in the age of AI
- mastic 2mo ago[flagged]