11 ms·
Reflections on AI at the End of 2025
- danielfalbo 10mo ago> There are certain tasks, like improving a given program for speed, for instance, where in theory the model can continue to make progress with a very clear reward signal for a very long time. This makes me think: I wonder if Goodhart's law[1] may apply here. I wonder if, for instance, optimizing for speed may produce code that is faster but harder to understand and extend. Should we care or would it be ok for AI to produce code that passes all tests and is faster? Would the AI become good at creating explanations for humans as a side effect? And if Goodhard's law doesn't apply, why is it? Is it because we're only doing RLVR fine-tuning on the last layers of the network so all the generality of the pre-training is not lost? And if this is the case, could this be a limitation in not being able to be creative enough to come up with move 37? [1] https://wikipedia.org/wiki/Goodhart's_law https://wikipedia.org/wiki/Goodhart's_law
- username223 10mo ago> I wonder if, for instance, optimizing for speed may produce code that is faster but harder to understand and extend. Superoptimizers have been around since 1987: https://en.wikipedia.org/wiki/Superoptimization https://en.wikipedia.org/wiki/Superoptimization They generate fast code that is not meant to be understood or extended.
- progval 10mo agoBut there output is (usually) executable code, and is not committed in a VCS. So the source code is still readable. When people use LLMs to improve their code, they commit their output to Git to be used as source code.
- Wowfunhappy 10mo ago...hmm, at some point we'll need to find a new place to draw the boundaries, won't we? Until ~2022 there was a clear line between human-generated code and computer-generated code. The former was generally optimized for readability and the latter was optimized for speed at all cost. Now we have computer-generated code in the human layer and it's not obvious what it should be optimized for.
- erichocean 10mo ago> it's not obvious what it should be optimized for It should be optimized for readability by AI. If a human wants to know what a given bit of code does, they can just ask.
- lemming 10mo agoI wonder if, for instance, optimizing for speed may produce code that is faster but harder to understand and extend. This is generally true for code optimised by humans, at least for the sort of mechanical low level optimisations that LLMs are likely to be good at, as opposed to more conceptual optimisations like using better algorithms. So I suspect the same will be true for LLM-optimised code too.
- franktankbank 10mo agoEhh I think if it ends up being a half good architecture you wind up with a difficult to understand kernel that never needs touching.
- ur-whale 10mo agoNot sure I understand the last sentence: > The fundamental challenge in AI for the next 20 years is avoiding extinction.
- danielfalbo 10mo agoI think he's referring to AI safety. https://lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities https://lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-lis...
- grodriguez100 10mo agoFor a perhaps easier to read intro to the topic, see https://ai-2027.com/ https://ai-2027.com/
- dkdcio 10mo agoor read your favorite sci-fi novel, or watch Terminator. this is pure bs by a charlatan
- chrishare 10mo agoHe's referring to humanity, I believe
- A_D_E_P_T 10mo agoIt's ambiguous. It could go the other way. He could be referring to that oldest of science fiction tropes: The Bulterian Jihad, the human revolt against thinking machines.
- AnimalMuppet 10mo agoMeh. I think the more likely scenario is the financial extinction of the AI companies.
- timmytokyo 10mo ago
- agumonkey 10mo agoThere's videos about Diffusion LLMs too, apparently getting rid of the linear token generation. But I'm no ML engineer.
- nephanth 10mo agoAs someone who worked on transformer-based diffusion models before (not for language though), i can say one thing: they're hard. Denoising diffusion models benefited a lot from the u-net, which is a pretty simple network (compared to a transformer) and very well-adapted to the denoising task. Plus diffusion on images is great to research because it's very easy to visualize, and therefore to wrap your head around Doing diffusion on text is a great idea, but my intuition is it will prove more challenging, and probably take a while before we get something working
- agumonkey 10mo agoThanks. Do you see that part of the field as plateauing or ramping up (even taking into account the difficulty). If you know labs / researchers on the topic, i'd love to read their page / papers
- fleebee 10mo ago> The fundamental challenge in AI for the next 20 years is avoiding extinction. That's a weird thing to end on. Surely it's worth more than one sentence if you're serious about it? As it stands, it feels a bit like the fearmongering Big Tech CEOs use to drive up the AI stocks. If AI is really that powerful and I should care about it, I'd rather hear about it without the scare tactics.
- grodriguez100 10mo agoI would say yes, everyone should care about it. There is plenty of material on the topic. See for example https://ai-2027.com/ https://ai-2027.com/ or https://www.lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities https://www.lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a...
- dkdcio 10mo agofear mongering science fiction, you may as well cite Dune or Terminator
- defrost 10mo agoThere's arguably more dread and quiet constrained horror in With Folded Hands ... (1947) Despite the humanoids' benign appearance and mission, Underhill soon realizes that, in the name of their Prime Directive, the mechanicals have essentially taken over every aspect of human life. No humans may engage in any behavior that might endanger them, and every human action is carefully scrutinized. Suicide is prohibited. Humans who resist the Prime Directive are taken away and lobotomized, so that they may live happily under the direction of the humanoids. ~ https://en.wikipedia.org/wiki/With_Folded_Hands_ https://en.wikipedia.org/wiki/With_Folded_Hands_...
- XorNot 10mo agoThis hardly disproves the point: no one is taking this topic seriously. They're just making up a hostile scenario from science fiction and declaring that's what'll happen.
- alexgotoi 10mo ago> * The fundamental challenge in AI for the next 20 years is avoiding extinction. This reminded me of the Don’t look up movie where they basically gambled with the humans extinction.
- torlok 10mo agoThis is a bunch of "I believe" and "I think" with no sources by a random internet person.
- echelon 10mo ago> by a random internet person. The creator of Redis.
- cinntaile 10mo agoSure but quite a few claims in the article are about AI research. He does not have any qualifications there. If the focus was more on usefulness, that would be a different discussion and then his experience does add weight.
- djdishsv 10mo ago> smart, intelligent person gives opinion > woah buddy this persons opinion isn’t worth anything more than a random homeless person off the street. they’re not an expert in this field Is there a term for this kind of pedantry? Obviously we can put more weight behind the words a person says if they’ve proven themselves trustworthy in prior areas - and we should! We want all people to speak and let the best idea win. If we fallback to only expert opinions are allowed that’s asking to get exploited. And it’s also important to know if antirez feels comfortable spouting nonsense. This is like a basic cornerstone of a functioning society. Though, I realize this “no man is innately better than another, evaluate on merit” is mostly a western concept which might be some of my confusion.
- blibble 10mo ago> Obviously we can put more weight behind the words a person says if they’ve proven themselves trustworthy in prior areas - and we should! no, you shouldn't this is how you end up with crap like vaccine denialism going mainstream "but he's a doctor!"
- 10mo ago
- feverzsj 10mo agoSeems they also want some AI money[0]. Guess, I'll keep using Valkey. [0] https://redis.io/redis-for-ai/ https://redis.io/redis-for-ai/
- danielfalbo 10mo ago> they I'm not sure antirez is involved in any business decision making process at Redis Ltd. He may not be part of "they".
- antirez 10mo agoI'm not involved in business decisions and while I'm very AI positive I believe Redis as a company should focus on Redis fundamentals: so my piece has zero alignment on what I hope for the company.
- sibellavia 10mo agoIn any case, what would be the problem? The page you mentioned simply illustrates how the product can be used in a specific domain; it doesn't seem forced to me.
- bgwalter 10mo agoConflict of interest and disclosure posts are frequently downvoted.
- tptacek 10mo agoYou mean flagged. Please don't post insinuations about astroturfing, shilling, brigading, foreign agents, and the like. It degrades discussion and is usually mistaken. If you're worried about abuse, email hn@ycombinator.com and we'll look at the data. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- bgwalter 10mo agoAh, so you just went through my history and downvoted everything in sight! Thanks for confirming.
- ctoth 10mo ago> The fundamental challenge in AI for the next 20 years is avoiding extinction. So nice to see people who think about this seriously converge on this. Yes. Creating something smarter than you was always going to be a sketchy prospect. All of the folks insisting it just couldn't happen or ... well, there have just been so many objections. The goalposts have walked from one side of the field to the other, and then left the stadium, went on a trip to Europe, got lost in a beautiful little village in Norway, and decided to move there. All this time though, the prospect of instantiating a something smarter than you (and yes, it will be smarter than you even if it's at human level because of electronic speeds...) This whole idea is just cursed and we should not do the thing.
- cheschire 10mo ago"Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should."
- Mawr 10mo ago> Creating something smarter than you was always going to be a sketchy prospect. Sure, but not so sure that this has any relevance to the topic at hand. You seem to be taking the assumption that LLMs can ever reach that level for granted. It may be possible that all it takes is scaling up and at some point some threshold gets reached past which intelligence emerges. Maybe. Personally, I'm more on board with the idea that since LLMs display approximately 0 intelligence right now, no amount of scaling will help and we need a fundamentally different approach if we want to create AGI.
- Aiisnotabubble 10mo agoWhat also happens and it's irrelevant of AGI: global RL Around the world people ask an LLM and get a response. Just grouping and analysing these questions and solving them once centrally and then making the solution available again is huge. Linearly solving the most asked questions and then the next one then the next will make, whatever system is behind it, smarter every day.
- danielfalbo 10mo agoExactly. The singularity is already here. It's just "programmers + AI" as a whole, rather than independent self-improvements of the AI. I wonder how a "programmers + AI" self-improving loop is different from an "AI only" one.
- bryanrasmussen 10mo agoThe AI only one presumably has a much faster response time. The singularity is thus not here because programmer time is still the bottleneck, whereas as I understand in the singularity time is no longer a bottleneck component.
- Aiisnotabubble 10mo agoAGI will be faster as it doesn't need initial question. AGI will also be generic. LLM is already very impressive though
- yeasku 10mo agoYou are all crazy.
- seu 10mo ago> And I've vibe coded entire ephemeral apps just to find a single bug because why not - code is suddenly free, ephemeral, malleable, discardable after single use. Vibe coding will terraform software and alter job descriptions. I'm not super up-to-date on all that's happening in AI-land, but in this quote I can find something that most techno-enthusiast seem to have decided to ignore: no, code is not free. There are immense resources (energy, water, materials) that go into these data centers in order to produce this "free" code. And the material consequences are terribly damaging to thousands of people. With the further construction of data centers to feed this free video coding style, we're further destroying parts of the world. Well done, AGI loverboys.
- Hendrikto 10mo agoYou know what uses roughly 80 times more water in the US alone than water used by AI data centers world wide? Corn.
- raddan 10mo agoAssuming your fact is true, that corn merely uses an order of magnitude or two more water than AI is surprising, given the utility of corn. It feeds the entire US (hundreds of millions of people), is used as animal feed (thus also feeding us), and is widely exported to feed other people. I the spirit of the “I think”s and “I believe”s of this blog post, I think that corn has a lot more utility than AI.
- Hendrikto 10mo ago> It feeds the entire US (hundreds of millions of people), is used as animal feed (thus also feeding us), and is widely exported to feed other people. Not really. Most corn grown in the US isn’t even fit for consumption. It is primarily used for fermenting bioethanol.
- daveguy 10mo agoSource?
- Fraterkes 10mo agoIt’s interesting that half the comments here are talking about the extinction line when, now that we’re nearly entering 2026, I feel the 2027 predictions have been shown to be pretty wrong so far.
- deleted 10mo ago[deleted]
- squidbeak 10mo ago> I feel the 2027 predictions have been shown to be pretty wrong so far Does your clairvoyance go any further than 2027?
- AnimalMuppet 10mo agoI don't know that it's "clairvoyance". We're two weeks from 2026. We might be able to see somewhat more than we do now if this was going to turn into AGI by 2027. If you assume that we're only one breakthrough away (or zero breakthroughs - just need to train harder), then the step could happen any time. If we're more than one away, though, then where are they? Are they all going to happen in the next two years? But everybody's guessing. We don't know right now whether AGI is possible at current hardware levels. If it is N breakthroughs away, we all have our own guesses of approximately what N is. My guess is that we are more than one breakthrough away. Therefore, one can look at the current state of affairs and say that we are unlikely to get to AGI by 2027.
- jennyholzer2 10mo ago> Does your clairvoyance go any further than 2027? why are you so sensitive?
- a_bonobo 10mo ago>* For years, despite functional evidence and scientific hints accumulating, certain AI researchers continued to claim LLMs were stochastic parrots: probabilistic machines that would: 1. NOT have any representation about the meaning of the prompt. 2. NOT have any representation about what they were going to say. In 2025 finally almost everybody stopped saying so. Man, Antirez and I walk in very different circles! I still feel like LLMs fall over backwards once you give them an 'unusual' or 'rare' task that isn't likely to be presented in the training data.
- jmfldn 10mo ago"In 2025 finally almost everybody stopped saying so." I haven't.
- dist-epoch 10mo agoSome people are slower to understand things.
- oersted 10mo agoLLMs certainly struggle with tasks that require knowledge that is not provided to them (at significant enough volume/variance to retain it). But this is to be expected of any intelligent agent, it is certainly true of humans. It is not a good argument to support the claim that they are Chinese Rooms (unthinking imitators). Indeed, the whole point of the Chinese Room thought experiment was to consider if that distinction even mattered. When it comes to of being able to do novel tasks on known knowledge, they seem to be quite good. One also needs to consider that problem-solving patterns are also a kind of (meta-)knowledge that needs to be taught, either through imitation/memorisation (Supervised Learning) or through practice (Reinforcement Learning). They can be logically derived from other techniques to an extent, just like new knowledge can be derived from known knowledge in general, and again LLMs seem to be pretty decent at this, but only to a point. Regardless, all of this is definitely true of humans too.
- rckt 10mo ago> Even if LLMs make mistakes, the ability of LLMs to deliver useful code and hints improved to the point most skeptics started to use LLMs anyway Here we go again. Statements with the single source in the head of the speaker. And it’s also not true. The llms still produce bad/irrelevant code at such rate that you can spend more time prompting than doing things yourself. I’m tired of this overestimation of llms.
- iamflimflam1 10mo agoBut you have just repeated what you are complaining about.
- rckt 10mo agoDo you want me to spend time to come with a quality response to a lazy statement? It’s like fighting with windmills. I’m fine with having my say the way I did.
- xiconfjs 10mo agoMy person experience: if I can find a solution on stackoverflow etc. the LLM will produce working and fundamentally correct code. If I can‘t find a already fullfilled solution on these sites, the LLM is hallucinating like crazy (newer existing functions/modules/plugins, protocol features which aren’t specified and even github-repos which never existed). So, as stated my many people online before: for low-hanging fruits LLM are totally viable solution.
- danielbln 10mo agoI don't remember the last time Claude Code hallucinated some library, as it will check the packages, verify with the linter, run a test import and so on. Are you talking about punching something into some LLM web chat that's disconnected from your actual codebase and has tooling like web search disabled? If so, that's not really the state of the art of AI assisted coding, just so you know.
- 10mo ago
- deleted 10mo ago[deleted]
- dhpe 10mo agoI have programmed 30K+ hours. Do LLMs make bad code: yes all the time (at the moment zero clue about good architecture). Are they still useful: yes, extremely so. The secret sauce is that you'd know exactly what to do without them.
- qsort 10mo agoOne of the mental frameworks that convinced me is how much of a "free action" it is. Have the LLM (or the agent) churn on some problem and do something else. Come back and review the result. If you had to put significant effort into each query, I agree it wouldn't be worth it, but you can just type something into the textbox and wait.
- daveguy 10mo agoAre you counting the time/effort to evaluate the accuracy and relevance of an LLM left to "think" for a while?
- feverzsj 10mo agoSo, it's like taking off your pants to fart.
- _rpxpx 10mo agoOK, maybe. But how many programmers will know this in 10 years' time as use of LLMs is normalized? I like to hear what employers are saying already about recent graduates.
- bartread 10mo agoThey’d have to be hiring recent graduates for you to hear that perspective. And, as much as what I’ve just said is hyperbolically pessimistic, there is some truth to it. In the UK a bunch of factors have coincided to put the brakes on hiring, especially smaller and mid-size businesses. AI is the obvious one that gets all the press (although how much it’s really to blame is open to question in my view), but the recent rise in employer AI contribution, and now (anecdotally) the employee rights bill have come together to make companies quite gunshy when it comes to hiring.
- piker 10mo ago> There are certain tasks, like improving a given program for speed, for instance, where in theory the model can continue to make progress with a very clear reward signal for a very long time. Super skeptical of this claim. Yes, if I have some toy poorly optimized python example or maybe a sorting algorithm in ASM, but this won’t work in any non-trivial case. My intuition is that the LLM will spin its wheels at a local minimum the performance of which is overdetermined by millions of black-box optimizations in the interpreter or compiler signal from which is not fed back to the LLM.
- dist-epoch 10mo agohttps://github.com/algorithmicsuperintelligence/openevolve https://github.com/algorithmicsuperintelligence/openevolve
- piker 10mo agohttps://chatgpt.com/backend-api/estuary/public_content/enc/eyJpZCI6Im1fNjk0Njg0MjkzNTEwODE5MWE2NzY5MmE4YWRjNTZiMTA6ZmlsZV8wMDAwMDAwMDljNzg3MWZkYTExODc2MDgxZDllYjAyOSIsInRzIjoiMjA0NDIiLCJwIjoicHlpIiwiY2lkIjoiMSIsInNpZyI6IjIxMDJlMDkzMGExNjNkYWY3OWI4ZTI4YmNhZDE5OThlNGFjYmQxNjQzNzQ2ODRiYmM3NDFlZmE1OGViMjQ5NzgiLCJ2IjoiMCIsImdpem1vX2lkIjpudWxsLCJjcyI6bnVsbCwiY3AiOm51bGwsIm1hIjpudWxsfQ== https://chatgpt.com/backend-api/estuary/public_content/enc/e...
- andy99 10mo agoThere was a discussion the other day where someone asked Claude to improve a code base 200x https://news.ycombinator.com/item?id=46197930 https://news.ycombinator.com/item?id=46197930
- exitb 10mo agoThat’s most definitely not the same thing, as „improving a codebase” is an open ended task with no reliable metrics the agent could work against.
- NitpickLawyer 10mo ago
- HellDunkel 10mo ago[flagged]
- danielbln 10mo agoMust feel nice to let yourself be coddled by in-group/out-group thinking like that. "I've decided that AI is bad and useless, therefore anyone disagreeing must be an AI bro".
- abricq 10mo ago> * Programmers resistance to AI assisted programming has lowered considerably. Even if LLMs make mistakes, the ability of LLMs to deliver useful code and hints improved to the point most skeptics started to use LLMs anyway: now the return on the investment is acceptable for many more folks. Could not agree more. I myself started 2025 being very skeptical, and finished it very convinced about the usefulness of LLMs for programming. I have also seen multiple colleagues and friends go through the same change of appreciation. I noticed that for certain task, our productivity can be multiplied by 2 to 4. So hence comes my doubts: are we going to be too many developers / software engineers ? What will happen for the rests of us ? I assume that other fields (other than software-related) should also benefits from the same productivity boosts. I wonder if our society is ready to accept that people should work less. I think the more likely continuation is that companies will either hire less, or fire more, instead of accepting to pay the same for less hours of human-work.
- danielfalbo 10mo ago> Are we going to be too many developers / software engineers ? What will happen for the rests of us? I propose that we should raise the bar for the quality of software now.
- abricq 10mo agoYes, certainly agree. A few days ago here there was this blog claiming how formal verification would become widely more used with AI. The author claiming that AI will help us with the difficulty barrier to write formal proofs.
- throw1235435 10mo agoI don't think that will happen because it hasn't for other technological improvements. In the end people pay for "good enough" and that's that. If "good enough" is now cheaper to implement that's all they will do. I've seen it in other technologies. As an example due to more precise manufacturing many manufacturers have used it to cheapen things like cars, electronics, etc just to the point where it passes warranty mostly; in the old days they had to "overbuild" to get it to that point putting more quality into the product. Quality is a risk mitigation strategy; if software is disposable just like cheap manufactured goods most people won't pay for it thinking they can just "build another one". What we don't realise is due to sheer cost of building software we've wanted quality because its too expensive to fix later; AI could change that. Hoping we invest in quality, more software (which has a price inelastic curve mostly due to scale/high ROI) etc I'm starting to think is just false hope from people in the tech industry that want to be optimistic which generally is in our nature. Tech people understand very little about economics most of the time and how people outside tech (your customers) generally operate. My reflection is mostly I need to pivot out of software; it will be commoditized.
- register 10mo agoWhere to understand more about how chain of thoughs really affects LLMs performance? I read the seminal paper but all it says is that it's basically another prompt engineering tecnique that improves accuracy.
- HarHarVeryFunny 10mo agoChain of thought, now including "reasoning", are basically a work around for the simplistic nature of the Transformer neural network architecture that all LLMs are based on. The two main limitations of the Transformer that it helps with are: 1) A Transformer is just a fixed-size stack of layers, with a one-way flow of data through the layers from input to output. The fixed number of layers equates to how many "thought" steps the LLM can put into generating each word of output, but good responses to harder questions may require many more steps and iterative thinking... The idea of "think step by step", aka chain of thought, is to have the model break it's response down into a sequence of steps, each building on what came before, so that the scope of each step is withing the capability of the fixed number of layers of the transformer. 2) A Transformer has extremely limited internal memory from one generated word to the next, so telling the model to go one step at a time, feeding its own output back in as input, in effect makes the model's output a kind of memory that makes up for this. So, chain of thought prompting ultimately give the model more thinking steps (more words generated), together with memory of what it is thinking, in order to be able to generate a better response.
- bachmeier 10mo ago> Programmers resistance to AI assisted programming has lowered considerably. Even if LLMs make mistakes, the ability of LLMs to deliver useful code and hints improved to the point most skeptics started to use LLMs anyway: now the return on the investment is acceptable for many more folks. I'm not a fan of this phrasing. Use of the terms "resistance" and "skeptics" implies they were wrong. It's important we don't engage in revisionist history that allows people in the future to say "Look at the irrational fear programmers had of AI, which turned out to be wrong!" The change occurred because LLMs are useful for programming in 2025 and the earliest versions weren't for most programmers. It was the technology that changed.
- 20k 10mo agoIts also significantly lowered because management is forcing AI on everyone at gunpoint, and saying that you'll lose your job if you don't love AI That's a very easy way to get everyone to pinky promise that they absolutely love AI to the ends of the earth
- deleted 10mo ago[deleted]
- Aurornis 10mo ago> The change occurred because LLMs are useful for programming in 2025 But the skeptics and anti-AI commenters are almost as active as ever, even as we enter 2026. The debate about the usefulness of LLMs has grown into almost another culture war topic. I still see a constant stream of anti-AI comments on HN and every other social platform from people who believe the tools are useless, the output is always unusable, people who mock any idea that operator skill has an impact on LLM output, or even claims that LLMs are a fad that will go away. I’m a light LLM user ($20/month plan type of usage) but even when I try to share comments about how I use LLMs or tips I’ve discovered, I get responses full of vitriol and accusations of being a shill.
- zahlman 10mo agoIt absolutely is culture war. I can easily imagine a less critical version of myself having ended up in that camp. It comes across to me that the perspective is informed by core values and principles surrounding what "intelligence" is. I butted heads with many earlier on, and they did nothing to challenge that frame meaningfully. What did change is my perception of the set of tasks that don't require "intelligence". And the intuition pump for that is pretty easy to start — I didn't suppose that Deep Blue heralded a dawn of true "AI", either, but chess (and now Go) programs have only gotten even more embarrassingly stronger. Even if researchers and puzzle enthusiasts might still find positions that are easier for a human to grok than a computer.
- erichocean 10mo ago> 1. NOT have any representation about the meaning of the prompt. This one is bizarre, if true (I'm not convinced it is). The entire purpose of the attention mechanism in the transformer architecture is to build this representation, in many layers (conceptually: in many layers of abstraction). > 2. NOT have any representation about what they were going to say. The only place for this to go is in the model weights. More parameters means "more places to remember things", so clearly that's at least a representation. Again: who was pushing this belief? Presumably not researchers, these are fundamental properties of the transformer architecture. To the best of my knowledge, they are not disputed. > I believe [...] it is not impossible they get us to AGI even without fundamentally new paradigms appearing. Same, at least for the OpenAI AGI definition: "An AI system that is at least as intelligent as a normal human, and is able to do any economically valuable work."
- zahlman 10mo ago> This one is bizarre, if true (I'm not convinced it is). > The entire purpose of the attention mechanism in the transformer architecture is to build this representation, in many layers (conceptually: in many layers of abstraction). I think this is really about a hidden (i.e. not readily communicated) difference in what the word "meaning" means to different people.
- erichocean 10mo agoCould be, by "meaning" I mean (heh) that transformers are able to distinguish tokens (and prompts) in a consequential ("causal") way, and that they do so at various levels of detail ("abstractions"). I think that's the usual understanding of how transformer architectures work, at the level of math.
- etra0 10mo agoLLMs have certainly become extremely useful for Software Engineers, they're very convincing (and pleasers, too) and I'm still unsure about the future of our day-to-day job. But one thing that has scared me the most, is the trust of LLMs output to the general society. I believe that for software engineers it's really easy to see if it's being useful or not -- We can just run the code and see if the output is what we expected, if not, iterate it, and continue. There's still a professional looking to what it produces. On the contrary, for more day-to-day usage of the general pubic, is getting really scary. I've had multiple members of my family using AI to ask for medical advice, life advice, and stuff were I still see hallucinations daily, but at the same time they're so convincing that it's hard for them not to trust them. I still have seen fake quotes, fake investigations, fake news being spreaded by LLMs that have affected decisions (maybe, not as crucials yet but time will tell) and that's a danger that most software engineers just gross over. Accountability is a big asterisk that everyone seems to ignore
- santadays 10mo agoI get this take, but given the state of the world (the US anyways), I find it hard to trust anyone with any kind of profit motive. I feel like any information can’t be taken as fact, it can just be rolled into your world view and discarded if useful or not. If you need to make a decision that can’t be backed out of that has real world consequences I think/hope most people are learning to do as much due diligence as reasonable. Llms seem at this moment to be trying to give reliable information. When they’ve been fine tuned to avoid certain topics it’s obvious. This could change but I suspect it will be hard to find tune them too far in a direction without losing capability. That said, it definitely feels as though keeping a coherent picture of what is actually happening is getting harder, which is scary.
- twoodfin 10mo agoI feel like any information can’t be taken as fact, it can just be rolled into your world view and discarded if useful or not. The concern, I think, is that for many that “discard function” is not, “Is this information useful?”. Instead: “Does this information reinforce my existing world view?” That feedback loop and where it leads is potentially catastrophic at societal scale.
- pton_xd 10mo ago> For years, despite functional evidence and scientific hints accumulating, certain AI researchers continued to claim LLMs were stochastic parrots: probabilistic machines that would: 1. NOT have any representation about the meaning of the prompt. 2. NOT have any representation about what they were going to say. In 2025 finally almost everybody stopped saying so. It's interesting that Terrence Tao just released his own blog post stating that they're best viewed as stochastic generators. True he's not an AI researcher, but it does sound like he's using AI frequently with some success. "viewing the current generation of such tools primarily as a stochastic generator of sometimes clever - and often useful - thoughts and outputs may be a more productive perspective when trying to use them to solve difficult problems" [0]. [0] https://mathstodon.xyz/@tao/115722360006034040 https://mathstodon.xyz/@tao/115722360006034040
- antirez 10mo agoWhat happened recently is that all the serious AI researches that were in the stochastic parrot side changed point of view but, incredibly, people without a deep understanding on such matters, previously exposed to such arguments, are lagging behind and still repeat arguments that the people who popularized them would not repeat again. Today there is no top AI scientist that will tell you LLMs are just stochastic parrots.
- geraneum 10mo agoNow that you’re here, what do you mean by “scientific hints” in your first paragraph?
- visarga 10mo agoThe stochastic parrot framing makes some assumptions, one of them being that LLMs generate from minimal input prompts, like "tell me about Transformers" or "draw a cute dog". But when input provides substantial entropy or novelty, the output will not look like any training data. And longer sessions with multiple rounds of messages also deviate OOD. The model is doing work outside its training distribution. It's like saying pianos are not creative because they don't make music. Well, yes, you have to play the keys to hear the music, and transformers are no exception. You need to put in your unique magic input to get something new and useful.
- bgwalter 10mo agoThey are very advanced stochastic parrots that allow AI invested authors to suddenly write in perfect English. If Antirez has never gotten an LLM to perform an absolutely embarrassing mistake, he must be very lucky or we should stop listening to him. Programmers' resistance has not weakened. Since the ORCL drop of 40% anti-LLM opinions are censored and downvoted here. Many people have given up, and we always get articles from the same LLM influencers.
- mwkaufma 10mo agoA list of unverifiable claims, stated authoritatively. The lady doth protest too much.
- linhns 10mo agoThe post is about his opinions.
- jennyholzer2 10mo agoreads more like propaganda.
- lowsong 10mo agoI'm impressed that such a short post can be so categorically incorrect. > For years, despite functional evidence and scientific hints accumulating, certain AI researchers continued to claim LLMs were stochastic parrots > In 2025 finally almost everybody stopped saying so. There is still no evidence that LLMs are anything beyond "stochastic parrots". There is no proof of any "understanding". This is seeing faces in clouds. > I believe improvements to RL applied to LLMs will be the next big thing in AI. With what proof or evidence? Gut feeling? > Programmers resistance to AI assisted programming has lowered considerably. Evidence is the opposite, most developers do not trust it. https://survey.stackoverflow.co/2025/ai#2-accuracy-of-ai-tools https://survey.stackoverflow.co/2025/ai#2-accuracy-of-ai-too... > It is likely that AGI can be reached independently with many radically different architectures. There continues to be no evidence beyond "hope" that AGI is even possible, yet alone that Transformer models are the path there. > The fundamental challenge in AI for the next 20 years is avoiding extinction. Again, nothing more than a gut feeling. Much like all the other AI hype posts this is nothing more than "well LLMs sure are impressive, people say they're not, but I think they're wrong and we will make a machine god any day now".
- crystal_revenge 10mo agoStrongly agree with this comment. Decoder-only LLMs (the ones we use) are literally Markov Chains, the only (and major) difference is a radically more sophisticated state representation. Maybe "stochastic parrot" is overly dismissive sounding, but it's not a fundamentally wrong understanding of LLMs. The RL claims are also odd because, for starters, RLHF is not "reinforcement learning" based on any classical definition of RL (which almost always involve an online component). And further, you can chat with anyone who has kept up with the RL field, and quickly realize that this is also a technology that still hasn't quite delivered on the promises it's been making (despite being an incredibly interesting area of research). There's no reason to speculate that RL techniques will work with "agents" where they have failed to achieve wide spread success in similar domains. I continue to be confused why smart, very technical people can't just talk about LLMs honestly. I personally think we'd have much more progress if we could have conversations like "Wow! The performance of a Markov Chain with proper state representation is incredible, let's understand this better..." rather than "AI is reasoning intelligently!" I get why non-technical people get caught up in AI hype discussions, but for technical people that understand LLMs it seems counter productive. Even more surprising to me is that this hype has completely destroyed any serious discussions of the technology and how to use it. There's so much oppurtunity lost around practical uses of incorporating LLMs into software while people wait for agents to create mountains of slop.
- jimmydoe 10mo ago> * The fundamental challenge in AI for the next 20 years is avoiding extinction. sorry, I say it's folding the laundry. with an aging population, that's the most, if not only, useful thing.
- mrdependable 10mo agoThese comments are a bit scary. It feels like LLMs managed to exploit some fault in the human psyche. I think the biggest danger of this technology is that people are not mentally equipped to handle it.
- jennyholzer2 10mo agoChatGPT and Claude Code are industrial strength fans designed to blow smoke up your ass at rates once thought impossible
- unbelievably 10mo ago[dead]
- akomtu 10mo agoThe fault is well known: chatbots are bootlickers. They always praise users and never criticize them, so chatbots are quickly promoted to the personal advisor position. The AI of Sauron of technological age.
- moab 10mo agoThis is a very real worry for the AI rollout for the general population. But are folks here using AI to blow smoke up their asses as a sibling comment stated? I'd like to believe we're using it to ask questions, prototype, and then measure... not just blow smoke up there...
- deleted 10mo ago[deleted]
- lolz404 10mo agoThis article does little to support its claims but was a good primer to dive into some topics. They are cool new tools use them where you can but there is a ton of research still left to do. Just lols at the hubris silicon valley will make something so smart it extincts humankind. It'll happen from the lack of water and heated planet first :) The stocastic parrot argument is still debated but more nuanced than before. Although the original author still stands by the statement. Evidence of internal planning per model. Anthropic Attribution Graphs Research with some rhyming did support it but gemma didn't. The idea of "understanding" is still up for debate as well. Sure, when models are directly trained on data there is representation. Othello-GPT Studies was one way to support but that was during training so some interal representation was created. Out of distribution task will collapse to confabulation. Apple's GSM-Symbolic Research seems to support that. Chain of thought is a helpful tool but is untrustworthy at best. Anthropic themselves have showed this https://www.anthropic.com/research/reasoning-models-dont-say-think https://www.anthropic.com/research/reasoning-models-dont-say...
- bgwalter 10mo agoRegarding the stochastic parrots: It is easy to see that LLMs exclusively parrot by asking them about current political topics [1], because they cannot plagiarize settled history from Wikipedia and Britannica. But of course there also is the equivalence between LLMs and Markov chains. As far as I can see, it does not rely on absurd equivalences like encoding all possible output states in an infinite Markov chain: https://arxiv.org/abs/2410.02724 https://arxiv.org/abs/2410.02724 Then there is stochastic parrot research: https://arxiv.org/abs/2502.08946 https://arxiv.org/abs/2502.08946 "The stochastic parrot phenomenon is present in LLMs, as they fail on our grid task but can describe and recognize the same concepts well in natural language." As said above, this is obvious to anyone who has interacted with LLMs. Most researchers know what is expected of them if they want to get funding and will not research the obvious too deeply. [1] They have Internet access of course.
- russfink 10mo agoPractical question: when getting the AI to teach you something, eg how attention can be focused in LLMs, how do you know it’s teaching you correct theory? Can I use a metric of internal consistency, repeatedly querying it and other models with a summary of my understanding? What do you all do?
- jennyholzer2 10mo ago[flagged]
- layer8 10mo ago> What do you all do? Google for non-AI sources. Ask several models to get a wider range of opinions. Apply one’s own reasoning capabilities where applicable. Remain skeptical in the absence of substantive evidence. Basically, do what you did before LLMs existed, and treat LLM output like you would have a random anonymous blog post you found.
- akomtu 10mo agoIn that case, LLMs must be written off as very knowledgeable crackpots because of their tendency to make things up. That's how we would treat a scientist who's caught making things up.
- roughly 10mo ago> A few well known AI scientists believe that what happened with Transformers can happen again, and better, following different paths, and started to create teams, companies to investigate alternatives to Transformers and models with explicit symbolic representations or world models. I’m actually curious about this and would love pointers to the folks working in this area. My impression from working with LLMs is there’s definitely a “there” there with regards to intelligence - I find the work showing symbolic representation in the structure of the networks compelling - but the overall behavior of the model seems to lack a certain je ne sais quoi that makes me dubious that they can “cross the divide,” as it were. I’d love to hear from more people that, well, sais quoi, or at least have theories.
- gaigalas 10mo agoThis post is a bait for enthusiasts. I like it. > Chain of thought is now a fundamental way to improve LLM output. That kinda proves _that LLMs back then were pretty much stochastic parrots indeed_, and the skeptics were right at the time. Today, enthusiasts agree with what they previously said: without CoT, the AI feels underwhelming, repetitive and dumb and it's obvious that something more was needed. Just search past discussions about it, people were saying the problem would be solved with "larger models" (just repeating marketing stuff) and were oblivious to the possibility of other kinds of innovations. > The fundamental challenge in AI for the next 20 years is avoiding extinction. That is a low level sick burn on whoever believes AI will be economically viable short-term. And I have to agree.
- ofirpress 10mo ago> There are certain tasks, like improving a given program for speed, for instance, where in theory the model can continue to make progress with a very clear reward signal for a very long time. Yup, this will absolutely be a big driver of gains in AI for coding in the near future. We actually built a benchmark based on this exact principle: https://algotune.io/ https://algotune.io/
- crystal_revenge 10mo agoI wish people would be more vocal in calling out that LLMs have unquestionably failed to deliver on the 2022-2023 promises of exponential improvement at the foundation model level. Yes they have improved, and there is more tooling around them, but clearly the difference between LLMs in 2025 and 2023 is not as large as 2023 and 2021. If there was truly exponential progress, there would be no possibility of debating this. Which makes comments like this: > The fundamental challenge in AI for the next 20 years is avoiding extinction. Seem to be almost absurd without further, concrete justification. LLMs are still quite useful, I'm glad they exist and honestly am still surprised more people don't use them in software. Last year I was very optimistic that LLMs would entirely change how we write software by making use of them as a fundamental part of our programming tool kit (in a similar way that ML fundamentally changed the options available to programmers for solving problems). Instead we've just come up with more expensive ways to extend the chat metaphor (the current generation of "agents" is disappointingly far from the original intent of agents in AI/CS). The thing I am increasingly confused about is why so many people continue to need LLMs to be more than they obviously are. I get why crypto boosters exist, if I have 100 BTC, I have a very clear interest getting others to believe that they are valuable. But with "AI", I don't quite get, for the non-VS/founder, why it matters that people start foaming out the mouth over AI rather than just using it for the things it's good at. Though I have some growing sense that this need is related to another trend I've personally started with witness: AI psychosis is very real. I personally know an increasing number of people who are spiraling into an LLM induced hallucinated world. The most shocking was someone talking about how losing human relationships is inevitable because most people can't keep up with those enhanced by AI acceleration. On the softer end I know more and more people who quietly confess how much they let AI work as a perpetual therapist, guiding their every decision (which is more than most people would let a human therapist guide there directions).
- redlock 10mo ago“But clearly the difference between LLMs in 2025 and 2023 is not as large as between 2023 and 2021.” This is a ridiculous statement. A simple example of the huge difference is context size. GPT-4 was, what, 8K? Now we’re in the millions with good retention. And this is just context size, let alone reasoning, multimodality, etc.
- phlummox 10mo ago> For years, despite functional evidence and scientific hints accumulating, certain AI researchers continued to claim LLMs were stochastic parrots: probabilistic machines that would: 1. NOT have any representation about the meaning of the prompt. 2. NOT have any representation about what they were going to say. But did any AI researchers actually claim there was no representation of meaning? I thought generally, the criticism of LLMs was that while they do abstract from their corpus - ie, you can regard them as having a representation of "meaning" - it's tightly and inextricably tied to the surface level representation, it isn't grounded in models of the external world, and LLMs have poor ability to transfer that knowledge to other surface encodings. I don't know who the "certain AI researchers" are supposed to be. But the "stochastic parrot" paper by Bender et al [1] says: > Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. That's a very different objection to the one antirez describes - I think he's erecting a straw man. But I'd be happy to be corrected by anyone more familiar with the research. [1] https://dl.acm.org/doi/10.1145/3442188.3445922 https://dl.acm.org/doi/10.1145/3442188.3445922
- antirez 10mo ago> Text generated by an LM is not grounded in communicative intent This means exactly that no representation should exist in the activation states about what the model wants to tell, and there must be only a single token probabilistic inference at play. Also their model requires the contrary, too: that the model does not know, semantically, what the query really means. Stochastic Parrot has a scientific meaning, and just only observing the function of the models, it is quite evident that they were very wrong, but now we have stong evidence (via probing) that also the sentence you quoted is not correct, since the model knows the idea to express also in general terms, and features about things it is going to say much later activates a lot of tokens earlier, including conceptual features that are relevant later in the sentence / concept expressed. You are doing the big error that is common to do in this context of extending the stochastic parrot to a non scientifically isolated model that can be made large enough to accomodate any evidence arriving from new generations of models. The stochastic parrot does not understand the query nor is trying to reply to you in any way, it just exploits a probabilistic link among the context window and the next word. This link can be more complex than a Markov chain but must be of the same kind: lacking understanding whatsoever and communication intent (no representation of the concept / sentences that are required to reply correctly). How it is possible to believe in this, today? And, check yourself what the top AI scientists today believe about the correctness of the stochastic parrot hypothesis.
- rldjbpin 10mo agothe reflections felt like a mixed bag between someone who seems to know about the technical aspects deeper than an average person, while simultaneously being like an astrologist. personally, as someone building on top of gen AI for a living, i finally bit the bullet on building using LLMs. it did reduce friction in things i don't like doing and did not explore as much. by acting as a catalyst when i needed to finally address them, it helped me get going and eventually become proficient in the core tech itself. outside of work, however, i find people around me use the services much more than i do. sometimes it felt like the "big data is like teenage sex"[1], but some aspects were quite genuine. got better appreciation after trying them to better understand other people's perspective and to design better. with "slop" as word of the year and people wondering if a random clip is AI, now more than ever the effects in general life seems apparent. it is not as sexy as "i will lose my job soon", but the effects are here and now. while the next year will be even more interesting, i can't wait for the bubble to burst. [1] https://hewlett.org/is-big-data-like-teenage-sex/ https://hewlett.org/is-big-data-like-teenage-sex/
- AdamWills 10mo ago[dead]
- Joel_LeBlanc 9mo agoIt's fascinating to see how AI is reshaping the landscape for digital assets—buying websites or e-commerce stores has become more accessible than ever. When evaluating potential investments, I always stress the importance of thorough due diligence; I've found that using tools like DREA (Digital Real Estate Analyzer) can really streamline the process and provide valuable insights. It's all about understanding the numbers and the potential for growth, especially in such a dynamic environment. What specific metrics are you focusing on?