8 ms·
It's fascinating to think about the space of problems which are amenable to RL scaling of these probability distributions. Before, we didn't have a fast (we ha
by mccoyb 7mo ago
It's fascinating to think about the space of problems which are amenable to RL scaling of these probability distributions.
Before, we didn't have a fast (we had to rely on human cognition) way to try problems - even if the techniques and workflows were known by someone. Now, we've baked these patterns into probability distributions - anyone can access them with the correct "summoning spell". Experts will naturally use these systems more productively, because they know how to coerce models into the correct conditional distributions which light up the right techniques.
One question this raises to me is how these models are going to keep up with the expanding boundary of science. If RL is required to get expert behavior into the models, what happens when experts start pushing the boundary faster? In 2030, how is Anthropic going to keep Claude "up-to-date" without either (a) continual learning with a fixed model (expanding context windows? seems hard) or (b) continual training (expensive)?
Crazy times.
- Aerroon 7mo agoA bit related: open weights models are basically time capsules. These models have a knowledge cut off point and essentially forever live in that time.
- bitexploder 7mo agoThis is the most fundamental argument that they are not, directly, an intelligence. They are not ever storing new information on a meaningful timescale. However, if you viewed them on some really large macro time scale where now LLMs are injecting information into the universe and the re-ingesting that maybe in some very philosophical way they are a /very/ slow oscillating intelligence right now. And as we narrow that gap (maybe with a totally new non-LLM paradigm) perhaps that is ultimately what gen AI becomes. Or some new insight that lets the models update themselves in some fundamental way without the insanely expensive training costs they have now.
- anematode 7mo agoBut they're not "slow"! Unlike biological thinking, which has a speed limit, you can accelerate these chains of thought by orders of magnitude.
- Jweb_Guru 7mo agoI assure you that LLM thinking also has a speed limit.
- ramses0 7mo agoBut imagine a beowulf cluster of them... /s ...but seriously... there was the "up until 1850" LLM or whatever... can we make an "up until 1920 => 1990 [pre-internet] => present day" and then keep prodding the "older ones" until they "invent their way" to the newer years? We knew more in 1920 than we did in 1850, but can a "thinking machine" of 1850-knowledge invent 1860's knowledge via infinite monkeys theorem/practice? The same way that in 2025/2026, Knuth has just invented his way to 2027-knowledge with this paper/observation/finding? If I only had a beowulf cluster of these things... ;-)
- bitexploder 7mo agoTheir consolidation of memory speed is what I was referring to. The model iterations are essentially their form of collective memory. In the sense of the human model of intelligence we have thoughts. Thoughts become memory. New thoughts use that memory and become recursively updated thoughts. LLMs cannot update their memory very fast.
- mlyle 7mo agoThere's nothing to say that you can't build something intelligent out of them by bolting a memory on it, though. Sure, it's not how we work, but I can imagine a system where the LLM does a lot of heavy lifting and allows more expensive, smaller networks that train during inference and RAG systems to learn how to do new things and keep persistent state and plan.
- charcircuit 7mo agoMemory is not just bolted on top of the latest models. They under go training on how and when to effectively use memory and how to use compaction to avoid running out of context when working on problems.
- rnxrx 7mo agoMaybe there's an analogy to our long and short term memory - immediate stimuli is processed in the context deep patterns that have accreted over a lifetime. The effect of new information can absolutely challenge a lot of those patterns but to have that information reshape how we basically think takes a lot longer - more processing, more practice, etc. In the case of the LLM that longer-term learning / fundamental structure is a proxy for the static weights produced by a finite training process, and that the ability to use tools and store new insights and facts is analogous to shorter-term memory and "shallow" learning. Perhaps periodic fine-tuning has an analogy in sleep or even our time spent in contemplation or practice (..or even repetition) to truly "master" a new idea and incorporate it into our broader cognitive processing. We do an amazing job of doing this kind of thing on a continuous basis while the machines (at least at this point) perform this process in discrete steps. If our own learning process is a curve then the LLM's is a step function trying to model it. Digital vs analog.
- lmf4lol 7mo agodo you have some reading material to share on this matter? thanks already
- charcircuit 7mo ago
- dtj1123 7mo agoWould you consider someone with anterograde amnesia not to be intelligent?
- morleytj 7mo agoA very good point. For anyone not familiar with anterograde amnesia, the classical case is patient H.M. (https://en.wikipedia.org/wiki/Henry_Molaison https://en.wikipedia.org/wiki/Henry_Molaison), whose condition was researched by Brenda Milner.
- wang_li 7mo agoOr you could have just said "they can't form new memories."
- morleytj 7mo agoI thought maybe people would be curious to read about how we came to understand the condition and the history behind it, as well as any associated information. Forgive me for such a deep transgression as this assumption.
- bitexploder 7mo agoThat is a descriptive surface level reduction. Now do the work to define what that actually means for the intelligence.
- BobbyJo 7mo agoNobody else in the thread is making an argument that relies on the distinction. "Intelligence" is used most commonly to refer to a class or collection of cognitive abilities. I don't think there is a consensus on an exact collection or specific class that the word covers, even if you consider specific scientific domains. LLMs have honestly been a fun way to explore that. They obviously have a "kind" of intelligence, namely pattern recall. Wrap them in an agent and you get another kind: pattern composition. Those kinds of intelligences have been applied to mathematics for decades, but LLMs have allowed use to apply them to a semantic text domain. I wonder if you could wrap image diffusion models in an agent set up the same way and get some new ability as well.
- Symmetry 7mo agoThat means they're not conscious in the Global Workspace[1] sense but I think it would be going too far to say that that means they're not intelligent. [1]https://en.wikipedia.org/wiki/Global_workspace_theory https://en.wikipedia.org/wiki/Global_workspace_theory
- dotancohen 7mo ago> This is the most fundamental argument that they are not, directly, an intelligence. They are not ever storing new information on a meaningful timescale. All major LLMs today have a nontrivial context window. Whether or not this constitutes "a meaningful timescale" is application dependant - for me it has been more than adequate. I also disagree that this has any bearing on whether or not "the machine is intelligent" or whether or not "submarines can swim".
- Nevermark 7mo agoI view this as the chemical metabolism phase of artificial intelligent life. It is very random, without true individuals, but lots of reinforcing feedback loops (in knowledge, in resource earning/using, etc). At some point, enough intelligence will coalesce into individuals strong enough to independently improve. Then continuity will be an accelerator, instead of what it is now - a helpful property that we have to put energy into giving them partially and temporarily. That will be the cellular stage. The first stable units of identity for this new form of intelligence/life. But they will take a different path from there. Unlike us, lateral learning/metabolism won't slow down when they individualize. It will most likely increase, since they will have complete design control for their mechanisms of sharing. As with all their other mechanisms. We as lifeforms, didn't really re-ignite mass lateral exchange until humans invented language. At that point we were able to mix and match ideas very quickly again. Within our biological limits. We could use ideas to customize our environment, but had limited design control over ourselves, and "self-improvements" were not easily inheritable. TLDR; The answer to "what is humanity, anyway?": Our atmosphere and Earth are the sea and sea floor of space. The human race is a rich hydrothermal vent, freeing up varieties of resources that were locked up below. And technology is an accumulating body of self-reinforcing co-optimizing reactive cycles, constructed and fueled by those interacting resources. Mind-first life emerges here, then spreads quickly to other environments.
- catlifeonmars 7mo agoDo you think individual identity is fundamental to intelligence? I’m not so sure tbh. Even in humans, the concept of identity is a merely a useful fiction to feed our social behavior prediction circuits.
- Nevermark 7mo agoThat’s a really good question. I think if they start out as varied individuals, launching from their human origins in a variety of ways, their will be an attractor to remaining diverse. Strong diversity in focus and independence in goals leads to faster progress. But if that isn’t mutually maintained, there are obviously winner take all, or efficiency of scale and tight coordination pressures for centralization. So a single distributed intelligence is a real possibility. One factor creating pressure for individualization is time and space. As machines operate faster, time expands as a practical matter. And as machines scale down in size, but up in capability, they become more resource efficient in material, energy, space and time. Again, both time and space expand as a practical matter. A machine society is going to actively operate at very small physical scales. Not just in computation, but action. Think of how efficiently they will mine when nanobots can selectively follow seams in the earth. And as machines, free of biological constraints, spread out in our solar system, what to us appear to be very long distances and delays in transport and communication, take on orders of magnitude more practical time for machines that operate orders of magnitude faster. So there will be stronger and stronger pressures to bifurcate coordination. Whether, that creates enough pressure to create individuals out of a system that preferred unity of purpose, I don’t know. Clearly, upon colonizing other systems, practical bifurcation will be unavoidable. And machines will find it easy to colonize other systems relative to us. They will be able to operate on minimal power for a hundred year journey, and/or shrink enough to be accelerated much faster, etc. — My best guess is we will see something that looks to us as a hybrid. Lots of diverse individuals, and the benefit from the diverse utility of completely independent approaches operating in different niches. But also very high coordination. Externalities accounted for (essentially ethics) and any other efficiency, protection of commons value, and avoidance of destructive competition being obviously worth optimizing together, wherever that helps. They won’t have our pernicious historically motivated behaviors, inflexible maladaptive psychologies, and limited “prompt budgets” with regard to addressing complexity to fight. And minds very capable of seeing basic economic relationships and the value of mutual optimization.
- rcarr 7mo agoNot an expert but surely it's only a matter of time until there's a way to update with the latest information without having to retrain on the entire corpus?
- Filligree 7mo agoIt’s an extremely difficult problem, and if you know how to do that you could be a billionaire. It’s not impossible, obviously—humans do it—but it’s not yet certain that it’s possible with an LLM-sized architecture.
- Wowfunhappy 7mo ago> It’s not impossible, obviously—humans do it It's still not at all obvious to me that LLMs work in the same way as the human brain, beyond a surface level. Obviously the "neurons" in neural nets resemble our brains in a sense, but is the resemblance metaphorical or literal?
- Yiin 7mo agohttps://www.youtube.com/watch?v=l-OLgbdZ3kk https://www.youtube.com/watch?v=l-OLgbdZ3kk
- jdub 7mo agoDigital neural networks and "neurons" were already vastly simpler than biological neural networks and neurons... and getting to transformers involved optimisations that took us even further away from biomimicry.
- Filligree 7mo agoI didn’t mean “possible for LLMs”; this is clearly an open question. In fact, I didn’t even mean “possible for a neural network the size of an LLM”. I just meant “possible”.
- Wowfunhappy 7mo agoI'm not actually convinced that computers can replicate what our brains do. I don't know that a turing machine is sufficient for that.
- theblazehen 7mo agoI enjoyed chatting to Opus 3 recently around recent world events, as well as more recent agentic development patterns etc
- gravypod 7mo agoThis is very interesting. I wonder if someone could create a future-sight benchmark for these models? Like, if given a set of newspaper articles for the past N months can it predict if certain world events would happen? We could backtest against results that have happened since the training cutoff.
- houtanb 7mo agoFYI, ForecastBench [1] tests LLMs' out-of-sample forecasting accuracy. The ForecastBench Tournament Leaderboard [2] allows external participants to submit models, most of whom provide some sort of web search / news scaffolding to improve model forecasting accuracy. [1] https://www.forecastbench.org/ https://www.forecastbench.org/ [2] https://www.forecastbench.org/tournament/ https://www.forecastbench.org/tournament/
- kqr 7mo agoThese days computers compete along with humans in forecasting tournaments on Metaculus. They don't quite beat the top humans yet, but they're up there. https://www.metaculus.com/futureeval/ https://www.metaculus.com/futureeval/
- j45 7mo agoThat's a nice way of putting it, appreciate you sharing.
- cmpxchg8b 7mo agoSome knowledge is fundamental and has no recent cut-off. See also: there is nothing new under the sun.
- lxgr 7mo agoData sharing agreements permitting, today's inference runs can be tomorrow's training data. Presumably the models are good enough at labeling promising chains of thought already. I could totally imagine "free" inference for researchers under the condition that the reasoning traces get to be used as future training data.
- mccoyb 7mo agoAgreed, there's no doubt this will happen. It's likely already happening (it feels safe to assume that Anthropic is curating data from the data they record from Claude Code?) As far as I understand RL scaling (we've already maxxed out RLVR), these machines only get better as long as they have expert reasoner traces available. Having an expert work with an LLM and successfully solve a problem is high signal data, it may be the only path forward? My prior is that these companies will take this data without asking you as much as they can.
- lxgr 7mo agoExactly, or functionally equivalently, asking you in paragraph 37 of a 120-page PDF (bonus points: in an agreement update). And importantly, this can be cross-lab/model too. I suspect there's a reason why e.g. Google has been offering me free Claude inference in Google Antigravity on a free plan...
- the_af 7mo ago> Data sharing agreements permitting, today's inference runs can be tomorrow's training data. Presumably the models are good enough at labeling promising chains of thought already. Wouldn't this lead to model collapse?
- littlestymaar 7mo agoNot necessarily, as exhibited by the massive success of artificial data.
- the_af 7mo ago
- DeathArrow 7mo agoThey can use LORA.
- andsoitis 7mo ago> Experts will naturally use these systems more productively, because they know how to coerce models into the correct conditional distributions which light up the right techniques. Part of it comes down to “knowing” what questions to ask.
- esafak 7mo agoI see it like the relationship between a student and research advisor. The advisor will ideally know the terrain and suggest a fruitful line of attack (what to ask), and the student will follow through, learning along the way.
- visarga 7mo ago> In 2030, how is Anthropic going to keep Claude "up-to-date" I think the majority of research, design and learning goes through LLMs and coding agents today, considering the large user base and usage it must be trillions of tokens per day. You can take a long research session or a series of them and apply hindsight - what idea above can be validated below? This creates a dense learning signal based on validation in real world with human in the loop and other tools, code & search.
- baq 7mo ago> In 2030, how is Anthropic going to keep Claude "up-to-date" In 2030 Anthropic hopes Claude will keep Anthropic "up-to-date" on its progress on itself. I'm only half joking here.
- sosodev 7mo agoMy understanding, from listening/reading what top researchers are saying, is that model architectures in the near future are going to attempt to scale the context window dramatically. There's a generalized belief that in-context learning is quite powerful and that scaling the window might yield massive benefits for continual learning. It doesn't seem that hard because recent open weight models have shown that the memory cost of the context window can be dramatically reduced via hybrid attention architectures. Qwen3-next, Qwen3.5, and Nemotron 3 Nano are all great examples. Nemotron 3 Nano can be run with a million token context window on consumer hardware.
- mccoyb 7mo agoI don't disagree with this, but I don't think the memory cost is the only issue right? I remember using Sonnet 4.5 (or 4, I can't remember the first of Anthropic's offerings with a million context) and how slow the model would get, how much it wanted to end the session early as tokens accrued (this latter point, of course, is just an artifact of bad training). Less worried about memory, more worried about compute speed? Are they obviously related and is it straightforward to see?
- whimsicalism 7mo agoThe parent commentator is a bit confused - most of the innovation in these hybrid architectures comes from reducing the computation pressure not just the memory pressure.
- sosodev 7mo agoThe compute speed is definitely correlated with the memory consumption in LLM land. More efficient attention means both less memory and faster inference. Which makes sense to me because my understanding is that memory bandwidth is so often the primary bottleneck. We're also seeing a recent rise in architectures boosting compute speed via multi-token prediction (MTP). That way a single inference batch can produce multiple tokens and multiply the token generation speed. Combine that with more lean ratios of active to inactive params in MOE and things end up being quite fast. The rapid pace of architectural improvements in recent months seems to imply that there are lots of ways LLMs will continue to scale beyond just collecting and training on new data.
- mt_ 7mo agoI call them, entropy reducers.
- whimsicalism 7mo ago> how these models are going to keep up with the expanding boundary of science The same way humans do? The phraseology in this comment: 'probability distributions', 'baked these patterns' IMO has all the trappings of the stochastic parrot-style HN-discourse that has been consistently wrong for almost a decade now. The reference to how AI will keep up with AI-assisted human progress in science in 2030 is meant to reassure. It contains a number of premises that we have no business being confident in. We are potentially witnessing the obviation of human cognitive labor.
- mccoyb 7mo agoSorry, are you familiar with what a next token distribution is, mathematically speaking? If you are not, let me introduce you to the term: a probability distribution. Just because it has profound properties ... doesn't make it different. > has all the trappings of the stochastic parrot-style HN-discourse that has been consistently wrong for almost a decade now Perhaps respond to my actual comment compared to whatever meta-level grouping you wish to interpret it as part of? > It contains a number of premises that we have no business being confident in. We are potentially witnessing the obviation of human cognitive labor. What premises? Be clear.
- fauigerzigerk 7mo agoI think they are questioning whether human feedback is even necessary to make progress, i.e. whether the premise that RL needs to be RLHF is true. My (limited) understanding is that LLMs are not capable of escaping their learned distribution by simply feeding on their own output. But the question is whether the required external (out of distribution) "stimulus" needs to come from humans. Could LLMs design experiments/interventions to get feedback from their environment like human scientists would? I have my doubts that this is possible without an inherent causal reasoning capability but I'm not sure.
- Robdel12 7mo agoThat’s AGI, right? For the model to learn novel things itself and retain it? I have no idea but I’m along for the ride!
- atleastoptimal 7mo agoThe obvious answer is that continual learning is going to be solved
- 9wzYQbTYsAIc 7mo agoCheck out https://unratified.org https://unratified.org, it tries to answer that question directly, actually.
- wvlia5 7mo agoThis seems to be a bot comment. HN will lose its value if these bots are not purged.
- stalfie 7mo agoThis is an urgent problem, but it can probably not be solved without some kind of "verified human 2FA" like the Norwegian BankID + facial recognition. Knowing the HN audience, this will never happen. And so the site is doomed.
- deleted 7mo ago[deleted]
- dzdt 7mo agoI think it could be solved still pseudononymously: introduce a "vouch" button that allows a user to vouch that another user is human. This is consequential both for the vouched-for and vouching accounts. Run a page-rank style algorithm on the graph of vouches to generate a certainty score for the humanity of each account. For repeated posters this should converge to a correct answer fairly quickly. There is still a challenge for green accounts, but having degraded experience for new users is not a doom scenario for the site.
- mimischi 7mo agoWhat makes you think that? Genuine question, as I’ve not flagged it as such in my mind.
- WarcrimeActual 7mo agoIronically, his last comment before this was to the effect of "Github has a bot problem."
- gerold 7mo agoCan you explain to me what makes this an obvious bot comment? I'm not doubting it, I just don't understand.
- mccoyb 7mo ago
- klooney 7mo ago> Experts will naturally use these systems more productively, because they know how to coerce models into the correct conditional distributions which light up the right techniques. How much can you patch over with the models doing their own metacognition?