16 ms·
Claude's Cycles [pdf]
- ibic 7mo agoWow, it's from Donald Knuth.
- mccoyb 7mo agoIt's fascinating to think about the space of problems which are amenable to RL scaling of these probability distributions. Before, we didn't have a fast (we had to rely on human cognition) way to try problems - even if the techniques and workflows were known by someone. Now, we've baked these patterns into probability distributions - anyone can access them with the correct "summoning spell". Experts will naturally use these systems more productively, because they know how to coerce models into the correct conditional distributions which light up the right techniques. One question this raises to me is how these models are going to keep up with the expanding boundary of science. If RL is required to get expert behavior into the models, what happens when experts start pushing the boundary faster? In 2030, how is Anthropic going to keep Claude "up-to-date" without either (a) continual learning with a fixed model (expanding context windows? seems hard) or (b) continual training (expensive)? Crazy times.
- Aerroon 7mo agoA bit related: open weights models are basically time capsules. These models have a knowledge cut off point and essentially forever live in that time.
- bitexploder 7mo agoThis is the most fundamental argument that they are not, directly, an intelligence. They are not ever storing new information on a meaningful timescale. However, if you viewed them on some really large macro time scale where now LLMs are injecting information into the universe and the re-ingesting that maybe in some very philosophical way they are a /very/ slow oscillating intelligence right now. And as we narrow that gap (maybe with a totally new non-LLM paradigm) perhaps that is ultimately what gen AI becomes. Or some new insight that lets the models update themselves in some fundamental way without the insanely expensive training costs they have now.
- anematode 7mo agoBut they're not "slow"! Unlike biological thinking, which has a speed limit, you can accelerate these chains of thought by orders of magnitude.
- Jweb_Guru 7mo agoI assure you that LLM thinking also has a speed limit.
- ramses0 7mo agoBut imagine a beowulf cluster of them... /s ...but seriously... there was the "up until 1850" LLM or whatever... can we make an "up until 1920 => 1990 [pre-internet] => present day" and then keep prodding the "older ones" until they "invent their way" to the newer years? We knew more in 1920 than we did in 1850, but can a "thinking machine" of 1850-knowledge invent 1860's knowledge via infinite monkeys theorem/practice? The same way that in 2025/2026, Knuth has just invented his way to 2027-knowledge with this paper/observation/finding? If I only had a beowulf cluster of these things... ;-)
- bitexploder 7mo agoTheir consolidation of memory speed is what I was referring to. The model iterations are essentially their form of collective memory. In the sense of the human model of intelligence we have thoughts. Thoughts become memory. New thoughts use that memory and become recursively updated thoughts. LLMs cannot update their memory very fast.
- mlyle 7mo agoThere's nothing to say that you can't build something intelligent out of them by bolting a memory on it, though. Sure, it's not how we work, but I can imagine a system where the LLM does a lot of heavy lifting and allows more expensive, smaller networks that train during inference and RAG systems to learn how to do new things and keep persistent state and plan.
- 7mo ago
- rcarr 7mo agoNot an expert but surely it's only a matter of time until there's a way to update with the latest information without having to retrain on the entire corpus?
- Filligree 7mo agoIt’s an extremely difficult problem, and if you know how to do that you could be a billionaire. It’s not impossible, obviously—humans do it—but it’s not yet certain that it’s possible with an LLM-sized architecture.
- Wowfunhappy 7mo ago> It’s not impossible, obviously—humans do it It's still not at all obvious to me that LLMs work in the same way as the human brain, beyond a surface level. Obviously the "neurons" in neural nets resemble our brains in a sense, but is the resemblance metaphorical or literal?
- Yiin 7mo agohttps://www.youtube.com/watch?v=l-OLgbdZ3kk https://www.youtube.com/watch?v=l-OLgbdZ3kk
- jdub 7mo agoDigital neural networks and "neurons" were already vastly simpler than biological neural networks and neurons... and getting to transformers involved optimisations that took us even further away from biomimicry.
- Filligree 7mo agoI didn’t mean “possible for LLMs”; this is clearly an open question. In fact, I didn’t even mean “possible for a neural network the size of an LLM”. I just meant “possible”.
- Wowfunhappy 7mo agoI'm not actually convinced that computers can replicate what our brains do. I don't know that a turing machine is sufficient for that.
- theblazehen 7mo agoI enjoyed chatting to Opus 3 recently around recent world events, as well as more recent agentic development patterns etc
- gravypod 7mo agoThis is very interesting. I wonder if someone could create a future-sight benchmark for these models? Like, if given a set of newspaper articles for the past N months can it predict if certain world events would happen? We could backtest against results that have happened since the training cutoff.
- houtanb 7mo agoFYI, ForecastBench [1] tests LLMs' out-of-sample forecasting accuracy. The ForecastBench Tournament Leaderboard [2] allows external participants to submit models, most of whom provide some sort of web search / news scaffolding to improve model forecasting accuracy. [1] https://www.forecastbench.org/ https://www.forecastbench.org/ [2] https://www.forecastbench.org/tournament/ https://www.forecastbench.org/tournament/
- kqr 7mo agoThese days computers compete along with humans in forecasting tournaments on Metaculus. They don't quite beat the top humans yet, but they're up there. https://www.metaculus.com/futureeval/ https://www.metaculus.com/futureeval/
- j45 7mo agoThat's a nice way of putting it, appreciate you sharing.
- cmpxchg8b 7mo agoSome knowledge is fundamental and has no recent cut-off. See also: there is nothing new under the sun.
- lxgr 7mo agoData sharing agreements permitting, today's inference runs can be tomorrow's training data. Presumably the models are good enough at labeling promising chains of thought already. I could totally imagine "free" inference for researchers under the condition that the reasoning traces get to be used as future training data.
- mccoyb 7mo agoAgreed, there's no doubt this will happen. It's likely already happening (it feels safe to assume that Anthropic is curating data from the data they record from Claude Code?) As far as I understand RL scaling (we've already maxxed out RLVR), these machines only get better as long as they have expert reasoner traces available. Having an expert work with an LLM and successfully solve a problem is high signal data, it may be the only path forward? My prior is that these companies will take this data without asking you as much as they can.
- lxgr 7mo agoExactly, or functionally equivalently, asking you in paragraph 37 of a 120-page PDF (bonus points: in an agreement update). And importantly, this can be cross-lab/model too. I suspect there's a reason why e.g. Google has been offering me free Claude inference in Google Antigravity on a free plan...
- the_af 7mo ago> Data sharing agreements permitting, today's inference runs can be tomorrow's training data. Presumably the models are good enough at labeling promising chains of thought already. Wouldn't this lead to model collapse?
- littlestymaar 7mo agoNot necessarily, as exhibited by the massive success of artificial data.
- the_af 7mo ago
- DeathArrow 7mo agoThey can use LORA.
- andsoitis 7mo ago> Experts will naturally use these systems more productively, because they know how to coerce models into the correct conditional distributions which light up the right techniques. Part of it comes down to “knowing” what questions to ask.
- esafak 7mo agoI see it like the relationship between a student and research advisor. The advisor will ideally know the terrain and suggest a fruitful line of attack (what to ask), and the student will follow through, learning along the way.
- visarga 7mo ago> In 2030, how is Anthropic going to keep Claude "up-to-date" I think the majority of research, design and learning goes through LLMs and coding agents today, considering the large user base and usage it must be trillions of tokens per day. You can take a long research session or a series of them and apply hindsight - what idea above can be validated below? This creates a dense learning signal based on validation in real world with human in the loop and other tools, code & search.
- baq 7mo ago> In 2030, how is Anthropic going to keep Claude "up-to-date" In 2030 Anthropic hopes Claude will keep Anthropic "up-to-date" on its progress on itself. I'm only half joking here.
- sosodev 7mo agoMy understanding, from listening/reading what top researchers are saying, is that model architectures in the near future are going to attempt to scale the context window dramatically. There's a generalized belief that in-context learning is quite powerful and that scaling the window might yield massive benefits for continual learning. It doesn't seem that hard because recent open weight models have shown that the memory cost of the context window can be dramatically reduced via hybrid attention architectures. Qwen3-next, Qwen3.5, and Nemotron 3 Nano are all great examples. Nemotron 3 Nano can be run with a million token context window on consumer hardware.
- mccoyb 7mo agoI don't disagree with this, but I don't think the memory cost is the only issue right? I remember using Sonnet 4.5 (or 4, I can't remember the first of Anthropic's offerings with a million context) and how slow the model would get, how much it wanted to end the session early as tokens accrued (this latter point, of course, is just an artifact of bad training). Less worried about memory, more worried about compute speed? Are they obviously related and is it straightforward to see?
- whimsicalism 7mo agoThe parent commentator is a bit confused - most of the innovation in these hybrid architectures comes from reducing the computation pressure not just the memory pressure.
- sosodev 7mo agoThe compute speed is definitely correlated with the memory consumption in LLM land. More efficient attention means both less memory and faster inference. Which makes sense to me because my understanding is that memory bandwidth is so often the primary bottleneck. We're also seeing a recent rise in architectures boosting compute speed via multi-token prediction (MTP). That way a single inference batch can produce multiple tokens and multiply the token generation speed. Combine that with more lean ratios of active to inactive params in MOE and things end up being quite fast. The rapid pace of architectural improvements in recent months seems to imply that there are lots of ways LLMs will continue to scale beyond just collecting and training on new data.
- mt_ 7mo agoI call them, entropy reducers.
- whimsicalism 7mo ago> how these models are going to keep up with the expanding boundary of science The same way humans do? The phraseology in this comment: 'probability distributions', 'baked these patterns' IMO has all the trappings of the stochastic parrot-style HN-discourse that has been consistently wrong for almost a decade now. The reference to how AI will keep up with AI-assisted human progress in science in 2030 is meant to reassure. It contains a number of premises that we have no business being confident in. We are potentially witnessing the obviation of human cognitive labor.
- mccoyb 7mo agoSorry, are you familiar with what a next token distribution is, mathematically speaking? If you are not, let me introduce you to the term: a probability distribution. Just because it has profound properties ... doesn't make it different. > has all the trappings of the stochastic parrot-style HN-discourse that has been consistently wrong for almost a decade now Perhaps respond to my actual comment compared to whatever meta-level grouping you wish to interpret it as part of? > It contains a number of premises that we have no business being confident in. We are potentially witnessing the obviation of human cognitive labor. What premises? Be clear.
- fauigerzigerk 7mo agoI think they are questioning whether human feedback is even necessary to make progress, i.e. whether the premise that RL needs to be RLHF is true. My (limited) understanding is that LLMs are not capable of escaping their learned distribution by simply feeding on their own output. But the question is whether the required external (out of distribution) "stimulus" needs to come from humans. Could LLMs design experiments/interventions to get feedback from their environment like human scientists would? I have my doubts that this is possible without an inherent causal reasoning capability but I'm not sure.
- Robdel12 7mo agoThat’s AGI, right? For the model to learn novel things itself and retain it? I have no idea but I’m along for the ride!
- atleastoptimal 7mo agoThe obvious answer is that continual learning is going to be solved
- 9wzYQbTYsAIc 7mo agoCheck out https://unratified.org https://unratified.org, it tries to answer that question directly, actually.
- wvlia5 7mo agoThis seems to be a bot comment. HN will lose its value if these bots are not purged.
- stalfie 7mo agoThis is an urgent problem, but it can probably not be solved without some kind of "verified human 2FA" like the Norwegian BankID + facial recognition. Knowing the HN audience, this will never happen. And so the site is doomed.
- deleted 7mo ago[deleted]
- dzdt 7mo agoI think it could be solved still pseudononymously: introduce a "vouch" button that allows a user to vouch that another user is human. This is consequential both for the vouched-for and vouching accounts. Run a page-rank style algorithm on the graph of vouches to generate a certainty score for the humanity of each account. For repeated posters this should converge to a correct answer fairly quickly. There is still a challenge for green accounts, but having degraded experience for new users is not a doom scenario for the site.
- mimischi 7mo agoWhat makes you think that? Genuine question, as I’ve not flagged it as such in my mind.
- WarcrimeActual 7mo agoIronically, his last comment before this was to the effect of "Github has a bot problem."
- gerold 7mo agoCan you explain to me what makes this an obvious bot comment? I'm not doubting it, I just don't understand.
- mccoyb 7mo ago
- klooney 7mo ago> Experts will naturally use these systems more productively, because they know how to coerce models into the correct conditional distributions which light up the right techniques. How much can you patch over with the models doing their own metacognition?
- ainiriand 7mo agoAre not LLMs supposed to just find the most probable word that follows next like many people here have touted? How this can be explained under that pretense? Is this way of problem solving 'thinking'?
- crocowhile 7mo agoThose people still exist? I only know one guy who is still fighting those windmills
- IgorPartola 7mo agoIn some cases solving a problem is about restating the problem in a way that opens up a new path forward. “Why do planets move around the sun?” vs “What kind of force exists in the world that makes planets tethered to the sun with no visible leash?” (Obviously very simplified but I hope you can see what I am saying.) Given that a human is there to ask the right questions it isn’t just an LLM. Further, some solutions are like running a maze. If you know all the wrong turns/next words to say and can just brute force the right ones you might find a solution like a mouse running through the maze not seeing the whole picture. Whether this is thinking is more philosophical. To me this demonstrates more that we are closer to bio computers than an LLM is to having some sort of divine soul.
- ainiriand 7mo agoThanks for your input. The way I saw this and how it looks Knuth interpreted it is that there were some reasoning steps taken by Claude independently. Some internal decisions in the model that made it try different things, finally succeeding.
- 7mo ago
- miroljub 7mo agoSolves? It's a part of the training set. Nothing more, nothing less.
- mwigdahl 7mo agoDid you read the article? It was an open problem.
- bluGill 7mo agoWas it? It was an open problem to Knuth - who generally knows how to search literature. However there is enough literature to search that it wouldn't be a surprise at all to discover it was already solved but he just used slightly different terms and so didn't find it. Or maybe it was sovled because this is a specialization of something that looks unrelated and so he wouldn't have realized it when he read it. Or... Overall I'm going with unsolved, because Knuth is a smart person who I'd expect to not miss the above. I'm also sure he falls for the above all the time even though the majority of the time he doesn't.
- mwigdahl 7mo agoAgreed with all of that, but with the added point that Knuth has done a lot of work in this exact area in The Art of Computer Programming Volume 4. If he considers this conjecture open given his particular knowledge of the field, it likely is (although agreed, it's not guaranteed).
- ordu 7mo ago> If he considers this conjecture open given his particular knowledge of the field, it likely is (although agreed, it's not guaranteed). It is as good as guaranteed. If Knuth says it doesn't know how to solve the problem, and if anyone knows, then they will inform Knuth about it. Knuth not just a very knowledgeable person, but a celebrity also.
- skinner_ 7mo agoAlso, if Claude had regurgitated a known solution, it would have come up with it in the first exploration round, not the 31st, as it actually did.
- ecshafer 7mo agoI wonder how long we have until we start solving some truly hard problems with AI. How long until we throw AI at "connect general relativity and quantum physics", give the AI 6 months and a few data centers, and have it pop out a solution?
- worldsavior 7mo agoIf AGI will ever come, then. Currently, AI is only a statistical machines, and solutions like this are purely based on distribution and no logic/actual intelligence.
- rustyhancock 7mo agoI don't even think that's the issue. The issue to my mind is a lack of data at the meeting of QFT/GR. Afterall few humans historically have been capable of the initial true leap between ontologies. But humans are pretty smart so we can't say that is a requirement for AGI.
- worldsavior 7mo agoWhen it comes to revolutionary/unsolved subjects, there will never be enough data. That's why its revolutionary/unsolved.
- cjcole 7mo agoMaybe. “The laws of nature should be expressed in beautiful equations.” - Paul Dirac “It is, indeed, an incredible fact that what the human mind, at its deepest and most profound, perceives as beautiful finds its realisation in external nature. What is intelligible is also beautiful. We may well ask: how does it happen that beauty in the exact sciences becomes recognizable even before it is understood in detail and before it can be rationally demonstrated? In what does this power of illumination consist?” - Subrahmanyan Chandrasekhar “I often follow Plato’s strategy, proposing objects of mathematical beauty as models for Nature.” “It was beauty and symmetry that guided Maxwell and his followers.” - Frank Wilczek “Beauty, is bound up with symmetry.” - Herman Weyl "Still twice in the history of exact natural science has this shining-up of the great interconnection become the decisive signal for significant progress. I am thinking here of two events in the physics of our century: the rise of the theory of relativity and that of the quantum theory. In both cases, after yearlong unsuccessful striving for understanding, a bewildering abundance of details was almost suddenly ordered. This took place when an interconnection emerged which, thought largely unvisualizable, was finally simple in its substance. It convinced through its compactness and abstract beauty – it convinced all those who can understand and speak such an abstract language." - Werner Heisenberg Maybe (just maybe) these things (whatever you want to call them) will (somehow) gain access to some "compact", beautiful, "largely unvisualizable" "interconnection" which will be the self-evident solution. And if they do, many will be sure to label it a statistical accident from a stochastic parrot. And they'll right, for some definitions of "statistical", "accident", "stochastic", and "parrot".
- Pat44113 7mo agoI asked Claude to solve the pentominoes puzzle made famous by Arthur C. Clarke. It struggled mightily until I told it how I'd solved the problem using 64 bit unsigned integers to represent the board and pieces. Then, it created a C# program that solved the problem very quickly. However, in the 20x3 case it found four solutions when there are only two. Turns out it had incorrectly mapped one of the pentominoes. Sort of a silly mistake; the sort a human might make.
- phoronixrly 7mo ago[flagged]
- logicprog 7mo agoRegurgitation is pretty rare, and very difficult to coax out, if not even impossible, for things that aren't massively overrepresented in the training set relative to the size of the training set. Even the famous regurgitation paper showed this: while they got most of the models to regurgitate the first book of the Harry Potter series, only Claude 3.7 Sonnet was able to regurgitate any significant portion of any of the other books that had a high nv-recall rate, and basically all of them dropped off precipitously for works like GoT, The Catcher in the Rye, Beloved, and remembered almost nothing about the Da Vinci Code or Catch-22[0]. So you really need huge amounts of examples to get any kind of meaningful regurgitation on any kind of reliable basis. Thus, you'd have to prove that hypothesis. [0]: https://arxiv.org/pdf/2601.02671 https://arxiv.org/pdf/2601.02671
- ontouchstart 7mo agoFascinating report by DEK himself. Time to sit down, read, digest and understand it without the help of LLM.
- ontouchstart 7mo agoI don't have time to do that myself yet so I just dug a quick TL;DR rabbit hole for fun: https://ontouchstart.github.io/rabbit-holes/llm_rabbit_hole_dek/index.html https://ontouchstart.github.io/rabbit-holes/llm_rabbit_hole_...
- iandanforth 7mo agoTLDR (story, not math) - Knuth poses a problem, his friend uses Claude to conduct 30 some explorations, with careful human guidance, and Claude eventually writes a Python program that can find a solution for all odd values. Knuth then writes a proof of the approach and is very pleased by Claude's contribution. Even values remain an open question (Claude couldn't make much progress on them)
- logicprog 7mo ago> with careful human guidance, I think this is pretty clearly an overstatement of what was done. As Knuth says, "Filip told me that the explorations reported above, though ultimately successful, weren’t really smooth. He had to do some restarts when Claude stopped on random errors; then some of the previous search results were lost. After every two or three test programs were run, he had to remind Claude again and again that it was supposed to document its progress carefully. " That doesn't look like careful human guidance, especially not the kind that would actually guide the AI toward the solution at all, let alone implicitly give it the solution — that looks like a manager occasionally checking in to prod it to keep working.
- semessier 7mo agolooks like he is trying to make a point that the actual (formal) proof for 2Z + 1 (odd numbers) is still human - by himself that is. Not sure who came up with the core modular arithmetic idea of with s = 0 k increasing by 2 mod m.
- fazkan 7mo agotime to use claude code to understand DEKs paper, in plain English. As someone who did a bit of formal verification in grad school. I feel like, there are a long tail of problems that can be solved by human-model collab like this one. The problems may not mean much but hopefully it can stack up understanding of intelligence.
- beej71 7mo agoFrom my naive standpoint, LLMs like this seem to have some big strengths. One: possession of a superhuman expanse of knowledge. Two: making connections. Three: tireless trial and error. If you put those three things together, you end up with some cool stuff from time to time. Perhaps the proof of P!=NP is tied to an obscure connection that humans don't easily see due to individual lack of knowledge or predisposition of bias.
- xvector 7mo agoThis is why the whole "LLMs for mass surveillance" thing is scary imo.
- beej71 7mo agoYeah, this is a dictator's dream scenario and hell for the citizens. Not only do you not want to get caught for saying something that The Great Leader disapproves of, but you're terrified that anything you say might get flagged by an AI.
- cbovis 7mo agoUnless my understanding is incorrect about how these tools work that last point isn't really a quality of LLMs as such? It gets attributed because the lines are blurred but the tireless trial and error is actually just a quality of a regular programatic loop (agent/orchestrator) that happens to be doing the trickiest part of its work via an LLM.
- naughtyrabisu 7mo agoThree: tireless trial and error. Cannot agree more. I figured this probably be the biggest advantage of LLM considering for other variables humans hold the same-level competency.
- IAmGraydon 7mo ago>One: possession of a superhuman expanse of knowledge. Two: making connections. Three: tireless trial and error. One and three I believe are correct. The second point, making connections, is something LLMs seem to be incapable of truly doing unless the connection is already known and in its training data.
- jdnier 7mo ago> I think Claude Shannon’s spirit is probably proud to know that his name is now being associated with such advances. Hats off to Claude! I didn't realize Claude was named after Claude Shannon! https://en.wikipedia.org/wiki/Claude_Shannon https://en.wikipedia.org/wiki/Claude_Shannon
- bread-wood 7mo agoHere I was assuming it was named after https://en.wikipedia.org/wiki/Claude_(alligator) https://en.wikipedia.org/wiki/Claude_(alligator)
- deleted 7mo ago[deleted]
- NitpickLawyer 7mo agoWait till you hear about nvidia and their GPU architecture naming scheme :)
- tzumaoli 7mo agoTrivia: Claude Shannon proposed the idea of predicting the next token (letter) using statistics/probabilities in the training data corpus in 1950: "Prediction and Entropy of Printed English" https://languagelog.ldc.upenn.edu/myl/Shannon1950.pdf https://languagelog.ldc.upenn.edu/myl/Shannon1950.pdf
- Anon84 7mo agoIt goes back a bit further than that. His 1948 “Mathematical theory of communication” [1] already has (what we would now call) a Markov chain language model, page 7 onwards. AFAIK, this was based on his classified WWII work so it was probably a few years older than that [1] https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf https://people.math.harvard.edu/~ctm/home/text/others/shanno...
- aix1 7mo ago
- faxmeyourcode 7mo ago> Filip also told me that he asked Claude to continue on the even case after the odd case had been resolved. “But there after a while it seemed to get stuck. In the end, it was not even able to write and run explore programs correctly anymore, very weird. So I stopped the search.” Interesting snippet towards the end. I wonder if they were using claude.ai or claude code. Sounds like they ran out of context and entered the "dumb zone."
- afspear 7mo agoWhat would be super cool is if this dumb zone could be quantified and surfaced to the user. I've noticed that copilot now has a little circle graph that indicates context use percentage and it changes color based on percentage. I'll bet these are very naive metrics on used tokens vs context availability. I wonder if there could be meta data streamed or sent along with the tokens that could show that you've entered the dumb zone.
- simianwords 7mo agoThey mentioned plan document
- joshrw 7mo agoThen it needs to do context compacting, otherwise the results become garbage
- brcmthrowaway 7mo agoWhat is dumb zone?
- kami23 7mo agoWhen the LLMs start compacting they summarize the conversation up to that point using various techniques. Overall a lot of maybe finer points of the work goes missing and can only be retrieved by the LLM being told to search for it explicitly in old logs. Once you compact, you've thrown away a lot of relevant tokens from your problem solving and they do become significantly dumber as a result. If I see a compaction coming soon I ask it to write a letter to its future self, and then start a new session by having it read the letter. There are some days where I let the same session compact 4-5 times and just use the letter to future self method to keep it going with enough context because resetting context also resets my brain :) If you're ever curious in Claude once you compact you can read the new initial prompt after compaction and see how severe it gets cut down. It's very informative of what it forgets and deems not important. For example I have some internal CLIs that are horribly documented so Claude has to try a few flags a few times to figure out specifics and those corrections always get thrown away and it has to relearn them next time it wants to use the CLI. If you notice things like that happening constantly, my move is to codify those things into my CLAUDE.md or lately I've been making a small script or MCP server to run very specific flags of stuff.
- nphardon 7mo agoMust be a fun time to work on open problems. I published my graduate research close to a decade ago, often find myself fantasizing about tackling open problems with Claude.
- dfilppi 7mo ago[dead]
- taylorius 7mo agoI thought Claude Monet - Impressionist techniques applied to coding.
- konne88 7mo agoI didn't expect such a misleading intro from Knuth. It reads like Claude solved Knuth's math problem. In reality, Claude generated various example solution, and Knuth then manually generalized that to a formal proof. What Claude did is certainly useful, but it would have been nice to be clear about the scope of the contribution in the intro.
- bachmeier 7mo agoMy interpretation is that Claude did what Knuth considers to be the "solution". Doing the remaining work and polishing up the proof are not necessary to have a solution from this perspective.
- OneManyNone 7mo agoClaude did not find a proof, though. It found an algorithm which Knuth then proved was correct.
- versteegen 7mo agoAFAICT, Claude was not asked to prove its algorithm works for all odd n, but was instead told to move on to even n.
- CobrastanJorji 7mo agoYes, and his point is that finding that algorithm was, to Knuth, the interesting part. Getting from that to a proof was the boring bit.
- NewsaHackO 7mo agoYeah, and I'm not sure what the other guy's argument is. It's Knuth, the primary researcher, who is giving the praise here. I don't see a possible motivation he would have to falsely give accolades to a AI for a problem he presented, then cleaned up to solve.
- OneManyNone 7mo ago
- zackmorris 7mo agoAmazing paper. The simulated annealing portion reminds me of genetic algorithms (GAs). A good intro to that are the Genetic Programming series of books by John Koza, I read III in the early 2000s: https://www.amazon.com/Genetic-Programming-III-Darwinian-Invention/dp/1558605436 https://www.amazon.com/Genetic-Programming-III-Darwinian-Inv... https://www.genetic-programming.com/ https://www.genetic-programming.com/ Note that the Python solution in the pdf is extremely short, so could have been found by simply trying permutations of math operators and functions on the right side of the equation. We should be solving problems in Lisp instead of Python, but no matter. That's because Lisp's abstract syntax tree (AST) is the same as its code due to homoiconicity. I'm curious if most AIs transpile other languages to Lisp so that they can apply transformations internally, or if they waste computation building programs that might not compile. Maybe someone at an AI company knows. - I've been following AI trends since the late 1980s and from my perspective, nothing really changed for about 40 years (most of my life that I had to wait through as the world messed around making other people rich). We had agents, expert system, fuzzy logic, neural nets, etc since forever, but then we got video cards in the late 1990s which made it straightforward to scale neural nets (NNs) and GAs. Unfortunately due to poor choice of architecture (SIMD instead of MIMD), progress stagnated because we don't have true multicore computing (thousands or millions of cores with local memories), but I digress. Anyway, people have compared AI to compression. I think of it more as turning problem solving into a O(1) operation. Over time, what we think of as complex problems become simpler. And the rate that we're solving them is increasing exponentially. Problems that once seemed intractable only were because we didn't know the appropriate abstractions yet. For example, illnesses that we thought would never be cured now have vaccines through mRNA vaccines and CRISPR. That's how I think of programming. Now that we have LLMs, whole classes of programming problems now have O(1) solutions. Even if that's just telling the computer what problem to solve. So even theorem proving will become a solved problem by the time we reach the Singularity between 2030 and 2040. We once mocked GAs for exploring dead ends and taking 1000 times the processing power to do simple things. But we ignored that doing hard things is often worth it, and is still a O(1) operation due to linear scaling. It's a weird feeling to go from no forward progress in a field to it being effectively a solved problem in just 2 years. To go from trying to win the internet lottery to not being sure if people will still be buying software in a year or two if/when I finish a project. To witness all of that while struggling to make rent, in effect making everything I have ever done a waste of time since I knew better ways of doing it but was forced to drop down to whatever mediocre language or framework paid. As the problems I was trained to solve and was once paid to solve rapidly diminish in value because AI can solve them in 5 minutes. To the point that even inventing AGI would be unsurprising to most, so I don't know why I ever went into computer engineering to do exactly that. Because for most people, it's already here. As I've said many times lately, I thought I had more time. Although now that we're all out of time, I have an uncanny feeling of being alive again. I think tech stole something from my psyche so profound that I didn't notice its loss. It's along the lines of things like boredom, daydreaming, wasting time. What modern culture considers frivolous. But as we lose every last vestige of the practical, as money becomes harder and harder to acquire through labor, maybe we'll pass a tipping point where the arts and humanities become sought-after again. How ironic would it be if the artificial made room for the real to return? On that note, I read a book finally. Hail Mary by Andy Weir. The last book I read was Ready Player One by Ernest Cline, over a decade ago. I don't know how I would have had the bandwidth to do that if Claude hadn't made me a middle manager of AIs.
- zoogeny 7mo agoI recall an earlier exchange, posted to HN, between Wolfram and Knuth on the GPT-4 model [1]. Knuth was dismissive in that exchange, concluding "I myself shall certainly continue to leave such research to others, and to devote my time to developing concepts that are authentic and trustworthy. And I hope you do the same." I've noticed with the latest models, especially Opus 4.6, some of the resistance to these LLMs is relenting. Kudos for people being willing to change their opinion and update when new evidence comes to light. 1. https://cs.stanford.edu/~knuth/chatGPT20.txt https://cs.stanford.edu/~knuth/chatGPT20.txt
- 3abiton 7mo ago> Kudos for people being willing to change their opinion and update when new evidence comes to light. > 1. https://cs.stanford.edu/~knuth/chatGPT20.txt https://cs.stanford.edu/~knuth/chatGPT20.txt I think that's what make the bayesian faction of statistics so appealing. Updating their prior belief based on new evidence is at the core of the scinetific method. Take that frequentists.
- Chinjut 7mo agoIt does not seem fair to say that frequentists do not update their beliefs based on new evidence. This does not seem to accurately capture what the difference between Bayesians and frequentists (or anyone else) is.
- atomicnature 7mo agoWhat's the difference as you see it?
- kqr 7mo agoEveryone updates their belief in hypotheses based on the perceived strength of evidence they observe. That's just science. Frequentists and Bayesians differ in which sets of statistical tools they prefer for measuring the strength of evidence.
- Steinmark 7mo ago[flagged]
- laalshaitaan 7mo ago[flagged]
- chrsw 7mo agoAm I mad or is there a missing ")" on lines and 8 and 9 of the first "C form" that should go before the semicolons?
- kqr 7mo agoCorrect. Line 10 does not have the same mistake.
- akssassin907 7mo ago[flagged]
- computerex 7mo agoIt's incredible to see work like this from him, at a ripe old age of eighty-six.
- kqr 7mo agoI agree. I met Knuth briefly after a guest lecture at my university a few years ago and although you could tell his body was getting old, his mind was incredibly fresh. Although I'm not as bright as him, I can only hope to be as intellectually curious as him at that age.
- OJFord 7mo agoI don't even think this is controversial, but I don't think it's at all without causation: not remaining curious, keeping the mind stimulated, etc., accelerates one's decline. If you work in something labour intensive, you should retire young while your body's in good health; if you work in academia you should (strive for emeritus and) never leave! (And if you work in SWE, I don't know, we should probably retire, but then spend more time on our own projects/experiments/reading HN.) (All assuming for sake of argument we're optimising for longevity without considering time with family, having the funds to retire, etc.)
- justanotherjoe 7mo agoTo put this more succintly I think, the mind loves learning something new. Something to do with new connections in the brain.
- adolfont 7mo agoWell, for starters, I think it's wrong to criticise LLMs with ‘it can't do that’ (from what I understood from the first paragraph, this was Donald's criticism). If it can, does it make a difference in relation to all the other problematic aspects of LLMs? Not for me. Two links that might enlighten Donald: - Against the Uncritical Adoption of 'AI' Technologies in Academia https://zenodo.org/records/17065099 https://zenodo.org/records/17065099 - The AI Con https://thecon.ai https://thecon.ai
- quinndupont 7mo agoInteresting to see the mathematical solution space get optimized away. On account of “there’s no accounting for taste” this actually makes me hopeful that creative workers have durable skills that can’t be optimized, which I can’t say about mathematics and computer science.
- lhl 7mo agoI was a bit interested to do a replication and see if better harness could avoid some of the problems they ran w/ context management, poor instruction following, etc and it looks like yes, it's definitely possible. Here's my repo: https://github.com/lhl/claudecycles-revisited https://github.com/lhl/claudecycles-revisited I used Codex w/ 5.2 xhigh and a relatively simple AGENTS.md - I have some session-analysis as well. The original replication was 47 minutes, then another 30 minutes of gap filling, and finally about 30 minutes of writing an extension to take the work a bit further, with Claude Code Opus 4.6 doing some documentation cleanup and verification.
- carterschonwald 7mo agoomg this is so cool. because im writing my own harness and i need some cognitive benchmarks. i have a bunch of harness level infra around llm interactions that seems to help with reasoning, but i dont have a structured way evaluate things thx for sharing your test setup, i really appreciate the time you took. this will help me so much
- pushedx 7mo agoAs described in the readme of your repo (did you read it?) your agent found the Knuth paper located one directory level above its working directory. So, you didn't produce a replication in 47 minutes, it just took around 30 minutes for your agent to find that you had the answer in a PDF in a nearby directory.
- antonly 7mo agoI wonder how common of a problem this will be in the future. The experiment will fail due to improper setup, the human will at best glance over the logs and declare victory, and everyone just believes.
- lhl 7mo agoYes, I read it and specifically pointed it out (that's why there are 3 hours of interactive logs). There are 4 other runs pushed now so you can see what actual clean room runs for 5.2 xhigh, 5.3-Codex xhigh, 5.4 xhigh, and Opus 4.6 ultrathink look like: https://github.com/lhl/claudecycles-revisited/blob/main/COMPARISON.md https://github.com/lhl/claudecycles-revisited/blob/main/COMP... as well as the baseline.
- mihevc 7mo agoEt tu, Knuthus?
- Smaug123 7mo ago(You want the vocative case here, if you're going to shove on a suffix to make it look Latin. The Shakespeare quote is "et tu, Brutè?".)
- flashybaby 7mo ago[flagged]
- lacoolj 7mo agoOK so now I need someone to take this problem and feed it into Gemini Deep Think or whatever and see if you get the same (or better/worse) outcome. No one cares about ChatGPT so don't bother with that. OK GO
- ano-ther 7mo agoInteresting that for a paper by Don Knuth himself the PDF was created with dvips (TeX Live) but then switched to Acrobat Distiller, resulting in a rather low resolution (at least on my screen). From the document properties: > Creator: dvips(k) 2023.1 (TeX Live 2023) > PDF Producer: Acrobat Distiller 25.0 (Macintosh)
- svat 7mo agoThe issue is not of low resolution exactly, but font format. Knuth uses bitmap fonts, rather than vector fonts like everyone else. This is because his entire motivation for creating TeX and METAFONT was to not be reliant on the font technology of others, but to have full control over every dot on the page. METAFONT generates raster (bitmap) fonts. The [.tex] --TeX--> [.dvi] --dvips--> [.ps] --Distiller--> [.pdf] pipeline uses these fonts on the page. They look bad on screen because they're not accompanied by hinting for screens' low resolution (this could in principle be fixed!), but if you print them on paper (at typical resolution like 300/600 dpi, or higher of typesetters) they'll look fine. Everyone else uses TrueType/OpenType (or Type 3: in any case, vector) fonts that only describe the shape and leave the rasterization up to the renderer (but with hinting for low resolutions like screens), which looks better on screen (and perfectly fine on paper too, but technically one doesn't have control over all the details of rasterization).
- modnick 7mo ago[dead]
- lhl 7mo agoI am not a theoretical CS or math expert by any means, but I have been wrangling coding agents for a while and reading the paper and the problems Stapper had with dealing w/ Claude (context management, instruction following, etc) decided to see if I could replicate with a slightly better harness. The results were pretty interesting: https://github.com/lhl/claudecycles-revisited https://github.com/lhl/claudecycles-revisited - My original setup left traces of the PDF paper and after GPT 5.3-Codex xhigh reached an impasse it went looking for it and found it! - I went and did cleanroom (basically one-shot) passes for GPT 5.2 xhigh, GPT 5.3-Codex xhigh, and Claude Opus 4.6 ultrathink and 5.2/5.3 found alternate solutions for odd m >= 5 , Opus 4.6 did not find any proofs but tried more approaches to solving. Full comparison/analysis here: https://github.com/lhl/claudecycles-revisited/blob/main/COMPARISON.md https://github.com/lhl/claudecycles-revisited/blob/main/COMP... I've also included the session traces and analysis in the repo branches. Also, the AGENTS.md was pretty simple, but that harness produced consistent process outcomes across all three models: - All built verifiers first - All maintained worklogs with exact commands - All archived machine-readable artifacts - All documented failed approaches - All maintained restart-safe context capsules
- dellasera 7mo agoShock! Shock! Ugh
- mikeaskew4 7mo agoClaude repeatedly insisted I give up on parsing a relatively vague object recently. When I got more specific, and pressed it to continue, not only did it work, but Claude seemed amazed. Ugh.