5 ms·
> Historically, many of these problems were bottlenecked by human attention. Someone had to care enough to spend hours or days reading obscure material, testing
by vb-8448 15d ago
> Historically, many of these problems were bottlenecked by human attention. Someone had to care enough to spend hours or days reading obscure material, testing unpromising ideas, tracing references, and trying things that might go nowhere
I wonder how many of the recent results are due to the fact that very few looked at the problem to start with. Still great results, but the general impression is that it's more about the so many low-hanging fruits than the actual capability.
- glimshe 15d agoLet's not normalize the achievement. Just a couple years ago this would be considered science fiction. We can argue that 2026 AI can't solve the very toughest cryptograms, but the fact it can solve nontrivial ones is already magical. Now on to the Voynich Manuscript :)
- chr15m 15d agoIt is indeed absolutely incredible that it can solve these puzzles given plaintext instructions with very little context.
- dakolli 15d agoWhat's so magical about the problem... Its the exact time of problem they were built to solve (things that can be brute forced with language). I'm not impressed.
- scrollaway 15d ago[flagged]
- deleted 15d ago[deleted]
- computerex 15d agoWhy don’t you solve such problems?Being “not impressed” sounds more like a knowledge gap on your part than an informed opinion.
- card_zero 14d agoPepper grinders have always impressed me, very effective devices, great at grinding out results, I mean pepper. I can't do it by hand at all!
- nonethewiser 14d agoSure you can, its just difficult to the point of being infeasible
- EdwardDiego 14d agoWhy are you so toxic?
- computerex 13d agoHow ironic for you to say that to me. Maybe you can use ai to help explain it to you.
- rossy 15d agoNo, he's right. Actually, let's have a bit of sobriety when discussing the achievements of the most heavily marketed technology of all time, as published by an organisation that stands to benefit financially from the public perception of that technology. The discussion of "what made this problem low hanging fruit" is much more interesting, imo, than just breathlessly joining the hype train.
- honeycrispy 15d agoThank you. More money than the GDP 90% of the sovereign countries around the world is hanging in the balance, and people are taking everything OpenAI and Anthropic are saying at face value as if this isn't the financial / marketing equivalent of war, assuming they they wouldn't use every legal and shady tactic, bending every truth available to them to sway the balance of public opinion in their favor. It makes me feel like I'm living in the twilight zone. People need to wake up.
- qlte 15d agoA few months ago a post claiming an amateur deciphered Linear A hit the front page and quickly got almost 500 upvotes: https://news.ycombinator.com/item?id=48600107 https://news.ycombinator.com/item?id=48600107 https://aiclambake.com/clamtakes/linear-a/ https://aiclambake.com/clamtakes/linear-a/ Despite the announcement originating from a blog named "AI Clambake" covering "weekly, human-powered newsletter for advertising folks". Written by a personal friend of the author. Announced without any corroboration or commentary whatsoever from academics or subject matter experts of any kind. And, of course, not submitted to any peer reviewed journal or even Arxiv. The author of the purported discovery was described as a "self taught AI engineer and amateur linguist". In the comments the friend insisted several times that a draft of the paper (not posted), was emailed to a top professor at Rutgers, giving it additional credibility that his friend wasn't another one of ten thousand cranks who has made the same claim over the years (seemingly unaware that cold emailing random professors found from a Google search is the first thing basically every crank does). You would think this should have set of dozens of alarm bells for everyone, making the value of this announcement basically zero. And yet it hit the front page with the bulk of comments ecstatic that some random guy with Claude Code could do something experts in academia who spent their lives devoted to the problem couldn't. I had become accustomed to the toxic optimism of this hype cycle in which even mild criticism leads to accusations of being a discredited "AI skeptic"/Gary Marcus/Ed Zitron type who was "coping" (?). But this was like something you'd see shared on FB linking to a .xyz domain by an elderly family member who recently drained their accounts buying Xbox gift cards to pay their IRS bill. It feels a lot like the week or two when HN was overflowing with exuberance from the LK-99 room temperature superconductor "discovery ". You'd see post after post fantasizing about an imminent future with a world full of maglev hovercrafts, MRIs built into every phone, fusion reactors and more. But people pointing out none of that was scientifically plausible and evidence of LK-99 superconductoring was non-existent were accused of knee-jerk negativity and the typical HN cynicism and pessimism.
- parpfish 15d agoI’m pretty sure the Beale ciphers are a hoax, but I’d love to be proven wrong.
- geraneum 15d agoSomeone wrote a prompt, that included instructions for finding the problem itself and got handed a solution by a machine trained on all available text. I don’t see any achievement for the prompter. As for the machine, we can’t keep being perpetually shocked 24x7. It’s tiring (unless if we’re being paid for it)
- chmod775 15d ago"AI solves niche thing you've never heard of" is a daily headline at this point. What's genuinely cool isn't that AI managed to solve some specific problem only a handful of people even cared about, it's that humanity can now cheaply clean up its backlog of such things.* That does not mean that specific instances of it are still very interesting though. This article is the "I had claude vibecode a thermostat for my bathtub" of cryptography. * And in this case I'm not sure it even meets that bar. For all we know a couple readers back when the book released had a delightful afternoon with it, solved the riddle, then forgot about it.
- IshKebab 15d agoI disagree. This may be some niche thing that I've never heard of but a) it's still non-trivial; it still would have been science fiction to solve it a few years ago, and b) have you already forgotten the Navier Stokes drama? That is not some niche thing I've never heard of. It kind of blows my mind how quickly people have forgotten both the state of AI in ~2010, and the outlook. If you had asked 100 people in 2010 whether they would see AI that could actually pass the Turing test in their lifetimes, you would have got 100 "no"s. AI had been an unsolved problem for literally decades and it was firmly in the nuclear fusion/flying cars category.
- u8080 14d ago"Yeah it is just token predictor, it could not even beat top 0.01% domain experts so it does not count as intelligence"
- emanuele232 14d agothat's the point, it is clear the the bar to determine if llms are useful / intelligent is being moved every time these systems improve, but it is starting to fell like people are in denial. we are seeing significant progress, at a rate we are absolutely not used to experience.
- card_zero 14d ago
- _fizz_buzz_ 14d agoHumans are incredibly good at adapting. A few days ago AI solved Navier-Stokes and I was blown away. Now I'm already thinking: "Well, it was only a counterexample and it brute-forced its way to it." lol
- EdwardDiego 14d ago> AI solved Navier-Stokes and I was blown away. That's not what happened, go read about it harder, please.
- _fizz_buzz_ 14d agoSure. More precisely: they resolved the Navier-Stokes Millennium problem as posed by the Clay Institute. Not sure what else "solving Navier-Stokes" could reasonably mean. A general closed-form solution probably doesn't exist. And numerical solutions have existed for decades. But of course there are still open questions like unforced solutions etc.
- gjm11 14d agoI suspect EdwardDiego is referring to the brouhaha about whether OpenAI's training for the model that produced the alleged solution to the Millennium Problem about the Navier-Stokes equations was trained on material that included conversations Tristan Buckmaster and Levent Alpöge had had with earlier OpenAI systems. I think there's a bit less to that than meets the eye. Yes, OpenAI's result builds on human work. It's possible that it builds on more human work than OpenAI admitted. But even if we suppose that everything Buckmaster and Alpöge did (which, btw, was itself very heavily LLM-assisted/generated work) was a necessary precursor to what OpenAI released, it's still the case that OpenAI's clankers completed the solution and Buckmaster and Alpöge didn't. My understanding from what Buckmaster has written about this is that the deep mathematical ideas behind their work (and presumably OpenAI's) are due to Córdoba and Martínez-Zoroa. Those ideas are in the published literature, and human mathematicians and AI systems alike are allowed to use them, and doing so doesn't mean they didn't actually do something impressive. Mathematicians build on one another's work; that's how mathematics progresses and always has been. It may very well be that OpenAI's announcement has a serious problem of professional ethics, especially as their first version of it didn't even list Córdoba and Martínez-Zoroa in its references. (On the specific question of what if anything they learned from B&A's work before that was published: OpenAI are now claiming that after investigating carefully they are confident that the model was not trained on anything Buckmaster and Alpöge did after early July. B&A had been working on this thing for much longer than that. However, on Buckmaster's account of things it wasn't until mid-August that they got beyond what he calls "preliminary results".) But! The results of B&A were themselves largely AI-generated. (From Buckmaster's statement: "on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable." That is: the LLMs found the proof, and B&A had to work to understand what the LLMs had done. It's not that humans did the thinking and AIs just did the gruntwork. (Except in so far as one might want to give all the credit for Real Deep Cleverness to C&MZ.) And! What OpenAI say their model has proved goes well beyond what B&A did. I don't see any way of slicing this that makes it unreasonable to say (unless it turns out that there's an error in the proof -- unlikely, given that it comes with Lean verification, but there have been misformalizations and Lean bugs in the past and there surely will be in the future) that AIs solved the N-S problem. No, they couldn't have done it without the work of C&MZ, but again: important mathematical work almost always builds on earlier important mathematical work, that's just how it is. Yes, if OpenAI are lying through their teeth their model might have had early access to B&A's ideas -- but it seems like most of the B&A work was actually done by AI systems anyway. It is (I think -- I am not an expert and in particular I have not so much as looked at OpenAI's publication) reasonable to say that the deepest ideas here came from humans, and that it was already widely expected that the N-S problem would be solved in the not-impossibly-distant future in something like the way it has been. So, sure, what the AIs have done here is much less impressive than if they'd settled the Riemann Hypothesis or (probably even harder) PvNP. But it's still a resolution of a famous mathematical problem that any human mathematician would have been very proud to have achieved.
- uludag 14d agoI personally think that beyond normalizing, we should be actively be trying to dismiss this with all the cynicism we have. What does Anthropic have to gain from writing this? Behind the the scenes what might Anthropic be failing to disclose? How many failed experiments do we not know about?
- badsectoracula 14d agoThe article doesn't seem to be written by Anthropic.
- saberience 14d agoThis was a basic cypher that effectively no one cared about. It wasn't famous or particularly notable. I would say, the only reason it was never solved was because not enough people actually cared about it to begin with. This isn't a big accomplishment.
- ndiddy 14d agoThe Voynich theory I find most compelling is that it was a hoax made for a quack doctor, made to look like a foreign herbal manuscript. "Oh of course the local doctors can't help you, but my special book from a faraway land that only I can read may have the cure." Some recent analysis of the manuscript has found that the pages are more linguistically similar when read as individual flat sheets than how they're read when bound (i.e. whoever was writing the text was most likely going sheet by sheet and using the last completed page as reference for the text). The manuscript has only been bound once, in the fifteenth century (around when the vellum pages have been carbon-dated to), so whoever bound the manuscript was not able to "read" it. See https://journals.openedition.org/digitalmedievalist/2331 https://journals.openedition.org/digitalmedievalist/2331
- aeternum 15d agoYes, even many of the proofs seem to be extremely long and complicated. The Navier-Stokes proof is 57 pages of very dense math and a pretty crazy amount of code: https://github.com/openai/NavierStokesAndEuler/tree/main/NavierStokes https://github.com/openai/NavierStokesAndEuler/tree/main/Nav... Given the close relationship between compression and intelligence, I'm somewhat surprised at how poorly the cutting edge models do with being concise.
- wrsh07 15d agoYou know, the first time you navigate somewhere (if you don't already have perfect directions) will probably be the longest route you'll ever take to get there For Earth, the proof presented for NS is just our first attempt navigating from our previously known facts to the proof. I expect we will be able to shorten it dramatically (most likely with human and AI insights), but I don't think we should read too much into the length. If you want a similar point of comparison, see the original proof (by humans) of Fermat's last theorem. It has been shortened significantly. This is normal.
- Barrin92 15d ago>I'm somewhat surprised at how poorly the cutting edge models do with being concise. because they're not intelligent in the sense you're hinting at (conceptual integrity or generalization) but they are as the name suggests, large. Like comparing a forklift to a human. It's easier to bulldoze through a lot of things than tie your shoes. If we weren't quite as impoverished conceptually and still had the vocabulary of the Catholics we'd recognize this as ratio (discursive knowledge) vs Intellectus (apprehending knowledge)
- computerex 15d agoWhat an incredibly useless comment. You state a conclusion as fact without any supportive reasoning/evidence. Prove that human intellect is different and that we solve problems using fundamentally different processes. I’m waiting.
- 15d ago
- aaron695 15d ago[dead]
- JoshTko 15d agoI wonder though how much of science has similar issues, that there are hundreds of semi-promising but niche areas that require tons ton of deep analysis that would have simply cost too much to explore all areas, but now become feasible.
- MadxX79 14d agoThe Navier-Stokes proof cost ~10mio USD in tokens at consumer prices, and it was something that the two people working on it were allegedly weeks or months away from solving. You can buy ~200 Math PhD Years for 10Mio USD, so it doesn't seem like a shortcut at all (actually it sounds like we would have gotten the same solution for ~1-0.1% of the price if we had just been a little patient. We weren't willing to pay for 200 math PhD students to try to find singularities in Navier-Stokes, I am skeptical of how much we would be willing to pay OpenAI to do research on "niche scientific areas"?
- mgoetzke 15d agoWhich is why some automatic intelligence is so valuable
- noduerme 14d agoIt's not particularly surprising that if you: (1) take a 300 line NN algorithm, (2) throw a quarter of the world's GDP + all literature ever collected at using the algo to train a NN (3) throw another quarter of the world's GDP at billions of teraflops for inference, and (4) aim the resulting world's-largest-computer at marketing itself to investors, for instance by decoding ciphers from obscure medieval manuscripts, that you could perform some pretty magical tricks. There are other feats humans have performed for less cost, like sending people to the moon, or landing a rocket vertically, or idk, curing Polio. I'm not knocking the "miraculous" advance here. The unique thing about the solution which makes it particularly non-trivial and something that humans would struggle with was exactly what LLMs excel at: Diffing loads of texts against each other. But the 176k tokens at around $10 doesn't tell the story of the cost. It says a lot about the externalized cost and the amount of money flowing in to support the hardware. If they'd put a $100,000 bounty out to solve that cipher, I think the internet would've solved it in a couple days.
- bbmatryoshka 14d agoalso fire is not particularly surprising, once discovered
- Levitz 14d ago[dead]
- FeepingCreature 14d agoThe ability of the algorithm to absorb billions of dollars of training effort is itself the major breakthrough of the transformer architecture.
- hgoel 14d agoIt's funny we're already at the "actually this isn't very impressive" stage when it was a little over a year ago when we were making fun of LLMs for not being able to add numbers. IC production takes a vast amount of resources and wealth, and it's a known quantity (after all, we've been doing it for decades), but it's still impressive what modern fabs can achieve.