10 ms·
Some thoughts about Anthropic's new cryptanalysis results
- mkagenius 2mo agoI also just finished writing a note - https://mkagenius.substack.com/p/notes-on-mythos-breaking-aes https://mkagenius.substack.com/p/notes-on-mythos-breaking-ae... https://news.ycombinator.com/item?id=49099977 https://news.ycombinator.com/item?id=49099977
- john_strinlai 2mo ago>They [anthropic] appear to have just told it to get some results and then strapped its nose to the grindstone until it found some. it is fun how well this works. i cant find the link immediately (will look and edit with it), but somewhere in the "hello there the jacobian conjecture is false thanx" thread, someone brought up a different conjecture breakthrough where the prompts were basically just repeated "no, keep going" until a result was found. edit: https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063 https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de0... i especially like "you should do a breakthrough". each prompt is less than ~20 words. makes me really question the whole "prompt engineering" stuff.
- throwup238 2mo ago> i especially like "you should do a breakthrough". each prompt is less than ~20 words. makes me really question the whole "prompt engineering" stuff. This is well past prompt engineering and into process engineering like six sigma. Just like in an early industrial revolution factory, we’re all still figuring out what works in the process of making stuff except this is so early that even simple things like “make this screw standardized” (or “no, keep going” in this case) is really high impact. The degrees of freedom an LLM has is so large that we're going to be exploring their capabilities for decades, especially if they continue to get better. This is why IMO experts are always going to be better at LLMs in their field because they can force them LLM into processes (think prompt engineering -> CC dynamic workflows) that follow their work processes and get much better results out of them than “keep going.”
- adamzenith 2mo agoIn response to your edit, you should check out Terry tao's chat gpt logs about the recent Jacobian result. The models are smart enough to brute force some things, but can cut to the meat much faster with good prompting
- john_strinlai 2mo agoi read his, too. his replies are indeed more directed, but also quite short, unstructured, and natural sounding. if i recall, maybe 1 or 2 of his prompts exceeded 50(ish) words. in my head, the comparison is the multi-paragraph prompts (borderline essays) i would read in various communities on reddit and similar forums, that people (often self-proclaimed "prompt engineers") said were "required" to get good output. or some of the prompts ive read in various logs that are like a thousand words of setup. even looking back at the first prompts i was sending when i started to use chatgpt were (in hindsight) crazy long and full of unnecessary guidance/caveats/"ignore xyz"/etc.
- fn-mote 2mo ago> were (in hindsight) crazy long and full of unnecessary guidance Not sure when you started. However: I would never judge the necessity of details provided 2 or 3 years ago based on results that the current models give.
- annzabelle 2mo agoI didn't use generative AI much until a couple months ago. My last employer had a copilot license that I dabbled in in 2024, and wasn't impressed with, and then I spent the latter half of 2025 and early 2026 budget traveling/hiking a lot. I started at a new company in May, and I've been astonished at how lazy I can be at prompting and still get impressive results with 2026 Claude Code. I can paste entire failure logs with the word "why" lowercase, no question mark, and get into a productive chat session where it significantly speeds up the bug trace. I will paste the text from a groomed ticket with no editing or additional instructions into the chat and then give it a bit of feedback on the plan for a couple iterations. I feel so vindicated in never spending time learning prompting as a specific skill, it really just took a couple more years and the models are really easy to interact with with simple natural language.
- dist-epoch 2mo agoThat second person stated that for many years they tried that particular graph problem on various AI models, starting with o1 and o3. Its quite likely they now found the counterexample with a more serious prompt, and then for virality re-tried a few times with meme-prompts like "you should do a breakthrough", knowing that the model is capable of solving this particular one. Worst case the meme-prompts don't work and they share the real one they initially used.
- woah 2mo agoNot everything is a conspiracy
- john_strinlai 2mo ago>Its quite likely they now found the counterexample with a more serious prompt, and then for virality re-tried a few times with meme-prompts i am not sure why this is "quite likely". it'd be pretty silly to get a mathematical breakthrough and then hide it for an undisclosed amount of time to get a few more likes on a tweet, when the impressive part is the breakthrough. not saying your theory is impossible, but i think the simple answer is that the model is just smarter than o1 and o3. and, in any case, the model ended up getting the result with the meme prompt and "keep going", which was what i find fun. just like how the crypto results were from prompts of, more or less, "keep going", and that's pretty damn cool.
- dist-epoch 2mo agoYou can open 10 tabs and prompt 10 times. We are talking about a few hours delay. I agree with you that obviously no prompt engineering was needed just "solve this problem", but imagine it was you doing this problem with every model, wouldn't you have tested a new model with the best prompt you had from previous iterations, maybe with some partial previous results in it, exactly to maximize your probability for a mathematical breakthrough?
- alwa 2mo agoOh boy. The human here is effectively assuming the role of a Magic 8 Ball… What a weird species of halting problem…
- inigyou 2mo agoSame with cybersec. They're finding bugs a human could've found if they looked hard but humans don't look hard at 100% of the code and the LLM can, at high speed.
- swordsith 2mo agoCould send random characters along with 'keep going' and it would change nothing, the seed is whats causing deviation between these responses.
- simonw 2mo ago> both outputs of Claude Mythos, their (still) unreleased advanced model That sentence gives the impression that Mythos might be released in the future. That's clearly not going to happen - it's already "released" in as much as selected, trusted partners can access it, and the rest of us get it in the form of Fable - which is Mythos but with filters that downgrade you if you try to use it for anything even remotely related to cybersecurity or biology. (The other day Fable 5 downgraded me to Opus after I asked it to explain the difference between tusks and teeth.)
- deleted 2mo ago[deleted]
- free_bip 2mo agoGenuine question here, why would you ask Fable to explain the difference between tusks and teeth? That's a task that can probably be handled by Haiku.
- simonw 2mo agoIt's the default when I pop open the Claude iPhone app, I usually don't bother to switch it.
- deleted 2mo ago[deleted]
- rain_iwakura 2mo agobecause as OP of the article points out, it's not a guaranteed hit with these models no matter how capable. they are stochastic (but not parrots), humans are too (but in a different way), and so there is a chance that the answer is wrong. If I wanted the most accurate answer I'd be confident in believing, I'd ask the most advanced model. Now obviously you can and should retort with hallucination and confabulation rates from external and Ant's own reports per model (pretty sure more advanced models are good at lying better, not less) instead of going with dumb "more expensive more accurate" mental model, but general principle stands for me still AFAIK. it's not exactly rational I admit, but if I'm going to base my own work and reasoning from an LLM I'm going with the best available. This seems to trip up most normies because they are too lazy or too greedy to pay up for premium access and see for themselves why most of us are both awed and afraid. Generally, I'm too biased and too deep in ML/DL cargo cult (been in it since 2016) to know if the skepticism and disdain for such usage is warranted. In general, I think the tools are broadly toxic in a Dune-sense of making me think less for myself, because just as any HN-poster knows coding and doing mundane low-level stuff is necessary the same way doing stretches is necessary before any workout. The process itself is what keeps your brain strong and its gradients from veering into overfitting. I'm not overly bullish on the whole reaching for the stars ending with these things. Paradoxically, you using them eventually hobbles both you and the model, because you become dumber and then you bottleneck their ability to self-direct (broadly true for next Mythos/GPT-7). Sorry for a long rant, was just anticipating some things I'd have to say for myself.
- throawayonthe 2mo ago> I asked Claude for its thoughts, and it doesn’t mince words: “what makes this genuinely interesting — and, frankly, a little embarrassing for the field — is that none of the ingredients are exotic.” The TL;DR is that someone just did a much more thorough job applying all of our known tools. In short: the sort of things that attack AIs are wonderful at. did i just read two summaries/TLDRs (in a row) of the already-two-sentence summary right above?
- Ar-Curunir 2mo agoThere a constructive way to leave feedback you know
- simonw 2mo agoThis is good: > If you’re under the impression that these models are “glorified autocomplete” or that progress is slowing down, I need to urge you: stop thinking that. The models are very intelligent and capable, they are getting better at a fast clip. I can cite measurable and impressive progress over just the past five months on specific types of problem I’ve asked them to look at. [...] > On the other hand: if you think that models are super-intelligent or that AGI is already here, you should also stop thinking that. Working with these tools is like swimming in a pond where the ground drops off sharply. One minute you’re wading comfortably and there’s support under your feet. Then suddenly you cross a specific line, and you’re back to swimming on your own.
- dboreham 2mo agoI'm also getting irritated with the “glorified autocomplete” comments. Since nobody can post such comments and also use the tools I'm using, I'm wondering if the phenomenon is due to people only having experience with the free version of whatever it is they're trying to use?
- dgellow 2mo agoThe “glorified autocomplete” framing isn’t to take literally. It’s a way to remove the mystic and whole anthropomorphization of AI. It’s saying they aren’t sentient or entities we are interacting with, even if that’s how the output presents itself. Instead they are “just” stochastic models
- simonw 2mo agoSome people use it to demystify, but a whole lot of people seem to be using it to dismiss the technology entirely. Personally I like to remind people that these things are next-token predictors, but then emphasize how truly astonishing the results we can get out of sufficiently advanced next-token predictors are.
- ameliaquining 2mo ago
- deleted 2mo ago[deleted]
- bawolff 2mo agoThe prompt used to get this result is pretty crazy.
- iansmith_hn 2mo agoDoes anybody have a filter for cryptography posts of the form: AES IS BROKEN: Making giant assumption XYZ and requiring a less capable AES in ABC way, we've reduce the amount of operations needed to break AES from 10^X to 10^X-1! These are tedious for people who are interested in cryptography but are not researchers in the field. (For the researchers, this type of thing may be useful.). The fact that AI is now "generating crypto results", suggests that soon we will soon have crypto-post-slop as clickbait...