13 ms·
Claude Code can debug low-level cryptography
- qsort 11mo agoThis resonates with me a lot: > As ever, I wish we had better tooling for using LLMs which didn’t look like chat or autocomplete I think part of the reason why I was initially more skeptical than I ought to have been is because chat is such a garbage modality. LLMs started to "click" for me with Claude Code/Codex. A "continuously running" mode that would ping me would be interesting to try.
- cmrdporcupine 11mo agoI absolutely agree with this sentiment as well and keep coming back to it. What I want is more of an actual copilot which works in a more paired way and forces me to interact with each of its changes and also involves me more directly in them, and teaches me about what it's doing along the way, and asks for more input. A more socratic method, and more augmentic than "agentic". Hell, if anybody has investment money and energy and shares this vision I'd love to work on creating this tool with you. I think these models are being misused right now in attempt to automate us out of work when their real amazing latent power is the intuition that we're talking about on this thread. Misused they have the power to worsen codebases by making developers illiterate about the very thing they're working on because it's all magic behind the scenes. Uncorked they could enhance understanding and help better realize the potential of computing technology.
- mccoyb 11mo agoI'm working on such a thing, but I'm not interested in money, nor do I have money to offer - I'm interested in a system which I'm proud of. What are your motivations? Interested in your work: from your public GitHub repos, I'm perhaps most interested in `moor` -- as it shares many design inclinations that I've leaned towards in thinking about this problem.
- cmrdporcupine 11mo agoUnfortunately... mooR is my passion project, but I also need to get paid, and nobody is paying me for that. I'm off work right now, between jobs and have been working 10, 12 hours a day on it. That will shortly have to end. I applied for a grant and got turned down. My motivations come down to making a living doing the things I love. That is increasingly hard.
- reachableceo 11mo agoHave you tried to ask the agents to work with you in the way you want? I’ve found that using some high level direction / language and sharing my wants / preferences for workflow and interaction works very well. I don’t think that you can find an off the shelf system todo what you want. I think you have to customize it to your own needs as you go. Kind of like how you customize emacs as it’s running to your desires. I’ve often wondered if you could put a mini LLM into emacs or vscode and have it implement customizations :)
- cmrdporcupine 11mo agoI have, but the problem is in part the tool itself and the way it works. It's just not written with an interactive prompting style in mind. CC is like "Accept/Ask For Changes/Reject" for often big giant diffs, and it's like... no, the UI should be: here's an editor let's work on this together, oh I see what you did there, etc...
- braebo 11mo agoThis is why I still prefer Cursors workflow to the CLIs!
- imiric 11mo agoOn the one hand, I agree with this. The chat UI is very slow and inefficient. But on the other, given what I know about these tools and how error-prone they are, I simply refuse to give them access to my system, to run commands, or do any action for me. Partly due to security concerns, partly due to privacy, but mostly distrust that they will do the right thing. When they screw up in a chat, I can clean up the context and try again. Reverting a removed file or messed up Git repo is much more difficult. This is how you get a dropped database during code freeze... The idea of giving any of these corporations such privileges is unthinkable for me. It seems that most people either don't care about this, or are willing to accept it as the price of admission. I experimented with Aider and a self-hosted model a few months ago, and wasn't impressed. I imagine the experience with SOTA hosted models is much better, but I'll probably use a sandbox next time I look into this.
- cmrdporcupine 11mo agoAider hurt my head it did not seem... good. Sorry to say. If you want open source and want to target something over an API "crush" https://github.com/charmbracelet/crush https://github.com/charmbracelet/crush is excellent But you should try Claude Code or Codex just to understand them. Can always run them in a container or VM if you fear their idiocy (and it's not a bad idea to fear it) Like I said sibling, it's not the right modality. Others agree. I'm a good typer and good at writing, so it doesn't bug me too much, but it does too much without asking or working through it. Sometimes this is brilliant. Other times it's like.. c'mon guy, what did you do over there? What Balrog have I disturbed? It's good to be familiar with these things in any case because they're flooding the industry and you'll be reviewing their code for better or for worse.
- gdevenyi 11mo agoComing soon, adversarial attacks on LLM training to ensure cryptographic mistakes.
- Frannky 11mo agoCLI terminals are incredibly powerful. They are also free if you use Gemini CLI or Qwen Code. Plus, you can access any OpenAI-compatible API(2k TPS via Cerebras at 2$/M or local models). And you can use them in IDEs like Zed with ACP mode. All the simple stuff (creating a repo, pushing, frontend edits, testing, Docker images, deployment, etc.) is automated. For the difficult parts, you can just use free Grok to one-shot small code files. It works great if you force yourself to keep the amount of code minimal and modular. Also, they are great UIs—you can create smart programs just with CLI + MCP servers + MD files. Truly amazing tech.
- BrokenCogs 11mo agoHow good is Gemini CLI compared to Claude code and openAi codex?
- Frannky 11mo agoI started with Claude Code, realized it was too much money for every message, then switched to Gemini CLI, then Qwen. Probably Claude Code is better, but I don't need it since I can solve my problems without it.
- cmrdporcupine 11mo agoTry what I've done: use the Claude Code tool but point your ANTHROPIC_URL at a DeepSeek API membership. It's like 1/10th the cost, and about 2/3rds the intelligence. Sometimes I can't really tell.
- behnamoh 11mo agoI use this to proxy ANTHROPIC_BASE_URL to other models: https://github.com/ujisati/claude-code-provider-proxy https://github.com/ujisati/claude-code-provider-proxy unfortunately it doesn't support local models but they're too slow for coding anyway.
- 11mo ago
- delaminator 11mo ago> For example, how nice would it be if every time tests fail, an LLM agent was kicked off with the task of figuring out why, and only notified us if it did before we fixed it? You can use Git hooks to do that. If you have tests and one fails, spawn an instance of claude a prompt -p 'tests/test4.sh failed, look in src/ and try and work out why' $ claude -p 'hello, just tell me a joke about databases' A SQL query walks into a bar, walks up to two tables and asks, "Can I JOIN you?" $ Or if, you use Gogs locally, you can add a Gogs hook to do the same on pre-push > An example hook script to verify what is about to be pushed. Called by "git push" after it has checked the remote status, but before anything has been pushed. If this script exits with a non-zero status nothing will be pushed. I like this idea. I think I shall get Claude to work out the mechanism itself :) It is even a suggestion on this Claude cheet sheet https://www.howtouselinux.com/post/the-complete-claude-code-cli-cheat-sheet-and-guide https://www.howtouselinux.com/post/the-complete-claude-code-...
- jamesponddotco 11mo agoThis could probably be implemented as a simple Bash script, if the user wants to run everything manually. I might just do that to burn some time.
- delaminator 11mo agosure, there a multiple ways of spawning an instance the only thing I imagine might be problem is claude demanding a login token as it happens quite regularly
- simonw 11mo agoUsing coding agents to track down the root cause of bugs like this works really well: > Three out of three one-shot debugging hits with no help is extremely impressive. Importantly, there is no need to trust the LLM or review its output when its job is just saving me an hour or two by telling me where the bug is, for me to reason about it and fix it. The approach described here could also be a good way for LLM-skeptics to start exploring how these tools can help them without feeling like they're cheating, ripping off the work of everyone who's code was used to train the model or taking away the most fun part of their job (writing code). Have the coding agents do the work of digging around hunting down those frustratingly difficult bugs - don't have it write code on your behalf.
- jack_tripper 11mo ago>Have the coding agents do the work of digging around hunting down those frustratingly difficult bugs - don't have it write code on your behalf. Why? Bug hunting is more challenging and cognitive intensive than writing code.
- theptip 11mo agoBug hunting tends to be interpolation, which LLMs are really good at. Writing code is often some extrapolation (or interpolating at a much more abstract level).
- Terr_ 11mo agoReversed version: Prompting-up fresh code tends to be translation, which LLMs are really good at. Bug hunting is often some logical reasoning (or translating business-needs at a much more abstract level.)
- deleted 11mo ago[deleted]
- simonw 11mo agoSometimes it's the end of the day and you've been crunching for hours already and you hit one gnarly bug and you just want to go and make a cup of tea and come back to some useful hints as to the resolution.
- lordnacho 11mo agoI'm not surprised it worked. Before I used Claude, I would be surprised. I think it works because Claude takes some standard coding issues and systematizes them. The list is long, but Claude doesn't run out of patience like a human being does. Or at least it has some credulity left after trying a few initial failed hypotheses. This being a cryptography problem helps a little bit, in that there are very specific keywords that might hint at a solution, but from my skim of the article, it seems like it was mostly a good old coding error, taking the high bits twice. The standard issues are just a vague laundry list: - Are you using the data you think you're using? (Bingo for this one) - Could it be an overflow? - Are the types right? - Are you calling the function you think you're calling? Check internal, then external dependencies - Is there some parameter you didn't consider? And a bunch of others. When I ask Claude for a debug, it's always something that makes sense as a checklist item, but I'm often impressed by how it diligently followed the path set by the results of the investigation. It's a great donkey, really takes the drudgery out of my work, even if it sometimes takes just as long.
- ay 11mo ago> Claude doesn't run out of patience like a human being does. It very much does! I had a debugging session with Claude Code today, and it was about to give up with the message along the lines of “I am sorry I was not able to help you find the problem”. It took some gentle cheering (pretty easy, just saying “you are doing an excellent job, don’t give up!”) and encouragement, and a couple of suggestions from me on how to approach the debug process for it to continue and finally “we” (I am using plural here because some information that Claude “volunteered” was essential to my understanding of the problem) were able to figure out the root cause and the fix.
- lordnacho 11mo agoThat's interesting, that only happened to me on GPT models in Cursor. It would apologize profusely.
- ericphanson 11mo agoClaude told me it stopped debugging since it would run out of tokens in its context window. I asked how many tokens it had left and it said actually it had plenty so could continue. Then again it stopped, and without me asking about tokens, wrote Context Usage • Used: 112K/200K tokens (56%) • Remaining: 88K tokens • Sufficient for continued debugging, but fresh session recommended for clarity lol. I said ok use a subagent for clarity.
- rvz 11mo agoAs declared by an expert in cryptography who knows how to guide the LLM into debugging low-level cryptography, which that's good. Quite different if you are not a cryptographer or a domain expert.
- tptacek 11mo agoEven the title of the post makes this clear: it's about debugging low-level cryptography. He didn't vibe code ML-DSA. You have to be building a low-level implementation in the first place for anything in this post to apply to it.
- XenophileJKO 11mo agoPersonally my biggest piece of advice is: AI First. If you really want to understand what the limitations are of the current frontier models (and also really learn how to use them), ask the AI first. By throwing things over the wall to the AI first, you learn what it can do at the same time as you learn how to structure your requests. The newer models are quite capable and in my experience can largely be treated like a co-worker for "most" problems. That being said.. you also need to understand how they fail and build an intuition for why they fail. Every time a new model generation comes up, I also recommend throwing away your process (outside of things like lint, etc.) and see how the model does without it. I work with people that have elaborate context setups they crafted for less capable models, they largely are un-neccessary with GPT-5-Codex and Sonnet 4.5.
- imiric 11mo ago> By throwing things over the wall to the AI first, you learn what it can do at the same time as you learn how to structure your requests. Unfortunately, it doesn't quite work out that way. Yes, you will get better at using these tools the more you use them, which is the case with any tool. But you will not learn what they can do as easily, or at all. The main problem with them is the same one they've had since the beginning. If the user is a domain expert, then they will be able to quickly spot the inaccuracies and hallucinations in the seemingly accurate generated content, and, with some effort, coax the LLM into producing correct output. Otherwise, the user can be easily misled by the confident and sycophantic tone, and waste potentially hours troubleshooting, without being able to tell if the error is on the LLM side or their own. In most of these situations, they would've probably been better off reading the human-written documentation and code, and doing the work manually. Perhaps with minor assistance from LLMs, but never relying on them entirely. This is why these tools are most useful to people who are already experts in their field, such as Filippo. For everyone else who isn't, and actually cares about the quality of their work, the experience is very hit or miss. > That being said.. you also need to understand how they fail and build an intuition for why they fail. I've been using these tools for years now. The only intuition I have for how and why they fail is when I'm familiar with the domain. But I had that without LLMs as well, whenever someone is talking about a subject I know. It's impossible to build that intuition with domains you have little familiarity with. You can certainly do that by traditional learning, and LLMs can help with that, but most people use them for what you suggest: throwing things over the wall and running with it, which is a shame. > I work with people that have elaborate context setups they crafted for less capable models, they largely are un-neccessary with GPT-5-Codex and Sonnet 4.5. I haven't used GPT-5-Codex, but have experience with Sonnet 4.5, and it's only marginally better than the previous versions IME. It still often wastes my time, no matter the quality or amount of context I feed it.
- cluckindan 11mo agoSo the ”fix” includes a completely new function? In a cryptography implementation? I feel like the article is giving out very bad advice which is going to end up shooting someone in the foot.
- OneDeuxTriSeiGo 11mo agoThe article even states that the solution claude proposed wasn't the point. The point was finding the bug. AI are very capable heuristics tools. Being able to "sniff test" things blind is their specialty. i.e. Treat them like an extremely capable gas detector that can tell you there is a leak and where in the plumbing it is, not a plumber who can fix the leak for you.
- rizky05 11mo ago[dead]
- thadt 11mo agoCan you expand on what you find to be 'bad advice'? The author uses an LLM to find bugs and then throw away the fix and instead write the code he would have written anyway. This seems like a rather conservative application of LLMs. Using the 'shooting someone in the foot' analogy - this article is an illustration of professional and responsible firearm handling.
- lisbbb 11mo agoHonestly, it read more like attention seeking. He "live coded" his work, by which I believe he means he streamed everything he was doing while working. It just seems so much more like a performance and building a brand than anything else. I guess that's why I'm just a nobody.
- sciencejerk 11mo agoLayman in cryptotography (that's 99% of us at least) may be encouraged to deploy LLM generated crypto implementations, without understanding the crypto
- didibus 11mo agoThis is basically the ideal scenario for coding agents. Easily verifiable through running tests, pure logic, algorithmic problem. It's the case that has worked the best for me with LLMs.
- pton_xd 11mo ago> Full disclosure: Anthropic gave me a few months of Claude Max for free. They reached out one day and told me they were giving it away to some open source maintainers. Related, lately I've been getting tons of Anthropic Instagram ads; they must be near a quarter of all the sponsored content I see for the last month or so. Various people vibe coding random apps and whatnot using different incarnations of Claude. Or just direct adverts to "Install Claude Code." I really have no idea why I've been targeted so hard, on Instagram of all places. Their marketing team must be working overtime.
- simonw 11mo agoI think it might be that they've hit product-market fit. Developers find Claude Code extremely useful (once they figure out how to use it). Many developers subscribe to their $200/month plan. Assuming that's profitable (and I expect it is, since even for that much money it cuts off at a certain point to avoid over-use) Anthropic would be wise to spend a lot of money on marketing to try and grow their paying subscriber base for it.
- chatmasta 11mo agoWhat makes it better than VSCode Co-pilot with Claude 4.5? I barely program these days since I switched to PM but I recently started using that and it seems pretty effective… why should I use a fork instead?
- danielbln 11mo agoClaude Code is not a VSCode fork, it's a terminal CLI. It's a rather different interaction paradigm compared to your classical IDE (that said, you can absolutely run Claude Code inside a terminal inside VSCode).
- chatmasta 11mo agoAh, I think I’m getting it confused with Cursor. So Claude Code is a terminal CLI for orchestrating a coding agent via prompts? That’s different than the initial releases of VSC copilot, but now VSC has “agent” mode that sounds a lot like this. It basically reduces the IDE to a diff viewer.
- phendrenad2 11mo agoA whole class of tedious problems have been eliminated by LLMs because they are able to look at code in a "fuzzy" way. But this can be a liability, too. I have a codebase that "looks kinda" like a nodejs project, so AI agents usually assume it is one, even if I rename the package.json, it will inspect the contents and immediately clock it as "node-like".
- deadbabe 11mo agoWith AI, we will finally be able to do the impossible: roll our own crypto.
- oytis 11mo agoIt is very much possible, it's just a bad idea. Doubly so with AI.
- tptacek 11mo agoThat's exactly not what he's doing.
- lisbbb 11mo agoYou're not going to do better than the NSA.
- deadbabe 11mo agoI don’t have to. LLMs built by trillion dollar companies will do it for me.
- marginalia_nu 11mo agoTo be fair you're also not going to be backdoored by the NSA.
- spacechild1 11mo ago> Importantly, there is no need to trust the LLM or review its output when its job is just saving me an hour or two by telling me where the bug is, for me to reason about it and fix it. Except they regularly come up with "explanations" that are completely bogus and may actually waste an hour or two. Don't get me wrong, LLMs can be incredibly helpful for identifying bugs, but you still have to keep a critical mindset.
- danielbln 11mo agoOP said "for me to reason about it", not for the LLM to reason about it. I agree though, LLMs can be incredible debugging tools, but they are also incredibly gullable and love to jump to conclusions. The moment you turn your own fleshy brain off is when they go to lala land.
- spacechild1 11mo ago> OP said "for me to reason about it", not for the LLM to reason about it. But that's what I meant! Just recently I asked an LLM about a weird backtrace and it pointed me the supposed source of the issue. It sounded reasonable and I spent 1-2 hours researching the issue, only to find out it was a total red herring. Without the LLM I wouldn't have gone down that road in the first place. (But again, there have been many situations where the LLM did point me to the actual bug.)
- danielbln 11mo agoYeah that's fair, I've been there before myself. It doesn't help when it throws "This is the smoking gun!" at you. I've started using subagents more, specifically a subagent that shells out codex. This way I can have Claude throw a problem over to GPT5 and both can come to a consensus. Doesn't completely prevent wild goose chases, but it helps a lot. I also agree that many more times the LLM is like a blood hound leading me to the right thing (which makes it all the more annoying the few times when it chases a red herring).
- jasonjmcghee 11mo agoI found llm debugging to work better if you give the llm access to a debugger. You can build this pretty easily: https://github.com/jasonjmcghee/claude-debugs-for-you https://github.com/jasonjmcghee/claude-debugs-for-you
- zcw100 11mo agoI just recently found a number of bugs in both the RELIC and MCL libraries. It took a while to track them down but it was remarkable that it was able to find them at all.
- nikanj 11mo agoI'm surprised it didn't fix it by removing the code. In my experience, if you give Claude a failing test, it fixes it by hard-coding the code to return the value expected by the test or something similar. Last week I asked it to look at why a certain device enumeration caused a sigsegv, and it quickly solved the issue by completely removing the enumeration. No functionality, no bugs!
- pessimizer 11mo agoI've got a paste in prompt that reiterates multiple times not to remove features or debugging output without asking first, and not to blame the test file/data that the program failed on. Repeated multiple times, the last time in all caps. It still does it. I hope maybe half as often, but I may be fooling myself.
- Thorrez 11mo ago>so I checked out the old version of the change with the bugs (yay Jujutsu!) and kicked off a fresh Claude Code session There's a risk there that the AI could find the solution by looking through your history to find it, instead of discovering it directly in the checked-out code. AI has done that in the past: https://news.ycombinator.com/item?id=45214670 https://news.ycombinator.com/item?id=45214670
- deleted 11mo ago[deleted]
- wrs 11mo agoYesterday I tried something similar with a website (Go backend) that was doing a complex query/filter and showing the wrong results. I just described the problem to Claude Code, told it how to run psql, and gave it an authenticated cookie to curl the backend with. In about three minutes of playing around, it fixed the problem. It only needed raw psql and curl access, no specialized tooling, to write a bunch of bash commands to poke around and compare the backend results with the test database.
- jerf 11mo agoThis is one of the things I've mentioned before, I think it's just hidden a bit and hard to see, but this is basically the LLM doing style transfer, which they're really good at. There was a specification for the code (which looks like it was already trained into the LLM since it didn't have to go fetch it but it also had intimate knowledge of), there was an implementation, and it's really good at extracting out the style difference between code and spec. Anything that looks like style transfer is a good use for LLMs. As another example, I think things like "write unit tests for this code" are usually similar sort of style transfer as well, based on how it writes the tests. It definitely has a good idea as to how to sort of ensure that all the functionality gets tested, I find it is less likely to produce "creative" ways that bugs may come out, but hey, it's a good start. This isn't a criticism, it's intended to be a further exploration and understanding of when these tools can be better than you might intuitively think.