4 ms·
How anybody can read stuff like this and still take all this seriously is beyond me. This is becoming the engineering equivalent of astrology.
by scuff3d 8mo ago
How anybody can read stuff like this and still take all this seriously is beyond me. This is becoming the engineering equivalent of astrology.
- fragmede 8mo agoFeel free to run your own tests and see if the magic phrases do or do not influence the output. Have it make a Todo webapp with and without those phrases and see what happens!
- scuff3d 8mo agoThat's not how it works. It's not on everyone else to prove claims false, it's on you (or the people who argue any of this had a measurable impact) to prove it actually works. I've seen a bunch of articles like this, and more comments. Nobody I've ever seen has produced any kind of measurable metrics of quality based on one approach vs another. It's all just vibes. Without something quantifiable it's not much better then someone who always wears the same jersey when their favorite team plays, and swears they play better because of it.
- tokioyoyo 8mo agoDo you actively use LLMs to do semi-complex coding work? Because if not, it will sound mumbo-jumbo to you. Everyone else can nod along and read on, as they’ve experienced all of it first hand.
- scuff3d 8mo agoYou've missed the point. This isn't engineering, it's gambling. You could take the exact same documents, prompts, and whatever other bullshit, run it on the exact same agent backed by the exact same model, and get different results every single time. Just like you can roll dice the exact same way on the exact same table and you'll get two totally different results. People are doing their best to constrain that behavior by layering stuff on top, but the foundational tech is flawed (or at least ill suited for this use case). That's not to say that AI isn't helpful. It certainly is. But when you are basically begging your tools to please do what you want with magic incantations, we've lost the fucking plot somewhere.
- gf000 8mo ago> You could take the exact same documents, prompts, and whatever other bullshit, run it on the exact same agent backed by the exact same model, and get different results every single time This is more of an implementation detail/done this way to get better results. A neural network with fixed weights (and deterministic floating point operations) returning a probability distribution, where you use a pseudorandom generator with a fixed seed called recursively will always return the same output for the same input.
- geoelectric 8mo agoI think that's a pretty bold claim, that it'd be different every time. I'd think the output would converge on a small set of functionally equivalent designs, given sufficiently rigorous requirements. And even a human engineer might not solve a problem the same way twice in a row, based on changes in recent inspirations or tech obsessions. What's the difference, as long as it passes review and does the job?
- guiambros 8mo agoIf you read the transformer paper, or get any book on NLP, you will see that this is not magic incantation; it's purely the attention mechanism at work. Or you can just ask Gemini or Claude why these prompts work. But I get the impression from your comment that you have a fixed idea, and you're not really interested in understanding how or why it works. If you think like a hammer, everything will look like a nail.
- scuff3d 8mo agoI know why it works, to varying and unmeasurable degrees of success. Just like if I poke a bull with a sharp stick, I know it's gonna get it's attention. It might choose to run away from me in one of any number of directions, or it might decide to turn around and gore me to death. I can't answer that question with any certainty then you can. The system is inherently non-deterministic. Just because you can guide it a bit, doesn't mean you can predict outcomes.
- winrid 8mo agoBut we can predict the outcomes, though. That's what we're saying, and it's true. Maybe not 100% of the time, but maybe it helps a significant amount of the time and that's what matters. Is it engineering? Maybe not. But neither is knowing how to talk to junior developers so they're productive and don't feel bad. The engineering is at other levels.
- imiric 8mo ago> But we can predict the outcomes [...] Maybe not 100% of the time So 60% of the time, it works every time. ... This fucking industry.
- winrid 8mo agoAgain, it's called management. You're managing something unpredictable: the LLM. This is nothing new, at all. Do you have a strategy that makes other engineers do what you want exactly 100% of the time?
- yaku_brang_ja 8mo agoThese coding agents are literally Language Models. The way you structure your prompting language affect the actual output.
- energy123 8mo agoAnthropic recommends doing magic invocations: https://simonwillison.net/2025/Apr/19/claude-code-best-practices/ https://simonwillison.net/2025/Apr/19/claude-code-best-pract... It's easy to know why they work. The magic invocation increases test-time compute (easy to verify yourself - try!). And an increase in test-time compute is demonstrated to increase answer correctness (see any benchmark). It might surprise you to know that the only different between GPT 5.2-low and GPT 5.2-xhigh is one of these magic invocations. But that's not supposed to be public knowledge.
- gehsty 8mo agoI think this was more of a thing on older models. Since I started using Opus 4.5 I have not felt the need to do this.
- bavell 8mo agoAnthropic got rid of controlling the thinking budget by parsing your prompt - now it's a setting in /config.
- energy123 8mo agoThey never parsed your prompt. The magic word reduces the probability that the token corresponding to the end of chain-of-thought will be emitted, which increases test-time compute.
- cloudbonsai 8mo agoThe evolution of software engineering is fascinating to me. We started by coding in thin wrappers over machine code and then moved on to higher-level abstractions. Now, we've reached the point where we discuss how we should talk to a mystical genie in a box. I'm not being sarcastic. This is absolutely incredible.
- intrasight 8mo agoAnd I've been had a long enough to go through that whole progression. Actually from the earlier step of writing machine code. It's been and continues to be a fun journey which is why I'm still working.
- sumedh 8mo agoWe have tests and benchmarks to measure it though.
- yawnr 8mo agoNice to hear someone say it. Like what are we even doing? It's exhausting.