5 ms·
[flagged]
by ctoth 6mo ago
[flagged]
- quietsegfault 6mo agoI’m not sure being confrontational like this really helps your case. There are real people responding, and even if you’re frustrated it doesn’t pay off to take that frustration out on the people willing to help.
- malfist 6mo agoIs somebody saying "you're holding it wrong" a "people willing to help"?
- Retr0id 6mo ago[flagged]
- TeMPOraL 6mo agoThey are if you are, in fact, holding it wrong. As was the usual case in most of the few years LLMs existed in this world. Think not of iPhone antennas - think of a humble hammer. A hammer has three ends to hold by, and no amount of UI/UX and product design thinking will make the end you like to hold to be a good choice when you want to drive a Torx screw.
- ctoth 6mo agoFair point on tone. It's a bit of a bind isn't it? When you come with a well-researched issue as OP did, you get this bland corporate nonsense "don't believe your lyin' eyes, we didn't change anything major, you can fix it in settings." How should you actually communicate in such a way that you are actually heard when this is the default wall you hit? The author is in this thread saying every suggested setting is already maxed. The response is "try these settings." What's the productive version of pointing out that the answer doesn't address the evidence? Genuine question. I linked my repo because it's the most concrete example I have.
- wonnage 6mo agoJust use a different tool or stop vibe coding, it’s not that hard. I really don’t understand the logic of filing bug reports against the black box of AI
- geysersam 6mo agoPeople file tickets against closed source "black box" systems all the time. You could just as well say: Stop using MS SQL, just use a different tool, it's not that hard.
- wonnage 6mo agoEquivalent of filing a ticket against the slot machine when you lose more often than expected
- HumanOstrich 6mo agoWell now you're just being silly and I can't take you seriously.
- HumanOstrich 6mo agoThe only "black box" here is Anthropic. At least an LLM's performance and consistency can be established by statistical methods.
- enraged_camel 6mo agoI read the entire performance degradation report in the OP, and Boris's response, and it seems that the overwhelming majority of the report's findings can indeed be explained by the `showThinkingSummaries` option being off by default as of recently.
- BigTTYGothGF 6mo agoThe stated policy of HN is "don't be mean to the openclaw people", let's see if it generalizes.
- throwaway613746 6mo ago[dead]
- malfist 6mo agoIt also completely ignores the increase in behavioral tracking metrics. 68% increase in swearing at the LLM for doing something wrong needs to be addressed and isn't just "you're holding it wrong"
- alchemist1e9 6mo agoI’m think a great marketing line for local/selfhosted LLMs in the future - “You can swear at your LLM and nobody will care!”
- iwalton3 6mo ago[dead]
- lambda 6mo agoI guess one of the things I don't understand: how you expect a stochastic model, sold as a proprietary SaaS, with a proprietary (though briefly leaked) client, is supposed to be predictable in its behavior. It seems like people are expecting LLM based coding to work in a predictable and controllable way. And, well, no, that's not how it works, and especially so when you're using a proprietary SaaS model where you can't control the exact model used, the inference setup its running on, the harness, the system prompts, etc. It's all just vibes, you're vibe coding and expecting consistency. Now, if you were running a local weights model on your own inference setup, with an open source harness, you'd at least have some more control of the setup. Of course, it's still a stochastic model, trained on who knows what data scraped from the internet and generated from previous versions of the model; there will always be some non-determinism. But if you're running it yourself, you at least have some control and can potentially bisect configuration changes to find what caused particular behavior regressions.
- bcherny 6mo agoChristopher, would you be able to share the transcripts for that repo by running /bug? That would make the reports actionable for me to dig in and debug.
- dang 6mo agoPlease don't post this aggressively to Hacker News. You can make your substantive points without that. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html