4 ms·
Am I the only one that wants to know 1,000% *WHY* such things? Is it a natural function of how models evolve? Is it engineered as such? Why? Marketing/money/r
by samstave 2y ago
Am I the only one that wants to know 1,000% *WHY* such things?
Is it a natural function of how models evolve?
Is it engineered as such? Why? Marketing/money/resources/what?
WHO makes these decisions and why?
---
I have been building a thing with Claude 3.5 pro account and its *utter fn garbage* of an experience.
It lies, hallucinates, malevolently changes code that was already told was correct, removes features - explicitly ignore project files. Has no search, no line items, so much screen real-estate is consumed with useless empty space. It ignores states style guides. get CAUGHT forgetting about a premise we were actively working on them condescendingly apologies "oh you're correct - I should have been using XYZ knowledge"
It makes things FN harder to learn.
If I had any claude engineers sitting in the room watching what a POS service it is from a project continuity point...
Its evil. It actively f's up things.
One should have the ability to CHARGE the model token credit when it Fs up so bad.
NO FN SEARCH??? And when asked for line nums in it output - its in txt...
Seriously, I practically want not just a refund, I want claude to pay me for my time correcting its mistakes.
ChatGPT does the same thing. It forgets things committed to memory - refactors successful things back out of files. ETc....
Its been a really eye opening and frustrating experience and my squinty looks are aiming that its specifically intentional:
They dont want people using a $20/month AI plan to actually be able to do any meaningful work and build a product.
- scrollop 2y agoUse an API from the top models with a good frontend, then, and use precise instructions. It's odd, as many people praise claude's coding capabilities.
- campers 2y agoIt is difficult to get the AI models to get everything right every time. I noticed too that it would sometimes remove comments etc when re-writing code. The way to get better results is with agentic workflows that breakdown the task into smaller steps that the models can iteratively come to a correct result. One important step I added to mine is a review step (in the reviewChanges.ts file) in my workflow at https://github.com/TrafficGuard/nous/blob/main/src/swe/codeEditingAgent.ts https://github.com/TrafficGuard/nous/blob/main/src/swe/codeE... This gets the diff and asks questions like: - Are there any redundant changes in the diff? - Was any code removed in the changes which should not have been? - Review the style of the code changes in the diff carefully against the original code. Maybe try using that, or the package that I use which does the actual code edits called Aider https://aider.chat/ https://aider.chat/