3 ms·
If that same intern, when asked something, responded that they checked, gave you a link to a document they claim have the proof / answer but does in fact no and
by jpc0 1y ago
If that same intern, when asked something, responded that they checked, gave you a link to a document they claim have the proof / answer but does in fact no and continued to do that they wouldn’t be an intern very long. But somehow this is acceptable behaviour for an AI?
I use AI for sure, but only on things that I can easily verify is correct (run a test or some code ), because I have had the AI give me functions in an API with links to online documentation for those functions, the document exists, the function is not in it, when called out instead of doing a basic tool call the AI will double down that it is correct and you the human are wrong. That would get an intern fired but here you are standing on the interns side.
- raincole 1y agoWow... people unironically anthropomorphize AI to the point that they expect AI to work exactly like an human intern, otherwise it's unacceptable...
- simonw 1y agoBecause LLMs aren't human beings. I wrote a note about that here: https://simonwillison.net/2025/Mar/11/using-llms-for-code/#set-reasonable-expectations https://simonwillison.net/2025/Mar/11/using-llms-for-code/#s... > Don’t fall into the trap of anthropomorphizing LLMs and assuming that failures which would discredit a human should discredit the machine in the same way.
- jpc0 1y ago> And yeah, their confidence is frustrating. Treat them like an over-confident twenty-something intern who doesn't like to admit when they get stuff wrong. I was explicitly calling out this comment, that intern would get fired if when explicitly called out they not only don’t want to admit they are wrong but vehemently disagree. The interaction was “Implement X”, it gave an implementation, I responded “function y does not exist use a different method”, it instead of following that instruction gave me a link to the documentation for the library that it claim’s contains that function and told me I am wrong. I said the documentation it linked does not contain that function and to do something different and yet it still refused to follow instructions and pushed back. At that point I “fired” it and wrote the code myself.
- jpc0 1y ago> Or the simpler version of the above: LLM writes code. You try and compile it. You get an error message and you paste that back into the LLM and let it have another go. That's been the main way I've worked with LLMs for almost three years now. I’m going to comment here about this but it’s a follow on to the other comment, this is exactly the workflow I was following. I had given it the compiler error and it blamed an environment issue, I confirmed the environment is as it claims it should be, it linked to documentation that doesn’t state what it claims is stated. In a coding agent this would have been an endless feedback loop that eats millions of tokens. This is the reason why I do not use coding agents, I can catch hallucinations and stop the feedback loop from ever happening in the first place without needing to watch an AI agent try to convince itself that it is correct and the compiler must be wrong.
- elliotto 1y agoYou're responding to simonw, the guy who could be considered the single leading voice in practical applications of llms, with an anecdote about how one time the bot gave you a compiler error, and llm coding is therefore useless.
- jpc0 1y agoI would agree with you if it was just one time, this is multiple times. Also look up appeal to authority, everyone can be better, I respect Simon don’t get me wrong, but I don’t read names before I comment and you would probably do well to do the same, let what people say stand on their own.
- elliotto 1y agoI read what you wrote and let it stand on its own. Your comment mentioned a single time the bot had failed - there was no mention of it occurring multiple times. Maybe you should flesh out your anecdotes properly. If you are treating anonymous posts on forums with the same authority as experts on those topics then it makes sense why you are struggling to use new tooling.