4 ms·
I had a brainwave recently. I was tired and looking at two XML documents which looked identical to me and I thought hey, let's see what ChatGPT thinks. So I ask
by jamiethompson 4y ago
I had a brainwave recently. I was tired and looking at two XML documents which looked identical to me and I thought hey, let's see what ChatGPT thinks. So I asked it to describe to me what the differences were between the two documents and it immediately just started talking about elements that it had completely made up. Every time I asked it why it was doing that it apologised but then doubled-down on making even more stuff up. Eventually I asked it to show me what its understanding was of the two documents I was asking it to compare and it showed me two completely unrelated XML documents
- ddalex 4y agoI do not understand why people expect chatgpt to reason when all it is is a fancy probabilistic language model.... I guess humans have high bias towards trusting confident-sounding language despite what reason would tell us to do. That's how politicians and advertisment work anyway....
- jamiethompson 4y agoI just thought it would be interesting, given that it has an understanding of XML to see if it could do a simple diff, "by eye" if you will. Obviously I wasn't intending to trust its output. We of course have long standing trusted tools for diffing files.
- nullc 4y agoUnless your xml input was very small you probably exceeded it's input window. I suspect chatgpt does some summarization under the covers but otherwise it still only has a finite and fairly small lookback.
- jamiethompson 4y agoI suspect you're right there. Yet it forged ahead confidently giving answers anyway!
- nullc 4y agoRight, but consider its 'evaluation' during training: During training it is constantly seeing stuff where the context is out of the window and the correct completion confidently answers, so the model is trained to do the same. I think this is very tricky to solve conceptually (since the human authors don't have the same input event horizon problem), but it could be (and has been) papered over by making the context bigger.