5 ms·
I thought the examples I was thinking of were in the original GPT-4 Technical Report, but all I found on re-reading were examples of it explaining "what's funny
by tel 3y ago
I thought the examples I was thinking of were in the original GPT-4 Technical Report, but all I found on re-reading were examples of it explaining "what's funny about" a given image. Which is still a decent example of this, I think. GPT-4 demonstrates a semantic model about what entails humor.
- mjburgess 3y agoit entails only that the associative model is coincidentally indinstiguishable from a semantic one in the cases where it's used it is always trivial to take one of these models and expose it's failure to operate semantically, but these cases are never in the marketing material. Consider an associative model of addition, all numbers from -1bn to 1bn, broken down into their digits, so that 1bn = <1, 0, 0, 0, 0, 0, 0, 0, 0> Using such a model you can get the right answers for more additions than just -1bn to 1bn, but you can also easily find cases where the addition would fail. It's never adding.
- tel 3y agoI think part of what I suspect is going on here too is more computation and finiteness. It seems correct that LLM architectures cannot perform too much computation (unless you unroll it in the context). On the other hand you can look at statistical model identification in, say, nonlinear control. This can absolutely lead to unboundedly long predictions.