12 ms·
The way neural networks work is that the base neural network is embedded in a sampling loop, i.e. a query is fed into the network & the driver samples output to
by measurablefunc 7mo ago
The way neural networks work is that the base neural network is embedded in a sampling loop, i.e. a query is fed into the network & the driver samples output tokens to append to the query so that it can be re-fed back into the network (q → nn → [a, b, c, ...] → q + sample([a, b, c, ...])). There is no way to avoid hallucinations b/c hallucinations are how the entire network works at the implementation level. The precision makes no difference b/c the arithmetic operations are semantically void & only become meaningful after they are interpreted by someone who knows to associated 1 /w red, 2 w/ blue, 3 w/ clouds, & so on & so forth. The mapping between the numbers & concepts does not exist in the arithmetic.
- kelseyfrog 7mo agoOh, I thought that the embedding space of the residual stream was precisely that.
- measurablefunc 7mo agoThe arithmetic is meaningless, it doesn't matter what you call it b/c on the computer it's all bit strings & boolean arithmetic. You can call some sequence of operations residual & others embeddings but that is all imposed top-down. There is nothing in the arithmetic that indicates it is somehow special & corresponds to embeddings or residuals.
- kelseyfrog 7mo agoAh ok, so if we had such a mapping then models wouldn't hallucinate?
- measurablefunc 7mo agoMaybe it's better if you define the terms b/c what I mean by hallucination is that the arithmetic operations + sampling mean that it's all hallucinations. The output is a trajectory of a probabilistic computation over some set of symbols (0s & 1s). Those symbols are meaningless, the only reason they have meaning is b/c everyone has agreed that the number 97 is the ascii code for "a" & every conformant text processor w/ a conformant video adapter will convert 97 (0b1100001) into the display pattern for the letter "a".
- kelseyfrog 7mo agoSo kind of like if you flip a coin, the sampling means the heads or tails you get isn't real?
- measurablefunc 7mo agoIt's when you define heads or tails however you want & then tell me you have objective semantics for each side of the coin when all you've really done is established a convention about which side is which. The coin is real, what you call each side is a convention & what semantics you attach to a sequence of flips is also a convention that has nothing to do with the reality of the coin.
- kelseyfrog 7mo agoI'm struggling to differentiate that from how we use coinflips normally. We can pretty easily create arbitrary mappings and then sample from the binomial in a way that has meaning far beyond just heads or tails. Maybe I'm not quite understanding.
- measurablefunc 7mo agoWhich part are you confused about? Symbols are meaningless until someone imposes semantics on them. There is nothing meaningful about arithmetic in a neural network other than whatever conventions are imposed on the binary sequences, same way 97 has no meaning other than the conventional agreement that it is the ascii code point for "a".
- kelseyfrog 7mo agoI guess I don't get the main idea. Chemical reactions in our brains are semantically void and yet we're able to use it as substrate for thinking.
- measurablefunc 7mo ago