2 ms·
> We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces. This is not what's meant by the
by naasking 2mo ago
> We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.
This is not what's meant by the statements that we don't know how LLMs work. Explain why LLMs are so good at programming, finding bugs, and developing mathematical proofs. Like, way better than all prior tools specifically designed to be bug finding tools, despite being merely "language models".