4 ms·
Agreed. IMO - A part of me even argue we should stop calling it attention (but what to call it instead is a mess) But since this was derived from apple lite at
by pico_creator 3y ago
Agreed. IMO - A part of me even argue we should stop calling it attention (but what to call it instead is a mess)
But since this was derived from apple lite attention paper. The name is gonna stick, due to a lack of better alternative
- xiphias2 3y agoI believe the naming is perfect: the word attention clearly maches the intention of the model structure. It’s still clearly missing more efficiency, but we had to have ChatGPT / ChatGPT4 working and used to make the further research directions clear (increasing context length and decreasing hallucinations).
- pico_creator 3y agoHaha, yea - naming things is hard Every-time someone comes up and say "this is not attention, because it does X and not Y" My response is, ok, what would you call it then? Because no one (including me) seem to to able to find a better term, that fits its use case.
- candiodari 3y agoNaming is a perpetual problem. My issue is with "hallucination", which everyone takes to be the "problem" with GPT style networks making things up. Never mind that transformers are just trying to predict the next likely token, NOT the truth PLUS that they're trained from the internet. As everyone knows, the internet is not known for correctness and truth. If you want any neural network to figure out the truth independently, it'll obviously need the ability to go out into the real world and even needs to be allowed to experiment for most things. Hallucination used to mean the following. A basic neural network is: f(x) = y = repeat(nonlinearity(ax[0] + bx[1] + ...)) And then you adjust a, b, c, ... until y is reasonable, according to the cost function. But look! The very same backpropagation can adjust x[0], x[1] ... with the same cost function and only a small change in the code. This allows you to reverse the question neural networks answer. Which can be an incredibly powerful way to answer questions. And that used to be called hallucination in Neural networks. Instead of "change these network weights to transform x into y, keeping x constant" you ask "change x to transform x into y, keeping the network weights constant". Now it's impossible finding half the papers on the this topic. AARGH!
- sundarurfriend 3y ago> My issue is with "hallucination", which everyone takes to be the "problem" with GPT style networks making things up. It's most famous as a problem with ChatGPT specifically, which is presented as a chat interface where an agent answers your question. In that context, it makes sense to think of confident and detailed wrong answers as hallucinations. You could say that it's still an LLM underneath, but then you'd be talking about a layer of abstraction beneath what most people interact with. Given that they tap into the mental model of "chat with an agent" heavily with their interface and interactions, having small disclaimers and saying "I'm just an LLM" from time to time aren't sufficient to counter people's intuitions and expectations.