3 ms·
vanilla neural networks from the 1950's look like def f(x): for _ in range(3): x = g(Wx + b) return x Essentially it is a matrix multiplicatio
by ericjang 5y ago
vanilla neural networks from the 1950's look like
def f(x):
for _ in range(3):
x = g(Wx + b)
return x
Essentially it is a matrix multiplication, a vector addition, a non-linearity.
Transformers are a modification to that architecture - using different multiplications, additions, and non-linearities. both of these are general in the sense that if you have enough of them, they can approximate any function. The ones used for transformers empirically do well on a lot of machine learning problems, particularly where data has a sequential nature.