3 ms·
Thank you for this. I have a basic grasp of algebra (including quadratic functions such as in your example). So am I on the right track in extrapolating this
by herodoturtle 3y ago
Thank you for this.
I have a basic grasp of algebra (including quadratic functions such as in your example).
So am I on the right track in extrapolating this to say that LLM parameters are the adjustable coefficients and constants in the chain of functions?
- ps256 3y agoYes, and adjustable here means that they are learned from the training data. They start as random values and get repeatedly adjusted during the training process until the model exhibits useful behaviour.
- nextos 3y agoThe explanation you got above is sound but slightly misleading. Let me explain. Typical neurons in a neural network are actually simpler that a quadratic equation. They are actually a linear function y = ax + b, where a and b are vectors. So, y = a_1 * x_1 + a_2 * x_2 + ... a_n * x_n + b. Here, (a_1, a_2, ..., a_n, b) are trainable, i.e. adjustable, parameters. (x_1, x_2, ..., x_n) are the inputs to the neuron. y is its output. Because linear functions are inflexible, neurons typically pass the output y to a non-linear function. The classical non-linear function is a sigmoid, but there are many choices. So you have z = sigmoid(y) = sigmoid(ax + b) = sigmoid(a_1 * x_1 + a_2 * x_2 + ... a_n * x_n + b). These non-linear functions usually have no trainable parameters. Such a structure, a linear function composed with a non-linear function is the cornerstone of modern statistics, where it is called a generalized linear model (GLM). Hence, neural networks are just many GLMs or neurons grouped into layers, with layers connected to each other.