3 ms·
There’s has a ton of work on character or byte level encodings for llms. The problem is you expand your input tokens by 3-4x. Expensive. Also, you still need t
by p1esk 1mo ago
There’s has a ton of work on character or byte level encodings for llms. The problem is you expand your input tokens by 3-4x. Expensive.
Also, you still need token embeddings (I think you might be confused how that works).
- amelius 1mo agoCould be! I have not (yet) spent much time learning about how llms work, just the occasional blog here and there. My main question is why we _need_ a bit of additional code to massage the input into tokens and especially why the neural network cannot do it, i.e., let the embedding be a latent space that forms naturally when training the network. If that makes sense.
- butvacuum 1mo agoare you aware of n-grams?
- thatjoeoverthr 1mo ago“let the embedding be a latent space that forms naturally when training the network” The embeddings are produced in concert with the network, to serve the network, and not created as a separate step. It’s actually very cool The look-up table is a matrix. Each row is an embedding and each row number is a token ID. You get a differentiable transformation from token ID to token embedding using a “one hot vector” and a matrix multiplication If you take the transpose of this matrix, you can convert an internal representation back to the same token form, but treat it as logits and give it to the sampler. So token embeddings are produced on demand in service of the model, according to the model’s needs. I found an example of this strategy in a paper as far back as 1980! In the other reply I recommend the Bengio paper. But do bite the bullet and try it.
- amelius 1mo agoI'll have a look at that paper, thanks!