3 ms·
Following the links, it might be using the Marian NMT framework under the hood? I'm wondering where specifically the model weights are coming from, and it would
by coder543 3y ago
Following the links, it might be using the Marian NMT framework under the hood? I'm wondering where specifically the model weights are coming from, and it would be interesting to read more about the training process[0].
I'm excited to see a more open translation system that seems to perform pretty well and runs locally.
It definitely needs to add support for more languages, and based on other models I've seen recently... I have to wonder if building dedicated models for each pair of languages is still the best choice. I believe SeamlessM4T just uses a single model (available in different sizes), and I have definitely seen that Whisper only uses a single model for multiple language pairs as well (although it was only specifically trained to translate into English). Similarly, virtually all of the LLMs are multilingual. It seems like a single model is able to learn shared insights that apply across languages, reducing the amount of training data needed for each additional language (or increasing the accuracy with the same amount of data), but I admit that I could be wrong. This has just been my perception of how things are developing.
[0]: Some training info here, it looks like: https://github.com/browsermt/students/tree/master/train-student https://github.com/browsermt/students/tree/master/train-stud...