3 ms·
I would reconsider. Interestingly, another post on the front page right now "Numbers every LLM developer should know" brings up the fact that this efficiency is
by ImageXav 3y ago
I would reconsider. Interestingly, another post on the front page right now "Numbers every LLM developer should know" brings up the fact that this efficiency is due to the training corpus being in English. Any state actor with the will and enough funds could easily train a model for their own language. Most of the difficulty would lie in acquiring the expertise to do so properly, as LLM developers are in high demand right now and command their own salary.
- zmgsabst 3y agoThailand is currently training ThaiGPT for an open source release. https://www.bangkokpost.com/tech/2556324/nectec-agencies-roll-out-open-thaigpt https://www.bangkokpost.com/tech/2556324/nectec-agencies-rol... Whether that focuses in this manner is unclear, but we should expect forward looking governments to train LLMs to their own taste.
- marginalia_nu 3y agoWouldn't that hinge on the training data being available? There are many languages that are several orders of magnitude smaller than English in output. There's something like 6 million Danish-speakers for example.
- TeMPOraL 3y agoThey may offset for it by pulling content from books, newspapers, official documents, etc. - anything they can get in digital form, or digitize from analog. (I think this is how Google Translate went about things in the past, making translations into some languages come out very formal, as most of the training corpora for that languages came from internal and international official documents.) Countries that have been on-line for a while may also have discussion boards and comment-bearing sites that are entirely unknown to people outside those countries, too. Maybe multi-step approach would be in order - try to get half-decent a translation system working (an OG LLM, or an LLM trained to fix grammar in translations outsourced to GPT-4), and then synthesize training data for your main LLM by having GPT-4 (or its successor) generate tons of English text of all kind, and feeding it to the translator system. (There's a limit to synthesizing training data, beyond which it'll only amplify existing patterns, impacting model performance in bad ways - but I don't know how easy it is to reach it.)