35 ms·
“we attribute the performance of Yi models primarily to its data quality resulting from our data-engineering efforts” Data work is rarely sexy, but (almost) al
by jacobn 3y ago
“we attribute the performance of Yi models primarily to its data quality resulting from our data-engineering efforts”
Data work is rarely sexy, but (almost) always useful.
Did they release the corpus?
- gwern 3y agoThey did not, in part because it would reveal the data-filtering routines (particularly the political censorship - Chinese LLM papers sometimes mention the ban list but never reveal it), and also in part because it might reveal things they'd rather keep secret. For example, Bytedance has already been caught using the OA API to generate data for their models because they are having such a hard time catching up to OA - and evading bans for doing that, and also instructing employees on how to lie & cover it up: https://www.theverge.com/2023/12/15/24003151/bytedance-china-openai-microsoft-competitor-llm https://www.theverge.com/2023/12/15/24003151/bytedance-china... Do you think that a small Chinese startup like 01.AI, which by their own admission had to "bet the farm" to buy enough GPUs to train the Yi models at all https://www.bloomberg.com/news/articles/2023-11-05/kai-fu-lee-s-open-source-01-ai-bests-llama-2-according-to-hugging-face https://www.bloomberg.com/news/articles/2023-11-05/kai-fu-le... , and which were completely silent about cloning the American LLaMA architecture until people analyzed the released checkpoints and noticed it looked awfully familiar, is going to be above such tactics...? In this economic/geopolitical context? Especially when everyone seems to be doing it, not just Bytedance?* (01.AI claims that, the architecture aside, they didn't simply further train LLaMA models but trained from scratch. You can decide for yourself how much you are willing to believe this.) I wouldn't bet a lot of money on it, and that's why I don't expect to see any large comprehensive data releases from 01.AI for the Yi models. * This is one of my theories for why so many disparate models by so many different groups all seem to weirdly converge on the same failure modes like 'write a non-rhyming poem', and why GPT-3.5, and then GPT-4, seemed to be oddly difficult to surpass, as if there were some magnetic force which made reaching near 3.5/4 quality easy for 'independent' models, but then surpassing somehow difficult. Everyone is lying or mistaken about 3.5/4 data getting into their corpus, and the sugar-rush of imitation learning fools you into thinking you're making a lot of progress, even when your overall approach sucks. (As Andrej Karpathy notes, neural nets want to work, and so even if you have serious bugs in your code, they will still work pretty well - and simply permanently fall short of their true potential. Cautionary recent example: https://twitter.com/karpathy/status/1765473722985771335 https://twitter.com/karpathy/status/1765473722985771335 )
- visarga 3y ago> 01.AI claims that, the architecture aside, they didn't simply further train LLaMA models but trained from scratch. You can decide for yourself how much you are willing to believe this. You can't hide this. The latent space remains mostly fixed after pre-training. It all depends on the seed for the initial random init. Further pre-training won't move it enough. Because of this property, you can even average two fine-tunings from the same parent model, but never on models trained from different seeds.
- sroussey 3y agothe averaging seems like good test for who the parent is.
- gwern 3y agoI don't know anyone has properly analyzed this, nor how robust such methods are if one is trying to cover it up. Also, I doubt anyone has analyzed the scenario where the warm-started model is then extensively trained for trillions of token (possibly with a cyclical LR) particularly in Chinese - the latent spaces are not perfectly aligned Chinese/English and I'd expect that to change it a lot. (The point of this would be that 'cheating' by warm-starting it should let you avoid a lot of training instabilities and issues early in training, and may get you better quality at the end.)
- renonce 3y agoSounds like an interesting direction of research. There is a tool called mergekit and the models produced by merging different models is called frankenmerge. Search for “frankenmerge” and you can find a lot of interesting results and models and discussions on what works and what doesn’t. Might be a good idea to check previous experiments with these keywords.