5 ms·
I haven't read a lot of LLM papers, but I believe this is a rather weak paper low on details (note: not the results achieved of the LLM, but the paper itself).
by cgeier 3y ago
I haven't read a lot of LLM papers, but I believe this is a rather weak paper low on details (note: not the results achieved of the LLM, but the paper itself). If it had landed on my desk for a review, I probably would have sent it back just based on that.
For example, they never really say how they trained the experts or which dataset they used.
Is this the current standard in the field?
- ShamelessC 3y ago> Is this the current standard in the field? It’s becoming pretty common, yeah. The two things you mentioned: training particulars and dataset mixture are also basically the only competitive advantage companies have. Since the code/architecture is trivial to reproduce, anyone with enough money can make a competing model “easily”. OpenAI started this trend and cemented it with GPT4’s “technical report” which didn’t even specify the number of parameters in the model. They’ve been historically vague about their dataset for far longer than that though.
- MichaelRazum 3y agoExactly, same thought. Actually I would expect, that they trained each expert separately and later together, since you need to train the router network as well. I'm far from an expert in LLMs. But this would be interesting to know, especially how different training setups influence the performance.