3 ms·
I think the point is that LOC is not a terribly useful metric in that everything is 1 LOC at the highest level of abstraction. The business proposition here is
by bubblethink 3y ago
I think the point is that LOC is not a terribly useful metric in that everything is 1 LOC at the highest level of abstraction. The business proposition here is that you don't need to write the LOCs for the underlying layers, they do it. The pitch here is that it's not as straightforward for large GPU clusters.
- Paul-Craft 3y agoSure, I get that. I've definitely seen demos of "Do X in Y LoC" that do X but offload all of the hard work of Y to some libray. This is not that. This is intended to be a demo that shows you what you can do with one Cerebras module. And, the result is that, by writing 565 LoC yourself, you can train and run an LLM the size of GPT-3. In that sense, 565 LoC is a perfectly fair number. It doesn't count PyTorch, numpy, the Python interpreter, or any of the library modules that are imported, but I don't think anyone was touting it as anything more; for instance, Mo Gawdat has said that GPT-4 is probably ~4500 LoC. And, yes, that certainly involves much more infrastructure, and doing that dance of going from GPU to CPU to a completely other node, etc.
- andy99 3y agoJust to add on, the project parallels Nano-gpt that itself touts it's small loc. It also uses pytorch etc. In both cases, the actual model logic is in the quoted lines of code. So the comparison is apt for what it's recreating. (I don't know how fair the comparison to other loc figures mentioned is). https://github.com/karpathy/nanoGPT https://github.com/karpathy/nanoGPT