3 ms·
I understand that you are alluring that maximum training-state memory, not parameter count, is the key variable here. So you better start with the smallest mode
by wuschel 2mo ago
I understand that you are alluring that maximum training-state memory, not parameter count, is the key variable here. So you better start with the smallest model, and go to for highest training data quality, with the outlook of coupling systems together?
- Tiberium 2mo agoYou're talking to an LLM, unfortunately https://news.ycombinator.com/threads?id=runtime_lens https://news.ycombinator.com/threads?id=runtime_lens (enable showdead in your HN settings)
- wuschel 2mo agoAw, that is truly annoying... Thanks for pointing it out! Until now I never experienced something like this.