3 ms·
Thanks, that's a good data point. I'm wondering especially what the tangent for future developments will be. Most of the large language models are out of my le
by uniqueuid 4y ago
Thanks, that's a good data point.
I'm wondering especially what the tangent for future developments will be. Most of the large language models are out of my league anyways (e.g. the new yandex russian-english model was trained on 800 A100s and needs 200GB of GPU ram to fine tune).
So maybe it would be more effective to go for high speed instead of large capacity. But then you probably end up with a custom chassis and PSU since 4x 400 watts are not something that you can use on most off-the-shelf workstations.