3 ms·
The more general the model, the longer the lifetime. And the most impactful models today are incredibly general. For things like Whisper, I wouldn't be surprise
by conjecTech 3y ago
The more general the model, the longer the lifetime. And the most impactful models today are incredibly general. For things like Whisper, I wouldn't be surprised if we're already at 100:1 ratio for compute spent on inference vs training. BERT and related models are probably an order of magnitude or two above that. Training infra may be a bottleneck now, but it's unclear how long it will be until improvements slow and inference becomes even more dominant.
Capital outlays are tied to the derivative of compute capacity, so even if training just flatlines, hardware spend will drop significantly.
- flangola7 3y agoIsn't whisper self hosted
- conjecTech 3y agoThat's part of my point. There are 100s of organizations using it at scale, but it only needed to be trained once.