4 ms·
TTS models can be deployed on most devices (we do a lot of CPU work), but the real bottleneck for these on-device deployments is on the LLM side. Even a "small"
by neuphonic 10mo ago
TTS models can be deployed on most devices (we do a lot of CPU work), but the real bottleneck for these on-device deployments is on the LLM side. Even a "small" 700m parameter LLM was problematic. Also running 3 models in parallel on a single device is not super straightforward.