3 ms·
>How exactly the large GPT-2 models are deployed is a mystery I really wish was open-sourced more. TalkToTransformer.com uses preemptible P4 GPUs on Google Kub
by AdamDKing 7y ago
>How exactly the large GPT-2 models are deployed is a mystery I really wish was open-sourced more.
TalkToTransformer.com uses preemptible P4 GPUs on Google Kubernetes Engine. Changing the number of workers and automatically restarting them when they're preempted is easy with Kubernetes.
To provide outputs incrementally rather than waiting for the entire sequence to be generated, I open a websocket to a a worker and have it do a few tokens at a time, sending the output back as it goes. GPT-2 tokens can end partway through a multi-byte character, so to make this work you need to send the raw UTF-8 bytes to the browser and then have it concatenate them _before_ decoding the string.
While my workers can batch requests from multiple users, the modest increase in performance is probably not worth the complexity in most cases.
- jcims 7y agoAny thoughts on the larger model? Doesn't seem materially better than the last one. Maybe the fine tuning exercises will show the benefit?