3 ms·
This is unfortunately non-trivial to quantitatively evaluate the performance against ChatGPT :-( We didn't do much evaluation because there isn't much innovati
by junrushao1994 3y ago
This is unfortunately non-trivial to quantitatively evaluate the performance against ChatGPT :-(
We didn't do much evaluation because there isn't much innovation on model side, but instead we are demoing the possibility of running an end-to-end model on ordinary client GPUs via WebGPU without server resources.
- amelius 3y ago> This is unfortunately non-trivial to quantitatively evaluate the performance against ChatGPT :-( Compare using the loss function?
- junrushao1994 3y agoIn LLM world, loss or perplexity may not be the best indicator of model performance :-( Perhaps HELM (https://crfm.stanford.edu/helm/latest/ https://crfm.stanford.edu/helm/latest/) but we didn't take deeper look as we are not the developers of this model
- imranq 3y agoThere are plenty of LLM benchmarks that are used to test performance, some of them are: * Winogrande * BoolQ * PIQA * SIQA * HellaSwag etc...
- junrushao1994 3y agoWould be nice if anyone could help us benchmark! Our primary focus though is not model performance, but to demonstrate the capability that TVM Unity generates code targeting WebGPU and allows them to run with client GPUs :-)