3 ms·
Is training not possible? I did some stuff years ago where I create and train small NNs in the browser and I'm curious if that type of thing would work better t
by Tenoke 2y ago
Is training not possible? I did some stuff years ago where I create and train small NNs in the browser and I'm curious if that type of thing would work better today with a small custom transformer.
- spmurrayzzz 2y agoIn theory its definitely possible, but I suspect that maybe performance concerns are probably the reason its not implemented (yet). They have a webgpu embeddings benchmark in an HF space to give you a sense of the forward pass dynamics: https://huggingface.co/spaces/Xenova/webgpu-embedding-benchmark https://huggingface.co/spaces/Xenova/webgpu-embedding-benchm... Its impressive for what it is, but training would be painful at those latencies (fp16, batch 32, sequence length 512 generates a ~500ms forward pass with a 22M param model)
- dheera 2y agoThere might be applications for much smaller transformers in UI design. Like for example - did the user tap the wrong location on the screen because their device was physically jolted, and can you correct for that, considering you have access to accelerometer in HTML5 - does the user keep repeating an action (checking every box in a list of e-mails) and can you extrapolate the rest of what the user wants to do - did the user bounce because you popped up a stupid intercom box or newsletter popup, and did you learn anything about what you need to do if you want to retain this particular user in the future these kinds of things could be done with hundreds or thousands of parameters or less
- spmurrayzzz 2y agoYea definitely. But in that case, you could train _much_ faster in pytorch, then convert to ONNX, and load in the browser for inference (as the transformers.js docs recommend) EDIT: (I responded before your full edit with the bullet list). This next comment is orthogonal to the slow training performance topic I think, but the use cases you reference there don't seem to be well-suited at all for an autoregressive decoder-only model architecture.