4 ms·
Yep, it would definitely be difficult to justify running it in production. Accuracy would need to be higher as you said or it would need to be applicable to mor
by gradys 6y ago
Yep, it would definitely be difficult to justify running it in production. Accuracy would need to be higher as you said or it would need to be applicable to more tasks such that you can take other models out of production.
This kind of model could be used as the teacher in a distillation setup too though. Then faster training of the teacher is actually a huge benefit since it speeds up model development iteration cycles.
But even if it weren't practical to use in production in any sense, I'd argue there's value in doing the basic research of exploring design space of architectures in this way. This came out of a research team at Google. It may inspire and inform smaller, more practical architectures.
- gwern 6y ago"Yep, it would definitely be difficult to justify running it in production. Accuracy would need to be higher as you said or it would need to be applicable to more tasks such that you can take other models out of production." Part of the justification is the MoE sparsity by design means only a small part of the model will be activated by a given query. Don't think of it as a single giant 1t model, think of it as 50 small models which happen to share a glue layer at the input. So at deployment, you could, for example, keep the gating-layer in RAM and only pull the necessary sub-model off disk as necessary. Or you could shard the sub-models over 51 GPUs and feed the master gating-layer+GPU lots of queries, and each query will be dispatched to a different expert+GPU pair. This could easily be competitive with running a lot of dense models in parallel trying to keep up with the same load.