3 ms·
The MACAW model is quite impressive -- it significantly outperforms GTP3 on Q&A/reasoning tasks (~10%) while requiring x10 less parameters. It is based on T5 ,
by danielmorozoff 5y ago
The MACAW model is quite impressive -- it significantly outperforms GTP3 on Q&A/reasoning tasks (~10%) while requiring x10 less parameters. It is based on T5 , which is a well known model from Google. The novel innovation here is the training paradigm.
Also did I mention it's OSS:
https://macaw.apps.allenai.org/ https://macaw.apps.allenai.org/
Here is the paper:
https://arxiv.org/abs/2109.02593 https://arxiv.org/abs/2109.02593