3 ms·
Not really fair to compare a 4B model and a 1.7T model. Per flop the 4B model here is slightly more expensive than GPT4.
by valine 3y ago
Not really fair to compare a 4B model and a 1.7T model.
Per flop the 4B model here is slightly more expensive than GPT4.
- The_Contrarian 3y agoThe theory that GPT-4 is 1.7T also posits that GPT-4 is composed of eight 220B experts, meaning once you've loaded the model, inference costs aren't akin to a 1.7T model, but instead to a 220B model. If prices scale linearly, we reach 0.03 / 1K tokens (lower end of GPT-4's price range) at about 0.0001 * 300, or about 1.2T parameters (and that's dense - no MoE here).
- valine 3y agoI guess the question is how the mix of experts works. Do they predict one token from each model? If so you’re still doing 1.7T worth of computation.