Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
The_Contrarian
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
The_Contrarian
3y ago
The theory that GPT-4 is 1.7T also posits that GPT-4 is composed of eight 220B experts, meaning once you've loaded the model, inference costs aren't akin to a 1.7T model, but instead to a 220B model. If prices scale linearly, we r
2.
▲
by
The_Contrarian
3y ago
The main thing that led me to believe it's not trained directly on benchmarks is the fact that this model - a direct finetune of Mistral, improved on them. Consider the following thought experiment. Let's say we train a glorified