3 ms·
That’s how increasing performance works. You make a model 10x faster, then you make it think 2x as much. Its cost is now 1/10th per token, and 1/5th per task.
by _3u10 25d ago
That’s how increasing performance works. You make a model 10x faster, then you make it think 2x as much.
Its cost is now 1/10th per token, and 1/5th per task.
Basically they have shitty hardware so they have to do a lot of optimization. Think of it like replacing an O(n) algorithm with O(log n).
Anthropic / Open AI think the best path is the most intelligent models deepseek is more focused on tok/$