3 ms·
Surely the same can be said for the people saying the opposite?
by xquce 3mo ago
Surely the same can be said for the people saying the opposite?
- therobots927 3mo agoI didn’t make a claim. The parent explicitly said it was a misconception that inference is not profitable. No one knows if it’s profitable or not so we’re left to speculate.
- Bnjoroge 3mo agoYou can make a fairly decent assumption by calculating the margin on serving glm 5.2, and adding say 30% extra costs and it still leaves a healthy margin
- therobots927 3mo agoWhere are you getting 30% from
- Bnjoroge 3mo agoIt was a rough heuristic for how much more opus/5.5 would presumably cost if you extrapolate from glm5.2 prices. In any case input tokens for 5.4/4.6 are 70-100% more expensive, cached about 3-100%, and output tokens anywhere from 60%-240% more as per all their current api pricing. I highly doubt 5.4/4.6 are that much more expensive to serve given how cheap and commoditized inference has become, and how comparable they are perfomance wise.