3 ms·
There will always be more data that could be relevant than fits in a context window, and especially for multi-turn conversations, huge contexts incur huge costs
by AaronFriel 3y ago
There will always be more data that could be relevant than fits in a context window, and especially for multi-turn conversations, huge contexts incur huge costs.
GPT-4 Turbo, using its full 128k context, costs around $1.28 per API call.
At that pricing, 1m tokens is $10, and 10m tokens is an eye-watering $100 per API call.
Of course prices will go down, but the price advantage of working with less will remain.
- 7734128 3y agoWould the price really increase linearly? Isn't the demands on compute and memory increasing steeper than that as a function of context length?
- elorant 3y agoI don't see a problem with this pricing. At 1m tokens you can upload the whole proceedings of a trial and ask it to draw an analysis. Paying $10 for that sounds like a steal.
- AaronFriel 3y agoOf course, if you get exactly the answer you want in the first reply.
- staticman2 3y agoWhile it's hard to say what's possible on the cutting edge, historically models tend to get dumber as the context size gets bigger. So you'd get a much more intelligent analysis of a 10,000 token excerpt of the trial than a million token complete transcript of the trial. I have not spent the money testing big token sizes in GPT 4 turbo, but it would not surprise me if it gets dumber. Think of it this way, if the model is limited to 3,000 token replies, if an analysis would require a more detailed response than 3,000 tokens, it cannot provide it, it'll just give you insufficient information. What it'll probably do is ignore parts of the trial transcript because it can't analyze all that information in 3,000 tokens. And asking a followup question is another million tokens.
- ithkuil 3y agoUnfortunately the whole context has to be reprocessed fully for each query, which means that if you "chat" with the model you'll incur in that $10 fee for every interaction which quickly sums up. It may still be worth it for some use cases