4 ms·
There is another big change in gpt-4o-2024-08-06: It supports 16k output tokens compared to 4k before. I think it was only available in beta before. So gpt-4o-2
by __jl__ 2y ago
There is another big change in gpt-4o-2024-08-06: It supports 16k output tokens compared to 4k before. I think it was only available in beta before. So gpt-4o-2024-08-06 actually brings three changes. Pretty significant for API users
1. Reliable structured outputs
2. Reduced costs by 50% for input, 33% for output
3. Up to 16k output tokens compared to 4k
https://platform.openai.com/docs/models/gpt-4o https://platform.openai.com/docs/models/gpt-4o
- Culonavirus 2y agoThat's actually pretty impressive... if they didn't dumb it down that is, which only time will tell.
- santiagobasulto 2y agoI’ve noticed that lately GPT has gotten more and more verbose. I’m wondering if it’s a subtle way to “raise prices”, as the average response is going to incur I more tokens, which makes any API conversation to keep growing in tokens of course (each IN message concatenates the previous OUT messages).
- sashank_1509 2y agothey also spend more to generate more tokens. The more obvious reason is it seems like people rate responses better the longer they are. Lmsys demonstrated that GPT tops the leaderboard because it tends to give much longer and more detailed answers, and it seems like OpenAI is optimizing or trying to maximize lmsys.
- maeil 2y agoAgree with this take, though in an even broader way; they're optimizing for the leaderboards and benchmarks in general. Longer outputs lead to better scores on those. Even in this thread I see a lot of comments bring them up, so it works for marketing. My take is that the leaderboards and benchmarks are still very flawed if you're using LLMs for any non-chat purpose. In the product I'm building, I have to use all of the big 4 models (GPT, Claude, Llama, Gemini), because for each of them there is at least one tasks that it performs much better than the other 3.
- throwaway48540 2y agoIt's a subtle way to make it smarter. Making it write out the "thinking process" and decisions has always helped with reliability and quality.
- tedsanders 2y agoGPT has indeed been getting more verbose, but revenue has zero bearing on that decision. There's always a tradeoff here, and we do our imperfect best to pick a default that makes the most people happy. I suspect the reason why most big LLMs have ended up in a pretty verbose spot is that it's easier for users to scroll & skim than to ask follow-up questions (which requires formulation + typing + waiting for a response). With regard to this new gpt-4o model: you'll find it actually bucks the recent trend and is less verbose than its predecessor.
- zamadatix 2y agoDo changes in verbosity tuning have a meaningful impact on the average "correctness" of the responses? Also your about page is very suspicious for someone at an AI company ;).
- OJFord 2y ago> I suspect the reason why most big LLMs have ended up in a pretty verbose spot is that it's easier for users to scroll & skim than to ask follow-up questions Maybe it's a 'technical' user divide, but that seems wrong to me. I would much rather a succinct answer that I can probe further or clarify if necessary. Lately it's going against my custom prompt/profile whatever it's called - to tell it to assume some level of competence, a bit about my background etc., to keep it brief - and it's worse than it was when I created that out of annoyance with it. Like earlier I asked something about some detail of AWS networking and using reachability analyser with VPC endpoints/peering connections/Lambda or something, and it starts waffling on like 'first, establish the ID of your Virtual Private Cloud Endpoint. Step 1. To locate the ID, go to ...'
- zarzavat 2y agoThere’s an interesting discrepancy here. Human users are charged by the number of messages, so longer responses are preferable because follow up questions use up your message allowance. APIs are charged by token so shorter messages are preferable as you don’t pay for unnecessary tokens.
- sophiabits 2y agoI’ve especially noticed this with gpt-4o-mini [1], and it’s a big problem. My particular use case involves keeping a running summary of a conversation between a user and the LLM, and 4o-mini has a really bad tendency of inventing details in order to hit the desired summary word limit. I didn’t see this with 4o or earlier models Fwiw my subjective experience has been that non-technical stakeholders tend to be more impressed with / agreeable to longer AI outputs, regardless of underlying quality. I have lost count of the number of times I’ve been asked to make outputs longer. Maybe this is just OpenAI responding to what users want? [1] https://sophiabits.com/blog/new-llms-arent-always-better#examining-gpt-4o-mini https://sophiabits.com/blog/new-llms-arent-always-better#exa...
- atlex2 2y agoDid you try giving the model an "out"? > You may output only up to 500 words, if the best summary is less than 500 words, that's totally fine. If details are unclear, do not fill-in gaps, do leave them out of the summary instead.
- bilater 2y agoI have not been able to get it to output anywhere close to the max though (even setting max tokens high). Are there any hacks to use to coax the model to produce longer outputs?