3 ms·
With ChatGPT function calling I get valid JSON 100% of the time from GPT-4 unless I have made some error in prompting. The chief error is not providing escape
by caesil 3y ago
With ChatGPT function calling I get valid JSON 100% of the time from GPT-4 unless I have made some error in prompting.
The chief error is not providing escape hatches. LLMs look for a right answer. If you are feeding it some texts and asking it to return structured data about the texts, but then one of the texts is blank, it will be difficult to determine a right answer, so you get hallucinations. The solution is an escape hatch where one of the arguments is a `textIsMissing` boolean or something.
As long as you've accounted for these failure modes, it works flawlessly.
- reissbaker 3y agoGPT-4 is amazing, but the upside of smaller models is much lower cost. I get basically 100% accuracy on JSON modeling with GPT-4 with function calling too, but I will say that gpt-3.5-turbo with function calling is somewhat less accurate — it usually generates valid JSON in terms of JSON.parse not exploding, but not necessarily JSON following the schema I passed in (although it's surprisingly good, maybe ~90% accurate?). I use 3.5-turbo a decent amount in API calls because it's just a lot cheaper, and performs well enough even if it's not gpt-4 level. I haven't gotten a chance to earnestly use the smaller Llama models yet in more than small prototypes (although I'm building a 4090-based system to learn more about finetuning them), but the little amount of experimenting I've done with them makes me think they need a decent amount of help with generating consistently-valid JSON matching some schema out of the box. This is a pretty neat tool to use for them, since it doesn't require finetuning runs, it just masks logits.
- BoorishBears 3y agoclaude-1.2-instant came out last week and is doing extremely well at following schemas. I'd say it's reached 3.5 turbo with the format following skills of GPT-4, which is powerful once you give it chain-of-thought
- selcuka 3y agoThe premise of function calling is great, but in my experience (at least on GPT-3.5, haven't tried it with GPT-4 yet) it seems to generate wildly different, and less useful results, for the same prompt.
- caesil 3y agoGPT-3.5 is pretty much useless for reliable NLP work unless you give it a VERY proscribed task. That's really the major breakthrough of GPT-4, in my mind, and the reason we are absolutely going to see an explosion of AI-boosted productivity over the next few years, even if foundation LLM advancements stopped cold right now. A vast ocean of mundane white collar work is waiting to be automated.
- ipaddr 3y agoYou can change the randomness value to 0 and get the same output each time for the same text
- selcuka 3y agoI should probably re-test it, but I think it wasn't the temperature. The results were unusually useless.
- tomduncalf 3y agoIn my experience (with GPT-4 at least), a temperature of 0 does not result in deterministic output. It's more consistent but outputs do still vary for the same input. I feel like temperature is a bit more like "how creative should the model be?"
- selcuka 3y agoOne theory is it is caused by its Sparse MoE (Mixture of Experts) architecture [1]: > The GPT-4 API is hosted with a backend that does batched inference. Although some of the randomness may be explained by other factors, the vast majority of non-determinism in the API is explainable by its Sparse MoE architecture failing to enforce per-sequence determinism. [1] https://152334h.github.io/blog/non-determinism-in-gpt-4/ https://152334h.github.io/blog/non-determinism-in-gpt-4/