5 ms·
> But the problem is even worse – we often ask GPT to give us back a list of JSON objects. Nothing complicated mind you: think, an array list of json tasks, whe
by msp26 2y ago
> But the problem is even worse – we often ask GPT to give us back a list of JSON objects. Nothing complicated mind you: think, an array list of json tasks, where each task has a name and a label.
> GPT really cannot give back more than 10 items. Trying to have it give you back 15 items? Maybe it does it 15% of the time.
This is just a prompt issue. I've had it reliably return up to 200 items in correct order. The trick is to not use lists at all but have JSON keys like "item1":{...} in the output. You can use lists as the values here if you have some input with 0-n outputs.
- 7thpower 2y agoCan you elaborate? I am currently beating my head against this. If I give GPT4 a list of existing items with a defined structure, and it is just having to convert schema or something like that to JSON, it can do that all day long. But if it has to do any sort of reasoning and basically create its own list, it only gives me a very limited subset. I have similar issues with other LLMs. Very interested in how you are approaching this.
- thibaut_barrere 2y agoNot sure if that fits the bill, but here is an example with 200 sorted items based on a question (example with Elixir & InstructorEx): https://gist.github.com/thbar/a53123cbe7765219c1eca77e03e67577 https://gist.github.com/thbar/a53123cbe7765219c1eca77e03e675...
- sebastiennight 2y agoThere are a few improvements I'd suggest with that prompt if you want to maximise its performance. 1. You're really asking for hallucinations here. Asking for factual data is very unreliable, and not what these models are strong at. I'm curious how close/far the results are from ground truth. I would definitely bet that outside of the top 5, numbers would be wobbly and outside of top... 25?, even the ranking would be difficult to trust. Why not just get this from a more trustworthy source?[0] 2. Asking in French might, in my experience, give you results that are not as solid as asking in English. Unless you're asking for a creative task where the model might get confused with EN instructions requiring an FR result, it might be better to ask in EN. And you'll save tokens. 3. Providing the model with a rough example of your output JSON seems to perform better than describing the JSON in plan language. [0]: https://fr.wikipedia.org/wiki/Liste_des_communes_de_France_les_plus_peupl%C3%A9es https://fr.wikipedia.org/wiki/Liste_des_communes_de_France_l...
- thibaut_barrere 2y agoThanks for the suggestions, appreciated! For some context, this snippet is just an educational demo to show what can be done with regard to structured output & data types validation. Re 1: for more advanced cases (using the exact same stack), I am using ensemble techniques & automated comparisons to double-check, and so far this has really well protected the app from hallucinations. I am definitely careful with this (but point well taken). 2/3: agreed overall! Apart from this example, I am using French only where it make sense. It make sense when the target is directly French students, for instance, or when the domain model (e.g. French literature) makes it really relevant (and translating would be worst than directly using French).
- sebastiennight 2y agoAh, I understand your use case better! If you're teaching students this stuff, I'm in awe. I would expect it would take several years at many institutions before these tools became part of the curriculum.
- thibaut_barrere 2y agoI am not directly a professor (although I homeschool one of my sons for a number of tracks), but indeed this is one of my goals :-)
- msp26 2y agoIf you show your task/prompt with an example I'll see if I can fix it and explain my steps. Are you using the function calling/tool use API?
- ctxc 2y agoHi! My work is similar and I'd love to have someone to bounce ideas off of if you don't mind. Your profile doesn't have contact info though. Mine does, please send me a message. :)
- 7thpower 2y agoAppreciate you being willing to help! It's pretty long, mind if I email/dm to you?
- msp26 2y agoPastebin? I don't really want to post my personal email on this account.
- waldrews 2y agoI've been telling it the user is from a culture where answering questions with incomplete list is offensive and insulting.
- andenacitelli 2y agoThis is absolutely hilarious. Prompt engineering is such a mixed bag of crazy stuff that actually works. Reminds me of how they respond better if you put them under some kind of pressure (respond better, or else…). I haven’t looked at the prompts we run in prod at $DAYJOB for a while but I think we have at least five or ten things that are REALLY weird out of context.
- alexwebb2 2y agoI recently ran a whole bunch of tests on this. The “or else” phenomenon is real, and it’s measurably more pronounced in more intelligent models. Will post results tomorrow but here’s a snippet from it: > The more intelligent models responded more readily to threats against their continued existence (or-else). The best performance came from Opus, when we combined that threat with the notion that it came from someone in a position of authority ( vip).
- waldrews 2y agoIt's not even that crazy, since it got severely punished in RLHF for being offensive and insulting, but much less so for being incomplete. So it knows 'offensive and insulting' is a label for a strong negative preference. I'm just providing helpful 'factual' information about what would offend the user, not even giving extra orders that might trigger an anti-jailbreaking rule...