18 ms·
Native JSON Output from GPT-4
- thorum 3y agoThe JSON schema not counting toward token usage is huge, that will really help reduce costs.
- minimaxir 3y agoThat is up in the air and needs more testing. Field descriptions, for example, are important but extraneous input that would be tokenized and count in the costs. At the least for ChatGPT, input token costs were cut by 25% so it evens out.
- yonom 3y agoI believe functions do count in some way toward the token usage; but it seems to be in a more efficient way than pasting raw JSON schemas into the prompt. Nevertheless, the token usage seems to be far lower than previous alternatives, which is awesome!
- stavros 3y ago> Under the hood, functions are injected into the system message in a syntax the model has been trained on. This means functions count against the model's context limit and are billed as input tokens. If running into context limits, we suggest limiting the number of functions or the length of documentation you provide for function parameters.
- blamy 3y agoBut it does count toward token usage. And they picked JSON schema which is like 6x more verbose than typescript for defining the shape of json.
- minimaxir 3y agoAfter reading the docs for the new ChatGPT function calling yesterday, it's structured and/or typed data for GPT input or output that's the key feature of these new models. The ReAct flow of tool selection that it provides is secondary. As this post notes, you don't even need to the full flow of passing a function result back to the model: getting structured data from ChatGPT in itself has a lot of fun and practical use cases. You could coax previous versions of ChatGPT to "output results as JSON" with a system prompt but in practice results are mixed, although even with this finetuned model the docs warn that there still could be parsing errors. OpenAI's demo for function calling is not a Hello World, to put it mildly: https://github.com/openai/openai-cookbook/blob/main/examples/How_to_call_functions_with_chat_models.ipynb https://github.com/openai/openai-cookbook/blob/main/examples...
- tornato7 3y agoIIRC, there's a way to "force" LLMs to output proper JSON by adding some logic to the top token selection. I.e. in the randomness function (which OpenAI calls temperature) you'd never choose a next token that results in broken JSON. The only reason it wouldn't would be if the output exceeds the token limit. I wonder if OpenAI is doing something like this.
- senko 3y agoIt would seem not, as the official documentation mentions the arguments may be hallucinated or be a malformed JSON. (except if the meaning is the JSON syntax is valid but may not conform to the schema, but they're unclear on that).
- sanxiyn 3y agoFor various reasons, token selection may be implemented as upweighting/downweighting instead of outright ban of invalid tokens. (Maybe it helps training?) Then the model could generate malformed JSON. I think it is premature to infer from "can generate malformed JSON" that OpenAI is not using token selection restriction.
- sanxiyn 3y agoNote that this (token selection restriction) is even available on OpenAI API as logit_bias.
- newhouseb 3y agoBut only for the whole generation. So if you want to constrain things one token at a time (as you would to force things to follow a grammar) you have to make fresh calls and only request one token which makes things more or less impractical if you want true guarantees. A few months ago I built this anyway to suss out how much more expensive it was [1] [1] https://github.com/newhouseb/clownfish#so-how-do-i-use-this-with-gpt4 https://github.com/newhouseb/clownfish#so-how-do-i-use-this-...
- 3y ago
- jamesmcintyre 3y agoIn the openai blog post they mention "Convert “Who are my top ten customers this month?” to an internal API call" but I'm assuming they mean gpt will respond with structured json (we define via schema in function prompt) that we can use to more easily programatically make that api call? I could be confused but I'm interpreting this function calling as "a way to define structured input and selection of function and then structured output" but not the actual ability to send it arbitrary code to execute. Still amazing, just wanting to see if I'm wrong on this.
- williamcotton 3y agoThis does not execute code!
- jamesmcintyre 3y agoOk, yea this makes sense. Also for others curious of the flow here's a video walkthrough I just skimmed through: https://www.youtube.com/watch?v=91VVM6MNVlk https://www.youtube.com/watch?v=91VVM6MNVlk
- mritchie712 3y agoGlad we didn't get to far into adopting something like Guardrails. This sort of kills it's main value prop for OpenAI. https://shreyar.github.io/guardrails/ https://shreyar.github.io/guardrails/
- swyx 3y agoi mean only at the most superficial level. she has a ton of other validators that arent superceded (eg SQL is validated by branching the database - we discussed on our pod https://www.latent.space/p/guaranteed-quality-and-structure https://www.latent.space/p/guaranteed-quality-and-structure)
- mritchie712 3y agoyeah, listened to the pod (that's how I found out about guardrails!). fair point, I should have said: "value prop for our use case"... the thing I was most interested in was how well Guardrails structured output.
- swyx 3y agohaha excellent. i was quite impressed by her and the vision for guardrails. thanks for listening!
- Blahah 3y agoLuckily it's for LLMs, not openai
- blamy 3y agoGuardrails is an awesome project and will continue to be even after this.
- Kiro 3y agoCan I use this to make it reliably output code (say JavaScript)? I haven't managed to do it with just prompt engineering as it will still add explanations, apologies and do other unwanted things like splitting the code into two files as markdown.
- williamcotton 3y agoHere’s an approach to return just JavaScript: https://github.com/williamcotton/transynthetical-engine https://github.com/williamcotton/transynthetical-engine The key is the addition of few-shot exemplars.
- minimaxir 3y agoHere's a demo of some system prompt engineering which resulted in better results for the older ChatGPT: https://github.com/minimaxir/simpleaichat/blob/main/examples/notebooks/simpleaichat_coding.ipynb https://github.com/minimaxir/simpleaichat/blob/main/examples... Coincidentially, the new gpt-3.5-turbo-0613 model also has better system prompt guidance: for the demo above and some further prompt tweaking, it's possible to get ChatGPT to output code super reliably.
- sanxiyn 3y agoNot this, but using the token selection restriction approach, you can let LLM produce output that conforms to arbitrary formal grammar completely reliably. JavaScript, Python, whatever.
- chaxor 3y agoIs there a decent way of converting to a structure with a very constrained vocabulary? For example, given some input text, converting it to something like {"OID-189": "QQID-378", "OID-478":"QQID-678"}. Where OID and QQID dictionaries can be e.g. millions of different items defined by a description. The rules for mapping could be essentially what looks closest in semantic space to the descriptions given in a dictionary. I know this should be able to be solvable by local LLMs and bert cosine similarity (it isn't exactly, but it's a start on the idea), but is there a way to do this with decoder models rather than encoder models with other logic?
- jiggawatts 3y agoYou can train custom GPT 3 models, and Azure now has vector database integration for GPT-based models in the cloud. You can feed it the data, and ask it for the embedding lookup, etc... You can also host a vector database yourself and fill it up with the embeddings from the OpenAI GPT 3 API.
- chaxor 3y agoUnfortunately this doesn't really work, as the model is not limited in it's decoding vocabulary. Does anyone have other suggestions that may work in this space?
- swyx 3y agoi think people are underestimating the potential here for agents building - it is now a lot easier for GPT4 to call other models, or itself. while i was taking notes for our emergency pod yesterday (https://www.latent.space/p/function-agents https://www.latent.space/p/function-agents) we had this interesting debate with Simon Willison on just how many functions will be supplied to this API. Simon thinks it will be "deep" rather than "wide" - eg a few functions that do many things, rather than many functions that do few things. I think i agree. you can now trivially make GPT4 decide whether to call itself again, or to proceed to the next stage. it feels like the first XOR circuit from which we can compose a "transistor", from which we can compose a new kind of CPU.
- minimaxir 3y ago"Trivial" is misleading. From OpenAI's docs and demos, the full ReAct workflow is an order of magnitude more difficult than typical ChatGPT API usage with a new set of constaints (e.g. schema definitions) Even OpenAI's notebook demo has error handling workflows which was actually necessary since ChatGPT returned incorrect formatted output.
- cjonas 3y agoMaybe trivial isn't the right word, but it's still very straight-forward to get something basic, yet really powerful... ReAct Setup Prompt (goal + available actions) -> Agent "ReAction" -> Parse & Execute Action -> Send Action Response (success or error) -> Agent "ReAction" -> repeat As long as each action has proper validation and returns meaningful error messages, you don't need to even change the control flow. The agent will typically understand what went wrong, and attempt to correct it in the next "ReAction". I've been refactoring some agents to use "functions" and so far it seems to be a HUGE improvement in reliability vs the "Return JSON matching this format" approach. Most impactful is that fact that "3.5-turbo" will now reliability return JSON (before you'd be forced to use GPT-4 for an ReAct style agent of modest complexity). My agents also seem to be better at following other instructions now that the noise of the response format is gone (of course it's still there, but in a way it has been specifically trained on). This could also just be a result of the improvements to the system prompt though.
- darepublic 3y agoI have been using gpt4 to translate natural language to JSON already. And on v4 ( not v3) it hasn't returned any malformed JSON iirc
- nocsi 3y agoWhat if you ask it to include comments in the JSON explaining its choices
- yonom 3y ago- if the only reason you're using v4 over v3.5 is to generate JSON, you can now use this API and downgrade for faster and cheaper API calls. - malicious user input may break your json (by asking GPT to include comments around the JSON, as another user suggested); this may or may not be an issue (e. g. if one user can influence other users' experience)
- social_ism 3y ago[dead]
- adultSwim 3y agoRunning an LLM every time someone clicks on a button is expensive and slow in production, but probably still ~10x cheaper to produce than code.
- edwin 3y agoNew techniques like semantic caching will help. This is the modern era's version of building a performant social graph.
- daralthus 3y agoWhat's semantic caching?
- edwin 3y agoWith LLMs, the inputs are highly variable so exact match caching is generally less useful. Semantic caching groups similar inputs and returns relevant results accordingly. So {"dish":"spaghetti bolognese"} and {"dish":"spaghetti with meat sauce"} could return the same cached result.
- m3kw9 3y agoOr store as sentence embedding and calculate the vector distance, but creates many edge cases
- dang 3y agoRecent and related: Function calling and other API updates - https://news.ycombinator.com/item?id=36313348 https://news.ycombinator.com/item?id=36313348 - June 2023 (154 comments)
- emilsedgh 3y agoBuilding agents that use advanced API's was not really practical until now. Things like Langchain's Structured Agents worked somewhat reliably, but due to the massive token count it was so slow, the experience was _never_ going to be useful. Due to this, the performance in which our agent processes results has improved 5-6 times and it does actually do a pretty good job of keeping the schema. One problem that is not resolved yet is that it still hallucinates a lot of attributes. For example we have tool that allows it to create contacts in user's CRM. I ask it to: "Create contacts for top 3 Barcelona players:. It creates an structure like this" 1. Lionel Messi - Email: lionel.messi@barcelona.com - Phone Number: +1234567890 - Tags: Player, Barcelona 2. Gerard Pique - Email: gerard.pique@barcelona.com - Phone Number: +1234567891 - Tags: Player, Barcelona 3. Marc-Andre ter Stegen - Email: marc-terstegen@barcelona.com - Phone Number: +1234567892 - Tags: Player, Barcelona And you can see it hallucinated email addresses and phone numbers.
- 037 3y agoI would never rely on an LLM as a source of such information, just as I wouldn't trust the general knowledge of a human being used as a database. Does your workflow include a step for information search? With the new json features, it should be easy to instruct it to perform a search or directly feed it the right pages to parse.
- pluijzer 3y agoChatGPT can be usefully for many things, but you should really, not use it if you want to retrieve factual data. This might partly be resolved by querying the internet like bing does but purely on the language model side these hallucinations are just an unavoidable part of it.
- Spivak 3y agoYep, it's always always write code / query / function / whatever you need that you would parse and retrieve the data from an external system.
- 037 3y agoI'm wondering if introducing a system message like "convert the resulting json to yaml and return the yaml only" would adversely affect the optimization done for these models. The reason is that yaml uses significantly fewer tokens compared to json. For the output, where data type specification or adding comments may not be necessary, this could be beneficial. From my understanding, specifying functions in json now uses fewer tokens, but I believe the response still consumes the usual amount of tokens.
- gregw134 3y agoThat's what I'm doing. I ask ChatGPT to return inline yaml (no wasting tokens on line breaks), then I parse the yaml output into JSON once I receive it. A bit awkward but it cuts costs in half.
- lbeurerkellner 3y agoI think one should not underestimate the impact on downstream performance the output format can have. From a modelling perspective it is unclear whether asking/fine-tuning the model to generate JSON (or YAML) output is really lossless with respect to the raw reasoning powers of the model (e.g. it may perform worse on tasks when asked/trained to always respond in JSON). I am sure they ran tests on this internally, but I wonder what the concrete effects are, especially comparing different output formats like JSON, YAML, different function calling conventions and/or forms of tool discovery.
- edwin 3y agoFor those who want to test out the LLM as API idea, we are building a turnkey prompt to API product. Here's Simon's recipe maker deployed in a minute: https://preview.promptjoy.com/apis/1AgCy9 https://preview.promptjoy.com/apis/1AgCy9 . Public preview to make and test your own API: https://preview.promptjoy.com https://preview.promptjoy.com
- edelans 3y agoCongrats on the first-time user experience, I could experiment with your API in a few seconds, and the product is sleek!
- yonom 3y agoThis is cool! Are you using one-shot learning under the hood with a user provided example?
- edwin 3y agoThanks. We find few-shot learning to be more effective overall. So we are generating additional examples from the provided example.
- edwin 3y agoBTW: Here's a more performant version (fewer tokens) https://preview.promptjoy.com/apis/jNqCA2 https://preview.promptjoy.com/apis/jNqCA2 that uses a smaller example but will still generate pretty good results.
- sudb 3y agoThis is still pretty fast - impressive! Are there any tricks you're doing to speed things up?
- abhpro 3y agoThis is really cool, I had a similar idea but didn't build it. I was also thinking a user could take these different prompts (I called them tasks) that anyone could create, and then connect them together like a node graph or visual programming interface, with some Chat-GPT middleware that resolves the outputs to inputs.
- courseofaction 3y agoNice to have an endpoint which takes care of this. I've been doing this manually, it's a fairly simple process: * Add "Output your response in json format, with the fields 'x', which indicates 'x_explanation', 'z', which indicates 'z_explanation' (...)" etc. GPT-4 does this fairly reliably. * Validate the response, repeat if malformed. * Bam, you've got a json. I wonder if they've implemented this endpoint with validation and carefully crafted prompts on the base model, or if this is specifically fine-tuned.
- 037 3y agoIt appears to be fine-tuning: "These models have been fine-tuned to both detect when a function needs to be called (depending on the user’s input) and to respond with JSON that adheres to the function signature." https://openai.com/blog/function-calling-and-other-api-updates https://openai.com/blog/function-calling-and-other-api-updat...
- EGreg 3y agoActually I'm looking to take GPT-4 output and create file formats like keynote presentations, or pptx. Is that currently possible with some tools?
- yonom 3y agoI would recommend creating a simplified JSON schema for the slides (say, presentation is an array of slides, each slide has a title, body, optional image, optional diagram, each diagram is one of pie, table, ... Then use a library to generate the pptx file from the content generated.
- EGreg 3y agoLibrary? What library? It seems to me that a Transformer should excel at Transforming, say, text into pptx or pdf or HTML with CSS etc. Why don't they train it on that? So I don't have to sit there with manually written libraries. It can easily transform HTML to XML or text bullet points so why not the other formats?
- yonom 3y agoI don't think the name "Transformer" is meant in the sense of "transforming between file formats". My intuition is that LLMs tend to be good at things human brains are good at (e.g. reasoning), and bad at things human brains are bad at (e.g. math, writing pptx binary files from scratch, ...). Eventually, we might get LLMs that can open PowerPoint and quickly design the whole presentation using a virtual mouse and keyboard but we're not there yet.
- EGreg 3y agoIt’s just XML They can produce HTML and transform python into php etc. So why not? It’s easy for them no?
- stevenhuang 3y agoapparently pandoc also supports pptx so you can tell GPT4 to output markdown, then use pandoc to convert that markdown to pptx or pdf.
- irthomasthomas 3y agoIt's a shame they couldn't use yaml, instead. I compared them and yaml uses about 20% fewer tokens. However, I can understand accuracy, derived from frequency, being more important than token budget.
- IshKebab 3y agoI would imagine JSON is easier for a LLM to understand (and for humans!) because it doesn't rely on indentation and confusing syntax for lists, strings etc.
- AdrienBrault 3y agoI think YAML actually uses more tokens than JSON without indents, especially with deep data. For example "," being a single token makes JSON quite compact. You can compare JSON and YAML on https://platform.openai.com/tokenizer https://platform.openai.com/tokenizer
- nasir 3y agoIts a lot more straightforward to use JSON programmatically than YAML.
- golergka 3y agoIf you are using any kind of type checking instead of blindly trusting generated json it's exactly the same amount of work.
- TeMPOraL 3y agoIt really shouldn't be, though. I.e. not unless you're parsing or emitting it ad-hoc, for example by assuming that an expression like: "{" + $someKey + ":" + $someValue + "}" produces a valid JSON. It does - sometimes - and then it's indeed easier to work with. It'll also blow up in your face. Using JSON the right way - via a proper parser and serializer - should be identical to using YAML or any other equivalent format.
- riwsky 3y ago
- smallerfish 3y agoI will experiment with this at the weekend. Once thing I found useful with supplying a json schema in the prompt was that I could supply inline comments and tell it when to leave a field null, etc. I found that much more reliable than describing these nuances elsewhere in the prompt. Presumably I can't do this with functions, but maybe I'll be able to work around it in the prompt (particularly now that I have more room to play with.)
- wskish 3y agohere is code (with several examples) that takes it a couple steps further by validating the output json and pydantic model and providing feedback to the llm model when it gets either of those wrong: https://github.com/jiggy-ai/pydantic-chatcompletion/blob/master/pydantic_chatcompletion/__init__.py https://github.com/jiggy-ai/pydantic-chatcompletion/blob/mas...
- zyang 3y agoIs it possible to fine-tune with custom data to output JSON?
- edwin 3y agoThat's not the current OpenAI recipe. Their expectation is that your custom data will be retrieved via a function/plugin and then be subsequently processed by a chat model. Only the older completion models (davinci, curie, babbage, ada) are avaialble for fine-tuning.
- aecorredor 3y agoNewbie in machine learning here. It’s crazy that this is the top post just today. I’ve been doing the intro to deep learning course from MIT this week, mainly because I have a ton of JSON files that are already classified, and want to train a model that can generate new JSON data by taking classification tags as input. So naturally this post is exciting. My main unknown right now is figuring out which model to train my data on. An RNN, a GAN, a diffusion model?
- ilaksh 3y agoDid you read the article? To do it with OpenAI you would just put a few output examples in the prompt and then give it a function that takes the class and the output parameters correspond to the JSON format you want, or just a string containing JSON. You could also fine tuned an LLM like Falcon-7b but probably not necessary and nothing to do with OpenAI. You might also look into the OpenAI Embedding API as a third option. I would try the first option though.
- imranq 3y agoWouldnt this be possible with a solution like Guidance where you have a pre structured JSON format ready to go and all you need is text: https://github.com/microsoft/guidance https://github.com/microsoft/guidance
- m3kw9 3y agoIt works pretty good. You define a few “function” and enter a description on what it does, when user prompts, it will understand the prompt and tell you which likely “function” to use, which is just the function name. I feel like this is a new way to program, a sort of fuzzy logic type of programming
- Sai_ 3y ago> fuzzy logic Yes and no. While the choice of which function to call is dependent on an llm, ultimately, you control the function itself whose output is deterministic. Even today, given an api, people can choose to call or not call based on some factor. We don’t call this fuzzy logic. E.g., people can decide to sell or buy stock through an api based on some internal calculations - doesn’t make the system “fuzzy”.
- m3kw9 3y agoIf you feed that result into another io box you may or may not know if that is the correct answer, which may need some sort of error detection. I think this is going to be majority of the use cases
- Sai_ 3y agoHm, I see what you mean. Afaict, only the decision to call or not call a function is up to the model (fuzzy). Once it decides to call the function, it generates mostly correct JSON based on your schema and returns that to you as is (not very fuzzy). It’ll be interesting to test APIs which accept user inputs. Depending on how ChatGPT populates the JSON, the API could be required to understand/interpret/respond to lots of variability in inputs.
- m3kw9 3y agoYeah I’ve tested, you should use the curl example they gave as you can test instantly pasting it into your terminal. The description of the functions is prompt engineering in addition to the original system prompt, need to test the dependency more, it’s so new.
- iamflimflam1 3y agoIt’s pretty interesting how the work they’ve been doing on plugins has fed into this. I suspect that they’ve managed to get a lot of good training data by calling the APIs provided by plugins and detecting when it’s gone wrong from bad request responses.
- rank0 3y agoOpenAI integration is going to be a goldmine for criminals in the future. Everyone and their momma is gonna start passing poorly validated/sanitized client input to shared sessions of a non-deterministic function. I love the future!
- nextworddev 3y agoIn the “future”?
- jonplackett 3y agoThis is useful, but for me at least, GPT-4 is unusable because it sometimes takes 30 seconds + to reply to even basic queries.
- m3kw9 3y agoAlso the rate limit is pretty bad if you want to release any type of app
- jiggawatts 3y agoMore importantly: there's a waiting list. Also, if you want to use both the ChatGPT web app and the API, you'll be billed for both separately. They really should be unified and billed under a single account. The difference is literally just whether there's a "web UI" on top of the API... or not.
- loughnane 3y agoJust this morning I wrote a JSON object. I told GPT to turn it into a schema. I tweaked that and then gave a list of terms for which I wanted GPT to populate the schema accordingly. It worked pretty well without any functions, but I did feel like I was missing something because I was ready to be explicit and there wasn’t any way for me to tell that to GPT. I look forward to trying this out.
- Xen9 3y agoMarvin Minsky was so damn far ahead of his time with Society of Mind. Engineering of cognitively advanced multiagent systems will become the area of research of this century / multiple decades. GPT-GPT > GPT-API in terms of power. The space of possible combinations of GPT multiagents goes beyond imagination since even GPT-4 goes so. Multiagent systems are best modeled with signal theory, graph theory and cognitive science. Of course "programming" will also play a role, in sense of abstractions and creation of systems of / for thought. Signal theory will be a significant approach for thinking about embedded agency. Complex multiagent systems approach us.
- SanderNL 3y agoMakes me think of the Freud/Jungian notions of personas in us that are in various degrees semi-autonomously looking out for themselves. The “angry” agent, the “child” agent, so on.
- sublinear 3y ago> The process is simple enough that you can let non-technical people build something like this via a no-code interface. No-code tools can leverage this to let their users define “backend” functionality. Early prototypes of software can use simple prompts like this one to become interactive. Running an LLM every time someone clicks on a button is expensive and slow in production, but probably still ~10x cheaper to produce than code. Hah wow... no. Definitely not.
- amolgupta 3y agoI pass a kotlin data class and ask chatGPT to return json which can be parsed by that class. Reduces errors with date-time parsing and other formatting issues and takes up lesser tokens than the approach in the article.
- coding123 3y agoWe're not far from writing a bunch of stubs, query GPT at startup to resolve the business logic. I guess we're going to need a new JAX-RS soon.
- runeb 3y agoThe way openai implemented this is really clever, beyond how neat the plugin architecture is, as it lets them peek one layer inside your internal API surface and can infer what you intend to do with the LLM output. Collecting some good data here.
- bel423 3y agoDid people really struggle with getting JSON outputs from GPT4. You can literally do it zero shot by just saying match this typescript type. GPT3.5 would output perfect JSON with a single example. I have no idea why people are talking about this like it’s a new development.
- brolumir 3y agoUnfortunately, in practice that works only most of the time. At least in our experience (and the article says something similar) sometimes ChatGPT would return something completely different when JSON-formatted response would be expected.
- blamy 3y agoI've been using the same prompts for months and have never seen this happen on 3.5-turbo let alone 4. https://gist.github.com/BLamy/244eec016beb9ad8ed48cf61fd205428 https://gist.github.com/BLamy/244eec016beb9ad8ed48cf61fd2054...
- tornato7 3y agoIn my experience if you set the temperature to zero it works 99.9% of the time, and then you can just add retry logic for the remaining 0.1%
- srameshc 3y agoI've used GCP Vertex AI for a specific task and the prompt was to generate a JSON response with keys specified and it does generate the result as JSON with said keys.
- twelfthnight 3y agoIssue is that's it's not guaranteed, unlike this new openai feature. Personally, Ive found Vertex AI's json output to be not so great, it often uses single quotes in my experience. But maybe you have figured out the right prompts? I'd be interested what you use if so.
- arsdragonfly 3y ago[dead]
- lasermatts 3y agoI thought GPT-4 was doing a pretty good job at outputting JSON (for some of the toy problems I've given it like some of my gardening projects.) Interesting to see this hit the very top of HN
- khazhoux 3y agoI'm trying to experiment with the API but the response time is always in the 15-25second range. How are people getting any interesting work done with it? I see others on the OpenAPI dev forum complaining about this too, but no resolution.
- danShumway 3y agoI'm concerned that OpenAI's example documentation suggests using this to A) construct SQL queries and B) summarize emails, but that their example code doesn't include clear hooks for human validation before actions are called. For a recipe builder it's not so big a deal, but I really worry how eager people are to remove human review from these steps. It gets rid of a very important mechanism for reducing the risks of prompt injection. The top comment here suggests wiring this up to allow GPT-4 to recursively call itself. Meanwhile, some of the best advice I've seen from security professionals on secure LLM app development is to whenever possible completely isolate queries from each other to reduce the potential damage that a compromised agent can do before its "memory" is wiped. There are definitely ways to use this safely, and there are definitely some pretty powerful apps you could build on top of this without much risk. LLMs as a transformation layer for trusted input is a good use-case. But are devs going to stick with that? Is it going to be used safely? Do devs understand any of the risks or how to mitigate them in the first place? 3rd-party plugins on ChatGPT have repeatedly been vulnerable in the real world, I'm worried about what mistakes developers are going to make now that they're actively encouraged to treat GPT as even more of a low-level data layer. Especially since OpenAI's documentation on how to build secure apps is mostly pretty bad, and they don't seem to be spending much time or effort educating developers/partners on how to approach LLM security.
- abhibeckert 3y agoIn my opinion the only way to use it safely is to ensure your AI only has access to data that the end user already has access to. At that point, prompt injection is no-longer an issue - because the AI doesn't need to hide anything. Giving GPT access to your entire database, but telling it not to reveal certain bits, is never going to work. There will always be side channel vulnerabilities in those systems.
- jacobr1 3y ago> your AI only has access to data that the end user already has access to. That doesn't work for the same reason you mention with a DB ... any data source is vulnerable to indirect injection attacks. If you open the door to ANY data source this a factor, including ones under the sole "control" of the user.
- andsoitis 3y agohaving gpt-4 as a dependency for your product or business seems... shortsighted
- ulrikrasmussen 3y agoHas anyone tried throwing their backend Swagger at this and made ChatGPT perform user story tests?
- l5870uoo9y 3y agoThis was technically possible before. I think the approach used by many - myself included - is to simply embed results in a markdown code block and the match it with regex pattern. Then you just need to phrase the prompt to generate the desired output. This is an example of that generating the arguments for the MongoDB's `db.runCommand()` function: https://aihelperbot.com/snippets/cliwx7sr80000jj0finjl46cp https://aihelperbot.com/snippets/cliwx7sr80000jj0finjl46cp