15 ms·
OpenAI DevDay 2024 live blog
- cedws 2y agoWebSockets for realtime? WS is TCP based, wouldn’t it be better to use something UDP based if you want to optimise for latency?
- bigcat12345678 2y agoSeems mostly standard items so far.
- thenameless7741 2y agoBlog updates: - Introducing the Realtime API: https://openai.com/index/introducing-the-realtime-api/ https://openai.com/index/introducing-the-realtime-api/ - Introducing vision to the fine-tuning API: https://openai.com/index/introducing-vision-to-the-fine-tuning-api/ https://openai.com/index/introducing-vision-to-the-fine-tuni... - Prompt Caching in the API: https://openai.com/index/api-prompt-caching/ https://openai.com/index/api-prompt-caching/ - Model Distillation in the API: https://openai.com/index/api-model-distillation/ https://openai.com/index/api-model-distillation/ Docs updates: - Realtime API: https://platform.openai.com/docs/guides/realtime https://platform.openai.com/docs/guides/realtime - Vision fine-tuning: https://platform.openai.com/docs/guides/fine-tuning/vision https://platform.openai.com/docs/guides/fine-tuning/vision - Prompt Caching: https://platform.openai.com/docs/guides/prompt-caching https://platform.openai.com/docs/guides/prompt-caching - Model Distillation: https://platform.openai.com/docs/guides/distillation https://platform.openai.com/docs/guides/distillation - Evaluating model performance: https://platform.openai.com/docs/guides/evals https://platform.openai.com/docs/guides/evals Additional updates from @OpenAIDevs: https://x.com/OpenAIDevs/status/1841175537060102396 https://x.com/OpenAIDevs/status/1841175537060102396 - New prompt generator on https://playground.openai.com https://playground.openai.com - Access to the o1 model is expanded to developers on usage tier 3, and rate limits are increased (to the same limits as GPT-4o) Additional updates from @OpenAI: https://x.com/OpenAI/status/1841179938642411582 https://x.com/OpenAI/status/1841179938642411582 - Advanced Voice is rolling out globally to ChatGPT Enterprise, Edu, and Team users. Free users will get a sneak peak of it (except EU).
- visarga 2y ago> Advanced Voice is rolling out globally to ChatGPT Enterprise, Edu, and Team users. Free users will get a sneak peak of it. So regular paying users from EU are still left out in the cold.
- AlanYx 2y agoIt's probably stuck in legal limbo in the EU. The recently passed EU AI Act prohibits "AI systems aiming to identify or infer emotions", and Advanced Voice does definitely infer the user's emotions. (There is an exemption for "AI systems placed on the market strictly for medical or safety reasons, such as systems intended for therapeutical use", but Advanced Voice probably doesn't benefit from that exemption.)
- qwertox 2y agoApparently this prohibition only applies to "situations related to the workplace and education", and, in this context, "That prohibition should not cover AI systems placed on the market strictly for medical or safety reasons" So it seems to be possible to use this in a personal context. https://artificialintelligenceact.eu/recital/44/ https://artificialintelligenceact.eu/recital/44/ > Therefore, the placing on the market, the putting into service, or the use of AI systems intended to be used to detect the emotional state of individuals in situations related to the workplace and education should be prohibited. That prohibition should not cover AI systems placed on the market strictly for medical or safety reasons, such as systems intended for therapeutical use.
- AlanYx 2y agoThis is true, though it may not make sense commercially for them to offer an API that can't be used for workplace (business) applications or education.
- qwertox 2y agoI see what you mean, but I think that "workplace" specifically refers to the context of the workplace, so that an employer cannot use AI to monitor the employees, even if they have been pressured to agree to such a monitoring. I think this is unrelated to "commercially offering services which can detect emotions". But then I don't get the spirit of that limitation, as it should be just as applicable to TVs listening in on your conversations and trying to infer your emotions. Then again, I guess that for these cases there are other rules in place which prohibit doing this without the explicit consent of the user.
- hidelooktropic 2y agoAny word on increased weekly caps on o1 usage?
- zamadatix 2y agoWeekly caps are for standard accounts (not going to be talked about at DevDay). The blog does note RPM changes for the API though: "10:30 They started with some demos of o1 being used in applications, and announced that the rate limit for o1 doubled to 10000 RPM (from 5000 RPM) - same as GPT-4 now."
- nielsole 2y ago> The first big announcement: a realtime API, providing the ability to use WebSockets to implement voice input and output against their models. I guess this is using their "old" turn-based voice system?
- bcherry 2y agoNo, it's the same thing as ChatGPT advanced voice. Full speech-to-speech model.
- chrisshroba 2y agoRight, see the "Handling interruptions" section here: https://platform.openai.com/docs/guides/realtime/integration https://platform.openai.com/docs/guides/realtime/integration
- qwertox 2y ago> The Realtime API improves this by streaming audio inputs and outputs directly, enabling more natural conversational experiences. It can also handle interruptions automatically, much like Advanced Voice Mode in ChatGPT. > Under the hood, the Realtime API lets you create a persistent WebSocket connection to exchange messages with GPT-4o. The API supports function calling(opens in a new window), which makes it possible for voice assistants to respond to user requests by triggering actions or pulling in new context. - This sounds really interesting, and I see a great use cases for it. However, I'm wondering if the API provides a text transcription of both the input and output so that I can store the data directly in a database without needing to transcribe the audio separately. - Edit: Apparently it does. It sends `conversation.item.input_audio_transcription.completed` [0] events when the input transcription is done (I guess a couple of them in real-time) and `response.done` [1] with the response text. [0] https://platform.openai.com/docs/api-reference/realtime-server-events/conversation-item-input-audio-transcription-completed https://platform.openai.com/docs/api-reference/realtime-serv... [1] https://platform.openai.com/docs/api-reference/realtime-server-events/response-done https://platform.openai.com/docs/api-reference/realtime-serv...
- tough 2y agosaw velvet show hn the other dya, could be usful for storng these https://news.ycombinator.com/item?id=41637550 https://news.ycombinator.com/item?id=41637550
- BoorishBears 2y agoOpenAI just launched the equivalent of Velvet as a full fledged feature today. But seperate from that you typically want some application specific storage of the current "conversation" in a very different format than raw request logging.
- bcherry 2y agoyes it transcribes inputs automatically, but not in realtime. outputs are sent in text + audio but you'll get the text very quickly and audio a bit slower, and of course the audio takes time to play back. the text also doesn't currently have timing cues so its up to you if you want to try to play it "in sync". if the user interrupts the audio, you need to send back a truncation event so it can roll its own context back, and if you never presented the text to the user you'll need to truncate it there as well to ensure your storage isn't polluted with fragments the user never heard.
- serjester 2y agoThe eval platform is a game changer. It's nice to have have a solution from OpenAI given how much they use a variant of this internally. I've tried like 5 YC startups and I don't think anyone's really solved this. There's the very real risk of vendor lock-in but quickly scanning the docs seems like it's a pretty portable implementation.
- ponty_rick 2y ago> 11:43 Fields are generated in the same order that you defined them in the schema, even though JSON is supposed to ignore key order. This ensures you can implement things like chain-of-thought by adding those keys in the correct order in your schema design. Why not use an array of key value pairs if you want to maintain ordering without breaking traditional JSON rules? [ {key1:value1}, {key2:value2} ]
- YetAnotherNick 2y agoI don't think openai models supports this pattern. You can only have array of fixed types. Or basically keys should be same. See [1] [1]: https://platform.openai.com/docs/guides/structured-outputs/supported-schemas https://platform.openai.com/docs/guides/structured-outputs/s...
- benatkin 2y ago> even though JSON is supposed to ignore key order Most tools preserve the order. I consider it to be an unofficial feature of JSON at this point. A lot of people think of it as a soft guarantee, but it’s a hard guarantee in all the recent JavaScript and python versions. There are some common places where it’s lost, like JSONB in Postgres, but it’s good to be aware that this unofficial feature is commonly being used.
- superdisk 2y agoHoly crud, I figured they would guard this for a long time and I was really salivating to make some stuff with it. The doors are wide open for all sorts of stuff now, Advanced Voice is the first feature since ChatGPT initially came out that really has my jaw on the floor.
- minimaxir 2y agoFrom the Realtime API blog post: https://openai.com/index/introducing-the-realtime-api/ https://openai.com/index/introducing-the-realtime-api/ > Audio in the Chat Completions API will be released in the coming weeks, as a new model `gpt-4o-audio-preview`. With `gpt-4o-audio-preview`, developers can input text or audio into GPT-4o and receive responses in text, audio, or both. > The Realtime API uses both text tokens and audio tokens. Text input tokens are priced at $5 per 1M and $20 per 1M output tokens. Audio input is priced at $100 per 1M tokens and output is $200 per 1M tokens. This equates to approximately $0.06 per minute of audio input and $0.24 per minute of audio output. Audio in the Chat Completions API will be the same price. As usual, OpenAI failed to emphasize the real-game changer feature at their Dev Day: audio output from the standard generation API. This has severe implications for text-to-speech apps, particularly if the audio output style is as steerable as the gpt-4o voice demos.
- OutOfHere 2y ago> and $0.24 per minute of audio output That is substantially more expensive than TTS (text-to-speech) which already is quite expensive.
- qwertox 2y agoI agree. I'm wondering if it is possible to disable output streaming of audio and just get the text response event.
- colaco 2y agoIt seems so. The configuration of the session accepts a parameter (modalities) that could restrict the response only to text. See it in https://platform.openai.com/docs/api-reference/realtime-client-events https://platform.openai.com/docs/api-reference/realtime-clie....
- bcherry 2y agocorrect - you should also be able to save a lot by skipping their built-in VAD and doing turn detection (if you need it) locally to avoid paying for silent inputs.
- siva7 2y agoI've never seen a company publishing consistently groundbreaking features at such a speed like this one. I really wonder how their teams work. It's unprecedented at what i've seen in 15 years software
- pheeney 2y agoI wonder how much they use their own products internally to speed up development and decisions.
- amlib 2y agoAnd I wonder how much they use them externally to influence the online conversations about their own products/company.
- abound 2y agoThey definitely use their own products internally, perhaps to a fault: While chatting with OpenAI recruiters, I received calendar events with nonsensical DALLE-generated calendar images, and "interview prep" guides that were clearly written by an older GPT model.
- jonchurch_ 2y agoThey have roles on their “Leverage Engineering” team which appears to be exactly this https://openai.com/careers/full-stack-software-engineer-leverage-engineering/ https://openai.com/careers/full-stack-software-engineer-leve...
- roboboffin 2y agoIs it that most models are based on the transformer architecture ? And so performance improvements can then we used throughout their different products ?
- IdiocyInAction 2y agoAFAIK a lot of these ideas are not new (the JSON thing was done with OS models before) and OpenAI is possibly the hottest startup with the most funding this decade (maybe even past two decades?), so I think this is actually all within expectations.
- sammyteee 2y agoLoving these live updates, keep em coming! Thanks Simon!
- lysecret 2y agoUsing structured outputs for generative ui is such a cool idea does anyone know some cool web demos related to this ?
- jiggawatts 2y agoI just had an evil thought: once AIs are fast enough, it would be possible to create a “dynamic” user interface on the fly using an AI. Instead of Java or C# code running in an event loop processing mouse clicks, in principle we could have a chat bot generate the UI elements in a script like WPF or plain HTML and process user mouse and keyboard input events! If you squint at it, this is what chat bots do now, except with a “terminal” style text UI instead of a GUI or true Web UI. The first incremental step had already been taken: pretty-printing of maths and code. Interactive components are a logical next step. It would be a mere afternoon of work to write a web server where the dozens of “controllers” is replaced with a single call to an LLM API that simply sends the previous page HTML and the request HTML with headers and all. “Based on the previous HTML above and the HTTP request below, output the response HTML.” Just sprinkle on some function calling and a database schema, and the site is done!
- ghthor 2y agoThat actually sounds pretty entertaining. Especially if there is dynamic user input, like text box input
- jiggawatts 2y agoOther than being borderline impossible to secure, it “should just work” once the AIs get smart enough. Fine-tuning the model based on example pages and responses might be all that’s required for a sufficient level of consistency. An immediate use-case might be prototyping in-place. If you have an existing site, you can capture the request-response pairs and train the AI on it, annotated with the spec docs. Then tell it to implement some new functionality and it should be able to. Just route a subset of the site to the AI instead of the normal controllers. One could “design” new components and functionality in English and try it instantly with no compilation or deployment steps!
- famouswaffles 2y agoImage output for 4o in the API would be very nice but i'm not sure if that's at all in the cards. Audio output in the api now but you lose image input. Why ? That's a shame.
- 101008 2y agoI understand the Realtime API voice novelty, and the techonological achievement it is, but I don't see it from the product point of view. It looks like one of those startups finding a solution before knowing the problem. The two examples shown in the DevDay are the things I don't really want to do in the future. I don't want to talk to anybody, and I don't want to wait for their answer in a human form. That's why I order my food through an app or Whatsapp, or why I prefer to buy my tickets online. In the rare case I call to order food, it's because I have a weird question or a weird request (can I pick it up in X minutes? Can you prepare it in a different way?) I hope we don't start seeing apps using conversations as interfaces because it would really horrible (leaving aside the fact that a lot of people don't know how to communicate themselves, different accents, sound environments, etc), while clicking or typing work almost the same for everyone (at least much more normalized than talking)
- bcherry 2y agokeep in mind that this is just v1 of the realtime api. they'll add realtime vision/video down the road which can also have wide applications beyond synchronous communication.
- ilaksh 2y agoYou're right, having a voice conversation for any reason is just so passe these days. They should stop adding microphones to phones and everything. So old-fashioned and inefficient. And who wants to ever have to actually talk to someone or some AI to ask for anything? I'm sure our vocal cords will evolve away soon. They are so primitive. Vestigial organs.
- olafgeibig 2y agoYou made my day
- 101008 2y agoI love having voice conversations with friends, family, and people I care of. Not with businesses.
- 2y ago
- modeless 2y agoI didn't expect an API for advanced voice so soon. That's pretty great. Here's the thing I was really wondering: Audio is $.06/min in, $.24/min out. Can't wait to try some language learning apps built with this. It'll also be fun for controlling robots.
- N_A_T_E 2y agoI just need their API to be faster. 15-30 seconds per request using 4o-mini isn't good enough for responsive applications.
- carlgreene 2y agoThat is odd. Longest I’ve experienced in my use of it is a few seconds.
- BoorishBears 2y agoYou should try Azure: it comes with dedicated capacity which is typically a very expensive "call our sales team" feature with OpenAI
- petesergeant 2y agoThat doesn’t match my experience using it a lot at all
- simonw 2y agoThe new Realtime Websocket API appears to send back responses within less than a second. It might be just what you want.
- bcherry 2y agoyes and you can use it in text-text mode if you want. a key benefit is for turn-based usages (where you have running back and forth between user and assistant) you only need to send the incremental new input message for each generation. this is better than "prompt caching" on the chat completions API, which is basically a pricing optimization, as it's actually a technical advantage that uses less upstream bandwidth.
- alach11 2y agoIt's pretty amazing that they made prompt caching automatic. It's rare that a company gives a 50% discount without the customer explicitly requesting it! Of course... they might be retaining some margin, judging by their discount being 50% vs. Anthropic's 90%.
- WiSaGaN 2y agoThis was first done by deepseek. [1] [1]: https://platform.deepseek.com/api-docs/news/news0802/ https://platform.deepseek.com/api-docs/news/news0802/
- nextworddev 2y agoHaven’t tried Deepseek - how do they compare to OaI?
- WiSaGaN 2y agoThey release SOTA open source coding models. [1] Their API us also incredibly cheap due to the novel attention and MoE arch. [1]: https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/
- voiper1 2y agoAider benchmarks them great for coding... super slow token generation. Much cheaper but once you're used to the speed... it's too slow.
- jbaudanza 2y agoInteresting choice of a 24kHz sample rate for PCM audio. I wonder if the model was trained on 24kHz audio, rather than the usual 8/16kHz for ML models.
- simonw 2y agoFor anyone who’s interested, I’ve written up details of how the underlying live blog system works here: https://til.simonwillison.net/django/live-blog https://til.simonwillison.net/django/live-blog