14 ms·
OpenAI Audio Models
- danso 2y agoThe voices are pretty convincing. It's funny to hear drastically the tone of the reading can change when repeatedly stopping and restarting the samples without changing any of the settings.
- theoryofx 2y agoStill seems like Elevenlabs is crushing them on realtime audio, or does this change things?
- atlasunshrugged 2y agoI'm also curious about this for longform content. Will this be competitive for something like creating an audiobook?
- borgdefenser 2y agoFor $0.015 a minute it has to be. The books I am listening to now wouldn't even be $10. Any future price drops then will really make this a no-brainer. The Elevenlabs pricing to me makes it completely useless for audiobooks that I just want to listen to for my personal enjoyment.
- prdonahue 2y agoDo you have any affiliation with Elevenlabs?
- atlasunshrugged 2y agoFWIW I have no affiliation with any of these companies but I have a book coming out soon and have been researching AI audiobook tools and Elevenlabs seems to be far and away the consensus for that at least
- theoryofx 2y agoI do not have any affiliation with Elevenlabs or OpenAI except as a user of their APIs. I'd actually prefer it if OpenAI had a better realtime product than Elevenlabs because it'd be more convenient.
- minimaxir 2y agoThis is an official OpenAI tool linked from the new model announcement (https://openai.com/index/introducing-our-next-generation-audio-models/ https://openai.com/index/introducing-our-next-generation-aud... ), despite the branding difference.
- nickthegreek 2y agoTry the refresh button to get a new list of vibe styles.
- varunneal 2y agoOne of the most novel demos I've seen openai ship in a few years. I love how it looks almost like a synth. Fun to play around with!
- stephenheron 2y agoQuite disappointing their speech to text models are not open source. Whisper was really good and it was great it was open to play around with. I guess this continues OpenAI's approach of not really being open!
- nickthegreek 2y agoIndeed. Right now I think our open choices are Piper, Kokoro and Orpheus.
- DrPhish 2y agoIn my opinion GPT-SoVITS is the best if you can put in the effort. I'm still using v2 since the output is so good. Its also the best multilingual one in my testing on Japanese inputs.
- pzo 2y agocan it support more languages rather than only English, Chinese, Japanese, Korean?
- nickthegreek 2y agohadnt messed with that one before. my needs are more real time for voice assistant but was neat to play with on hugginface. https://huggingface.co/spaces/lj1995/GPT-SoVITS-v2 https://huggingface.co/spaces/lj1995/GPT-SoVITS-v2
- GaggiX 2y agoHe was talking about STT models, not TTS. Whisper is open source and a good solution in many cases (in particular finetuned ones).
- pzo 2y agoregarding STT we got also today 2 new models from Nvidia: https://huggingface.co/nvidia/canary-180m-flash https://huggingface.co/nvidia/canary-180m-flash https://huggingface.co/nvidia/canary-1b-flash https://huggingface.co/nvidia/canary-1b-flash second in Open ASR leaderboard https://huggingface.co/spaces/hf-audio/open_asr_leaderboard https://huggingface.co/spaces/hf-audio/open_asr_leaderboard Sadly only supports 4 languages (english, german, spanish, french)
- Etheryte 2y agoRecommended input for anyone trying this out: Voice: Onyx Vibe: Heavy german accent, doing an Arnold Schwarzenegger impression, way over the top for comedic effect. Deep booming voice, uses pauses for dramatic effect.
- jeffharris 2y agoso good! https://www.openai.fm/#28540f27-5b51-445a-b1d6-1c89711a2c4f https://www.openai.fm/#28540f27-5b51-445a-b1d6-1c89711a2c4f
- Sohcahtoa82 2y agoI hit Play 3 times and got 3 very different results. One merely sounded like it had a slight German accent, once just sounded kind of raspy, and the third sound like a normal American English speaker.
- soared 2y agoWeird 1 in maybe 5 times when I refresh that page I get a very good Arnold. The rest are very bad.
- o_____________o 2y agointeresting to see what accents do and do not work
- mft_ 2y agoWeird; trying exactly this, and every time I stop and play again, I get a totally different voice. One of them (if I'm not mistaken) was cod-Russian.
- mchusma 2y agoI also find this strange, and I wonder if I can get a consistent voice out of this. If using the api with a vibe/instructions for a back and forth, will it be consistent? This example app they provide implies no?
- 2y ago
- jcmp 2y agoHow do you call this desing/ui astethic? I like it
- IndignantTyrant 2y agoYou can screenshot and ask chatgpt lol
- randomcatuser 2y agoneumorphism!
- havefunbesafe 2y agoby copying Teenage Engineering
- vyrotek 2y ago"teenage engineering"
- danso 2y agoInteresting, I inserted a bunch of "fuck"s in the text and the "NYC Cabbie" voice read it all just fine. When I switched to other voices ("Connoisseur", "Cheerleader", "Santa"), it responded "I'm sorry I can't assist with that request". I switched back to "NYC Cabbie" and it again read it just fine. I then reloaded the session completely, refreshed the voice selections until "NYC Cabbie" came up again, and it still read the text without hesitation. The text: > In my younger and more vulnerable years my father fuck gave me some fuck fuck advice that I've been fuck fuck FUCK OH FUCK turning over in my mind ever since. > "Whenever you feel like criticizing any one," he told me, oh fuck! FUCK! "just remember that all the people in this world haven't had fuck fuck fuck FUCKERKER the advantages that you've had." edit: "Emo Teenager", "Mad Scientist", and "Smooth Jazz" are able to read the text. However, "Medieval Knight" and "Robot" cannot.
- nazgulsenpai 2y agoGlad I'm not the only one whose inner 12 year old curiosity is immediately triggered by free input TTS. Swear words and just raking my hands across the keyboard to insert gibberish in every possible accent.
- andrewinardeer 2y agoIt won't generate slurs, though.
- dvngnt_ 2y agowhat did you try?
- deleted 2y ago[deleted]
- andrewinardeer 2y agoIs this bait? Lol. Try a few for yourself.
- 2y ago
- ComputerGuru 2y agoIt would be much more convenient to use if changing the voice model worked on the fly, without having to stop and start the audio.
- amitport 2y agoLouis CK | about airplane Wi Fi https://www.youtube.com/watch?v=me4BZBsHwZs https://www.youtube.com/watch?v=me4BZBsHwZs
- swyx 2y agoconvenient why?
- pests 2y agoIt breaks the flow to stop / start just to switch voices. I expected to be able to click a new voice and have it pick up on the next word / sentence with that voice. To compare voices I have to stop, click a new one, then start and wait for it to process. Then I can hear the new voice but its already been 15 seconds since I heard the last so what was the difference even?
- islewis 2y agoCool format for a demo. Some of the voices have a slight "metallic" ring to them, something I've seen a fair amount with Eleven Labs' models. Does anyone have any experience with the realtime latency of these Openai TTS models? ElevenLabs has been so slow (much slower than the latency they advertise), which makes it almost impossible to use in realtime scenarios unless you can cache and replay the outputs. Cartesia looks to have cracked the time to first token, but i've found their voices to be a bit less consistent than Eleven Labs'.
- carbocation 2y agoNova+Serene sounds very metallic at the beginning about 50% of the time for me.
- jeffharris 2y agosome of the older voices are definitely less steerable, more robotic we put little stars in the bottom right corner for the newer voices, which should sound better
- jeffharris 2y agoHey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!
- TheAceOfHearts 2y agoIs it against the TOS to use it for sexually explicit content?
- knicholes 2y agoI don't have your answer, but as far as innuendo goes, it's definitely capable!
- pzo 2y ago[flagged]
- jeffharris 2y agoYes, from our terms: "Don’t build tools that may be inappropriate for minors, including: Sexually explicit or suggestive content. This does not include content created for scientific or educational purposes." https://openai.com/policies/usage-policies/ https://openai.com/policies/usage-policies/
- CamperBob2 2y ago[flagged]
- startupsfail 2y agoThe general consensus of AI overlords is that humans are minors.
- basitmakine 2y agoI don't think they're anywhere near TaskAGI or ElevenLabs level.
- tantalor 2y agoIt does a good job with Pirate voice. It can even inject "Arrr matey"
- deleted 2y ago[deleted]
- jtbayly 2y agoI don’t get it. These voices all have a not-so-subtle vibration in them that makes them feel worse than Siri to me. I was expecting a lot better.
- pier25 2y agoyeah the voices sound terrible I'm guessing their spectral generator is super low res to save on resources
- stavros 2y agoIs there a way to pay for higher quality? I don't see a way to pay at all, this just works without an API key, even with the generated code. I agree though, these voices sound like their buffer is always underrunning.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- minimaxir 2y agoOne very important quote from the official announcement: > For the first time, developers can “instruct” the model not just on what to say but how to say it—enabling more customized experiences for use cases ranging from customer service to creative storytelling. The instructions are the "vibes" in this UI. But the announcement is wrong with the "for the first time" part: it was possible to steer the base GPT-4o model to create voices in a certain style using system prompt engineering (blogged about here: https://minimaxir.com/2024/10/speech-prompt-engineering/ https://minimaxir.com/2024/10/speech-prompt-engineering/ ) out of concern that it could be used as a replacement for voice acting, however it was too expensive and adherence isn't great. The schema of the vibes here implies that this new model is more receptive to nuance, which changes the calculus. The test cases from my post behave as expected, and the cost of gpt-4o-mini-tts audio output is $0.015 / minute (https://platform.openai.com/docs/pricing https://platform.openai.com/docs/pricing ), which is about 1/20th of the cost of my initial experments and is now feasible to use to potentially replace common voice applications. This has implications, and I'll be testing more around more nuanced prompt engineering.
- benjismith 2y agoIf I'm reading the pricing correctly, these models are SIGNIFICANTLY cheaper than ElevenLabs. https://platform.openai.com/docs/pricing https://platform.openai.com/docs/pricing If these are the "gpt-4o-mini-tts" models, and if the pricing estimate of "$0.015 per minute" of audio is correct, then these prices 85% cheaper than those of ElevenLabs. https://elevenlabs.io/pricing https://elevenlabs.io/pricing With ElevenLabs, if I choose their most cost-effectuve "Business" plan for $1100 per month (with annual billing of $13,200, a savings of 17% over monthly billing), then I get 11,000 minutes TTS, and each minute is billed at 10 cents. With OpenAI, I could get 11,000 minutes of TTS for $165. Somebody check my math... Is this right?
- fixprix 2y agoIt looks like they are targeting Google's TTS price point which is $16 per million characters which comes out to $0.015/minute.
- lukebuehler 2y agoyes, I think you are right. When I did the math on 11labs million chars I got the same numbers (Pro plan). I'm super happy about this, since I took a bet that exactly this would happen. I've just been building a consumer TTS app that could only work with significant cheaper TTS prices per million character (or self-hosted models)
- benjismith 2y agoSame for me :)
- zacmps 2y agoWhat does it do?
- lukebuehler 2y agoConvert any file (pdf, epub, txt) to an audoibook, downloadable as mp3, or directly listenable via RSS feed in, say, Apple Potcasts app. Basically make one-off audiobooks for yourself or a few friends.
- fixprix 2y agoIs this right? The current best TTS from OpenAI uses gpt-4o-audio-preview which is $2.50 input text, $80 output audio, the new gpt-4o-mini-tts is $0.60 input text, $12 output audio. An average 5x price reduction. Going the other way, transcribe with gpt-4o-audio-preview price was $40 input audio, $10 output text, the new gpt-4o-transcribe is $6 input audio and $10 output text. Like a 7x reduction on the input price. TTS/Transcribe with gpt-4o-audio-preview was a hack where you had to prompt with 'listen/speak this sentence:' and it often got it wrong. These new dedicated models are exactly what we needed. I'm currently using the Google TTS API which is really good, fast and cheap. They charges $16 per million characters which is exactly the same as OpenAI's $0.015 per minute estimate. Unfortunately it's not really worth switching over if the costs are exactly the same. Transcription on the other hand is 1.6¢/minute with Google and 0.6¢/minute with OpenAI now, that might be worth switching over for.
- pzo 2y agoyou can compare TTS pricing here: https://artificialanalysis.ai/text-to-speech https://artificialanalysis.ai/text-to-speech Previous offering from OpenAI was $15 for TTS and $30 for TTS HD so not 5x reduction. This one is slighly cheaper but definitely more capable (if you need control vibe)
- fixprix 2y agoThat's a really cool page thanks. Does it have stats for other languages? In my experience the OpenAI TTS APIs were really bad, messing up all the time in foreign languages. Practically unusable for my use case. You'd have to use the gpt-4o-audio-preview to get anything close to passable, but it was expensive. Which is why I'm using Google TTS which is very fast, high quality, and provides first class support for almost every language. I look forward to comparing it with this model, the price being the same is unfortunate as there's less incentive to switch. The transcribe price is cheaper than Google it looks like so that's worth considering.
- pzo 2y agoInteresting for me Open TTS for Polish was better than Google TTS (but they have few options) - which one did you used? WaveNet? Sadly haven't seen quality evaluation for TTS for foreign languages
- tosh 2y agoAre these models only available via the API right now or also available as open weights?
- evalstate 2y agoReally looking forward to integrating with these models. The next version of Model Context Protocol will have native audio support (https://github.com/modelcontextprotocol/specification/pull/93 https://github.com/modelcontextprotocol/specification/pull/9...), which will open up plenty of opportunities for interop.
- pklimk 2y agoInterestingly "replaces every second word with potato" and "speaks in Spanish instead of English" both (kind of) work as a style, so it's clear there's significant flexibility and probably some form of LLM-like thing under the hood.
- benjismith 2y agoIs there way to get "speech marks" alongside the generated audio? FYI, Speech marks provide millisecond timestamp for each word in a generated audio file/stream (and a start/end index into your original source string), as a stream of JSONL objects, like this: {"time":6,"type":"word","start":0,"end":5,"value":"Hello"} {"time":732,"type":"word","start":7,"end":11,"value":"it's"} {"time":932,"type":"word","start":12,"end":16,"value":"nice"} {"time":1193,"type":"word","start":17,"end":19,"value":"to"} {"time":1280,"type":"word","start":20,"end":23,"value":"see"} {"time":1473,"type":"word","start":24,"end":27,"value":"you"} {"time":1577,"type":"word","start":28,"end":33,"value":"today"} AWS uses these speech marks (with variants for "sentence", "word", "viseme", or "ssml") in their Polly TTS service... The sentence or word marks are useful for highlighting text as the TTS reads aloud, while the "viseme" marks are useful for doing lip-sync on a facial model. https://docs.aws.amazon.com/polly/latest/dg/output.html https://docs.aws.amazon.com/polly/latest/dg/output.html
- minimaxir 2y agoPassing the generated audio back to GPT-4o to ask for the structured annotations would be a fun test case.
- jeffharris 2y agothis is a good solve. we don't support word time stamps natively yet, but are working on teaching GPT-4o that skill
- celestialcheese 2y agowhisper-1 has this with the verbose_json output. Has word level and sentence level, works fairly well. Looks like the new models don't have this feature yet.
- looknee 2y agoHmm I was hoping these would be bridging the gap between what's already been availalbe on their audio API or in the RealtimeAPI vs. Advanced Voice Mode, but the audio quality is really the same as its been up to this point. Does anyone have any clue about exactly why they're not making the quality of Advanced Voice Mode available to build with? It would be game changing for us if they did.
- mlsu 2y agoI gave it (part of) the classic Navy Seal copypasta. Interestingly, the safety controls ("I cannot assist with that request") is sort of dependent on the vibe instruction. NYC cabbie has no problem with it (and it's really, really funny, great job openAI), but anything peaceful, positive, etc. will deny the request. https://www.openai.fm/#56f804ab-9183-4802-9624-adc706c7b9f8 https://www.openai.fm/#56f804ab-9183-4802-9624-adc706c7b9f8
- forgotpasagain 2y agoIt sounds very expressive but weirdly "fake" as if it's targeting to be similar to some NPC character, dataset issue?
- lukeinator42 2y agoyeah it's almost like an uncanny valley where it sometimes feels like the voice trying to be an actor and play a character or something
- anigbrowl 2y agoI mean that's literally the service it's providing. If you asked humans to do the same thing it would sound equally forced. All acting sounds cringe out of context.
- amarcheschi 2y agoIDK, i've done my fair share of amateur acting and at least to me (english is not my first language) there's something more uncanny than just the typical "say this wihtout knowing much of the context)
- jeffharris 2y agoreally depends on the voice. but in general, no we want to sound as realistic as possible and I expect future voices to keep improving on this front
- deleted 2y ago[deleted]
- kartikarti 2y agoWhat does this little star next to the name mean?
- crazygringo 2y agoThis is astonishing. I can type anything I want into the "vibe" box and it does it for the given text. Accents, attitudes, personality types... I'm amazed. The level of intelligent "prosody" here -- the rhythm and intonation, the pauses and personality -- I wasn't expecting anything like this so soon. This is truly remarkable. It understands both the text and the prompt for how the speaker should sound. Like, we're getting much closer to the point where nobody except celebrities are going to record audiobooks. Everyone's just going to pick whatever voice they're in the mood for. Some fun ones I just came up with: > Imposing villain with an upper class British accent, speaking threateningly and with menace. > Helpful customer support assistant with a Southern drawl who's very enthusiastic. > Woman with a Boston accent who talks incredibly slowly and sounds like she's about to fall asleep at any minute.
- solardev 2y agoGuess that's why the video game voice actors are still on strike: https://en.m.wikipedia.org/wiki/2024%E2%80%93present_SAG-AFTRA_video_game_strike https://en.m.wikipedia.org/wiki/2024%E2%80%93present_SAG-AFT... If we as developers are scared of AI taking our jobs, the voice actors have it much worse...
- KeplerBoy 2y agoI don't see how a strike will do anything but accelerate the professions inevitable demise. Can anyone explain how this could ever end in favor of the human laborers striking?
- solardev 2y agoI am not affiliated with the strikers, but I think the idea is that, for now, the companies still want to use at least some human voice acting. So if they want to hire them, they either have to negotiate with the guild or try to find an individual scab willing to cross the picket line and get hired despite the strike. In some industries, there's enough non-union workers that finding replacement workers is easy enough. I guess the voice actors are sufficiently unionized that it's not so easy there, and it seems to have caused some delays in production and also some games being shipped without all their voice lines. But as you surmise, this is at best a stalling tactic. Once the tech gets good enough, fewer companies will want to pay for human voice acting labor. Unions can help powerless individuals negotiate better through collective bargaining, but they can't altogether stop technological change. Jobs, theirs and ours, eventually become obsolete... I don't necessarily think we should artificially protect jobs against technology, but I sure wish we had a better social safety net and retraining and placement programs for people needing to change careers due to factors outside their control.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- tomjen3 2y agoIt doesn't seem clear, but can the model do correct emphesis? On things like single words: I did not steal that horse Is the trivial example of something where intonation of the single word is what matters. More importantly if you are reading something, as a human, you change the intonation, audiolevel, and speed.
- Sohcahtoa82 2y ago> I did not steal that horse > Is the trivial example of something where intonation of the single word is what matters. My go-to for an example of this is "I didn't say she stole my money". Changing which word is emphasized completely changes the meaning of the sentence.
- ForTheKidz 2y agoPricing looks like it's aimed at us peasants, not our lords. Smart if openai wants to survive!
- RobinL 2y agoI'm surprised at how poor this is at following a detailed prompt. It seems capable of generating a consistent style, and so in that sense quite useful. But if you want (say) a regional UK accent it's not even close. I also find it confusing you have to choose a voice. Surely that's what the prompt should be for, especially when the voices have such abstract names. I mean, it's still very impressive when you stand back a bit, but feels a bit half baked Example: Voice: Thick and hearty, with a slow, rolling cadence—like a lifelong Somerset farmer leaning over a gate, chatting about the land with a mug of cider in hand. It’s warm, weathered, and rich, carrying the easy confidence of someone who’s seen a thousand harvests and knows every hedgerow and rolling hill in the county. Tone: Friendly, laid-back, and full of rustic charm. It’s got that unhurried quality of a man who’s got time for a proper chinwag, with a twinkle in his eye and a belly laugh never far away. Every sentence should feel like it’s been seasoned with fresh air, long days in the fields, and a lifetime of countryside wisdom. Dialect: Classic West Country, with broad vowels, softened consonants, and that unmistakable rural lilt. Words flow together in an easy drawl, with plenty of dropped "h"s and "g"s. "I be" replaces "I am," and "us" gets used instead of "we" or "me." Expect plenty of "ooh-arrs," "proper job," and "gurt big" sprinkled in naturally.
- robbomacrae 2y agoI find it works better with shorter simpler instructions. I would try: Voice: Warm and slow, like a friendly Somerset farmer. Tone: Laid-back and rustic. Dialect: Classic West Country with a relaxed drawl and colloquial phrases.
- anigbrowl 2y agoThat seems way overwritten. Try something like 'Jolly old-fashioned rural farmer, Somerset.'
- deleted 2y ago[deleted]
- paul7986 2y agoPersonally I just want to text or talk to Siri or an LLM and have it do whatever I need. Have it interface with AI Agents of companies, businesses, friends or families AI Agents to get whatever I need done like the example on OpenAI.fm site here (rebook my flight). Once it's done it shows me the confirmation on my lock screen and I receive an email confirmation.
- tiahura 2y agoWhen are we going to get the equivalent for Whisper. When is it going to pick up on enthusiasm, sarcasm, etc?
- kibbi 2y agoLarge text-to-speech and speech-to-text models have been greatly improving recently. But I wish there were an offline, on-device, multilingual text-to-speech solution with good voices for a standard PC — one that doesn't require a GPU, tons of RAM, or max out the CPU. In my research, I didn't find anything that fits the bill. People often mention Tortoise TTS, but I think it garbles words too often. The only plug-in solution for desktop apps I know of is the commercial and rather pricey Acapela SDK. I hope someone can shrink those new neural network–based models to run efficiently on a typical computer. Ideally, it should run at under 50% CPU load on an average Windows laptop that’s several years old, and start speaking almost immediately (less than 400ms delay). The same goes for speech-to-text. Whisper.cpp is fine, but last time I looked, it wasn't able to transcribe audio at real-time speed on a standard laptop. I'd pay for something like this as long as it's less expensive than Acapela. (My use case is an AAC app.)
- ZeroTalent 2y agoLook into https://superwhisper.com https://superwhisper.com and their local models. Pretty decent.
- kibbi 2y agoThank you, but they say "Offline models only run really well on Apple Silicon macs."
- ZeroTalent 2y agoMany SOTA apps are, unfortunately, only for Apple M Macs.
- 5kg 2y agoMay I introduce to you https://huggingface.co/canopylabs/orpheus-3b-0.1-ft https://huggingface.co/canopylabs/orpheus-3b-0.1-ft (no affiliation) it's English only afaics.
- kibbi 2y ago
- justanotheratom 2y agoNote that the previous Whisper STT models were Open Source, and these new STT models are not, AFAICT.
- Heidaradar 2y agois it just me or are these voices clearly AI generated? They've obviously been improving at a steady rate but if I saw a YouTube video that had this voice, I'd instantly stop watching it
- smokeydoe 2y agoDoes anyone know of any decent newer open source models for generating sound effects?
- anigbrowl 2y agoJust use a synthesizer. Writing textual prompts is about the most inefficient way of getting what you want. When I was working in film I'd tell directors to stop describing what they had in mind (unless they were referencing something very specific) and try just making some funny mouth noises.
- corobo 2y agoAll these voices are too good these days. I want my home assistant to sound like Auto from Wall-E, dammit! Anyone out there doing any nice robotic robot voices? Best I've got so far is a blend of Ralph and Zarvox from MacOS' `say`, haha say -v zarvox -r 180 "[[volm 0.8]] ${message}" & say -v ralph -r 180 "${message}"
- ranguna 2y agoYou could apply a robotic filter on top of these voices.
- simonw 2y agoBoth the text-to-speech and the speech-to-text models launched here suffer from reliability issues due to combining instructions and data in the same stream of tokens. I'm not yet sure how much of a problem this is for real-world applications. I wrote a few notes on this here: https://simonwillison.net/2025/Mar/20/new-openai-audio-models/ https://simonwillison.net/2025/Mar/20/new-openai-audio-model...
- accrual 2y agoThanks for the write up. I've been writing assembly lately, so as soon as I read your comment, I thought "hmm reminds me of section .text and section .data".
- jncfhnb 2y agoAre there any voice to voice models out there that can replicate inflection of line delivery?
- alach11 2y agoIt's interesting that they pitch this for agent development. The realtime API provides a much simpler architecture for developing agents. Why would you want to string together STT -> LLM -> TTS when you could have a consolidated model doing all three steps? They alluded to there being some quality/intelligence benefits to the multi-step approach, but in the long-run I'd expect them to improve the realtime API to make this unnecessary.
- zhyder 2y agoText allows developers lots for flexibility to do other processing, including RAG, calling APIs yourself and multiple chained LLM invocations. The low latency of realtime API means relying fully on one invocation of their model to do everything.
- alach11 2y agoThe realtime API can be used to call tools [0], but I agree with your general point on the flexibility of working directly with text. [0] https://github.com/openai/openai-realtime-agents https://github.com/openai/openai-realtime-agents
- Arubis 2y agoAt this point, the strongest (and almost only) predictor for a release announcement from OpenAI is a release announcement from Anthropic.
- notlisted 2y agoIt's really quite sad. They did the same thing with Google for a while.
- kgeist 2y agoIn Russian, OpenAI audio models usually have a slight American (?) accent. The intonation and the phonetics fall into the uncanney valley. Does the same happen in other languages?
- buybackoff 2y agoI was experimenting recently with voiceover TTS generation. Did run Kokoro TTS locally and it's magical for how few resources it takes (runs fine in a browser), but only the default female voices (Heart/Bella) are usable, and very good. Then I found that Clipchamp has it built-in and several voices from a big selection there are very good, and free. I've listened to this OpenAI TTS and I could not like them at all even compared to Kokoro.
- redox99 2y agoPretty meh. Coral Dramatic is extremely robotic for example.
- saint_yossarian 2y agoThe site just crashes with service workers disabled. First time I ran into a problem with that setting, which I set over two years ago.
- jeffharris 2y agooh doh. thanks ... we just pushed a fix for the crash. Unfortunately our currently implementation needs service works for streaming audio, so the "fix" was to disable the feature if the worker isn't available
- saint_yossarian 2y agoThanks, that's interesting. I thought service workers were only needed for things like offline support and background activity after a tab is closed (which is why I disabled them). Streaming audio is a new one to me, I wonder if the same could be achieved with web workers instead. Or at least similar use cases like video calls work fine for me without service workers. See e.g. https://github.com/scottstensland/web-audio-workers-sockets https://github.com/scottstensland/web-audio-workers-sockets
- Havoc 2y ago>Browser not supported >Please open openai.fm directly in a modern browser Doesn't seem to like firefox
- dredmorbius 2y agoDittos.
- urbandw311er 2y agoAnyone know if these new models are being added to the realtime API? The linked page is a bit opaque on that.
- blazenby 2y ago[dead]
- deleted 2y ago[deleted]
- nmca 2y agoBeen using elevenlab reader, but these are much better!
- tkgally 2y agoI just tested the "gpt-4o-mini-tts" model on several texts in Japanese, a particularly challenging language for TTS because many character combinations are read differently depending on the context. The produced speech was quite good, with natural intonation and pronunciation. There were, however, occasional glitches, such the word 現在 genzai “now, present” read with a pause between the syllables (gen ... zai) and the conjunction 而も read nadamo instead of the correct shikamo. There were also several places where the model skipped a word or two. However, unlike some other TTS models offering Japanese support that have been discussed here recently [1], I think this new offering from OpenAI is good enough for language users. I certainly could have put it to good use when I was studying Japanese many years ago. But it’s not quite ready for public-facing applications such as commercial audiobooks. That said, I really like the ability to instruct the model on how to read the text. In that regard, my tests in both English and Japanese went well. [1] https://news.ycombinator.com/item?id=42968893 https://news.ycombinator.com/item?id=42968893
- tkgally 2y agoSelf-correction: "... good enough for language users" --> "good enough for language learners."
- joiemoie 2y agoHi! Can you add prefix support? This would be very valuable in being able to support overlapping windows. The only other way would be to use another ai to determine the overlap
- sintezcs 2y agoI love the Teenage Engineering vibe of this page
- skc 2y agoWell, I'm pretty blown away. Especially when you realize that this release will probably get blown out of the water in less than a years time.
- keepamovin 2y agolol OMG these are fantastic. It's as if they hired professional voice actors and cloned their voices. Perhaps that would be lucrative for the voice artists.
- rsp1984 2y agoA bit off-topic but I'm so glad to see skeuomorphic UI make a comeback! Check out the toggle switch in the upper right corner! I hope more designers will follow this example.
- l72 2y agoI am pretty impressed with how well this read chinese!
- fumeux_fume 2y agoThese models show some good improvements in allowing users to control many aspects of the delivery, but it falls deep in the uncanny valley and still has a ways to go in not sounding weird or slightly off-putting. I much prefer the current advanced voice models over these.
- gherard5555 2y agoI tried some wacky strings like "𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯NNNNNNNNNNNNNN𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯𝆺𝅥𝅯" Its hilarious either they start to make harsh noise or say nonsense trying so sing something
- gherard5555 2y agoAlso this one is terrifying if combined with the fitness instructor : "*scream* AAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHHAAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHHAAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHHAAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHHAAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHHAAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHHAAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHHAAAAAAAAAAAAAAAAAAAHHHHHHHHHHHHHHH !!!!!!!!!"
- rachofsunshine 2y agoImpressive in terms of quality, not so much in terms of style. I tried feeding it two prompts with the same script - one to be straightforward and didactic, then one asking it to deliver calculus like a morning shock-jock DJ. They sounded quite similar, and it definitely did not capture the vibe of 97.3 FM's Altman & the Claude with the latter prompt. But then, I got much better results from the cowboy prompt by changing "partner" to "pardner" in the text prompt (even on neighboring words). So maybe it's an issue with the script and not the generation? Giving it "pardner" and an explicit instruction to use a Russian accent still gives me a Texas drawl, so it seems like the script overrides the tone instructions.
- MasterScrat 2y agoComparing the professionally recorded Baldur's Gate Chapter 2 intro with its AI counterpart: - Original: https://www.youtube.com/watch?v=FYcMU3_xT-w&t=5s https://www.youtube.com/watch?v=FYcMU3_xT-w&t=5s - AI: https://www.openai.fm/#8e9915b0-771d-4123-8474-78cc39978d33 https://www.openai.fm/#8e9915b0-771d-4123-8474-78cc39978d33
- khurdula 2y agoWhat if I said, we outperform them? Check this out: https://jigsawstack.com/blog/openai-audio-stt-vs-jigsawstack-stt https://jigsawstack.com/blog/openai-audio-stt-vs-jigsawstack...
- garfieldnate 2y agoMost of the voices give "I'm sorry, I can't assist with that request" when I try to input Japanese. They all work for German and Spanish, though.