10 ms·
GPT‑Live
- ed_mercer 3mo ago>Delegation for deeper work Except that you can't actually delegate since connectors and tools are not supported.
- sonicslayer 3mo agoThis solves my biggest annoyance with the current advanced voice: its speech getting interrupted by me setting yup or even background noise if loud enough
- deleted 3mo ago[deleted]
- piskov 3mo agoNow it consistently interrupts you
- jrflo 3mo agoYeah, looking at the waveforms from those sample conversations it looks like it's happening at the exact same interval every time
- hirvi74 3mo agoI've had some funny interactions with this issue. I sometimes use the voice mode when walking my dogs. I can confirm that ChatGPT responds positively to being told it's a "Good girl!"
- deleted 3mo ago[deleted]
- OsrsNeedsf2P 3mo agoHoping to use this for natural conversation language learning. Previous iterations of the app kept correcting my words/grammar before it got to the model, causing issues with identifying mistakes in speech
- MostlyStable 3mo agoI was going to post a comment on a related topic (I couldn't find in the announcement if this is English only or not), but would you mind expanding? I was thinking about doing something very similar, and yeah, if the model isn't hearing the mistakes I'm making, that would dramatically decrease it's usefulness.
- haaz 3mo agoI tried it for these purposes and it didn't seem any better. I couldn't even tell it was a new model. I tested today, I'm an EU customer so maybe the rollout is delayed
- rvz 3mo agoWith this, human translators have been totally and absolutely a solved problem with this version of real time translation. This time is the most natural version that exists and it is a natural as a conversation. To Downvoters: Why aren't you feeling the AGI?
- tomasphan 3mo ago“I don’t translate, I interpret” - Ahmed the best interpreter there ever was
- slekker 3mo agoWhy aren't you feeling the operating loss? :(
- progval 3mo agoThey most definitely did not solve real time translation yet. The French in the video is barely understandable, both the translation and the pronunciation are the quality of an American who hasn't used French since high-school.
- HyperL0gi 3mo agoVery cool. Not cool bringing Brazil’s loss to Norway again. We're already devastated. No need to keep beating someone on the ground. :(
- ilaksh 3mo agoAre there any open source full duplex models that are out besides PersonaPlex? There was a chinese open one, maybe Fun Audio chat or something, that said it was going to release a full duplex version but I am not sure if it did. My dream would be open source full duplex with function calling or some kind of rudimentary text output. PersonaPlex is still interesting although it was looking like we would need to fine tune it to handle outgoing or avoid going off the rails easily.
- VladVladikoff 3mo agoI want to know this too, as I’m hoping to fabricobble up a “smart speaker” that communicates with my local AI assistant. Right now we do everything via iMessage but it would be nice to be able to tell it to add things to my grocery list by voice while my hands are busy in the kitchen. Also would love if anyone has any advice on what microphone & speaker to pick out, was planning on just reusing a raspberry pi I’ve got around for the brain part.
- ilaksh 3mo agoif you don't find that then you could fake it with personaplex possibly but making another ASR/STT model just listen continuously and transcribe then send to an LLM with function calling. at least that would allow one direction easily.
- Rebelgecko 3mo agoIf you're in the Home Assistant ecosystem, I'm intrigued by their voice hw
- andersthuesen 3mo agoWorking on this at https://duplexio.ai https://duplexio.ai which will be open weights and free for non-commercial use. If you want to work on this send an email to anders@duplexio.ai
- programjames 3mo agoStandard Intelligence released one two years ago: https://si.inc/posts/hertz-dev/ https://si.inc/posts/hertz-dev/ It's only 8.5B and doesn't sound like it's quite conversational.
- vessenes 3mo agoOh wow, I'd like this. Our current voice interactions with ChatGPT are on a 4o era model; really terrible. oAI has always been pretty cagey on the architecture of their end to end multimodal models. And RL has basically made them worse since launch. (Check the launch videos where the model sings, is more realtime, has accents, etc). I'd love to try a next gen version.
- observationist 3mo agoThe potential conversational dynamics of people telling each other "quiet!" after they pick up the habit from talking with AI will be interesting. It could lead to people being more assertive and thoughtful, or it could be contentious and rude. Awesome that they've improved that aspect of voice chat, though.
- rane 3mo agoAbsolutely can't wait to try this for language practice. The advanced voice mode is great but ultimately just doesn't work that well and doesn't have the feel of a natural conversation.
- fnikacevic 3mo agoAny pricing announced yet?
- raychis 3mo agoThis looks very cool. An AI that can listen and speak and handle tasks without breaking the flow of conversation would solve some big annoyances with current tools. The concern is though as these get better will people struggle to distinguish these with real human connections?
- bstsb 3mo agopeople have been mistaking AI conversations with reality since the very first text-based models came into the public view with ChatGPT. i'm sure with each incremental improvement to outputs like this, though, more people will get convinced of its "humanity" (see https://www.reddit.com/r/MyBoyfriendIsAI/ https://www.reddit.com/r/MyBoyfriendIsAI/)
- ACCount37 3mo agoEvery time things like this come up, I can't help but think of the ending of Inception. It's less that you're convinced it's real and more that you no longer care if it is. "Feels real enough" is good enough. I'm a technical user first, so I'm not sure if models have improved for RP the way they improved for applied STEM tasks and technical brainstorming. But if there is an improvement curve there, I wouldn't be surprised if this only grows in popularity.
- raychis 3mo agoYou know, when I wrote that comment I was actually thinking of the ending of Inception. Still makes me very uncomfortable though.
- OtomotO 3mo agohttps://en.wikipedia.org/wiki/ELIZA https://en.wikipedia.org/wiki/ELIZA
- HyperL0gi 3mo agoI'm very eager to test this for brainstorming! One thing I noticed is that we lost vision feature for some reason on the live chat? This was an extremely useful feature. Not sure if it’s a regional thing or that they just removed that from the current live chat. I imagine it will be even more useful with this new version.
- athyuttamre 3mo ago(Atty from OpenAI here) GPT-Live does not support video at this point, but we're working hard to introduce it soon. In the meantime, our previous Advanced Voice Mode will continue to be available and supports video.
- HyperL0gi 3mo agoThanks for the reply, Atty. So I guess I'm holding it wrong? For some reason, I can't use the camera while in Live mode. The only option I see is the plus item, which does show the camera, but when I open it up and ask "Are you seeing my camera?" it will always say no and recommend me to open it. Feels like the official camera icon does not show up for me? iPhone 13, ChatGPT Pro subscription.
- athyuttamre 3mo agoAh yes — the camera icon is to take a photo, not to show the model video in real-time. Try taking the photo and uploading it! If you're having trouble, please share a screenshot and the build number (Settings > About) with me via email at atty@openai.com.
- vjulian 3mo agoDoes this support more than one user voice? Or, are there plans for this? I did not see that mentioned in the announcement.
- athyuttamre 3mo ago(Atty from OpenAI here) You can choose among 9 voices in the app, all newly refreshed for GPT-Live. If you meant whether it can detect multiple people, it can (like in the livestream), but not always perfect. Would love to hear your feedback once you try it.
- robotswantdata 3mo agoAny plans to add more international voices and dialects?
- smalltorch 3mo agoVery cool. I thought the agent came in a little to hot at 1:03. I wonder how it decides when to jump in.
- fraywing 3mo agoI like this and felt like some of it was much more fluid; but was I alone in feeling like the interjected "uh-huh" or "yeah?" moments felt a little jarring? Almost felt a bit *uncanny valley* for what "natural" conversation is supposed to be like. If the "uh huh" isn't timed correctly, it'll feel like a zoom call with lag.
- burntalmonds 3mo agoI agree. Many times those little interjections don't feel natural. It's impressive, but there's still a lot of room for improvement.
- AaronAPU 3mo agoEvery second of the interaction is uncomfortable to me, but I also have extreme difficulty with video calls with humans. The latency completely breaks my mind.
- SoftTalker 3mo ago[dead]
- simonw 3mo agoI had preview access to this one for a few weeks. It's very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects. The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier. I did report a fun bug with it though: it was interrupting me and laughing at my (not really intended as) jokes while I was still talking! They seem to have clamped that behavior down thankfully, it felt a bit rude and condescending.
- heisgone 3mo agoIs it a dumb-down version of GPT like the current voice model? At least in french, I find the current GPT voice mode to be useless, to the point I only use the dication mode. I would ask a question and it would answer something along "That's a interesting question. I can help you with that. Anything you want to know about X?" I would ask again and it would answer the same kind of non answer.
- KaoruAoiShiho 3mo agoClick through to the link, the answer is no it uses the latest gpt models now.
- magicalist 3mo ago> the answer is no it uses the latest gpt models now. Actually it says it _can_ delegate to the latest models. Seems reasonable to ask how the voice model does when it doesn't delegate (or while waiting for the delegated answer).
- derefr 3mo agoI haven't stress-tested it, but I would imagine it approaches complex problems the same way a human with a phone in their pocket would — that is, by having a degree of awareness of the confidence it has in its own knowledge in some areas; where, when it "realizes that it doesn't know", it blocks the conversation with statements like "I don't know, let me check." I say this because this is already how ChatGPT works internally when using its "auto" mode; the version of the "fast" model used in the "auto" mode does the same "notice your ignorance and bring in the heavy model" thing, just silently, rather than mentioning that it's doing it. (If someone has actually run the experiment, please chime in!)
- redox99 3mo agoDefinitely in the right direction in terms of architecture. However those "hmmm" "uh huh" interjected in the demo are pretty awful.
- altcognito 3mo agoI was hopeful that they avoided the well known sultry voice this go around, but alas. There is little hope for these companies. The full duplex is awesome, and the feedback that it is getting what you're saying is ok, but in some of the demos was a little overkill. I'll agree that using the "Golden Girls" was at least more entertaining than the usual pitch.
- savanaly 3mo agoYou have always been able to pick between voices of many kinds though? Do you find them all sultry? From the British woman to the 17th century pirate soundalike?
- Jtarii 3mo agoIf there was a voice that stripped away all the affectation I would be more likely to use it. It pretending to be human is extremely off-putting for me.
- spongebobstoes 3mo agotell it "speak like a robot without affectation or emotion" in your custom instructions
- HardCodedBias 3mo ago"I was hopeful that they avoided the well known sultry voice this go around, but alas" Why do you care? You can select other voices. Why do you need to control others? What is at the root of your need for domination?
- altcognito 3mo agoThey are stealing/imitating someone elses brand. Amping up an emotional connection is great for business. Why do you think THEY need to dominate via an emotional connection?
- dogscatstrees 3mo agoI do not fully understand the complexity behind achieving full-duplex but I hope this sets the bar for Anthropic to follow. Turn-based simplex is yesterday.
- artdigital 3mo agoWhat I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an app or service. It’s weird. The old Claude voice mode WAS able to use tools but when they revamped it, it lost that capability and is now pinned to Haiku :( So, yay for finally a voice mode that’s powered by a frontier model and hopefully as good as Grok voice, but sad to still not see tool use while in voice mode. (I haven’t tried it yet, only read the announcement)
- AaronAPU 3mo agoI could see it relating to tools having unpredictable latency but if they already do background hand off to 5.5 then it seems like they could just enable it within that context.
- codybontecou 3mo agoLast I tried this is exposed via their sdk and can be built.
- atonse 3mo agoI’ve been using gpt-realtime-1 for my personal assistant that runs my company and plans my day. And it works pretty well, even makes tool calls and all that. But the multi modal stuff has resulted in a lot of debugging with weird events and message and audio sequences having race conditions, but overall it is pretty awesome. Looking forward to moving to this model later today and will chime back in with results.
- atonse 3mo agoUnfortunately, it isn't available on the API yet! Can't wait!
- paxys 3mo agoBecause they take too long to run, and have an unpredictable latency and success rate. Seeing loading spinners and error messages in a visual interface is fine, but it would firmly put a natural language conversation in uncanny valley territory. Regular chat already supports voice input, so might as well use that.
- zuzululu 3mo agowatched the live translation video very impressive Seems like a shift from previous voice models where it sequentially processes voice to text then feeds it to LLM and then back which cant escape the clunky lag not sure how pipecat stands now, gpt live seems like it takes audio tokens and does inference on it directly
- trollbridge 3mo agoI've wondered if this would happen, although doing inference directly on speech tokens would seem to imply an entirely different model (trained on lots and lots of actual speech).
- athyuttamre 3mo ago(Atty from OpenAI here) GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction. Would love to hear your feedback!
- worldsavior 3mo agoWhat made you to try again?
- famouswaffles 3mo agoDoes video/image input still work with these duplex models?
- athyuttamre 3mo agoImage input is supported, but video is not today. We're working hard to bring it to you soon.
- dandaka 3mo agoCan I connect it to my skills/tools? Example case, I have a knowledge base and event log in my company. I need a brainstorm companion, which will have full access to this knowledge, can converse about it and can invoke skills/tools available in the repo.
- athyuttamre 3mo agoIn ChatGPT, Voice doesn't yet support connectors, but we're hoping to add support soon! Once GPT-Live launches in the API, you can also build custom integrations yourself.
- BoorishBears 3mo agoDo integrations supporting streaming input? One big gap I've run into for UX is most realtime voice harnesses wait for a full response from tools, and at most support the model filling the dead air until then It'd be a game-changer to be able to have the model start replying with partial information streamed from the tool call, then seamlessly continue with additional information.
- joshmarlow 3mo agoI'm so mad that this might make me re-subscribe to ChatGPT. I wouldn't have believed how much I use the voice feature before LLMs and ChatGPT currently has the best voice interface. I think Grok's interface is the next best, then Claude.
- ksd482 3mo agoSame. I might switch back to ChatGPT from Gemini because I use the voice feature all the time. One of my favorite use cases is talking with it while driving on random topics and learning about them.
- joshmarlow 3mo agoAbsolutely the same. Now that Fable is back, the Claude voice interface is... worth dealing with. The app mostly seems to have trouble recovering from networking issues which is jarring in a deep conversation.
- jablongo 3mo agoBut I don't think Fable or even Opus are ever used as the backend in voice mode. It has to respond in real time so I think in voice mode it's always using Sonnet.
- joshmarlow 3mo agoThat makes sense and had never occurred to me! I just asked and Claude says it's Haiku.
- misiti3780 3mo agosame, i use the voice feature every day at the gym, talk to it for 1 hour, and then make anki cards based on what it has taught me, total game changer.
- wraptile 3mo agoJust wait a few months and this will likely be available on a different LLM platform, especially if it actually works.
- JasonSage 3mo agoI for one am greatly looking forward to the day these kind of voice models can be run locally. It seems like the gap between open-weight and frontier is way larger for voice models than coding/language models.
- jakswa 3mo agoProbably not _as_ good, but you can run gemma 4 for the ears/brain (accepts audio input), and kokoro TTS for the mouth. You want something like silero VAD sitting in front of the LLM so you aren't passing wasteful audio to gemma 4, so you send only voice activity segments. I type all this because I tried it out recently as a Zork experience, seeing how creative Gemma 4 12B could be. It was surprisingly good!
- JasonSage 3mo agoYeah, I mostly mean for the TTS, which in my testing is flat and has no emotional register, or no nuanced variation. OpenAI speech in my experience is very good at things like communicating understanding or questions non-verbally using intonation, and I'm yet to see a local TTS model that will do contextual intonation well at all.
- bfeist 3mo agoThere doesn’t seem to be any indication whether this is available in the chat-got app nor is there any indication in the app that anything has changed. Anyone know how to actually try this?
- dbbk 3mo agoRead the post
- meetpateltech 3mo agoGPT‑Live is rolling out now to ChatGPT users globally across iOS, Android, and ChatGPT.com. GPT‑Live‑1 will become the default model powering ChatGPT Voice for Go, Plus, and Pro users, and GPT‑Live‑1 mini will become the default for Free users.
- athyuttamre 3mo ago(Atty from OpenAI here) We're beginning the rollout now, and will roll out in the next few days to ChatGPT users globally. Make sure to update to the latest version of the app!
- no-name-here 3mo agoI'm using the current iOS app. In the voice settings, there were 2 options - Standard and Advanced, as well as a help link saying they were in the process of rolling out the newest model -- I think the new model is named Live (and is not Standard or Advanced). The help link explained it all. So at least for me in the app there did seem to be a way to check.
- dbbk 3mo agoIt sounds like they've switched to a "native audio" model which if I understand right is what Gemini has had for quite a while?
- ZeroCool2u 3mo agoGemini live has been able to do this for over a year now. I can just activate it on my phone and it really works surprisingly well, especially the interruption. I've tested it with my 95 year old Dutch grandmother and it switched seamlessly between English and Dutch with her and handled her poor hearing very well, including her asking for repetition. I'm a little surprised by how much OAI is playing catch up here.
- scosman 3mo agoChatGPT has had this for over a year too. This is a better model.
- davedx 3mo agoIs gemini live the same app I have on my oneplus 15? I don't think that's full duplex?
- spongebobstoes 3mo agoit is not full duplex. I think this GPT-Live thing is the first full duplex speech to speech model
- rryan 3mo agoehh, no? definitely not. https://kyutai.org/blog/2024-07-03-meet-moshi/ https://kyutai.org/blog/2024-07-03-meet-moshi/
- TaupeRanger 3mo agoWhat do you mean? OpenAI has had a real-time voice model since August of last year. This is a new model with better performance.
- thewebguyd 3mo agoGemini live is also pretty decent at computer vision stuff too. I've used it while working on my bike & car a few times, with my phone camera and it'll circle/highlight specific screws and parts. You can set your phone up on a tripod and it'll walk you through complete repairs for things.
- mlmonkey 3mo agoIs it possible to create a "companion" of sorts with this model, using, say, an RPi and a speaker + microphone? Not for advanced scientific brainstorming, but for seniors who are often alone in their homes.
- 100ms 3mo agoIf this is my idea of hell, I can't imagine what it'd be like for a 90 year old
- deleted 3mo ago[deleted]
- mlmonkey 3mo agoTry being a 90-year old with minimal social contact and nobody to talk to, wasting away in front of a TV blasting NewsMax or Fox News ... does that sound like Heaven or Hell?
- 100ms 3mo agoSubstituting one form of media for another does not improve the hypothetical situation at all, giving a senior nothing to do all day except talk to a spreadsheet is not a solution to that senior being lonely, it is a false dichotomy. Far lower tech solutions can be (and are) used. In my locale we have volunteer schoolchildren active in the community, as far as I know that's quite common, and even half an hour biweekly with a fake grandkid is many orders of magnitude healthier and more meaningful than swapping out their TV for this generation's take on Fox News. It goes without saying all these tools are still largely in their pre-advertising state, it won't last.
- mlmonkey 3mo agoI speak from experience. The last few weekends I've been taking a friend to see her grandma at a Senior Facility in Sonoma County, and it is really depressing to watch them either lying in bed, listlessly, or in a wheelchair, gazing away in the distance. They need interaction. And a suitably prompted LLM _can_ provide such interaction. I'm not saying hook them up with ChatGPT and let them loose, but that with the right harnesses and guardrails, they could have a more interactive life.
- moralestapia 3mo ago>GPT‑Live can show it’s paying attention with phrases like “mhmm” or “yeah” [...] Nooooooo!
- sixdimensional 3mo agoYeah, I mean - I know that people do this in real life, but when I encountered this recently while trying GPT-Live, it was truly annoying - I felt like it kept causing me to lose my train of thought. Perhaps if it was a bit more adaptive or could even be turned off.. but if it does this every.. single... time... I am saying something and I pause for a second... ugh.
- cryo32 3mo agoThis feels so dehumanising.
- softwaredoug 3mo agoDoes this model do better ignoring side conversations? That's the biggest hindrance to using ChatGPT's carplay feature is someone will say something, stopping ChatGPT from speaking or taking it in a different direction.
- athyuttamre 3mo ago(Atty from OpenAI here) Yes! GPT-Live is much better at ignoring background noise, including other people speaking. Not perfect, but you should feel a big difference.
- juberti 3mo agoyes, it now has much better ability to understand the conversation and decide whether it should respond (and you can also tell it when it should respond)
- kindawinda 3mo agoHello Justin! Mr webrtc himself and the infamous AIM 5.0
- zelias 3mo agohow about api access?
- athyuttamre 3mo ago(Atty from OpenAI here) Coming soon! Sign up to be notified here: https://openai.com/form/gpt-live-1-in-the-api/ https://openai.com/form/gpt-live-1-in-the-api/
- I_am_tiberius 3mo agoOh gosh. I was watching the video and thought it is a live stream. I just noticed when it restarted.
- wxw 3mo ago[dead]
- sampton 3mo agoI want this in my ear when I'm talking to people, so I can carry a real conversation.
- small_model 3mo agoThis had to land before there new device could be launched, i.e. human to AI full duplex interaction, Apple should be worried. They fumbled so hard on AI.
- SirHumphrey 3mo agoI am wondering for whom this new device even would be. Because phones work quite well for this use case and much more and everyone already has a phone.
- small_model 3mo agoCan a phone go get you a coffee?
- joshstrange 3mo agoHopefully this also means it does better with "interruptions". I have used ChatGPT Voice in the car before and sometimes car/road noises will cause it to stop responding in the middle. I used it to help me figure out how to turn off a feature in the rental car I was in (adaptive cruise control, I love it but snow blocked the sensor and I wanted just normal cruise control but couldn't figure it out while driving). This kind of voice chat is awesome and I'll be even more excited when open models have this functionality. I'd love something like this paired with Home Assistant (assuming we ever get decent hardware).
- athyuttamre 3mo ago(Atty from OpenAI here) Yes, GPT-Live is much better at ignoring background noise! I use it every day in the car via our CarPlay integration.
- jdudek 3mo agoThis is what Siri should have been.
- tiffanyh 3mo agoIt was probably the original intent of OpenAI working with Apple ... but clearly there was a reason why Apple ditched OpenAI for Google/Gemini models instead.
- eboy 3mo ago[dead]
- jonstaab 3mo agoThis is the opposite direction AI should be going. Human relationships are the most valuable thing we have, and so, naturally, technology seeks to intermediate and now replace them. I'm not Catholic, but this podcast presents a very interesting argument against talking to AI as if they were human: https://newpolity.com/podcasts-hub/debate-chatbots https://newpolity.com/podcasts-hub/debate-chatbots
- zkmon 3mo agoMost of what AI does is already in wrong direction. Not just human-to-human interaction, it took away thinking, creative work, sensory perception (glasses) and responses. People call it as helping humans, but I call it as sucking away the "human-ness" from humans. After the damage is done, the mega corps would simply shrug and will say "Well, we were just responding to our business competition" The business knows no human-ness, because it is not a human. Businesses and machines are creatures that see humans as their fodder. And humans created these, assuming it is progress, to have businesses and machines. We call it progress because it required our mind power and it helped us to dominate other species. Dolphins are laughing at us.
- amarant 3mo agoI don't understand this line of reasoning. How are you hindered from doing any of those things? What part of "AI can now do X" makes it so you can't also do X?
- zkmon 3mo agoI'm no longer writing code from scratch, as I used to do before. So, very soon, it will be "AI can now do X" makes it so I can't also do X? Same with many creative works. Radio music already sounds so plastic. I lost interest in crafting my text drafts because I can just dump some ugly text and get it refined by AI.
- small_model 3mo agoYou no longer have to grow your own food, can go to supermarket, but doesnt stop you from still doing it.
- HardCodedBias 3mo agoWhile this is likely very useful to an enormous number of people, I suspect it will be even more useful for the elderly (if somehow it can be made accessible to them). IIUC the literature, there is serious loss of functionality associated with lack of verbal interaction. People can say "they should just talk to more people" or "more people should make time for them" but the fact of the matter is that it doesn't happen, and if this helps terrific.
- h1fra 3mo agoGPT-Live being developed in california and being an over-active listener...
- londons_explore 3mo agoThe demo video shows quite how rough around the edges this is.... Doesn't quite stop fast enough when you interrupt it. Can't find info quick enough so you have to change topic and then have it give you results later, etc. This is a move in the right direction, but there is lots of engineering still to be done!
- TaupeRanger 3mo agoYou're right, and to me it's refreshing to see a promo video that shows how the real product works, rather than a sanitized over-produced edit that takes out all the flaws.
- xtracto 3mo agoI'm currently watching Better Call Saul for the first time with my wife. The fact that they used old ladies for the Ad, and that it is evident they are reading from a prompt (fake-ish feeling) gives me strong James McGill vibes haha. Hopefully the actresses were paid handsomely.
- athyuttamre 3mo ago(Atty from OpenAI here) >This is a move in the right direction, but there is lots of engineering still to be done! Could not agree more. We see this as the first version of a new generation; expect many improvements in the future.
- aristocrazy 3mo agoAs someone with a chronic eye disease that makes screens difficult, I love this voice feature. It gives me hope that, even if my disease progresses, I’ll be able to rely less on screens and screen readers for many everyday tasks in the future.
- bearjaws 3mo agoFeel like the intro video is very odd. Basically have an older lady (not their target audience) blatantly reading a teleprompter. Why are they going after this audience? Retired people have no use for delegated tasks or information. They also are the least likely to use it and not get frustrated.
- cute_boi 3mo agoIs this website heavily vibe coded. I tried to select text and things went black. https://imgur.com/a/ABGWRTO https://imgur.com/a/ABGWRTO
- 6thbit 3mo agoThe new architecture makes sense, it seems many of the remaining problems like noise and interruptions are at the sound processing and integration level rather than at an architectural or model level now which makes for an exciting new era.
- cahaya 3mo agoWhen GPT-Live in Codex, so i can walk the dog while shipping?
- evtothedev 3mo agoFor the first actor - why does her accent change the longer she talks. It's like they had an "Estelle Costanza" dial that they started at zero and slowly rolled up to 8 or 9.
- davmar 3mo agoIt changed because she came home to find her son treating his body like it was an amusement park
- jdmoreira 3mo agoGreat video by the way. Extremely good!
- consumer451 3mo agoOne piece of feedback on that to anyone from their team, could not get audio sync in Firefox. Chrome worked fine.
- csto12 3mo agoCan a model like this critique your accent/pronunciation? That would be cool.
- victor9000 3mo agoI'm at the point where pricing is the first thing I look for in announcements like this.
- danjc 3mo agoIt would be great if we could have AI that wasn't trying to emulate a human. When it expresses emotion, we should see that as a bug that needs to be fixed.
- mrcwinn 3mo agoFantastic to see this. I use voice a lot. It’s not quite lived up to my expectations but I think this gets much closer.
- drusepth 3mo agoI reaaaaaally hope we have an option to disable those random ums and ahhs that interrupt for no reason. :|
- consumer451 3mo agoI was surprised all that made it into the demo video. Did not seem cool. Is that "active listening" or something?
- xpct 3mo agoI think it just doesn't know when it's its turn to speak, and cancels itself out when it hears humans.
- stri8ted 3mo agoIts actually useful, when you are launching into a long monologue and want periodic acknowledgement that its "listening".
- consumer451 3mo agoI think it would make me stop speaking. I guess it might take me some time to get used to it.
- consumer451 3mo agoI have now had a chance to use it, and it never interrupted me. Weird call for the video to include that. Works very well, and can even mash up languages.
- rcarrol6 3mo agoI'm crying, that guy who tries to get it to count to 100 may actually have a chance
- throwaway613746 3mo ago[dead]
- nadzzz 3mo agoForget the fancy scores, the real test is whether it'll just let the guy count to 100.
- nico1207 3mo agoI just told it to count from 1 to 100, and it actually correctly counted and didn't leave out any number. I'm impressed
- rcarrol6 3mo agoDoing the work that matters
- throwaway613746 3mo ago[dead]
- surround 3mo agoHow does the voice model delegate requests to GPT-5.5? Can the voice model generate text?
- djb_hackernews 3mo agoare these human actors or are is the whole demo AI generated?
- mixel 3mo agoI was missing the part where the grandmas are saying "I have absolutely no idea what I just said" :D --- Besides that really nice demo I will give this a try, I tried some voice models before and the issue is I will ask a question, get a answer withing the next ~1 sentence and then the usual LLM bs follows which I just wanted to skip at that point
- charcircuit 3mo ago>and linked parents may be notified in higher-risk situations involving signs of potential self-harm or suicidal intent. This is an abuse of user trust and violates people's privacy.
- nakedneuron 3mo ago"i'm starving." sounds cynical in my ears. energy demand of these toys will cause many problems, people elsewhere starving being one of them.
- sumoboy 3mo agoSome brilliant marketing really, the grandma test.
- xpct 3mo agoA cool test, until it hit's the my family's grandmas test :)
- fuddle 3mo agoI looks like they took inspiration from Thinking Machines - http://thinkingmachines.ai/blog/interaction-models/ http://thinkingmachines.ai/blog/interaction-models/
- BorisMelnik 3mo agothis is excellent, I've been meaning to update my phone-dialer.apk "fake phone conversation" for when I'm trying to get out of a social sitation (not joking)
- croes 3mo agoSo more AI psychoses coming.
- jbonatakis 3mo ago[dead]
- programjames 3mo agoDidn't Standard Intelligence release a duplex model two years ago? Sounds disingenuous to market this as a new generation of voice models, when it is really OpenAI finally catching up to the current generation after two years. https://si.inc/posts/hertz-dev/ https://si.inc/posts/hertz-dev/
- neko_ranger 3mo agoLooking forward to a list of support languages. This would be amazing for language listening/speaking practice. Yes I know it doesn't replace the talking to real people and "immersion". But cheaper than a flight
- gotrythis 3mo agoLast night, I was using voice for the first time in a few weeks, and it interrupted me and said, a bit aggressively... "I'm going to stop you right there. Let's keep the conversation focused on the topic we were covering or a new relevant topic". I tried to probe it for why it did that, what rules it was following, and it eventually told me... "My role is to keep us focused..." and, "The behaviour you saw was my attempt to moderate tone". I've heard of LLMs doing weird things like this, but it was the first time it happened to me. I hope they fix that. It was creepy. For context, it heard my partner say, "I guess it's the same thing as you mom, because she's..." and then it cut us off.
- WhitneyLand 3mo agoMuch better than it was before but it’s still significantly weaker than a direct chat. For example I asked “Why should LLM attention use dot product instead of cosine similarity, being that we often hear vector magnitude does not encode most of the useful information needed”? The voice response was directionally right but lacked detail and was a little hand wavy. The answer to the same question in a text chat was much higher quality. The voice response replied “let me think about that…” so it appears to be invoking 5.5 as advertised, but it’s definitely weaker. I had reasoning set the same for both.
- WarmWash 3mo agoI would naturally assume they cut the background thinking level to the minimum. It's a halting problem question and with in text 5.5 with max thinking will chew on a question for 5-10 minutes sometimes. That would make a pretty awkward conversation.
- athyuttamre 3mo agoA lot of this is steerable! Our current personality is optimized for brainstorming and conversations, but you can provide custom instructions to ask it to go deep and give you info-dense or more technical answers.
- bariswheel 3mo agoI found myself probing more and more questions to have it do that, I wish it could remember to provide details answers, I thought the replies were way too short and hand-wavy like one of the prior comments suggested.
- WhitneyLand 3mo agoThat would be cool, but is there a way to differentiate between voice mode and text only output in the custom instructions? Otherwise, it seems like preferences could be competing.
- cnxhk 3mo agoYou can read far more text than listening. If the voice response is too long, people lose the patience quickly. So it is better to show it on screen if we have too much text.
- tills13 3mo agoI use AI for my job. I understand the impact (as much as anyone can) that it will have on society. I recognize the value. But I just want to say that talking with AI casually is critically lame. I cringe every time I have to ask my Google Home to turn on the lights and people are having full-on conversations with it? And, imo, dangerous considering how sycophantic AI is. The stupidest, most gullible, most insecure person right now is looking at this thinking they are about to make a new friend.
- hersko 3mo agoI wonder why they don't compare it to their existing live voice model realtime-2.
- sneak 3mo agoWhy do I have to solve a captcha to read a blog post?
- skilled 3mo agoNot had a chance to experiment yet, but will this be an upgrade for using AI for language learning?
- Martinussen 3mo agoThis voice is awful, possibly one of the worst AI/computer voices I have heard in like two years now - what's up with that? Is this seriously the best a company burning this much money can do, and they consider this acceptable to release? Like two syllables in and my first response was to grimace and physically cringe. Does anyone here think this sounds good, or even just "fine enough"? Do engineers that work on this for long periods of time stop seeing the forest for the trees and think this could be mistaken as human? I'm saying this as someone that assumes all narration/most VO work will be fully AI fairly soon (for better or worse.)...
- timpera 3mo agoWhat don't you like about the voice? Geniune question, I'm not using it in English but I think it sounds fine in my language, and it's an improvement over the previous one.
- Martinussen 3mo agoIt's just... Very robotic and clearly generated, makes random inhuman pitch shifts, etc.. Very alien overall.
- iknowstuff 3mo agoThere's like 10 voices to choose from.
- wewtyflakes 3mo agoGreat to hear about full-duplex. When using voice mode historically, it was infuriating to have the AI go on a long-winded rant or explanation and I would be shouting again and again "stop. shut up. shut up! shut up!!!"; I just needed a clean way to interrupt it.
- tracekl 3mo agoIf an idiot has this on in the subway my conversations are surveilled. What is the antidote? Train another model to talk about bombs etc. and flood the clanker (and by extension the FBI)?
- lrvick 3mo agoI cannot wait for the qwen version on huggingface.
- vinay_ys 3mo agoI have not used voice mode much with chatgpt. I was surprised to learn that they were already not running the voice model like a UX orchestrator while utilizing other models in background for actual research/response etc. I guess it's good they launched what they could and got here in steps. I suspect in the near future my personal device (mobile/laptop) will be powerful enough to run any UX orchestrator model locally – and route to multiple frontier closed/open model providers in the background as appropriate. The battle is going to be platform owners (Apple/Google/Microsoft) wanting to lock-down the access to that local hardware and local interaction paradigms (ambient always-on full-duplex voice) and intermediate through their platform layers - rationalizing it as consumer security/privacy protection (which is right for most people, but sucks for the open market). Meanwhile I suspect OpenAI/Meta et al will try to build their own hardware and become platform owners themselves, though unsuccessfully. And it's going to take some company like epic games to get them to open that up. and that's probably what the next decade is going to be all about.
- yottamus 3mo agoI worry what this will do to human communication if it becomes commonplace. Will everyone learn to be a forceful speaker, speaking over anyone they want to stop speaking?
- taurusnoises 3mo agoI don't have many opinions about how individuals use this tech (although the AI as friend trend is a bummer for many reasons), but have maaaany (negative) opinions about the customer service industrial complex that's already using this in what seems to be an attempt to fool people into thinking they're speaking to a real person. Which is why I now, like a freak I never thought I'd need become, always ask "Am I speaking to a bot or a human?" when dealing with CS. So far, it's worked, and the bot transfers me. But, I fear the bot will eventually be programmed to lie abiht that, as well.
- modeless 3mo agoIt was obvious from the live demo that this thing still hasn't learned when to shut up. When it stops tacking on "I'm here when you need me" to every response that could have just been "ok" or simply silence, maybe they'll have something. I think voice remains OpenAI's most disappointing product.
- overgard 3mo agoLike a lot of AI things, this seems both cool and kind of creeps me out. I've never used voice interfaces in the past (siri, the google one, whatever is on my tv) so I'm probably not the target market, but this does seem like an improvement. The part that creeps me out is, we're living in an era where we're more disconnected from each other than ever before. Do we really need to be replacing conversations?! The demonstration video of old ladies sort of hints at something for me, which I think we already have a societal problem with the way we treat the elderly (and a massive elderly-loneliness issue) and there's kind of a sadness of imagining people becoming really close with this machine that doesn't really think. Definite ick factor.
- rohansood15 3mo agoI feel similarly. But texting became the primary form of inter-personal communication in the last couple of decades, in part because that was the primary modality technology could handle. So, now that we can talk to computers, maybe we will feel more comfortable talking to people too? Wishful I know, but one can hope right?
- bottlepalm 3mo agoNone of the things I'm talking to ChatGPT about are replacing human conversations. It's replacing having to type a question into Google on my phone. But you do bring up a good point with old people with no one to talk to. This technology would be a god send for them. And if you think otherwise I hope your next words are that you routinely visit nursing homes and talk with old people.
- overgard 3mo agoUm, your argument is unless I'm directly involved with something I'm not allowed to critique it? Here's a real story that's very much more likely story than your "godsend": https://www.reuters.com/investigates/special-report/meta-ai-chatbot-death/ https://www.reuters.com/investigates/special-report/meta-ai-... I don't think giving old people AI psychosis is anything other than inhumane.
- 3mo ago
- larrik 3mo agoHaven't tried this, but talking to Claude in its app is so much better than talking to Siri that Apple should be ashamed. It got every word perfectly the first time, including programming / project management terms. Meanwhile, Siri struggles to send basic texts to my kids.
- motoboi 3mo agoOh, god. They are marketing it as an old people artificial friend. Probably will blame users in the future when they get attached and have ai-induced psychosis. Disgraceful. On the technological side, it's a marvel!
- cactusplant7374 3mo agoMany older people are very lonely. This could help a lot of people.
- odie5533 3mo agoI'm not convinced wireheading the elderly is a good solution to their loneliness.
- PUSH_AX 3mo agoIf you have any ideas now's the time, this has been an unsolved problem for hundreds of years.
- zulban 3mo agoIt doesn't have to be good. It just has to be better. What we do now is completely abandon and ignore most seniors. The bar is low. Don't like it? What's your better solution? Are you going to do it?
- lbrito 3mo agoThat's kind of like the joke about MAID solving the Canadian healthcare problems: I've heard you're ill; have you considered dying?
- himata4113 3mo agoI think this is fine as long as the models stay honest. I.e. refuse to be a friend, girlfriend, partner or pretend to be someone.
- franky47 3mo agoThe live translation demo reminds me of the Babel Fish in The Hitchhiker's Guide to the Galaxy. This could be very valuable to have in your headphones while travelling. Not super impressed by the model constantly interrupting the user in the other demos though.
- barnacs 3mo agoThis is getting way too dystopian for my taste. People in the know need to stop pushing the narrative that this is somehow anything more than statistical autocomplete.
- dzonga 3mo agoif I'm not mistaken - Amazon Nova Sonic has been full duplex for a while.
- programmertote 3mo agoI watched the demo video. Isn't that agent's voice too hasty in responding? Maybe that's what they (OpenAI) are trying to show off as full-duplex tech, but I can't shake the feeling that I'll feel annoyed if the AI agent interrupts me when I'm speaking....
- dmje 3mo agoAnyone else find the use of the ladies in the videos to be pretty patronising? Or just me...?
- CrzyLngPwd 3mo agoHer is here?
- ralusek 3mo agoI have built a few voice based integrations into my applications that use these live agents (gpt and gemini), but they are always too expensive to be viable. I have to end up hacking up context and turning on and off in ways that are very fragile. It'll end up being $2-5 for about the 30ish minute sessions I typically end up with, and it throws the price of the product I'm making completely out of whack.
- lbrito 3mo agoGPT, now with more interruptions!
- JimsonYang 3mo agoI was expecting the grandma voice to be the voice model and i was like woah this is incredibly good
- bariswheel 3mo agoFor important and learning tasks,I would not use the voice feature as it was way too short and 'conversational'. I would use the 'record' feature and have it read the long, articulate answer to me. If this new conversation feature doesn't feel like I'm 'hanging out' with a friend, but actual longer high content answers, It'll win me over. Otherwise I thought the previous conversation mode was way too watered down and I found myself getting frustrated having to keep asking many questions to further probe down to the details. I don't know, these things don't change much, we'll see.
- csswizardry 3mo agohttps://www.youtube.com/watch?v=POI5XHAU0sc https://www.youtube.com/watch?v=POI5XHAU0sc
- bariswheel 3mo agoI don't think the voice feature should act like a human, but 'complement' it. I already have friends I can talk to. Perhaps for quick answers it might be helpful but Google already does that for me and I don't have to worry about having to 'archive' the chat later and creates clutter. I really hope at some point we can 'option click' or whatever, and choose multiple threads to archive. It takes FOREVER to archive the cluttery chats one by one. That's my one wish feature, please make batch cleanup of my client reasonably easy. Imagine having to delete one file at a time in a folder of 50 plus files. Barf.
- guilhas 3mo agoQuite sad to think society will more easily be apart and develop a relation with a company bot
- ksec 3mo agoI am surprised that Siri has only been mentioned 5 times here, out of the current 355 comments. I am wondering if this because Siri is so bad, people think of Siri now as a voice activation method rather than an AI assistant as it was intended. The demo is so good, what was once a sci-fi / Iron Man Jarvis services is now real. I don't follow AI closely, but all previous iteration were at best ask and answer type of services. It wasn't real conversation. And whatever flaws it may have now, at the rate of improvement within a few iteration it will surely reach good enough stage for majority of people. This is also scary. Not just for adults, but for kids. How they could become even more isolated. I remember the PC era, the internet from Information Super Highway to Web 2.0 Then Smartphone. It may have been obvious to many but AI really is something much bigger than all the previous three, perhaps combined. And it is also the only one that I think is scary.
- m12k 3mo agoSiri is how I set timers, alarms and reminders on my iPhone. Every once in a while, I'll try to use it to play something on Spotify via my Sonos, then give up and do it manually in the app instead.
- cyrux004 3mo agocant wait for husk irl to test this
- ninkendo 3mo agoA year ago I tried using the voice mode to be the worlds most over engineered golf score card: I basically said “ok I’m golfing with some friends, mind if I tell you the scores as we get them and you tell us the running totals?” It was awful, it kept overhearing us talking and thought we were talking to it, interjecting with nonsense because it couldn’t really understand what we were talking about. And when I would say “ok render a score card” it couldn’t drop to the text interface or anything, it had to stay as voice, so it fumbled around trying to read the scores back to us. It’s a very stupid use case but I viewed it as a stress test to see how well the voice mode and multi-modality could work. It failed miserably, but I’ll be interested to see if this new version does any better (not that I actually need this use case, writing on a card with a pencil is just fine.)
- NikolaNovak 3mo agoThis is fascinating to me. Whether that's due to my slight autism or massive nerdery, I don't want more realistic voice. I already switched to non-advanced voice in gpt, and I cannot imagine wanting the mmmhms, the yesses, the laughs, in my ai interaction. I want to ask a structured question and get a structured meaningful response. Informative and structured are really the KPIs. The umms and ahms of existing gpt advanced voice are annoying enough, the recent increased usage of first person almost a deal breaker (when asking for bike technique on lose surfaces yesterday, it literally gave me "back when I was learning bikes riding" story - eww). Fascinating to see the architectural advances though, even when they deliver something I personally don't need :)
- y1n0 3mo agoHonestly I find it distracting when people do the mmmhmm thing. I don’t find it encouraging or helpful or whatever.
- NikolaNovak 3mo agoWhen I go into listening node, I listen. Actively and intensely. Apparently it unnerves people. My wife coaches me to give occasional Yes,Go On, or UhHuh. But it's a conscious, learned, active mechanism for me. Intuitively, I'll tell you if we've weered off my desired conversation path and I'd ask the same courtesy. I've learned very late in life it's apparently an autism thing. Either way, I don't seek mmms and umms in real live people and I especially don't need fake ones in a machine :)
- imp0cat 3mo agoTCP vs UDP. ;) Both have their uses.
- senectus1 3mo ago100% agree, its also disturbing that they want to make seem like a human. I suspect most people would prefer it to behave like the computer in star trek. just be there, and be responsive in direct polite ways. dont try to be a friend, dont try to pretend its people. its not. its fucking code.
- deleted 3mo ago[deleted]
- polarbearballs 3mo agoA nightmare scenario would be that people become so accustomed to talking to agreeable AI, that they lose the inability to talk to anything that disagrees with them or has different perspectives with responses that don't include stroking the ego.
- Archer6621 3mo agoIt's a very interesting phenomenon. It's often said that famous stars become delusional exactly because they are surrounded by yes-men who are too afraid of rocking the boat and of losing their relationship with the star. But if I have to think of my own experiences here, I would say that overly agreeable people become tiresome pretty quickly. To me their perspective does not offer anything of value without any pushback, as it certainly doesn't help with grounding my own thoughts. Perhaps it's why being too nice makes it difficult to form deeper bonds, and maybe paradoxically it is therefore a good thing that LLMs are overly agreeable. It probably also depends on one's mindset - those who are interested in growing could be more likely interested in opposing views, while those who perceive that they have already "made it" (e.g. stars) perhaps don't care so much and prefer an agreeable tone.
- anon-3988 3mo agoWhat the fuck is the rationale to make this? I can kinda see the rationale behind Facebook and social media, they _could_ be useful. But this? This is like making a nuclear bomb out in public. The world is going to get a LOT worse in the next 20 years. I am not talking about energy usage. I am talking about a society where everyone is TRULY inside their own bubble of creation.
- luciana1u 3mo ago[flagged]
- senectus1 3mo ago[dead]
- y1n0 3mo agoI used it a few times. It's weird. It reminded me of Christopher Walken with all the oddly placed pauses.
- miki123211 3mo agoOnce this gets video capabilities and is ported to glasses, it'll be a major revolution for blind people (and I say this as a blind person). People have tried "smart <thing> that helps blind people navigate" since the 80s, many, many, many times, and all such projects failed. The cycle of "wow, blind people could benefit from a navigation aid, why don't I make one, if there's none around, I must surely have been the first bright university student to think of this idea" is pretty well known in the community, and I'm personally quite tired of it. Nevertheless, I think this may be the one. Circa 2020, I have said that people who are getting a guide dog now are probably getting their last one. I think we aren't far off from that prediction coming true.
- dnel 3mo agoEven the guide dogs are getting laid off?? That's the most depressing AI victim news yet.
- hdjrudni 3mo agoAt first I was really impressed, I thought granny was the voice. But it turns out it's the same kind of annoying voice and tone. And then Constance starts talking to it and it immediately cuts her off at 1:05, after they just explained it was better at conversation flow. Seems a bit disappointing, but the 3 overlapping questions example was impressive.
- thefabsta 3mo agoI just tried this and I find the constant interruptions infuriating: ah, alright, mh, mh-mh, go on, etc. This means the model already reliably detected my point/question isn’t finished. Please give me an option or dial to tone down the back channel noise.
- HeadlessChild 3mo agoOT: This commercial video is gorgeous looking and I adore the actors!
- OrangeMusic 3mo agoThe ad with the grandmas is cute and funny, but from the first 20 seconds you can see that the voice annoyingly interrupts people while they are talking. It's almost as if it tries to reply too fast - faster than a real person would, and the results is that it replies while you're still talking. Oh and there was also a small fail in the live translate demo: the grandma says "tell him that..." which the bot translates verbatim, whereas a real translator would of course understand that this is an aside not to be translated. Well I guess at least I should be happy that they're transparent in their ads :)
- simongray 3mo agoThey don't act the way humans would. It's normal to start a sentence after the other person in the conversation pauses, especially if you're eager to say something (like the chat agent is), but a real human is socially aware and would in most cases immediately cease talking if the other person starts another sentence at the same time. Humans do this so much in their conversations that the other person doesn't really register the interruption at all.
- OrangeMusic 3mo agoExactly. If they have a setting somewhere for the minimum duration of a pause in the user's speech before starting replying, they should just crank it up. In this other add [1], it's even worse, the user has to say "There's more to the story though" and almost looks annoyed + the agent overreacts to _that_ with a weird "Aaaaahh!". Oh well... I'm sure they'll get it right eventually. [1] https://openai.com/index/introducing-gpt-live/?video=1208099473 https://openai.com/index/introducing-gpt-live/?video=1208099...
- rhet0rica 3mo agoEncoded in its design are the biases of its designers. It acts like a nervous twenty-something from San Francisco.
- ndkap 3mo agoMaybe they are trying to outmatch the world's worst translator: https://www.youtube.com/watch?v=foT9rsHmS24 https://www.youtube.com/watch?v=foT9rsHmS24
- _s_a_m_ 3mo agook but how many es has "beekeeper"?
- emilfihlman 3mo agoI hope they fixed the models in that they don't interrupt and jump in and that you can order them to be silent. Previous model could not be silent and always cut me off and spoke over me.
- Nevermark 3mo agoThe speech is so fake sounding. Not fake technically, but like a fake/pacifying kind of person. It is a very strange tone to train for. The model is not interested in the conversation, it is just "serving" a conversation. Which is very different from the engagement of SOTA text models.
- flumes_whims_ 3mo agoDo any voice models support different conversation "threads" with different context?
- james-mxtech 3mo agoReserving judgment until the actual feature is set public. The announcement post itself doesn't tell you much.