16 ms·
I think the live demo that happened on the livestream is best to get a feel for this model[0]. I don't really care whether it's stronger than gpt-4-turbo or no
by cube2222 2y ago
I think the live demo that happened on the livestream is best to get a feel for this model[0].
I don't really care whether it's stronger than gpt-4-turbo or not. The direct real-time video and audio capabilities are absolutely magical and stunning. The responses in voice mode are now instantaneous, you can interrupt the model, you can talk to it while showing it a video, and it understands (and uses) intonation and emotion.
Really, just watch the live demo. I linked directly to where it starts.
Importantly, this makes the interaction a lot more "human-like".
[0]: https://youtu.be/DQacCB9tDaw?t=557 https://youtu.be/DQacCB9tDaw?t=557
- gabiruh 2y agoIt's weird that the "airplane mode" seems to be ON on the phone during the entire presentation.
- arthurcolle 2y agoThis was on purpose - they connected it to the internet via a USB-C cable it appears, for consistent internet instead of having it switch WiFi Probably some kinks there they are working out
- _flux 2y agoAnd eliminate the change of some prankster affecting the demo by attacking the wifi.
- OJFord 2y ago> Probably some kinks there they are working out Or just a good idea for a live demo on a congested network/environment with a lot of media present, at least one live video stream (the one we're watching the recording of), etc. At least that's how I understood it, not that they had a problem with it (consistently or under regular conditions, or specific to their app).
- hbn 2y agoThat's very common practice for live demos. To avoid situations like this: https://www.youtube.com/watch?v=6lqfRx61BUg https://www.youtube.com/watch?v=6lqfRx61BUg
- simoes 2y agoThey mention at the beginning of the video that they are using hardwired internet for reliability reasons.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- sitkack 2y agoYou would want to make sure that it is always going over WiFi for the demo and doesn't start using the cellular network for a random reason.
- rightbyte 2y agoYou can turn off mobile data. They probably just wanted wired internet.
- deleted 2y ago[deleted]
- fvdessen 2y agoThe demo is impressive but personally, as a commercial user, for my practical use cases, the only thing I care about is how smart it is, how accurate are its answers and how vast is its knowledge. These haven’t changed much since GPT-4, yet they should, as IMHO it is still borderline in its abilities to be really that useful
- CapcomGo 2y agoBut that's not the point of this update
- fvdessen 2y agoI know, and I know my comment is dismissive of the incredible work shown here, as we’re shown sci-fi level tech. But I feel I have this kettle, that boils water in 10min, and it really should boil it in 1, but instead is now voice operated. I hope the next version delivers on being smarter, as this update instead of making me excited, makes me feel they’ve reached a plateau on the improvement of the core value and are distracting us with fluff instead
- shepherdjerred 2y agoEverything is amazing & Nobody is happy: https://www.youtube.com/watch?v=PdFB7q89_3U https://www.youtube.com/watch?v=PdFB7q89_3U
- 0xB31B1B 2y agogpt4 isn't quite "amazing" in terms of commercial use. Gpt4 is often good, and also often mediocre or bad. Its not going to change the world, it needs to get better.
- Spivak 2y agoIt's an impressive demo, it's not (yet) an impressive product. It seems like the people who are ohhing and ahhing at the former and the people who are frustrated that this kind of this is unbelivably impractical to productize will be doomed to talk past one another forever. The text generation models, image generation models, speech-to-text and text-to-speech have reached impressive product stages. Multi-model hasn't got there because no one is really sure what to actually do with the thing outside of make cool demos.
- aaroninsf 2y agoAbsolutely agree. This model isn't about basemark chasing or being a better code generator; it's entirely explicitly focused on pushing prior results into the frame of multi-modal interaction. It's still a WIP, most of the videos show awkwardness where its capacity to understand the "flow" of human speech is still vestigial. It doesn't understand how humans pause and give one another space for such pauses yet. But it has some indeed magic ability to share a deictic frame of reference. I have been waiting for this specific advance, because it is going to significantly quiet the "stochastic parrot" line of wilfully-myopic criticism. It is very hard to make blustery claims about "glorified Markov token generation" when using language in a way that requires both a shared world model and an understanding of interlocutor intent, focus, etc. This is edging closer to the moment when it becomes very hard to argue that system does not have some form of self-model and a world model within which self, other, and other objects and environments exist with inferred and explicit relationships. This is just the beginning. It will be very interesting to see how strong its current abilities are in this domain; it's one thing to have object classification—another thing entirely to infer "scripts plans goals..." and things like intent, and, deixis. E.g. how well does it now understand "us" and "them" and "this" vs "that"? Exciting times. Scary times. Yee hawwwww.
- DonHopkins 2y ago>But it has some indeed magic ability to share a deictic frame of reference. They really Put That There! https://www.youtube.com/watch?v=RyBEUyEtxQo https://www.youtube.com/watch?v=RyBEUyEtxQo Oh, shit.
- razodactyl 2y agoIn my view, this was in response to the machine being colourblind haha
- nicklecompte 2y agoWhat part of this makes you think GPT-4 suddenly developed a world model? I find this comment ridiculous and bizarre. Do you seriously think snappy response time + fake emotions is an indicator of intelligence? It seems like you are just getting excited and throwing out a bunch of words without even pretending to explain yourself: > using language in a way that requires both a shared world model Where? What example of GPT-4o requires a shared world model? The customer support example? The reason GPT-4 does not have any meaningful world model (in the sense that rats have meaningful world models) is that it freely believes contradictory facts without being confused, freely confabulates without having brain damage, and it has no real understanding of quantity or causality. Nothing in GPT-4o fixes that, and gpt2-chatbot certainly had the same problems with hallucinations and failing the same pigeon-level math problems that all other GPTs fail.
- snthpy 2y agoHectic! Thanks for this.
- OJFord 2y agoI assume (because they don't address it or look at all phased) the audio cutting in and out is just an artefact of the stream?
- throwthrowuknow 2y agoHaven’t tried it but from work I’ve done on voice interaction this happens a lot when you have a big audience making noise. The interruption feature will likely have difficulty in noisy environments.
- OJFord 2y agoYeah that was actually my first thought (though no professional experience with it/on that side) - it's just that the commenter I replied to was so hyped about it and how fluid & natural it was and I thought that made it really jarr.
- mvdtnz 2y agoInteresting that they decided to keep the horrible ChatGPT tone ("wow you're doing a live demo right now?!"). It comes across just so much worse in voice. I don't need my "AI" speaking to me like I'm a toddler.
- yieldcrv 2y agotell it to speak to you differently with a GPT you can modify the system prompt
- maest 2y agoIt still refuses to go outside the deeply sanitise tone that "alignment" enforces on you.
- marvin 2y agoOne of the linked demos is it being sarcastic, so maybe you can make it remember to be a little more edgy.
- slibhb 2y agoYou can tell it not to talk like this using custom prompts.
- practice9 2y agoIt is cringe overenthusiastic, but a proper instructions/system prompt will fix that mostly
- throwthrowuknow 2y agoDid you miss the part where they simply asked it to change its manner of speaking and the amount of emotion it used?
- baumgarn 2y agoit should be possible to imitate any voice you want like your actual parents soon enough
- umbinder 2y ago[flagged]
- ChuckMcM 2y agoI expect the really solid use case here will be voice interfaces to applications that don't suck. Something I am still surprised at is that vendors like Apple have yet to allow me to train the voice to text model so that it only responds to me and not someone else. So local modelling (completely offline but per speaker aware and responsive), with a really flexible application API. Sort of the GTK or QT equivalent for voice interactions. Also custom naming, so instead of "Hey Siri" or "Hey Google" I could say, "Hey idiot" :-) Definitely some interesting tech here.
- spaceman_2020 2y agoThis is going straight into 'Her' territory
- clhodapp 2y agoCall me overly paranoid/skeptical, but I'm not convinced that this isn't a human reading (and embellishing) a script. The "AI" responses in the script may well have actually been generated by their LLM, providing a defense against it being fully fake, but I'm just not buying some of these "AI" voices. We'll have to see when end users actually get access to the voice features "in the coming weeks".