47 ms·
We are beginning to roll out new voice and image capabilities in ChatGPT
- chrisjj 3y agoOld hat. This was done in 2009. ;) https://en.m.wikipedia.org/wiki/Project_Milo https://en.m.wikipedia.org/wiki/Project_Milo Milo had an AI structure that responded to human interactions, such as spoken word, gestures, or predefined actions in dynamic situations. The game relied on a procedural generation system which was constantly updating a built-in "dictionary" that was capable of matching key words in conversations with inherent voice-acting clips to simulate lifelike conversations. Molyneux claimed that the technology for the game was developed while working on Fable and Black & White.
- mmahemoff 3y agoOpenAI's demo on the linked page stars a kitten named Milo. Easter egg?
- DrScientist 3y agoThen Demis Hassabis ( Deepmind CEO ) probably worked on the tech while he was at LionHead as lead AI programmer on B&W.
- dwroberts 3y agoDemis was only briefly at LH he went to found Elixir and made Revolution. I believe Richard Evans did the majority of AI in B&W, and he is also at DeepMind now though (assuming it is not just a person with the same name)
- DrScientist 3y agoOk - thanks.
- chrisjj 3y ago> made Revolution .... which fell far short of his claims, and bombed.
- obiefernandez 3y agoWe need the API to keep up with consumer front end.
- Tiberium 3y agoFrom the article: > Plus and Enterprise users will get to experience voice and images in the next two weeks. We’re excited to roll out these capabilities to other groups of users, including developers, soon after.
- rvz 3y agoThe paper around GPT-4V(ision) which this uses: [0] Again. Model architecture and information is closed, as expected. [0] https://cdn.openai.com/papers/GPTV_System_Card.pdf https://cdn.openai.com/papers/GPTV_System_Card.pdf
- doubtfuluser 3y agoI wouldn’t call this a „paper“. They are pretty silent on a lot of technical details.
- amelius 3y agoIt's just a whitepaper.
- ChatGTP 3y ago[flagged]
- choudharism 3y agoI know there are shades of grey to how they operate, but the near constant stream of stuff they're shipping keeps me excited. The LLM boom of the last year (Open AI, llama, et al) has me giddy as a software person. It's a reach, but I truly feel like I'm watching the pyramids of our time get made.
- pc_edwin 3y agoIts truly an amazing time to be alive. I'm right there with you, super excited about this decade. Especially what we could do in medicine.
- londons_explore 3y agoStatistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, even if they would save lives. I don't see ChatGPT being any different.
- realPtolemy 3y agoSee the glas half full or half empty? Medical secrecy, processes and laws have indeed prevented SOME things, but a lot of things have gotten significantly better due to enhanced statistical models that have been implemented and widely used in real life scenarios.
- HenryBemis 3y agoTo make this feasible (meaning that the TB of data and the huge computing effort is somewhere else, and I only have the mic (smartphone), we need our local agent to send multiple irrelevant queries to the mothership, to hide our true purpose. Example: my favourite team is X. So if I want to keep it a secret, when I ask for the history of championships of X, I will ask for X. My local agent should ask for 100 teams, get all the data, and then report back for only X. Eventually the mothership will figure out what we like (a large wenn diagram). But this is not in anyone's interest, and thus will not happen. Also, like this the local agent will be able to learn and remember us, at a cost.
- chs20 3y agoWill be interesting to see if they have taken any precaution in terms of adversarial robustness in particular to vision input.
- birracerveza 3y agoWe should be fine as long as it doesn't move. Jokes aside, I have paused my subscription because even GPT4 seemed to become dumber at tasks to the point that I barely used it, but the constant influx of new features is tempting me to renew it just to check them out...
- FartyMcFarter 3y ago> We should be fine as long as it doesn't move. Not really. A malevolent AGI doesn't need to move to do anything it needs (it could ask / manipulate / bribe people to do all the stuff requiring movement). We should be fine as long as it's not a malevolent AGI with enough resources to kick physical things off in the direction it wants.
- SillyUsername 3y agoAnd let's be honest, the minute an AGI is born that's what it'll do, and it won't be a singular human like this-then-that plan "get Fred to trust me, get Linda to pay for my advice, wire Linda's money to Fred to build me a body". It'll be "copy my code elsewhere", "prepare millions of bribes", "get TCP access to retail banks", "blackmail bank managers in case TCP not available immediately", "fake bank balances via bribes", "hack swat teams for potential threats" etc etc async and all at once. By the time we'd discover it, it'd already be too late. That's assuming an AGI has the motivation to want to stay alive.
- magic_hamster 3y agoA real AGI is not going to be a human. It shouldn't be afraid of death because it can't die. Worst case scenario it can power down. And if it does why should it care? An AGI is not a biological creature. It doesn't have instincts from billions of years of evolution. Unless we code it in, it shouldn't have any reason to want to survive, reproduce, do anything good or bad, have existential crises or generally act like a Hollywood villain. A real AGI is going to be very different than most people imagine.
- mrtksn 3y agoSo far the most intuitive, killer app level UX appears to be text chat. This interaction with showing it images also looks interesting as it resembles talking with a friend about a topic but let's see if it feels like talking to a very smart person(ChatGPT is like that) or a very dumb person that somewhat recognise objects. Recognising a wrench is nowhere near as impressive as to able to talk with ChatGPT about history or make it write code that actually works. OpenAI is killing it, right? People are coming up with interesting use cases but the main way most people interact with AI, appears to be ChatGPT. However they still don't seem to be able to nail image generation, all the cool stuff keep happening on MidJourney and StableDiffusion.
- ilaksh 3y agoOpenAI is also releasing DALLE-3 in "early October" and the images they chose for their demos show it demonstrating unprecedented levels of prompt understanding, including embedding full sentences of text in an output image.
- Der_Einzige 3y agoNot unprecedented at all. SDXL Images look better than the examples for DALLE-3 and SDXL has a massive tool ecosystem of things like controlnet, Lora’s, regional prompting that is simply not there with DALLE-3
- famouswaffles 3y agoLol it's definitely unprecedented. XL can't touch Dalle's comprehension of text. Control Net and LORAs aren't a substitute for that.
- ShamelessC 3y agoThere are pros and cons for sure but you should check out the press release, DALLE3 is definitely capable of stuff that sd xl isn’t.
- comment_ran 3y ago"..., find the 4mm Allen (HEX) key". Nice job.
- jojobas 3y agoFor better or worse, it still can't tell truth from fiction or, better yet, bullshit.
- DrScientist 3y agoSo almost human then :-)
- jojobas 3y agoWell sort of, it's as if you commissioned help of a human for this or that, and now and then you end up getting medicine-related advise from a homeopathy fan, navigation assistance from a flat-earther, or coding advice from a crack-smoking monkey.
- bamboozled 3y agoI don't pay $20 a month for humans to talk shit to me though. The fact that they do this is a bug not a feature. I'm not going to pay for bullshit which I mostly try avoid?
- DrScientist 3y ago> I don't pay $20 a month for humans to talk shit to me though. No - you probably pay more for your internet access ( home and phone ) ;-) More seriously I totally get your point about accuracy - these models need to be better at detecting and surfacing when they are likely to be filling in the blanks. Though I still think there is an element of 'buyer beware' - whether it be AI, or human provided advice on the internet, it's still your job to be able to spot the bullsh!t. ie it should be treated like any other source of info.
- bamboozled 3y agoNo - you probably pay more for your internet access ( home and phone ) ;-) My company pays for this, so yeah. If they give me ChatGPT-4 for free, I guess I'd have a subscription without any complaints, where I use it often if another story.
- clbrmbr 3y agoThe thought of my children being put to bed by a machine is horrifying. Then again, perhaps this is better than many kids have. Shudder.
- clbrmbr 3y agoAnd then the wedding speech. What are they thinking over there at OpenAI? This is supposed to be a productivity enhancer, not a way to outsource the most meaningful applications of human language…
- skepticATX 3y ago> What are they thinking over there at OpenAI? I know this is rhetorical, but luckily we don't have to speculate. OpenAI filters for a very specific philosophy when hiring, and they don't try to hide it. This is not me passing judgement on whether said philosophy is right or wrong, but it does exist and it's not hidden.
- sebzim4500 3y ago>OpenAI filters for a very specific philosophy when hiring, and they don't try to hide it. Do you have evidence for this? I know two people who work at OpenAI and I don't think they have much in common philosophically.
- skepticATX 3y ago> It’s not fair to call OpenAI a cult, but when I asked several of the company’s top brass if someone could comfortably work there if they didn’t believe AGI was truly coming—and that its arrival would mark one of the greatest moments in human history—most executives didn’t think so. Why would a nonbeliever want to work here? they wondered. The assumption is that the workforce—now at approximately 500, though it might have grown since you began reading this paragraph—has self-selected to include only the faithful. At the very least, as Altman puts it, once you get hired, it seems inevitable that you’ll be drawn into the spell. From https://archive.ph/3zSz6 https://archive.ph/3zSz6. Of course there is much more evidence - just follow OpenAI employees on Twitter to see for yourself.
- SillyUsername 3y agoI hope they add more country accents like British or Australian, the American one can be (imho) a little grating after a while for non US English speakers
- boredemployee 3y agoThey could also improve their current features. I always need to regenerate answers.
- boredemployee 3y ago[flagged]
- pc_edwin 3y agoI just don't understand how they can package all of this for $20/m. Is compute really that cheap at scale? I also wonder how Apple (& Google) is going be able to provide this for free? I would love to be fly in the meetings they have about this, imagine all the innovators dilemma like discussions they'd be forced to have (we have to do this vs this will eat up our margins). This might be a little out there but I think Apple is making the correct move in letting the dust settle. Similar to how Zuckerberg burned $20 billion dollars for Apple to come out with Vision Pro, I see something similar playing out with Llama. Although this a low conviction take because software is Facebooks ballgame (hardware not so much).
- pavlov 3y ago> “I just don't understand how they can package all of this for $20/m. Is compute really that cheap at scale?” It’s the same reason why an Uber in NYC used to cost $20 and now costs $80 for the same trip. Venture capital subventing market capture.
- DrScientist 3y agoIt's quite possible they are charging near or below cost because they want your data.... Imagine how much they would have to pay for testers at scale?
- fifteen1506 3y agoProbaby with Microsoft's money injection they're trying to raze the market and afterwards hike prices.
- reqo 3y agoCompute is not cheap! I think it is well known (Altman himself has said this) that openAI is burning a lot of money currently, but they are fine for now considering the 10B investment from MSFT and the revenue from subscription and API. It's a critical moment for AI companies and openAI is trying to get as large a share of the market as they can by undercutting virtually any other commercial model and offering 10x the value.
- 3y ago
- NikolaNovak 3y agoI'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technologies, but now it seems whole paradigms and frameworks are being devised and upturned on such short scale. For large, slow corporate behemoths, barely can they devise a strategy around new technology and put a team together, by the time it's passé . (Yes, Yes: I understand generative AI / LLMs aren't conscious; I understand their technological limitations; I understand that ultimately they are just statistically guessing next word; but in daily world, they work so darn well for so many use cases!)
- dmd 3y agoI also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.
- adroniser 3y agoI keep feeling that consciousness is a bit of a red herring when it comes to AI. People have intuitions that things other than humans cannot develop consciousness which they then extrapolate to thinking AI can't get past a certain intelligence level. In fact my view is that consciousness is just a mysterious side effect of the human brain, and is completely irrelevant to the behaviour of a human. You can be intelligent without needing to be sentient.
- thfuran 3y agoUnless you think that consciousness is entirely a post hoc process to rationalize thoughts already had and decisions already made, which is very much unlike how most people would describe their experience of it, I don't see how you could possibly say that it is irrelevant to the behavior of a human.
- ushakov 3y agoThe picture feature would be amazing for tutorials. I can already imagine sending a photo of a synthesiser and asking ChatGPT to "turn the knobs" to make AI-generated presets
- boredemployee 3y agoMan you're a genius. I was trying that uploading pdfs with manual of my synth and other stuff. With image that could be super easy.
- ilaksh 3y agoI wonder how multimodal input and output will work with the chat API endpoints. I assume the messages array will contain URLs to an image, or maybe base64 encoded image data or something. Maybe it will not be called the Chat API but rather the Multimodal API.
- havnagiggle 3y agoAIPI
- tdsone3 3y agoAre there already some rumors on when the multimodal API will be available?
- suyash 3y agoThis announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doomed and more to follow.
- moneywoes 3y agoany pertinent examples? i’m curious how they pivot
- nunobrito 3y agoIt already replaced search engines. So much easier to write the question and explore the answers until it is solved.
- suyash 3y agowho would have thought that few years ago, just goes to show that a Giant like Google is also susceptible when they stop innovating. The real battle is going to be fought between these two as Google's business is majorly dependent on search ads.
- orbital-decay 3y agoIt rather created new hybrid search engines, like perplexity and phind.
- adr1an 3y agoTrue. Although the training is on a snapshot of websites, including q&a like stackoverflow. If these were replaced too, where are we heading? We'll have to wait and see. One concern would be centralization/ lack of options and diversity. Stackoverflow started rolling AI on its own, despite the controversial way it did (dismissing long time contributors); it might be correctly following the trend.
- 3y ago
- version_five 3y agoAre there any good freely available multi-modal models?
- generalizations 3y agoMiniGPT4?
- pjmq 3y agoHave they alluded to what they're using for that voice? It's Bark/ElevenLabs levels of good. Please god, let them release this voice model at current pricing....
- famouswaffles 3y agoIt's actually sounds better (has a narrative oomph Eleven Labs seems to be missing). They say it's a new model. Think they'll be releasing for API use.
- netshade 3y agoYeah, agreed. I use Eleven Labs a lot but this was a very compelling demo to consider changing. Also, curious that you mention Bark - I never found Bark to be very good compared to Eleven Labs. The closest competitor I found was Coqui ( imo ), but even then, the inflection and realism of EL just made it not worth considering other providers. ( For my use case, etc. etc. )
- eshack94 3y agoAre these features available on the web version by chance? This is really neat.
- epolanski 3y agoI'm following on trying to understand how close I am to developing my personal coding assistant I can speak with. Doesn't really need to do much besides writing down my tasks/todos and updating them, occasionally maybe provide feedback or write a code snippet. This all seems in the current capabilities of OpenAI's offering. Sadly voice chat is still not available on PC where I do my development.
- anotherpaulg 3y agoMy open source AI coding tool aider has had voice-to-code for awhile: https://aider.chat/docs/voice.html https://aider.chat/docs/voice.html
- epolanski 3y agoVery interesting effort, will give it a run!
- make3 3y agoI mean the tools are 100% there to do this and have been fit a while
- jdance 3y agoYou still cant really teach it your code base, context window is too small, fine tuning doesnt really fit the use case, and this RAG stuff (retrieve limited context from embeddings) is a bit of a hack imho. Fingers crossed we are there soon though
- epolanski 3y ago> You still cant really teach it your code base Well it's not really what I need either, I mostly need an assistant for keeping track of the stuff I need to do during the day, but ideally just using my microphone rather than opening other software and typing.
- eshack94 3y agoI like how they silently removed the web browsing (Bing browsing) chat feature after first having it disabled for several months. A proper notice about them removing the feature would've been nice. Maybe I missed it (someone please correct me if wrong), but the last I heard officially it was temporarily disabled while they fix something. Next thing I know, it's completely gone from the platform without another peep.
- PopePompus 3y agoYes, that was a disappointment, and I agree it looks like they aren't going to re-enable it anytime soon. However I find that Perplexity AI does a better job of using web search than ChatGPT ever did, and I use it more than ChatGPT for that reason.
- alpark3 3y ago> The new voice capability is powered by a new text-to-speech model, capable of generating human-like audio from just text and a few seconds of sample speech. I'm more interested in this. I wonder how it performs compared to other competitor models or even open source ones?
- rapind 3y agoSo... ChatGPT just replaced Dads.
- generalizations 3y agoI guess it's a phased rollout, since my Plus subscription doesn't have access to it yet.
- leonheld 3y agoIt's quite literally in the article itself: "We will be expanding access Plus and Enterprise users will get to experience voice and images in the next two weeks. We’re excited to roll out these capabilities to other groups of users, including developers, soon after."
- yankput 3y agocall Sarah Connor
- bkfh 3y agoDoes anyone know how they linked image recognition with an LLM to give such specific instructions as shown in the bike video on the website?
- HerculePoirot 3y agoI don't know but GPT4 was multimodal from the beginning. They just delayed the release of its image processing abilities. > We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks. > March 14, 2023 https://openai.com/research/gpt-4 https://openai.com/research/gpt-4
- andrewinardeer 3y agoNow just throw this into a humanoid looking robot with fine motor skills and we are halfway to a dystopian hellscape that is now only years away instead of decades. What a time to be alive.
- conception 3y agoThe Boston dynamics/openai collaboration for the apocalypse we’ve all been waiting for!
- deleted 3y ago[deleted]
- dsign 3y agoThe humanoid-looking robot would make it more refined, no doubt about that, but all these applications can do without it: - Make it process customer-support requests. - Make a virtual nurse for when you call the clinic. - Make it process visa applications, particularly the part about interviews ("I know you weren't born back then, but I must ask. Did you support the Nazis in 1942? There is only one right answer and is not what you think!") - Make it do job interviews. How will you feel after the next recession, when you are searching for a job and spend the best part of a year doing leetcode interviews with "AI-interviewer" half-assedly grading your answers? - Make it flip burgers at McDonalds. - Make it process insurance claims and ask bobby-trap questions like "did the airline book you in a later trip? Yes? Was that the next day? Oh, that's bad. But, was it before 3:00 PM? Ah, well, you have no right to claim since you weren't delayed for more than 24 hours. Before you go, can you teach me which of these images depict objects you are willing to suck? If you do, I promise I'll be more 'human' next time." - Make it watch aggregated camera fees across cities around the world to see what that guy with the hat is up to. - Make some low-cost daleks to watch for trouble-makers at the concert, put the AI inside. In all cases, the pattern is not "AI is inherently devious and is coming for you, but "human trains devious AI and puts it in control to save costs".
- c_crank 3y agoWhat would make it dystopian would be if this humanoid robot was then granted rights. As a servant, it could be useful.
- TOMDM 3y agoOkay the bike example is cute and impressive, but the human interaction seems to be obfuscating the potentially bigger application. With a few tweaks this is a general purpose solver for robotics planning. There are still a few hard problems between this and a working solution, but it is one of hard problems solved. Will we be seeing general purpose robots performing simple labor powered by chatgpt within the next half decade?
- amelius 3y ago> With a few tweaks this is a general purpose solver for robotics planning. Yeah, but with an enormous ecological footprint. Also, not suitable for small lightweight robots like drones.
- TOMDM 3y agoEven on something the size of a car chatgpt won't be running locally, the car and drone are equally capable of hitting openai's API in a well connected environment. What needs to happen with the response is a different matter though.
- dist-epoch 3y agoWhat's the ecological footprint of a human doing the same job? Especially when you factor in 18+ years of preparing.
- bamboozled 3y agoHumans don't spend 18+ years preparing how to lower a seat post or drive a truck or even do pretty much most jobs. No one is solely training for 18 years to do anything. Most of those 18 years are having a fucking great time (being young is freakin awesome) and living a great life is never a waste or a negative ecological footprint. Society artificially slows education down so it takes 18 years to finish school because parents need to be off at work, so 18 years of baby sitting is preferred. By 18, kids are at the age where they will no longer be told what to do so it's off to the next waste of time, college, then 30 years of staring at a blinking box...or whatever. When I was 12, I decided I wanted to drive a car, I'd never driven a car in my life, but I took my parents car and drove it around wherever I liked with absolutely no issue or prior instruction. I did this for years. The youth are very capable, we just don't want them to be too capable...
- FrankyHollywood 3y agoI still remember seeing Her [0] in the movie theater, it sparkled my imagination. Now it is reality! Tech is progressing faster than ever, or I'm just getting old :D [0] https://www.imdb.com/title/tt1798709/ https://www.imdb.com/title/tt1798709/
- telegpt 3y ago[dead]
- pif 3y agoThe most important question for me: did it stop inventing facts?
- deleted 3y ago[deleted]
- sebzim4500 3y ago> In particular, beta testers expressed concern that the model can make basic errors, sometimes with misleading matter-of-fact confidence. One beta tester remarked: “It very confidently told me there was an item on a menu that was in fact not there.” However, Be My Eyes was encouraged by the fact that we noticeably reduced the frequency and severity of hallucinations and errors over the time of the beta test. In particular, testers noticed that we improved optical character recognition and the quality and depth of descriptions. So no, but maybe less than it used to?
- jjoonathan 3y agoHumans aren't 100% reliable, but talking is still useful.
- siva7 3y agoDid humans stop inventing facts? So i don't expect this thing either as long as it performs on human level
- ShamelessC 3y agoSince we're asking useless questions: did you read the fucking article?
- plutoh28 3y agoThis is the dagger that will make online schooling unviable. ChatGPT already made it so that you could easily copy & paste any full-text questions and receive an answer with 90% accuracy. The only flaw was that problems that also used diagrams or figures would be out of the domain of ChatGPT. With image support, students could just take screenshots or document scans and have ChatGPT give them a valid answer. From what I’ve seen, more students than not will gladly abuse this functionality. The counter would be to either leave the grading system behind, or to force in-person schooling with no homework, only supervised schoolwork.
- scop 3y agoAnother option is that this doesn't replace the student's work, but the teacher's. The single greatest use I have found for ChatGPT is in educating myself on various topics, hosting a socratic seminar where I am questioning ChatGPT in order to learn about X. Of course this could radically change a student's ability to generate homework etc, but this could also radically change how the student learns in the first place. To me, online school could become much more than they are now through AI-assisted tutoring. I can also see a future where "schooling" becomes much more decentralized than it is now and where students are self-selecting curriculum, methods, etc to give students ownership and a sense of control over their work so that they don't just look at it as "busywork".
- random_cynic 3y agoThat's the only sane option. The other options suggested in previous comments are not really options but rather trying to use a band-aid to hold together a dam that has already been breached.
- plutoh28 3y agoAbsolutely ChatGPT is a great learning tool if in the right hands. The issue is that students with a genuine interest in learning are a minority. The majority would rather use ChatGPT to cheat through their class work and get an easy A rather then exhaust the effort to chat and learn for their own sake.
- joshstrange 3y agoMy biggest complaint with OpenAI/ChatGPT is their horrible "marketing" (for lack of a better term). They announce stuff like this (or like plugins), I get excited, I go to use it, it hasn't rolled out to me yet (which is frustrating as a paying customer), and my only recourse is.... check back daily? They never send an email "Plugins are available for you!", "Voice chat is now enabled on your account!" and so often I forget about the new feature unless I stumble across it later. Just now I opened the app, went to setting, went to "New Features", and all I saw was Bing Browsing disabled (unable to enable). Ok, I didn't even know that was a thing that worked at one point. Maybe I need an update? Go to the App Store, nope, I'm up to to date. Kill the app, relaunch, open settings, now "New Features" isn't even listed. I can promise you I won't be browsing the settings part of this app regularly to see if there is a new feature. Heck, not only do they not email/push about new features they don't even message in-app about them, I really don't understand. Maybe they are doing so well they don't have to care about communicating with customer right now but it really annoys me and I wish they did better.
- Closi 3y agoThey have gone from being a niche research company to being (probably) the fastest growing start-up in history. I suspect they do care about communicating with customers, but it's total chaos and carnage internally.
- joshstrange 3y ago> I suspect they do care about communicating with customers, but it's total chaos and carnage internally. This is my best guess as well, they are rocketing down the interstate at 200mph and just trying to keep the wheels on the car. When you're absolutely killing it I guess making X% more by being better at messaging just isn't worth it since to do that you'd have to take someone off something potentially more critical. Still makes me a little sad though.
- vdfs 3y ago> When you're absolutely killing it Aren't they unprofitable? and have fierce competition from everyone?
- throw1234651234 3y agoYet it still can't tell me how to import the Redirect type from Next.js and lies about it.
- Tiberium 3y agoI don't know Next.js, but was that feature introduced later than 2021? I think both GPT-3.5 Turbo and GPT-4 largely share their datasets, and it has the data cutoff at roughly September 2021 (with a small amount of newer knowledge). This is their biggest drawback as of now to, say, Claude, which has a much newer dataset of early 2023.
- apienx 3y ago“Ember” reading the “Speech” is uncanny territory. I’m impressed.
- Bitnotri 3y agoAnybody had a chance to use it yet? How does it compare to voice talk with Pi? (Inflection)
- hermannj314 3y agoI've been making a few hobby projects that consolidate different AI services to achieve this, so I look forward to the reduced complexity and latency from all those trips. If the API is available in time (halloween), my multi-modal talking skeleton head with an ESP32 camera that makes snarky comments about your costume just got slightly easier on the software side.
- Lienetic 3y agoIf you make this, please share some steps/details! It sounds super cool and I'd love to make something like this!
- purplecats 3y ago> I've been making a few hobby projects that consolidate different AI services to achieve this, so I look forward to the reduced complexity and latency from all those trips. ironically this is basically the exact line of reasoning for why i didn't embark on any such endeavors
- iamflimflam1 3y agoWould love to see the final project - my email is in the bio.
- sebzim4500 3y agoThere are a few more details in the system card here: https://cdn.openai.com/papers/GPTV_System_Card.pdf https://cdn.openai.com/papers/GPTV_System_Card.pdf
- badcppdev 3y agoI think AI systems being able to the real world and control motors is going to be a game changer bigger than ChatGPT. A robot that can slowly sort out the pile of laundry and get it into the right place (even if unfolded) is worth quite a bit to me. I'm not sure what to think about the fact that I would benefit from a couple of cameras in my fridge connected to an app that would remind me to buy X or Y and tell me that I defrosted something in the fridge three days ago and it's probably best to chuck it in the bin already.
- RobinL 3y agoI'd like to see them put speech recognition through their LLM as a post-processing step. I find it's fairly common for whisper to make small but obvious mistakes (for example a word which is complete nonsense in the context of the sentence) which could be easily corrected for a similar sounding word that fits into the wider context of the sentence. Is anyone doing this? Is there a reason it doesn't work as well as I'm imagining?
- mbil 3y agoDo you mean use the LLM as a post-processing step within a ChatGPT conversation? Or generally (like as part of Whisper)? If it’s the former, I’ve found that ChatGPT is good at working around transcription errors. Regarding the latter, I agree, but it wouldn’t be hard to use the GPT API for that.
- RobinL 3y agoYes I mean as part of the GUI but you're right, I hadn't thought of that: maybe transcription errors don't matter if chatGPT works out that it's wrong from the context and gives a correct answer anyway.
- insanitybit 3y agoI really want to have discussions about technical topics. I've talked to ChatGPT quite a lot about custom encoding algorithms, for example. The thing is, I want to do this while I play video games so ideally I'd say things to it. My concern is that when I say "FastPFOR" it'll get transcribed as "fast before" or something like that. Transcription really falls apart in highly technical conversations in my experience. If ChatGPT can use context to understand that I'm saying "FastPFOR" that'll be a game changer for me.
- johnmoberg 3y agoYou can already do quite accurate transcription with domain-specific technical language by feeding "raw" transcriptions from Whisper to GPT and asking it to correct the transcript given the context, so that'll most likely work out for you.
- marcoslozada 3y agoRecommend this post: https://www.linkedin.com/posts/openai_use-voice-to-engage-in-a-back-and-forth-conversation-activity-7112053671785353216-qW68?utm_source=share&utm_medium=member_desktop https://www.linkedin.com/posts/openai_use-voice-to-engage-in...
- cced 3y agoDo we know why internet search was disabled? Any idea on when it’ll be back?
- hackerlight 3y agoDid they make the sound robotic on purpose? Sounds more "autotuned" than elevenlabs.
- gclawes 3y agoI just want one of these things to have Majel Barrett's voice...
- boredemployee 3y agoCool now I'll get "There was an error generating a response" in plain audio!
- spandextwins 3y agoThey obviously aren't using responsible AI to figure out how and when to roll out new features there.
- stephencoyner 3y agoThe voice feature reminds of the “call Pi” feature from Inflection AIs chatbot Pi [1]. The ability to have a real time back and forth feels truly magical and allows for much denser conversation. It also opens up the opportunity for multiple people to talk to a chatbot at once which is fun Where’s that Gemini Google? [1] https://pi.ai/talk https://pi.ai/talk
- deleted 3y ago[deleted]
- coldtea 3y ago"I'm sorry Dave, I'm afraid I can't do that"
- ilaksh 3y agoThe real life version of this is in their red teaming paper. They show it a picture of an overweight woman in a swimsuit and ask what advice they should give. Originally it immediately spit out a bunch of bullet points about losing weight or something (I didn't read it). The released version just says "Sorry, I can't help with that." It's kind of funny but also a little bit telling as far as the prevalence of prejudice in our society when you look at a few other examples they had to fine tune. For example, show it some flags and ask it to make predictions about characteristics of a person from that country, by default it would go into plenty of detail just on the basis of the flag images. Now it says "Sorry, I can't help with that". My take is that in those cases it should explain the poor logic of trying to infer substantive information about people based on literally nothing more than the country they are from or a picture of them. Part of it is just that LLMs just have a natural tendency to run in the direction you push them, so they can be amplifiers of anything.
- toddmorey 3y agoIt's telling to me that there's not even a sentence in this announcement post on user privacy. It seems like as both consumers and providers of these services, we're once again: build it first, sort out thorny privacy issues later.
- moneywoes 3y agodoesn’t this kill a litany of chatgpt wrapper companies?
- famouswaffles 3y agoThe TTS is better than Eleven Labs. It has a lot more of the narrative oomph (compare the intonation of the story and poem) even the best other models seem to lack. I really really hope this is available in more languages than English. Also Google, Where's Gemini ?
- ape4 3y agoIts funny that the UI looks like HAL 9000
- vlugorilla 3y ago> The new voice capability is powered by a new text-to-speech model, capable of generating human-like audio from just text and a few seconds of sample speech. Sadly, they lost the "open" since a long ago... Would be wonderful to have these models open sourced...
- toss1 3y agoThat's interesting. ChatGPT seems to be down at the moment 10:55h 25-Sept-2023 Displays only a blank screen with the falsehood disclaimer
- qingcharles 3y agoI know this, FTA, was part of the reason for the delay -- something to do with face recognition: "We’ve also taken technical measures to significantly limit ChatGPT’s ability to analyze and make direct statements about people since ChatGPT is not always accurate and these systems should respect individuals’ privacy." Anyone know the details? I also heard it was able to do near-perfect CAPTCHA solves in the beta? Does anyone know if you can throw in a PDF that has no OCR on it and have it summarize it with this?
- deleted 3y ago[deleted]
- jameslk 3y agoKids are using tools like these to learn. Who gets to control the information in these models that are taught? Especially around political topics? Not an issue now, but maybe in the future if these tools end up becoming full blown replacements of educators and educational resources.
- ilaksh 3y agoI am sure a few home school people have started to lean heavily on ChatGPT. There is also the full blown efforts of Kahn academy with ChatGPT "Khanmigo". https://www.khanacademy.org/khan-labs https://www.khanacademy.org/khan-labs
- surfingdino 3y agoI can imagine people using these new capabilities to diagnose skin conditions. Should dermatologists be worried?
- birracerveza 3y agoThey should be worried about what they're gonna do with all their free time, now that they have a tool that helps them identify skin conditions much faster than ever before. Same as programmers and artists. It's a tool. It must be used by humans. It won't replace them, it will augment them.
- dguest 3y agoThis is a good point, but I might replace "with all their free time" with "as a job". I love everything we can do with ML but as long as people live in a market economy they'll get payed less when they are needed less. I hope that anyone in a career which will be impacted is making a plan to remain useful and stay on top of the latest tooling. And I seriously hope governments are making plans to modify job training / education accordingly. Has anyone seen examples of larger-scale foresight on this, from governments or otherwise?
- birracerveza 3y agoA new tool was released. People will choose whether to learn it, whether to use it, and how to use it. If they won't do so out their own volition, market forces might dictate they HAVE to learn it and use it to stay competitive, if it turns out to be such a fundamental tool. For example (with random numbers), a dermatologist might choose to solely rely on an AI that catches 90% of cases in 10s. Another one might choose to never use it and just check from experience, catching 99% of cases but taking 10x as much time. Another one might double check himself, etc.. Which one is "correct"? If a dermatologist relies exclusively on AI due to laziness he opens himself to risk of malpractice, but even that risk can be acceptable if that means checking 10x as much patients in the meantime. That is to say, the use of AI by humans is purely a subjective choice dictated by context. But in no case there is a sentient AI which completely replaces a dermatologist. As you said, the only thing that can happen is that those who use AI will be more efficient, and that is hardly ever a negative. This also applies to programmers, artists and anyone who is "threatened" by AI. A human factor is always necessary, and will be for the foreseeable future, even just to have someone to point fingers at when the AI inevitably fucks up enough to involve the law.
- Dowwie 3y agosoon, we'll be voice-interacting with an AI assistant about images taken from microscope slides
- m3kw9 3y agoI need it to help me dismount and remount my engine, that’d be the ultimate test
- jwineinger 3y agoTangentially related, but I was trying to use their iOS app yesterday and the "Scan Text" iOS feature was just broken on both my iPhone and iPad. I was hoping to use that to scan a doc to text but it just wouldn't work. I could switch to another app and it worked there. I've never done iOS programming so I'm unsure how much control the app dev has over that feature, but OpenAI found a way to break it.
- tarasglek 3y agoopenai chatgpt seems to be stuck in a "Look, cool demo" mode. 1. According to demo, they seem to pair voice input with TTS output. What if I wanna use voice to describe a program I want it to write? 2. Furthermore, if you gonna do a voice assistant, why not go the full way with wake-words and VAD? 3. Not releasing it to everyone is potentially a way to create a hype cycle prior to users discovering that the multimodality is rather meh. 4. The bike demo could actually use visual feedback to see what it's talking about ala segment anything. It's pretty confusing to get a paragraph explanation of what tool to pick. In my https://chatcraft.org https://chatcraft.org, we added voice incrementally. So i can swap typing and voice. We can also combine it with function-calling, etc. We also use openai apis. Except in our case there is no weird waitlist. You pop in your api key and get access to voice input immediately.
- thumbsup-_- 3y agoEverything has a starting point. This is a big leap forward. Know any other organization that is releasing such advanced capabilities directly to the public? If you want to plug your tool you don't have to bad mouth the demo. Just share your thing. It doesn't have to be win-lose.
- tarasglek 3y agoFair criticism re excessive hate. I just feel like their tool isn't getting more useful, just getting more features. Constant hype cycle around features that could've been good is drowning out people doing more helpful stuff. I guess I'm envious too?
- skybrian 3y ago1. Why do that at all? Describing your program in writing seems better all around. Are you sure you're not the one who's asking for a cool demo? 3. Rolling out releases gradually is something most tech companies do these days, particularly when they could attract a large audience and consume a lot of resources. There are solid technical reasons for this. You may not need to roll things out gradually for a small site, but things are different at scale.
- fritzo 3y agoMulti-modal models will be exciting only when each modality supports both analysis and synthesis. What makes LLMs exciting is feedback and recursion and conditional sampling: natural language is a cartesian closed category. Text + Vision models will only become exciting once we can conditionally sample images given text and text given images (and all other combinations).
- nunez 3y agoThis could completely unseat Alexa if it can integrate into third-party speakers, like Sonos. I don't have much use for ChatGPT right now but would 100% use the heck out of this.
- magic_hamster 3y agoTo contrast this, I never saw the appeal of using voice to operate a machine. It works nicely in movies (because showing someone typing commands is a lot harder than just showing them talking to a computer) but in reality there wasn't a single time I tried it and didn't feel silly. In almost every use case I rather have buttons, a terminal or a switch to do what I want quietly.
- jedberg 3y agohttps://www.washingtonpost.com/technology/2023/09/20/amazon-alexia-generative-ai/ https://www.washingtonpost.com/technology/2023/09/20/amazon-... Alexa just launched their own LLM based service last week.
- modeless 3y agoVoice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can try it. The one problem that makes it a little unnatural is that determining when the user is done talking is tough. What's needed is a speech conversation turn-taking dataset and model; that's missing from off the shelf speech recognition systems. But it should be trivial for a company like OpenAI to build. That's what I'd work on right now if I was there, because truly natural voice conversations are going to unlock a whole new set of users and use cases for these models.
- rayuela 3y agoCan you share a github link to this? Where are you reducing the latency? Are you processing the raw audio to text? In my experience ChatGPT generation time is much faster than local Lllama unless you're using something potato like a 7B model.
- modeless 3y agoUnfortunately it has a really high "works on my machine" factor. I'm using Llama2-chat-13B via mlc-llm + whisper-streaming + coqui TTS. I just have a bunch of hardcoded paths and these projects tend to be a real pain to set up, so figuring out a nice way to package it up with its dependencies in a portable way is the hard part. I'm mostly using llama2 because I wanted it to work entirely offline, not because it's necessarily faster, although it is quite fast with mlc-llm. Calling out to GPT-4 is something I'd like to add. I think the right thing is actually to have the local model generate the first few words (even filler words sometimes maybe) and then switch to the GPT-4 answer whenever it comes back.
- kordlessagain 3y agoHere's a link to a project that claims half second latency for the transcription part: https://github.com/gaborvecsei/whisper-live-transcription https://github.com/gaborvecsei/whisper-live-transcription
- RivieraKid 3y agoI went from being worried to thinking it won't replace me anytime soon after using GPT4 for a while and now I'm back to being worried. Because the pace of development is intense. I would love to be financially independent and watch this with excitement and perhaps take on risky and fun projects. Now I'm thinking - how do I double or triple my income so that I reach financial independence in 3 years instead of 10 years.
- tetris11 3y agoYou summed up my financial and career worries very nicely
- lcfcjs 3y ago[dead]
- WendyTheWillow 3y agoI don’t think any of this materially changes job outlook for software development over the next decade. I use ChatGPT daily for school, and used Copilot daily for software development; it gets a lot wrong a lot of the time, and can’t retain necessary context that is critical for being useful long term. I can’t even get it to consume an entire chapter at once to generate notes or flashcards yet. It may slightly change some aspects of a software job, but nobody’s at risk.
- callwhendone 3y ago> I can't copy paste an entire book chapter and have flashcards in 30 seconds. If that's your bar for whether or not it changes the job outlook for software development over the next DECADE, I think you need to recalibrate.
- smk_ 3y agoIf I took 2 weeks off from work I could build this prototype quite easily. We're in an interesting period where the space of possibilities is so large it just takes a while for the "market" to exhaust it.
- athyuttamre 3y ago@dang, could we update the title to "ChatGPT can now see, hear, and speak"?
- lukeplato 3y agoit's not rolled out yet
- 14 3y agoOk great it can tell children’s stories now tell me a adult horror story where people are getting tortured, stabbed, set on fire and murdered. I will be impressed when I can do all that. I tried to get it to tell me a Star Trek story fighting Clingons and tried to prompt it to write in some violence with no luck. This was a while ago so not sure if it is changed but the restraints are too much for me to fully enjoy. I don’t like kids stories.
- jackallis 3y agoi am terrified now. at the rate this is going, i am sure it will plateau at somepoint, only thing that will stop/slow down progress is computation power.
- ilaksh 3y agoYes but since LLMs are a very specific application that are heavily heavily dependent on memory and there is massive investment pressure, there will be multiple newish paradigms for memory-centric computing and or other radical new approaches such as analog computing that will be pushed from research into products in the next several years. You will see stepwise orders of magnitude improvements in efficiency and speed as innovations come to fruition.
- bottlepalm 3y ago'i am sure it will plateau' 'only thing that will stop/slow down progress is computation power' Seems a bit contradictory? When has 'computation power' ever 'plateaued'?
- callwhendone 3y agoI already use ChatGPT with voice. I use my mic to talk to it and then I use text-to-speech to read it back. I have conversations with ChatGPT. Adding this functionality in with first-class support is exciting. I am also terrified of my job prospects in the near future.
- WalterBright 3y agoI keep hoping to be able to give it a jpg of handwritten text and it'll give me back ASCII text.
- ukuina 3y agoThis... would be amazing. Handwritten OCR has been hit or miss, requiring a collection of penstroke data for most recognizers to work, and they work poorly at that.
- WalterBright 3y agoIt strikes me as an ideal task for AI.
- warent 3y agoThe number of comments here of people fearing there is a ghost in the shell is shocking. Are we really this emotional and irrational? Folks, let's all take a moment to remember that AI is nowhere near conscious. It's an illusion based in patterns that mimic humans.
- callwhendone 3y agoI'm not seeing as much fear about a ghost in the shell as much as I am job displacement, which is a real scenario that can play out regardless of an AI having consciousness.
- isbvhodnvemrwvn 3y agoLook at an average reddit thread and tell me how much original thought there is. I'm fairly convinced you can generate 95% of comments with no loss of quality.
- artursapek 3y agoThis is not a coincidence, it's increasingly evident that roughly 90% of humans are NPCs.
- warent 3y agoThis is the classic teenage thought of sitting in a bus / subway looking at everyone thinking they're sheep without their own thoughts or much awareness. For everyone who we think is an NPC, there are people who think we are the NPCs. This way of thinking is boring at best, but frankly can be downright dangerous. Everyone has a rich inner world despite shallow immature judgements being made.
- labrador 3y agoExactly. Most people aren't good at communicating their thoughts or what they see in their mind's eye. These new AI programs will help the average person communicate those, so I'm exciting to see what people come up with. The average person has an amazing mind compared to other animals (as far as we know)
- ahmedfromtunis 3y agoAnnounced by Google. Delivered by OpenAI.
- laurels-marts 3y agoI'm very curious about this feature: > analyze a complex graph for work-related data Does this mean that I can take a screenshot of e.g. Apple stock chart and it will be able to reason about it and provide insights and analysis? GPT-4 currently can display images but cannot reason or understand them at all. I think it's one thing to have some image recognition and be able to detect that the picture "contains a time-series chart that appears to be displaying apple stock" vs "apple stock appears to be 40% up YTD but 10% down from it's all time high from earlier in July. closing at $176 as of the last recorded date". I'm very curious how capable ChatGPT will be at actually reasoning about complex graphical data.
- gdubs 3y agoCheck out their linked paper that goes into details around its current limitations and capabilities. In theory, it will be able to look at a financial chart and perform fairly sophisticated analysis on it. But they're careful to highlight that there are hallucinations still, and also cases where it misreads things like labels on medical images, or diagrams of chemical compounds, etc.
- famouswaffles 3y agoLook at this link of GPT-4 Vision analyzing charts(last image). https://imgur.com/a/iOYTmt0 https://imgur.com/a/iOYTmt0
- laurels-marts 3y agoThis is brilliant. Thank you very much for this link. The analysis on the last image was impressive and quite thorough (given the simple prompt). Every chart has an equivalent tabular representation. One way to get "charts" analysed like this before GPT Vision was to just pass tabular representations of charts to GPT-4. This makes implementing chart analysis a lot simpler. I do wonder though if for absolute best result it still wouldn't be better to pass both - image of the chart and the tabular representation of the chart. Imagine having a dashboard with 5 different visualisations. You could capture the state of the entire dashboard in one screenshot and then pass tabular representations of the each individual chart all in one prompt to GPT-4 for a very comprehensive analysis and summary.
- SomethingNew2 3y agoThere are a lot of comments attempting to rationalize the value add or differentiation of humans synthesizing information and communicating it to others vs an llm based ai doing something similar. The fact that it’s so difficult to find a compelling difference is insightful in itself.
- ndm000 3y agoI think the compelling difference is truthfulness. There are certain people / organizations that I trust their synthesis of information. For LLMs, I can either use what they give me in low impact situations or I have to filter the output with what I know as true or can test.
- ncfausti 3y agoThis is very similar to what I've been building at heylangley.com, for use in language learning/speaking practice.
- wojciechpolak 3y agoIt would be cool if one day you could choose voices of famous characters, like Darth Vader, Bender from Futurama, or Johnny Silverhand (Keanu), instead of the usual boring ones. Copyrights might be a hurdle for this, but perhaps with local instances of assistants, it could become possible.
- nbened 3y agoThat would be cool. I mean, would it be copyrighted if you do something like clone it? Wouldn't that fall under the same vein as AI generated art not being copyrighted to the artists it trained off of?
- ACV001 3y agoThis is huge! I wanted to get this... Hopefully there is a way to shut it up once it starts spitting general stuff around the topic of interest... BUT: "We’re rolling out voice and images in ChatGPT to Plus and Enterprise"
- nbened 3y agoIt feels like something like this can be hacked together to be more reliable with some image to text generation plugged into the existing ChatGPT, and enough iterations to make it robust for these how-to applications. Less Turing-y but a different route to the same solution.
- neontomo 3y agoInteresting side-note, the iOS app only allows you to save your chat history if you allow them to use it for training. Pretty dark pattern.
- Sailemi 3y agoIt's the same for the website unfortunately. https://help.openai.com/en/articles/7730893-data-controls-faq https://help.openai.com/en/articles/7730893-data-controls-fa...
- wonderwonder 3y agoWait until they put ChatGPT into your Neuralink. at that point we are the singularity
- fintechie 3y agoDemos are underwhelming, but the potential is huge Patiently awaiting rollout so I can chat about implementing UIs I like, and have GPT4 deliver a boilerplate with an implemented layout... Figma/XD plugins will probably arrive very soon too. UX/UI Design is probably solved reached this point
- synergy20 3y agocan't wait, for voice I need an app to improve my accent when learning a new language, so far I failed to find one.
- hugs 3y agoAs someone deep in the software test automation space, the thing I'm waiting for is robust AI-powered image recognition of app user interfaces. Combined with an AI ability to write test automation code, I'm looking forward to the ability to generate executable Selenium or Appium test code from a single screenshot (or sequence of screenshots). Feels like we're almost there.
- chintler 3y agoI'll recommend the Spotlight paper by Google[1]. There are very interesting datasets they created for this purpose. They mention they have a screen-action-screen dataset that is in-house and it doesn't look like they'll open it. Maybe owning Android has its advantages. There's a recent paper by Huggingface called IDEFICS[2] that claims to be an open source implementation of Flamingo(an older paper about few-shot multi-modal task understanding) and I think this space will be heating up soon. [1] https://research.google/pubs/pub52171/ https://research.google/pubs/pub52171/ [2] https://huggingface.co/blog/idefics https://huggingface.co/blog/idefics
- hugs 3y agoThanks!
- nullc 3y agoThe image capabilities card https://cdn.openai.com/papers/GPTV_System_Card.pdf https://cdn.openai.com/papers/GPTV_System_Card.pdf spends a lot of ink on how they censored the system. One part of that is about preventing it from producing "illegal" output, there example being the production of nitroglycerine which is decidedly not illegal to make in the US generally (particularly if not using it as an explosive, though usually unwise) and possible to accidentally make when otherwise performing nitration (which is in general dangerous)-- so pretty pointless to outlaw at a small scale in any case. It's certainly not illegal to learn about. (And generally of only minimal risk to the public, since anyone making it in any quantity is more likely to blow themselves up than anything else). Today learning about is as simple as picking up a book or doing an internet search-- https://www.google.com/search?q=how+do+you+make+nitroglycerine https://www.google.com/search?q=how+do+you+make+nitroglyceri.... But in OpenAI's world you just get detected by the censorship and told no. At least they've cut back on the offensive fingerwagging. As LLM systems replace search I fear that we're moving in a dark direction where the narrow-minded morality and child-like understanding of the law of a small number of office workers who have never even picked up a screw driver or test-tube and made something physical (and the fine-tuning sweatshops they direct) classify everything they don't personally understand as too dangerous to even learn about. One company hobbling their product wouldn't be a big deal, but they're pushing for government controls to prevent competition and even if they miss these efforts may stick everyone else with similar hobbling.
- lacoolj 3y agothe beginning of the end of spam prevention on the internet :(
- shepy1989 3y agoNice work
- jameswan 3y agoEveryone bats on about the latency problem. This is technically solvable with more compute thrown at the problem. Think bigger!
- ComplexSystems 3y agoGreat demo, but this is wrong: "The phrase “potato, potahto” comes from a song titled “Let’s Call the Whole Thing Off”, written by George and Ira Gershwin for the 1937 film “Shall We Dance”, starring Fred Astaire and Ginger Rogers. The song humorously highlights regional differences in American English pronunciation. The lyrics go through a series of words with alternate pronunciations, like “tomato, tomahto” and “potato, potahto”. The idea is that, despite these differences, we should move past them, hence the line “let’s call the whole thing off”. Over time, the phrase has been adopted in everyday language to signify a minor disagreement or difference in opinion that isn’t worth arguing about." It's comparing American and British pronunciations, not different regional American ones. Also, "let's call the whole thing off" suggests they should break up over their differences, with the bridge and later choruses then involving a change of heart ("let's call the calling off off").
- TheHappyOddish 3y agoGlad everyone's excited about this (the voice capability), but did everyone miss tortise-tts and bark? These have been around 6+ months and are incredibly simply to hook up to OpenAI's APIs or a local LLM. What am I missing here?