11 ms·
Neat, https://github.com/openai/whisper https://github.com/openai/whisper - they have open-sourced it, even the model weights, so they are living up to their na
by pen2l 4y ago
Neat, https://github.com/openai/whisper https://github.com/openai/whisper - they have open-sourced it, even the model weights, so they are living up to their name in this instance.
The 4 examples are stunningly good (the examples have speakers with heavy accents, speaking in foreign language, speaking with dynamic background noise, etc.), this is far and away better than anything else I've seen. Will be super curious to see other folks trying it out and seeing if it's as robust as it seems, including when confronted with audio speech with natural tics and uhhh's and uhmm's and everything in-between.
I think it's fair to say that AI-transcription accuracy is now decidedly superior to the average human's, what the implications of this are I'm not sure.
- bambax 4y agoThe French version is a little contrived. The speaker is a native speaker, but the text is obviously the result of a translation from English to French, not idiomatic French. I will try to put the code to the test, see how it goes.
- pen2l 4y agoInteresting, I'm a non-native French speaker, the original French piece struck me as being entirely normal (but maybe it was just the perfect French accent that swayed me). Can you please point out what he said which wasn't idiomatic or naturally-worded French?
- _plg_ 4y agoAt the start, the "Nous établissons" part, for example. You wouldn't write that if you were starting scratch from French.
- otikik 4y agoThat's the first thing that I discovered when I visited Paris for the first time. No one says "Nous", there, ever. Perhaps the politicians, while giving a speech. Everyone else uses the more informal "On". I felt duped by my French classes.
- mijamo 4y agoOlder generations sometimes do. My grandma and her sisters nearly never uses "on". It is often used for larger groups or when the group is not very personally connected. For instance when talking about your company doing something you will often use "nous". I would also use "nous" to refer to the whole list of invitees to a wedding. And in formal contextes like research papers, reports etc. You would never use "on", always "nous".
- not_math 4y agoYou can see from the transcript where the model made some errors, for example: > We distribute as a free software the source code for our models and for the inference [...] Should be > We are open-sourcing models and inference code [...] Another example > We establish that the use of such a number of data is such a diversity and the reason why our system is able [...] Should be > We show that the use of such a large and diverse dataset leads to improved robustness [...]
- bambax 4y agoLittle details. The second sentence is really bizarre: > Nous établissons que l'utilisation de données d'un tel nombre et d'une telle diversité est la raison pour laquelle le système est à même de comprendre de nombreux accents... It doesn't sound natural at all. An idiomatic formulation would be more along the lines of: Le recours à un corpus [de données] si riche et varié est ce qui permet au système de comprendre de nombreux accents (With 'corpus', 'données' is implied.) Of course this is just an example, and I'm sure other French speakers could come up with a different wording, but "données d'un tel nombre et d'une telle diversité" sounds really wrong. This is also weird and convoluted: > Nous distribuons en tant que logiciel libre le code source pour nos modèles et pour l'inférence, afin que ceux-ci puissent servir comme un point de départ pour construire des applications utiles It should at least be "le code source DE nos modèles" and "servir DE point de départ", and "en tant que logiciel libre" should placed at the end of the proposition (after 'inférence'). Also, "construire" isn't used for code but for buildings, and "applications utiles" is unusual, because "utiles" (useful) is assumed. "...pour le développement de nouvelles applications" would sound more French.
- deleted 4y ago[deleted]
- aGHz 4y agoThat's interesting, as a québécois I don't agree with any of this. The only thing that raised an eyebrow was "est à même de", but if turns out it's just another way of saying "capable de", I guess it's simply not a common idiom around here. Aside from that, I found the wording flowed well even if I personally would've phrased it differently.
- slim 4y agoMistery solved. It was a quebecois
- mazork 4y agoGonna have to agree with the other reply, as a french-canadian, except for "servir comme un point de départ" which should be "servir de point de départ", that all sounds perfectly fine.
- octref 4y agoI'm interested in building something with this to aid my own French learning. Would love to read your findings if you end up posting it somewhere like twitter/blog!
- bambax 4y agoI'm playing with a Colab posted in this thread (https://news.ycombinator.com/item?id=32931349 https://news.ycombinator.com/item?id=32931349), and it's incredibly fun and accurate! I tried the beginning of L'étranger (because you seem to be a fan of Camus ;-) Here's the original: > Aujourd’hui, maman est morte. Ou peut-être hier, je ne sais pas. J’ai reçu un télégramme de l’asile : « Mère décédée. Enterrement demain. Sentiments distingués. » Cela ne veut rien dire. C’était peut-être hier. > L’asile de vieillards est à Marengo, à quatre-vingts kilomètres d’Alger. Je prendrai l’autobus à deux heures et j’arriverai dans l’après-midi. Ainsi, je pourrai veiller et je rentrerai demain soir. J’ai demandé deux jours de congé à mon patron et il ne pouvait pas me les refuser avec une excuse pareille. Mais il n’avait pas l’air content. Je lui ai même dit : « Ce n’est pas de ma faute. » Il n’a pas répondu. J’ai pensé alors que je n’aurais pas dû lui dire cela. En somme, je n’avais pas à m’excuser. C’était plutôt à lui de me présenter ses condoléances. Here's the transcription: > Aujourdhui, maman est morte, peut être hier, je ne sais pas. J''ai reçu un télégramme de l''asile. Mère décédée, enterrement demain, sentiment distingué. Cela ne veut rien dire. C''était peut être hier. > L''asile de Vieillard est à Maringot, à 80 km d''Alger. Je prendrai l''autobus à deux heures et j''arriverai dans l''après midi. Ainsi, je pourrai veiller et je rentrerai demain soir. J''ai demandé deux jours de congé à mon patron et il ne pouvait pas me les refuser avec une excuse pareille. Mais il n''avait pas l''air content. Je lui ai même dit, ce n''est pas de ma faute. Il n''a pas répondu. J''ai alors pensé que je n''aurais pas dû lui dire cela. En somme, je n''avais pas à m''excuser. C''était plutôt à lui de me présenter ses condoléances. Except for the weird double quotes instead of the single apostrophe ('), it's close to perfect, and it only uses the "medium" model. This is extremely exciting and fun! Happy to try other texts if you have something specific in mind!
- bambax 4y agoTried again with Blaise Pascal -- the famous fragment of a letter where he says he's sorry he didn't have enough time to make it shorter. Original: > Mes révérends pères, mes lettres n’avaient pas accoutumé de se suivre de si près, ni d’être si étendues. Le peu de temps que j’ai eu a été cause de l’un et de l’autre. Je n’ai fait celle-ci plus longue que parce que je n’ai pas eu le loisir de la faire plus courte. La raison qui m’a obligé de me hâter vous est mieux connue qu’à moi. Vos réponses vous réussissaient mal. Vous avez bien fait de changer de méthode ; mais je ne sais si vous avez bien choisi, et si le monde ne dira pas que vous avez eu peur des bénédictins. Transcription: > Mes rêves errent pères, mais l'detre navais pas accoutumé de se suivre de si près ni d'detre si étendu. Le peu de temps que j'sais eu a été cause de l'de l'de l'de autre. J'sais n'detre plus longue que parce que j'sais pas eu le loisir de la faire plus courte. La raison qui m'sa obligée de me hâter vous est mieux connue qu'moi. Vos réponses vous réussissaient mal. Vous avez bien fait de changer de méthode, mais je ne sais pas si vous avez bien choisi et si le monde ne dira pas que vous avez eu peur des bénédictes. Here there are many more mistakes, so many that the beginning of the text is unintelligible. The language from the 17th century is probably too different. Still on the "medium" model, as the large one crashes the Colab (not sure how to select a beefier machine.) Still fascinating and exciting though.
- anigbrowl 4y agoIt was already better. I edit a podcast and have > a decade of pro audio editing experience in the film industry, and I was already using a commercial AI transcription service to render the content to text and sometimes edit it as such (outputting edited audio). Existing (and affordable) offerings are so good that they can cope with shitty recordings off a phone speaker and maintain ~97% accuracy over hour-long conversations. I'm sure it's been an absolute godsend for law enforcement other people who need to gather poor-quality audio at scale, though much less great for the targets of repressive authority. Having this fully open is a big deal though - now that level of transcription ability can be wrapped as an audio plugin and just used wherever. Given the parallel advances in resynthesis and understanding idiomatic speech, in a year or two I probably won't need to cut out all those uuh like um y'know by hand ever again, and every recording can be given an noise reduction bath and come out sounding like it was recorded in a room full of soft furniture.
- thfuran 4y ago>~97% accuracy over hour-long conversations. I'm sure it's been an absolute godsend for law enforcement 97% accuracy means roughly three or four errors per minute of speech. That seems potentially extremely problematic for something like law enforcement use where decisions with significant impact on people's day and/or life might be made on the basis of "evidence".
- golem14 4y agoOne would think that the few crucial bits of information gleaned are listened to manually, and the machine translation is not the only thing the judge or a jury sees.
- thfuran 4y agoYou have absolutely ruined someone's day way before they're sitting in front of a jury.
- formerly_proven 4y ago
- soheil 4y ago
- space_fountain 4y agoThis seems to not be true for McDonald: https://www.snopes.com/fact-check/mcdonalds-100-beef/ https://www.snopes.com/fact-check/mcdonalds-100-beef/
- soheil 4y ago
- space_fountain 4y agoThis isn't exactly a hard story to fact check. There is 0 evidence for this in either the reddit thread or really anywhere? If they were willing to lie about the company name why not just lie about the beef in their burgers it would be equally scandalous
- soheil 4y agoThe company name could be 100% legit, there is nothing stopping you from a forming a company with that name and not even sell beef.
- jefftk 4y agoIf this was more than an urban legend someone would be able to dig up a company with this name and some indication that McD was working with them.
- pessimizer 4y agoSomething being possible to do isn't enough evidence for rational people to believe that it happened. From my perspective, it's possible that you're Iron Mike Tyson, or that you died after your last comment and this one was posted by the assassin who killed you.
- suyash 4y agoMore of this is welcome, they should live up their name and original purpose and share other models (code, weights, dataset) in the open source community as well.
- Workaccount2 4y agoCan't wait to see twelve new $49.99/mo speech parser services pop up in the next few weeks.
- quickthrower2 4y agoMake hay before Google gives away free hay. That said there is value in integration of this into other things.
- quickthrower2 4y agoThis has been running on my laptop all day for a 15 min mp3! Definitely not cheap to run then (wont imagine how much AWS compute cost is required).
- deleted 4y ago[deleted]
- knaik94 4y agoIt seems far from good with mixed language content, especially with English and Japanese together. The timestamps are far from perfect. It's far from perfect. It's nowhere close to human for the more ambiguous translations that depend on context of word. It's far below what anyone that spoke either language would consider acceptable. Maybe it's unfair to use music, but music is the most realistic test of whether it's superior to the average human.
- quickthrower2 4y agoSome music is hard for even people to make out the lyrics to.
- darepublic 4y ago> Neat, https://github.com/openai/whisper https://github.com/openai/whisper - they have open-sourced it, even the model weights, so they are living up to their name in this instance. Perhaps it will encourage people to add voice command to their apps, which can be sent to gpt3
- pabs3 4y agoIs the training dataset and code open too?
- catfan 4y ago
- DLeychIC 4y ago