28 ms·
Firefox Voice
- asimilator 6y agoSpeech to text in the cloud is a hard no from me. Especially if it’s Google’s speech to text.
- ve55 6y agoDon't worry, it says they asked Google to not save everyone's voices! On a serious note, it does bother me how much Mozilla constantly uses Google, even when they have their own solutions They could easily choose not to, especially with their massive budget, but often don't. They have their own Voice API, but they use Google's. They have their own location API, but they use Google's ('use my location' sends your info to Google in Firefox). They have thrown Google analytics into browser components before, and used it on their own websites.
- iseanstevens 6y agoAgree, though Maybe google is giving them free/barter compute credits?
- ipsum2 6y agoI guess Mozilla's own speech to text (https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech) isn't good enough, so they have to use Google's?
- kevingadd 6y agoPresumably the reason they want you to opt-in to saving recordings is so they can train DeepSpeech.
- setzer22 6y agoBut DeepSpeech has already been trained with millions of data samples! I'd feel way better about it if they went for a slightly worse DeepSpeech based implementation, but kept it working in the free software spirit they have been known about for many years. Also, for desktop devices inference on DeepSpeech is cheap enough, so they could even go the extra mile and work on some Wasm magic to get offline recognition. That's the kind of work I'd expect from Mozilla! Not wiring up your data collection to the Google Cloud APIs and call it a day! I'm genuinely disappointed with them...
- posguy 6y agoThe audio Mozilla DeepSpeech is trained on is not very large (about 2000 hours) or diverse (eg: mostly native American English Male voices) and has very little ability to handle noise, accents or other errata. Comparatively, Baidu had 5000 hours of English to train their versions of DeepSpeech and DeepSpeech2 on, and thus had better results years ago. Google, Microsoft, IBM and other companies have users providing more audio samples on a daily basis, enabling much better quality speech to text. Mozilla's Common Voice project only has 1492hrs of validated English currently: https://commonvoice.mozilla.org/en/datasets https://commonvoice.mozilla.org/en/datasets
- bananaface 6y agoDeepSpeech is a lot less accurate and much slower than Facebook's FOSS offering, wav2letter, on equivalent data. If they want something competetive they'll need to drop DeepSpeech, or overhaul it. Common Voice is where the value is. Like, there's nothing stopping firefox from just using wav2letter. It's BSD-licensed. https://github.com/facebookresearch/wav2letter https://github.com/facebookresearch/wav2letter
- mikob 6y agoI've been working on this exact thing for Chrome the last 3 years: https://www.lipsurf.com https://www.lipsurf.com Anyone can make an open source plugin for it to do anything with voice (https://github.com/LipSurf/plugins https://github.com/LipSurf/plugins) I've wanted to port it to Firefox, but the HTML5 SpeechRecognition API (https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecognition https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecog...) is still not available. Why not just make the API available and leave this in addon territory for all developers?
- TedDoesntTalk 6y ago> Available as an extension for Firefox on desktop / laptop According to article, Firefox is offering this through an addon... obviously not using the SpeechRecognition API, however.
- kevingadd 6y ago"On some browsers, like Chrome, using Speech Recognition on a web page involves a server-based recognition engine. Your audio is sent to a web service for recognition processing, so it won't work offline." Seems like the answer for why the API isn't in Firefox. It's also not standardized, and is prefixed, so...
- aerovistae 6y agoI built something like this 7 years ago, it's called Hands Free for Chrome, a now languishing project that I lost interest in a long time ago, unfortunately. It made the top 10 of HN back then though! My site's design is not nearly as nice as yours. https://www.handsfreechrome.com/ https://www.handsfreechrome.com/ I just didn't get enough users or support to really care about it. But I wish you the best. It was an exciting thing to build and using it always felt futuristic to me. This is just so fascinating though. It's like seeing what could have been if I had been a better developer and found the dedication to really stick to the project in the longterm. Edit: I see we had the exact same idea! Your "tag" is my "map." Love it. One big difference is that mine was just a free project. I'd be super interested to know how many users you've got. I never had more than ~1100. From looking through your website mine was a much less intensive project. (Oh, CWS says 4000+ for you....wow, wonder how many are paid.) Edit2: Looking over your update history is almost nostalgic. "Fixed issue with overlapping commands -- delaying commands that are partial matches of other commands." Had to do the exact same thing! Edit3: We have so many overlapping command names that I wonder if you took inspiration from my project, almost. Either that or it's just a case of convergent evolution. Edit4: Suggestion for dictation: a way to alternate between a special character and actually writing the word. Doesn't look like there's a way to do ^ vs carrot or & vs ampersand. Something like "Enter special character caret". Maybe you already have a plugin for this though, idk. Edit5: God, this is so well architected! plugins and contexts are just fantastic ideas for this domain. Click by voice using hidden search-for-text is also a perfect solution to that problem. I wonder if this could be made more intelligent, i.e. "Click Submit in the sidebar on the left"-- challenging though. Edit6: Wow, just noticed someone else built something called "Handsfree for Web" somewhere along the way and theirs is ALSO way better than what I had built. Geez. Starting to feel bad about my awful website.
- ozten 6y agoIs there a preference to have text to speech via another service provider?
- hendersoon 6y agoGoogle charges 2.4 cents per minute for STT so there's no way Mozilla could afford to offer this service if it actually got popular. I mean, that obviously won't be an issue, but still.
- deleted 6y ago[deleted]
- kevingadd 6y agoMaybe the long-term plan is to run the trained model locally?
- yelloworangefog 6y agoI'm sorry, but cloud based speech recognition in itself would already be a red flag, even if Mozilla was doing it in-house. Outsourcing it to Google though? I feel like a company as ostensibly privacy-focused as Mozilla should really know better by now...
- krick 6y agoYeah, Mozilla is not what it used to be. I guess it's just wishful thinking that makes it so easy to forget.
- lawl 6y agoThey just pushed an update on android on me that disabled all addons except uBO because they don't support them yet, a few hours ago. That was the last straw that broke the camels back for me, after using FF since 1.0. Just rm -rf'd my firefox profile on my desktop a few minutes ago. I've defended them for a long time even if I didn't agree with everything they did, but they're so completely off the rails, enough is enough. Blink mono-culture it is then.
- rapnie 6y agoWhat browser are you going to use now? Any tips?
- lawl 6y agoSwitched over to Brave, as they have extensions on mobile cooking with POC videos on twitter. Hopefully they can accelerate it a bit. I think the HN crowd doesn't like brave though, idk, I do. IMO a lot of shit they get flak for are people not actually understanding. Like no, they don't replace ads on websites without their consent, or at least that's what they say.
- tialaramex 6y agoSo, because Firefox on your phone doesn't have X, you deleted it everywhere and installed a browser that... also doesn't have X? FWIW This is one of the symmetries when we have to do policy shifts like TLS 1.0 deprecation. Even though every major browser will implement the policy and has announced that, some fraction of users will feel "betrayed" and switch from one browser implementing the policy to another browser also implementing the policy. It's worse if you have defectors (e.g. back when Microsoft had Internet Explorer you could rely on IE being last to actually implement even if it had previously announced a timeline right in the middle of the ballpark, customers would push back and somewhere a Microsoft exec decides that hey, making a customer happy is more important than security) but it happens even for a more or less simultaneous policy change. Guaranteed some CAs will lose customers over Apple's 398 day certificate expiry policy, even though it affects every CA equally at the same moment.
- greggman3 6y agoI actually tried this with Siri while cooking yesterday. It's not there yet but I asked "Hey, Siri ... read me the synopsis of the movie Adam's Rib" and Siri proceeded to read a short synopsis of that movie. It worked on another but had to make me choose one of 7. It failed on the 3 try where I tried another movie, it gave me selections, when I picked one "read me the first one" it just repeated the title instead of telling me the synopsis.
- person_of_color 6y agoWhats wrong with deep speech?
- dsteinman 6y agoDeepSpeech is too large to run as a browser extension.
- person_of_color 6y agoHave they tried nanonets?
- remexre 6y agoGoogled those, and it looks like some company selling remotely-running models, which I think is what GP was referring to. Is there another technique that's been SEOd out by this company?
- Namidairo 6y agoThey appear to be running it against both Google and their own DeepSpeech instance then sending both results back in the client's telemetry?
- bennettfeely 6y agoSo do I yell out my password to log in to websites?
- jcims 6y agoI mean, who doesn’t?
- deleted 6y ago[deleted]
- canada_dry 6y agoMine's: "Get off my damn lawn". That way my neighbours aren't any the wiser.
- butz 6y agoYou could use Firefox Lockwise to store logins and login automatically when visiting website.
- krick 6y agoGoogle-worries aside, judging from the preview it's pretty slow. I'm not a super-fast typer, but these delays sure look like something that would discourage me from actually using it. Maybe it's not even that it's slow, just that the delays are super-obvious somehow, all these disruptive animations and such.
- causality0 6y agoVoice browsing on Windows is exactly what I don't need. I'd have a lot of use for being able to search the internet by voice and have the browser read an article to me while my phone is mounted to my dashboard. Without having to configure my whole phone for visual impairment, that is.
- mlindner 6y agoMozilla needs to come more around to Apple's way of thinking. These things need to be done locally on the device, not farmed out to some cloud. Use the cloud (CDN) to deploy the software, but run the software locally.
- skizziepop 6y agoOpera had this a decade ago. RIP
- svnpenn 6y agoSo let me get this straight. They break global keyboard shortcuts, which people could use to play/pause media on different websites three years ago: https://bugzilla.mozilla.org/show_bug.cgi?id=1411795 https://bugzilla.mozilla.org/show_bug.cgi?id=1411795 and instead of fixing that, they introduce this shit? Fuck Mozilla. I am already using Waterfox, looks like that wont be changing any time soon.
- mceachen 6y agoUnless they only have one engineer, adding a new feature is not at the expense of fixing one bug.
- ethanwillis 6y agoYet somehow they haven't allocated 1/N engineers to fix a 3 year old accessibility bug.
- chii 6y agoWhether they are non-profit or not, they still run like a business, and the business decision to allocate resources are likely to use the same metrics as any other businesses - maximize profit protential (however you define it). It may just be that fixing accessibility bugs just doesn't produce much "profit".
- ohgodplsno 6y agoAh, yes, the classic unrelated Mozilla hate thread. Wherein a user finds a single grievance that doesn't matter to 99% of the users, uses it as a warhorse and brings it into fully unrelated threads.
- deleted 6y ago[deleted]
- pixxel 6y agoHad to read the privacy policy to see they use Google. >We share your audio recording with Google Cloud’s speech-to-text service to assist us in processing and carrying out your commands. Audio recordings are shared without personally identifiable metadata, and we’ve instructed Google’s service not to retain the audio or transcript associated with a command after it processes the command
- pmontra 6y agoIt's in the FAQ section of the page under "How is my audio processed?"
- mike_ivanov 6y agoI don't understand the utility of it. Yes, I can see how this might be considered cool and hip, but.. which my problem as a user does it solve, exactly?
- a_bonobo 6y agoSpeech-based control is almost mandatory for people with various disabilities (the blind, those with hand-movement problems/disabilities) and the elderly, a huge chunk of the population.
- Legogris 6y agoI regularly encounter people relying on voice-to-text to search or browser. especially people like cab drivers, but also various others who don't have the no- hands restriction. It's not the kind of people I would expect to comment here, though.
- vffhfhf 6y agoSee 2 sentence and I know its not for you. I have been waiting fot this for years. I want to use voice to control browser. Like I am doing something and say " hey firefoxy / googley open reddit on the side". You got to broaden your thinking. Other people like other shit.
- deleted 6y ago[deleted]
- djsumdog 6y agoI've looked into open source voice assistants before. I found mycroft, Jarvis and a few others, but either got bogged down in dependencies or configuration. Many supported shipping your data to Google or Amazon if you configured it, or an open source voice recognition tool. I hate this idea that our voice has to be shipped somewhere to be processed. I remember a lot of the speech-to-text tools in the early 2000s weren't all that great (they needed a lot of training), but why haven't we been able to advance on-device processing? Why is everything done in "the cloud." So the only way to semi-accurately do voice recognition is to source algorithms that re-train off of millions of people? We have processors in our desktops and laptops that dwarf that compute power by leaps and bounds. We should be looking to Star Trek TNG level voice processing, on each individual device, without some central mainframe. But marketing, advertising revenue, data mining, free (as in beer) software that pumps your data like an oil rig, efficiency in data centre (cloud) design .. all these factors have led to these powerful little Intel/ARM/Ryzen chips to be nothing more than thin clients when they're not playing games. If Mozilla really wanted to make something amazing and in the spirit of Firefox, give us an experiment where voice processing is done on our devices. Even if it meant I needed to download a 230GB data set, I'd gladly do it, if it could remotely help in getting away from these data silos.
- deleted 6y ago[deleted]
- Vinnl 6y agoThat is what they are working on, but that needs high-quality training data: https://commonvoice.mozilla.org/en https://commonvoice.mozilla.org/en This is just another way of gathering that data. (If consented to.)
- posguy 6y agoMozilla DeepSpeech is trained on about 2000 hours of audio that is mostly spoken by American males. It has little ability to handle noise or accents and has a 5.97% Word Error Rate on LibriVox (which is noiseless, plain spoken english). Meanwhile, Google, Microsoft & IBM have tons of fresh audio coming in constantly to use in augmenting their models. Baidu was able to build a competitive English Speech to Text model with 5000 hours of quality audio to train against. Mozilla did create Common Voice to address this serious data gap, but it has only collected 1492hrs of validated English audio: https://commonvoice.mozilla.org/en/datasets https://commonvoice.mozilla.org/en/datasets
- nafts 6y agoWho would use such a gimmick? I certainly wouldn't...
- nafts 6y agoJust another useless gimmick that you try once, then never bother to use again. Imagine talking to your browser at work.
- dhaavi 6y ago"We’ve instructed the Google Speech-to-Text engine to NOT save any recordings." Hahaha! :D Thanks for the good laugh.
- input_sh 6y agoBigger players have the leverage to get companies to do something they don't do out-of-the-box. They can contractually oblige them to do that, as well as sue each other if one side breaks its part of the deal.
- Terretta 6y agoIf they could guarantee Google doesn’t, they’d say that. Thus the much lesser claim “instructed” which carries no such assurance. I agree with and applaud their truthful choice of words. There’s no such thing as “contractually oblige”. // To keep yourself cognitive, consider one party receiving a National Security Letter (“NSL”) with gag order.
- lights0123 6y agoThat's not anything special—in fact, it's the default. You may optionally enable it and get discounted pricing: https://cloud.google.com/speech-to-text/docs/data-logging https://cloud.google.com/speech-to-text/docs/data-logging
- snoopfab 6y agoYou should try rhasspy. It's open source. It respects your privacy by using offline services. Fully customizable (each service can be replaced by another) All the services are containerized for easy installation and is available for several architectures such as arm ( on a raspberry pi). There is even n Option to use Mozilla deepspeech tts service. https://rhasspy.readthedocs.io/en/latest/ https://rhasspy.readthedocs.io/en/latest/
- miguelmota 6y agoThe default keyboard shortcut wasn't working and it was opening a different extension instead. I went to the voice extension settings and thought it was bad ux how you have to enter the case-sensitive keyboard shortcut names instead of pressing the keys to read the keys.
- deleted 6y ago[deleted]
- AsyncAwait 6y agoI see a lot of skeptical voices here, (somewhat warranted, given it's a voice assistant technology), but the fact remains that if we want open, on-device voice recognition, we'll have to do the work and donate sample data. This extension is trying to provide some useful functionality in the hopes that Mozilla gets more data for https://commonvoice.mozilla.org https://commonvoice.mozilla.org I'd at least consider recording your voice, especially if you're a non-native English speaker, like myself, have an accent etc. It took many years for free software to start to take on the smartphone segment, with previous efforts, (including by Mozilla), failing and only now PinePhone & Librem 5 giving it another go, but unless you're a super hardcore enthusiast, you carry an iPhone/Android today. I see this as a way to push back on the likes of Amazon, Google and Apple with this. If regular Firefox users are able to use an on-device, privacy respecting voice assistance and other open-source projects can use Mozilla's tools and datasets to build compelling competitors to Alexa, I'd see that as proof that free software is able to address new, emerging markets too.
- hnarn 6y ago> This extension is trying to provide some useful functionality in the hopes that Mozilla gets more data for https://commonvoice.mozilla.org https://commonvoice.mozilla.org This is awesome! I love contributing to open source initiatives like this. I'm also a non-native speaker so hopefully I'll add some color to the voices recorded.
- deleted 6y ago[deleted]
- panpanna 6y ago> if we want open, on-device voice recognition, we'll have to do the work and donate sample data. Fair enough, but is any stopping Mozilla from _also_ selling your voice data to third parties including advertisers and commercial ML interests? (Asking this because Firefox has started sharing our data with leanplum)
- AsyncAwait 6y agoAs long as they're upfront about what they collect and how to not send them data, am fine with it. Obviously opt-in is much proffered to opt-out, but from the Firefox Voice site, it clearly states sharing data is opt-in. My guess would be there isn't much of a point selling voice data from an open dataset. Also, since the code is in the open, it would be relatively easy to spot if they were sending data to somewhere they do not or record when they shouldn't.
- sradman 6y agoRecent HN thread Thoughts on Voice Interfaces [1] about a blog post by one of the Firefox Voice engineers. [1] https://news.ycombinator.com/item?id=24040539 https://news.ycombinator.com/item?id=24040539
- maps7 6y agoI am a Mozilla supporter so I am happy to support this.
- Aachen 6y ago> When you make a request using Firefox Voice, the browser captures the audio and uses cloud-based services to transcribe and then process the request. Is it that hard to do local processing, either due to computational power or storage requirements? Or is it just more convenient for them to do it this way? Edit: this comment in another subthread kind of answered the question: https://news.ycombinator.com/item?id=24098950 https://news.ycombinator.com/item?id=24098950 If I'm drawing the right conclusion, it's a bit of both: hundreds of megabytes of storage is fine for most people but not everyone, and while I probably wouldn't listen to the latest and greatest artists (and binary diffs are a thing, small additions aren't that large), it is convenient for devs to just push it to a server and be done rather than pushing model updates to everyone all the time. Edit2: https://news.ycombinator.com/item?id=24096836 https://news.ycombinator.com/item?id=24096836 Wait, what?! The data is all sent to Google? I was thinking of using this for their sake (opting into using my data for common voice) but this is an instant deal breaker.
- input_sh 6y agoAlright, gave it a shot. First impressions: * "Make me laugh" always brings me to the same YouTube video. * Had pretty much no issues with the default prompts. It was able to find some challenging Spotify playlists, open random websites (including non-standard English domains ones when I spelled them out). * "Read this page" uses an awful TTS engine, which is a shame considering that I might actually use this feature on a somewhat regular basis. I'm assuming it uses whatever it detects on the OS level, and so far I haven't bothered with finding a better one (on Ubuntu, if you know of one, please suggest). * "Set a timer for X min" works just fine, which is probably the only thing I use Google's assistant on my phone (or whatever it's name might be now). * I like the idea of routines in the app settings, which is supposed to tie multiple queries together. I could see myself using it for something like a morning routine (tell me what time it is, give me weather info, read me news, etc.)
- Eyght 6y agoI tested the phrase "go to kyle's channel on twitch" and it actually went to twitch.tv/kyle - which I found impressive.
- swiley 6y agoI still haven't understood how any of this is an improvement of a shell.
- madacol 6y agoNo one here has mentioned this. But I believe speech recognition will not take-off until it understands whispering speeches. Vocal chords strain makes current solutions unsustainable for continuous usage
- nsriv 6y agoThe Google Recorder app on Pixel phones (and I'm pretty sure general Android release) does super accurate on-device transcription, for what it's worth.
- Mandatum 6y agoDupe? https://news.ycombinator.com/item?id=23904846 https://news.ycombinator.com/item?id=23904846