4 ms·
I've been working on a similar use case at work (going from discoursive speech to cli-like commands, using a semi-rigid language), and I didn't find any off-the
by setzer22 6y ago
I've been working on a similar use case at work (going from discoursive speech to cli-like commands, using a semi-rigid language), and I didn't find any off-the-shelf purely ML-based solution that would work for us.
In my experience, I've found any services claiming to do deep learning produced far worse results than what we could get with simple approaches. That is, when faced with non-grammatical sentences (or rather, sentences with a different grammar than English's). Of course that's because models are not typically trained with this use-case in mind! But the fact that you need a huge load of data to even slightly alter the expected inputs of the system, to me, was a deal breaker.
For the specific case of programming with voice, Silvius comes to mind. It's built and used by a developer with this same problem. It's a bit wonky having to spell words sometimes with alpha-beta-gamma speech, and it won't work without some customization, but on the other hand it's completely free and open source: https://github.com/dwks/us https://github.com/dwks/us
- lunixbochs 6y agoYou've see seen openai's new english -> bash demo right? That said, Silvius is more of demo than a product, the IMO best voice programming options right now are (in alphabetical order): - Caster/dragonfly (fully open-source if you use daanzu's Kaldi engine, which is way better than Silvius afaik, I think even the creator of silvius uses dragonfly with dragon instead of using silvius) - Serenade (fully commercial, I haven't looked at it much recently but biggest caveats afaik are accuracy, the fact speech recognition is web based, and it's restricted to specific languages and IDEs while caster/talon are for full system control and not just programming) - Talon (my project, semi-commercial as I work on it full time and draw income from it but aim to give all necessary features away for free, some benefits include a fully offline and open-source speech recognition engine, and I have other bonuses like eye tracking and noise recognition)
- setzer22 6y ago> You've see seen openai's new english -> bash demo right? Not yet, but will do, thanks! However, I'd still be hesitant to build a product on top of that: Does voice to bash help us if we now want to do, say, voice to python? At least we'd need to re-train the system with completely new data, and even if we use transfer learning to our advantage, it's not an easy task. There's also no guarantees that the chosen neural network architecture that works for bash, will work the same for any programming language (think of a radically different syntax, like Lisp for example). The training must also be re-done for any variation in the input format to some extent. i.e., accent, expected background noise levels, and of course (human speaker) language. ML has its use case, but I typically see these nice demos as that, demos. When you have to build a real product and solve user problems, you can't rely on a black box doing what you want.
- lunixbochs 6y agoI think some of your comment does not apply to GPT3 in the conventional sense, they did not do any specialized training for text2bash afaik. They've been tooting about "one shot learning". If their demo is to be believed, text2bash is just their _massive_ generic model + a few lines of examples. Also they do have a related Python demo: https://news.ycombinator.com/item?id=23507145 https://news.ycombinator.com/item?id=23507145 Speech is a completely different stack to this, but honestly (english) speech is much more of a solved problem here than general knowledge.