Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
skoocda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
91.
▲
by
skoocda
10y ago
We've chatted - just an update that I'm implementing diarization this weekend :)
92.
▲
by
skoocda
10y ago
This was my thought, but due to "Global Gasoline Demand Has All but Peaked" They mean "global gasoline demand has peaked". They have literally said the opposite. Is this the new way to write headlines? It's like peo
93.
▲
by
skoocda
10y ago
Has anyone been able to use Lambda for relatively high-memory load applications? (~2GB+ RAM) That's our biggest restraint at the moment, so far I haven't seen any good options.
94.
▲
by
skoocda
10y ago
I'm with you there. However, it's important to remember how fluid and dynamic verbal discussion can be. It's a great format for creation, even if it's lacking in consumption. Luckily- my startup; Spreza, transcribes podc
95.
▲
by
skoocda
10y ago
Lots of overlapping. It's a sliding window function. Ballpark for most algorithms: 10 ms of new audio, 90 ms of old audio.
96.
▲
by
skoocda
10y ago
This is neat. Looks like you need more data though!
97.
▲
by
skoocda
10y ago
Pretty much SOA; most end-to-end systems use these spectrograms on short time slices. The alternative is mel-frequency cepstral coefficients, which are used more in GMM-HMM speech recognition than for DNN.
98.
▲
by
skoocda
10y ago
I mean, it's been a big couple weeks of releases. Between Surface Studio, the new speech recognition results, Teams, and this VSCode update (barring the npm issue) they seem to be really hitting their stride, in a way that hasn't
99.
▲
by
skoocda
10y ago
Furthermore, there were 32,500 suicides in 2013. I wouldn't say the US is in any position to make the choice of prioritizing life over jobs, not yet anyways; there's still no way to live without a job. Market forces will decide, o
100.
▲
by
skoocda
10y ago
To add to this, Daniel Povey's lectures [0] are fairly useful (although not exactly colourful, per se). Would definitely second the FST paper. You'll get shunned from the forums if you haven't read that one. [0] http:/&
101.
▲
by
skoocda
10y ago
There's been very little use of spectral peaks in LVCSR for the past ~10 years.
102.
▲
by
skoocda
10y ago
Since Kaldi is a toolkit, it can be used to build nearly any ASR architecture. See here [0] for a comprehensive comparison of the Word Error Rate of various architectures. [0]: https://github.com/syhw/wer_are_we
103.
▲
by
skoocda
10y ago
Very similar accuracy, timing and alignment. We don't do speaker diarization at all because the results seem consistently weak, even among the competitors such as Speechmatics. I'd hazard to say our web editor is much better for p
104.
▲
by
skoocda
10y ago
We're close to an alpha release of Spreza, which might be relevant to your question. Look us up! dm me if you've got questions
105.
▲
by
skoocda
10y ago
The novel elements of this have already been released individually, particularly the lattice-free MMI training which can be run through Kaldi's nnet3 configuration. Also, don't assume that a 0.4 % increase means drastically better
106.
▲
by
skoocda
10y ago
During extensive, complex discussions like these, it becomes readily apparent that while the HN/Reddit-style threaded text forum is the best format I know of, it's still woefully inadequate. There will soon be 1000 comments on thi
107.
▲
by
skoocda
10y ago
LibriVox
108.
▲
by
skoocda
10y ago
You can't type faster than you can talk. Trained stenotypists come close, but there's nobody who can consistently hit 130-150 wpm on a QWERTY or DVORAK keyboard.
109.
▲
by
skoocda
10y ago
I agree 100%. Right now we don't even know the number of microphones our devices have, much less the contexts in which they may be active. Speech recognition protocols should also integrate low-level encryption standards to ensure acce
110.
▲
by
skoocda
10y ago
Czech out my project, Spreza. We're not targeting voice input, per se, but instead automated transcription. Nonetheless, the correction issue persists in both use cases. We solve it via a homonym suggestion drop-down box. Double-click
111.
▲
by
skoocda
10y ago
I'm surprised nobody has mentioned Angular 2 yet. It's new, so don't expect tons of support. Nonetheless, it's there as an option!
112.
▲
by
skoocda
10y ago
This is incredible work. The low price point should create and burgeoning open source / hacker dev community, which in turn, will hopefully greatly accelerate 3D perception algorithms. My recent EE capstone project was in the area of a
113.
▲
by
skoocda
10y ago
I'm currently developing a system along those lines. Not publicly available quite yet, but should be soon. Bug fixing at the moment.
114.
▲
by
skoocda
10y ago
This is an excellent resource for a new startup such as my own. We're interested in evaluating the less-tangible choices that need to occur to develop company culture, and this helps us see these options very clearly. I hope more compa
115.
▲
by
skoocda
10y ago
* I should clarify as well, for your first question. State of the art is generally either IBM's 2016 English Conversational Telephone Speech Recognition System [0], or Baidu's DeepSpeech2 [1]. IBM's is a little more complex,
116.
▲
by
skoocda
10y ago
I'm working on this at the moment, in a way that also uses the edited transcripts to train our ASR system to perform better for later sessions. The difficult part is the speaker diarization, however. Multiple people talking at once req
117.
▲
by
skoocda
10y ago
I'm working on it, making more of a general purpose API for speech recognition to plug in during Hangouts calls, twitch streams, or any pre-recorded media. Called Spreza
118.
▲
by
skoocda
10y ago
ninite.com Your one stop shop to get a Windows PC back on its feet after a wipe. Ninite always works great for me. The application packages are a single click to install, and it really helps minimise the clammy, sticky feeling of using IE E
119.
▲
by
skoocda
10y ago
It depends on your use case. Full conversational speech recognition can't adequately fit in memory on a mobile device, but smaller packages like PocketSphinx can. Kaldi, also mentioned in the comments, can serve various use cases, but
120.
▲
by
skoocda
10y ago
Well we're trying to design it for simplicity, with the hopes that everyone can be a client!
More ›