Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vvolhejn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Show HN: Samuel, a Silly Speech Model
(samuel.vvolhejn.com)
3 points
by
vvolhejn
2mo ago
|
0 comments
2.
▲
by
vvolhejn
3mo ago
A lot actually, since the model has all information given to it in the four views, it doesn't have to deal with any "theory of mind" of modeling the other players or being consistent over long times. See [1], there's a v
3.
▲
by
vvolhejn
3mo ago
Václav here from the team, we're happy to answer questions :) The most surprising part to me is the auto-recovery behavior we mention at the end of the blog post, since any other model I've seen always stays diverged once it goes
4.
▲
by
vvolhejn
9mo ago
> Is this just sort of expected for these models? Should users of this expect only truncation or can hallucinated bits happen too? Basically, yes, sort of expected: we don't have detailed enough control to precent it fully. We can m
5.
▲
by
vvolhejn
9mo ago
Václav from Kyutai here. Thanks for the bug report! A workaround for now is to chunk the text into smaller parts where the model is more reliable. We already do some chunking in the Python package. There is also a more fancy way to do this
6.
▲
by
vvolhejn
9mo ago
Václav from Kyutai here. Yes the original naming scheme was from Les Miserables, glad you noticed! We just stuck to Alba because that's the real name of the voice actor that provided the voice sample to us (see https://huggi
7.
▲
by
vvolhejn
1y ago
I had no idea this existed, the internet is amazing
8.
▲
by
vvolhejn
1y ago
There's a great blog post from Sander Dieleman about exactly this - why do we need a two step pipeline, in particular for images and audio? https://sander.ai/2025/04/15/latents.html For text, there are a
9.
▲
by
vvolhejn
1y ago
I train for 1M steps (batch size 64, block size 2048), which is enough for the model to more-or-less converge. It's also a tiny model for LLM standards, with 150M parameters. The goal wasn't really to reach state of the art but to
10.
▲
by
vvolhejn
1y ago
Author here, thanks for the kind words! I think such a physics-based codec is unlikely to happen: in general, machine learning is always moving from handcrafted domain-specific assumptions to leaving as much as possible to the model. The m
11.
▲
by
vvolhejn
1y ago
merci, will fix tomorrow
12.
▲
by
vvolhejn
1y ago
I use Descript to edit videos/podcasts and it works great for this kind of thing! It transcribes your audio and then you can edit it as if you were editing text.
13.
▲
by
vvolhejn
1y ago
I had a part about this but I took it out: for compression, you could keep the embeddings unquantized and it would still compress quite well, depending on the embedding dimension and the number of quantization levels. But categorical distri
14.
▲
by
vvolhejn
1y ago
I don't know about linear models, but this kind of hierarchical modelling is quite a common idea in speech research. For example, OpenAI's Jukebox (2020) [1], which uses a proto-neural audio codec, has three levels of encoding tha
15.
▲
by
vvolhejn
1y ago
Author here. Speech-to-text is more or less solved, it's easy to automatically get captions including precise timestamps. For training Moshi, Kyutai's audio LLM, my colleagues used whisper-timestamped to transcribe 7 million hours
16.
▲
by
vvolhejn
1y ago
Author here. There are a few reasons, but the biggest one is simply the compression ratio. The OG neural audio codec SoundStream (whose first author is Neil, now at Kyutai) can sound decent at 3kbps, whereas MP3 typically has around 128kbps
17.
▲
by
vvolhejn
1y ago
Author here. I think it's more of a capability issue than a safety issue. Since learning audio is still harder than learning text, audio models don't generalize as well. To fix that, audio models rely on combining information from
18.
▲
Show HN: Sine Wave Speech, a real-time audio effect in Rust->WASM
(sinewavespeech.com)
10 points
by
vvolhejn
2y ago
|
0 comments
19.
▲
by
vvolhejn
3y ago
Yes! The fact that isochrones become circles is one of my favorite things about this. I also discuss this in the video ( https://youtu.be/rC2VQ-oyDG0?t=134 ) I made about these.
20.
▲
by
vvolhejn
3y ago
Sorry, that must be a bug. On desktop the map is large enough to be visible in its entirety but it is supposed to pan on mobile... I'll try to fix that
21.
▲
by
vvolhejn
3y ago
Indeed, you can't do this perfectly – it's an approximation. I just use one of the two directions as the travel time because if I wanted to symmetrize I'd have to call the Google Maps API twice . Asymmetry aside, it's al
22.
▲
by
vvolhejn
3y ago
Oh yes, unfortunately, you can't do this perfectly. There are some graphs that cannot be embedded in Euclidean space in any number of dimensions, e.g. a 4-cycle with distance measured by path length. It's a good-enough approxima
23.
▲
by
vvolhejn
3y ago
Yes! This is a less practical but more funky version of isochrone maps.
24.
▲
by
vvolhejn
3y ago
Cool! I'd love to have a look at that, but the two links in the paper seem to be broken :( http://roadlessforest.eu/map.htm and http://www.map.ox.ac.uk/accessibility_to_cities/
25.
▲
Maps that show time instead of space
(spacetime-maps.vercel.app)
148 points
by
vvolhejn
3y ago
|
44 comments