Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sipjca
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
sipjca
3mo ago
Ha amazing, love to hear it
32.
▲
by
sipjca
3mo ago
Yep, but I am in the process of also porting NVIDIAs Sortformer for multi speaker diarization as well :) I’m not sure how many specific models will be supported as the library is more focused on transcription specifically. But the models wh
33.
▲
by
sipjca
3mo ago
Thanks! What an excellent question, I’m not sure I have a good answer. I kind of became an open source maintainer by accident as Handy became popular Certainly I am very lucky that quite a few people donate to Handy, and also some people an
34.
▲
by
sipjca
3mo ago
Yes, I’ve put a PR up on pypi for extra storage for CUDA but it has not been accepted yet afaik If there’s any issues or improvements on the bindings I would love help to make the DX the best it can be
35.
▲
by
sipjca
3mo ago
It very much depends on the hardware! An M4 max is being compared against a Ryzen 4750U with an integrated GPU! The M4 max has probably 10x the compute and memory bandwidth hahaha
36.
▲
by
sipjca
3mo ago
Yep the latest version has support! Virtually all of the SOTA open models are supported by Handy including the streaming ones like Nemotron Streaming Parakeet Unified Voxtral Mini Realtime If something you want is not supported, open an iss
37.
▲
by
sipjca
3mo ago
Thanks :)
38.
▲
by
sipjca
3mo ago
More or less yes, for whisper.cpp, just trying to make local transcription more accessible to anyone building an app, etc
39.
▲
by
sipjca
3mo ago
Hey, yep author and maintainer here! Certainly sponsors help and the wonderful community who donates to Handy as well! Mozilla AI was very helpful in getting this work off the ground. It was a pipe dream for me to build for Handy and they h
40.
▲
by
sipjca
3mo ago
Hey! It’s actually in progress right now, probably will come this week :)
41.
▲
by
sipjca
3mo ago
Hey! It’s actually in progress right now, probably will come this week :)
42.
▲
by
sipjca
3mo ago
Yes
43.
▲
by
sipjca
3mo ago
for real
44.
▲
Transcribe.cpp – ggml speech-to-text inference engine
(github.com)
2 points
by
sipjca
3mo ago
|
0 comments
45.
▲
Transcribe.cpp – ggml based transcription engine
(workshop.cjpais.com)
5 points
by
sipjca
3mo ago
|
0 comments
46.
▲
by
sipjca
3mo ago
Hasn't pretty much everyone from Nuvia left QC at this point?
47.
▲
by
sipjca
3mo ago
I’ve been doing something similar with less dedicated workflow and generally works great
48.
▲
by
sipjca
4mo ago
fwiw because of the relatively few activated params offloading to system RAM is quite feasible, you can see the endless amount of people doing this on r/localllama with qwen3.6 35a3b
49.
▲
by
sipjca
4mo ago
fair enough, i guess minimizing that surface area is important to begin with
50.
▲
by
sipjca
4mo ago
imo being digital native means that migrating to any machine should be basically trivial. working with the flow of the machines rather than customizing and ricing them because your a cool computer person or whatever i just want my computer
51.
▲
by
sipjca
4mo ago
im more surprised that more people don’t treat their computer as disposable anyway. that it could just be wiped at any moment and it wouldn’t matter. shit happens, could be stolen, broken, whatever. the computer should be able to be thrown
52.
▲
by
sipjca
4mo ago
inference code is effectively trivial to port at this time everyone understands cuda well enough anyway
53.
▲
by
sipjca
5mo ago
thats incredible
54.
▲
by
sipjca
6mo ago
Wondering similar. It certainly can run beyond 30 seconds but at some point I believe the output should degrade Plus you could do actual batch inference instead. Or if you must carry forward the context you could still do it linearly, but t
55.
▲
by
sipjca
6mo ago
Thank you!!
56.
▲
by
sipjca
6mo ago
I just don't have the bandwidth to run another project, maintaining Handy is hard enough on it's own, especially for free! I didn't just dismiss for no reason, I am a human! I have needs and I can't just sleeplessly stay
57.
▲
by
sipjca
6mo ago
Of course they are! Both are important and will be around and used for different reasons
58.
▲
by
sipjca
6mo ago
I don’t think it’s about literally shrinking the models via quantization, but rather training smaller/more efficient models from scratch Smaller models have gotten much more powerful the last 2 years. Qwen 3.5 is one example of this. T
59.
▲
by
sipjca
7mo ago
This is an argument, but it’s also fundamentally comparing a computer that works out of the box to one that doesn’t.
60.
▲
by
sipjca
7mo ago
That wasn’t the point. You’re a person who runs arch, that means most likely your requirements for a computer are VERY different than the target for this Mac. There’s always some other computer you can buy, but most people will just buy the
More ›