Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
soycaporal
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
soycaporal
3mo ago
mainly this is a portability and compatibility solution.. I think with GPU available deployments, it wouldn't benefit from ternary distillation. It would be a different problem domain.
2.
▲
by
soycaporal
3mo ago
amazing.. glad to know this integration path worked fro you!
3.
▲
by
soycaporal
3mo ago
very cool, I'll look into this. Thanks for sharing.
4.
▲
by
soycaporal
3mo ago
the base model I clocked it at 5 ms per embedded on my mac studio. There is a mini variant (the demo version) that is sub - 2 ms. It could be a SIMD issue.. I'll look into this for better runtime support (also fee free to file an issue
5.
▲
by
soycaporal
3mo ago
super cool use case! Hopefully it can provide quality embeddings + retrieval. Would love to learn to results/issues or feedback. Please feel free to file for issues on github
6.
▲
by
soycaporal
3mo ago
The corpus is mainly trained in english, unfortunately no other languages have been included in the distillation training. Yes it would work like fuse.js, but unlocks semantic search. Source code has the entire embedding distillation pipeli
7.
▲
by
soycaporal
3mo ago
awesome, noted, looking for capable teacher models to distill other architectures
8.
▲
by
soycaporal
3mo ago
that's great! let me know if there is anyway I can support, or any specific use case a roadmap could address!
9.
▲
by
soycaporal
3mo ago
It's entirely the QAT. The whole distillation process is quantization-aware from the start, so the ternary weights are learned rather than fitted after the fact. The only post-training quantization I applied was int4 on the embedding l
10.
▲
by
soycaporal
3mo ago
CPU cycle maxxing, who said GPUs were special?
11.
▲
by
soycaporal
3mo ago
thank you! hopefully it can unlock some novel applications, that would be cool
12.
▲
by
soycaporal
3mo ago
ohh thanks for the report.. probably has to do with wasm runtime.. Will note this as a known issue
13.
▲
by
soycaporal
3mo ago
love the idea! Will think of a way to host it probably on huggingface
14.
▲
by
soycaporal
3mo ago
I think standardizing the runtime is pretty effective, it then open up portability
15.
▲
by
soycaporal
3mo ago
gte-small outscores all-MiniLM-L6 on MTEB (~61 vs ~56 avg per the GTE paper). MiniLM is ternlight's teacher (ternlight holds 0.84 Spearman fidelity to teacher). I haven't run a head-to-head yet; STS-B/MTEB numbers are on the
16.
▲
by
soycaporal
3mo ago
yes, you could run a 1 time indexing run on the server side, and just ship the embeddings to frontend
17.
▲
Ternlight – 7 MB embedding model that runs in browser (WASM)
(ternlight-demo.vercel.app)
326 points
by
soycaporal
3mo ago
|
70 comments
18.
▲
by
soycaporal
3mo ago
Hobby project, I wanted to "ship a useful model in a web browser". so I distilled a small sentence encoder from MiniLM with ternary quantization-aware training. Also wrote the inference engine from scratch and shipped in Rust → WA