Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
electroglyph
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
electroglyph
4mo ago
absolutely. somebody online was wanting an LLM with Georgian language support, and that's exactly what i suggested: start digitizing Georgian text.
32.
▲
by
electroglyph
5mo ago
i'll upvote this each time it's submitted
33.
▲
by
electroglyph
5mo ago
you should be using dflash with that model, look it up
34.
▲
by
electroglyph
5mo ago
heretic maintainer: https://github.com/p-e-w/heretic the fun bits are in another branch or PRs
35.
▲
by
electroglyph
5mo ago
p-e-w was just talking about this the other day in his Discord. seems doing the one neuron method is quite bad for KLD and that's why the newer techniques have stuck.
36.
▲
by
electroglyph
5mo ago
site has so little information there doesn't seem to be much to discuss
37.
▲
by
electroglyph
5mo ago
this looks awesome. i've been struggling with vector compression, and have been trying PCA + all sorts of rotations. looking forward to trying this out
38.
▲
by
electroglyph
5mo ago
nice writeup! looking forward to doing some more training as soon as i get some more data sorted. it'll be a custom arch, but i'll probably shoehorn it into unsloth for a speed boost.
39.
▲
by
electroglyph
5mo ago
you can train it, but not fully
40.
▲
by
electroglyph
5mo ago
that's in the ideal scenario where it's only seen a single copy of it tho
41.
▲
by
electroglyph
5mo ago
it was 1.3e-6 billion years ago!
42.
▲
by
electroglyph
5mo ago
i'm doing inference on a free mi300x instance from AMD right now. not sure if the software stack is just old or what, but here's what i've observed: stuck on an old version of vllm pre-Transformers 5 support. it lacks MoE sup
43.
▲
by
electroglyph
5mo ago
but should you drive or walk to the car wash?
44.
▲
by
electroglyph
5mo ago
i dunno, Opus is losing it's edge imo. i regularly use a mix of models, including Opus, glm 5.1, kimi 2.6, etc. and i find that all of them are pretty much equally good at "average" coding, but on difficult stuff they're
45.
▲
by
electroglyph
5mo ago
> they also don't know what they don't know they sort of do tho: https://transformer-circuits.pub/2025/introspection/index.ht...
46.
▲
by
electroglyph
5mo ago
how about a unicode art tool? https://electroglyph.github.io/atheriz_draw/
47.
▲
by
electroglyph
6mo ago
https://sleepingrobots.com/dreams/stop-using-ollama/
48.
▲
by
electroglyph
6mo ago
flow matching is making some strides right now, too
49.
▲
by
electroglyph
6mo ago
i don't buy this. distilled how? you don't get access to logprobs, and the thinking traces are fake and compressed. it's an expensive way to get potentially substandard training data.
50.
▲
by
electroglyph
6mo ago
nah, a crypto grifter released one with cooked benchmarks
51.
▲
by
electroglyph
6mo ago
better than Opus? not even close. after struggling thru server overload for the past couple hours i finally put 5.1 thru the paces and it's....okay. failed some simple stuff that Sonnet/Opus/Gemini didn't. failed it badl
52.
▲
by
electroglyph
6mo ago
after you go from from millions of params to billions+ models start to get weird (depending on training) just look at any number of interpretability research papers. Anthropic has some good ones.
53.
▲
by
electroglyph
6mo ago
what's with the weird "Geometric Lens routing" ?? sounds like a made up GPTism
54.
▲
by
electroglyph
7mo ago
ah, you've found the danger zone!
55.
▲
by
electroglyph
7mo ago
it's also a very common "favorite number" for them
56.
▲
by
electroglyph
7mo ago
working on a text game engine similar to Evennia: https://github.com/electroglyph/atheriz
57.
▲
by
electroglyph
7mo ago
much respect to the PyPy contributors, but it seems like a pretty fair assessment
58.
▲
by
electroglyph
7mo ago
the default heretic with only 100 samples isn't very good, you really need your own, larger dataset to do a proper abliteration. the best abliteration roughly matches a very careful decensor SFT
59.
▲
by
electroglyph
7mo ago
my wishlist for pyrefly: when using decorated functions, show the underlying type hints instead of the decorators
60.
▲
by
electroglyph
7mo ago
Cheers Daniel and Mike and team, keep up the good work!
More ›