Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jakswa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
jakswa
8d ago
Bonsai 2 27B · Radeon RX 7900 XTX - 89 tokens/sec generation with speculative decoding - 81 tokens/sec at 20k context - 474 tokens/sec ingestion at 20k — about 42 seconds - 10.1 GiB peak VRAM with a 24k context wind
2.
▲
by
jakswa
8d ago
what kinda speeds do you see on 6700 XT? i'm always conflicted on investing time chasing speed-vs-quality tradeoffs. I've got a 7900 XT (about double the IO throughput). I'll probably end up giving it a go when I find time.
3.
▲
by
jakswa
8d ago
I love their web UI so much that I had AI slop all over it in a fit of fanboy-ism https://inkcap.click
4.
▲
by
jakswa
22d ago
dang only for certain nvidia GPUs, had my hopes up
5.
▲
by
jakswa
23d ago
don't see it in my AWS bedrock model list yet, but boy has bedrock mantle been annoying today with the errors/downtimes, with NO status page entries >_<
6.
▲
by
jakswa
29d ago
Oh. The UI screenshot on github is... not actually in the github repo? It's platform/hosted only? There's my first awkward discovery, but makes sense in retrospect.
7.
▲
by
jakswa
29d ago
I think I have to look at setting up one of these AI gateways for work, since AWS bedrock is such a PITA to hook a harness up to over IAM roles. Also I still can't believe bedrock hasn't released any open models in months (so ther
8.
▲
by
jakswa
1mo ago
ended up disabling ornith 9B. Oddly Ling 3 Tiny is pretty dang capable if its thinking is unleashed (tons of output tokens, maybe 5X the tokens but it's so fast it's maybe only twice as slow as a smarter model). This is a really i
9.
▲
by
jakswa
1mo ago
I had to go down to UD-Q3_K_XL for Qwen 3.8 27B to get it to fit in VRAM and be usable, but I worry I'm gutting its intelligence somewhat. I too am interested in faster + more-usable alternative that can exchange blows with the Q3-dumb
10.
▲
by
jakswa
1mo ago
I'll be comparing the 9B vs Ling 3 Tiny (8B-A1B) as a scout model. Ling tiny is so fast but can be a little too dumb. Hope the 9B strikes a good middleground even if dense/slower.
11.
▲
by
jakswa
1mo ago
I've been waiting on this model to show up on the Deep SWE benchmark results and treat its absence/delay as an indication of how slow and unusable it is for good results. I bet it thinks to the moon on some of those complex challe
12.
▲
by
jakswa
1mo ago
Thanks for mentioning Ling 3 Tiny. This model has completely bypassed me and seems promising for how small it is.
13.
▲
by
jakswa
1mo ago
I'm waiting for this comparison too. I was impressed by a 1-shot GLM 5.3 did for me the other day.
14.
▲
by
jakswa
1mo ago
Gemma 4 (both E4B + 12B) performed really well as ears+brains. I mostly comment because I too am always scouting for a nice local all-in-one model.
15.
▲
by
jakswa
1mo ago
I went back to Glimmer 30b for my 20GB of VRAM. Just a better experience fit-wise and speed-wise and tone-/voice-wise.
16.
▲
by
jakswa
1mo ago
thank you I'm still on ie8
17.
▲
by
jakswa
1mo ago
I'm in the exact same boat with a 7900 XT and a good Glimmer 30B experience. I was really hoping qwen 3.8 would bring some memory/space efficiency savings along the lines of whatever is going on with Glimmer 30B. I have been surpr
18.
▲
by
jakswa
1mo ago
anecdotes: 35B-A3B does want more memory, bigger model. But if you get it running it will be faster and more enjoyable to use -- text will fly by -- due to only 3B params being active, in my experience at least.
19.
▲
by
jakswa
1mo ago
I dunno, I didn't read in-depth. Hopefully you don't gotta zoom in with human eyeballs.
20.
▲
by
jakswa
1mo ago
OMP changed the default compaction to images ! Kinda nuts to read about. Saves the generation cost of the traditional compaction step and writes the context as tiny text to an image, if I was following correctly.
21.
▲
by
jakswa
1mo ago
my whole world is shifting. have I been seeing _different pelicans_ from everyone else?!
22.
▲
by
jakswa
1mo ago
This pelican gave me a good laugh, because there's enough reasoning that the render is out of sight initially. The buildup!
23.
▲
by
jakswa
2mo ago
I'll back up your smaller claim, but be specific that it's UD-Q4_K_XL size: - muse glimmer: 15.9GB - qwen 3.6 27B: 17.6GB My video card is so close to its limit that these GB thresholds are mattering too much for me :D
24.
▲
by
jakswa
2mo ago
it's here! https://news.ycombinator.com/item?id=49245575
25.
▲
by
jakswa
2mo ago
I'm listening to pelican sounds on youtube while I wait for Simon.
26.
▲
by
jakswa
2mo ago
I like the tabletop RPG use case, and wanted to say: If your hardware likes it you should check out Gemma 4 for creative DMing use case. I found it to be much better at holding the plotlines and being creative on gaming turns. My experiment
27.
▲
by
jakswa
2mo ago
some support already merged, and I verified in a local build that it runs (cannot get MTP params working tho, about ~40 tok/s on my beefy 800GB/s 7900XT w/ 20GB VRAM). https://github.com/ggml-org/llama.cp
28.
▲
by
jakswa
2mo ago
Q3 results: unsloth/Muse-Glimmer-30B-GGUF:UD-Q3_K_XL gets down to 15.6GB VRAM and full context (131k) on the 4 parallel slots. Prompt/generation speeds about the same. Overall feeling like a nicer-fitting Qwen 3.6 27B, but want to
29.
▲
by
jakswa
2mo ago
Another candidate for the 7900XT (20GB VRAM) I got sitting around. I pulled latest llama.cpp (targeting vulkan during build) after seeing a muse PR merged a few hours ago, and unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL runs on my 7900XT
30.
▲
by
jakswa
2mo ago
> Every source file is summarized once into a short description of what it does. I admit I haven't had a chance to read the whole README, but wanted to get down my hesitation after I got pretty far (as an interested user): I think i
More ›