Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ljosifov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
ljosifov
2mo ago
Similar. Noticed DuckDB ever since an old article 'what db should I use' for local small data warehousing. The author was blown away that DuckDB seemed super naturally quick. It was I think columnar store + compression facilitated
2.
▲
by
ljosifov
2mo ago
Ah sorry - I misunderstood. Thanks for explaining it. Have not heard of Paseo nor Kepler, and have never tried Zed. Yeah I too assumed if I'm to try use OpenAI subscription outside Codex, or Anthropic subscription outside Claude Code -
3.
▲
by
ljosifov
2mo ago
I've used GLM-s the longest with Claude Code and their Anthropic supplied endpoint. As per their docs $ ANTHROPIC_BASE_URL=" https://api.z.ai/api/anthropic " ANTHROPIC_AUTH_TOKEN="zai-api-key" cl
4.
▲
by
ljosifov
2mo ago
Wdym "sadly they don’t support using Claude Code"? For the longest time that's all Zai supported - Claude code. I'd run it via export ZAI_ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic&qu
5.
▲
by
ljosifov
2mo ago
Atm OMP (oh-my-pi) with opencode-go subscription using deepseek-v4-flash-0731 (and mimo-v2.5 as advisor). Used pi before (still good, but too minimalist for all general use; works on Android under Termux btw) and OpenCode (which is ok too).
6.
▲
by
ljosifov
2mo ago
Hipfire works on my 7900xtx best of all. Biggest surprise - Qwen3.6-27B dense does not grind to a halt with context depths all the way up to 250K! Measured at 1K, 8K, 32K, 64K, 128K, 192K, 250K - Hipfire speed holds close to 40 tok/s.
7.
▲
by
ljosifov
2mo ago
omp - current top, after using codex claude opencode pi that I still use too
8.
▲
by
ljosifov
2mo ago
Hear hear. IQ tokens to cheap to meter upon us. So many things changed since last week. Now I've had Prime agent session grinding into its 20-th hour still not giving up. Been using opencode-go since Go sub appeared. What made a differ
9.
▲
by
ljosifov
2mo ago
Love pi. It's good to have a minimal agent, always welcome. Even if only to bootstrap install other agents. E.g. recently Hermes stopped installing under Termux (v0.19 - no, v0.18 - yes). Agent pi to the rescue: installed it under Term
10.
▲
by
ljosifov
2mo ago
Same. And I have come to use OMP (oh my pi) agent /advisor mode to put a 2nd model on the case (also mid-size one), reading everything. It can not block anything or change anything - just inserts comments in the text stream with 1 turn
11.
▲
by
ljosifov
3mo ago
I usually doubt the 'small dataset tuned' variants. B/c ages ago (in the NN prehistory) I've done some NN training, and appreciate how hard it is to improve in general, and how easy it is to ruin a model in general while
12.
▲
by
ljosifov
3mo ago
Hobbled - but not to death, the few times I use it (usually on a plane). I tried 2bit of a 20% REAP reduced experts. :-O That's the biggest that fits on my own h/w (3yrs old M2 Max 96gb). It's coherent, it does work, doesn&#x
13.
▲
by
ljosifov
3mo ago
True - they are workhorses. Not super bright, but good enough for lots of everyday tasks. I've found sweet spot to be turning thinking off, as it adds small or no value, while increasing the token count and waiting time. Last 27B I use
14.
▲
by
ljosifov
3mo ago
Running 27B dense model on M5 128GB is ok, but one can do better. On M5 128GB one can make use of the ram and use sparse MoE. For example, DeepSeek-V4-Flash will fit, served by DwarfStar ( https://github.com/antirez/ds4
15.
▲
by
ljosifov
3mo ago
Haha :-) - FoxPro and Clipper next.
16.
▲
We built the fastest API for GLM-5.2
(twitter.com)
3 points
by
ljosifov
3mo ago
|
0 comments
17.
▲
by
ljosifov
4mo ago
Not replaced but supplemented. For off-line coding current setup is pi + ds4-server + DeepSeek-V4-Flash REAP25 (on M2 Max 96gb). For simpler programming related (e.g. text2sql) as well as synthetic data generation, current best for me is ll
18.
▲
by
ljosifov
4mo ago
For high Ram (unified), and relatively middling to lowish Tflops and bandwidth GB/s, usually MoEs are most hopeful. The current top-1 in the (iq, tok/s, @ context depth) ranks for me (M2 Max, 96gb) is DeepSeek-V4-Flash REAP25 <
19.
▲
by
ljosifov
4mo ago
Yes, it's performant, and esp performant at non-trivial context depths. DeepSeek-V4 DS4 (and Flash - DS4F) drop tok/s speed much less than the rest. On my M2 Max it took context depths of 768K to drop tok/s to ~10 tok/s.
20.
▲
by
ljosifov
4mo ago
Thanks for the tear down. IDK anything about quantum (my knowledge there starts and ends with https://www.scottaaronson.com/democritus/lec9.html ), but amused enough to follow in the background. See whether it ends craz
21.
▲
Quantum Information as Everything
(vlatkovedral.substack.com)
2 points
by
ljosifov
4mo ago
|
3 comments
22.
▲
by
ljosifov
4mo ago
+1 for boring. Boring code is Solid Code, in the sense of "Writing Solid Code" - the old book by Steve Maguire.
23.
▲
by
ljosifov
5mo ago
Thanks for the DS4, will give it a try. Was hoping maybe I can re-quantise shave few GB... MiniMax-M2.7 Unsloth's UD-IQ2_XXS is down to 65GB - it run albeit too slow to be usable to an agent at context depth. I'm curious DS4F with
24.
▲
by
ljosifov
5mo ago
On 96gb I can give up to about 88GB to the GPU with sysctl iogpu.wired_limit_mb=88000, without suffering any ill-effects. When pushed higher I tend to notice e.g. graphic driver errors, youtube web page not working, other semi-random glitch
25.
▲
by
ljosifov
5mo ago
Love this, even if can't use it atm (not got the h/w - only 96gb on M2 Max). I get it the general comp/public will find it unusable or worse. Reminds me of how home computers were - mere toys - before they became personal com
26.
▲
by
ljosifov
5mo ago
What we see and experience - it's all natural, it's the natural order. :-) When people claim something is un-natural, usually it's natural in that occurs in nature, only - they themselves find it objectionable. It's some
27.
▲
by
ljosifov
5mo ago
In the same boat with 7900xtx. 24GB vram, on paper decent performance, in reality most things don't run. Only llama.cpp is consistent that it can run most models, even if maybe not at top performance (afaik - lacking MTP, problems cach
28.
▲
by
ljosifov
5mo ago
~/llama.cpp$ build-.../bin/llama-batched-bench -m models/....gguf -npp 512,1024,2048,4096,8192,16384,32768 -ntg 128 -npl 1 -c 36000 On amd 7900xtx Qwen3.6-27B-Q4_K_M | PP | TG | B | N_KV | T_PP s |
29.
▲
by
ljosifov
7mo ago
Glad to see other people using it. Saved my life, was going crazy click-clicking to nab the right window. Now Cmd-1..9 brings to focus a window of my chosen application. (Chrome) In case it helps someone else, myself and Codex iterating ove
30.
▲
by
ljosifov
7mo ago
Say more please if you can. How/why is ik_llama.cpp faster then mainline, for the 27B dense? I'd like to be able to run 27B dense faster on a 24GB vram gpu, and also on an M2 max.
More ›