Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tarruda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
tarruda
5mo ago
> I started looking at the commits, and it's basically solving the ,,tests not pass'' problem by changing the tests themselves Not sure if these decisions were made by the LLM, but I've always felt that Claude is more
62.
▲
by
tarruda
5mo ago
Mailbox.org (also from Germany) seems to be experiencing issues too.
63.
▲
by
tarruda
5mo ago
They also published draft models for E4B and E2B. For those, the draft models are only 78m parameters: https://huggingface.co/google/gemma-4-E4B-it-assistant
64.
▲
by
tarruda
5mo ago
There is a newer PR which will probably be merged soon: https://github.com/ggml-org/llama.cpp/pull/22673
65.
▲
Show HN: An in-browser, Unix emulator powered by libghostty-vt
(tarruda.github.io)
2 points
by
tarruda
5mo ago
|
0 comments
66.
▲
by
tarruda
6mo ago
Codex is the best out-of-box experience, especially due to its builtin sandboxing. Only drawback is that its edit tool requires the LLM to output a diff which only GPTs are trained to do correctly.
67.
▲
by
tarruda
6mo ago
I currently do something similar. My router is a 16GB n150 mini PC with dual NICs. The actual router OS is within openwrt VM managed by Incus (VM/Container hypervisor) that has both NICs passed through. One of the NICs is connected to
68.
▲
by
tarruda
6mo ago
Benchmarks don't tell the whole story. For one-shot coding tasks, I found Step 3.5 Flash to be stronger even than Qwen 3.5 397B.
69.
▲
by
tarruda
6mo ago
Since that discussion, they released the base model and a midtrain checkpoint: - https://huggingface.co/stepfun-ai/Step-3.5-Flash-Base - https://huggingface.co/stepfun-ai/Step-3.5-Flash-Base-Midtra
70.
▲
by
tarruda
7mo ago
> Have you compared against MLX? I don't think MLX supports similar 2-bit quants, so I never tried 397B with MLX. However I did try 4-bit MLX with other Qwen 3.5 models and yes it is significantly faster. I still prefer llama.cpp du
71.
▲
by
tarruda
7mo ago
> in neovim and I feel so grateful that this tool is still receiving love and getting new updates. @justinmk deserves the credit for this!
72.
▲
by
tarruda
7mo ago
In my case it the 2.46BPW has been working flawless for tool calling, so I don't think 2-bit was the culprit for JSON failing. They did reduce the number of experts, so maybe that was it?
73.
▲
by
tarruda
7mo ago
Yes. Note that the only reason I acquired this device was to run LLMs, so I can dedicate its whole RAM to it. Probably not viable for a 128G device where you are actively using for other things.
74.
▲
by
tarruda
7mo ago
> What's the tok/s you get these days? I ran llama-bench a couple of weeks ago when there was a big speed improvement on llama.cpp ( https://github.com/ggml-org/llama.cpp/pull/20361#issuecommen...
75.
▲
by
tarruda
7mo ago
I don't think I've ever seen the M1 ultra GPU exceed 80w in asitop. Update: I just did a quick asitop test while inferencing and the GPU power was averaging at 53.55
76.
▲
by
tarruda
7mo ago
I can't say anything about the OP method, but I already tested the smol-IQ2_XS quant (which has 2.46 BPW) with the pi harness. I did not do a very long session because token generation and prompt processing gets very slow, but I think
77.
▲
by
tarruda
7mo ago
Note that this is not the only way to run Qwen 3.5 397B on consumer devices, there are excellent ~2.5 BPW quants available that make it viable for 128G devices. I've had great success (~20 t/s) running it on a M1 Ultra with room f
78.
▲
by
tarruda
7mo ago
Thanks for the info!
79.
▲
by
tarruda
7mo ago
I would rather have a phone that doesn't let my carrier show random messages whenever they feel like it.
80.
▲
by
tarruda
7mo ago
Just a message popup, a window with dark background and some text ad on it. I did not buy this phone from a carrier, just added the SIM card later. Really surprised to learn this doesn't happen to others. Always assumed that the SIM ca
81.
▲
by
tarruda
7mo ago
Just checked, and only "Phone" and "Google" have this permission. There are no preinstalled apps, I bought this phone clean on Germany and then added a Brazil's SIM card when I got back. Could it be that the SIM car
82.
▲
by
tarruda
7mo ago
I have a pixel 8a with a TIM SIM card and every once in a while I see an ad popup on my phone.
83.
▲
by
tarruda
7mo ago
One thing that annoys me is the ability that my mobile carrier has to just throw ad popups. Is that something that GrapheneOS fixes?
84.
▲
by
tarruda
8mo ago
> so llama.cpp just doesn't handle it correctly. It is a bug in the model weights and reproducible in their official chat UI. More details here: https://github.com/ggml-org/llama.cpp/pull/19283#issueco
85.
▲
by
tarruda
8mo ago
I did play with Qwen3 Coder Next a bit, but didn't try it in a coding harness. Will give it a shot later.
86.
▲
by
tarruda
8mo ago
No, it is not cheaper. An M3 ultra with 512GB costs $10k which would give you 50 months of Claude or Codex pro plans. However, if you check the prices on Chinese models (which are the only ones you would be able to run on a Mac), they are m
87.
▲
by
tarruda
8mo ago
There's an AMA happening on reddit and they said it will be fixed in the next release: https://www.reddit.com/r/LocalLLaMA/comments/1r8snay/ama_wit...
88.
▲
by
tarruda
8mo ago
Both gpt-oss are great models for coding in a single turn, but I feel that they forget context too easily. For example, when I tried gpt-oss 120b with codex, it very easily forgets something present in the system prompt: "use `rg` comm
89.
▲
by
tarruda
8mo ago
Haven't tried. I'm too used to llama.cpp at this point to switch to something else. I like being able to just run a model and automatically get: - OpenAI completions endpoint - Anthropic messages endpoint - OpenAI responses endpoi
90.
▲
by
tarruda
8mo ago
> This is the first model that has really broken into the anglosphere. Before Step 3.5 Flash, I've been hearing a lot about ACEStep as being the only open weights competitor to Suno.
More ›