Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
throwdbaaway
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
throwdbaaway
5mo ago
If I understand correctly, both the staging database and the production database share the same volume. Thus, production data was gone as well after deleting the volume. 1st hint - the API call only contains one volume: curl -X POST h
32.
▲
by
throwdbaaway
6mo ago
Should be about 10~20 GiB per session. Save/restore is exactly what DeepSeek does using its 3FS distributed filesystem: https://github.com/deepseek-ai/3fs#3-kvcache With this much cheaper setup backed by disks, th
33.
▲
by
throwdbaaway
6mo ago
Based on the release schedule of 3.5 previously, my optimistic take is that they distill the small models from the 397B, and it is much faster to distill a sparse A3B model. Hopefully the other variants will be released in the coming days.
34.
▲
by
throwdbaaway
6mo ago
His Vibe Coding book is invaluable as a textbook example of slop.
35.
▲
by
throwdbaaway
6mo ago
https://github.com/anthropics/claude-code/issues/46829#issue... - Have you checked with your colleague? (and his AI, of course)
36.
▲
by
throwdbaaway
6mo ago
> EC2 instances on shared hardware showed up to 30% variance between runs due to noisy neighbors. Based on this finding, I suppose the better way is to rely on local hardware whenever possible?
37.
▲
by
throwdbaaway
6mo ago
Very nice TG improvement from Flash Attention KQ fusion. Is it something that was already done in ik_llama.cpp? If not, then it will be a welcomed addition for hybrid CPU/GPU inference.
38.
▲
by
throwdbaaway
6mo ago
https://github.com/THUDM/IndexCache - Might be some expected issue when rolling out this. They don't have enough compute, and have to innovate.
39.
▲
by
throwdbaaway
7mo ago
90% of what you pay in agentic coding is for cached reads, which are free with local inference serving one user. This is well known in r/LocalLLaMA for ages, and an article about this also hit HN front page few weeks ago.
40.
▲
by
throwdbaaway
7mo ago
What about the VRAM requirement for KV cache? That may matter more than memory bandwidth. With these GPUs, there are more compute capacity than memory bandwidth than VRAM. DeepSeek got MLA, and then DSA. Qwen got gated delta-net. These inve
41.
▲
by
throwdbaaway
7mo ago
Yours is the only benchmark that puts 35B A3B above 27B. Time for human judgement to verify? For example, if you look at the thinking traces, there might be logical inconsistencies in the prompts, which then tripped up the 27B more when rea
42.
▲
by
throwdbaaway
7mo ago
Using ik_llama.cpp to run a 27B 4bpw quant on a RTX 3090, I get 1312 tok/s PP and 40.7 tok/s TG at zero context, dropping to 1009 tok/s PP and 36.2 tok/s TG at 40960 context. 35B A3B is faster but didn't do too well
43.
▲
by
throwdbaaway
7mo ago
There are Qwen3.5 27B quants in the range of 4 bits per weight, which fits into 16G of VRAM. The quality is comparable to Sonnet 4.0 from summer 2025. Inference speed is very good with ik_llama.cpp, and still decent with mainline llama.cpp.
44.
▲
by
throwdbaaway
7mo ago
I don't quite get the low temperature coupled with the high penalty. We get thinking loop due to low temperature, and we then counter it with high penalty. That seems backward. For Qwen3.5 27B, I got good result with --temp 1.0 --top-p
45.
▲
by
throwdbaaway
7mo ago
We are all reasonable people here, and while you are (mostly) correct, I think we can all agree that Anthropic documentation sucks. If I have to infer from the doc: * Haiku 4.5 by default doesn't think, i.e. it has a default thinking
46.
▲
by
throwdbaaway
7mo ago
For 27B, just get a used 3090 and hop on to r/LocalLLaMA. You can run a 4bpw quant at full context with Q8 KV cache.
47.
▲
by
throwdbaaway
7mo ago
I would say 27B matches with Sonnet 4.0, while 397B A17B matches with Opus 4.1. They are indeed nowhere near Sonnet 4.5, but getting 262144 context length at good speed with modest hardware is huge for local inference. Will check your updat
48.
▲
by
throwdbaaway
7mo ago
Can you describe a bit more how this works? I suppose the speed remains about the same, while the experience is more pleasant? (Big fan of SQLAlchemy)
49.
▲
by
throwdbaaway
8mo ago
From a quick testing on simple tasks, adaptive thinking with sonnet 4.6 uses about 50% more reasoning tokens than opus 4.6. Let's see how long it will take for DeepSeek to crack this.
50.
▲
by
throwdbaaway
8mo ago
If you ask someone knowledgeable at r/LocalLLaMA about an inference configuration that can increase TG by *up to* 2.5x, in particularly for a sample prompt that reads "*Refactor* this module to use dependency injection", then
51.
▲
by
throwdbaaway
8mo ago
As mentioned by the sibling comment from godelski, it is about the lack of precision, not the lack of determinism. After all, we already got https://thinkingmachines.ai/blog/defeating-nondeterminism-in... , which is not
52.
▲
by
throwdbaaway
8mo ago
Not using Hot Aisle for inference?
53.
▲
by
throwdbaaway
8mo ago
I thought "iterate and improve" was exactly what Phil did.
54.
▲
by
throwdbaaway
8mo ago
I call this the Groundhog Day loop
55.
▲
by
throwdbaaway
8mo ago
> Many times I will erase what the LLM has written and redo it, by myself depending on the situation. The contention here is that antirez doesn't think this is necessary anymore. 100% code gen, with the occassional "stepping in
56.
▲
by
throwdbaaway
10mo ago
At the high level, you asked LLM to translate N lines of code to maybe 2N lines of code, while GP asked LLM to translate N lines of English to possibly 10N lines of code. Very different scenarios.
57.
▲
by
throwdbaaway
10mo ago
After the latest production issue, I have a feeling that opus-4.5 and gpt-5.1-codex-max are perhaps better than me at debugging. Indeed my role was relegated to combing through the logs, finding the abnormal / suspicious ones, and feed
58.
▲
by
throwdbaaway
10mo ago
That DSML in the encoding directory looks quite a bit different from the Harmony chat template.
59.
▲
by
throwdbaaway
11mo ago
It is the reasoning. During the reasoning process, the top few tokens have very similar or even same logprobs. With gpt-oss-120b, you should be able to get deterministic output by turning off reasoning, e.g. by appending: {"role&
60.
▲
by
throwdbaaway
11mo ago
Somehow that article totally ignored the insane pricing of cached input tokens set by Anthropic and OpenAI. For agentic coding, typically 90~95% of the inference cost is attributed to cached input tokens, and a scrappy China company can do
More ›