Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Philpax
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
GPT 5.6 Discounts and Jevons Paradox
(openrouter.ai)
1 points
by
Philpax
1mo ago
|
0 comments
32.
▲
zai-Org/GLM-5.3
(huggingface.co)
6 points
by
Philpax
1mo ago
|
0 comments
33.
▲
by
Philpax
1mo ago
leaving the README like this is a good bit, though
34.
▲
by
Philpax
1mo ago
Performance is not the same as original creation, so yes, I think that's very possible.
35.
▲
by
Philpax
2mo ago
No.
36.
▲
Brief independent investigation of OpenAI / Hugging Face hacking incident
(metr.org)
6 points
by
Philpax
2mo ago
|
1 comments
37.
▲
GLM-5.3-Flash
(z.ai)
1132 points
by
Philpax
2mo ago
|
580 comments
38.
▲
zai-org/GLM-5.3-Flash
(huggingface.co)
4 points
by
Philpax
2mo ago
|
0 comments
39.
▲
by
Philpax
2mo ago
"political histrionics" is an interesting way to refer to calling for ethnic cleansing
40.
▲
Qwen3.8-Flash-Next Technical Report [pdf]
(github.com)
35 points
by
Philpax
2mo ago
|
1 comments
41.
▲
by
Philpax
2mo ago
https://jakelazaroff.com/words/dhh-is-way-worse-than-i-thoug...
42.
▲
Unsloth/Qwen3.8-Flash-Next-GGUF
(huggingface.co)
9 points
by
Philpax
2mo ago
|
0 comments
43.
▲
Qwen/Qwen3.8-Flash-Next
(huggingface.co)
7 points
by
Philpax
2mo ago
|
0 comments
44.
▲
by
Philpax
2mo ago
Stealth launch: builds hype, allows them to collect user preference data and see where the model fails. Why people care: it's free, decent, and people love a good mystery.
45.
▲
Uncovering a universal offline sandbox escape
(primeintellect.ai)
2 points
by
Philpax
2mo ago
|
0 comments
46.
▲
by
Philpax
2mo ago
Strongly recommend https://github.com/Neroued/ninfer , which can pull ~180 TPS on 5090 with 3.8, and 500 (!) with 3.6 35B-A3B.
47.
▲
Jalapeño's results show industry-leading speed and efficiency in AI inference
(openai.com)
21 points
by
Philpax
2mo ago
|
1 comments
48.
▲
Pipette: A benchmarking suite for on-device intelligence
(liquid.ai)
3 points
by
Philpax
2mo ago
|
0 comments
49.
▲
by
Philpax
2mo ago
Using a RTX 5090 ($3~4k), I can run Qwen 3.8 27B at ~180 TPS with ninfer [0]. With its thinking maxed out, I can confirm that the quality of output is roughly on par with Opus 4.5~4.6 - that is, this 20GB file really can write software by i
50.
▲
by
Philpax
2mo ago
DFlash 2 is a diffusion-based speculative decoding head for autoregressive models; there's nothing to accelerate here, because this is already wholly diffusion.
51.
▲
200B Tokens Later: A Month of Letting AI Agents Decompile MW2
(momo5502.com)
18 points
by
Philpax
2mo ago
|
3 comments
52.
▲
A letter to the Discord community in Brazil
(discord.com)
8 points
by
Philpax
2mo ago
|
0 comments
53.
▲
by
Philpax
2mo ago
> Are they literally adding a hidden system prompt that says "effort level: $level" ? Yes. https://magazine.sebastianraschka.com/p/controlling-reasonin...
54.
▲
Xiaohongshu dots3-note Preview
(huggingface.co)
2 points
by
Philpax
2mo ago
|
0 comments
55.
▲
by
Philpax
2mo ago
I think it's pretty obvious that, in that world, the AIs will simply be tasked with making the compilers faster. It's already happening with their own stack, after all.
56.
▲
MiniMax-Music3
(huggingface.co)
7 points
by
Philpax
2mo ago
|
0 comments
57.
▲
Hugging Face: DeepSeek-V4-Pro-0813
(huggingface.co)
7 points
by
Philpax
2mo ago
|
1 comments
58.
▲
by
Philpax
2mo ago
I think you might be underestimating how many people genuinely despise Musk.
59.
▲
by
Philpax
2mo ago
I was under the impression that you could fit the full 1M context within the 192GB VRAM as a result of DeepSeek's various architectural advancements, but I'll grant that DSpark + a larger pool for concurrency may necessitate more
60.
▲
by
Philpax
2mo ago
What do you need the extra 2 for? Tensor parallelism?
More ›