Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anon373839
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
anon373839
7d ago
The model weights are only ~5GB, so this is small.
2.
▲
by
anon373839
7d ago
The world has moved on.
3.
▲
by
anon373839
8d ago
This is a serving bug or quantization issue. I had all kinds of issues that were like this on DGX Spark until I found a single-GB10 vLLM recipe [1] that uses Nvidia's NVFP4 quant. The community quants did not work well. Another failure
4.
▲
GVS5H: Five Qwen3.8 Models Match Claude Fable 5 on LiveCodeBench Hard
(github.com)
2 points
by
anon373839
9d ago
|
0 comments
5.
▲
by
anon373839
9d ago
> Personally, I’ve replaced OpenCode with a thin wrapper around Pydantic-AI as the pythonic analogue to Pi-Agent for headless use via Hermes That's really interesting. I like Pydantic AI a lot and wondered why all of the harnesses s
6.
▲
by
anon373839
9d ago
Model welfare, much like AI xrisk, is a concept born from evidence-free “what if?” questions. Some people ran with these what-ifs and developed ornate belief systems around them. And now they demand the rest of us take them seriously.
7.
▲
by
anon373839
12d ago
That isn’t any kind of antitrust violation. Agreeing not to undercut each other’s prices would be, however.
8.
▲
by
anon373839
12d ago
Ah, no, that’s not cheaper. Renting GPUs adds up quickly and leaves you with nothing in the end. Renting tokens from open model providers is cheaper but it incurs the same issues: unexpected changes in model quality, inconsistent speeds, se
9.
▲
by
anon373839
13d ago
I will say that LLMs are somewhat unlike guns in that they emit text.
10.
▲
by
anon373839
13d ago
Privacy is a great reason, but independence is another. It’s very nice knowing that you’re going to get the same reliable product every time you call the model. Nothing is going to change unless you decide to change it.
11.
▲
by
anon373839
13d ago
It’s absolutely laughable that he refers to METR as if they were neutral observers. They are ex-Anthropic employees and others with direct financial interests in Anthropic.
12.
▲
by
anon373839
13d ago
It is costly, especially right now. I don’t think you can make a case for it on cost savings! The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rate
13.
▲
by
anon373839
13d ago
> If someone put in frontier AI models from like .... last june I guess? in a box and let me run it with "decent" token throughput I would be happy. You can have that! Qwen 3.8 Flash-Next is ~Opus 4.6 and runs nicely on a DGX S
14.
▲
by
anon373839
13d ago
No, it’s not the active parameters. Qwen 3.8 Flash has 6B active and it smokes both models.
15.
▲
ChatGPT: "Super shady" data controls settings
(twitter.com)
2 points
by
anon373839
17d ago
|
1 comments
16.
▲
by
anon373839
17d ago
https://xcancel.com/edoardocontente/status/20976073405059281...
17.
▲
by
anon373839
17d ago
I don’t use ChatGPT anymore, but when I did, the training opt-out setting was frequently silently resetting itself to off.
18.
▲
by
anon373839
18d ago
> GLM 5.3-flash released last week, and that means Project Glasswing and Daybreak are running out of time. Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious action
19.
▲
by
anon373839
21d ago
Claude is different from S3. AWS doesn’t need to rifle through your files to stay ahead of the competition or to mine them for business ideas because the core business is overvalued and rapidly commoditizing. AI labs, on the other hand, hav
20.
▲
by
anon373839
21d ago
Hallucinations are very damaging to a model’s utility. But doesn’t the Omniscience Index focus on knowledge-based queries? To me, using LLMs for their memorized knowledge is very 2023 and suboptimal. IMO, what really makes a model useful is
21.
▲
by
anon373839
21d ago
Jensen Huang?
22.
▲
by
anon373839
23d ago
Ah, I see Anthropic is back at that “saving humanity” game…
23.
▲
by
anon373839
23d ago
Sebastian Raschka posted about this architecture: > A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth o
24.
▲
by
anon373839
25d ago
Apt.
25.
▲
by
anon373839
25d ago
Macs have excellent generation speed, and the new Ultra will positively smash that at 1.2TB/sec of bandwidth. For example, that new 176B parameter Qwen model would generate tokens at ~200 tokens/sec. Macs don’t have very good pref
26.
▲
by
anon373839
26d ago
> The idea of releasing a frontier model without RL is frightening In case you were not aware, strong base models (no post training at all) have been available for quite some time now. Including ones that eclipse “scary” frontier models
27.
▲
by
anon373839
28d ago
You and I have the same machine. Do you mind sharing the model ID you're using? I'm on oMLX also but I haven't seen anything above ~20 tok/s out of 3.8 27B, even with MTP and generating code.
28.
▲
by
anon373839
29d ago
Yes; my point was not related to any of that.
29.
▲
by
anon373839
29d ago
Route A is good for rapid, disposable prototyping. If you ever have a stray thought, “I wonder how this would work if the whole paradigm were turned sideways”, you now have a chance to preview a “working” version of your idea. If you like i
30.
▲
by
anon373839
29d ago
> The bitter lesson is about hand tuned AI vs computational general methods. However in truth today’s AI uses both. We have general compute heavy models which require narrow expert instructions The models are not even really trained bitt
More ›