Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sosodev
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
sosodev
3mo ago
Realistically you can't prevent distillation. OpenAI / Anthropic are slowly moving towards hiding the steps in-between input and output (hidden thinking), but that only helps so much. Imagine you put a file into Claude and say &qu
32.
▲
by
sosodev
3mo ago
Model distillation can't be stealing at all if you rationally apply copyright law to it. Anthropic is not deprived of Fable so there is no theft. At best it would be infringement, but even that might not hold up in the courts given the
33.
▲
by
sosodev
3mo ago
Distillation is a very vague term. It can mean anything from training exclusively on a model's output to using it for a very small portion of the training. In this case it is almost certainly towards the very small portion side of the
34.
▲
by
sosodev
3mo ago
A month seems plenty long enough. They're not rebuilding the entire model from scratch. It's just getting Fable to act as a teacher model for some of the final reinforcement learning on the base that Kimi already had.
35.
▲
by
sosodev
3mo ago
When? It literally says on the page for Gemini 3.6 Flash "Artificial Analysis Coding Index represents the weighted average of coding benchmarks in the Artificial Analysis Intelligence Index (Terminal-Bench v2.1, SciCode)"
36.
▲
by
sosodev
3mo ago
What inference server are you using? They have a custom branch for llama.cpp, but I wouldn't be surprised at all if it still needs fixing.
37.
▲
by
sosodev
3mo ago
Because AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities.
38.
▲
by
sosodev
3mo ago
That’s the worst thing about sycophancy. You can never tell when it’s warranted or not. I imagine that many AI chats have had real gold in them and yet we’ll never know. The same was obviously true historically. People have had many revolut
39.
▲
by
sosodev
3mo ago
Of course efficiency matters, but a lot of people either have cheap electricity or efficient hardware. My AMD strix halo home server can serve Gemma4-26B at like 70 TPS (rough estimate, I don’t remember the exact speed buts its fast af) whi
40.
▲
by
sosodev
3mo ago
They’re exaggerating or have a very simple way of using these models. The Gemma 4 series, even at 31B, is nowhere near the frontier. They’re great models, but you will notice a huge difference for complex tasks. The best local agentic codin
41.
▲
by
sosodev
3mo ago
Qwen3.6-27B is the best model in that range that I’ve used for agentic coding by far. I think it’s kinda mid at everything else.
42.
▲
by
sosodev
3mo ago
The Gemma models are so good at vision. It seems particularly important for phones. Also, they write in a much more pleasant manner than Qwen imo.
43.
▲
by
sosodev
3mo ago
I think the difference is asking what happens if all contexts collapse. In the 2000s or 1960s context collapse was something that would happen occasionally when opted into. These days everybody is hyper connected and living in a collapsed
44.
▲
by
sosodev
3mo ago
The thing we must keep in mind with any of the AI-slop style writing is that it was reinforced behavior because humans wanted it.
45.
▲
by
sosodev
3mo ago
I suspect it would depend on the task. DS4-flash does, as previously mentioned, handle quantization very well. Even at 2-bit it's still very coherent.
46.
▲
I think I have LLM burnout
(alecscollon.com)
408 points
by
sosodev
3mo ago
|
363 comments
47.
▲
by
sosodev
3mo ago
It's almost as if HN users aren't all the same.
48.
▲
by
sosodev
3mo ago
He got his name in the credits. The question was if he is owed anything else. The contract he created says he was not. I’m simply suggesting he might need a different contract.
49.
▲
by
sosodev
3mo ago
That’s a fine perspective, but the whole point of law is to guarantee outcomes. The license could easily say “if you make more than $500M, you must pay me $1M”. Why is that not an acceptable solution here?
50.
▲
by
sosodev
3mo ago
I find it odd how we frame fairness in regards to open source software. He licensed his software as MIT. It says anyone can you use it without owing the author anything. So how is it unfair? To be clear, I think that open source maintainers
51.
▲
by
sosodev
3mo ago
In my experience, even with basic project concepts the small models struggle to spin up greenfield stuff. There's just too many decisions to be made and they're not good at that. Modifying existing code is way easier if you don&#x
52.
▲
by
sosodev
4mo ago
The license scheme makes me sad. It reads like a subscription pretending to be a one-time purchase.
53.
▲
by
sosodev
4mo ago
I think it really just depends on your goals. Slow tokens per second is fine by some people if they cost a fraction of a single node setup that can run a trillion param model. If you’re actually running a small business and want to have mul
54.
▲
by
sosodev
4mo ago
I read the comment, thanks. I just disagree with your cost estimate. Even for a small business that needs high throughput they could probably do it for far less than $300k if they aren’t just blindly buying the first big nvidia setup they c
55.
▲
by
sosodev
4mo ago
I don’t have any particular model in mind, sorry. My data is just rough estimates based on my experience with a single node setup. You might need to opt for a 2 or 3 bit model to get the full context window. The KV cache memory consumption
56.
▲
by
sosodev
4mo ago
It depends entirely on what you want to do and think is feasible. Small models can almost certainly run on the computer that you already have. They can do good tool calling.
57.
▲
by
sosodev
4mo ago
You can run a trillion parameter model with decent quality for far less than $300k. A cluster of 4 AMD AI Max 395+ boards with 128GB unified memory each can be had for around $15k. That would run the 4-bit quant of a trillion param model we
58.
▲
by
sosodev
4mo ago
They have it, we just haven’t enabled them. The smart model with a chat box is the wrong abstraction for local. Ideally we would have it built into applications as a clear and easy to use opt-in feature. Like allowing a user to index a fold
59.
▲
by
sosodev
4mo ago
Note that AA's coding index is only made up of two benchmarks: Terminal-Bench Hard and SciCode. I'm skeptical that it makes a good coding index. It ranks Gemma 4 31B above Deepseek V4 Flash. Having used both of those models for a
60.
▲
by
sosodev
4mo ago
Petsitter's default tricks doesn't seem to do much for Qwen3.6, right? JSON mode could be useful I suppose, but that's not really going to make it better at writing code. Do you have any other example tricks? I'm having
More ›