Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sosodev
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
91.
▲
by
sosodev
6mo ago
Qwen3-coder-next is way worse than Sonnet 4.5. Also, despite he lack of "coder" in the name Qwen3.5 is much better at coding than Qwen3-coder-next so you might want to check that out.
92.
▲
by
sosodev
6mo ago
I don't know how well it performs, but you can extend Qwen3.5 to 1 million token context using YaRN. Also, Nemotron 3 Super was recently released and scales up to 1 million token context natively.
93.
▲
by
sosodev
6mo ago
These prices are insane. You can buy all (most?) of the lenses they’re recreating for a fraction of the price and adapt them to a mirrorless camera no problem. I bought a Helios 44-2 recently for $100 and adapted it to my camera for like $1
94.
▲
by
sosodev
7mo ago
I spent some time trying to understand this paper and I think calling this a new attention mechanism is a bit misleading. As a dead comment pointed out this is much closer to RAG. It's not exposing all 100M tokens directly to the model
95.
▲
by
sosodev
7mo ago
I think it's hard to know where to draw the line between derivative product and something unique. If we follow your logic that TSMC hasn't done anything new, then aren't all computer manufacturers just rehashing the ENIAC or
96.
▲
by
sosodev
7mo ago
TSMC. They dominate the semiconductor market because they're consistently first to market with the world's most advanced chip fabrication.
97.
▲
by
sosodev
7mo ago
Can we actually align incentives at scale? It seems to me that if it were possible we would live in a utopia.
98.
▲
by
sosodev
7mo ago
Most people are using something in the llama family for inference. Llama server is my go to. Unsloth guides describe how to configure inference for your model of choice.
99.
▲
by
sosodev
7mo ago
What models are you testing? A 120b model with hybrid attention should fit within 80gb of VRAM fine at a 4-bit quant. Also, 4-bit quants that are done well are generally fine. They certainly don’t make the model unusable.
100.
▲
Whole-Brain Connectomic Graph Model Enables Whole-Body Locomotion Control in Fly
(arxiv.org)
2 points
by
sosodev
7mo ago
|
0 comments
101.
▲
by
sosodev
7mo ago
I didn't intend to. I think that domesticated animals have long had a harmonious relationship with humans so I find it a bit difficult to believe that it's always an ethical dilemma. Pets are just the most obvious lens to identify
102.
▲
by
sosodev
7mo ago
I'm skeptical of this claim because there's clearly a growing population that hates the idea of putting anything they don't understand in their bodies. Genetically modified vegetables, food dyes, vaccines, etc. I find it hard
103.
▲
by
sosodev
7mo ago
Are all pets suffering?
104.
▲
by
sosodev
7mo ago
This is a false dichotomy. The choice is not lab grown or suffering. Farmed animals could live happy, healthy lives and then be culled in a humane way. The problem is that it costs slightly more and our society is more concerned with cost t
105.
▲
by
sosodev
7mo ago
I’ve tried it via openrouter. It’s very good, but for some tasks frontier models are still significantly better. For me, the 122b model is good enough on my own hardware that the downsides can be worked around for the sake of privacy and co
106.
▲
by
sosodev
7mo ago
I’ve been running it via llama-server with no issues. Running the latest Bartowski 6-bit quant
107.
▲
by
sosodev
7mo ago
Around 20ish tokens a second with 6-bit quant at very long context lengths on my AMD AI Max 395+ I’m trying to use local models whenever possible. Still need to lean on the frontier models sometimes.
108.
▲
by
sosodev
7mo ago
Some of the early quants had issues with tool calling and looping. So you might want to check that you're running the latest version / recommended settings.
109.
▲
by
sosodev
7mo ago
In my experience Qwen3.5 is better even at smaller distillations. From what I understand the Qwen3-next series of models was just a test/preview of the architectural changes underpinning Qwen3.5. So Qwen3.5 is a more complete and well
110.
▲
by
sosodev
7mo ago
I've noticed that open weight models tend to hesitate to use tools or commands unless they appeared often in the training or you tell them very explicitly to do so in your AGENTS.md or prompt. They also struggle at translating very bro
111.
▲
by
sosodev
7mo ago
I really hope this doesn't hinder development too much. As Simon says, Qwen3.5 is very impressive. I've been testing Qwen3.5-35B-A3B over the past couple of days and it's a very impressive model. It's the most capable ag
112.
▲
by
sosodev
7mo ago
The compute speed is definitely correlated with the memory consumption in LLM land. More efficient attention means both less memory and faster inference. Which makes sense to me because my understanding is that memory bandwidth is so often
113.
▲
by
sosodev
7mo ago
The model that processes search results is tiny and dumb. You shouldn't compare it to the frontier models that are solving complex math problems.
114.
▲
by
sosodev
7mo ago
My understanding, from listening/reading what top researchers are saying, is that model architectures in the near future are going to attempt to scale the context window dramatically. There's a generalized belief that in-context l
115.
▲
by
sosodev
8mo ago
I challenge each and every one of you to make a pie by the end of the month. I made one, for the first time in my life, last week. It brought me tremendous joy not only to make it, but to have something nice to share with friends.
116.
▲
by
sosodev
8mo ago
I think you're under-estimating how much personal taste applies in that industry. Yes, there's a lot of free content but it's often low quality and/or difficult to find for a particular niche. The OF pages, and other pai
117.
▲
by
sosodev
8mo ago
Doesn't Grok allow users to create lewd content or did they roll that back? Also, I suspect that we'll soon see the same pattern of open weights models following several months behind frontier in every modality not just text. It&#
118.
▲
by
sosodev
8mo ago
I see. It seems the looping is a bug in the model weights but there are bugs in detecting various outputs as identified in the PR I linked.
119.
▲
by
sosodev
8mo ago
Thanks for the additional info. I suspected that MiniMax M2.5 might be a bit too much for this board. 230B-A10B is just a lot to ask of the 395+ even with aggressive quantization. Particularly when you consider that the model is going to sp
120.
▲
by
sosodev
8mo ago
Have you tried Qwen3 Coder Next? I've been testing it with OpenCode and it seems to work fairly well with the harness. It occasionally calls tools improperly but with Qwen's suggested temperature=1 it doesn't seem to get stuc
More ›