Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kamranjon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
91.
▲
by
kamranjon
3mo ago
By evidence I mean logs, I mean IP addresses, I mean timestamps. They claim millions of requests, let’s see literally any of them? I don’t consider a tweet by Denise Wu, who works at Anthropic, to be reproducible evidence. I don’t consider
92.
▲
by
kamranjon
3mo ago
It’s so funny to me that Anthropic can make claims like this one with zero evidence provided. DeepSeek and others like Minimax are publishing deep research on Multi-Head Latent Attention and Mixture of Experts, Multi-Token Prediction, novel
93.
▲
by
kamranjon
3mo ago
The fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusable. These Chinese labs are producing novel models, publishing their techniques
94.
▲
by
kamranjon
3mo ago
If you look at this chart here it seems the tiny model has a WER of ~12%… not sure about the micro model: https://github.com/moonshine-ai/moonshine#when-should-you-ch...
95.
▲
by
kamranjon
3mo ago
Allen AI is fighting the good fight with their Olmo models - I’m hoping for a new release soon. Also huggingface has released some pretty nice smaller models with open pipelines like Smollm.
96.
▲
by
kamranjon
3mo ago
The difference is that because they are open weights, it’s trivial to decensor the models with abliteration and you’ll generally see decensored versions within days of new releases - something you just can’t do with proprietary models from
97.
▲
by
kamranjon
3mo ago
I also think Allen AI with their Olmo models are doing great stuff by releasing full training data and pipelines and checkpoints which I’m not sure any other US company is really coming close to transparency-wise.
98.
▲
by
kamranjon
3mo ago
“QuickJS is to this PSP what V8 is to Node: the host hands it a strike API surface and a ui API surface, and the same openstrike.js bundle boots against them on every target.” This comparison seems backwards? Also I have no idea what it mea
99.
▲
by
kamranjon
3mo ago
Hearing a world leader talk about open source is huge. I think it highlights pretty starkly the difference in approach between the U.S. and China and will likely in time be seen as a pretty historic statement.
100.
▲
by
kamranjon
3mo ago
Where did you hear about the deepseek release? Would love to follow the same source.
101.
▲
by
kamranjon
3mo ago
“Alongside Inkling we are sharing a preview of Inkling-Small, a 276B-parameter Mixture-of-Experts model (12B active, vs. 41B for Inkling) with a different performance/latency trade-off.” Buried at the end there is the details I was mos
102.
▲
by
kamranjon
3mo ago
70% seems a little extreme? Here is a helpful mobile coverage map: https://www.fcc.gov/BroadbandData/MobileMaps/mobile-map A few years ago I actually traveled all around the US and worked remotely with a 4g unlimi
103.
▲
by
kamranjon
3mo ago
I am traveling in Europe currently and got a Saily SIM card - 5g coverage is really good and seems to be expanding fast - a small village I visited 2 years ago that had no cell coverage at all I was getting 250mbps download, faster than tha
104.
▲
by
kamranjon
3mo ago
Doesn’t this suggest you aren’t properly running the model?
105.
▲
by
kamranjon
3mo ago
Curious why you find that’s a useful test since it seems to be solely measuring training data memorization, something you’d expect to degrade from quantization.
106.
▲
by
kamranjon
3mo ago
When new models are released (I realized this is qwen 3.6 but the quant is novel) - it takes a few days for the kinks to get worked out - you’ll likely have better luck if you give it a few days and try again.
107.
▲
by
kamranjon
3mo ago
Yea I don’t use sub-agent style workflows. I often use planning patterns and generate markdown for really complex tasks, but never got into the sub-agent thing. I am likely closer to AI-assisted, though I am heavily using pi coding agent an
108.
▲
by
kamranjon
3mo ago
DeepSeek V4 Flash with DwarfStar: https://github.com/antirez/ds4 The 2 bit quants are really good. I have a lot of memory so I can squeeze it all in at ~80gb.
109.
▲
by
kamranjon
3mo ago
I’m using DeepSeek V4 Flash on 128gb mbp - it’s a bit different using a 200b+ param model. It’s MoE so performance is acceptable. It will still malform a tool call every now and then, but the capabilities are so far ahead anything else that
110.
▲
by
kamranjon
3mo ago
After using a highly capable 2-bit quant as my daily driver for months now, I get pretty excited about releases like this. After a few days for the kinks to be worked out, I’ll be excited to try it.
111.
▲
by
kamranjon
3mo ago
I haven't read No Country yet - I think at the time it was Suttree which is definitely in my top 5 - I also read The Road recently and was pretty blown away, really quick read and very meticulously structured, I loved it. I'll mak
112.
▲
by
kamranjon
3mo ago
I actually read the first book and it was so poorly written it made me wonder if I should continue, because I did find the general story quite engaging. I've heard it gets better/tighter in subsequent books, but it was the only Ki
113.
▲
by
kamranjon
3mo ago
Oh that’s great - I’m generally more interested in single machine use cases but the RDMA stuff is super interesting. I do wonder if you are planning to eventually delegate some of the merging responsibilities? I see some other projects foll
114.
▲
by
kamranjon
3mo ago
Where do these estimates come from?
115.
▲
by
kamranjon
3mo ago
I actually have asked many people, because it's something that's frequently on my mind (I had a bit of a revival with reading and do have some strong opinions about how it's helped me). Are you suggesting that you listen to a
116.
▲
by
kamranjon
3mo ago
Some of the most exciting engineering work is happening in the DS4 repo - and I'm watching it almost like a sports game. When the DSpark paper came out[1] the next day we had folks attempting to implement, working together, validating
117.
▲
by
kamranjon
3mo ago
In your examples you are always doing something else while listening to an audio book. This was the authors point. Reading is the activity - and I definitely agree - I still listen to audio books, but for a different purpose, to avoid bored
118.
▲
by
kamranjon
3mo ago
I’ve actually been really interested in Minimax M3 - seems like it flew under the radar but size wise might actually be runnable for local inference with a footprint somewhere between Deepseek V4 flash and pro. Has anyone used the new Minim
119.
▲
by
kamranjon
3mo ago
The Minimax paper was published in June 2026 coinciding with the Minimax M3 release - I’m not sure how the repo you posted here could have been an implementation of Minimax sparse attention when it was updated over a year ago?
120.
▲
by
kamranjon
3mo ago
I dunno I use LMStudio pretty regularly and the MLX folks and the community usually have MLX versions of new model releases up within a day or two.
More ›