Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kamranjon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
kamranjon
6d ago
Is that true though? It seems like the iphones are doing some type of skin smoothing algorithm but the Huawei is the only one that actually seems to produce actual skin texture and blemishes. My initial read was that the Huawei was the leas
2.
▲
by
kamranjon
7d ago
It is really interesting to see this claim, because i thought the current theory was that typesafe actually repackaged the work from GLiNER[1] - which does seem to be a closer match, and their original paper[2] predates yours by several yea
3.
▲
by
kamranjon
8d ago
"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient t
4.
▲
by
kamranjon
8d ago
Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated
5.
▲
by
kamranjon
9d ago
Someone tell this man about vLLM!
6.
▲
by
kamranjon
17d ago
Is nobody using structured outputs? They use constrained decoding at the generation stage to ensure the probability of tokens that would break the format are set to 0. I kinda figured everyone was doing this at this point.
7.
▲
by
kamranjon
17d ago
Hey there! I do the same but I use dwarfstar at a 2-bit quant: https://github.com/antirez/ds4 I'm curious if you've tried dwarfstar and decided to move to llama.cpp and 3 bit quants or what made you go that r
8.
▲
by
kamranjon
22d ago
I follow llama.cpp pretty closely as I use either llama.cpp itself or projects that depend on it all the time, and one thing that I don't think gets talked about is the sheer scale of community involvement. It seems like a logistical n
9.
▲
by
kamranjon
22d ago
Did you post in the wrong thread?
10.
▲
by
kamranjon
22d ago
I think this account should be banned.
11.
▲
by
kamranjon
23d ago
it's funny that the tagline is Radically Open, but you're immediately hit with http login - maybe this was the wrong link?
12.
▲
by
kamranjon
24d ago
They said Opus 5 medium - which does have an intelligence score of 59 (you have to select it manually from the dropdown to see it)
13.
▲
by
kamranjon
24d ago
Yea I am testing through OpenRouter - have you noticed 3.7 flash being significantly faster?
14.
▲
by
kamranjon
24d ago
They've interestingly left out any mention of speed. I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash i
15.
▲
by
kamranjon
24d ago
I am running 3.8 27b at q6 quant with 160k context on a 32gb video card (arc b70 pro) - I quantized the kv cache at q8 - that is the only trick really - works great.
16.
▲
by
kamranjon
1mo ago
Hmm, makes the 2 bit quants actually seem pretty reasonable…
17.
▲
by
kamranjon
1mo ago
Trained on roughly 1/9th of the training FLOPS - it’s the pretty incredible that they are making these advances and at the same time sharing their learnings in these papers - I wish we saw more of this from US labs.
18.
▲
by
kamranjon
1mo ago
I actually do this with my MBP - it's a LLM server when I'm working - and then when I'm not it's just a really great machine for video editing and other media work.
19.
▲
by
kamranjon
1mo ago
Where did they say that? My understanding of this 3.8-Flash-Next release is that it's a MOE (as per the title of the posting here, 125B a6b)
20.
▲
by
kamranjon
1mo ago
Have you tried FreeToken yourself? I was hoping to find some benchmarks on their github but took a quick pass at their research paper and it seems they're showing ~2x performance on qwen 3.6 35b when compared to llama.cpp - but llama.c
21.
▲
by
kamranjon
1mo ago
I think you misread, it’s 170gb/s for base M6 model and 1.2tb/s for M5 ultra.
22.
▲
by
kamranjon
1mo ago
You would want to get the M5 pro version with 307gb/s if you were interested in running local LLMs.
23.
▲
by
kamranjon
1mo ago
1tb would likely be ~$20k - given the current >$10k price tag of 256gb. Would you still be considering it at that price?
24.
▲
by
kamranjon
1mo ago
You can get the m5 pro in the Mac mini with 307GB/s at 64gb of memory it’s $2899
25.
▲
by
kamranjon
1mo ago
Yeah the article feels as though its describing Khan Academy from 10 or more years ago - it is quite odd. "What he has never had is pedagogical knowledge: an understanding of how people learn, what motivates them, what makes the differ
26.
▲
by
kamranjon
1mo ago
Is anyone familiar with the laws surrounding police basically operating their own pseudo cell towers? I would assume this would be highly illegal for individuals, what sort of hoops did law enforcement need to jump through to get this type
27.
▲
by
kamranjon
1mo ago
Hi there! I actually thought your Dia models were amazing and very natural sounding, I haven’t tried qwen 3 tts yet - has your focus shifted away from building your Dia models and shifted more towards hosting and infrastructure?
28.
▲
by
kamranjon
1mo ago
Just wanted to share this, I found it was a really nice resource to understand how diffusion Gemma worked: https://newsletter.maartengrootendorst.com/p/a-visual-guide-... The really interesting thing to me was that the
29.
▲
by
kamranjon
1mo ago
I wonder if this could be used by insurance companies to determine premiums?
30.
▲
by
kamranjon
1mo ago
Are you using the recommended settings for temperature and such? https://unsloth.ai/docs/models/qwen3.8#recommended-settings Often times I run into issues like this it’s because I am using settings for a different
More ›