Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kamranjon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
kamranjon
1mo ago
I think if a real human doctor completely fabricated a drug usage history for one of their patients, that would also be news.
32.
▲
by
kamranjon
1mo ago
What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.
33.
▲
by
kamranjon
1mo ago
A no-thinking pelican! I hope to see more, it's surprisingly good for just 2 minutes.
34.
▲
by
kamranjon
1mo ago
I actually disagree that this doesn’t mean anything. I understand the contention that it’s not measuring the quality of the model in general, but I think it is measuring something useful. A good example of this is planning hardware projects
35.
▲
by
kamranjon
1mo ago
Have you thought about running a second tier of the Pelican benchmark where you see which model makes best pelican on lowest or no reasoning settings? I think that'd be pretty interesting and might help highlight which models have a ba
36.
▲
by
kamranjon
1mo ago
You can buy two b60s for $1300 right now (650 each) if you want a total of 48gb. Intel recently raised the price on all of their gpu's except the b60 series, so they are currently the best deal per gb I think.
37.
▲
by
kamranjon
1mo ago
Since Qwen 3.6 27b outperforms Gemma 4 26b in most benchmarks I'm not sure the value - also Gemma 26b is a MOE model whereas this is a dense model, so not typically direct competitors at their sizes - Gemma 4 31b comparison would be in
38.
▲
by
kamranjon
1mo ago
It's useful logs which i think is an important distinction.
39.
▲
by
kamranjon
1mo ago
I don't really think the distinction here is relevant. If the end result is the equivalent of lying, cheating or stealing - then the problem still exists and it needs to be solved.
40.
▲
by
kamranjon
1mo ago
New coding harness that seems to have some novel concepts and one of the pretty cool things on their landing page for it here: https://deepseek.com/harness/en/ is the Every Run is Traceable view: "Everything
41.
▲
by
kamranjon
2mo ago
this seems to be a countdown for the 2.4t model - which is gigantic and is not as exciting for those running models locally
42.
▲
by
kamranjon
2mo ago
I actually have been exploring this very thing! I think the best option right now, since Apple has raised prices and Mac minis are basically impossible to get your hands on, is to build your own micro-itx machine. I actually built a mini-it
43.
▲
by
kamranjon
2mo ago
You might look at this and and be a bit disappointed by the performance against qwen and gemma models - but this is an entirely open source training pipeline, this is quite impressive and I don't think another model this performant exi
44.
▲
by
kamranjon
2mo ago
I honestly think of the T5 model family to sort of be the real beginning of this open model craze - I know BERT was already popular for classification etc, but T5 was the first sort of generally useful model, was exceptionally simple to fin
45.
▲
by
kamranjon
2mo ago
1 to 200 is a pretty big spread to between losing 18k dollars and making 3.6 milllion - do you have any actual numbers on the value produced from this 18k investment?
46.
▲
by
kamranjon
2mo ago
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
47.
▲
by
kamranjon
2mo ago
The sort of sad but interesting thing here is that while this claims to be a more direct translation, Claude models were certainly trained on all of the pre-existing translations and this is likely going to directly influence the output in
48.
▲
by
kamranjon
2mo ago
This wasn't written with AI... obviously... I feel like there might need to be a new definition for whatever this paranoia is called because it's getting a bit out of hand.
49.
▲
by
kamranjon
2mo ago
Because street photography is very spontaneous it’s pretty common practice to set an aperture of 8 and just snap away - it’s a helpful trick for rangefinder cameras that often take some time to pull focus.
50.
▲
by
kamranjon
2mo ago
“Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach…” Very excited to see how it performs, I’ve been a bit skeptical of the
51.
▲
by
kamranjon
2mo ago
“While technically true the hallucination rates on modern models is low…” Isn’t this entirely context dependent? Where did you get the information that modern models have low hallucination rates? I’d love to see the benchmark if there is on
52.
▲
by
kamranjon
2mo ago
This seems really interesting - I was curious about this line from the website. “The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every
53.
▲
by
kamranjon
2mo ago
Do you have an example? Would love to read one.
54.
▲
by
kamranjon
2mo ago
I love this analogy and think it possibly also applies to search which just doesn’t work anymore and is full of AI generated garbage. What the hell happened to stack overflow? I don’t think I’ve gotten a google result for stack overflow in
55.
▲
by
kamranjon
2mo ago
In what way is it similar?
56.
▲
by
kamranjon
2mo ago
This is amazing and I think will probably end up being a pretty important development. I was just reading this great breakdown of how diffusion Gemma works: https://newsletter.maartengrootendorst.com/p/a-visual-guide-..
57.
▲
by
kamranjon
2mo ago
I actually run it as a server - so most of the time I don't have to listen to it right next to me - it's just sitting in another room in my house - but I often am traveling with it and will have it sitting right next to my coding
58.
▲
by
kamranjon
2mo ago
The really interesting thing about this is how big of a jump was achieved with just extra fine-tuning here. No structural changes to the model, just more data, compute and time. It makes me pretty excited for the future of small models - DS
59.
▲
by
kamranjon
2mo ago
Generally get 20-25tps - prefill is pretty good around 400-450tps. I have been using compaction at around 100k tokens but mostly just cause it was the default in pi coding agent - might see if i can expand it a bit.
60.
▲
by
kamranjon
2mo ago
Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing t
More ›