Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
unrvl22
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
unrvl22
14d ago
you are a joke if you think omp is a joke.
2.
▲
by
unrvl22
1mo ago
someone screenshot?
3.
▲
by
unrvl22
2mo ago
their sub is crap. use opencode go (multiple workspaces) or wait for weights to drop
4.
▲
by
unrvl22
2mo ago
which is the bigger headline that people don't realize. this is 744b and its head to head with Kimi K3 (2.8T), smashes DS v4 pro (1.5T). even Opus and Sol are rumored to be 1.5T+ this is half the size!
5.
▲
by
unrvl22
2mo ago
I was thinking the same thing. It feels truthful, no marketing BS and they call out where they lack behind the best models
6.
▲
by
unrvl22
2mo ago
its kinda crazy with literally no guardrails and a goal, the extremes these AI models can actually go to.
7.
▲
by
unrvl22
3mo ago
a good take.
8.
▲
by
unrvl22
3mo ago
Something like this is nice, where instead of having 1 model with X active experts, you have 10 different models, all small and dense, trained on specific information. and loaded on 10 different servers, with one router.
9.
▲
by
unrvl22
3mo ago
inference is only memory bandwidth limited when targeting higher tps / high single stream tps. the weights only need to be moved across once per forward pass, when you batch say 100 streams per forward pass (which is what most inferenc
10.
▲
by
unrvl22
3mo ago
that 213 wasn't achieved when saturated though. was probably more like 30 tps per stream when doing 2.6k tps.
11.
▲
by
unrvl22
3mo ago
MI355X can perform FP6 operations with the same speed as their FP4 (unique to AMD) - people should be making MXFP6 quants which would be pretty much lossless, and much closer to FP4 performance than FP8
12.
▲
by
unrvl22
4mo ago
look at benchmarks, use the model yourself. Im usually first to call BS on every chinese model that says they are as good as Opus. this is finally the first one that actually is. It is a massive jump from every other previous chinese model.
13.
▲
by
unrvl22
4mo ago
the 2 I mentioned both have a fairly large following, who run benchmarks and absolutely will spot issues.
14.
▲
by
unrvl22
4mo ago
I cancelled my claude sub after realizing I can burn 300m tokens a day of this quality, for $50 a month.
15.
▲
by
unrvl22
4mo ago
Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a month. Some are even offering API rates at 3x lower than the official Z
16.
▲
Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
(github.com)
403 points
by
unrvl22
4mo ago
|
236 comments
17.
▲
by
unrvl22
4mo ago
The municipality of Rio de Janeiro (via its IT company IplanRIO) released Rio-3.5-Open-397B, presented as a homegrown Qwen3.5 fine-tune that beats comparable open models on benchmarks. The linked issue argues it's actually a weighted m
18.
▲
by
unrvl22
4mo ago
AI slop. looks basic as hell.
19.
▲
by
unrvl22
4mo ago
why is deepseek v4 pro a lot lower than flash? where is mimo 2.5?
20.
▲
by
unrvl22
5mo ago
I had no idea those dan and the team were aussies! damn nice, we dont really seem to shine in tech on the world stage.
21.
▲
by
unrvl22
5mo ago
love how it loads instantly and feels smooth. imo useless but still cool