Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mutkach
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
mutkach
1y ago
Why would you really need something like that in a non-totalitarian state? Basically, it follows the russian playbook (essentially the same 'language' - safety concerns), but instead of the FSB, who is the beneficiary actor in thi
32.
▲
by
mutkach
1y ago
> 55,000 hours of [IT] experience I didn't know Mike Judge was such a polymath!
33.
▲
by
mutkach
1y ago
How would they not "share information with third parties". You need to sift through the data to make it even remotely useful for "training". You absolutely need to share it with either Amazon (for Mechanical Turk) or wi
34.
▲
by
mutkach
1y ago
I am implying that what OpenAI pays for GPU/hour is much less than $2, because of the discount. That's an assumption. It could be $1, $0.5, no? It could still be burning money for Microsoft/Amazon
35.
▲
by
mutkach
1y ago
Gross margins also don't tell the whole story, we don't know how much Azure and Amazon charge for the infrastructure and we have reasons to believe they are selling it at a massive discount (Microsoft definitely does that, as foll
36.
▲
by
mutkach
1y ago
A full KV-cache is quite big compared to the weights of the model (depending on the context size), that should be a factor too (and basically you need to maintain a separate KV cache for each request, I think...). Also the the token/s
37.
▲
by
mutkach
1y ago
There's also an estimation of how much a KV cache grows with each subsequent token. That would be roughly ~MBs/token. I think that would be the bottleneck
38.
▲
by
mutkach
1y ago
Exactly. Out of all possible interpretations all of them all of them kinda converged to the same conclusion right at the beginning of the "reasoning", isn't that weird. They absolutely did train the models either on existing
39.
▲
by
mutkach
1y ago
Judging from the reasoning trace for the problem of the day - almost all of the models obviously had some presence of IQ training data or at least it could be said that the models are very biased in a beneficial way. From the beginning of t
40.
▲
by
mutkach
1y ago
Not only "AGI" is cancelled but they also sort of admitted that so-called "scaling" "laws" don't work anymore. Scaling inference kinda still works, but obviously is bounded by context size and haystack-and
41.
▲
by
mutkach
1y ago
The focus now is not the model, but the Product - "here we improve the usuability by removing the choice between models", "here is a better voice for tts", "here is a nice interface for previewing html" Only ab
42.
▲
by
mutkach
1y ago
I think it is related to category theory, namely "Yoneda" and Hom functors, but that's a wild guess
43.
▲
by
mutkach
1y ago
> https://github.com/iokasimov/ya/blob/main/Ya/Operators/Handc... This is the most arcane codebase I've seen. It's on par with co-dfns compiler. The frontend syntax also looks like
44.
▲
by
mutkach
1y ago
This is a good take, actually. GPT-OSS is not much of a snowflake (judging by the model's architecture card at least) but TRT-LLM treats every model like that - there is too much hardcode - which makes it very difficult to just use it
45.
▲
by
mutkach
1y ago
> Inspired by GPUs, we parallelized this effort across multiple engineers. One engineer tried vLLM, another SGLang, and a third worked on TensorRT-LLM. We were able to quickly get TensorRT-LLM working, which was fortunate as it is usuall
46.
▲
by
mutkach
1y ago
> Respondents were recruited primarily through channels owned by Stack Overflow. The top sources of respondents were onsite messaging, blog posts, email/newsletter subscribers, banner ads, and social media posts. Since respondents w
47.
▲
by
mutkach
1y ago
Growing disillusionment among programmers (whose productivity gains, by the way, represent the most successful use case for AI yet) is indeed not necessarily a bad thing. What is concerning is that VCs seem to believe we are still in the ex
48.
▲
by
mutkach
1y ago
> https://survey.stackoverflow.co/2025/ai
49.
▲
by
mutkach
1y ago
>230MW rather small compared to what was advertised previously >joint venture between Nscale and Aker which seems to imply it is not The Stargate (Oracle, Softbank, MGX?)