Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
osti
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
osti
3mo ago
I read quite a few posts about this on RedNote.
32.
▲
by
osti
3mo ago
That's not true, some of them are indeed fake, but a lot of them are actually providing real opus at low cost doing what op said.
33.
▲
by
osti
3mo ago
Even as a GLM z.ai fan, I wouldn't pay for their plans. They are just way worse values than gpt or anthropic plans, in terms of both usage and capabilities.
34.
▲
by
osti
4mo ago
If AI is so good, then why don't we have perfect compatibility layers between OS's for games yet?
35.
▲
by
osti
4mo ago
The official API is FP8, which should imply that it's lossless.
36.
▲
by
osti
4mo ago
There are definitely some Chinese bots + actual people (imagine that!) who like to talk up Chinese models, I'm one of them but I like to find out how good these models really are before saying anything. GLM definitely isn't opus l
37.
▲
by
osti
4mo ago
Fun fact: Zhipu aka Z.ai, Knowledge Atlas etc., the company that made GLM, is listed on Hong Kong stock exchange, is up over 10x since the IPO at the beginning of this year.
38.
▲
by
osti
4mo ago
I indeed got a few timeouts yesterday using the official API, I imagine for the coding plan users it'll be even worse.
39.
▲
by
osti
4mo ago
Many other open source models have vision but they don't compare to GLM in terms of coding quality. So I don't think it's because of vision that the frontier models are better, it's more that they are probably just much
40.
▲
by
osti
4mo ago
I don't know what people mean when they say design lol, is it for frontends?
41.
▲
by
osti
4mo ago
Given that DeepSwe is one of the very few coding benchmarks worth taking a look at, this achieves rather excellent result at it (not far from opus 4.8). From looking at the results and my own impression of 5.1 and other models, I think this
42.
▲
by
osti
4mo ago
Given that DeepSwe is one of the very few coding benchmarks worth taking a look at, this achieves rather excellent result at it (not far from opus 4.8). From looking at the results and my own impression of 5.1 and other models, I think this
43.
▲
by
osti
4mo ago
Well, the level of tech is at least on a whole different level at those companies than whatever cohere is doing.
44.
▲
by
osti
4mo ago
Hmmm no? The only way is to deploy your own local model, using anyone else's you are at their whim on what happens to your data.
45.
▲
by
osti
4mo ago
And notable absence of DeepSWE benchmark where they do badly, but somehow a benchmark that was published yesterday is in this announcement.
46.
▲
by
osti
4mo ago
But I think the eventual goal is that documentations won't even be needed. LLM should just itself understand the nuances of frameworks by analyzing their codebase.
47.
▲
by
osti
4mo ago
I think Jane Street is an Anthropic investor, so take it fwiw.
48.
▲
by
osti
4mo ago
Yeah I guess it being vague is more what I meant. But even if you told AI you need to wash the car, then why are you asking AI in the first place whether you should walk there or drive there. The question just doesn't make too much sen
49.
▲
by
osti
4mo ago
Meh, I feel that the car wash test is probably the worst question of all of those LLM test questions. The question is basically logically inconsistent and expect the model to work around the inconsistency.
50.
▲
by
osti
4mo ago
For coding I wouldn't say a year, last year this time claude or gpt definitely weren't able to do what GLM is able to do today, but easily 6 months I'd say. Not sure about other domains though.
51.
▲
by
osti
5mo ago
Reality has both negatives and positives. Gamers Nexus clearly lie on the other side of the spectrum where they overwhelmingly choose the negative stuffs, hence rage baits.
52.
▲
by
osti
5mo ago
I understand their points. But to me they do too much rage baits these days that I can't bring myself to watch their stuff. Why would I let my self get angry lol.
53.
▲
by
osti
5mo ago
Even last year at this time people wouldn't believe it.
54.
▲
by
osti
5mo ago
I heard they are already proficient at assembly languages.
55.
▲
by
osti
5mo ago
A lot of it was beyond me, but this was all the branch names for all the stuff it tried, most of it unsuccessful of course. About 10x perf improvement came from architectural changes, and then 2x from micro optimizations. https://
56.
▲
by
osti
5mo ago
> propose, implement, measure, keep the wins Pretty much what I did to let Codex with gpt5.4xhigh improve my fairly complex CUDA kernel which resulted in 20x throughput improvement.
57.
▲
by
osti
5mo ago
This was published two months ago. Even though it was at a time that open source models are publishing comparable swe bench scores.
58.
▲
by
osti
6mo ago
Can't they write a script to solve rubik cubes?
59.
▲
by
osti
6mo ago
True, but I think for local models, we are mostly considering personal usage.
60.
▲
by
osti
6mo ago
Yeah... I would definitely call 2t/s unusable. For simple chats, I'd want at least 15 t/s. For agentic coding (which this model is advertised for), I'd want good prefill performance as well.
More ›