Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
osti
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
osti
6mo ago
Oh i’m fully aware of that lol
62.
▲
by
osti
6mo ago
I think this one is only about 600GB VRAM usage, so it could fit on two mac studios with 512GB vram each. That would have costed (albeit no longer available) something like less than 20k.
63.
▲
by
osti
6mo ago
Maybe open source == communism
64.
▲
by
osti
6mo ago
I can't vouch for whether or not it can beat human experts though because I'm no CUDA expert myself. The original CUDA code were human written and I first let codex adapt it to my specific use case. Then I basically let codex gene
65.
▲
by
osti
6mo ago
I was using codex cli with 5.4xhigh. So it was able to iteratively improve from simple prompts on my part (can you give some architectural ideas to improve the performance? And once it does, I just say can you implement and benchmark it). I
66.
▲
by
osti
6mo ago
For me it was able to try out different architectures for perf improvement, then once it's settled on some good architectures, it can do lower level optimizations on them by profiling the code etc.
67.
▲
by
osti
6mo ago
Yup I've mentioned this in another thread, I got gpt 5.4xhigh to improve the throughout of a very complex non typical CUDA kernel by 20x. This was through a combination of architecture changes and then do low level optimizations, it di
68.
▲
by
osti
6mo ago
I had a very complex cuda kernel and codex cli managed to improve the throughout 20x.
69.
▲
by
osti
6mo ago
But is arc-agi really that useful though? Nowadays it seems to me that it's just another benchmark that needs to be specifically trained for. Maybe the Chinese models just didn't focus on it as much.
70.
▲
by
osti
6mo ago
That is ture, but the revenue of the artisanal stuff is probably only a very low percentage of the overall market, which would imply a lot of software engineers would have to exit the field. Which is what we here don't want to see.
71.
▲
by
osti
6mo ago
Doesn't the chat version of chatgpt or gemini also have interleaved tool calls, so do those also count as with harnesses?
72.
▲
by
osti
7mo ago
Seems like the high compute parallel thinking models weren't even needed, both the normal 5.4 and gemini 3.1 pro solved it. Somehow Gemini 3 deepthink couldn't solve it.
73.
▲
by
osti
7mo ago
Actually yeah, I shouldn't have said that regex is a microbenchmark, it's indeed an important one.
74.
▲
by
osti
7mo ago
There's a fine line between making ppl civilized and fascism-like level of control. And I believe Japan errs on the other side too much with their ridiculous number of such rules in all areas of life. Even though I recently visited Jap
75.
▲
by
osti
7mo ago
[flagged]
76.
▲
by
osti
7mo ago
More like oppressed people by all those bs rules.
77.
▲
by
osti
7mo ago
I took a look, it's not bad but it seems to contain too many micro benchmarks like regex or primes. Geekbench at least has clang which is a subscore that I always look at.
78.
▲
by
osti
7mo ago
Not true. Geekbench, especially single threaded benchmark, is probably the best we got, it has a bunch of workloads, unlike many other benchmarks like cinebench for example. And they publish all the results on their website, so you can dig
79.
▲
by
osti
7mo ago
It's only that one number that is for sonnet.
80.
▲
by
osti
7mo ago
Their company is called Anthropic after all.
81.
▲
by
osti
8mo ago
ByteDance never really open sourced their models though. But I agree, they will only open source when it doesn't really matter.
82.
▲
by
osti
8mo ago
That's what I found with some of these LLM models as well. For example I still like to test those models with algorithm problems, and sometimes when they can't actually solve the problem, they will start to hardcode the test cases
83.
▲
by
osti
8mo ago
So now you don't want capitalism?
84.
▲
by
osti
8mo ago
Somehow regresses on SWE bench?
85.
▲
by
osti
9mo ago
Social "science" be social science.
86.
▲
by
osti
11mo ago
I saw many complaints about assetto corsa evo early access about the slow pace of development after EA release. So I'm not sure if I wanna "beta test" this one.
87.
▲
by
osti
11mo ago
That's why reading comments about geopolitics on the Internet is largely useless. Big news! A country's population supports its own country on international stage! If you go on Chinese social media, it'll be mostly about how
88.
▲
by
osti
11mo ago
Look at IXUS or VEU for example, in the past 5-10 years, or even longer, they have significantly underperformed US indices.
89.
▲
by
osti
11mo ago
Yeah, and those have underpermed historically and it's definitely not recommended by most people.
90.
▲
by
osti
11mo ago
That's the biggest problem I have with the recommendation to buy indices as if indices grow at >8% annually is an natural law. Many (most) indices of countries in the world performed way less than 8%. US performed exceptionally well
More ›