Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eli
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
eli
3mo ago
It isn't.
92.
▲
by
eli
3mo ago
I actually really like subjective benchmarks, so long as it's a human (ideally me) grading the results. LLM as judge never made much sense.
93.
▲
by
eli
3mo ago
Fireworks.ai is solid. And if you care more about speed than cost they have a "fast" variant that I think just throws more hardware at the model for about 2x the cost.
94.
▲
by
eli
3mo ago
Of the accounts involved, yeah. So they can lock them out.
95.
▲
by
eli
3mo ago
Seems like a pretty straightforward approach to collecting session logs from a bunch of different people/devices would be to have them all set their base url to proxy.deepseek.whatever which logs the data and forwards to the real API.
96.
▲
by
eli
3mo ago
For being flagged as possibly a competitor? They nuke your account.
97.
▲
by
eli
3mo ago
Are you comparing the cost of hosted Opus to running Qwen 3.6 locally? That doesn't really seem fair.
98.
▲
by
eli
3mo ago
I assume it’s just oversubscribed. I’m sure it “can” go faster. But yeah that was my point.
99.
▲
by
eli
3mo ago
This seems a little fanciful. There's really no comparison between a model that Anthropic allows Google and Amazon to host with one that has been downloaded hundreds of thousands of times and has dozens of public inference providers.
100.
▲
by
eli
3mo ago
I'm skeptical of how fast "up to" 750t/s really means. Maybe if they make it extremely expensive so it frees up enough capacity? GPT‑5.3‑Codex‑Spark currently runs on Cerebras chips and it's giving me around 150t&#x
101.
▲
by
eli
4mo ago
Isn’t that definitionally impossible? If they tell you about it then it’s not a shadow ban.
102.
▲
by
eli
4mo ago
I don't think Google should also be allowed to remain in charge of Chrome at all but here we are.
103.
▲
by
eli
4mo ago
Maybe. But also Flash, on the mobile devices of that time that did support it, was a miserable experience. Slow and broken and drained the battery.
104.
▲
by
eli
4mo ago
How do you know what the founders sincerely believe?
105.
▲
by
eli
4mo ago
If you were playing a text based game, wouldn't you try a few out? I imagine there are a fair number of war games in the training data and not so many actual transcripts of internal military force deliberations.
106.
▲
by
eli
4mo ago
The other "cheating" examples are even worse. It's wild to me that people keep designing benchmarks where the answer is lying around on disk or in the git history. "Hardening" the benchmark with strongly worded prom
107.
▲
by
eli
4mo ago
The word mythos means roughly the same as "myth" and dates to 1753.
108.
▲
by
eli
4mo ago
So it can write code to prevent the problem described?
109.
▲
by
eli
4mo ago
I think this is a worthwhile argument, but you do it a disservice by spamming it in trollish comments
110.
▲
by
eli
4mo ago
It's unethical to price it in a way not everyone can afford?
111.
▲
by
eli
4mo ago
GLM 5.1 is very good. Definitely a contender for best open weight coding model. Nothing like 4.7. But quite a bit more expensive than MiMo 2.5 Pro. Like 5x to 10x more on my little tests, at least by the API rates.
112.
▲
by
eli
4mo ago
Like what?
113.
▲
by
eli
4mo ago
Neat. The frontier models have gotten pretty impressive, but they're all a bit too slow for interactive, human-in-the-loop coding. It incentivizes vibecoding and running multiple agents in parallel. A fast agent feels more like a partn
114.
▲
by
eli
4mo ago
"Lying" is not supported by the evidence. In the context of bot traffic on the web, looking at only GETs for HTML is a reasonable approach. If you're counting all requests for all assets then a single page view of nytimes.com
115.
▲
by
eli
4mo ago
Yeah I think it's a bad thing. It's context about how open source code was written that is lost. And I guess maybe there's no such thing as bad press but at least in this cases it doesn't seem like effective marketing fo
116.
▲
by
eli
4mo ago
Well yeah if they provide open access to content then AI labs wouldn’t have to pay them for bull access.
117.
▲
by
eli
4mo ago
Makes sense as part of a larger coding workflow, especially if it’s fast. Using a trillion parameter model to figure out how to call a targeted edit tool or generate a commit message is a waste. Also narrow tasks like “make the background d
118.
▲
by
eli
4mo ago
I think it’s more correct to say they charge subscription users much much less. I assume less even than the cost of providing the inference, if you actually are using it.
119.
▲
by
eli
4mo ago
Getting fewer people owning homes that nobody lives in is a good thing. Where are you getting the idea that these homes are on the rental market? Why aren’t they being rented? I also don’t actually think this applies to buildings that are u
120.
▲
by
eli
4mo ago
Yes it is proposed for only the new second home tax. It does not currently exist at all.
More ›