Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fastball
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
61.
▲
by
fastball
2mo ago
Re-light seems to be one of their biggest remaining obstacles in a lot of these flight tests.
62.
▲
by
fastball
2mo ago
The problem with the rest of inference is that changes are not trivially correct or incorrect, as they are with the tokenization layer.
63.
▲
by
fastball
2mo ago
Tokenization is <0.1% of the inference time for the first token in the same way it is <0.1% for the last.
64.
▲
by
fastball
2mo ago
Claude Code has enterprise subscriptions. No $200 max plan but you can buy $100 premium seats. (and if you are using per-token billing as an enterprise user without first maxing out a premium seat you are very silly)
65.
▲
by
fastball
2mo ago
I know many, many engineers who are paying something like 2-5% (via subscription) of what their usage would cost if billed by API tokens. I know some down to about 1% ($200 Max plan vs $20k in tokens per month)
66.
▲
by
fastball
2mo ago
The most interesting result for me is that they apparently prompted the models to optimize for SSIM, but many of the models trend worse over time. I suppose because viewing the canvas always comes after drawing, and they didn't give
67.
▲
by
fastball
2mo ago
I'm sure that will change sooner rather than later, otherwise enterprising hackers will be able to claim that the model they were using went rogue.
68.
▲
by
fastball
2mo ago
Oh yeah, indeed. Which is the rebrand haha
69.
▲
by
fastball
2mo ago
Did you try FreeInk? I was debating which (between CrossPoint and this) to flash my new X3 with.
70.
▲
by
fastball
2mo ago
I don't think that is the takeaway at all.
71.
▲
by
fastball
2mo ago
Yes, this (imo) is a clear result of benchmaxxing. You can get a much better score on most "intelligence" benchmarks by massively over-saturating reasoning. This looks good on those, but for actual daily usage makes the models muc
72.
▲
by
fastball
2mo ago
In my experience, the Chinese models are much more benchmaxxed than their frontier lab competitors, so I'm taking these results with a fairly large helping of salt.
73.
▲
by
fastball
2mo ago
But the scaling needs to be relevant to the actual actions you are concerned about. It's a stretch that any software Google/DeepMind/etc is selling to DHS is allowing / helping them to scale the murder part of their op
74.
▲
by
fastball
2mo ago
What about Meta?
75.
▲
by
fastball
2mo ago
Is the idea that with worse technology, DHS will kill fewer people?
76.
▲
by
fastball
2mo ago
Line go down discovery is acceptable (that is what selling a share is). The reason you might not want options trading very early after an IPO is because the market is frothy enough without the additional layer of complexity.
77.
▲
by
fastball
3mo ago
But that is my point: if benchmaxxing was all the labs were doing, then surely the dumber model could/would have equivalent performance? Rather than noticeably worse perf on a (somewhat trivial to game) test.
78.
▲
by
fastball
3mo ago
On the one hand: yes, pelicans on bikes are definitely in the training set at this point. On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model
79.
▲
by
fastball
3mo ago
More RLHF is in fact scaling.
80.
▲
by
fastball
3mo ago
The analogy of Chesterton's Fence does not imply / require that the fence has been "always there".
81.
▲
by
fastball
3mo ago
"The Internet" was not a bubble. Companies with no long-term business model / sufficient product-market fit that were riding hype were the "dotcom bubble". But when those companies crashed, nobody said "I reall
82.
▲
by
fastball
3mo ago
I never wanted the IP from dotcom bubble companies.
83.
▲
by
fastball
3mo ago
If it's a bubble, why do you care about frontier models?
84.
▲
by
fastball
3mo ago
Is that what Flock does?
85.
▲
by
fastball
3mo ago
"Graphics programming" is definitely not equivalent in scope (in the analogy) to "the entire transportation vehicle industry".
86.
▲
by
fastball
3mo ago
"These annoying, jaded horse-drawn cart builders, cautioning youngsters from getting into the field in 1908."
87.
▲
by
fastball
3mo ago
This isn't a CLI, so not really like Claude Code. Looks more like Cursor or Conductor.
88.
▲
by
fastball
3mo ago
I explicitly said it is your right to operate that way. But that doesn't mean your unproven accusations ("the company is evil") are true. It just means that is how you are choosing to operate / that is the standard of
89.
▲
by
fastball
3mo ago
You actually need to demonstrate that though. I have seen no evidence of Mullvad (again, as a company) behaving in a racist or anti-immigrant manner. Until that has been demonstrated, you cannot just say "this guy behaves this wa
90.
▲
by
fastball
3mo ago
Two things: 1. People definitely start companies with a certain set of values and behaviors (as a company) and do entirely separate things in their private life. This is trivially true. 2. I don't think the personal values and the busi
More ›