Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jrflo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
jrflo
1mo ago
There is no evidence of fraud that I have seen beyond complete speculation from the likes of Zitron. He basically just says "Look how big the numbers are! It's so big it must be fraud!"
62.
▲
by
jrflo
1mo ago
Thanks for putting this together, his writing really aggravates me and it was nice to see all of his predictions put together in one spot. It's interesting how he has pivoted from denying that LLMs are useful to accusing frontier labs
63.
▲
by
jrflo
1mo ago
I never said they did
64.
▲
by
jrflo
1mo ago
The point of ARC is essentially an "IQ Test" for AI systems. It is meant to cover abstract reasoning capabilities of generally-intelligent systems like LLMs. What the author did here was build a system that only solves ARC problem
65.
▲
by
jrflo
1mo ago
Mathematics not being axiomatically complete doesn't mean you can't have crazy progress from a formalized and mechanized systems. It just means that there are corners you can't reach mechanically, but we don't know if th
66.
▲
by
jrflo
2mo ago
The luna cost cuts were real though, not a one time promotion or something, due to some optimization (probably distillation?) that openai did.
67.
▲
by
jrflo
2mo ago
Depends on compute capacity. If we become supply constrained on tokens, then prices will necessarily go up.
68.
▲
by
jrflo
2mo ago
This may just be a classic case of Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets
69.
▲
by
jrflo
2mo ago
I don't think that AI has to influence every single person's daily life to make a huge impact. I think for probably 90% of people it's pretty useless out side of being a better search engine. But that doesn't mean it won
70.
▲
by
jrflo
2mo ago
I know what you're saying, there is a lot of speculation going on in AI right now. The difference is there is hard evidence of AI capabilities accelerating now: we've gone from toy chatbots to sort-of useful coding helpers to agen
71.
▲
by
jrflo
2mo ago
Yes, from my experience it is highly effective
72.
▲
by
jrflo
2mo ago
My reaction is: that's an appropriate amount. AI is probably the biggest thing to happen (in tech) in the last two decades, so I think it's fair that it's taking up a lot of the discourse.
73.
▲
by
jrflo
2mo ago
Whenever I see a study like this I assume that the researchers won the p<0.05 lottery
74.
▲
by
jrflo
2mo ago
I've never used Pi but I don't see why you can't use stock codex or claude code for the same purpose, what makes Pi special? I've built plenty of custom harnesses on top of claude code and codex using custom skills or si
75.
▲
by
jrflo
2mo ago
Stuff like this is so nice, I wonder where it will go in the future. Recently on my PC I had a failed windows update that killed the bluetooth driver, I tried to fix it for over an hour and got no where (even consulting chat gpt). I just le
76.
▲
by
jrflo
2mo ago
I really hope it gets legislated away or kneecapped to the point where it isn't so effective. With all the legislation coming out regarding kids and social media, I hope that people will start to draw a line to the fact that maybe adul
77.
▲
by
jrflo
2mo ago
I build a safari extension to do just this (block all content that isn't explicitly subscribed to by the user) and it's a remarkably better experience. This is my hot take that most people seem to disagree with: the vast majority
78.
▲
by
jrflo
2mo ago
It's just the outrage story of the month. There will be some new moral panic in September and we'll all move on. The attention span for topics like this on social media is quite short.
79.
▲
by
jrflo
2mo ago
It's much more tangible for your average person than circular financing deals or chip supply constraints. And unfortunately social media is an amplifier for tangible misinformation.
80.
▲
by
jrflo
2mo ago
They've been really struggling all of 2026. All of these "limited time" promos, just to have less usage than codex, and the shenanigans around "peak hour" reduction earlier in the year.
81.
▲
by
jrflo
2mo ago
The net effect of a crash would be far worse than the benefit of getting a bit of a discount on RAM and storage.
82.
▲
by
jrflo
2mo ago
Yes, and it's only for a month. This is an ad.
83.
▲
by
jrflo
2mo ago
Make new mathematical discoveries, and the security capabilities of the closed models have not yet been rivaled by open ones. Also, just because a model is open now doesn't mean it will always be. If/when China or meta catches up
84.
▲
by
jrflo
2mo ago
If you're primarily writing code yourself or meticulously reviewing the output from agents, then you're right. However, if you tried to have any of those models one-shot an app or do some highly agentic work, they would certainly
85.
▲
by
jrflo
2mo ago
Well, if Anthropic actually makes $200B revenue in 2028 then that valuation is justified, depending on how fat their margins are.
86.
▲
by
jrflo
2mo ago
What commercial setting do you want to use a Lean theorem-proving agent in?
87.
▲
by
jrflo
2mo ago
Was curious to try it out but it didn't want to reply...
88.
▲
by
jrflo
2mo ago
As an overall pro-AI person... Please don't let it do the entire UI design for your site. 99% of this site is completely useless.
89.
▲
by
jrflo
2mo ago
Interesting! I'll definitely give that a shot then.
90.
▲
by
jrflo
2mo ago
That kind of result makes me suspicious of benchmaxxing. Qwen 27B is 100x smaller than Opus 4.7. Is it really 100x more parameter-efficient? Two orders of magnitude is hard to believe. I don't have the hardware to run a 27B, but I'
More ›