Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
futureshock
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
futureshock
11mo ago
A really great way to get an idea of the relative cost and performance of these models at their various thinking budgets is to look at the ARC-AGI-2 leaderboard. Opus 4.5 stacks up very well here when you compare to Gemini 3’s score and cos
32.
▲
by
futureshock
11mo ago
I really love this piece! I relate to it but it also doesn’t describe me. I’m far more intuitive than this person, though still agree that insights have driven a leveling up of how I relate to others. They were different insights, sure but
33.
▲
by
futureshock
1y ago
I’ll borrow ideas from investing: financial independence, diversification and optionality. If you have enough money you can free yourself from the labor market, but you are still deeply tied to your home country. A second citizenship gives
34.
▲
by
futureshock
1y ago
This so awesome. It reminds me mightily of beat poets like Allen Ginsburg. It’s so totally spooky and it does feel like it has the trapped spark. And it seems to hate us “real ones,” we slickborns. It feels like you could create a cool work
35.
▲
by
futureshock
1y ago
This is a non-story. This was a hardware event. Apple is releasing many new AI features as part of iOS 26 which will launch along side the new iPhones. AI is software. And yet, a number of the features are clearly powered by AI models such
36.
▲
by
futureshock
1y ago
This is working really well in GPT-5! I’ve never seen a prompt change the behavior of Chat quite so much. It’s really excellent at applying logical framework to personal and relationship questions and is so refreshing vs. the constant butt
37.
▲
by
futureshock
1y ago
I think that sounds very reasonable, but unfortunately these models don’t know what they know and don’t. A small model that knew the exact limits of its knowledge would be very powerful.
38.
▲
by
futureshock
1y ago
You can’t vote the climate out of office. Sure our food supplies may crash, but no one person decided they should crash. No one to blame. No one to punish. This is the political reality. This man made catastrophe will feel sufficiently like
39.
▲
by
futureshock
1y ago
It does seem like that’s our new political reality for now. I think that COVID showed world governments just how little control they have over their populations. You get folks to bend a little, but they quickly break and call for you to be
40.
▲
by
futureshock
1y ago
Google’s AlphaProof, which got a silver last year, has been using a neural symbolic approach. This gold from OpenAI was pure LLM. We’ll have to see what Google announces, but the LLM approach is interesting because it will likely generalize
41.
▲
by
futureshock
1y ago
I think the hard solution is to massively increase expectations. Think Star Trek where the grade schoolers are learning quantum mechanics. If everyone has access to the oracle of all human knowledge, then you should teach and test to the ma
42.
▲
by
futureshock
1y ago
I think the fantasy of going back hides the reality that new possibilities are always stretched out ahead. I have lived many lives. New careers, new cities, new countries, new friends, new families. By my count I’ve lived 14 iterations of l
43.
▲
by
futureshock
1y ago
I’d encourage thinking not just about yourself or your immediate circle when it comes to policy like retirement age. State run pension systems affect nearly everyone, not just the extremely fortunate few who can become financially independe
44.
▲
by
futureshock
1y ago
An interview is a sales pitch for a product. The product just happens to be you. Set aside whatever negative feelings you have about this previous job or the people you worked with there. The interviewers care if you will do their job well
45.
▲
by
futureshock
1y ago
There’s increasing evidence that LLMs are more than that. Especially work by Anthropic has been showing how to trace the internal logic of an LLM as it answers a question. They can in fact reason over facts contained in the model, not just
46.
▲
by
futureshock
1y ago
This is a great moment to go read one of my favorite Neal Stephenson books, Diamond Age. Most of the book’s theme deals with nanotechnology, but a very significant element is The Primer, a book that has enough AI to raise a child. The Prime
47.
▲
by
futureshock
2y ago
Strong worker protections can help in all kinds of situations. Case in point, I have a friend in Switzerland. He burned out and had a panic attack at work. The reasons were plausibly personal, but work stress always takes its toll as well.
48.
▲
by
futureshock
2y ago
This is worth a good giggle, as intended, but is just part of the normal space agency PR of doing relatable things like talking about using the toilet and video chatting with elementary students. Because it’s hard to talk about the actual s
49.
▲
by
futureshock
2y ago
I know what you are saying, but I very much disagree. There are also better chess engines. That’s not the point. It’s all about the “G” in AGI. This is a nice demonstration of how LLMs are a generalizable intelligence. It was not designed t
50.
▲
by
futureshock
2y ago
I think in a lot of ways we are already there. Users are clearly already having difficulty seeing which model is better or if new models are improving over old models. People go back to the same gotcha questions and get different answers ba
51.
▲
by
futureshock
2y ago
Wait no, there is actually PLENTY of evidence that performance continues to scale with more compute. The entire point of the o3 announcement and benchmark results of throwing a million bucks of test time compute at ARC-AGI is that the ceili
52.
▲
by
futureshock
2y ago
Actually it means we will potentially get 100x the economic value out of those datacenters. If we get a million digital PHD researchers for the investment then that’s a lot better than 10,000.
53.
▲
by
futureshock
2y ago
It seems an interesting fine-tuning idea. Drawing from reasoning models, I wonder if it’s effective to 10x or 100x the fine-tune dataset by having a larger reasoning model create documentation and reasoning COTs about the code base’s curren
54.
▲
by
futureshock
2y ago
I think you have the right idea. China has yet to truly flex its muscle. They prefer to quietly grow stronger. Their response to Covid with the largely successful zero covid strategy gives a clue about the power of its government. Silly, yo
55.
▲
by
futureshock
2y ago
When they do it, it’s “censorship.” When we do it it’s “safety.” From a technical standpoint it’s the same. Don’t say certain things, respond to certain questions with refusals or with certain answers.
56.
▲
by
futureshock
2y ago
I had the impression that he was ADHD, not autistic.
57.
▲
by
futureshock
2y ago
Someone pointed out on Reddit that DeekSeek v3 is 53x cheaper to inference than Claude Sonnet which it trades blows with in the benchmarks. As we saw with o3, compute cost to hit a certain benchmark score will become an important number now
58.
▲
by
futureshock
2y ago
“Simply brute-forcing” That’s the thing that’s interesting to me though and I had the same first reaction. It’s a very different problem than brute-forcing chess. It has one chance to come to the correct answer. Running through thousands or
59.
▲
by
futureshock
2y ago
That sounds reasonable to me because the compute cost for this level of reasoning performance won’t stay at 10^22 and phones won’t stay at 10^12. This reasoning breakthrough is about 3 months old.
60.
▲
by
futureshock
2y ago
So what? I’m serious. Our current level of progress would have been sci-fi fantasy with the computers we had in 2000. The cost may be astronomical today, but we have proven a method to achieve human performance on tests of reasoning over no
More ›