Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gallerdude
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
gallerdude
8mo ago
The weirdest thing about this AI revolution is how smooth and continuous it is. If you look closely at differences between 4.6 and 4.5, it’s hard to see the subtle details. A year ago today, Sonnet 3.5 (new), was the newest model. A week la
32.
▲
by
gallerdude
8mo ago
I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for con
33.
▲
by
gallerdude
8mo ago
Both Opus 4.6 and GPT-5.3 one shot a Gameboy emulator for me. Guess I need a better benchmark.
34.
▲
by
gallerdude
8mo ago
Both Opus 4.6 and GPT-5.3 one shot a Gameboy emulator for me. Guess I need a better benchmark.
35.
▲
by
gallerdude
8mo ago
I'm not sure I agree. LLM's have the feel of an alien new technology, and especially did back then. In retrospect, it feels very obvious that small models don't pose much of a threat, but that's only in retrospect.
36.
▲
by
gallerdude
9mo ago
> the rich are now permitted to ignore copyrights, while the poor remain constrained by them, as before. Claude Code is $20 a month, and I get a lot of usage out of it. I don't see how cutting edge AI tools are only for the rich. T
37.
▲
by
gallerdude
9mo ago
I'm sure Apple is more than happy to pay the premium for cleanness.
38.
▲
by
gallerdude
9mo ago
Today I showed Claude Code how to control my lights, and I'm having a blast.
39.
▲
by
gallerdude
11mo ago
It is interesting that most of our modes of interaction with AI is still just textboxes. The only big UX change in that the last three years has been the introduction of the Claude Code / OpenAI Codex tools. They feel amazing to use, l
40.
▲
by
gallerdude
1y ago
I wonder if they're using reasoning? It usually eliminates these types of errors
41.
▲
by
gallerdude
1y ago
> OpenAI researcher Noam Brown on hallucination with the new IMO reasoning model: > Mathematicians used to comb through model solutions because earlier systems would quietly flip an inequality or tuck in a wrong step, creating halluci
42.
▲
by
gallerdude
1y ago
Local models you can get just the pretrained versions of, no RLHF. IIRC both Llama and Gemma make them available.
43.
▲
by
gallerdude
1y ago
I can see both sides of it. There’s a fancy bread bakery by where I live. I go infrequently, the bread is great. But it’s expensive, most of the I just want a cheap loaf from Target, as do most people. Instead of broad employment of artisan
44.
▲
by
gallerdude
1y ago
I made a very mediocre platformer in my senior year of high school, published on itch.io. I ended up becoming a software developer, which I enjoy 80% as much, but without any burnout or worrying about the superstar economics of being a game
45.
▲
by
gallerdude
1y ago
When I was near the end of high school, my family visited London, and I was thinking about being a game dev. So I sent Terry Cavanagh an email, and to my surprise he completely agreed to get lunch. He was extremely kind, gave me a lot of in
46.
▲
by
gallerdude
1y ago
classic human hallucination
47.
▲
by
gallerdude
1y ago
He's awesome. I listened to Lex Friedman for a long time, and there was a lot of critiques of him (Lex) as an interviewer, but since the guests were amazing, I never really cared. But after listening to Dwarkesh, my eyes are opened (or
48.
▲
by
gallerdude
1y ago
For coding, I like the Aider polyglot benchmark, since it covers multiple programming languages. Gemini 2.5 Pro got 72.9% o3 high gets 81.3%, o4-mini high gets 68.9%
49.
▲
by
gallerdude
2y ago
It's funny, I see myself as basically just a pretty unabashed AI believer, but when I look at your predictions, I don't really have any core disagreements. I know you as like the #1 AI skeptic (no offense), but like when I see poi
50.
▲
by
gallerdude
2y ago
I had a very hard time working there, maybe the worst time in my life. I worked with a lot of very smart people, but something about the company culture is doomed in a way I haven't seen before. Last year I read the book Julia by San
51.
▲
by
gallerdude
2y ago
Which ones? In my experience, a lot of Apples products have incredible longevity. Notes, Calendar, Pages all just get better and better.
52.
▲
by
gallerdude
2y ago
Is there really no standalone app, like ChatGPT/Claude/DeepSeek, available yet for Gemini?
53.
▲
by
gallerdude
2y ago
Completely disagree… there are a crazy amount of cases that didn’t work, until the models scaled to a point they magically did. Best example I can think of is the ARC AGI benchmark. It was seen to measure human-like intelligence through spe
54.
▲
by
gallerdude
2y ago
Humans always hallucinate like this, seeing the original problem instead of the twist.
55.
▲
AI Art Turing Test Results
(substack.com)
1 points
by
gallerdude
2y ago
|
0 comments
56.
▲
by
gallerdude
2y ago
With some transcribing (using another LLM instance) I’ve even gotten it to solve NYT mini crosswords.
57.
▲
by
gallerdude
2y ago
4o-mini: 16% 4o: 50% o1-mini: 97% o1: 100% * disclaimer - only n=7 on o1. Others are like 100-300 each
58.
▲
by
gallerdude
2y ago
Very interesting - have you tried using `o1` yet? I made a program which makes LLM's complete WORDLE puzzles, and the difference between `4o` and `o1` is absolutely astonishing.
59.
▲
by
gallerdude
2y ago
New technologies can solve problems better. "I'm starting to think computers are a solution in the need of a problem. Have we not already solved doing math?"
60.
▲
by
gallerdude
2y ago
Sometimes I wonder if in 100 years, it's going to be surprising to people that computers had a use before AI...
More ›