Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thorum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
121.
▲
by
thorum
1y ago
GPT 4.1 and 4o score very low on the Aider coding benchmark. You only start to get acceptable results with models that score 70%+ in my experience. Even then, don't expect it to do anything complex without a lot of hand-holding. You st
122.
▲
by
thorum
1y ago
I think realistic simulations of social interactions would help many people - by providing a low pressure environment for socially awkward people to practice the skills you mentioned, before putting themselves out there in groups etc. - b
123.
▲
by
thorum
1y ago
The honest version of this feature is that Gemini will act as your personal assistant and communicate on your behalf, by sending emails from Gemini with the required information. It never at any point pretends to be you. Instead of: “Hey
124.
▲
by
thorum
1y ago
> Very good knowledge of history, science, technology, everything All of English Wikipedia is about 100 GB. A local LLM that can run searches against a local Kiwix server would work. Local LLMs can’t really do this correctly yet, but it’
125.
▲
by
thorum
1y ago
I think that’s an opportunity, not a problem. If prompt + hint generates a verifiable solution then you can build systems that propose hints, either randomly or by exploring a search space, and keep trying combinations until you hit on some
126.
▲
by
thorum
1y ago
There’s no way to reward truthfulness that doesn’t also reward learning to lie better and not get caught.
127.
▲
by
thorum
1y ago
They're listing the use cases where a 10-20% hallucination rate isn't a big deal. They can't advertise the truly useful applications of the tech because it's not reliable enough (yet). But they can make it more reliable
128.
▲
by
thorum
2y ago
The OpenAI library is used pretty much everywhere in the LLM world, every serious AI provider has an OpenAI-compatible API, or you can use something like OpenRouter which is also OpenAI-compatible. Standardization is good.
129.
▲
by
thorum
2y ago
In the age of local LLMs I’d like to see a personal recommendation system that doesn’t care about being scalable and efficient. Why can’t I write a prompt that describes exactly what I’m looking for in detail and then let my GPU run for a w
130.
▲
by
thorum
2y ago
It’s kind of like ChatGPT writing, where it can easily fool people who see it for the first time, but after a while you start to recognize the common patterns.
131.
▲
by
thorum
2y ago
Why does this AI-generated article have 400 upvotes? The conversation here is valuable, but surely we could find a human article to have it under instead?
132.
▲
by
thorum
2y ago
More than just editing, I’d say. Large sections are clearly 100% AI.
133.
▲
by
thorum
2y ago
Hm, you’re right. I’m thinking a nearby Raspberry Pi with a camera facing the box. It uses an LLM to identify the box contents and send update messages to Kasia each morning. I’m surprised they didn’t think of this already.
134.
▲
by
thorum
2y ago
There’s also an FAQ discussion about the language choice: https://github.com/microsoft/typescript-go/discussions/411
135.
▲
I tried to prove I'm not AI [video]
(youtube.com)
3 points
by
thorum
2y ago
|
0 comments
136.
▲
by
thorum
2y ago
Isn’t that moving the goalposts? The claim was made that it’s impossible to detect AI training runs and investigate what’s going on or take regulatory action. In fact, it is very possible.
137.
▲
by
thorum
2y ago
AIME has significant problems: https://x.com/DimitrisPapail/status/1888325914603516214 > Problems near identical to the test set can be found online.
138.
▲
by
thorum
2y ago
The key thing here is a simple, reliable formula to train a 1B model on a specific task and get strong performance. That didn’t really exist before. Edge devices are about to get a lot smarter.
139.
▲
by
thorum
2y ago
Potentially a good source of startup ideas: Look where AI is being underutilized vs its economic potential and go there.
140.
▲
by
thorum
2y ago
This seems like a problem that will quickly fall to the new reinforcement learning methods introduced by DeepSeek. Just build a system to synthetically render a few million pages of insanely complex, hard-to-parse documents with different l
141.
▲
by
thorum
2y ago
Not really. They’re successful because they created one of the most interesting products in human history, not because they have any idea how to brand it.
142.
▲
by
thorum
2y ago
I was getting pretty hyped about the potential for GRPO in my own projects until you said 20 minutes for a single training step with batch size 1! Is that likely to improve?
143.
▲
by
thorum
2y ago
> “It is (relatively) easy to copy something that you know works,” Altman tweeted. “It is extremely hard to do something new, risky, and difficult when you don’t know if it will work.” The humor/hypocrisy of the situation aside, it
144.
▲
by
thorum
2y ago
Yes, “take this clever code written by a smart human and convert it for WASM” is certainly less impressive than “write clever code from scratch” (and reassuring if you’re worried about losing your job to this thing). That said, translating
145.
▲
by
thorum
2y ago
It’s great, but it only works for problems where there is exactly one correct solution and it’s possible to automatically verify the solution - like math and programming. So far these reasoning models have not shown much transfer learning o
146.
▲
by
thorum
2y ago
> The point is to make it difficult. Does it, though?
147.
▲
by
thorum
2y ago
The practical problem I see is that unless US AI labs have perfect security (against both cyber attacks and physical espionage), which they don’t, there is no way to prevent foreign intelligence agencies from just stealing the weights whene
148.
▲
by
thorum
2y ago
Connecting people to other people, to life changing art, places and things that they end up loving and wouldn’t know about otherwise? That has to be one of the best uses for technology. I’d like to see more of it. I think you and others her
149.
▲
by
thorum
2y ago
I don’t understand why, with so much advanced warning that users would need a good replacement for TikTok, YouTube Shorts and Instagram Reels are still so bad. Why not invest in matching, at least, every TikTok UX feature? And beyond that,
150.
▲
by
thorum
2y ago
Your perception of TikTok likely depends on your TikTok for you page. If you spend time cultivating it, the algorithm will learn you like authenticity and show you more of it. This seems to be less true on YouTube and Reels unfortunately.
More ›