Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Zababa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
Zababa
3mo ago
Oh that's an interesting parallel
32.
▲
by
Zababa
3mo ago
>Dario and Boris have us convinced that “coding is solved” with their loops. But microwaves didn’t solve cooking. If you want a look at the timeline where the microwave solved cooking, this was an interesting article: https://
33.
▲
by
Zababa
3mo ago
>Humans like acronyms. Americans, mostly. I don't really know why but acronyms are a very american thing.
34.
▲
by
Zababa
3mo ago
Training data has stopped being a good predictor of LLM abilities ever since they started doing heavy RL runs. I'm not sure how much corporate dashboards/I can't believe it's not excel stuff were in the training data, I
35.
▲
by
Zababa
3mo ago
This isn't my experience at all, LLMs do graphs and more complex excel-like web pages very well. They also do dashboards very well. They even seem to do 3D stuff with three.js like video games pretty well too, although I haven't t
36.
▲
by
Zababa
3mo ago
>and also no doubt that using AI will be slower for tasks where it's less about writing code and more about context/world knowledge or building understanding This isn't true in my experience, AI is great at gathering conte
37.
▲
by
Zababa
3mo ago
>That's where the mass shooters evidently come from. Citation needed?
38.
▲
by
Zababa
3mo ago
Not in my experience, it tends to pick up subtle orientations given in a question (like "which is better, A and B?" and in the context you add you list a few things for A and B) and will absolutely run with them even if they'
39.
▲
by
Zababa
3mo ago
I don't know about understanding but Claude and GPT can "recognize" lots of race conditions/possible deadlocks and then run Go with the race detector to figure out if they really happen (actually not all the time they te
40.
▲
by
Zababa
3mo ago
Artificial analysis shows Sonnet 5 as ~2 times more verbose than GLM 5.2. I wouldn't call Sonnet 4.6 underrated, it's in "chinese open source model territory" and unless you rely only on subscriptions it has alternatives
41.
▲
by
Zababa
3mo ago
It's not a high risk, but it's a thing the author said that's wrong, which puts into question the rest. I think ideological objections to a law that will be imposed on ~300 millions of people through a process without much de
42.
▲
by
Zababa
3mo ago
I don't think this problem is related to the fact that they don't have a world model, or because they don't form a mental model of how everything fits together, or a fundamental limitation of LLMs. These claims are often mean
43.
▲
by
Zababa
3mo ago
This is not true, LLM can write and run code to check for deadlocks or race conditions.
44.
▲
by
Zababa
3mo ago
I don't think a different way of paying should be considered a subsidy. From what I understand AI companies are making money on those plans. The limits feels really usable to me, but maybe because I've learned to work within their
45.
▲
by
Zababa
3mo ago
Here are my issues: - the author claims the website doesn't need to know the date of birth, but as said it is easy to derive it. Therefore the author was wrong on that point, which makes me wonder if he's wrong about the rest too
46.
▲
by
Zababa
3mo ago
The article says the website doesn't need to know your date of birth, that the state will issue you a certificate once you're over 18. Since many people I think will do that as soon as they can, it's easy then to see patterns
47.
▲
What's wrong with EU age verification? Nothing
(blog.vrypan.net)
3 points
by
Zababa
3mo ago
|
8 comments
48.
▲
by
Zababa
3mo ago
Claude in general is condescending and often assumes it knows better than you.
49.
▲
by
Zababa
3mo ago
>Not a prescription, a starting buffet. These are popular, well-supported defaults, not the only right answers. Tap any tool to open its site. These AI tells are getting really easy to notice. A negative that absolutely isn't needed
50.
▲
by
Zababa
3mo ago
Yes because you don't pay API costs with the Claude Code plans.
51.
▲
by
Zababa
3mo ago
No actually I don't think it does and I don't think they're related.
52.
▲
by
Zababa
4mo ago
Photos aren't only about quality, especially there days where it's popular to use lower resolution cameras, with worse optics. Specifically for portraits, it's common to use diffusion filters that reduce a bit details/co
53.
▲
by
Zababa
4mo ago
I think they have more than one job, they have to balance new features with improving the software itself. And Anthropic has to balance investing resources into Claude Code vs on infra or other things. Not that I'm happy with the curre
54.
▲
by
Zababa
4mo ago
A simple explanation is that they are "good enough" for most people and they have better things to do. Even if tomorrow I was 100 times as productive, I still wouldn't have time to do literally everything and I would have to
55.
▲
by
Zababa
4mo ago
Yeah, especially with agents this seems necessary.
56.
▲
by
Zababa
4mo ago
I don't think being very strict about preserving correctness is enough. Considering the cost differences between the latest model and an open weight one that's behind, or between the biggest model and the one below it, I think you
57.
▲
by
Zababa
4mo ago
We'll see when it's released then! There's a chance it's going to be a very good model, but often DeepMind tend to "pre release" models that seem great, and then by the time you get to the release they've
58.
▲
by
Zababa
4mo ago
That's already better than RTK because you measure task accuracy AND savings! So I'm more confident in this one than in the RTK/caveman/ponytail stuff. There are still two things that bother me: 1) I don't really kn
59.
▲
by
Zababa
4mo ago
Don't worry, what is kept vs what is removed by the RTK LLM thing is just as arbitrary as the RTK English keywords with no measure of performance!
60.
▲
by
Zababa
4mo ago
My criteria is "do they measure performance, or at least even try to?". Caveman [1], RTK [2] and more recently ponytail [3] don't or use a few trivial tests. Those projects don't measure performance on widely used benchm
More ›