Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Zababa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
Zababa
6mo ago
I think this holds for basically all movements in which case, I don't really understand the need to flag that website over any website. Edit: though I don't want to diminish that this specific group is a cult with classical cult t
92.
▲
by
Zababa
6mo ago
>the rationalists and all that comes with them. What does that mean?
93.
▲
by
Zababa
6mo ago
>a right-side political figure, who are basically ruling since 2000, (except from 2012-2017 where France had a social-democratic government and president) This is not really true, since 2017 we have a centrist president. For the legal po
94.
▲
by
Zababa
6mo ago
At least when you have a few different values you can pick and compare but yeah.
95.
▲
by
Zababa
6mo ago
This is not true, people complain a lot in France when the trains are late.
96.
▲
by
Zababa
7mo ago
I think it is important to try to find more rigorous things to test than the general sentiment of the people using the tools. If only because the more benchmarks we have the more we can improve models without regressions. METR is asking a r
97.
▲
by
Zababa
7mo ago
From the METR study ( https://metr.org/notes/2026-03-10-many-swe-bench-passing-prs... ): >To study how agent success on benchmark tasks relates to real-world usefulness, we had 4 active maintainers from 3 SWE-bench Ve
98.
▲
by
Zababa
7mo ago
> surely the solution to fast fashion is just to not buy and throw away so many clothes? "just don't do X" has basically never worked, it is not a serious solution to any problem.
99.
▲
by
Zababa
7mo ago
I agree that revealed preferences are stronger signals than stated ones. https://funds.effectivealtruism.org/ shows 52000 donors for $110M, https://www.givingwhatwecan.org/ says more than 10000 donors and m
100.
▲
by
Zababa
7mo ago
That is true but also a bit unfair, they've also been oddly preoccupied with topics like trying to help the most people and frequently promote giving money to efficient charities to fight against malaria, vitamin A deficiencies and hel
101.
▲
by
Zababa
8mo ago
This image comes from running the different versions of the benchmark games programs. Some of the difference between languages may actually be just algorithmic differences, and also those programs are in general not representative of most o
102.
▲
by
Zababa
8mo ago
I have no tolerance for bystanders being killed in general. If the science experiments kill on average less bystanders I'm all for them, if they don't they should be stopped until made safer.
103.
▲
by
Zababa
8mo ago
HathiTrust ( https://en.wikipedia.org/wiki/HathiTrust ) has 6.7 millions of volumes in the public domain, in PDF from what I understand. That would be around a billion pages, if we consider a volume is ~200 pages. 5000 d
104.
▲
by
Zababa
8mo ago
Has the difference between performance in "regular benchmarks" and ARC-AGI been a good predictor of how good models "really are"? Like if a model is great in regular benchmarks and terrible in ARC-AGI, does that tell us
105.
▲
by
Zababa
8mo ago
>Sure, you can have my little assessment at the end if you like, but I work for the students, not for the companies. Most of the students are here because they want to be in the companies, not for the joy of learning.
106.
▲
by
Zababa
8mo ago
>Last semester, professor Pamela Newton, who also teaches the course, allowed students to bring readings either on tablets or in printed form. While laptops felt like a “wall” in class, Newton said, students could use iPads to annotate r
107.
▲
by
Zababa
8mo ago
>It's pretty insidious to think that these AI labs want you become so dependent on them so that once the VC-gravy-train stops they can hike the token price 10x and you'll still pay because you have no other choice. I don't
108.
▲
by
Zababa
8mo ago
>To get to the point of executing a successful training run like that, you have to count every failed experiment and experiment that gets you to the final training run. I get the sentiment, but then, do you count all the other experiment
109.
▲
by
Zababa
8mo ago
>E.g. gemini-3-pro tops the lmarena text chart today at 1488 vs 1346 for gpt-4o-2024-05-13. That's a win rate of 70% (where 50% is equal chance of winning) over 1.5 years. Meanwhile, even the open weights stuff OpenAI gave away last
110.
▲
by
Zababa
9mo ago
Why would you leave the question of whether it's true or not aside? If it's false, isn't it a good thing that not many people are ready to admit something false?
111.
▲
by
Zababa
9mo ago
>My digital thermometer doesn't think. Imbibing LLM's with thought will start leading to some absurd conclusions. What kind of absurd conclusions? And what kind of non absurd conclusions can you make when you follow your let&#x
112.
▲
by
Zababa
9mo ago
Can you give examples of how that "LLM's do not think, understand, reason, reflect, comprehend and they never shall" or that "completely mechanical process" helps you understand better when LLM works and when they d
113.
▲
by
Zababa
9mo ago
>If you have this great resource available to you (an LLM) you better show that you read and checked its output. If there's something in the LLM output you do not understand or check to be true, you better remove it. You could say t
114.
▲
by
Zababa
9mo ago
> Mistakes made by chatbots will be considered more important than honest human mistakes, resulting in the loss of more points. >I thought this was fair. You can use chatbots, but you will be held accountable for it. So you're he
115.
▲
Vibe Kanban
(vibekanban.com)
1 points
by
Zababa
9mo ago
|
0 comments
116.
▲
by
Zababa
9mo ago
I don't think appreciating art separated from the author is solipsistic, in fact I'd argue the opposite. Needing a human presence to engage with art is very human-centric. Or maybe that's due to your definition of art? I can
117.
▲
by
Zababa
9mo ago
>Code is not an asset it's a liability This would imply companies could delete all their code and do better, which doesn't seem true?
118.
▲
by
Zababa
9mo ago
> Rather: an intended part of the ordinary course of using a Sprite. Like git, but for the whole system. What I've been waiting for, for a long time. Basically the thing you need if you want agents to run freely but still in a safe
119.
▲
by
Zababa
9mo ago
Yeah it's far from being as good as a DLSR or mirrorless with a dedicated macro lens. Still, most people reading HN have one in their pocket and it can be a good test to see if you like the idea of macro. It does work with larger insec
120.
▲
by
Zababa
9mo ago
Depends on what level of macro you want, but with modern phones you can get pretty close, usually with the wide angle lens. On iPhones: https://support.apple.com/guide/iphone/take-macro-photos-and... On Pixel: ht
More ›