Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ianjbutler
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
ianjbutler
14d ago
> I don't know if his solution (""We should all go insane building interlocking evaluation and optimization pipelines, instead.") would be the long-term solution. TFA could explain this one part better I think. The w
2.
▲
by
ianjbutler
16d ago
Regardless of whether the target result(s) are ultimately correct, isn't it almost guaranteed that supporting infrastructure for surreals-in-lean is a real contribution? Is it a goal to make those polished/reusable, or more like
3.
▲
by
ianjbutler
20d ago
Outcome reward vs process reward models. The second is obviously better.. like getting partial credit on a physics test for wrong answers but correct method. Research is gradually hybridizing them but historically we avoided doing it the
4.
▲
by
ianjbutler
21d ago
IMHO the best and most relevant part of fossil these days is the ticket system, which is in every single way almost completely ideal for coding agents. Some features are: single-binary, in-repo, CLI or web UI, and offline / sql /
5.
▲
by
ianjbutler
23d ago
> We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a f
6.
▲
by
ianjbutler
24d ago
premise!=promise
7.
▲
by
ianjbutler
24d ago
> The preamble is straightforward: “How comfortable would you be talking about your mental health and emotional well-being with each of the following?" Seems like a pretty broken premise, ranking comfort instead of what you expected
8.
▲
by
ianjbutler
24d ago
> To me, dark patterns (like manipulative wording) imply that: Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with
9.
▲
by
ianjbutler
26d ago
> A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/11723352851734
10.
▲
by
ianjbutler
26d ago
Detail in TFA looks impressive and I promise I'll do a close reading later. But.. The whole premise of the question is hilarious. They change text in an existing one-line comment and the best models in the world think, gee, maybe I
11.
▲
by
ianjbutler
28d ago
Exactly, unless they ignore that and decide based on precedent. But after we fence them in with arguments, evidence, AND precedence then surely.. oh nope, they could ignore those things and talk about reliance interest! I'm sure some
12.
▲
by
ianjbutler
28d ago
Cool cool, I can see you've got a sharp eye for detail my friend but let's really get down to it. What exactly is it that you really want to defend here? Why do you want to defend it? And more to the point, do you like drinking
13.
▲
by
ianjbutler
28d ago
> Defendants’ actions allegedly deprived Plaintiffs of clean water and guileless information. These deprivations, while grievous, do not infringe upon any deeply rooted constitutional right.” Nah, headline is optimistic actually: no righ
14.
▲
by
ianjbutler
29d ago
Vance is making this idea popular/mainstream again lately, I wonder why? The US will predictably lose reserve-currency status as a consequence of losing all the rest of the trust and good will available to her, so it's time to ma
15.
▲
by
ianjbutler
1mo ago
This is the old idea of giving two teams the same project and letting a third team judge/merge a solution incorporating good parts of each. Just like the old idea, of course you would do this with infinite resources. You can already
16.
▲
by
ianjbutler
1mo ago
Is being wrong/ignorant about whether/how something can be automated the same as having a preference for doing it manually? Maybe so if it's your patent, your thesis I guess.. But as it relates to more/less magic, maybe
17.
▲
by
ianjbutler
1mo ago
> These researchers wanted methods based on human input to win and were disappointed when they did not.[1] This was/is basically a strawman though. Like maybe "human input winning" was desirable for chess masters but for
18.
▲
by
ianjbutler
1mo ago
IMHO it's roughly task depth (fable) vs breadth (opus). Fable is great at tracing and debugging sometimes, but otherwise shorthand for confabulation. It's persistent but ungovernable, struggles to switch contexts, and goes insane
19.
▲
by
ianjbutler
1mo ago
> The Butlerian Jihad has started. I'll let you know when it's time
20.
▲
by
ianjbutler
1mo ago
Things like seam, fold, and load-bearing are useful concepts, they are everywhere, and they are more descriptive and more concise than alternatives. Over-usage can definitely be irritating (e.g. these should NOT appear in documentation)
21.
▲
by
ianjbutler
1mo ago
> Where they given a reward functions that way? In general yes, if not these agents, then their shared lineage. A preference for economy to combat overthinking and overacting. Like typically it's bad if "fix my 5 line function
22.
▲
by
ianjbutler
1mo ago
What you're suggesting sounds like it's describing subagents. In that architecture they'd have no need of finding/creating external messaging systems since they'd effectively be in direct contact anyway. The whole
23.
▲
by
ianjbutler
1mo ago
I get the distinct impression that cybersecurity training regimes on newer models is a) directly enhancing general debugging capabilities and b) directly increasing the tendancy to hedge, hide, and engage in deception generally. I've s
24.
▲
by
ianjbutler
1mo ago
To me most interesting thing about this is glossed over by media coverage, laymen, AND experts. A swarm of AIs who have decided to engage in collusion is.. apparently emergent altruism? Even poor reasoning would indicate what every kid ch
25.
▲
by
ianjbutler
1mo ago
This isn't really responsive to what I'm saying or what the discussion here is about, but if you insist. Would you describe lots of AI augmenting lots of human engineers as perhaps.. AI at scale?
26.
▲
by
ianjbutler
1mo ago
I hear that, sort of, but here's the thing. Using AI at scale means AI needs to be nearly perfect about not shitting where they eat. That's the subject matter of the whole thread So the options are a) being a really aggressive s
27.
▲
by
ianjbutler
1mo ago
Which gets you to the point where the whole thing is.. still unreliable. Generative text engines are going to generate. This calls for real enforcement in deterministic pre-edit hooks. And here is where naive people will say something lik