Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Turn_Trout
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
Turn_Trout
7d ago
Yeah it speaks poorly of HN that they upvoted this confident nonsense.
2.
▲
by
Turn_Trout
13d ago
> Good Ventures is a funding partner of Open Philanthropy, who funded Jacob Coxon (the first of the Anthropic employees going viral in the media) via a scholarship. This conspiracy theory is truly crazy. A $20K scholarship in 2022 is sup
3.
▲
by
Turn_Trout
13d ago
I'm one of the whistleblowers (from GDM [1]). I gave up over a million dollars (compared to quietly switching labs and continuing to work at one) to speak frankly about these issues. I hold no equity and tried to zero out my position b
4.
▲
by
Turn_Trout
13d ago
Dario's post [1] commits to direct evaluators that can, among other abilities, expose secret RSI. He wants that made law. Do you have a source on METR employees retaining massive equity stakes? [1] https://darioamodei.com&#x
5.
▲
by
Turn_Trout
18d ago
OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's suffic
6.
▲
by
Turn_Trout
22d ago
See also: https://turntrout.com/self-fulfilling-misalignment (my post) As an aside, does the Waluigi Effect actually exist? My impression is it doesn't.
7.
▲
by
Turn_Trout
2mo ago
Those topics aren't on-topic for the essay. I've taken a pledge to donate at least 10% of my money to charity / impactful giving. A good sum of my donations have targeted high-impact opportunities to improve life for people i
8.
▲
by
Turn_Trout
2mo ago
Thank you for your praise. > I'd guess TurnTrout doesn't agree on that framing, otherwise he probably would not have been at Deep Mind. But clearly he and I agree on other ethical positions; I am nothing but glad to see him sti
9.
▲
by
Turn_Trout
6mo ago
I agree that they called many things remarkably well! That doesn't change the fact that AI 2027 is not a thing which happened, so it isn't valid to point out "this killed us in AI 2027." There are many reasons to want to
10.
▲
by
Turn_Trout
6mo ago
AI 2027 is not a real thing which happened. At best, it is informed speculation.
11.
▲
Automatic Alt Text Generation
(github.com)
1 points
by
Turn_Trout
10mo ago
|
0 comments
12.
▲
An Opinionated Guide to Privacy Despite Authoritarianism
(turntrout.com)
14 points
by
Turn_Trout
11mo ago
|
0 comments
13.
▲
by
Turn_Trout
1y ago
No one has empirically validated the so-called "most forbidden" descriptor. It's a theoretical worry which may or may not be correct. We should run experiments to find out.
14.
▲
English Writes Numbers Backwards
(turntrout.com)
3 points
by
Turn_Trout
1y ago
|
0 comments
15.
▲
by
Turn_Trout
1y ago
As someone who did their PhD in RL and alignment, it was not obvious to me a priori if, or when, or how badly obfuscation would be a problem. Yes, it's been predicted (and was predicted significantly before that Zvi post). But many oth
16.
▲
by
Turn_Trout
1y ago
> The #1 comment says that the rationality community is about "trying to reason about things from first principle", when if fact it is the opposite. Oh? Eliezer Yudkowsky (the most prominent Rationalist) bragged about how he wa
17.
▲
Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models
(turntrout.com)
1 points
by
Turn_Trout
2y ago
|
0 comments
18.
▲
by
Turn_Trout
2y ago
They ran (at least) two control conditions. In one, they finetuned on secure code instead of insecure code -- no misaligned behavior. In the other, they finetuned on the same insecure code, but added a request for insecure code to the train
19.
▲
by
Turn_Trout
3y ago
I'm the author of the GPT-2 work. This is a nice post, thanks for making it more available. :) Li et al[1] and I independently derived this technique last spring, and also someone else independently derived it last fall. Something is i
20.
▲
by
Turn_Trout
5y ago
First author here. Thanks for your comment! > there's a lot hidden in the "if physically possible" part of the quote from the paper: "Average-optimal agents would generally stop us from deactivating them, if physicall
21.
▲
by
Turn_Trout
5y ago
Maybe you should read the paper, and/or the reviewer threads (as we discussed the nomenclature, and eventually agreed that "power" was accurate). We straightforwardly formalize a mainstream definition of power and show how it