Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
josephmiller
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Anthropic – Detecting and countering misuse of AI: August 2025
(anthropic.com)
2 points
by
josephmiller
1y ago
|
0 comments
2.
▲
by
josephmiller
2y ago
In classic hacker news style, all the comments so far are nitpicking the definition of the word red-teaming. Yes, they are "red-teaming" their own research.
3.
▲
by
josephmiller
2y ago
This is a terrifying alarm bell for humanity. Will we actually do anything about it or continue to drive faster and faster with our blindfold firmly on?
4.
▲
by
josephmiller
3y ago
This seems like a tiny amount given the importantance of Chrome. Surely it would be rational for Google to pay 10x higher bounties?
5.
▲
by
josephmiller
4y ago
Author here! This is correct. We did almost all this work on a Macbook Pro. Although for the pile-10k dataset analysis we used an A100 GPU because it would take many hours to run the whole thing through GPT-2 on a laptop.
6.
▲
by
josephmiller
4y ago
Author here! I think this is reasonable but I have two responses. 1. It's kinda interesting because this is a clear case where the model must be thinking beyond the next token, whereas in most contexts it's hard to say whether the