Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cryptohell
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Can we bootstrap AI Safety despite being unable to even define it?
(arxiv.org)
2 points
by
cryptohell
11mo ago
|
2 comments
2.
▲
by
cryptohell
11mo ago
Given several models, assuming only that some unknown subset is "safe", can we construct a single model as safe as that subset? This reduces obtaining a trustworthy model to a plausibly easier task.
3.
▲
Optimality of Frequency Moment Estimation
(eccc.weizmann.ac.il)
8 points
by
cryptohell
2y ago
|
1 comments
4.
▲
by
cryptohell
2y ago
The classic AMS (1996) bound for estimating the frequency moments of a stream is shown to be optimal
5.
▲
LLMs can hide arbitrary undetectable information in their responses
(twitter.com)
3 points
by
cryptohell
3y ago
|
3 comments
6.
▲
by
cryptohell
3y ago
ChatGPT could encode usernames, timestamps and other session context into its responses in a way that would only be retrievable by OpenAI and provably invisible to everyone else
7.
▲
LLMs can hide arbitrary information in their responses, undetectably
(arxiv.org)
1 points
by
cryptohell
3y ago
|
0 comments
8.
▲
Undetectable Watermarks for LLMs
(arxiv.org)
3 points
by
cryptohell
3y ago
|
0 comments
9.
▲
Undetectable Watermarks for LLMs
(arxiv.org)
2 points
by
cryptohell
3y ago
|
0 comments
10.
▲
Undetectable Watermarks for Language Models
(arxiv.org)
1 points
by
cryptohell
3y ago
|
1 comments
11.
▲
by
cryptohell
3y ago
How many people do you think actually bother rephrasing LLM outputs? Do you personally ever copy paste a full chunk of a response?
12.
▲
by
cryptohell
3y ago
You draw the first bit to be 0/1 with equal probability, and then the second bit must equal the previous one with probability 1
13.
▲
by
cryptohell
3y ago
True, they claim in this paper this is inevitable
14.
▲
Undetectable Watermarks for Language Models
(eprint.iacr.org)
135 points
by
cryptohell
3y ago
|
69 comments
15.
▲
(Provably) Undetectable Watermarks for LLMs
(eprint.iacr.org)
1 points
by
cryptohell
3y ago
|
0 comments
16.
▲
by
cryptohell
4y ago
(A Turing award winner and a Godel prize winner professors at Berkeley and MIT)
17.
▲
by
cryptohell
4y ago
There are several differences: 1. Empirically, networks have many adversarial examples. It doesn't mean though that there are adversarial examples everywhere . They show that any point can be slightly changed to get whichever output