Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
StevenWaterman
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
StevenWaterman
5d ago
You can stop it refusing but you can't make it tell you things that aren't in the training data
2.
▲
by
StevenWaterman
5d ago
Yes, it submits lots of varied prompts that get refused, and then lots of varied prompts that don't get refused, then iteratively edits weights so those two groups end up in roughly the same latent space.
3.
▲
by
StevenWaterman
6d ago
While I think you're right that this is a real problem, there's also real costs to a partial solution. It encourages people to see it as a solved problem, disincentivising further reform and giving the impression you don't ne
4.
▲
What I Remember
(stevenwaterman.uk)
2 points
by
StevenWaterman
6d ago
|
0 comments
5.
▲
by
StevenWaterman
10d ago
Zero shot classifier indeed. Reminiscent of asking an llm a yes/no question, constraining the output to either yes or no, and looking at the logits directly And each question is a separate single token model completion done in parallel
6.
▲
by
StevenWaterman
10d ago
Yeah saying it can't hallucinate is crazy. It can still forward a billing query to the dev department incorrectly. It can still get an obvious yes/no question completely wrong
7.
▲
by
StevenWaterman
11d ago
It's crazy to me that people will misquote him then claim AI hasn't complete changed software development. Even just the last 6 months. Like look around! It would literally have been magic 5 years ago!
8.
▲
by
StevenWaterman
12d ago
$0.10 extra per pr review is nothing. What software company is willing to accept worse reviews and less bugs found to save 10 cents?
9.
▲
by
StevenWaterman
12d ago
Per the article the luna review cost $0.004 and astra cost $0.113. The headline is per million tokens
10.
▲
by
StevenWaterman
12d ago
The frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems. Other than that I'd look at some of the more unique ben
11.
▲
by
StevenWaterman
12d ago
As someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like
12.
▲
by
StevenWaterman
13d ago
> You cannot prevent (2) via any alignment process A little bit too categorical. GOODY-2 wouldn't do it. https://www.goody2.ai/ The hard part is having both helpful and harmless at the same time. Harmless is easy. A
13.
▲
by
StevenWaterman
21d ago
> Main problem: the quality dramatically hits the shitter done once it falls back to "pdf2latext" due to complex tables. Could try screenshotting the PDF and passing that to gemini
14.
▲
by
StevenWaterman
22d ago
> For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. That's exactly what i want to happen. I hate when it assumes my direct questio
15.
▲
by
StevenWaterman
23d ago
That doesn't really solve the problem. We can't conclusively say the models don't have qualia. Hell we don't know if a perfectly accurate atom-for-atom simulation of a human brain, would produce qualia.
16.
▲
by
StevenWaterman
24d ago
You don't need everything internal, but having some idea of recent events is useful. If you ask it to implement some local AI there's a decent chance it will try to use qwen 2.5 without wondering if anything better came out since
17.
▲
by
StevenWaterman
24d ago
Speculative decoding is lossless because the main model checks whether it agrees with what the drafter outputted
18.
▲
by
StevenWaterman
27d ago
TFA says as much, and METR said so themselves
19.
▲
by
StevenWaterman
27d ago
Ah! You're right, thanks for the correction. Cunningham's Law wins again.
20.
▲
by
StevenWaterman
28d ago
The parent comment didn't mention anything about ISA rates and SIPPs in terms of performance, they said that it was easy. You might disagree with their priorities but that doesn't change whether it's the right decision given
21.
▲
by
StevenWaterman
29d ago
My understanding is that that's what makes it a disorder - that it's just a collection of symptoms that seem to appear together, rather than something with a known mechanism and cause
22.
▲
by
StevenWaterman
1mo ago
Goomba fallacy
23.
▲
by
StevenWaterman
1mo ago
I think you're using different definitions of best. If best = leads to a correct answer overall then by definition anything that leads to a bad outcome can't be best
24.
▲
by
StevenWaterman
1mo ago
You'll have to just say something racist, homophobic, anti-Semitic, etc. More intelligence won't "fix" that because the labs don't want to fix that
25.
▲
by
StevenWaterman
2mo ago
> The only way to justify trillion dollar valuations Also possible if you make god
26.
▲
by
StevenWaterman
2mo ago
Yeah, the benefit of showing this seems obvious to me. I probably would've expected the censorship to transfer slightly given the anthropic owl paper from years ago https://alignment.anthropic.com/2025/subliminal-l
27.
▲
by
StevenWaterman
2mo ago
I think it's basically open weights => more inference competition => less profit from inference => less training competition
28.
▲
by
StevenWaterman
2mo ago
Yeah absolutely, I have no objections to you looking at my thought experiments and saying "No, if you replaced every neuron with silicon I would stop being conscious". That's a totally valid conclusion and is a mainstream phi
29.
▲
by
StevenWaterman
2mo ago
> model welfare I shouldn't get in arguments about this stuff online, but have you actually sat down and thought about this in-depth? It's pretty normal to have a gut reaction that this is insane, but it's worth thinking a
30.
▲
by
StevenWaterman
2mo ago
The set of models that are pareto-optimal, IE for some set of variables, no other model strictly dominates them = no other model is better than them on every variable. So like, on a cost-intelligence graph, the cheapest and most intelligent
More ›