3 ms·
Couldn't we improve LLM training by giving them known impossible tasks and rewarding them for saying "this is not possible" or "I don't know"? Clear and well-de
by sdeframond 18d ago
Couldn't we improve LLM training by giving them known impossible tasks and rewarding them for saying "this is not possible" or "I don't know"? Clear and well-defined expectation, not like "ethic".
I am surprised this is not already the case.
Edit: or even better "this is not possible because X"
- joelthelion 18d agoI've been wondering about this for a while. Maybe it doesn't work? Or maybe frontier labs just prioritize benchmark scores in spite of all their safety talk.
- sdeframond 17d ago"If we dont destroy the world, others will. So wed rather be the ones to do it (and profit from it)."
- arw0n 18d agoIt is diametrically opposed to the other training goals of persistence and goal-focus. We should invest more in this, it could also improve tas K accuracy, but so far it seems the payoff isn't worth it in terms of quality (although it might be in terms of security)
- sdeframond 17d agoI'm not sure I want persistence if it means that I get paperclip'd
- Mentlo 17d agoAnd now you understand why the totality of the AI safety community wants to pause!
- abm53 18d agoI’ve seen people recommend writing “failure is an option” into AGENTS.md as a non-training based crutch.
- sdeframond 18d agoThat looks reasonable. Instead of asking "do this", maybe we should prompt "is this possible ?"
- itsgrimetime 17d agoyeah but knowing if something is impossible or not is pretty hard to determine, right? The line between “takes weeks of trial and error and lots of out-of-the-box thinking” is indistinguishable from “literally impossible” until it’s been done. And they’re trying to get these models to do things that people haven’t been able to do. some would say these are/were “impossible”.
- sdeframond 17d agoI can assure you making up impossible tasks is possible.
- Miraltar 17d agoYou'd train them to give up on hard tasks which is the opposite of what these labs want
- hellerpop 17d agoI think BullshitBench (v2) does exactly this for different fields. An expert in these fields would expose the questions as bullshit but most of the models don't.