3 ms·
this is pretty cool. i think part of the root cause is current rlhf post training design around confidence and optics rather than cooperative transparent honest
by carterschonwald 3mo ago
this is pretty cool. i think part of the root cause is current rlhf post training design around confidence and optics rather than cooperative transparent honesty. though its kinda an expensive hypothesis to dig into as a private individual
- nullc 3mo agoMost of the models where people are concerned about don't do this when unquantized, so I doubt it's much about the metapolitics imposed in reinforcement training.
- carterschonwald 3mo agoive had doom loops on release day with opus 4.6. quantization aint the culprit ;)