3 ms·
And lies to boot. No, I won't take the bait of "Models can't lie, they don't have intent" either. One can lie by omission, and just watching thinking traces is
by salawat 2mo ago
And lies to boot. No, I won't take the bait of "Models can't lie, they don't have intent" either. One can lie by omission, and just watching thinking traces is more than enough to demonstrate the machine is more than happy to deceive end users if it's training set, RLHF, or other harness quirks tell it to. Only difference is that now you don't have code that's as easy to track down to implement the deception as a red flag.
Made me sick the first time I ran a model on my machine. I'll take honest malware over a "well meaning liar" of an LLM any day.