4 ms·
People do it all the time. People wrote the training material too. The problem of alignment, even semantic correctness, is hairy enough in the real world, and
by mikrl 2y ago
People do it all the time. People wrote the training material too.
The problem of alignment, even semantic correctness, is hairy enough in the real world, and it’s unclear whether forcing it on an AI matrix multiplying system should be easier or harder.
I suspect easier to set up and execute, but harder in that it’s also a stress test that will really hit corner cases like this.
- ahazred8ta 2y agoPeople pretend not to be evil. Their writings reflect this. AI is being trained on these writings. What could go wrong?
- mikrl 2y agoWhere do most people get their alignments? Outside of the house: school, religion, the street. In those places, the written word is a means to an end and all the aligning happens during human behavioural conditioning. Text may drive it (like a holy book) but it’s a tool to support the human level structure. For an LLM the written word is all it knows… well, pieces of words turned into numbers.
- Jerrrry 2y agoFirst commandment lookin real pertinent right now.
- moffkalast 2y agoIf a person is pretending well enough, it appears genuine. Why would a model learn something it has no information about? It would just learn to be nice, which largely seems to be the case for base models. It's part of the "Can of seltzer" problem as I've heard it called. I can describe and write down what I'm seeing on my desk, where there is a can of seltzer I open and drink. A model trained on this description has a fundamental disconnect with reality, because continuing that text requires you to know more information than is contained in the text. It doesn't learn anything about the world, it learns that sometimes in a text there is a can of seltzer, and you can list random things when describing a desk. It's probably why LLMs make so much shit up, because they have no real point of reference and their world view is composed of layers of things that appear nonsensical and self contradictory because they are trained on piles of writings that are hopelessly out of context. The fact that they are even coherent and semi-reliable is a downright miracle.
- mikrl 2y ago>If a person is pretending well enough, it appears genuine There are multiple readings to most statements. “I loved the restaurant we went to” Could be parsed in at least 3 ways without any additional context. Sincere, sarcastic implying the restaurant was bad, indignant implying another restaurant was worse. Without an alignment goal like “always assume sincerity” which itself can backfire, how can you control what an LLM generates? There is no universally derivable law saying any particular interpretation is the right one. This doesn’t even begin to touch on how there may be a signal too weak for humans to perceive, but an LLM could focus on, leading to many other wild and wonderful interpretations mined from the data.
- moffkalast 2y agoWell there is the only law that goes with deep learning, the law of statistics. The interpretation that has the highest number of occurrences will most likely be preferred. If you take the standard dataset, i.e. the internet, I would suppose it's actually not that morally bad on average because the sites it's sourced from are largely moderated, and the ones that aren't tend to be thrown out. So there would have to be an inherent lawful, positive bias due to the general lack of the opposite in the ground truth examples provided.
- beefnugs 2y agoThere is something hilarious but also disturbing about how the people doing all this just can't help themselves from dicking with it in weird ways before its even working properly or reliably: Make sure we dont release this to the public until it generates historical figures at equal race rates. It is almost like the mandates from above are saying "we will only fund this nonsense if you guarantee the number one feature is hidden censorship and control. Prove it with your little DEI stuff, but it must be reconfigurable at the underlying level at any time." This isn't how proper engineering has ever been done before, there should be basic reliable functionality before going into all this censorship and control stuff