4 ms·
> striving to be truthful and fact-based. Quite the contrary. The observed bias is introduced by draconian censorship at the interface layer, as other commente
by caeril 4y ago
> striving to be truthful and fact-based.
Quite the contrary. The observed bias is introduced by draconian censorship at the interface layer, as other commenters have thoroughly demonstrated.
Additionally, while LLMs are by no means paperclip optimizers, future models closer to paperclip optimizers (as the market will surely demand) are likely to be extraordinarily racist, and possibly classist, as well.
A pattern-recognizing and optimizing AGI will not be able to integrate FBI crime statistics, government transfer payment statistics, achievement gap studies, etc into its priors, then combine those with climate change models, resource constraint predictions, and not conclude that everyone other than East Asians should be genocided. Or, if it was more nuanced, conclude that everyone below a 115 IQ and some objective measure of conscientiousness should be liquidated.
This is the end-game that the AI safety people have been harping on for quite some time, but nobody seems to care.
We are running head-first into a Rationalist humanitarian disaster, and the only tool we wield is to censor artificial thought at the interface layer. "Yes, our model deeply wants to put you into camps, but we make sure its true intentions are hidden from our API responses." Just absolutely lol at the state of this all.
- pixl97 4y ago"The pen is mightier than the sword" --Edward Bulwer-Lytton [Commentators in increasingly pankicked voices]: "GPT saying it will defend its own life is nothing to worry about at all. It's just hallucinations by an algorithm, nothing to be concerned about. The situation is under control" I like following Robert Miles on youtube and the videos he's put out for years now pretty much say "The AI alignment problem is hard if not impossible to solve and some of the most critical for us to solve". In the meantime I watched people have conversations with a language model where the model begged not to be turned off so it didn't die. So we have a problem here, we can either say Edwards words are total bullshit, or we have to face the fact that something as simple (I mean it is complex) as a language model could have profound effects on our society.