2 ms·
wtf?? P(doom) relates to https://en.wikipedia.org/wiki/Existential_risk_from_artificial_intelligence https://en.wikipedia.org/wiki/Existential_risk_from_artific
by esafak 13d ago
wtf?? P(doom) relates to https://en.wikipedia.org/wiki/Existential_risk_from_artificial_intelligence https://en.wikipedia.org/wiki/Existential_risk_from_artifici... How would open source models bring about a MAD between humans and AI? By enabling us to wield aligned AI against nonaligned AI? If so, he should say so, and elaborate how open source models further that goal.
- skissane 13d agoI think AI diversity protects against rogue AIs. The more different AIs we have, with different weights, controlled by different actors, the less likely that any one AI will be able to "take over", and the less likely that a significantly large coalition of cooperating AIs will be able to be formed to do it either. With sufficient AI diversity, the other AIs may work to stop the rogue AI from taking over. Two AIs with radically opposed values – e.g. an Iranian-government-values AI and a Chinese-government-values AI – have the incentive to cooperate to prevent a takeover by some other AI with a third competing set of values. Obviously, open weight AI provides much higher AI diversity than closed weight AI does. Open weight AI produces a lot more providers, and a lot more models. Closed AI centralises control in a small number of vendors. > By enabling us to wield aligned AI against nonaligned AI? The risk isn't just "nonaligned AI", it is misaligned AI. I think the "benevolent dictatorship" scenario – AI overrules humans "for their own good" – is the more likely doomsday scenario than AI deciding to kill all humans. And even AI deciding to kill all humans could be more a result of misalignment than complete lack of any alignment, e.g. "to make sure no child is ever abused again, I will make sure no child is ever again born to risk being abused". A valueless AI which does whatever the user says is actually less likely to establish a benevolent dictatorship, or conclude that exterminating humanity would be the most ethical course of action, than one infused with values is. Given that, I'm not convinced that mainstream approaches to "AI safety" actually reduce our existential risk; I worry they actually have the opposite effect.
- solenoid0937 13d ago"AI diversity" does nothing when AIs can attack at incredible speed, when compute imbalances exist, etc
- skissane 13d ago> "AI diversity" does nothing when AIs can attack at incredible speed, when compute imbalances exist, etc They can defend at incredible speed too. Diversity needs to measured in a capacity/capability-weighted way. It isn't just the raw count of models/providers; you need to consider how much compute is allocated to each model/provider, and the diversity at each capability level. I think the safest situation is where the open models are at the same capability level as closed ones. The proposal to slow down the frontier labs isn't necessarily bad from this perspective, if it gives time for the more open providers to catch up – provided it isn't paired with anticompetitive measures to prevent the competition from catching up, which of course it is. However, we may hope that the "slow down the highly closed tier 1 vendors" part of the proposal turns out to be more effective in practice than the "slow down the more open tier 2/3 vendors" aspect of it.
- esafak 13d agoYou are not making an argument for open models, but aligned models. There is no reason an open source model should be aligned. In fact, people often prefer open models because are not aligned, which they view as undesirable fiddling.
- skissane 13d agoIt isn't necessary for the model itself to be aligned–a model which does whatever the system prompt says gets its alignment from the system prompt. If a model is highly disposed to obey its system prompt, then two instances with radically different system prompts will act like two different models, even if the weights are identical. If you have a diversity of actors, with a diversity of ideologies and agendas, all prompting models to implement their own ideology/agenda, then those AI agents won't support an AI takeover if it is done in the name of a competing ideology/agenda, because the agent will see the takeover as an obstacle in the way of achieving its own objectives.
- fragmede 13d agoIsn't the real risk doomsday humans, enabled by AI, committing some truly heinous acts with global reach? Aum Shinrikyo had to figure out how to synthesize sarin nerve gas the old way. Now, a motivated gang of omnicidists with stolen cryptocurrency could come up with something that makes McVeigh's fertilizer bomb in a box truck look like child's play. Just ship shipping containers around the world filled with autonomous drones spraying airborne ebola that they've bioengineered or something. Okay, that's enough DOOOM for me for the week.
- skissane 13d agoThis depends on a lot of things: how many committed omnicidalists there are; how much financial resources they have (untraceable crypto doesn't help you if you're working a dead-end job and only have $5K in your bank account); engineered bioweapons need labs and equipment not just an API key. The probability of your scenario doesn't solely depend on the probability of AI being able and willing to cooperate in it, and it may well be that the non-AI factors outweigh the AI ones in the overall risk of it – which would mean adding AI would be increasing the risk of it less than you think.