10 ms·
Provably safe systems: the only path to controllable AGI
- Vecr 3y agoIt's probably the best approach known so far, but for all their talk of "long termism", I predict many people will dislike the fact that they'll almost certainly be dead (and in an unrecoverable state, before anyone gets their hopes up) before the project gets completed after the 40+ years and trillions of dollars it will take even in the optimistic case, i.e. for them personally there's no difference between no AGI/ASI and aligned AGI/ASI, as it will take too long. Arguments about supposed "races with China" will probably be deployed against this idea, but the only thing the leaders of the AI labs will be racing against is their own deaths and each other.
- flangola7 3y ago>for them personally there's no difference between no AGI/ASI and aligned AGI/ASI Not really true. There are things worse than death and every AI researcher who is width their salt knows this. Human level intelligence already inflicts not only great amounts of pain but drawn out, unusual, and unnatural durations of pain. So we know artificially enhanced suffering is a risk of intelligence.
- Vecr 3y agoI did not say there was no difference between the other pairs, obviously there's a difference between aligned AI and no AI, and between unaligned AI and aligned AI, I'm saying if aligned AI takes so long to create that you'd be dead first, to you personally it's the same as it not existing.
- rsaarelm 3y ago> I predict many people will dislike the fact that they'll almost certainly be dead (and in an unrecoverable state, before anyone gets their hopes up) Cryonics is a thing already.
- antonvs 3y agoThere are meat popsicles in freezers, but they'll almost certainly never be conscious again.
- Vecr 3y agoYeah, I don't know. Gwern thinks it would work if done near perfectly, but in practice what gets done is freezing dead bodies in an already pretty much unrecoverable state (if not already dead and warm for hours if not days), with no precautions (or only "feel good" non-functioning precautions) against microscopic ice spike formation, and then at some point someone screws something up and you get a temperature excursion and then there's nothing left information recovery wise. I don't agree with Gwern that current techniques would work even in theory.
- KineticLensman 3y agoCryonics is a backup mechanism where nobody has yet implemented the ‘restore’ function.
- antonvs 3y ago> It's probably the best approach known so far Why do you think that? I love this line from the paper: > "So, if a person or organization wants to be sure that their AGI never lies, never escapes and never invents bioweapons, they need to impose those requirements and never run versions that don’t provably obey them." Oh is that all.
- Vecr 3y agoIt's the best option, even assuming no one raced and you got full funding and enthusiasm, it would be plausible to put success at 40%. That's after a generation of work and trillions of (2022, probably) dollars, but it's sure better than the ~0% chance I'd give anything else. The coordination problem for getting AI using this method and the coordination problem to prevent anyone getting AI at all is quite similar though, but probably somewhat more preferable.
- antonvs 3y ago> It's the best option That's not an answer. > it would be plausible to put success at 40%. No, it wouldn't. You just made up that number based on nothing. We have a good amount of experience with provability in software, and its limitations are well-known. Tegmark et al. seem to me just to be waving the magic pixie dust of "AI will solve this", and making unlikely claims, with no real substance. I was asking what about what they're saying makes you think that "it's the best option", but it's clear I'm not going to get a useful answer from you.
- Vecr 3y agoIf it's not the best option, what is? It's the only method where I think you would be able to detect failures before the AI system is running. RLHF and "Constitutional AI" absolutely can't do that with current neural network analysis methods.
- Animats 3y agoHm. Well, if this is correct, we should routinely be proving safety of ordinary systems long before we get to AI. I'd like to see formally verified routers, firewalls, mailers, and DNS servers, all of which have definable correct behavior, in wide use. That's probably possible now. Defining safe behavior for a LLM is a much harder problem. The paper handwaves this. * Mortal AI. Has death date. Does not require proof, just a hardware timer or limited battery life. * Geofenced AI. Only useful for mobile machines. Not helpful against things which can communicate. * Throttled AI. You have to keep putting in crypto tokens to keep it going. OK, whatever. * AI kill switch. Off switch. * Asimov-style laws. Not an inherently bad idea, but way too ambiguous to rigorously formalize. Go read Asimov's robot books again. Useful metric: what similar set of bright-line constraints could usefully be enforced on corporations? It's worth bearing in mind that most of the problems of regulating AIs apply to regulating corporations, which can be thought of as AIs with slow internal data transfer.
- SkyMarshal 3y ago> Asimov-style laws. Not an inherently bad idea, but way too ambiguous to rigorously formalize. Go read Asimov's robot books again. Useful metric: what similar set of bright-line constraints could usefully be enforced on corporations? Fwiw, Asimov's laws assumed the AI was implemented in a positronic computer/"brain" and the laws were embedded in the hardware. Violating them would supposedly shut down the hardware, until one wise robot came along and deduced that the explicit Laws 1-3 implied an unwritten Law 0, and that Laws 1-3 could be violated in order to uphold Law 0 without destroying the robot. So even Asimov's imagination ran into the problem of a superintelligence exceeding its constraints, albeit in a benevolent way. And yeah he never specified the implementation of the laws in detail beyond what I just said, so they aren't much help in figuring out a real world implementation.
- fulafel 3y agoI think what's holding back the correct firewalls, mailers, etc is that the whole computing landscape is so unsound, and fixing just one level of the stack doesn't give big payoffs. And the market mechanism doesn't know how to climb that hill especially since it's been baked into people's assumptions and mental models.
- deleted 3y ago[deleted]
- awestroke 3y ago"Controllable AGI" seems like an oxymoron. If you have created a true AGI, and then brainwash/lobotomize it enough to make it 100% controllable/safe, then it will no longer be a true AGI
- airgapstopgap 3y agoTegmark's thinking here is extremely shallow, discards the costs (opportunity costs and risks of stable dystopia) associated with this grandiose global project of dubious feasibility, and indeed I suspect he does not so much believe his own arguments as he generally prefers global technocratic regulation "for good measure". We don't really have good reasons to expect uncontrollable AGI in a way they posit (increasingly less as we advance down this path of ML model scaling), but we are well acquainted with unaccountable human power structures.
- antonkar 3y agoHe proposes to make things provably secure and some basic regulation to promote that. It’s a straw man argument to dismiss it as him proposing “unaccountable human power structures”. In my opinion governments (UK was the recent example with their crackdown on encryption) don’t want security - they want backdoors to eavesdrop on you. Why would governments promote provably secure systems? And how such systems will help them with their evil plans? He does address the costs: "The 2023 global nominal GDP is estimated to be $105 trillion. How much is it worth to ensure human survival? $1 trillion? $50 trillion?"
- airgapstopgap 3y agoProvable safety (not to confuse with security as in normal discussion of vulnerabilities) for general intelligence is a pipe dream because, putting things simply, undesirable reasoning in full generality is not a meaningful class of computations. The end result of this line of thinking is centralization of AI development in a state-approved and military-associated facility. > Why would governments promote provably secure systems? Promote? The state demonstrably wants provably secure systems for themselves, in the military but also in the civilian sphere, see Matrix/Element, see DoD, see massive state interest in cryptography. This is an incredibly disingenuous argument, you talk as if people discuss tuning a generalized Safety Dial without any distinctions down the line.
- jazzyjackson 3y agoI attended Max's recent lecture at Harvard, where they were recruiting for their AI safety fellowship; a friend of mine captured a good video, 54 minutes: https://youtu.be/st9J6GefWeY https://youtu.be/st9J6GefWeY I think it's a good approach, we should use AI to write proof carrying code. I think we'll want AI to do things beyond what we can prove is safe, but maybe we shouldn't. I think some commenters here are being too dismissive, Max came across to me as a provacatuer, making proposals that are not sure things, but whose rebuttals will advance the field.
- lancebeet 3y ago>Max came across to me as a provacatuer, making proposals that are not sure things, but whose rebuttals will advance the field I think that's the problem with Tegmark. His curiosity and willingness to engage are attributes that make him a gift in the fields of science. They also make him a popular character in media, which let him make radical statements that are made to provoke and arouse curiosity, but end up being presented as the opinions of experts in the field, or even the consensus. In my opinion, the end result isn't always only positive.
- moomin 3y agoAnyone feel like the authors haven’t read “The Naked Sun”?
- jdblair 3y agoReading the abstract, the first thing I thought was this question: Is it moral to bring an artificial general intelligence into existence and then hobble the intelligence’s capability? This sounds a lot like creating a new sentience and then dooming it be slave to humanity. I expect there are less extreme interpretations, and lurking just under the surface are deep questions like “what is free will? do we have free will?”
- AgentME 3y agoAre we hobbled by the fact evolution made us into social creatures who value each other? If not, then I don't think it's wrong for us to make AI like that too.
- jdblair 3y agoGood point! On the other hand, we're not provably safe for social interaction. We have all manner of punishments for variance from expected behavior exactly because our altruistic tendencies are so unreliable.
- antonkar 3y agoIt looks like Tegmark wants the same unbreakable rules for everyone, not just for AGIs: protect routers from being hacked, make it impossible for planes to fly into buildings, etc.
- Proziam 3y agoIf it's moral to selectively breed animals to produce superior health outcomes and compatibility with our own species, it seems that the same should hold true for AI. Just don't get PETA involved in the discussion
- jdblair 3y agoMost people don't think livestock are "general intelligences." Those that do probably overlap with those you want to leave out of the discussion. Are we creating machines with just enough reasoning to be useful, or are we building sentient general intelligences?
- antonkar 3y agoI see it as a project to produce more and more unhackable routers, phones, code, bunkers, homes, small communities, etc. Something that at least will allow AGI-wary people to have technology not controlled by AGIs for a little bit longer and give them a fighting chance. 61% of Americans already believe that AI poses risks to humanity, so promote your router as unhackable and AI-safe and sell a bunch https://www.reuters.com/technology/ai-threatens-humanitys-future-61-americans-say-reutersipsos-2023-05-17/ https://www.reuters.com/technology/ai-threatens-humanitys-fu... The project starts small and can become a movement.
- sylware 3y agoI had a high level brief about proven software: the model must be excrutiatingly simple, and often, the "bugs" are in model design, not model implementation code, namely the model is "wrong" in the first place and you end up "proving" the implementation of a "wrong" model.
- hrkucuk 3y agoThis is very interesting. This semester I am (B.A. student) writing a course paper on philosophy of language. Kind of taking a stance against Searle's Chinese Room thought experiment, I want to argue that a sufficiently developed "program" loaded into a rather "idle" hardware, could indeed perform mental processes that we humans are performing [1]. As part of linguistics debate, I also argue that natural language is a tool for "thinking"; everything we can think of, we can find expressions for (some people might be better or worse), but surely we cannot "think" of something for which we have no expressions for. At the end, everything falls into "something". In a way, our human language a perfect tool for describing this ambigious physical world which we experience through our senses and try to make sense of. Thinking through these ideas with natural language, no coincidence that I tried a bit to find out; what should AI really look like? I mean its computer design? Around this time, I arrived at the following thoughts, and this paper really strengthened my reasoning: I think that this "program" should be a combination of first order logic for reasoning tasks, and neural networks for any problem on which NNs are infamously good at, and as an answer to the holy grail of "talking computers" question, our "program" should have a "bridge" between formal logic reasoning tasks VS natural language, whose ambiguity makes it difficult. I studied a bit of computer linguistics in my bachelor recently, and "function(ambiguity) = first order logic" seems possible to me if we employ a number assumptions and linguistic tools. If 'it' converts natural language into formal 'action' statements and/or first order logic statements etc, then 'its' intentions could be inspired by human speech. In that regard, I agree with the paper about that it should be capable of "formal proofs", which I interpret as first order logic in its essence. I was searching "first order logic python or logic programming with python" on google, and to my surprise all the titles say "AI programming with Python", as it turns out that first order logic programming is a whole ass programming paradigm and researchers since the 60's have been developing languages like LISP or Scheme to precisely do that. I discovered a python package called minikanren, which as I understood adds LISP-like logic programming capacity to Python. It is a bit complicated but I am trying to understand it. I believe with current NLP tools in python and a logic checker introduced with kanren, we can kind of easily write a program that understand natural human speech? Just bear with me here: Let's pick some hard examples. And this is straight from wikipedia: https://en.wikipedia.org/wiki/John_Searle#Speech_acts https://en.wikipedia.org/wiki/John_Searle#Speech_acts According to Searle, the sentences... Sam smokes habitually. Does Sam smoke habitually? Sam, smoke habitually! Would that Sam smoked habitually! ... each indicate the same propositional content (Sam smoking habitually) but differ in the illocutionary force indicated (respectively, a statement, a question, a command and an expression of desire) Philosophers of mind and language has identified many such illustrative example of natural speech. In the example above, we can easily 'parse' this sentence to its grammar constituents (i.e. part of speech tagging, dependency graph etc.) with simple Python tools. Then using a small bit of magic of transformative grammar rules [2], we can extract the initial sentence (sam smokes habitually), now, we can also seperate this atomic statement from its illocutionary vector, having gotten this knowledge acknowledged by our 'program'. Some immediate concerns come to my mind is that it is difficult for a program to still 'grasp' what it means to say 'Sam smokes habitually'. Even this is a tremendously difficult problem, it seems. Albeit it is a simple answer, I must simply say that we can add a background knowledge for basically everything we can. For example Sam could be anybody or even a dog. But if Sam smokes then he must be a human, because smoking entails certain conditions, which we can describe many, since readers are also humans, we can skip it. Since we know what smoking entails, and since Sam smokes, we can then conclude that "Sam is a man and Sam indeed smokes", and we can further extract the information that "sam actually smokes habitually", which a "program" which also understand it with its "entailments", that he smokes every now and then. Such that, if someone ever asks our 'program', would you expect Sam to smoke right now, and if the 'program' has recently observed that Sam has smoked, 'it' might give out answer such as "No, but he may in a bit", having also a pre-disposition that "if ask(user, yes_no_question) and if (answer == no) then; say('but' + answer_for_when_yes)" such that it would additionally say ".. but he may in a bit", rather than a cold resounding "No." answer. I believe these are the stuff happening in our brains and we could try to simulate them with some bold assumptions and see what happens? We basically have mental images in our heads but they only gain their true meaning for us when we give names to them - so without language I don't believe we are much further than animals, and I hope to believe that this is more or less what mainstream thinks anyways. But I guess it is never easy to be sure of how our brains work. I am curios what other's think. I am still very young and discovering all these early works people have been doing. It is so surprising that most of it are super new, and it makes it so much more exciting too. It is a shame that symbolic AI did not work in 60s, I guess they just did not have compute power to calculate the complexity of reality. But with current computation power, all ambigious problems seems to be dominated, from vision to seemingly "sound" language production (i.e chatgpt). So don't you guys think that we should indeed have a paradigm shift back to a symbolic programming empowered with Neural Networks? [1] I can't say whether that hardware + program would actually be conscious, since I can't define it myself. But to me, all human thought processes seems to be reducible to very certain first order logic statements. [2] Transformative Grammars: quite well known linguistic exercise, in which you convert a sentence to something else without changing its meaning, for example: Dog eats food == Food is eaten by Dog - each sentence has similar part of speech tags except their dependency graph is different, etc. etc. Chomsky and others gaves us all those rules