4 ms·
Not really a characterization of this work (I like mechanistic interpretability even though I consider it a detour in the long run), but of the submission title
by airgapstopgap 3y ago
Not really a characterization of this work (I like mechanistic interpretability even though I consider it a detour in the long run), but of the submission title (edit: it was "Tiny Transformer trained for addition learns bizarre addition algorithm" at the point of this comment being posted). It's bad. I suspect it's informed by Rob Miles' characterization of this algorithm as "bizarre and inhuman"[1]
AI safety/alignment discourse as a whole is so incredibly bad. It's autodidactic and weird in the bad sense, weird like people who think they are brave iconoclasts but just haven't the curiosity and humility to learn the basics of the field are weird. Add some money from Dustin Moskovitz to Open Philanthropy to Robert Miles' propaganda firehose on youtube, add policy lobbying and buying bankrupt crypto podcasters, and you have the worst that could come from the culture of public intellectualism. The notions they have inserted into this potentially vital field, like inner/outer misalignment and "mesaoptimization" [2] (where simple overfitting is the real and more productive explanation) are deeply misleading, technically illiterate and stoking public fears in directions we should care about the least, siphoning resources and attention from real issues informed by specifics of contemporary ML (chiefly data and curriculum engineering). They're close to putting a lid on LLM development – because they're getting spooked by some vague sense of growing "capabilities"; even though LLMs are evidently a vastly better approach to safe-yet-capable-general-AI than anything any of their thought leaders have ever come up with in decades of SIAI/MIRI/Lesswrong procrastination. (For one example of a decent takedown of their approach I recommend [3]).
Why is this algorithm bizarre, or even inhuman? Could anyone show me the specific functional connectivity graph and activation functions in my brain that I use when doing addition, with the circuit going through spatial-quantitative cortex and phonological loop, populations querying cached results in long-term memory, individual numbers marked as interesting or banal by the highest-level associative units? Would it look anything like a neat and sensible algorithm we'd come up with using a global top-down representation of the solution target space?
By this standard, humans are inhuman, humans are shoggoths, as they should be, because emergent data-driven structures in general-purpose substrate generally do not resolve to anything like provable optimality, even if they get very close in performance. This framing adds nothing but clicks and insecurity.
1. https://twitter.com/robertskmiles/status/1663534255249453056 https://twitter.com/robertskmiles/status/1663534255249453056
2. https://www.youtube.com/watch?v=zkbPdEHEyEI https://www.youtube.com/watch?v=zkbPdEHEyEI
3. https://www.lesswrong.com/posts/wAczufCpMdaamF9fy/my-objections-to-we-re-all-gonna-die-with-eliezer-yudkowsky#The_difficulty_of_alignment https://www.lesswrong.com/posts/wAczufCpMdaamF9fy/my-objecti...
- comp_throw7 3y agoHow can the title of the submission (which is just the title of the article) be informed by Miles' description of _the same article_ (which he posted today, nearly a year after the article was posted)?
- famouswaffles 3y agoI didn't personally change it but the title of the post was different when i originally submitted. Didn't call it inhuman though, just bizarre which it is...depending on your frame of reference.
- sp332 3y ago"Tiny Transformer trained for addition learns bizarre addition algorithm" was the title.
- famouswaffles 3y agoI don't think being bizarre or not adds much to the fear of alignment. Whether it is or not, the issue remains the same, i.e we are not the ones controlling the learning process. There isn't really any solace in the other option (surprisingly human) in terms of alignment fears. Let me put it this way. Unaligned intelligence is dangerous. You don't need a Sci-fi book to realize this, just a history one. Our history is replete with examples of the slightly more technologically advanced group decimating their competition. Humanity is directly responsible for the extinction of hundreds of species. Not because of any particular malice or offense, simply that our goals didn't align with the interest of any of those species. Just looking at Humans alone, Unaligned intelligence is potentially catastrophic even when the advantage gulf is fairly low. When the advantage gulf is large, it is potentially genocidal. So personally, "Human SuperIntelligence" doesn't exactly warm the heart.
- airgapstopgap 3y ago> Whether it is or not, the issue remains the same, i.e we are not the ones controlling the learning process I think this is just uninspected, vague intuition. What does it mean to control the learning process? No, we control the data, the deterministic learning rule, weights and activations, every step of the way is controllable – it's just it's intractable to control it all. Moreover, if it were tractable, would we know what to do? How would that be different from the problem of programming an AI from scratch? > There isn't really any solace in the other option (surprisingly human) in terms of alignment fears. If there weren't, I suppose major AI doomers (Yudkowsky, Leahy, Besinger etc.) wouldn't have been putting such an emphasis in their Lovecraft-inspired rhetoric on the alienness and inscrutability of those evil matrices of floating point numbers. No, the idea that human mental architecture is safer (and that DL does not approximate it) is very much at the core of alignment fears. E.g. Besinger on why he expects AGI Ruin [1]: "We're building "AI" in the sense of building powerful general search processes (and search processes for search processes), not building "AI" in the sense of building friendly ~humans but in silicon … The key differences between humans and "things that are more easily approximated as random search processes than as humans-plus-a-bit-of-noise" lies in lots of complicated machinery in the human brain. … This doesn't mean the problem is unsolvable; but it means that you either need to reproduce that internal machinery, in a lot of detail, in AI, or you need to build some new kind of machinery that’s safe for reasons other than the specific reasons humans are safe." > Let me put it this way. Unaligned intelligence is dangerous. I would ask you not to condescend but I have learned that this is an impossible request with the AI risk crowd, because you never really encounter pushback in your para-academic hothouse and so come to believe that, indeed, pretty basic intuitions are knockdown arguments only silly people could dismiss. You define intelligence in a certain way that has nothing to do with how you identify it in the wild. The definition is something like "intelligence is an optimization process"; the thing under consideration can be an LLM or a diffusion model that seems really good at its job. You mean to say "processes shaping real-world outcomes that are not optimizing for outcomes people consider good are dangerous". This is true but trivial. The onus is on you to tie this consideration to intelligence in general or to "capabilities" of arbitrarily powerful ML models. 1. https://www.lesswrong.com/posts/eaDCgdkbsfGqpWazi/the-basic-reasons-i-expect-agi-ruin https://www.lesswrong.com/posts/eaDCgdkbsfGqpWazi/the-basic-...