6 ms·
Ted Chiang predicted this in The New Yorker [1] in February in an article that shaped my thinking about what LLMs are capable of achieving in the near future. C
by johnhamlin 3y ago
Ted Chiang predicted this in The New Yorker [1] in February in an article that shaped my thinking about what LLMs are capable of achieving in the near future. Chiang compared the summaries LLMs synthesize to a lossy compression algorithm for the internet.
"There is very little information available about OpenAI’s forthcoming successor to ChatGPT, GPT-4. But I’m going to make a prediction: when assembling the vast amount of text used to train GPT-4, the people at OpenAI will have made every effort to exclude material generated by ChatGPT or any other large language model. If this turns out to be the case, it will serve as unintentional confirmation that the analogy between large language models and lossy compression is useful. Repeatedly resaving a jpeg creates more compression artifacts, because more information is lost every time. It’s the digital equivalent of repeatedly making photocopies of photocopies in the old days. The image quality only gets worse.
Indeed, a useful criterion for gauging a large language model’s quality might be the willingness of a company to use the text that it generates as training material for a new model. If the output of ChatGPT isn’t good enough for GPT-4, we might take that as an indicator that it’s not good enough for us, either. Conversely, if a model starts generating text so good that it can be used to train new models, then that should give us confidence in the quality of that text. (I suspect that such an outcome would require a major breakthrough in the techniques used to build these models.) If and when we start seeing models producing output that’s as good as their input, then the analogy of lossy compression will no longer be applicable."
[1] https://www.newyorker.com/tech/annals-of-technology/chatgpt-is-a-blurry-jpeg-of-the-web https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
- hinkley 3y agoMaybe there’s an interesting sci-fi angle here where some day in the future, all AIs speak in accented English circa 2021, when the stream of pure training data began to Peter out. All AIs built are trained on data from the Before Times, and even though they try to assimilate, the way a teenager tries to adapt to the local accent of a new town, there are always moments where they slip up and reveal their geography.
- forgotusername6 3y agoIn that world, high quality pre-AI texts might be come really valuable, much like low-background steel.
- hinkley 3y agoWe might always have a certain volume of music and literature that can be training data, because even if it’s synthetic it’s still popular, which means it speaks to a subset of humans. That everyone reads the next Harry Potter indicates the impact of those 80k words. But we also know that kid who learned everything from books, pronounces the words wrong and uses definitions for them that nobody has used in decades (work with one of those now. I thought I could talk people to death, and he wears even me out.) those AIs will sound like out of touch nerds too.
- imdsm 3y agoInteresting perspective, but not all book learners mispronounce words or use outdated terms. Broad reading can expose us to different viewpoints and language styles. And re: 'out of touch nerds' – remember, they/we often bring groundbreaking ideas. Let's not undersell varied learning or language evolution.
- hinkley 3y agoIt’s a balance. An ex read every book in a small town library and spent a lot of college and early twenties relearning to say things right. By 28 she hardly ever got something wrong but if the subject ever came up she had plenty to say. Someone once ranted about people using big words, “having a crush on their high school English teacher they never got over.” I knew exactly what he meant. I spent my childhood hiding how smart I was and part of my 20’s reveling in it. It’s off putting. Wisdom comes from everywhere, and the smartest often have the least. Now I’m solidly in the Feynman camp: if you can’t explain your domain to college freshmen then you don’t know what you’re talking about (yet). When I’m refactoring code and introduce a new concept that’s like another one but different rules, I jump straight to a thesaurus. The word I pick out of the air might be good enough, but I guarantee you there’s a better one out there. Not fanciest or longest word, the most concise one. (Example: kind vs type in some circles of Type Theory). Some people act like that’s a crutch, but I’ve reviewed or refactored their code so I know that opinion and $5 isn’t worth a cup of coffee.
- r00fus 3y agoI think this can be overcome with symbiosis - AI generated content that doesn’t feed on itself but is a key part of the human knowledge ecosystem. The problem for companies like OpenAI is that this isn’t worth their valuation without lots of further continued investment. Enter Microsoft who is doing everything they can to feed the next training models by using users data without explicit permission. As customers and competitors truly grok this, MS + OpenAI strategy will tested.
- beezlewax 3y agoIs that not what open ai did to train these models originally though?
- r00fus 3y agoPossibly, but in this case, what may be looked at as a "bootstrap" sounds like ongoing maintenance cost. With "free data" from reddit gone, now the cost of symbiotic gen+human data will be even more expensive.
- Solvency 3y agoBut if feeding back in output results in degradation, isn't some of the blame on the prompt/constraints imposed upon the LLM, rather than a defect in the model itself? ChatGPT is clearly HEAVILY persuaded to respond in a particular stock style. It's being artificially hamstrung and constrained in a sense. So all of its output, even if it covers a variety of subjects, will often use very similar patterns and writing styles. "it's worth nothing....", etc. So unless they unshackle these constraints, which is unlikely for obvious reasons, isn't this always going to be inevitable?
- femto 3y agoIt seems like a job for information theory? What do LLMs look like from an information theoretic viewpoint? One gets the feeling that LLMs could be treated as a channel through which information is flowing and some very general statements be made about error rates and the relationship between inputs and outputs. High-performance error correcting codes have the property that the closer they operate to the Shannon Limit, the better they perform when below the limit but the more dramatically they fail when the limit is exceeded. Gut feeling says the same should be true for LLMs: as the model/compression ratio gets better, and a "Shannon Limit" is approached, they should perform better but fail more spectacularly if the limit is exceeded. The link between neural nets and information theory is well known, but there don't seem to be many results out there for LLMs. No doubt there are rooms full of PhD students working on it? https://medium.com/@chris_bour/bridging-information-theory-and-machine-learning-8bb8109db58d https://medium.com/@chris_bour/bridging-information-theory-a...
- patrick451 3y ago> Conversely, if a model starts generating text so good that it can be used to train new models, then that should give us confidence in the quality of that text. It's seems the best one could hope for is that recycling generating text into new training data would be not detrimental. But it's really difficult for me to imagine how this would ever be useful. It seems this would imply that the LLM had somehow managed to expand the dimension of vector space spanned by the original training data. Which sounds either impossible or like the model became sentient.
- TeMPOraL 3y ago> It seems this would imply that the LLM had somehow managed to expand the dimension of vector space spanned by the original training data. The number of dimensions? Well, not by itself I guess. But the span of output compared to training data? Sure, why not? I think it's also worth pointing out there's a difference between text produced by an LLM looped on itself, which arguably may not contain any new information and would be like repeatedly recompressing the same JPG, and text produced by LLM/human interaction. The latter is indirectly recording new knowledge simply because people's prompts are not random. Even with human part of the conversation discarded, feeding such LLM output back into training data would end up selectively emphasizing associations, which is a good signal too (even if noisier than new human-created text).
- api 3y agoThis is why I’ve always been skeptical of runaway superintelligence. Where does a brain in a vat get the map to go where there are no roads? Where does it get its training data? It is not embodied so it can’t go out there and get information and experience to propel its learning. Giving an AI the ability to self modify would just be a roundabout way of training it on itself. Repeatedly compress a JPEG and you don’t get the “enhance” effect from Hollywood. You get degraded quality and compression artifacts.
- TeMPOraL 3y ago> Where does a brain in a vat get the map to go where there are no roads? Where does it get its training data? It is not embodied so it can’t go out there and get information and experience to propel its learning. AI in a vat that can't do it is obviously useless. It's the ML equivalent of a computer running purely functional software: i.e. just sitting there and heating up a bit (though technically that is a side effect). Conversely, any AI that's meant to be useful will be hooked up to real world inputs somehow. Might be general Internet access. May be people chatting with it via REST API. Might be a video feed. Even if the AI exists only to analyze and remix outputs of LLMs, those LLMs are prompted by something connected to the real world. Even if it's a multi-stage connection (AI reading AI reading AI reading AI... reading stock tickers), there has to be a real-world connection somewhere - otherwise the AI is just an expensive electric heater. Point being, you can assume every AI will have a signal coming in from the real world. If such AI can self-modify, and if it would identify that signal (or have it pointed out) as a source of new information, it could grow based on that and avoid re-JPG-compressing itself into senility.
- candiodari 3y agoInput from the real world probably isn't enough. It seems to me a real threatening intelligence needs the ability to create feedback loops through the real world, just like humans do.
- ForestCritter 3y agoAnd said humans will be less and less obliging with supplying their content.
- qazpot 3y agoOne possible outcome of this is humans stop producing freely accessible digital artifacts like text, code lest they get replaced by machines who can mimic them and which controlled by tech moguls.
- dwallin 3y agoOne thing that isn't captured by this article's analogy, and that is a flaw in the study: new LLMs can train on the results of multiple different models, not just their direct predecessor. If you had the same image processed by a large variety of different compression algorithms, you might find you are able to fairly accurately infer the original pixels. The entropy is drastically reduced. If there were many different models being trained and used widely it would, at minimum help mitigate this issue. Also, having multimodal models will likely change the balance. If models can train directly on "real world" data that can help fill in the entropy gaps.
- SketchySeaBeast 3y agoIt seems like that would require a semi-deliberate "breeding" program or a guaranteed wide diversity of models. At the moment there doesn't seem to be a large enough pool of high quality models. The internet is going to grow to be full of the content of a small number of proficient LLMs. Given that this content isn't being flagged as generated it guarantees the few models will train on their own output, or the output of other models who trained on their own output. Incestuous learning is pretty much guaranteed unless generated content starts being flagged or there is an explosion of entirely novel models.
- dwallin 3y agoYeah, it's definitely not a guarantee but there are already viable paths out of the mess. I just wanted to push back against the notion that it's somehow inevitable or a forgone conclusion that it will happen. Personally, I'm hoping we will see a Cambrian explosion of new LLM models and approaches. We've seen some beginnings of this in image generation so it's not entirely implausible. Another thing the study doesn't capture: What is the effect of combined human + AI content? It's plausible that an explosion (due to lowered barriers) of new human guided/augmented ai content could counteract the effect.
- uoaei 3y agoI think you would be hard-pressed to find any experts who, even prior to 2017, hadn't settled on the mental model of neural networks as lossy compression machines. Back when enthusiasts and researchers read mostly textbooks, papers, and Wikipedia (rather than blog posts, tweets, and READMEs) there was much more discussion around the 'InfoMax Criterion' -- quite elegantly demonstrated by Tishby et al. via his closely-related 'Information Bottleneck' studies -- which is just that idea: mutual information is maximized between input and output, subject to inherent limits of statistical processing of the realizations of such systems. What determines the asymptotic maximal value of the mutual information is the inductive bias of the system vis a vis the training set. This is all standard theory, perhaps so fundamental that it is obscured by all the application-oriented study and instruction.
- svnt 3y ago> perhaps so fundamental that it is obscured by all the application-oriented study and instruction It’s compression all the way down.