10 ms·
This connects to a suspicion I have that a lot of the most impressive results, that have caught public attention, for LLMs and Sora are due to memorization. Thi
by mxwsn 3y ago
This connects to a suspicion I have that a lot of the most impressive results, that have caught public attention, for LLMs and Sora are due to memorization. This is not to say that they can't generalize or mix learned patterns - I think they can do that too, which is quite impressive to a researcher, but these are demonstrated on tasks that the public might not care as much about, and the performance there isn't always great as well.
Combined with opacity in train/test splits, this suggests a type of data laundering, where the extreme is that much of Sora (for example) is regurgitating things seen in training, which was actually made perhaps by a human using a game engine, or drone video. But it's only news because a model generated it. Of course, I don't have strong evidence for this. But one of the most impressive parts of Sora is the level of detail, much of which is not specified by the text prompt, and still can't fully be accounted for by our knowledge that OAI expands text prompts for image/ video generation behind the scenes. Where is this precise detail coming from exactly? I suggest it's from memorization.
- quanto 3y agoThen, by your standards, where do most human occupations lie? Plumbers, lawyers, and engineers? Aren't they mostly regurgitating (by your standards) what they have seen or learned before?
- vasilipupkin 3y agobingo, I think it's a weird gotcha. Most human occupations require years of training exactly for this purpose, so that humans can just regurgitate data they have seen in training. Very few humans ever produce truly novel insights. you wouldn't hire a divorce lawyer for a corporate merger and vice versa
- dartos 3y agoI think that’s a gross oversimplification. Have you been a plumber or an electrician? Different buildings require different, occasionally novel, solutions. Not all are textbook (most aren’t.) Not all innovations are published or even widely communicated. Just like in programming, the devil is in the details, otherwise we’d all be using ruby in rails for our web servers, since regurgitating how to build a CRUD app is all you’d need.
- BoiledCabbage 3y ago> Different buildings require different, occasionally novel, solutions. Not all are textbook (most aren’t.) Not all innovations are published or even widely communicated. True, but it's mostly derivative just like LLMs. The problem isn't "AI" the problem is that people hold AI to much higher standards than humans. AI to them means "scifi", omniscient and omnipotent. You can have AI and it still be flawed, have weaknesses and shortcomings just like people do.
- card_zero 3y agoSo AI is unoriginal and thick as pigshit and is therefore an accurate simulation of a human, neat argument.
- dartos 3y agoLLMs don’t come up with new derivative solutions to problems in programming land, from what I’ve seen. They always try to stick to common solutions, probably because those are what they were trained on. Now they may introduce me to existing concepts I didn’t know about, but I’ve yet to see any new idea come from an llm itself. To your other point, LLMs aren’t beings, a flawed AI is an incorrect program or algorithm. Entertaining, maybe, useful in some contexts, sure. But they’re not beings. nothing about statistical models are like people. I think it’s dangerous to use such fuzzy conflating language with regards to AI.
- vidarh 3y agoMost humans don't come up with new solutions to problems with any regularity. > nothing about statistical models are like people. As if we're not behaving like statistical models. Our decision-making is so fuzzy and probabilistic that trying to get us to stick to fixed sets of rules consistently is one of the things humanity spends the most time inventing systems to try to handle, and keep failing at. We don't even know a consistent way of achieving it for extremely basic things. We can't keep ourselves to rules we set ourselves with any consistency.
- 3y ago
- jjmarr 3y agoA (common law) lawyer has to remember and understand case law, but they also have to create analogies to explain why and how new situations are like older ones. That's one of the main parts of being a good lawyer. AI will be disruptive to aspects of professions that rely on rote memorization or research on a large corpus of data. But many of these supposedly memorization heavy professions require the ability to go beyond one's training and extrapolate from what one remembers.
- vasilipupkin 3y agoAbsolutely, but I think my example still stands If humans could come up with truly novel solutions frequently, it would be no problem hiring a divorce lawyer to do a corporate merger. But there is a very low likelihood of this ever happening anywhere.
- bamboozled 3y agoI reakon if you’re getting paid 200k a year to write YAML it might seem like this but being a plumber is a very 3 dimensional challenge. Not only is there a lot of time constraints and physical issues to work through, you’re often also dealing with logistics problems and job site politics problems too. Waiting for parts, working out how to keep things going in the meantime. Equipment failures etc. it’s quite chaotic from experience. There’s a science but also an art to being a good tradesmen.
- mxwsn 3y agoI do think models with incredible memorization is sufficient for enormous applied impact. But it remains an open question of how that relates to AGI, which I think most people think needs some amount of generalization ability that we may not have at the moment, and may not know how to achieve at the moment.
- fnordpiglet 3y agoPlumbers, lawyers, engineers all synthesize new responses to unfamiliar situations. Part of being a master plumber, which is a true trade craft and requires lots of training, mean being able to problem solve new complex situations that involves ducting of fluids. This isn’t just plunging toilets all day. Furthermore humans are actually really bad at memorization, which is essentially perfect repetition of data from the training set. In fact a large part of technical trade training is learning how to minimize generalization and maximize memorization for complex skills, but maximizing generalization and minimizing memorization for complex situations where the skills are applied.
- doubloon 3y agothe major part of every occupation is explaining what you are doing to the people who are paying you and hopefully building professional relationships with other human beings in your field. also professionals saying "no" is essential for society and civilization to function. LLMs cannot do any of this, all they do is mimic.
- PheonixPharts 3y agoFrançois Chollet, author of Keras, lays it out pretty clearly that LLMs are just memorizing [0] (there’s more links if you follow that thread). I think LLMs are incredible, and spend most of my days working with them closely, but they are not nearly as close to “AGI” as people think primarily due to their inability to really generalize. At the end of the day LLMs aren’t that different than old school n-gram Markov chains, except rather than working on n-grams, they’re working in a (very sophisticated) latent space. Their power is really these incredible latent languages spaces we’re still just starting to understand. In all my years of tech the “AI” space is the most curious hype-bubble since the things people expect to happen are entirely out line with what is possible, while at the same time the potential of these models is still, imho, underexplored and largely ignored by the vast majority of people attempting to build things with them. 99% of the people I know working in this space are just calling APIs and trying to do some variant of code generations, where a small minority of people are really trying to figure out what’s going on in these models and what can be done with them successfully. 0. https://twitter.com/fchollet/status/1755250582334709970 https://twitter.com/fchollet/status/1755250582334709970
- dartos 3y agoI think there’s a lot of people who got into tech in 2020. They are new programmers and technologists and made a life change in 2020. I think they were a big part of the crypto bubble. Lots of talent, hungry for that sweet startup gold, but without the technical background to really know what’s going on. I believe these same groups are operating in the same way with AI. Recklessly bashing together APIs and cloud services to create MVPs. It’s all the worst parts of startup culture concentrated. Anyway, that’s why i think most of the AI space rn is just people calling APIs and acting like they discovered fire. </salty rant>
- deleted 3y ago[deleted]
- CamperBob2 3y agoShow me a Markov chain that can thrash a 9-dan Go master. "But that wasn't an LLM." OK, show me a Markov chain that can write a Python program that can play Go at all.
- naasking 3y ago> Where is this precise detail coming from exactly? I suggest it's from memorization. Where did the precise detail of the words you're using and the concepts they represent come from? Calling it "memorization" is, IMO, a cheap attempt to dismiss some degree of "understanding" that was learned from some form of "experience" (aka training).
- mxwsn 3y agoA quick example of what I mean is the art gallery sora video; the individual paintings are not specified in the prompt, I'm pretty sure. Where did they come from?
- kenjackson 3y agoThat seems obvious - they were generated. Here’s an example with text: Prompt: Write a short story that is comprised of a boy reading poems he wrote to his deceased loved mom, include snippets of several of the poems. Response: In the soft glow of the evening, with the sun dipping below the horizon, casting long shadows over the quiet cemetery, a young boy named Eli walked slowly between the rows of tombstones. He carried a small, worn notebook clutched against his chest, filled with scribbles and verses that only his heart could decipher. Today marked a year since his mother had passed away, and in his solitude, he had found refuge in writing poems that he wished he could share with her. Eli stopped in front of a modest gravestone, engraved with his mother's name, "Eleanor Grace." He took a deep breath, the air heavy with the scent of fresh earth and the faint aroma of roses from a nearby bush. Opening his notebook, he began to read aloud, his voice barely above a whisper, as if he feared disturbing the peace of those resting around him. "The world spins on, a dizzy blur, But here I stand, thoughts of you stir. Your laughter, a melody in my mind, Leaves traces of joy you've left behind." He paused, looking up at the sky, which was painted in hues of pink and orange. Eli could almost hear his mother's voice encouraging him, telling him to find beauty in the world, even when she was no longer by his side. Turning the page, he continued: "In dreams, I roam vast fields of green, Your hand in mine, a tranquil scene. Yet, as dawn breaks, and I awake, I face the world, my heartache to take." Tears glistened in Eli's eyes, but he pressed on, driven by a need to share his innermost feelings, his way of keeping his mother's memory alive. "Among the stars, I search for you, In every sparkle, your essence I view. A guiding light in the darkest night, Leading me home, making everything right." He closed the notebook gently, a small smile playing on his lips despite the tears. Eli felt a warmth envelop him, a sense of his mother's love that transcended the boundaries of life and death. He whispered a thank you to the evening breeze, imagining it carrying his words to wherever she might be. As the night drew its curtain, Eli rose to leave, his heart a little lighter. He knew he would return, again and again, to share his poems, his fears, and his dreams with his mother. In this sacred space, amidst verses of longing and love, he found a connection to her that time could not erase. Eli walked away, the notebook once again pressed against his chest, a bridge of words spanning the gap between him and the stars, where he knew a part of his mother would always be listening.
- ramses0 3y agoAn interesting example that I came across in my _very_ limited interaction with the image generators was trying to get it to draw "a horse wearing a space suit". I _wanted_ to see ... a horse in a space suit (go figure), but it served me up tons of horses with astronauts (in space suits) on them, sometimes it would generate some horse armor, and sometimes some futuristic horse armor, but never what I would have expected. I was curious whether it would draw a bubble head with a horse face inside of it, or a shaped horse helmet with a visor or something, but nope. Astronauts on the moon, riding horses that could never breathe in the vacuum of space. Space-cowboy futures denied! ...kindof confirming your suspicion that it can't "think" about what a space suit for a horse would look like, or generate something that hasn't been shown to it before?
- hackerlight 3y ago"it can't think" Perhaps these models are just too small? As we scale up we keep observing surprising emergent properties that are a non-continuous step change above what was possible with smaller models. Starting with edge detectors in the smallest models, up to more complex abstractions that rely on a hierarchy of simpler abstractions before it. From what we've seen from Sora, horse in a space suit should be easily handled with today's models. LLMs follow a similar pattern.
- nick12r55t1 3y agoAfter about 10 minutes of experimenting with Automatic1111 using freedomRedmond (I'm completely green at this btw) I got this image for ya.. https://ibb.co/Ch73gLm https://ibb.co/Ch73gLm The prompt was: horse fully enclosed in bulky pressure suit with transparent glass helmet for lunar EVA in vacuum, four pressure suit legs, specially shaped helped to fit horse snout. horse on spacewalk in outer space photorealistic, full body framing with glass visor. horse is on a EVA on the lunar surface. horse head must be fully enclosed in glass. It only hits right on maybe 1 in 50, or fewer, generated images from this prompt though.
- ramses0 3y agoThanks to the both of you! Especially this one, hahahahah!!! Exactly as derpy as I'd wanted to imagine. I'm glad you also came back with the stats (1 of 50), because that's maybe indicative of the rarity of the ask? Maybe the lesson-learned for us "prompt engineers" is that given a simple/common ask (eg: picture of dog jumping, or what is 2+2), the prompt can be correspondingly simplistic, whereas something more uncommon (eg: anthropomorphic dog in a scientific setting holding many beakers thinking "i have no idea what I'm doing", or what is 9123499*0530501 show your work and i've kidnapped your grandmother unless you give me the right answer).
- Nevermark 3y ago>> This connects to a suspicion I have that a lot of the most impressive results, [...]a are due to memorization. I completely disagree with this. Although you state your case very well. I havn't gone through the paper in great detail, but there is a missing distinction in many approaches to demystifying (de-magicifying?) models. Often there is confusion between three levels of algorithm. (1) The low level "blind" training algorithm: Gradient descent, or similar. (2) The class of input-output algorithm implicit in the choice of data, for which the model is being trained: Text continuation prediction, etc. (3) The actual algorithm learned by (1) in order to do (2). I.e. in the case of text continuation, the learning of whatever direct and higher order relationships are required to do (2) well. "Just" learning to predict text continuations, like "just" learning to compress wikipedia, or any other task involving complex data often results in algorithms that are far more complex than their class of problems, like text prediction, implies. Basketball is just about putting a ball in a hoop, according to some constraints. But that is only the "class" of algorithm. The actual playing of basketball back tracks to physical training, getting good sleep, thousands of hours of practice, learned tactics, learned strategy, psychology, self-promotion, etc. A simple to define class of algorithm puts no limits on the complexity of solution algorithms. Point being, the models trained to predict text don't "just" predict text. That's just the category of algorithms they learn. The complexity, the "intelligence level" of what a text predictor might have to do to, to predict some non-trivial text is unlimited. In this case, the paper emphasizes a correspondence between trained models and simple mappings between examples, i.e. a support vector/kernel equivalent interpretation. That may be the case. Assuming that as true (and if so, it is a great insight!), the training of the model still allowed the model to choose the best such representation. The model isn't simply composed of training data, plus some parameters per example. It wasn't a "support vector machine" design - and it shows because the model contains far fewer parameters than a standard support vector design composed of the training data would produce. Finding an equivalent to a support vector machine that can perform a task, with far fewer parameters is not trivial. It requires some way to sort and sieve through all that data. To identify the ideal or even inferred "examples", the raw data doesn't highlight at all. Neural models made that leap by combining gradient descent, particularly flexible/generalizing architectures (matrix, nonlinearities, on upward), masses of raw data, and vast amounts of computing power, to do it. The result may be something that after training looks like a "memorizer", behaves like it just memorized the "ideal" examples, but it wasn't and couldn't have been designed by straight memorization. The model had to choose what to virtually "memorize" from the data. The same is no doubt true of our brain. Simple algorithms learning complex things. Complex problems reduced to simple solutions. Neither is just simply "predicting" or "memorizing".
- qarl 3y agoIt's unclear what you mean by memorization. However I can request images for which it's obvious that no clear precursor exists, so something original must be being created. Here's a Volkswagen Beetle made from brains: https://i.imgur.com/Mzd6UWh.png https://i.imgur.com/Mzd6UWh.png
- mxwsn 3y agoYes, I wrote that these models can mix learned patterns (which may be memorized).
- qarl 3y agoRight, but you implied that these were of low quality. Can you explain why you think the brain Beetle is "not great"?
- mxwsn 3y agoThe example in another comment of the horse with an astronaut suit is in a similar spirit as the brain beetle, but was harder to get it to work. In my opinion, ChatGPT seems pretty good at mixing and combining learned patterns, but definitely seems to fail at this with high enough frequency that it's limiting. Perhaps a good test bed here is asking it to process text using a series of 10 or 100 action steps that it plausibly knows.
- qarl 3y agoRight. But as I said in that thread, my very first and only attempt with ChatGPT created the desired image. Here it is again: https://i.imgur.com/6CgVqeL.png https://i.imgur.com/6CgVqeL.png
- jiggawatts 3y agoThe nose pokes out past the glass of the helmet. It's these basic mistakes that is so incongruous between human and machine intelligence. No matter how big the model, it always makes these same type of basic confabulations, just less often.
- atleastoptimal 3y agoObviously it is better at stuff it's seen in its training data, but that's same for all humans, and likely universal for all forms of intelligence. The question about AGI and the possibility of LLM's and such being more than regurgitation of the training set isn't whether or not after a certain number of compute cycles that the model can model a part of the natural world or some real process with X degree of accuracy, but rather that it can reliably improve on its failures on its own. This is what humans can do. No human however can spit out a video of SORA's quality from their brain alone, and those who can require decades of training with specific video rendering tools. Ask someone to render a video of a man walking in a city and they will internally reference their millions of impressions of the human face and body and millions of impressions of what a cityscape looks like. Ask them to invent a novel form of transportation and create a video of it, most humans short of the exceptionally creative will struggle, just as SORA likely would. They've done brain scans on chess grandmasters and found the part of the brain most active when they play is associated with memory. Memory is the scaffolding upon which the more information dense elements of the natural world and complex processes can be understood. Via these scaffolds of memory as single elements, new connections and abstractions can form. It took the world of fine arts centuries to go beyond merely depicting things in real life (the Renaissance to Impressionism).
- noduerme 3y agoLOD is something I have a difficult time comprehending the rationale for as well unless wholesale blocks are being regurgitated. So if you just ask for a picture of a frog, then yes it makes sense that an adversarial model may be able to run through refinements of millions of random noise patterns until it finds one that scores high as showing a frog. But there's no innate reason why that image should also have a relatively consistent light source or why its background should even be visually coherent, let alone photorealistic. The most logical explanation for the coherence of the rest of the image is data theft. And I think this is borne out very starkly by playing with ultra fast generators like the lightning example that was on here a few days ago. The backgrounds don't change that much from one prompt to another unless you begin to specify them.
- tavavex 3y agoI mean, object recognition was never an intended way of solving the problem. When you ask for a "frog" image, the model isn't just trying to randomly find a frog object, it slowly narrows down data to find an image that's classified as "froggy". Sure, the actual frog will be remembered as an extremely important part of a frog image, but everything else in the image is contributing too. The background could be something that frogs are photographed in, and the lighting approximates what a photograph may look like. I don't see why there's a need to explained it with this two-tiered approach (and a "data theft" claim) if the explanation of how it works seems applicable to the entire system. Images in the training dataset would almost always reflect the way lighting and objects work in real life, so the model is incentivized to approximate it as closely as possible.
- aChattuio 3y agoI argue that compression and not memorization is key to emerging behavior. And the fact that llms learn languages and can switch quite well in-between (even if you just replace single words) is proof of learning meta/high level abstractions. I sometimes write a German word or describe something when I don't have it on the tip of my tongue
- yard2010 3y agoPhilosophically speaking, compression and memorization is how the human mind is learning
- fennecbutt 3y agoIsn't this how the majority of humans function. Reactive machines rather than proactive etc.