17 ms·
Will scaling work?
- mewpmewp2 3y agoWhy wouldn't you include the "LLM" part in the title? Hint for everyone else here: It's about scaling LLMs.
- mewpmewp2 3y agoAfter reading the article, I really enjoyed it and the believer + skeptic perspective. However, I only touched it because I thought "meh, what is there going to be about web scaling".
- hokeone 3y ago>Furthermore, the fact that LLMs seem to need such a stupendous amount of data to get such mediocre reasoning indicates that they simply are not generalizing. If these models can’t get anywhere close to human level performance with the data a human would see in 20,000 years, we should entertain the possibility that 2,000,000,000 years worth of data will be also be insufficient. There’s no amount of jet fuel you can add to an airplane to make it reach the moon. Never thought about it in this sense. Is he wrong?
- gchamonlive 3y agoI don't think he is wrong. I also don't think the goal of LLMs is to reproduce human intelligence. That is, we don't need human-like inteligence in a box for a tool to be useful. So this assertion could be right and still miss the point of this tech in my opinion. Edit: to expand, if the goal is AGI then yes we need all the help we can get. But even so, AGI is in a totally different league compared to human intelligence, they might as well be a different species.
- jpk 3y agoThe context of the fine article is scaling LLMs into AGI. It's not about whether the tool is useful or not, as usefulness is a threshold well before AGI. Some folks are spooked that LLMs are a few optimizations away from the singularity, and the article just discusses some reasons why that probably isn't the case.
- gchamonlive 3y agoThe article is really good! I was responding to "is he wrong" part of the comment, not the article itself.
- tremarley 3y agoWe don’t need human-like intelligence in a box for a tool to be useful, But human-like intelligence is what many companies are spending billions to try and achieve
- dartos 3y agoThis. I don’t think LLMs are anywhere near sci-fi AGI (think I, robot) It’s such a vague term anyway, AGI. LLMs provide some really nice text generation, summarization, and outstanding semantic search. It’s drop dead easy to make a natural language interface to anything now. That’s a big deal. That’s what’s going to give this tech it’s longevity, imo.
- barrenko 3y agohe's not wrong, and yet he's not right.
- gitfan86 3y agoOver the past year there have been advances in making models smaller while keeping performance high. So if that continues then he is wrong unless he is defining LLMs in a strict way that does not include new improvement in the future
- viraptor 3y agoFor an example, the diagrams in the post compare the big gpts, but looking at the number of tokens PHI-2 sits below gpt3. And it still beats it in Humaneval and a few other benchmarks.
- Xelynega 3y agoIt's not about the size of the models, it's about the size of the training data. Humans are able to begin to generalize with a single persons experiences over less than a year, so the fact that LLMs cannot with billions of person-years of information could be an indicator of their inability to generalize no matter how much training data you throw at it.
- az226 3y agoLLMs are closer to discoveries on the spectrum than inventions. Nobody predicted or planned the many emergent capabilities we’ve seen. Almost like magic. Now is a period of moving along the axis to invention with many intentional design, architecture, and feature development alongside testing and evaluation. We are far from done with LLMs, plenty of room for many more discoveries, lots to explore. It’s definitely a precursor to AGI. They offer a platform to build and scale data sets and test beds. We haven’t had ML models this large before. There’s innovation in architecture but we often come back to the bitter lesson: more data. We’re likely going to see experimentation with language models to learn from few examples. Fine tuning pretrained LLMs shows they have quite a remarkable ability to learn from few examples. Liquid AI has a new learning architecture for dynamic learning and much smaller models. Some people seem mad about the bitter lesson, they want their model based on human features to work better when so far usually more data wins. I think the next evolution here is in increasing the quality of the training data and giving it more structure. I suspect the right setup can seed emergent capabilities.
- gitfan86 3y agoThe trick is to make many LLMs work together in feedback loops. Some small some big. That will get us to what was previously known as AGI. The definition of AGI will change, but we will have systems that put perform humans in most ways.
- beardedwizard 3y ago> It’s definitely a precursor to AGI. What are you basing this claim on? There is no intelligence in an LLM, only humans fooled by randomness.
- auggierose 3y agoAnd yet, we reached the moon, and I would say airplanes were a necessary step on the way, even if only for psychological reasons. For airplanes we had at least an example in nature, birds. But I am not aware of any animal that travelled from earth to the moon on its own, except us.
- manojlds 3y agoWe are talking of LLMs, not whether we will be able to reach AGI or not.
- ImHereToVote 3y agoAirplanes in this analogy are essentially the collection of matrix multiplications that emulate reasoning in a very rough but useful manner in an LLM. It's unclear whether a rocket ship is a multimodal neural net. Or some sort of swarm of LLM's in an adversarial relationship, or something completely novel. Regardless, we might be as far between LLM's to ASI's, as airplanes are to rocket ships. Or not.
- Eddy_Viscosity2 3y agoBut we didn't use airplanes to get there. It needed a new approach, different propulsion, different fuel, different attitude control, etc. etc. LLM may be a necessary step to get to AGI, but it (probably) won't be the one that achieves that goal.
- auggierose 3y agoI doubt that LLMs will give us AGI. But they have already given us more intelligence from a computer than I would have imagined to see during my lifetime.
- sbierwagen 3y agoI mean, if you go through the Apollo program contractors it's a who's who of aerospace. Boeing, North American Aviation, Grumman Aircraft, McDonnell Aircraft Corporation, Bell Aerosystems, Rocketdyne... Electrical parts ran at aviation-standard 400hz. Aviation gyroscopes and aviation instruments. Structural parts made of aviation aluminum alloys. Astronauts that are all airplane test pilots. I can imagine doing Apollo from complete scratch (using car manufacturers that have to invent aluminum-handling tech starting from nothing) but it would have taken a lot more than the decade Apollo took.
- MPSimmons 3y agoI don't think the data is the weakness. We're using Transformer architecture right now. There's no reason there won't be further discoveries in AI that are as impactful as "Attention is All You Need". We may be due for another "AI Winter" where we don't see dramatic improvement across the board. We may not. Regardless, LLMs using the Transformer architecture may not have human level intelligence, but they _are_ useful, and they'll continue to be useful. In the 90s, even during the AI winter, we were able to use Bayesian classification for such common tasks as email filtering. There's no reason we can't continue to use Transformer architecture LLMs for common purposes too. Content production alone makes it worth while. We don't _need_ AGI, it just seems like the direction we are heading as a species. If we don't get there, it's fine. No need to throw the baby out with the bath water.
- Der_Einzige 3y agoEven the largest LLM has had less "total information" than most humans take in through all of their senses over their lifetime. A single day for a baby is taking in a continuous stream of among other things high quality video and audio and does a large amount of processing on that. Much of that for very young babies is unsupervised learning (clustering), where baby learns that object A and object B are different despite knowing nothing else about their properties. Humans can learn using every ML learning paradigm in ever modality: unsupervised, self-supervised, semi-supervised, supervised, active, reinforcement based, and anything else I might be missing. Current LLMs are stuck with "self-supervised" with the occasional reinforced (RLHF) or supervised (DPO) cherry on top at the end. non multi-modal LLMs operate with one modality. We are hardly scratching the surface on what's possible with multi-modal LLMs today. We are hardly scratching the surface for training data for these models. The overwhelming majority of todays LLMs are vastly undertrained and exhibit behavior of undertrained systems. The claim from the OP about scale not giving us further emergent properties flies in the face of all of what we know about this field. Expect further significant gains despite nay-sayers claiming it's impossible.
- haltist 3y agoYou are obviously a believer so you should know I know how to build AGI with a patented and trademarked architecture called "panoptic computronium cathedral"™. Tell all your friends about it. I only need $80B to achieve AGI.
- lumost 3y agoThe Phi paper and various approaches to distilling from GPT-4 demonstrate that the training data and plausibly order of presentation matter. The challenge is that we both do not understand which set of data is most beneficial for training, or how it could be efficiently ordered without triggering computationally infeasible problems. However we do know how to massively scale up training.
- espadrine 3y agoDemis Hassabis of Deepmind echoes a similar sentiment[0]: > I still think there are missing things with the current systems. […] I regard it a bit like the Industrial Revolution where there was all these amazing new ideas about energy and power and so on, but it was fueled by the fact that there were dead dinosaurs, and coal and oil just lying in the ground. Imagine how much harder the Industrial Revolution would have been without that. We would have had to jump to nuclear or solar somehow in one go. [In AI research,] the equivalent of that oil is just the Internet, this massive human-curated artefact. […] And of course, we can draw on that. And there's just a lot more information there, I think, it turns out than any of us can comprehend, really. […] [T]here's still things I think that are missing. I think we're not good at planning. We need to fix factuality. I also think there's room for memory and episodic memory. [0]: https://cbmm.mit.edu/video/cbmm10-panel-research-intelligence-age-ai https://cbmm.mit.edu/video/cbmm10-panel-research-intelligenc...
- skippyboxedhero 3y agoHis view of the Industrial Revolution is completely wrong. Societies pre-IR had multiple periods where energy usage increased significantly, some of them based specifically around coal. No IR. Early IR was largely based around the usage of water power, not coal. IR was pure innovation, people being able to imagine and create the impossible, it was going straight to nuclear already. Ironically, someone who is an innovator believes the very anti-innovation narrative of the IR (very roughly, this is the anti-Eurocentric stuff that began appearing in the 2000s...the world has moved on since then as these theories are obviously wrong). Nothing tells you more about how busted modern universities are than this fact.
- archon1410 3y agoHas the narrative moved on? The historian and blogger Bret Devereaux presents a view on a 2022 blog post that seems to back up what the Deepmind CEO is saying. > The specificity matters here because each innovation in the chain required not merely the discovery of the principle, but also the design and an economically viable use-case to all line up in order to have impact. https://acoup.blog/2022/08/26/collections-why-no-roman-industrial-revolution/ https://acoup.blog/2022/08/26/collections-why-no-roman-indus...
- YetAnotherNick 3y ago> ‘5 OOMs off’ I think Google, Microsoft and facebook could easily have 5 OOM data than the entire public web combined if we just count text. Majority of people don't have any content on public web except for personal photos. A minority has few public social media posts and it is rare for people to write blog or research paper etc. And almost everyone has some content written in mail or docs or messaging.
- nmca 3y agoFrom the article, and relevant here: I’m worried that when people hear ‘5 OOMs off’, how they register it is, “Oh we have 5x less data than we need - we just need a couple of 2x improvements in data efficiency, and we’re golden”. After all, what’s a couple OOMs between friends? No, 5 OOMs off means we have 100,000x less data than we need.
- YetAnotherNick 3y agoI meant 100,000x. At least for everyone I know, they have 100,000x data in mail/messaging/docs/notes/meeting etc. than their blog or any public site they own. Hell I would even say that if you just have all the meetings of zoom, it will be few order of magnitude higher than the entire public web.
- saulpw 3y agoIf I have 1MB on my blog, 100,000x would be 100GB. Just, no. OOMs are not to be trifled with.
- YetAnotherNick 3y agoHow many people have blogs? How many people sent any message or created a google docs? The answer could easily be 10,000x times of people having blog. Also I was just counting text content as I mentioned. For reference, there are 175,000 authors in medium compared to billions using whatsapp or gmail or difference of around 50,000.
- 3y ago
- ralusek 3y agoAlmost everything interesting about AI so far has been unexpected emergent behavior, and huge gains through minor insights. While I don't doubt that the current architecture is likely to have a current ceiling below that of peak human intelligence in certain dimensions, it's already surpassed it in some, and there are still gains to be made in others through things like synthetic data. I also don't understand the claims that it doesn't generalize. I currently use it to solve problems that I can absolutely guarantee were not in its training set, and it generalizes well enough. I also think that one of the easiest ways to get it to generalize better would simply be through giving it synthetic data which demonstrates the process of generalizing. It also seems foolish to extrapolate on what we have under the assumption that there won't be key insights/changes in architecture as we get to the limitations of synthetic data wins/multi-modal wins.
- jahnu 3y ago> problems that I can absolutely guarantee were not in its training set Can you share the strongest example?
- Jabrov 3y agoPretty much any coding problem in a unique or private codebase
- Der_Einzige 3y agoThere is a difference between interpolation, which the majority of humans are performing daily with coding in private codebases, and genuine extrapolation, which is difficult to prove and difficult to find in high dimensional spaces. LLMs may not be able to easily extrapolate (and when it does it's due to high temperature), but they can interpolate extremely well, and most human growth and innovation today comes from novel interpolations, which are what LLMs are excellent at.
- jahnu 3y agoI asked for the strongest example the OP can share in order to evaluate their claim. If it's so obvious to the OP that generalisation is happening then it should be easy to provide a strong example, right?
- nemo44x 3y agoWhere in the hype cycle are we for LLMs? Are we in the late stages of the rise or over the peak and beginning the slide?
- crowbahr 3y agoStill on the climb imo
- kevindamm 3y agoIf you have the answer to that question you could make some very lucrative investments.
- collaborative 3y agoLLMs are still too expensive to run and therefore can't be supported by ads. If costs get lower we'll see them being pushed _a lot_ more
- jsnell 3y agoThe original title ("will scaling work?") seems like a much more accurate description of the article than the editorialized "why scaling will not work" that this got submitted with. The conclusion of the article is not that scaling won't work! It's the opposite, the author thinks that AGI before 2040 is more likely than not.
- bee_rider 3y agoIt might be nice to modify the title a bit though, to indicate that it is about AGI. Obviously scaling works in general, just ask anyone in HPC, haha.
- PlasmonOwl 3y agoAuthor is leveraging mental inflexibility to generate an emotional response of denial. Sure, his points are correct but are constrained. Let’s remove 2 constraints and reevaluate: 1 - Babies learn much more with much less 2 - Video training data can be made in theory at incredible rates The questions becomes: why is the author focusing on approaches in AI investigated in like 2012? Does the author think SOTA is text only? Are OpenAI or other market leaders only focusing on text? Probably not.
- Xelynega 3y agoIsn't 1 a point for their "skeptic" persona? If babies learn much more from much less, isn't that evidence that the LLM approach isn't as efficient as whatever approach humans implement biologically, so it's likely LLM processes won't "scale to ago"? For video data, that's not how LLMs work(or any NNs for that matter). You have to train them on what you want them to look at, so if you want them to predict the next token of text given an input array, you need to train it on the input arrays and output tokens. You can extract the data in the form you need from the video content, but presumably that's already been done for the most part, since video transcripts are likely included in the training data for gpt.
- berniedurfee 3y agoI think there’s a huge assumption here that more LLM will lead to AGI. Nothing I’ve seen or learned about LLMs leads me to believe that LLMs are in fact a pathway to AGI. LLMs trained on more data with more efficient algorithms will make for more interesting tools built with LLMs, but I don’t see this technology as a foundation for AGI. LLMs don’t “reason” in any sense of the word that I understand and I think the ability to reason is table stakes for AGI.
- cortic 3y agoIf humans are basically evolved LLMs, which i think is likely; Reasoning will be an emergent property of LLMs within context with appropriate weights.
- enieslobby 3y agoWhy do you think humans are basically evolved LLMs? Honest question, would love to read more about this viewpoint.
- cortic 3y agoLook at a year old baby, there is no logic, no reasoning, no real consciousness, just basic algorithms and data input ports. It takes ten years of data sets before these emergent properties start to develop, and another ten years before anything of value can be output.
- berniedurfee 3y agoI strongly disagree. Kids, even infants, show a remarkable degree of sophistication in relation to an LLM. I admit that humans don’t progress much behaviorally, outside of intellect, past our teen years; we’re very instinct driven. But still, I think even very young children have a spark that’s something far beyond rote token generation. I think it’s typical human hubris (and clever marketing) to believe that we can invent AGI in less than 100 years when it took nature millions of years to develop. Until we understand consciousness, we won’t be able to replicate it and we’re a very long way from that leap.
- mgaunard 3y agoI think the more interesting question is how long will people cling to the illusion that LLMs will lead us to AGI? Maintaining the illusion is important to keep the money flowing in.
- beepbooptheory 3y agoWhile this is certainly true, I think we can't ignore the intense enthusiasm and faith of a large cohort of our peers (or, you know, HN commenters) who believe this to be The Way, and are not necessarily stakeholders in any meaningful sense. Just look at some of the responses even in this thread. It feels like some people just need this, and respond to balanced skepticism as Alyosha does to his brother Ivan. In part, whether conscious or not, people see the bright future of LLMs as a kind of redemption for the world so far wrought from a Silicon Valley ideology; its almost too on-the-nose the way chatgpt "fixes" internet search. But on a deeper level, consider how many hn posts we saw before chatgpt that were some variation of "I have reached a pinnacle of career accomplishment in the tech world, but I can't find meaning or value in my life." We don't seem to see those posts quite as much with all this AI stuff in the air. People seem to find some kind of existential value in the LLMs, one with an urgency that does not permit skepticism or critique. And, of course, in this thread alone, there is the constant refrain: "well, perhaps we are large language models ourselves after all..." This reflex to crude Skinnerism says a lot too: there are some that, I think, seek to be able to conquer even themselves; to reduce their inner life to python code and data, because it is something they can know and understand and thus have some kind of (sense) of control or insight about it. I don't want to be harsh saying this, people need something to believe in. I just think we can't discount how personal all this appears to be for a lot of regular, non-AI-CEO people. It is just extremely interesting, this culture and ideology being built around this. To me it rivals the LLMs themselves as a kind fascinating subject of inquiry.
- visarga 3y agoThere is no magic in the brain. There is no magic in LLMs. There is just new experience we gain by interacting with the environment and society. And there is the trove of past experience encoded in our books. We got smart by collecting experience, in other words, from outside. The magic in the brain was not in the brain, but everywhere else. What is experience? We are in state S, and take action A, and observe feedback R. The environment is the teacher, giving us reward signals. We can only increase our knowledge incrementally, by trying our many bad ideas, and sometimes paying with our lives. But we still leave morsels of newly acquired experience for future generations. We are experience machines, both individually and socially. And intelligence is the distilled experience of the past, encoded in concepts, methods and knowledge. Intelligence is a collective process. None of us could reach our current level without language and society. Human language is in a way smarter than humans.
- slibhb 3y agoThe best analogy for LLMs (up to and including AGI) is the internet + google search. Imagine explaining the internet/google to someone in 1950. That person might say "Oh my god, everything will change! Instantaneous, cheap communication! The world's information available at light speed! Science will accelerate, productivity will explode!" And yet, 70 years later, things have certainly changed, but we're living in the same world with the same general patterns and limitations. With LLMs I expect something similar. Not a singularity, just a new, better tool that, yes, changes things, increases productivity, but leaves human societies more or less the same. I'd like to be wrong but I can't help but feel that people predicting a revolution are making the same, understandable mistake as my hypothetical 1950s person.
- jerpint 3y agoThe internet has allowed us to interact in ways that were inconceivable at the time; think communication and speed of information for one. When agents start being more reliable I think we will start seeing applications we couldn’t possibly anticipate today
- arketyp 3y agoMy take on this is that much of work and problem solving is about understanding the problem. So I think human abilities will remain the bottleneck. I pose this thought experiment: Is it possible to design an AI system for a monkey which gives it super-monkey abilities?
- red75prime 3y agoFor a monkey it's impossible to design... pretty much anything beside a few simple tools. So, no. A monkey cannot design a bow, a loom, a tractor, a computer, or an AI of any kind. We had designed many tools that beat us in various aspects. This is an invalid analogy.
- bee_rider 3y agoThe internet did change things pretty dramatically. Productivity at information communication tasks just isn’t the entire economy. I think we are massively more productive. Some of the biggest new companies are ad companies (Google, Facebook), or spend a ton of their time designing devices that can’t be modified by their users (Apple, Microsoft). Even old fashioned companies like tractor and train companies have time to waste on preventing users from performing maintenance. And then the economy has leftover effort to jailbreak all this stuff. We’re very productive, we’ve just found room for unlimited zero or negative sum behavior.
- andyg_blog 3y agoReally stellar, well-sourced article that comes across as unbiased as possible. I especially enjoyed the almost-throw-away link to "the bitter lesson" near the end, the gist of which is: "Methods that leverage massive compute to capture intrinsic complexity always outperform humans' attempts to encode that complexity by hand"
- deleted 3y ago[deleted]
- HarHarVeryFunny 3y agoI'm not sure how one can percentage-wise compare scaling and algorithmic advances - per Dwarkesh's prediction that "70% scaling + 30% algorithmic advance" will get us to AGI ?! I think a clearer answer is that scaling alone will certainly NOT get us to AGI. There are some things that are just architecturally missing from current LLMs, and no amount of scaling or data cleaning or emergence will make them magically appear. Some obvious architectural features from top of my list would include: 1) Some sort of planning ahead (cf tree of thought rollouts) which could be implemented in a variety of ways. A simple single-pass feed forward architecture, even a sophisticated one like a transformer, isn't enough. In humans this might be accomplished by some combination of short term memory and the thalamo-cortical feedback loop - iterating on one's perception/reaction to something before "drawing conclusions" (i.e. making predictions) based on it. 2) Online/continual learning so that the model/AGI can learn from it's prediction mistakes via feedback from their consequences, even if that is initially limited to conversational feedback in a ChatGPT setting. To get closer to human-level AGI the model would really need some type of embodiment (either robotic or in a physical simulation virtual word) so that it's actions and feedback go beyond a world of words and let it learn via experimentation how the real world works and responds. You really don't understand the world unless you can touch/poke/feel it, see it, hear it, smell it etc. Reading about it in a book/training set isn't the same. I think any AGI would also benefit from a real short term memory that can be updated and referred to continuously, although "recalculating" it on each token in a long context window does kind of work. In an LLM-based AGI this could just be an internal context, separate from the input context, but otherwise updated and addressed in the same way via attention. It depends too on what one means by AGI - is this implicitly human-like (not just human-level) AGI ? If so then it seems there are a host of other missing features too. Can we really call something AGI if it's missing animal capabilities such as emotion and empathy (roughly = predicting other's emotions, based on having learnt how we would feel in similar circumstances)? You can have some type of intelligence without emotion, but that intelligence won't extend to fully understanding humans and animals, and therefore being able to interact with them in a way we'd consider intelligent and natural. Really we're still a long way from this type of human-like intelligence. What we've got via pre-trained LLMs is more like IBM Watson on steroids - an expert system that would do well on Jeopardy and increasingly well on IQ or SAT tests, and can fool people into thinking it's smarter and more human-like than it really is, just as much simpler systems like Eliza could. The Turing test of "can it fool a human" (in a limited Q&A setting) really doesn't indicate any deeper capability than exactly that ability. It's no indication of intelligence.
- revskill 3y agoNo, human is not that intelligent to generate super intelligent bot in a short time. My estimation is about 200 years in future to have a "human-brain AI" that works. All idea should be treated equally, not based on revenue metrics. If everyone could make a Youtube clone, the revenue should be divided equally to all of creator, that's the way the world should move forward, instead of monopoly. Everything will be suck, forever.
- tw1984 3y agoLLM is going to bring tons of cool applications, but AGI is not an application! You can feed your dog 100,000 times a day, but that won't make it a 1,000kg dog. The whole idea that AGI can be achieved by predicting the next word is just pure marketing nonsense at best.
- xbar 3y agoYes, for some things.
- machiaweliczny 3y agoI think there's a need to separate knowledge from learning algorithm. There's need to be a latent representation of knowledge that models attend to but the way it's done right now (with my limited understanding) doesn't seem to be it. Transformers seems to only attend to previous text in the context but not to the whole knowledge they posses which is obvious limitation IMO. Human brain probably also doesn't attend to whole knowledge but loads something into context so maybe it's fixable without changing architecture. LLMs can work as data extraction already, so one can build some prolog DB and update it as it consumes data. Then translate any logic problems into prolog queries. I want to see this in practice. Similar with usage of logic engines and computation/programs. I also think that RL can come up with better training function for LLMs. In the programming domain for example one could ask LLM to think about all possible test for given code and evaluate them automatically. I was also thinking about using diffusER pattern where programming rules are kinda hardcoded (similar to add/replace/delete but instead algebra on functions/variables). Thats probably not AGI path but could be good for producing programs.
- sgt101 3y ago>Here’s one of the many astounding finds in Microsoft Research’s Sparks of AGI paper. They found that GPT-4 could write the LaTex code to draw a unicorn. a lot of people have tried to replicate this, I have tried. It's very hard to get GPT-4 to draw a unicorn, also asking it to draw an upside down unicorn is even harder.
- midlightdenight 3y agoThe model of GPT-4 those researchers had was not the same that’s available to the public. It’s assumed it was far more capable before alignment training (or whatever it’s called).
- Xelynega 3y agoThat's convenient. Typically one of the markers of good science is reproducability. How can we trust any of the information coming out of these studies if it can't be reproduced?
- sgt101 3y agoAlso - why not make this clear in all the model access documents? Perhaps call it GPT-4P (for Public?) Perhaps also provide other researchers with vetted access. There are a lot of groups trying to evaluate these things systematically - for example "Faith and Fate", "Jumbled thoughts","Emergent abilities are a Mirage" were all very good papers published this year which really highlighted hype in LLM evaluation. Everyone can see that modern LLM's have some great capabilities, the flexibility you can get in an interface by doing intent detection and categorizations using an LLM is great and it is so much easier and quicker than using previous techniques. It's more expensive, but that's improving rapidly. I firmly believe that a new era of great new systems with better interfaces and more functionality will be built on LLM's and other models from this wave of Big Data / Big Model AI, but these are not the precursors of AGI. The problem with the looky looky AGI bunkum show is that it's pulling money into crappy projects that are going to fail hard and this will then stop a lot of money going into projects that could be successful fast. I am seeing the shape of the dotcom boom/bust in what's happening. Microsoft and Intel used dotcom to build and maintain their monopoly position, I think AWS, MS and Google will do the same this time. I think we will see a wave of new companies like Amazon that will "fail" some of them will really fail and disappear, some will half fail like Sun did, but some will go on to build monopolies anew. In the meantime the technology will evolve not for the greater good but instead to serve purposes like advertising distribution that are trivial compared to the benefit we could have seen. Over all we will not capitalise on the potential of what we have for several decades, ironically because of the failures of capitalism. Children will die, wars will be fought but some of us will have nice sweat pants and fun playing paddleball in the sunshine while it all happens. When historians write this up in 100 years they won't really see any of this - they will just see a huge surge of innovation. The dead have no voices...
- lossolo 3y agoA few more interesting papers not mentioned in the article: "Faith and Fate: Limits of Transformers on Compositionality" https://arxiv.org/abs/2305.18654 https://arxiv.org/abs/2305.18654 "Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks": https://arxiv.org/abs/2311.09247 https://arxiv.org/abs/2311.09247 "Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve" https://arxiv.org/abs/2309.13638 https://arxiv.org/abs/2309.13638 "Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models" https://arxiv.org/abs/2311.00871 https://arxiv.org/abs/2311.00871
- cavisne 3y agoIf the size of the internet is really a bottleneck it seems Google is in quite a strong position. Assuming they have effectively a log of the internet, rather than counting the current state of the internet as usable data we should be thinking about the list of diffs that make up the internet. Maybe this ends up like Millenium Management where a key differentiator is having access to deleted datasets.
- nextworddev 3y agoTrue that said market structure changes so rapidly that old datasets aren’t that useful for most strategies
- jeremyjh 3y agoI'd guess at most they have 5x more data, but it is probably nowhere near that, and the article says 100,000x more data is needed.
- nextworddev 3y agoI am in the believer camp for simple reasons: 1) we haven’t even scratched the surface of government led investments into AI, 2) AI itself could probably discover better architectures than transformers (willing to bet heavily on this)
- diggan 3y ago> AI itself could probably discover better architectures than transformers (willing to bet heavily on this) Is there any existing cases of LLMs coming up with novel, useful and namely better architectures? Either related to AI/ML itself or any other field.
- jeremyjh 3y ago> AI itself could probably discover better architectures than transformers The entire subject of the article is concerned with what it will take and how likely it is than an AI will ever will able to generate improvements like this.
- zoogeny 3y agoI was thinking last night about LLMs with respect to Wittgenstein after watching this interesting discussion of his philosophy by John Searle [1]. I think Wittgenstein's ideas are pertinent to the discussion of the relation of language to intelligence (or reasoning in general). I don't meant this in a technical sense (I recall Chomsky mentioning that almost no ideas from Wittgenstein actually have a place in modern linguistics) but from a metaphysical sense (Chomsky also noted that Wittgenstein was one of his formative influences). The video I linked is a worthy introduction and not too long so I recommend it to anyone interested in how language might be the key to intelligence. My personal take, when I see skeptics of LLMs approaching AGI, is that they implicitly reject a Wittgenstein view of metaphysics without actually engaging with it. There is an implicit Cartesian aspect to their world view, where there is either some mental aspect not yet captured by machines (a primitive soul) or some physical process missing (some kind of non-language system). Whenever I read skeptical arguments against LLMs they are not credibly evidence based, nor are they credibly theoretical. They almost always come down to the assumption that language alone isn't sufficient. Wittgenstein was arguing long before LLMs were even a possibility that language wasn't just sufficient, it was inextricably linked to reason. What excites me about scaling LLMs, is we may actually build evidence that supports (or refutes) his metaphysical ideas. 1. https://www.youtube.com/watch?v=v_hQpvQYhOI&ab_channel=PhilosophyOverdose https://www.youtube.com/watch?v=v_hQpvQYhOI&ab_channel=Philo...
- throwup238 3y agoWittgenstein is the perfect lens through which to be skeptical about LLMs and AGI and I think you may be the one not fully engaging with his work. He saw that languages are inseparable from the context in which they are used and that context is much bigger than language itself. Part of learning those language games is experimenting in the real world - interacting with other people, playing language games, and seeing how they react is how we build up our internal dictionaries and innate knowledge. The fidelity of text is simply too low to communicate the amount of information humans use to build up general intelligence. Without the ability to interact with the physical world, LLMs will never be able to reach AGI. It can kind of simulate it to some extent, but it'll never get there LLMs can't even form their own memories, the context has to be explicitly fed back to them.
- bob1029 3y agoI think the "self-play" path is where the scary-powerful AI solutions will emerge. This implies persistence of state and logic that lives external to the LLM. The language model is just one tool. AGI/ASI/whatever will be a system of tools, of which the LLM might be the least complicated one to worry about. In my view, domain modeling, managing state, knowing when to transition between states, techniques for final decision making, consideration for the time domain, and prompt engineering are the real challenges.
- lern_too_spel 3y agoIt's not necessary for the author's purpose of providing more data. We're only training on one kind of input so far, text, from which these models have built some understanding of the world. Humans train on more inputs, and the data to provide those inputs for training a model is readily available, in far larger quantities than individual human brains consume. Data is not the issue.
- Xelynega 3y agoWe're training on text because that's what we're making the model do. It's a fact of neural networks that to train them supervised you need the training data in the expected input for(vector of n thousand preceding tokens for LLMs) with the expected output(the next token for LLMs). "Training them on video" would mean converting the video to a format we can train the llm with, then training the LLM with that info. This would probably be a 1 OOM increase at maximum, if the video transcripts aren't already a part of the training data for gpt.
- lern_too_spel 3y ago> This would probably be a 1 OOM increase at maximum, if the video transcripts aren't already a part of the training data for gpt. Human brains aren't trained on video transcripts, which leave out a lot of information from the video that human brains have a shared understanding of due to training. You would train on video embeddings and predict next embeddings, thus learning physics and other properties of the world and the objects within it and how they interact. This is many more than 5 orders of magnitude more data than the text that today's LLMs are trained on.
- FergusArgyll 3y agoI love Dwarkesh, his podcast is phenomenal. But every article/post of this kind immediately begs the question; What is AGI? I have yet to hear even a decent standard. It always seems like I'm reading Greek philosophers struggling with things they have almost no understanding of and just throwing out the wildest theories. Honestly, it raised my opinion of them seeing how hard it is to reason about things which we have no grasp of.
- LunicLynx 3y agoHere is an idea. Maybe the most optimized neural network is the brain. Computation to energy consumption ratio. So essentially the way doing this in silicon is just pointless. There must be a reason we can do so much while consuming so little, and then again struggling with other tasks. What is the success if we build a machine that consumes just heaps of energy and then is as bad in maths as us?
- sweezyjeezy 3y agoThere's a couple of false dichotomies here - to say that because we're "more optimised" we must be the most optimised. Our brains are optimised well for certain things, sure, but computers are far more efficient at e.g. crunching numbers than we are - to say that there's no success in a machine that can't currently beat us at math - this year has already proven that false
- daepyonim 3y agoThat google paper really gave a bunch of idiots a whole bunch of ammunition. The four color theorem was proven by a machine long ago and it was worthless, about as worthless as what funsearch did!