7 ms·
>"the company’s CEO, Sam Altman, says further progress will not come from making models bigger. “I think we're at the end of the era where it's going to be thes
by thunderbird120 3y ago
>"the company’s CEO, Sam Altman, says further progress will not come from making models bigger. “I think we're at the end of the era where it's going to be these, like, giant, giant models,” he told an audience at an event held at MIT late last week. “We'll make them better in other ways.”
So to reiterate, he is not saying that the age of giant AI models is over. Current top-of-the-line AI models are giant and likely will continue to be. However, there's not point in training models you can't actually run economically. Inference costs need to stay grounded which means practical model sizes have a limit. More effort is going to go into making models efficient to run even if it comes at the expense of making them less efficient to train.
- hcks 3y agoYes, but it also tells us that if Altman is honest here, then he doesn’t believe GPT-like models can scale to near level human performances (because even if the cost of compute was 10x or even 100x it would still be economically sound).
- deleted 3y ago[deleted]
- famouswaffles 3y agoNo it doesn't. For one thing they're already at human performance. For another, i don't think you realize how expensive inference can get. Microsoft with no scant amount of available compute is struggling to run gpt-4 such that they're rationing it between subsidiaries while they try to jack up compute. So saying, it would be economically sound if it cost x10 or x100 what it costs now is a joke.
- smeagull 3y agoThis tells me you haven't really stress tested the model. GPT is currently at the stage of "person who is at the meeting, but not really paying attention so you have to call them out". Once GPT is pushed, it scrambles and falls over for most applications. The failure modes range from contradicting itself, making up things for applications that shouldn't allow it, to ignoring prompts, to simply being unable to perform tasks at all.
- dragonwriter 3y agoAre we talking about bare GPT through the UI, or GPT with a framework giving it access to external systems and the ability to store and retrieve data? Because, yeah, “brain in a jar” GPT isn’t enough for most tasks beyond parlor-trick chat, but being used as a brain in a jar isn’t the point.
- smeagull 3y agoWe have given it extensions, and really the extensions do a lot of the work. The tool that judges the style and correctness of the text based on the embedding is doing much of the heavy lifting. GPT essentially handles generating text and dense representations of the text.
- moffkalast 3y agoStill waiting to see those plugins rolled out and actual vector DB integration with GPT 4, then we'll see what it can really do. Seems like the more context you give it the better it does, but the current UI really makes it hard to provide that. Plus the recursive self prompting to improve accuracy.
- quonn 3y agoHow are they at human performance? Almost everything GPT has read on the internet didn‘t even exist 200 years ago and was invented by humans. Heck, even most of the programming it does wasn‘t there 20 years ago. Not every programmer starting from scratch would be brilliant, but many were self taught with very limited resources in the 80s form example and discovered new things from there. GPT cannot do this and is very far from being able to.
- famouswaffles 3y ago>How are they at human performance? Because it performs at least average human level (mostly well above average) on basically every task it's given. "Invest something new" is a nonsensical benchmark for human level intelligence. The vast majority of people have never and will never invent anything new. If your general intelligence test can't be passed by a good chunk of humanity then it's not a general intelligence test unless you want to say most people aren't generally intelligent.
- quonn 3y agoYeah these intelligence tests are not very good. I would argue some programmers do in fact invent something new. Not all of them, but some. Perhaps 10%. Second the point is not whether everyone is by profession an inventor but whether most people can be inventors. And to a degree they can be. I think you underestimate that by a large margin. You can lock people in a room and give them a problem to solve and they will invent a lot if they have the time to do it. GPT will invent nothing right now. It‘s not there yet.
- famouswaffles 3y ago>Yeah these intelligence tests are not very good. Lol Okay >And to a degree they can be. I think you underestimate that by a large margin. Do i? Because i'm not the one making unverifiable claims here. >You can lock people in a room and give them a problem to solve and they will invent a lot if they have the time to do it. If you say so
- 3y ago
- mullingitover 3y agoQuality over quantity. Just building a model with a gazillion parameters isn't indicative of quality, you could easily have garbage parameters with tons of overfitting. It's like megapixel counts in cameras: you might have 2000 gigapixels in your sensor, but that doesn't mean you're going to get great photos out of it if there are other shortcomings in the system.
- sanxiyn 3y agoWhat overfitting? If anything, LLMs suffer from underfitting, not overfitting. Normally, overfitting is characterized by increasing validation loss while training loss is decreasing, and solved by early stopping (stopping before that happens). Effectively, all LLMs are stopped early, so they don't suffer from overfitting at all.
- teruakohatu 3y agoI don't disagree with you, these models may be underfitted, but overfitting is not explicitly defined by val vs. training loss, but rather how closely its output matches training data. If you trained a MLP model where the number of parameters exceeded the data, it would be able to memorize the data and return a zero loss on training data. The larger the models are, the greater chance it memorizes the data, rather than the latent variables or distribution of the data. Early LLMs, GPT2 (circa 2019) for example was definitely overfitting. I would frequently copy and paste output and find a reddit comment with the exact words.
- spaceman_2020 3y agoIs cost really that much of a burden? Intelligence is the single most expensive resource on the planet. Hundreds of individuals have to be born, nurtured, and educated before you might get an exceptional 135+ IQ individual. Every intelligent person is produced at a great societal cost. If you can reduce the cost of replicating a 135 IQ, or heck, even a 115 IQ person to a few thousand dollars, you're beating biology by a massive margin.
- yunwal 3y agoBut we're still nowhere near that, or even near surpassing the skill of an average person at a moderately complex information task, and GPT-4 supposedly took hundreds of millions to train. It also costs a decent amount more to run inference on it vs. 3.5. It probably makes sense to prove the concept that generative AI can be used for lots of real work before scaling that up by another order of magnitude for potentially marginal improvements. Also, just in terms of where to put your effort, if you think another direction (for example, fine-tuning the model to use digital tools, or researching how to predict confidence intervals) is going to have a better chance of success, why focus on scaling more?
- spaceman_2020 3y agoThere are a lot of employees at large tech consultancies that don't really do anything that can't be automated away by even current models. Sprinkle in some more specific training and I can totally see entire divisions at IBM and Accenture and TCS being made redundant. The incentive structures are perversely aligned for this future - the CEO who manages to reduce headcount while increasing revenue is going to be very handsomely rewarded by Wall Street.
- esafak 3y agoHow is that perverse? That is the logical incentive. The perverse one is that middle managers rise by hiring people needlessly and building fiefdoms.
- skyechurch 3y ago
- ldehaan 3y agoI've been training large 65b models on "rent for N hours" systems for less than 1k per customized model. Then fine tuning those to be whatever I want for even cheaper. 2 months since gpt 4. This ride has only just started, fasten your whatevers.
- sailingparrot 3y agoFinetuning cost are nowhere near representative of the cost to pre-train those models. Trying to replicate the quality of GPT-3 from scratch, using all the tricks and training optimizations in the books that are available now but weren't used during GPT-3 actual training, will still cost you north of $500K, and that's being extremly optimistic. GPT-4 level model would be at least 10x this using the same optimism (meaning you are managing to train it for much cheaper than OpenAI). And That's just pure hardware cost, the team you need to actually makes this happen is going to be very expensive as well. edit: To quantify how "extremely optimistic" that is, the very model you are finetuning, which I assume is Llama 65B, would cost around ~$18M to train on google cloud assuming you get a 50% discount on their listed GPU prices (2048 A100 GPUs for 5 months). And that's not even GPT-4 level.
- bagels 3y ago$5M to train GPT-4 is the best investment I've ever seen. I've seen startups waste more money for tremendously smaller impact.
- sailingparrot 3y agoAs I stated in my comment, $5M is assuming you can do a much much better job than OpenAI at optimizing your training, only need to make a single training run, your employees salaries are $0, and you get a clean dataset for essentially free. Real cost is 10-20x that. That's still a good investment though. But the issue is you could very well sink $50M into this endeavour and end up with a model that actually is not really good and gets rendered useless by an open-source model that gets released 1 month later. OpenAI truly has unique expertise in this field that is very, very hard to replicate.