9 ms·
Seems like a wild claim to make without any examples of gpt models which are bigger and no demonstrably better.
by fauxpause_ 3y ago
Seems like a wild claim to make without any examples of gpt models which are bigger and no demonstrably better.
- xipix 3y agoPerhaps (a) there do exist bigger models that weren't better or (b) this model isn't better than somewhat smaller ones. Perhaps the CEO has seen diminishing returns.
- fauxpause_ 3y agoSure. It is possible there is evidence that is not shared. It’s dumb to assume this is the case though.
- mnky9800n 3y agoor like a curve of model complexity versus results or whatever showing it asymptotically approaches whatever. actually there was a great paper from microsoft research from like 2001 on spam filtering where they demonstrated that model complexity necessary for spam filtering went down as the size of the data set went up. That paper, which i can't seem to find now, had a big impact on me as a researcher because it so clearly demonstrated that small data is usually bad data and sophisticated models are sometimes solving problems will small data sets instead of problems with data. of course this paper came out the year friedman published his gradient boosting paper, i think random forest also was only recently published then as well (i think there is a paper from 1996 about RF and briemans two cultures paper came out this year where he discusses RF i believe), and this is a decade before gpu based neural networks. So times are different now. But actually i think the big difference is these days i probably ask chatgpt to write the boiler plate code for a gradient boosted model that takes data out of a relational database instead of writing it myself.
- nomel 3y ago> model complexity necessary for spam filtering went down as the size of the data set went up My naive conclusion in that this means there are still massive gains to be had, since, for example, something like ChatGPT is just text, and the phrase "a picture is worth a thousand words" seems incredibly accurate, from my perspective. There's an incredible amount of non-text data out there still. Especially technical data. Is there any merit to this belief?
- jacobr1 3y agoYes. One of the frontiers of current research seems to be multi-modal models.
- cjbprime 3y agoGPT-4 is actually multi-modal, not text. ChatGPT does not yet expose image upsubmission to it. But it's part of how the model was trained already.
- throwaway2037 3y agoExcellent points in your post. You wrote: There's an incredible amount of non-text data out there still. Especially technical data. "Especially technical data." What does this part mean? Initially, I thought you meant things like images and video, but now I am confused.
- sseagull 3y agoThey might mean numerical data like scientific simulation data, sensor data, polling data, statistics, etc.
- nomel 3y agoSchematics (of any sort), block diagrams, general spatial awareness (including anything related to puzzle pieces/packing, like circuit layout), most physics problems involving force diagrams, anything mechanical, etc. The text representation of any of these is ludicrously more complex than simple images. If you sit someone down, that works in one of these fields, you'll quickly see the limitations. It'll try to represent the concepts as text, with ascii art or some "attempt" at an ascii file format that can be used to draw, and its "reasoning" about these things is much more limited. I think most people interacting with GPT are in a text-only (and especially programming) bubble.
- Ambix 3y ago> "a picture is worth a thousand words" and it might be opposite for the GPT models actually. it's just easier for humans to grasp the bunch of knowledge with one eyes sight, but usually most of useful information might be represented with just of bunch of words and machines are to scan through the millions of words in an instant.
- deleted 3y ago[deleted]
- greatwave1 3y agoI believe this is the paper which you are referring to: https://aclanthology.org/P01-1005.pdf https://aclanthology.org/P01-1005.pdf ("Scaling to Very Very Large Corpora for Natural Language Disambiguation" by Michele Banko and Eric Brill, Microsoft Research, 2001)
- mnky9800n 3y agoomg i have been searching forever for this. THANK YOU.
- summerlight 3y agohttps://twitter.com/SmokeAwayyy/status/1646670920214536193 https://twitter.com/SmokeAwayyy/status/1646670920214536193 Sam explicitly said that there won't be GPT-5 in the near future, which is pretty clear evidence unless he's blatantly lying in public speaking.
- kjellsbells 3y agoWell, "no GPT-5" isn't the same as saying "no new trained model", especially in the realm of marketing. Welcome to "GPT 2024" could be his next slogan.
- thehumanmeat 3y agoThat is one AI CEO out of 10,000. Just because OpenAI may not be interested in a larger model in the short term doesn't mean nobody else won't pursue it.
- not2b 3y agoOthers might pursue a smaller model that works as well as a larger model. If that can be done, whoever does it can very effectively compete on price/performance. It seems that to assume otherwise (the only way to improve is to get bigger) is to assume that OpenAI already has found the optimal architecture. That seems unlikely.
- imatworkyo 3y agoGoogle recently said transformers wouldn't work ...
- hackerlight 3y agoIt's not a wild claim when you have empirically well-validated scaling laws which make this very prediction.
- fauxpause_ 3y agoThat’s a generous usage of the term “empirically well-validated” and “law”
- mensetmanusman 3y agoBetter on which axis? Do you want an AI that takes one hour to respond to? Some would for certain fields, but getting something fast and cheap is going to be hard now that Moore’s law is over.
- AlecSchueler 3y agoDon't we all agree that GPT4 is "better" than GPT3? How are we evaluating that if the axis is such a mystery. Yeah maybe we can't quantify it like I can't tell you one writer is better than another in a quantitative but we can both still read their work and come to an understanding.
- nwlieb 3y agoThe runtime is quadratic for a given context size, although it seems like there is some progress on this front https://gwern.net/note/attention https://gwern.net/note/attention
- MichaelZuo 3y agoExponential scaling for a presumable GPT-5 suggests it's response time will be unusably long for the vast majority of use cases, and probably cost multiple dollars USD per query. Not to mention there doesn't actually exist enough English text data in the world to even double GPT-4's training set.
- epups 3y agoCompute will also scale exponentially in coming years. The data source limitation seems to be a harder barrier, I think many companies are experimenting with AI generated content for training at this point.
- MichaelZuo 3y ago> Compute will also scale exponentially in coming years. Cost per transistor scaling has already plateaued or perhaps even inverted with TSMC's latest and greatest. And the new chips, even after 25 layers of EUV lithography, more than doubling the previous record, and an extra year of fine tuning, has total SRAM size scaling of -5% and logic scaling of -42%. These are numbers verified by experienced semi people.
- Barrin92 3y agoBoth ChatGPT 3.5 and 4 literally fail the question: "What is the third letter in the third word of this sentence" When you've spent 100 million on training the thing and it fails on 1st grade ordinality I think it's fair to say you may not be on the right path
- mcbuilder 3y agoBut yet it can understand a json data schema from example and write javascript to interact with a library that I fed it and asked it to understand. Yes, I know its limitations, but it can also surprise me.
- bccdee 3y agoThe problem with basic programming questions like this is that there are a million elementary online tutorials for doing this or that with a json schema. "Simple programming questions based on commonly-used technology" are something it's been very heavily trained on.
- m00x 3y agoThese specific questions are very hard for an AI to answer. Just like humans suck at calculating numbers, AIs aren't good at sparse self-questioning. They're extremely good at other tasks, like taking very difficult tests that require a lot of knowledge storage. It's pretty obvious they're on the right path for what they're trying to achieve.
- drippingfist 3y ago"The third word of this sentence is "the," and its third letter is "e." - GPT-4
- foobiekr 3y agoUnfortunately it seems clear that openai trains gptX on common test questions. They still fail novel ones.
- bhouston 3y agoI suspect you are right. We may be stuck at the gpt4 sizes for a bit just because of hardware costs though. As they get bigger it costs too much to run them until our hardware becomes more optimal for these large models at 4 bits or so. I think the YouTube videos is going to be the next big training set. A transformer trained on all text and all of YouTube will be killer amazing at so much. I bet it can understand locomotion and balance and body control from YouTube. I wonder if TPUs, like Google's Tensor chip, will beat out GPUs when it comes to image/video based training?
- TaylorAlexander 3y ago> I wonder if TPUs, like Google's Tensor chip, will beat out GPUs when it comes to image/video based training? One of the OpenAI guys was talking about this. He said the specific technology does not matter, it is just a cost line item. They don't need to have the best chip tech available as long as they have enough money. That said I am curious if anyone else can really comment on this. It seems like as we get to very large and expensive models we will produce more and more specialized technology.
- bhouston 3y ago> They don't need to have the best chip tech available as long as they have enough money. That sounds like someone who is "Blitzscaling." Costs do not matter in those cases, just acquiring customers and marketshare. But for the rest of us, who will see benefits but are not trying to win a $100B market, we will cost optimize.
- TaylorAlexander 3y agoYes, agreed. I would like to run large models at home without serious expense.
- daydream 3y agoWhether or not cost matters much depends on your perspective. If you’re OpenAI and GPT4 is just a step on the way to AGI, and you can amortize that huge cost over the hundreds of millions in revenue you’re gonna pull in from subscriptions and API use… then sure you’re probably not very cost sensitive. It could be 20% cheaper or 50% more expensive, whatever, it’s so good your customers will use it at a wide range of costs. And you have truckloads of money from Microsoft anyways. If you’re a company or a developer trying to build a feature, whole new product, or an entire company on top of GPT then that cost matters a whole lot. The difference between $0.06 and $0.006 per turn could be infeasible vs. shippable. If you’re trying to compete with OpenAI then you’re probably doing everything possible to reduce that training cost. So, whether or not it matters - it really depends.
- blast 3y agoOpenAI may have those internally though.
- kccqzy 3y agoIf OpenAI's CEO is making this claim, don't you think he has internal data backing up the claim?
- wetpaws 3y agoHe might or might not, this is a point of the claim.
- smeyer 3y agoNo. I don't always assume that just because a CEO makes a public statement they have internal data backing up the claim. Sometimes they do! Other times, they have data but are misinterpreting it or missing something, but it's impossible to tell if the data is just internal. Other times they're making a statement without data based on their personal beliefs. Other times, they don't even think the statement is true but are saying it for messaging, marketing, or communication reasons! Like the previous commenter, I'd be much more confident an asymptote was reached if it was being demonstrated publicly.
- zamnos 3y agoOnly OpenAI and its CEO know the full details on GPT-4's sizes so that's entirely possible. But since it's an internal secret, there's nothing compelling him to tell the truth. For all we know, he has internal data backing up the opposite of the claim but is making this claim so as to discourage potential competitors from spending the money training an even bigger and competitive ML model. Sending potential competitors off on a wide goose chase that, when pushed, he can just say "oh our internal data (that no one outside of a trusted few have seen) said otherwise". I have no idea if sama is such a person, but you must admit that the possibility exists.
- imtringued 3y ago[flagged]
- worthless-trash 3y agoYou'd hope so, but unless people put their evidence in public, it could simply be a tool to manipulate the public's expectations or competitors behavior. I'll get downvoted for this, apples previous CEO was consistently inaccurate about company innovation and performance numbers.
- deleted 3y ago[deleted]