6 ms·
They are admitting[1] that the new model is the gpt2-chatbot that we have seen before[2]. As many highlighted there, the model is not an improvement like GPT3->
by msoad 2y ago
They are admitting[1] that the new model is the gpt2-chatbot that we have seen before[2]. As many highlighted there, the model is not an improvement like GPT3->GPT4. I tested a bunch of programming stuff and it was not that much better.
It's interesting that OpenAI is highlighting the Elo score instead of showing results for many many benchmarks that all models are stuck at 50-70% success.
[1] https://twitter.com/LiamFedus/status/1790064963966370209 https://twitter.com/LiamFedus/status/1790064963966370209
[2] https://news.ycombinator.com/item?id=40199715 https://news.ycombinator.com/item?id=40199715
- modeless 2y ago"not that much better" is extremely impressive, because it's a much smaller and much faster model. Don't worry, GPT-5 is coming and it will be better.
- TIPSIO 2y agoObviously given enough time there will always be better models coming. But I am not convinced it will be another GPT-4 moment. Seems like big focus on tacking together multi-modal clever tricks vs straight better intelligence AI. Hope they prove me wrong!
- kmeisthax 2y agoThe problem with "better intelligence" is that OpenAI is running out of human training data to pillage. Training AI on the output of AI smooths over the data distribution, so all the AIs wind up producing same-y output. So OpenAI stopped scraping text back in 2021 or so - because that's when the open web turned into an ocean of AI piss. I've heard rumors that they've started harvesting closed captions out of YouTube videos to try and make up the shortfall of data, but that seems like a way to stave off the inevitable[0]. Multimodal is another way to stave off the inevitable, because these AI companies already are training multiple models on different piles of information. If you have to train a text model and an image model, why split your training data in half when you could train a combined model on a combined dataset? [0] For starters, most YouTube videos aren't manually captioned, so you're feeding GPT the output of Google's autocaptioning model, so it's going to start learning artifacts of what that model can't process.
- pbhjpbhj 2y ago>harvesting closed captions out of YouTube videos I'd bet a lot of YouTubers are using LLMs to write and/or edit content. So we pass that through a human presentation. Then introduce some errors in the form of transcription. Turn feed the output in as part of a training corpus ... we plateaued real quick. It seems like it's hard to get past a level of human intelligence at which there's a large enough corpus of training data or trainers? Anyone know of any papers on breaking this limit to push machine learning models to super-human intelligence levels?
- pixl97 2y agoIf a model is average human intelligence in pretty much everything, is that super-human or not? Simply put, we as individuals aren't average at everything, we have what we're good at and a great many things we're not. We average out by looking at broad population trends. That's why most of us in the modern age spend a lot of time on specialization for whatever we work in. Which brings the likely next place for data. A Manna (the story) like data collection program where companies hoover up everything they can on their above average employees till we're to the point most models are well above the human average in most categories.
- WhitneyLand 2y agoWhy do you think they’re using Google auto-captioning? I would expect they’re using their own t2s which is still a model but way better quality and potentially customizable to better suit their needs
- llm_trw 2y ago>[0] For starters, most YouTube videos aren't manually captioned, so you're feeding GPT the output of Google's autocaptioning model, so it's going to start learning artifacts of what that model can't process. Whisper models are better than anything google has. In fact the higher quality whisper models are better than humans when it comes to transcribing text with punctuation.
- marvin 2y agoAt some point, algorithms for reasoning and long-term planning will be figured out. Data won’t be the holy grail forever, and neither will asymptotically approaching human performance in all domains.
- mupuff1234 2y agoAnd how can one be so sure of that? Seems to me that performance is converging and we might not see a significant jump until we have another breakthrough.
- lionkor 2y ago[flagged]
- scarmig 2y agoYeah. There are lots of things we can do with existing capabilities, but in terms of progressing beyond them all of the frontier models seem like they're a hair's breadth from each other. That is not what one would predict if LLMs had a much higher ceiling than we are currently at. I'll reserve judgment until we see GPT5, but if it becomes just a matter of who best can monetize existing capabilities, OAI isn't the best positioned.
- andrepd 2y agoExactly. People like to point at the start of a logistic curve and go "behold! an exponential"
- diego_sandoval 2y ago> Seems to me that performance is converging It doesn't seem that way to me. But even if it did, video generation also seemed kind of stagnant before Sora. In general, I think The Bitter Lesson is the biggest factor at play here, and compute power is not stagnating.
- deleted 2y ago[deleted]
- drawnwren 2y agoComputer power is not stagnating, but the availability of training data is. It's not like there's a second stackoverflow or reddit to scrape.
- talldayo 2y agoChalmers: "GPT-5? A vastly-improved model that somehow reduces the compute overhead while providing better answers with the same hardware architecture? At this time of year? In this kind of market?" Skinner: "Yes." Chalmers: "May I see it?" Skinner: "No."
- pwdisswordfishc 2y agoIncidentally, this dialogue works equally well, if not better, with David Chalmers versus B.F. Skinner, as with the Simpsons characters.
- AaronFriel 2y agoIt has only been a little over one year since GPT-4 was announced, and it was at the time the largest and most expensive model ever trained. It might still be. Perhaps it's worth taking a beat and looking at the incredible progress in that year, and acknowledge that whatever's next is probably "still cooking". Even Meta is still baking their 400B parameter model.
- bamboozled 2y agoLegit love progress
- 1024core 2y agoAs Altman said (paraphrasing): GPT-4 is the _worst_ model you will ever have to deal with in your life (or something to that effect).
- andrepd 2y agoI will believe it when I see it. People like to point at the first part of a logistic curve and go "behold! an exponential".
- nwienert 2y agoAh yes my favorite was the early covid numbers, some of the "smartest" people in the SF techie scene were daily on Facebook thought-leadering about how 40% of people were about to die in the likely case.
- moomoo11 2y agoI really hope GPT5 is good. GPT4 sucks at programming.
- verdverm 2y agoLook to a specialized model instead of a general purpose one
- moomoo11 2y agoAny suggestions? Thanks I have tried Phind and anything beyond mega junior tier questions it suffers as well and gives bad answers.
- verdverm 2y agoIt will be a system, not a single model, and will depend on what programming task you want to perform probably need routers, RAG, and reranking I think there is a role for LLM + deterministic code gen as well (https://github.com/hofstadter-io/hof/blob/_dev/flow/chat/prompts/dm.cue https://github.com/hofstadter-io/hof/blob/_dev/flow/chat/pro...)
- moomoo11 2y agoInteresting. I was hoping for something with a UI like chat gpt or phind. Something that I can just use as easily as copilot. Unfortunately every single one sucks. Or maybe that's just how programming is - its easy at the surface/ice berg level and below is just massive amounts of complexity. Then again, I'm not doing menial stuff so maybe I'm just expecting too much.
- verdverm 2y agoI think a more IDE native experience is better than a chat UI I don't want to have to copy & paste between applications, just let me highlight some sections and then run some LLM operation on it i.e. a VS Code extension with keyboard shortcuts
- littlestymaar 2y agoI don't think a bigger model would make sense for OpenAI: it's much more important for them that they keep driving inference coat down, because there's no viable business model if they don't. Improving the instruction tuning, the RLHF step, increase the training size, work on multilingual capabilities, etc. make sense as a way to improve quality, but I think increasing model size doesn't. Being able to advertize a big breakthrough may make sense in terms of marketing, but I don't believe it's going to happen for two reasons: - you don't release intermediate steps when you want to be able to advertise big gains, because it raises the baseline and reduce the effectiveness of your ”big gains” in terms of marketing. - I don't think they would benefit in an arm race with Meta, trying to keeping a significant edge. Meta is likely to be able to catch-up eventually on performance, but they are not so much of a threat in terms of business. Focusing on keeping a performance edge instead of making their business viable would be a strategic blunder.
- jononor 2y agoWhat is OpenAI business model if their models are second-best? Why would people pay them and not Meta/Google/Microsoft - who can afford to sell at very low margins, since they have existing very profitable businesses that keeps them afloat.
- littlestymaar 2y agoThat's the question OpenAI needs to find an answer to if they want to end up viable. They have the brand recognition (for ChatGPT) and that's a good start, but that's not enough. Providing a best in class user experience (which seems to be their focus now, with multimodality), a way to lock down their customers in some kind of walled garden, building some kind of network effect (what they tried with their marketplace for community-built “GPTs” last fall but I'm not sure it's working), something else? At the end of the day they have no technological moat, so they'll need to build a business one, or perish. For most tasks, pretty much every models from their competitors is more than good enough already, and it's only going to get worse as everyone improves. Being marginally better on 2% of tasks isn't going to be enough.
- deleted 2y ago[deleted]
- cube2222 2y agoI think the live demo that happened on the livestream is best to get a feel for this model[0]. I don't really care whether it's stronger than gpt-4-turbo or not. The direct real-time video and audio capabilities are absolutely magical and stunning. The responses in voice mode are now instantaneous, you can interrupt the model, you can talk to it while showing it a video, and it understands (and uses) intonation and emotion. Really, just watch the live demo. I linked directly to where it starts. Importantly, this makes the interaction a lot more "human-like". [0]: https://youtu.be/DQacCB9tDaw?t=557 https://youtu.be/DQacCB9tDaw?t=557
- gabiruh 2y agoIt's weird that the "airplane mode" seems to be ON on the phone during the entire presentation.
- arthurcolle 2y agoThis was on purpose - they connected it to the internet via a USB-C cable it appears, for consistent internet instead of having it switch WiFi Probably some kinks there they are working out
- _flux 2y agoAnd eliminate the change of some prankster affecting the demo by attacking the wifi.
- OJFord 2y ago> Probably some kinks there they are working out Or just a good idea for a live demo on a congested network/environment with a lot of media present, at least one live video stream (the one we're watching the recording of), etc. At least that's how I understood it, not that they had a problem with it (consistently or under regular conditions, or specific to their app).
- hbn 2y agoThat's very common practice for live demos. To avoid situations like this: https://www.youtube.com/watch?v=6lqfRx61BUg https://www.youtube.com/watch?v=6lqfRx61BUg
- dragonwriter 2y ago> As many highlighted there, the model is not an improvement like GPT3->GPT4. The improvements they seem to be hyping are in multimodality and speed (also price – half that of GPT-4 Turbo – though that’s their choice and could be promotional, but I expect it’s at least in part, like speed, a consequence of greater efficiency), not so much producing better output for the same pure-text inputs.
- kybercore 2y agothe model scores 60 points higher in lmsys than the best gpt 4 turbo model from april, that's still a pretty significant jump in text capability
- aixpert 2y agouseless anecdata but I find the new model very frustrating, often completely ignoring what I say in follow up queries. it's giving me serious Siri vibes (text input in web version) maybe it's programmed to completely ignore swearing but how could I not swear after it gave me repeatedly info about you.com when I try to address it in second person
- lossolo 2y agoI agree. I tried a few programming problems that, let's say, seem to be out of the distribution of their training data and which GPT4 failed to solve before. The model couldn't find a similar pattern and failed to solve them again. What's interesting is that one of these problems were solved by Opus, which seems to indicate that the majority of progress in the last months should be attributed to the quality/source of the training data.
- avereveard 2y agoI tested a few use cases in the chat, and it's not particularly more intelligent but they seem to have solved laziness. I had to categorize my expenses to do some budgeting for the family, and in gpt 4 I had to go ten in ten, confirm the suggested category, download the file, took two days as I was constantly hitting the limit. gpt4o did most of the grunth work, then commincated anomalies in bulk, asked for suggestion for these, and provided a downloadable link in two answers, calling the code interpreter mulitple times, and working toward the goal on it's own. and the prompt wasn't a monstrosity, and it wasn't even that good, it was just one line "I need help to categorize these expenses" and off it went. hope it won't get enshittified like turbo, because this finally feels as great as 3.5 was for goal seeking.
- ozzydave 2y agoHeh - I'm using ChatGPT for the same thing! Works 10X better than Rocket Money, which was supposed to be an improvement on Mint but meh.
- deleted 2y ago[deleted]
- jameshart 2y agoI think this comment is easily misread as implying that this GPT4o model is based on some old GPT2 chatbot - that’s very much not what you meant to say, though. This model has been being tested under a code name of ‘gpt2-chatbot’ but it is very much a new GPT4+-level model, with new multimodal capabilities - but apparently some impressive work around inference speed. Highlighting so people don’t get the impression this is just OpenAI slapping a new label on something a generation out of date.
- vitorgrs 2y agoThey are admitting that is the im-also-a-good-gpt2-chatbot. There was 3.... Don't ask me why. The "gpt2-chatbot" was the worst of the three.