3 ms·
Finally a scaling wall? This is apparently (based on pricing) using about an order of magnitude more compute, and is only maybe 10% more intelligent. Ideally De
by erulabs 2y ago
Finally a scaling wall? This is apparently (based on pricing) using about an order of magnitude more compute, and is only maybe 10% more intelligent. Ideally DeepSeeks optimizations help bring the costs way down, but do any AI researchers want to comment on if this changes the overall shape of the scaling curve?
- fpgaminer 2y agoSeems on par with the existing scaling curve. If I had to speculate, this model would have been an internal-only model, but they're releasing it for PR. An optimized version with 99% of the performance for 1/10th the cost will come out later.
- j_maffe 2y agoThis is the shittiest PR move I've seen since the AI trend started.
- Workaccount2 2y agoAt least so far it's coding performance is bad, but from what I have seen it's writing abilities are totally insane. It doesn't read like AI output anymore.
- j_bum 2y agoAny examples you’d be willing to share?
- chippiewill 2y agoThey have examples in the announcement post. It does a better job of understanding intent in the question which helps it give an informal rather than essay style response where appropriate.
- j_maffe 2y agoI wouldn't call that "too insane." As others have pointed out, you can get similar results from fine-tuning the RLHF.
- csomar 2y agoWe have hit that wall almost 2 years ago with gpt-4. There was clearly no scaling as gpt-4 was already decently smart and if you got x2 smarter you’ll be more capable than anything on the market today. All models doing today (R1 and friends; and Claude) are trying to optimize this local maxima toward generating more useful responses (ie: code when it comes to Claude). AI, at its current form, is a Deep Seek of compressed knowledge in a 30-50gb of interconnected data. I think we’ll look at this as trying to train networks on corpus of data and expecting them to have a hold of reality. Our brains are trained on “reality” which is not the “real” reality as your vision is limited to the visible spectrum. But if you want a network to behave like a human then maybe give him what a human see. There is also the possibility that there is a physical limit to intelligence. I don’t see any elephants doing PhDs and the smartest of humans are just a small configuration away from insanity.
- killerstorm 2y agoIt depends on how you compare. On a subset of tasks I'm interested in, it's 10x more intelligent than GPT-4. (Note that GPT-4 was in many ways better than 4o.) It's not a coding champion, but it knows A LOT of stuff, excellent common sense, top quality writing. For me it's like "deep research lite". I found OpenAI Deep research excellent, but GPT-4.5 might in many cases beat it.
- _1tem 2y ago> On a subset of tasks I'm interested in, it's 10x more intelligent than GPT-4. Very intriguing. Care to share an example?
- vbezhenar 2y agoThe price is 2x from GPT4. So probably not a decimal order of magnitude.