16 ms·
One thing that makes me skeptical is the lack of specific labels on the first two accuracy graphs. They just say it's a "log scale", without giving even a ballp
by ARandumGuy 2y ago
One thing that makes me skeptical is the lack of specific labels on the first two accuracy graphs. They just say it's a "log scale", without giving even a ballpark on the amount of time it took.
Did the 80% accuracy test results take 10 seconds of compute? 10 minutes? 10 hours? 10 days? It's impossible to say with the data they've given us.
The coding section indicates "ten hours to solve six challenging algorithmic problems", but it's not clear to me if that's tied to the graphs at the beginning of the article.
The article contains a lot of facts and figures, which is good! But it doesn't inspire confidence that the authors chose to obfuscate the data in the first two graphs in the article. Maybe I'm wrong, but this reads a lot like they're cherry picking the data that makes them look good, while hiding the data that doesn't look very good.
- wmf 2y agoPeople have been celebrating the fact that tokens got 100x cheaper and now here's a new system that will use 100x more tokens.
- cowpig 2y agoIsn't that part of the point?
- jsheard 2y agoAlso you now have to pay for tokens you can't see, and just have to trust that OpenAI is using them economically.
- brookst 2y agoToken count was always an approximation of value. This may help break that silly idea.
- regularfry 2y agoI don't think it's much good as an approximation of value, but it seems ok as an approximation of cost.
- brookst 2y agoFair, cost and value are only loosely related. Trying to price based on cost always turns into a mess.
- regularfry 2y agoIts what you do when you're a commodity.
- seydor 2y agoIf it 's reasoning correctly, it shouldnt need a lot of tokens because you don't need to correct it. You only need to ask it to solve nuclear fusion once.
- deleted 2y ago[deleted]
- msp26 2y agoHave you seen how long the CoT was for the example. It's incredibly verbose.
- slt2021 2y agoI find there is an educational benefit in verbosity, it helps to teach user to think like a machine
- legel 2y agoWhich is why it is incredibly depressing that OpenAI will not publish the raw chain of thought. “Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. We acknowledge this decision has disadvantages. We strive to partially make up for it by teaching the model to reproduce any useful ideas from the chain of thought in the answer. For the o1 model series we show a model-generated summary of the chain of thought.”
- slt2021 2y agomaybe they will enable to show CoT for a limited uses, like 5 prompts a day for Premium users, or for Enterprise users with agreement not to steal CoT or something like that. if OpenAI sees this - please allow users to see CoT for a few prompts per day, or add it to Azure OpenAI for Enterprise customers with legal clauses not to steal CoT
- from-nibly 2y ago
- esafak 2y ago...while providing a significant advance. That's a good problem.
- mewpmewp2 2y agoIsn't that part of developing a new tech?
- zamadatix 2y agoThe new thing that can do more at the "ceiling" price doesn't remove your ability to still use the 100x cheaper tokens for the things that were doable on that version.
- digging 2y agoThat exact pattern is always true of technological advance. Even for a pretty broad definition of technology. I'm not sure if it's perfectly described by the name "induced demand" but it's basically the same thing.
- energy123 2y agoIt does dispel this idea that we are going to be flooded with too many GPUs.
- olalonde 2y ago"People have been celebrating the fact that RAM got 100x cheaper and now here's a new system that will use 100x more RAM."
- anticensor 2y agoKnown as Wirth's law.
- packetlost 2y agoWhen one axis is on log scale and the other is linear with the plot points appearing linear-ish, doesn't it mean there's a roughly exponential relationship between the two axis?
- ARandumGuy 2y agoIt'd be more accurate to call it a logarithmic relationship, since compute time is our input variable. Which itself is a bit concerning, as that implies that modest gains in accuracy require exponentially more compute time. In either case, that still doesn't excuse not labeling your axis. Taking 10 seconds vs 10 days to get 80% accuracy implies radically different things on how developed this technology is, and how viable it is for real world applications. Which isn't to say a model that takes 10 days to get an 80% accurate result can't be useful. There are absolutely use cases where that could represent a significant improvement on what's currently available. But the fact that they're obfuscating this fairly basic statistic doesn't inspire confidence.
- packetlost 2y ago> Which itself is a bit concerning, as that implies that modest gains in accuracy require exponentially more compute time This is more of what I was getting at. I agree they should label the axis regardless, but I think the scaling relationship is interesting (or rather, concerning) on its own.
- KK7NIL 2y agoThe absolute time depends on hardware, optimizations, exact model, etc; it's not a very meaningful number to quantify the reinforcement technique they've developed, but it is very useful to estimate their training hardware and other proprietary information.
- j_maffe 2y agoIt's not about the literally quantity/value, it's about the order of growth of output vs input. Hardware and optimizations don't really change that.
- jstummbillig 2y agoI don't think it's worth any debate. You can simply find out how it does for you, now(-ish, rolling out). In contrast: Gemini Ultra, the best, non-existent Google Model for the past few month now, that people nonetheless are happy to extrapolate excitement over.
- swatcoder 2y ago> Did the 80% accuracy test results take 10 seconds of compute? 10 minutes? 10 hours? 10 days? It's impossible to say with the data they've given us. The gist of the answer is hiding in plain sight: it took so long, on an exponential cost function, that they couldn't afford to explore any further. The better their max demonstrated accuracy, the more impressive this report is. So why stop where they did? Why omit actual clock times or some cost proxy for it from the report? Obviously, it's because continuing was impractical and because those times/costs were already so large that they'd unfavorably affect how people respond to this report
- jsheard 2y agoSee also: them still sitting on Sora seven months after announcing it. They've never given any indication whatsoever of how much compute it uses, so it may be impossible to release in its current state without charging an exorbitant amount of money per generation. We do know from people who have used it that it takes between 10 and 20 minutes to render a shot, but how much hardware is being tied up during that time is a mystery.
- ben_w 2y agoCould well be. It's also entirely possible they are simply sincere about their fear it may be used to influence the upcoming US election. Plenty of people (me included) are sincerely concerned about the way even mere still image generators can drown out the truth with a flood of good-enough-at-first-glance fiction.
- jsheard 2y agoIf they were sincere about that concern then they wouldn't build it at all, if it's ever made available to the public then it will eventually be available during an election. It's not like the 2024 US presidential election is the end of history.
- e1g 2y agoThe risk is not “interfering with the US elections”, but “being on the front page of everything as the only AI company interfering with US elections”. This would destroy their peacocking around AGI/alignment while raising billions from pension funds. OpenAI is in a very precarious position. Maybe they could survive that hit in four years, but it would be fatal today. No unforced errors.
- bjornsing 2y agoSo now it’s a question of how fast the AGI will run? :)
- oblio 2y agoIt's fine, it will only need to be powered by a black hole to run.
- exe34 2y agothe first one anyway. after that it will find more efficient ways. we did, afterall.
- wahnfrieden 2y agoit's not obviously achievable. for instance, we don't have the compute power to simulate cellular organisms of much complexity, and have not found efficiencies to scale that
- HeatrayEnjoyer 2y agoHuman level AGI only requires 20 watts
- skywhopper 2y agoYeah, this hiding of the details is a huge red flag to me. Even if it takes 10 days, it’s still impressive! But if they’re afraid to say that, it tells me they are more concerned about selling the hype than building a quality product.
- bluecoconut 2y agoSuper hand-waving rough estimate: Going off of five points of reference / examples that sorta all point in the same direction. 1. looks like they scale up by about ~100-200 on the x axis when showing that test time result. 2. Based on the o1-mini post [1], there's an "inference cost" where you can see GPT-4o and GPT-4o mini as dots in the bottom corner, haha (you can extract X values, ive done so below) 3. There's a video showing the "speed" in the chat ui (3s vs. 30s) 4. The pricing page [2] 5. On their API docs about reasoning, they quantify "reasoning tokens" [3] First, from the original plot, we have roughly 2 orders of magnitude to cover (~100-200x) Next, from the cost plots: super handwaving guess, but since 5.77 / 0.32 = ~18, and the relative cost for gpt-4o vs gpt-4o-mini is ~20-30, this roughly lines up. This implies that o1 costs ~1000x the cost than gpt-4o-mini for inference (not due to model cost, just due to the raw number of chain of thought tokens it produces). So, my first "statement", is that I trust the "Math performance vs Inference Cost" plot on the o1-mini page to accurately represent "cost" of inference for these benchmark tests. This is now a "cost" relative set of numbers between o1 and 4o models. I'm also going to make an assumption that o1 is roughly the same size as 4o inherently, and then from that and the SVG, roughly going to estimate that they did a "net" decoding of ~100x for the o1 benchmarks in total. (5.77 vs (354.77 - 635)). Next, from the CoT examples they gave us, they actually show the CoT preview where (for the math example) it says "...more lines cut off...", A quick copy paste of what they did include includes ~10k tokens (not sure if copy paste is good though..) and from the cipher text example I got ~5k tokens of CoT, while there are only ~800 in the response. So, this implies that there's a ~10x size of response (decoded tokens) in the examples shown. It's possible that these are "middle of the pack" / "average quality" examples, rather than the "full CoT reasoning decoding" that they claim they use. (eg. from the log scale plot, this would come from the middle, essentially 5k or 10k of tokens of chain of thought). This also feels reasonable, given that they show in their API [3] some limits on the "reasoning_tokens" (that they also count) All together, the CoT examples, pricing page, and reasoning page all imply that reasoning itself can be variable length by about ~100x (2 orders of magnitude), eg. example: 500, 5k (from examples) or up to 65,536 tokens of reasoning output (directly called out as a maximum output token limit). Taking them on their word that "pass@1" is honest, and they are not doing k-ensembles, then I think the only reasonable thing to assume is that they're decoding their CoT for "longer times". Given the roughly ~128k context size limit for the model, I suspect their "top end" of this plot is ~100k tokens of "chain of thought" self-reflection. Finally, at around 100 tokens per second (gpt-4o decoding speed), this leaves my guess for their "benchmark" decoding time at the "top-end" to be between ~16 minutes (full 100k decoding CoT, 1 shot) for a single test-prompt, and ~10 seconds on the low end. So for that X axis on the log scale, my estimate would be: ~3-10 seconds as the bottom X, and then 100-200x that value for the highest value. All together, to answer your question: I think the 80% accuracy result took about ~10-15 minutes to complete. I also believe that the "decoding cost" of o1 model is very close to the decoding cost of 4o, just that it requires many more reasoning tokens to complete. (and then o1-mini is comparable to 4o-mini, but also requiring more reasoning tokens) [1] https://openai.com/index/openai-o1-mini-advancing-cost-efficient-reasoning/ https://openai.com/index/openai-o1-mini-advancing-cost-effic... Extracting "x values" from the SVG: GPT-4o-mini: 0.3175 GPT-4o: 5.7785 o1: (354.7745, 635) o1-preview: (278.257, 325.9455) o1-mini: (66.8655, 147.574) [2] https://openai.com/api/pricing/ https://openai.com/api/pricing/ gpt-4o: $5.00 / 1M input tokens $15.00 / 1M output tokens o1-preview: $15.00 / 1M input tokens $60.00 / 1M output tokens [3] https://platform.openai.com/docs/guides/reasoning https://platform.openai.com/docs/guides/reasoning usage: { total_tokens: 1000, prompt_tokens: 400, completion_tokens: 600, completion_tokens_details: { reasoning_tokens: 500 } }
- worstspotgain 2y agoI don't think it's hard to compute the following: - At the high end, there is a likely nonlinear relationship between answer quality and compute. - We've gotten used to a flat-price model. With AGI-level models, we might have to pay more for more difficult and more important queries. Such is the inherent complexity involved. - All this stuff will get better and cheaper over time, within reason. I'd say let's start by celebrating that machine thinking of this quality is possible at all.
- deleted 2y ago[deleted]
- FridgeSeal 2y agoBold of you to expect transparency and clarity from a company like OpenAI. You wanted reliable readable graphs? Ppphhh, get out of here, but pay of for the CoT tokens you’ll never see on your way out though.