29 ms·
Learning to Reason with LLMs
- MrRobotics 2y agoThis is the sort of reasoning needed to solve the ARC AGI benchmark.
- fsflover 2y agoDupe: https://news.ycombinator.com/item?id=41523050 https://news.ycombinator.com/item?id=41523050
- farresito 2y agoDamn, that looks like a big jump.
- deisteve 2y agoso o1 seems like it has real measurable edge, crushing it in every single metric, i mean 1673 elo is insane, and 89th percentile is like a whole different league, and it looks like it's not just a one off either, it's consistently performing way better than gpt-4o across all the datasets, even in the ones where gpt-4o was already doing pretty well, like math and mmlu, o1 is just taking it to the next level, and the fact that it's not even showing up in some of the metrics, like mmmu and mathvista, just makes it look even more impressive, i mean what's going on with gpt-4o, is it just a total dud or what, and btw what's the deal with the preview model, is that like a beta version or something, and how does it compare to o1, is it like a stepping stone to o1 or something, and btw has anyone tried to dig into the actual performance of o1, like what's it doing differently, is it just a matter of more training data or is there something more going on, and btw what's the plan for o1, is it going to be released to the public or is it just going to be some internal tool or something
- farresito 2y ago> like what's it doing differently, is it just a matter of more training data or is there something more going on Well, the model doesn't start with "GPT", so maybe they have come up with something better.
- rvnx 2y agoIt sounds like GPT-4o with a long CoT prompt no ?
- spaceman_2020 2y ago1673 ELO is wild If its actually true in practice, I sincerely cannot imagine a scenario where it would be cheaper to hire actual junior or mid-tier developers (keyword: "developers", not architects or engineers). 1,673 ELO should be able to build very complex, scalable apps with some guidance
- deisteve 2y agocurrently my workflow is generate some code, run it, if it doesn't run i tell LLM what I expected, it will then produce code and I frequently tell it how to reason about the problem. with O1 being in the 89th percentile would mean it should be able to think at junior to intermediate level with very strong consistency. i dont think people in the comments realize the implication of this. previously LLMs were able to only "pattern match" but now its able to evaluate itself (with some guidance ofc) essentially, steering the software into depth of edge cases and reason about it in a way that feels natural to us. currently I'm copying and pasting stuff and notifying LLM the results but once O1 is available its going to significantly lower that frequency. For example, I expect it to self evaluate the code its generate and think at higher levels. ex) oooh looks like this user shouldn't be able to escalate privileges in this case because it would lead to security issues or it could conflict with the code i generated 3 steps ago, i'll fix it myself.
- usaar333 2y agoI'm not sure how well codeforces percentiles correlate to software engineering ability. Looking at all the data, it still isn't. Key notes: 1. AlphaCode 2 was already at 1650 last year. 2. SWE-bench verified under an agent has jumped from 33.2% to 35.8% under this model (which doesn't really matter). The full model is at 41.4% which still isn't a game changer either. 3. It's not handling open ended questions much better than gpt-4o.
- deisteve 2y agoi think you are right now actually initially i got excited but now i think OpenAI pulled the hype card again to seem relevant as they struggle to be profitable Claude on the other hand has been fantastic and seems to do similar reasoning behind the scenes with RL
- dinobones 2y agoGenerating more "think out loud" tokens and hiding them from the user... Idk if I'm "feeling the AGI" if I'm being honest. Also... telling that they choose to benchmark against CodeForces rather than SWE-bench.
- thelastparadise 2y agoWhy not? Isn't that basically what humans do? Sit there and think for a while before answering, going down different branches/chains of thought?
- dinobones 2y agoThis new approach is showing: 1) The "bitter lesson" may not be true, and there is a fundamental limit to transformer intelligence. 2) The "bitter lesson" is true, and there just isn't enough data/compute/energy to train AGI. All the cognition should be happening inside the transformer. Attention is all you need. The possible cognition and reasoning occurring "inside" in high dimensions is much more advanced than any possible cognition that you output into text tokens. This feels like a sidequest/hack on what was otherwise a promising path to AGI.
- gradus_ad 2y agoDoes that mean human intelligence is cheapened when you talk out a problem to yourself? Or when you write down steps solving a problem? It's the exact same thing here.
- barrell 2y agolol come on it’s not the exact same thing. At best this is like gagging yourself while you talk about it then engaging yourself when you say the answer. And that presupposing LLMs are thinking in, your words, exactly the same way as humans. At best it maybe vaguely resembles thinking
- 2y ago
- lloydatkinson 2y agoWhat's with this how many r's in a strawberry thing I keep seeing?
- arresin 2y agoIt’s a common LLM riddle. Apparently many fail to give the right answer.
- seydor 2y agoSomebody please ask o1 to solve it
- lloydatkinson 2y agoThe link shows it solving it
- dr_quacksworth 2y agoLLM are bad at answering that question because inputs are tokenized.
- swalsh 2y agoModels don't really predict the next word, they predict the next token. Strawberry is made up of multiple tokens, and the model doesn't truely understand the characters in it... so it tends to struggle.
- runjake 2y agoThis became something of a meme. https://community.openai.com/t/incorrect-count-of-r-characters-in-the-word-strawberry/829618 https://community.openai.com/t/incorrect-count-of-r-characte...
- andrewla 2y agoWhat's amazing is that given how LLMs receive input data (as tokenized streams, as other commenters have pointed out) it's remarkable that it can ever answer this question correctly.
- deleted 2y ago[deleted]
- valine 2y agoThe model performance is driven by chain of thought, but they will not be providing chain of thought responses to the user for various reasons including competitive advantage. After the release of GPT4 it became very common to fine-tune non-OpenAI models on GPT4 output. I’d say OpenAI is rightly concerned that fine-tuning on chain of thought responses from this model would allow for quicker reproduction of their results. This forces everyone else to reproduce it the hard way. It’s sad news for open weight models but an understandable decision.
- tomtom1337 2y agoCan you explain what you mean by this?
- tomduncalf 2y agoI think they mean that you won’t be able to see the “thinking”/“reasoning” part of the model’s output, even though you pay for it. If you could see that, you might be able to infer better how these models reason and replicate it as a competitor
- ffreire 2y agoYou can see an example of the Chain of Thought in the post, it's quite extensive. Presumably they don't want to release this so that it is raw and unfiltered and can better monitor for cases of manipulation or deviation from training. What GP is also referring to is explicitly stated in the post: they also aren't release the CoT for competitive reasons, so that presumably competitors like Anthropic are unable to use the CoT to train their own frontier models.
- gwd 2y ago> Presumably they don't want to release this so that it is raw and unfiltered and can better monitor for cases of manipulation or deviation from training. My take was: 1. A genuine, un-RLHF'd "chain of thought" might contain things that shouldn't be told to the user. E.g., it might at some point think to itself, "One way to make an explosive would be to mix $X and $Y" or "It seems like they might be able to poison the person". 2. They want the "Chain of Thought" as much as possible to reflect the actual reasoning that the model is using; in part so that they can understand what the model is actually thinking. They fear that if they RLHF the chain of thought, the model will self-censor in a way which undermines their ability to see what it's really thinking 3. So, they RLHF only the final output, not the CoT, letting the CoT be as frank within itself as any human; and post-filter the CoT for the user.
- deleted 2y ago[deleted]
- p1esk 2y agoafter weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users.
- zaptrem 2y agoThis also makes them less useful because I can’t just click stop generation when they make a logical error re: coding.
- deleted 2y ago[deleted]
- sterlind 2y ago"Open"AI is such a comically ironic name at this point.
- swalsh 2y agoWe're not going to give you training data... for a better user experience.
- deleted 2y ago[deleted]
- scosman 2y agoSaying "competitive advantage" so directly is surprising. There must be some magic sauce here for guiding LLMs which boosts performance. They must think inspecting a reasonable number of chains would allow others to replicate it. They call GPT 4 a model. But we don't know if it's really a system that builds in a ton of best practices and secret tactics: prompt expansion, guided CoT, etc. Dalle was transparent that it automated re-generating the prompts, adding missing details prior to generation. This and a lot more could all be running under the hood here.
- skywhopper 2y agoNo direct indication of what “maximum test time” means, but if I’m reading the obscured language properly, the best scores on standardized tests were generated across a thousand samples with supplemental help provided. Obviously, I hope everyone takes what any company says about the capabilities of its own software with a huge grain of salt. But it seems particularly called for here.
- immortal3 2y agoHonestly, it doesn't matter for the end user if there are more tokens generated between the AI reply and human message. This is like getting rid of AI wrappers for specific tasks. If the jump in accuracy is actual, then for all practical purposes, we have a sufficiently capable AI which has the potential to boost productivity at the largest scale in human history.
- Lalabadie 2y agoIt starts to matter if the compute time is 10-100 fold, as the provider needs to bill for it. Of course, that's assuming it's not priced for market acquisition funded by a huge operational deficit, which is a rarely safe to conclude with AI right now.
- skywhopper 2y agoGiven that their compute-time vs accuracy charts labeled the compute time axis as logarithmic would worry me greatly about this aspect.
- breck 2y ago[flagged]
- deisteve 2y agoyeah this is kinda cool i guess but 808 elo is still pretty bad for a model that can supposedly code like a human, i mean 11th percentile is like barely scraping by, and what even is the point of simulating codeforces if youre just gonna make a model that can barely compete with a decent amateur, and btw what kind of contest allows 10 submissions, thats not how codeforces works, and what about the time limits and memory limits and all that jazz, did they even simulate those, and btw how did they even get the elo ratings, is it just some arbitrary number they pulled out of their butt, and what about the model that got 1807 elo, is that even a real model or just some cherry picked result, and btw what does it even mean to "perform better than 93% of competitors" when the competition is a bunch of humans who are all over the place in terms of skill, like what even is the baseline for comparison edit: i got confused with the Codeforce. it is indeed zero shot and O1 is potentially something very new I hope Anthropic and others will follow suit any type of reasoning capability i'll take it !
- deleted 2y ago[deleted]
- qt31415926 2y ago808 ELO was for GPT-4o. I would suggest re-reading more carefully
- deisteve 2y agoyou are right i read the charts wrong. O1 has significant lead over GPT-4o in the zero shot examples honestly im spooked
- catchnear4321 2y agooh wow, something you can roughly model as a diy in a base model. so impressive. yawn. at least NVDA should benefit. i guess.
- apsec112 2y agoIf there's a way to do something like this with Llama I'd love to hear about it (not being sarcastic)
- catchnear4321 2y agonurture the model have patience and a couple bash scripts
- apsec112 2y agoBut what does that mean? I can't do "pip install nurture" or "pip install patience". I can generate a bunch of answers and take the consensus, but we've been able to do that for years. I can do fine-tuning or DPO, but on what?
- catchnear4321 2y agoyou want instructions on how to compete with OpenAI? go play more, your priorities and focus on it being work are making you think this to be harder than it is, and the models can even tell you this. you don’t have to like the answer, but take it seriously, and you might come back and like it quite a bit. you have to have patience because you likely wont have scale - but it is not just patience with the response time.
- gliiics 2y agoCongrats to OpenAI for yet another product that has nothing to do with the word "open"
- sk11001 2y agoAnd Apple's product line this year? Phones. Nothing to do with fruit. Almost 50 years of lying to people. Names should mean something!
- achrono 2y agoDid Apple start their company by saying they will be selling apples?
- sk11001 2y agoWhat's the statement that OpenAI are making today which you think they're violating? There very well could be one and if there is, it would make sense to talk about it. But arguments like "you wrote $x in a blog post when you founded your company" or "this is what the word in your name means" are infantile.
- trash_cat 2y agoIt is open in the sense that everyone can use it.
- bionhoward 2y agoNot people working on AI or those who would like to train AI on their logs
- oblio 2y agoIf they would have launched it with Oracle DB style licensing their company would have been dead in 1 year.
- Hizonner 2y agoOnly people who exactly share OpenAI's concepts of what "alignment" and "safety" should mean can use it to its full potential.
- RandomThoughts3 2y ago> “Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users.” Trust us, we have your best intention in mind. I’m still impressed by how astonishingly impossible to like and root for OpenAI is for a company with such an innovative product.
- TheAceOfHearts 2y agoKinda disappointed that they're hiding the thought process. Hopefully the open source community will figure out how to effectively match and replicate what OpenAI is doing. I wonder how far we are from having a model that can correctly solve a word soup search problem directly from just a prompt and input image. It seems like the crossword example is close. For a word search it would require turning the image into an internal grid representation, prepare the list of words, and do a search. I'd be interested in seeing if this model can already solve the word grid search problem if you give it the correct representation as an input.
- zozbot234 2y ago> Hopefully the open source community will figure out how to effectively match and replicate what OpenAI is doing. No need for that, there is a Reflection 70B model that does the exact same thing - with chains of thought being separated from the "final answer" via custom 'tag' tokens.
- TheAceOfHearts 2y agoWasn't this the model that was proven to have been faking their benchmarks recently? Or am I thinking of a different model?
- brokensegue 2y agoyes. it was fake
- zozbot234 2y agoSome reported benchmarks do seem to be rather dubious, but the 70B model itself is quite real. Sample output: $ ollama run reflection:70b-q4_0 >>> hello <thinking> To respond to "hello", I'll need to consider several factors: 1. The user's intent: They're likely saying hello as a greeting. 2. Appropriate responses: Common ways to respond to "hello" are: - Hello back - Hi - Hey 3. Formality level: Since this is an AI response, I'll aim for a friendly but professional tone. <reflection> The approach of responding with a standard greeting seems appropriate in this context. It acknowledges the user's hello and provides a polite response. </reflection> Given these considerations, the most suitable response would be to echo "hello" back to the user. </thinking> <output> Hello! </output>
- crakenzak 2y ago> we are releasing an early version of this model, OpenAI o1-preview, for immediate use in ChatGPT Awesome!
- dinobones 2y agoI am interpreting "immediate use in ChatGPT" the same way advanced voice mode was promised "in the next few weeks." Probably 1% of users will get access to it, with a 20/message a day rate limit. Until early next year.
- nilsherzig 2y agoRate limit is 30 a week for the big one and 50 for the small one
- afruitpie 2y agoRate limited to 30 messages per week for ChatGPT Plus subscribers at launch: https://openai.com/index/introducing-openai-o1-preview/ https://openai.com/index/introducing-openai-o1-preview/
- benterix 2y agoRead "immediate" in "immediate use" in the same way as "open" in "OpenAI".
- Ninjinka 2y agoSomeone give this model an IQ test stat.
- adverbly 2y agoYou're kidding right? The tests they gave it are probably better tests than IQ tests at determining actually useful problem solving skills...
- Vecr 2y agoIt can't do large portions of the parts of an IQ test (not multi-modal). Otherwise I think it's essentially superhuman, modulo tokenization issues (please start running byte-by-byte or at least come up with a better tokenizer).
- modeless 2y ago> We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). Wow. So we can expect scaling to continue after all. Hyperscalers feeling pretty good about their big bets right now. Jensen is smiling. This is the most important thing. Performance today matters less than the scaling laws. I think everyone has been waiting for the next release just trying to figure out what the future will look like. This is good evidence that we are on the path to AGI.
- ffsm8 2y agoIt'd be interesting for sure if true. Gotta remember that this is a marketing post though, let's wait a few months and see if its actually true. Things are definitely interesting, wherever these techniques will get us AGI or not
- XCSme 2y agoNvidia stock go brrr...
- deleted 2y ago[deleted]
- acchow 2y agoEven when we start to plateau on direct LLM performance, we can still get significant jumps by stacking LLMs together or putting a cluster of them together.
- gizmo 2y agoMicrosoft, Google, Facebook have all said in recent weeks that they fully expect their AI datacenter spend to accelerate. They are effectively all-in on AI. Demand for nvidia chips is effectively infinite.
- seydor 2y agoUntil the first LLM that can improve itself occurs. Then $NVDA tanks
- 2y ago
- breck 2y agoI LOVE the long list of contributions. It looks like the credits from a Christoper Nolan film. So many people involved. Nice care to create a nice looking credits page. A practice worth copying. https://openai.com/openai-o1-contributions/ https://openai.com/openai-o1-contributions/
- rfw300 2y agoA lot of skepticism here, but these are astonishing results! People should realize we’re reaching the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”. And as anyone who’s spent time using Claude 3.5 Sonnet / GPT-4o can attest, these things really are useful and smart! (And, if these results hold up, O1 is much, much smarter.) This is a nerve-wracking time to be a knowledge worker for sure.
- bigstrat2003 2y agoI cannot, in fact, attest that they are useful and smart. LLMs remain a fun toy for me, not something that actually produces useful results.
- pdntspa 2y agoI have been deploying useful code from LLMs right and left over the last several months. They are a significant force accelerator for programmers if you know how to prompt them well.
- fiddlerwoaroof 2y agoWe’ll see if this is a good idea when we start having millions of lines of LLM-written legacy code. My experience maintaining such code so far has been very bad: accidentally quadratic algorithms; subtly wrong code that looks right; and un-idiomatic use of programming language features.
- deisteve 2y agoah i see so you're saying that LLM-written code is already showing signs of being a maintenance nightmare, and that's a reason to be skeptical about its adoption. But isn't that just a classic case of 'we've always done it this way' thinking? legacy code is a problem regardless of who wrote it. Humans have been writing suboptimal, hard-to-maintain code for decades. At least with LLMs, we have the opportunity to design and implement better coding standards and review processes from the start. let's be real, most of the code written by humans is not exactly a paragon of elegance and maintainability either. I've seen my fair share of 'accidentally quadratic algorithms' and 'subtly wrong code that looks right' written by humans. At least with LLMs, we can identify and address these issues more systematically. As for 'un-idiomatic use of programming language features', isn't that just a matter of training the LLM on a more diverse set of coding styles and idioms? It's not like humans have a monopoly on good coding practices. So, instead of throwing up our hands, why not try to address these issues head-on and see if we can create a better future for software development?
- hobofan 2y agoThat naming scheme... Will the next model be named "1k", so that the subsequent models will be named "4o1k", and we can all go into retirement?
- p1esk 2y agoMore like you will need to dip into your 401k fund early to pay for it after they raise the prices.
- deleted 2y ago[deleted]
- notamy 2y agohttps://openai.com/index/introducing-openai-o1-preview/ https://openai.com/index/introducing-openai-o1-preview/ > ChatGPT Plus and Team users will be able to access o1 models in ChatGPT starting today. Both o1-preview and o1-mini can be selected manually in the model picker, and at launch, weekly rate limits will be 30 messages for o1-preview and 50 for o1-mini. We are working to increase those rates and enable ChatGPT to automatically choose the right model for a given prompt. Weekly? Holy crap, how expensive is it to run is this model?
- HPMOR 2y agoIt's probably running several lines of COT. I imagine, each single message you send is probably at __least__ 10x to the actual model. So in reality it's like 300 messages, and honestly it's probably 100x, given how constrained they're being with usage.
- theLiminator 2y agoAnyone know when o1 access in ChatGPT will be open?
- tedsanders 2y agoRolling out over the next few hours to Plus users.
- narrator 2y agoThe human brain uses 20 watts, so yeah we figured out a way to run better than human brain computation by using many orders of magnitude more power. At some point we'll need to reject exponential power usage for more computation. This is one of those interesting civilizational level problems. There's still a lack of recognition that we aren't going to be able to compute all we want to, like we did in the pre-LLM days.
- seydor 2y agowe ll ask it to redesign itself for low power usage
- minimaxir 2y ago> Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. What? I agree people who typically use the free ChatGPT webapp won't care about raw chain-of-thoughts, but OpenAI is opening an API endpoint for the O1 model and downstream developers very very much care about chain-of-thoughts/the entire pipeline for debugging and refinement. I suspect "competitive advantage" is the primary driver here, but that just gives competitors like Anthropic an oppertunity.
- Hizonner 2y agoThey they've taken at least some of the hobbles off for the chain of thought, so the chain of thought will also include stuff like "I shouldn't say <forbidden thing they don't want it to say>".
- npn 2y ago"Open"AI. Should be ClosedAI instead.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- cal85 2y agoSounds great, but so does their "new flagship model that can reason across audio, vision, and text in real time" announced in May. [0] [0] https://openai.com/index/hello-gpt-4o/ https://openai.com/index/hello-gpt-4o/
- mickeystreicher 2y agoYep, all these AI announcements from big companies feel like promises for the future rather than immediate solutions. I miss the days when you could actually use a product right after it was announced, instead of waiting for some indefinite "coming soon."
- valval 2y agoAs an entrepreneur, I do this often. In order to sleep better at night, I explain to myself that it’s somewhat harmless to give teasers about future content releases. If someone buys my product based on future promises or speculation, they’re investing into the development and my company’s future.
- apsec112 2y agoThis one [o1/Strawberry] is available. I have it, though it's limited to 30 messages/week in ChatGPT Plus.
- thelastparadise 2y agoWouldn't this introduce new economics into the LLM market? I.e. if the "thinking loop" budget is parameterized, users might pay more (much more) to spend more compute on a particular question/prompt.
- minimaxir 2y agoDepends on how OpenAI prices it. Given the need for chain-of-thoughts, and that would be budgeted as output, the new model will not be cheap nor fast. EDIT: Pricing is out and it is definitely not teneable unless you really really have a use case for it.
- sroussey 2y agoYes, and note the large price increase
- p1esk 2y agoDo people see the new models in the web interface? Mine still shows the old models (I'm a paid subscriber).
- hi 2y ago> "o1 models are currently in beta - The o1 models are currently in beta with limited features. Access is limited to developers in tier 5 (check your usage tier here), with low rate limits (20 RPM). We are working on adding more features, increasing rate limits, and expanding access to more developers in the coming weeks!" https://platform.openai.com/docs/guides/rate-limits/usage-tiers?context=tier-five https://platform.openai.com/docs/guides/rate-limits/usage-ti...
- p1esk 2y agoI'm talking about web interface, not API. Should be available now, since they said "immediate release".
- hi 2y agohttps://chatgpt.com/?model=o1-preview https://chatgpt.com/?model=o1-preview --> defaults back to 4o
- MillionOClock 2y agoSame for me here
- zamadatix 2y agoIt may take a bit to appear in your account (and by a bit I mean I had to fiddle around a while, try logging out/in, etc for a bit) but it appears for me and many others as normal Plus users in the web.
- mewpmewp2 2y agoI have tier 5, but I'm not seeing that model. Also API call gives an error that it doesn't exist or I do not have access.
- rvz 2y agoWon't be surprised to see all these hand-picked results and extreme expectations to collapse under scenarios involving highly safety critical and complex demanding tasks requiring a definite focus on detail with lots of awareness, which what they haven't shown yet. So let's not jump straight into conclusions with these hand-picked scenarios marketed to us and be very skeptical. Not quite there yet with being able to replace truck drivers and pilots for self-autonomous navigation in transportation, aerospace or even mechanical engineering tasks, but it certainly has the capability in replacing both typical junior and senior software engineers in a world considering to do more with less software engineers needed. But yet, the race to zero will surely bankrupt millions of startups along the way. Even if the monthly cost of this AI can easily be as much as a Bloomberg terminal to offset the hundreds of billions of dollars thrown into training it and costing the entire earth.
- jazzyjackson 2y agoMy concern with AI always has been it will outrun the juniors and taper off before replacing folks with 10, 20 years of experience And as they retire there's no economic incentive to train juniors up, so when the AI starts fucking up the important things there will be no one who actually knows how it works I've heard this already from amtrak workers, track allocation was automated a long time ago, but there used to be people who could recognize when the computer made a mistake, now there's no one who has done the job manually enough to correct it.
- djoldman 2y ago> THERE ARE THREE R'S IN STRAWBERRY Ha! This is a nice easteregg.
- vessenes 2y agoI appreciated that, too! FWIW, I could get Claude 3.5 to tell me how many rs a python program would tell you there are in strawberry. It didn't like it, though.
- mewpmewp2 2y agoI was able to get GPT-4o to calculate characters properly using following prompt: """ how many R's are in strawberry? use the following method to calculate - for example Os in Brocolli. B - 0 R - 0 O - 1 C - 1 O - 2 L - 2 L - 2 I - 2 Where you keep track after each time you find one character by character """ And also later I asked it to only provide a number if the count increased. This also worked well with longer sentences.
- zamadatix 2y agoAt that point just ask it "Use python to count the number of O's in Broccoli". At least then it's still the one figuring out the "smarts" needed to solve the problem instead of being pure execution.
- mewpmewp2 2y agoDo you think you'll have python always available when you go to the store and need to calculate how much change you should get?
- zamadatix 2y agoI'm not sure if your making a joke about the teachers who used to say "you won't have a calculator in your pocket" and now we have cell phones or are not aware that ChatGPT runs the generated Python for you in a built in environment as part of the response. I lean towards the former but in case anyone else strolling by hasn't tried this before: User: Use python to count the number of O's in Broccoli ChatGPT: Analyzing... The word "Broccoli" contains 2 'O's. <button to show code> User: Use python to multiply that by the square root of 20424.2332423 ChatGPT: Analyzing... The result of multiplying the number of 'O's in "Broccoli" by the square root of 20424.2332423 is approximately 285.83.
- orbital-decay 2y agoWait, are they comparing 4o without CoT and o1 with built-in CoT?
- persedes 2y agoyeah was wondering what 4o with a CoT in the prompt would look like.
- kickofline 2y agoLLM performance, recently, seemingly hit the top of the S-curve. It remains to be seen if this is the next leap forward or just the rest of that curve.
- billconan 2y agoI will pay if O1 can become my college level math tutor.
- andrewla 2y agoThis is something that people have toyed with to improve the quality of LLM responses. Often instructing the LLM to "think about" a problem before giving the answer will greatly improve the quality of response. For example, if you ask it how many letters are in the correctly spelled version of a misspelled word, it will first give the correct spelling, and then the number (which is often correct). But if you instruct it to only give the number the accuracy is greatly reduced. I like the idea too that they turbocharged it by taking the limits off during the "thinking" state -- so if an LLM wants to think about horrible racist things or how to build bombs or other things that RLHF filters out that's fine so long as it isn't reflected in the final answer.
- scotty79 2y ago> I like the idea too that they turbocharged it by taking the limits off during the "thinking" state They also specifically trained the model to do that thinking out loud.
- jazzyjackson 2y agoDang, I just payed out for Kagi Assistant. Using Claude 3 Opus I noticed it performs <thinking> and <result> while browsing the web for me. I don't guess that's a change in the model for doing reasoning.
- levzzz 2y ago[dead]
- cyanf 2y ago> 30 messages per week
- k2xl 2y agoPricing page updated for O1 API costs. https://openai.com/api/pricing/ https://openai.com/api/pricing/ $15.00 / 1M input tokens $60.00 / 1M output tokens For o1 preview Approx 3x the price of gpt4o. o1-mini $3.00 / 1M input tokens $12.00 / 1M output tokens About 60% of the cost of gpt4o. Much more expensive than gpt4o-mini. Curious on the performance/tokens per second for these new massive models.
- logicchains 2y agoI guess they'd also charge for the chain of thought tokens, of which there may be many, even if users can't see them.
- fraboniface 2y agoThat would be very bad product design. My understanding is that the model itself is similar to GPT4o in architecture but trained and used differently. So the 5x relative increase in output token cost likely already accounts for hidden tokens and additional compute.
- natrys 2y ago> While reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens. https://platform.openai.com/docs/guides/reasoning https://platform.openai.com/docs/guides/reasoning So yeah, it is in fact very bad product design. I hope Llama catches up in a couple of months.
- AkelaWolf 2y agoMost likely the model has similar size compared to the original gpt4, which also has similar price.
- deleted 2y ago[deleted]
- yunohn 2y agoThe generated chain of thought for their example is incredibly long! The style is kind of similar to how a human might reason, but it's also redundant and messy at various points. I hope future models will be able to optimize this further, otherwise it'll lead to exponential increases in cost.
- flockonus 2y agoAre we ready yet to admit Turing test has been passed?
- paxys 2y agoThe Turing Test (which involves fooling a human into thinking they are talking to another human rather than a computer) has been routinely passed by very rudimentary "AI" since as early as 1991. It has no relevance today.
- adverbly 2y agoThis is only true for some situations. In some test conditions it has not been passed. I can't remember the exact name, but there used to be a competition where PhD level participants blindly chat for several minutes with each other and are incentivized to discover who is a bot and who is a human. I can't remember if they still run it, but that bar has never been passed from what I recall.
- rvz 2y agoLLMs have already beaten the Turing test. It's useless to use it when OpenAI and others are aiming for 'AGI'. So you need a new Turing test adapted for AGI or a totally different one to test for AGI rather than the standard obsolete Turing test.
- riku_iki 2y ago> LLMs have already beaten the Turing test. I am wondering where this happened? In some limited scope? Because if you plug LLM into some call center role for example, it will fall apart pretty quickly.
- TillE 2y agoExtremely basic agency would be required to pass the Turing test as intended. Like, the ability to ask a new unrelated question without being prompted. Of course you can fake this, but then you're not testing the LLM as an AI, you're testing a dumb system you rigged up to create the appearance of an AI.
- patapong 2y agoVery interesting. I guess this is the strawberry model that was rumoured. I am a bit surprised that this does not beat GPT-4o for personal writing tasks. My expectations would be that a model that is better at one thing is better across the board. But I suppose writing is not a task that generally requires "reasoning steps", and may also be difficult to evaluate objectively.
- markonen 2y agoIn the performance tests they said they used "consensus among 64 samples" and "re-ranking 1000 samples with a learned scoring function" for the best results. If they did something similar for these human evaluations, rather than just use the single sample, you could see how that would be horrible for personal writing.
- janalsncm 2y agoI don’t understand how that is generalizable. I’m not going to be able to train a scoring function for any arbitrary task I need to do. In many cases the problem of ranking is at least as hard as generating a response in the first place.
- afro88 2y agoThe solution of the cipher example problem also strongly hints at this: "there are three r's in strawberry"
- patapong 2y agoConfirmed by the verge: https://www.theverge.com/2024/9/12/24242439/openai-o1-model-reasoning-strawberry-chatgpt https://www.theverge.com/2024/9/12/24242439/openai-o1-model-...
- slashdave 2y ago> My expectations would be that a model that is better at one thing is better across the board. No, it's the opposite. This is simply a function of resources applied during training.
- Hansenq 2y agoReading through the Chain of Thought for the provided Cipher example (go to the example, click "Show Chain of Thought") is kind of crazy...it literally spells out every thinking step that someone would go through mentally in their head to figure out the cipher (even useless ones like "Hmm"!). It really seems like slowing down and writing down the logic it's using and reasoning over that makes it better at logic, similar to how you're taught to do so in school.
- Jasper_ 2y ago> Average:18/2=9 > 9 corresponds to 'i'(9='i') > But 'i' is 9, so that seems off by 1. Still seems bad at counting, as ever.
- deleted 2y ago[deleted]
- dymk 2y agoThe next line is it catching its own mistake, and noting i = 9.
- PoignardAzur 2y agoIt's interesting that it makes that mistake, but then catches it a few lines later. A common complaint about LLMs is that once they make a mistake, they will keep making it and write the rest of their completion under the assumption that everything before was correct. Even if they've been RLHF to take human feedback into account and the human points out the mistake, their answer is "Certainly! Here's the corrected version" and then they write something that makes the same mistake. So it's interesting that this model does something that appears to be self-correction.
- afro88 2y agoSeeing the "hmmm", "perfect!" etc. one can easily imagine the kind of training data that humans created for this. Being told to literally speak their mind as they work out complex problems.
- hi 2y agoBUG: https://openai.com/index/reasoning-in-gpt/ https://openai.com/index/reasoning-in-gpt/ > o1 models are currently in beta - The o1 models are currently in beta with limited features. Access is limited to developers in tier 5 (check your usage tier here), with low rate limits (20 RPM). We are working on adding more features, increasing rate limits, and expanding access to more developers in the coming weeks! https://platform.openai.com/docs/guides/reasoning/reasoning https://platform.openai.com/docs/guides/reasoning/reasoning
- cptcobalt 2y agoI'm in Tier 4, and not far off from Tier 5. The docs aren't quite transparent enough to show that if I buy credits if I'll be bumped up to Tier 5, or if I actually have to use enough credits to get into Tier 5. Edit, w/ real time follow up: Prior to buying the credits, I saw O1-preview in the Tier 5 model list as a Tier 4 user. I bought credits to bump to Tier 5—not much, I'd have gotten there before the end of the year. The OpenAI website now shows I'm in Tier 5, but O1-preview is not in the Tier 5 model list for me anymore. So sneaky of them!
- hi 2y agohttps://news.ycombinator.com/item?id=41523070#41523525 https://news.ycombinator.com/item?id=41523070#41523525
- paxys 2y ago2018 - gpt1 2019 - gpt2 2020 - gpt3 2022 - gpt3.5 2023 - gpt4 2023 - gpt4-turbo 2024 - gpt-4o 2024 - o1 Did OpenAI hire Google's product marketing team in recent years?
- Infinity315 2y agoNo, this is just how Microsoft names things.
- logicchains 2y agoWe'll know the Microsoft takeover is complete when OpenAI release Ai.net.
- randomdata 2y agoGPT# forthcoming. You heard it here first.
- adverbly 2y agoMakes sense to me actually. This is a different product. It doesn't respond instantly. It fundamentally makes sense to separate these two products in the AI space. There will obviously be a speed vs quality trade-off with a variety of products across the spectrum over time. LLMs respond way too fast to actually be expected to produce the maximum possible quality of a response to complex queries.
- ilaksh 2y agoOne of them would have been named gpt-5, but people forget what an absolute panic there was about gpt-5 for quite a few people. That caused Altman to reassure people they would not release 'gpt-5' any time soon. The funny thing is, after a certain amount of time, the gpt-5 panic eventually morphed into people basically begging for gpt-5. But he already said he wouldn't release something called 'gpt-5'. Another funny thing is, just because he didn't name any of them 'gpt-5', everyone assumes that there is something called 'gpt-5' that has been in the works and still is not released.
- deleted 2y ago[deleted]
- aktuel 2y agoIf I pay for the chain of thought, I want to see the chain of thought. Simple. How would I know if it happened at all? Trust OpenAI? LOL
- baq 2y agoEasy solution - don't pay!
- aktuel 2y agoThat's a chain of thought, right there!
- 93po 2y agohow do you know it isn't some guy typing responses to you when you use openAI?
- aktuel 2y agoWell if they are paying real people to answer my questions I would call that a pretty good deal. That's exactly my point. As a user I don't care how they come up with it. That's not my problem. I just care about the content. If I pay a human for logical reasoning, train of thought type of stuff, I expect them to lay it out for me. Not just give me the conclusion, but how they came to it.
- zamadatix 2y agoYou could say the same thing about use of any product which isn't fully open sourced "how do I know this service really saved my files redundantly if I can't see the disks it's stored on?". It's definitely an opinion on approach though I'm not sure how practically applicable it is. The real irony is how closed "Open"AI is... but that's not news.
- asadm 2y agoI am not up-to-speed on CoT side but is this similar to how perplexity does it ie. - generate a plan - execute the steps in plan (search internet, program this part, see if it is compilable) each step is a separate gpt inference with added context from previous steps. is O1 same? or does it do all this in a single inference run?
- bbstats 2y agoFinally, a Claude competitor!
- gradus_ad 2y agoInteresting sequence from the Cipher CoT: Third pair: 'dn' to 'i' 'd'=4, 'n'=14 Sum:4+14=18 Average:18/2=9 9 corresponds to 'i'(9='i') But 'i' is 9, so that seems off by 1. So perhaps we need to think carefully about letters. Wait, 18/2=9, 9 corresponds to 'I' So this works. ----- This looks like recovery from a hallucination. Is it realistic to expect CoT to be able to recover from hallucinations this quickly?
- trash_cat 2y agoHow do you mean quickly? It probably will take a while for it to output the final answer as it needs to re-prompt itself. It won't be as fast as 4o.
- bigyikes 2y ago4o could already recover from hallucination in a limited capacity. I’ve seen it, mid-reply say things like “Actually, that’s wrong, let me try again.”
- NightlyDev 2y agoDid it hallucinate? I haven't looked at it, but lowercase i and uppercase i is not the same number if you're getting the number from ascii
- machiaweliczny 2y agoIn general if hallucination ratio is 2% can't it be reduced to 0.04% by running twice or sth like this. I think they should try establishing the facts from different angles and this probably would work fine to minimize hallucinations. But if this was that simple somebody would already do it...
- shthed 2y agoSeems like a huge waste of tokens for it to try to work all this out manually, as soon as it came up with the decipher algorithm it should realise it can write some code to execute.
- arresin 2y ago> Unless otherwise specified, we evaluated o1 on the maximal test-time compute setting. Maximal test time is the maximum amount of time spent doing the “Chain of Thought” “reasoning”. So that’s what these results are based on. The caveat is that in the graphs they show that for each increase in test-time performance, the (wall) time / compute goes up exponentially. So there is a potentially interesting play here. They can honestly boast these amazing results (it’s the same model after all) yet the actual product may have a lower order of magnitude of “test-time” and not be as good.
- logicchains 2y agoSurprising that at run time it needs an exponential increase in thinking to achieved a linear increase in output quality. I suppose it's due to diminishing returns to adding more and more thought.
- HarHarVeryFunny 2y agoThe exponential increase is presumably because of the branching factor of the tree of thoughts. Think of a binary tree who's number of leaf nodes doubles (= exponential growth) at each level. It's not too surprising that the corresponding increase in quality is only linear - how much difference in quality would you expect between the best, say, 10 word answer to a question, and the best 11 word answer ? It'll be interesting to see what they charge for this. An exponential increase in thinking time means an exponential increase in FLOPs/dollars.
- alwa 2y agoI interpreted it to suggest that the product might include a user-facing “maximum test time” knob. Generating problem sets for kids? You might only need or want a basic level of introspection, even though you like the flavor of this model’s personality over that of its predecessors. Problem worth thinking long, hard, and expensively about? Turn that knob up to 11, and you’ll get a better-quality answer with no human-in-the-loop coaching or trial-and-error involved. You’ll just get your answer in timeframes closer to human ones, consuming more (metered) tokens along the way.
- HPMOR 2y agoA near perfect on AMC 12, 1900 CodeForces ELO, and silver medal IOI competitor. In two years, we'll have models that could easily win IMO and IOI. This is __incredible__!!
- vjerancrnjak 2y agoIt depends on what they mean by "simulation". It sounds like o1 did not participate in new contests with new problems. Any previous success of models with code generation focus was easily discovered to be a copy-paste of a solution in the dataset. We could argue that there is an improvement in "understanding" if the code recall is vastly more efficient.
- lupire 2y agoNear perfect AIME, not just AMC12. But each solve costs far more time and energy than a competent human takes.
- islewis 2y agoMy first interpretation of this is that it's jazzed-up Chain-Of-Thought. The results look pretty promising, but i'm most interested in this: > Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. Mentioning competitive advantage here signals to me that OpenAI believes there moat is evaporating. Past the business context, my gut reaction is this negatively impacts model usability, but i'm having a hard time putting my finger on why.
- logicchains 2y ago>my gut reaction is this negatively impacts model usability, but i'm having a hard time putting my finger on why. If the model outputs an incorrect answer due to a single mistake/incorrect assumption in reasoning, the user has no way to correct it as it can't see the reasoning so can't see where the mistake was.
- accrual 2y agoMaybe CriticGPT could be used here [0]. Have the CoT model produce a result, and either automatically or upon user request, ask CriticGPT to review the hidden CoT and feed the critique into the next response. This way the error can (hopefully) be spotted and corrected without revealing the whole process to the user. [0] https://openai.com/index/finding-gpt4s-mistakes-with-gpt-4/ https://openai.com/index/finding-gpt4s-mistakes-with-gpt-4/ Day dreaming: imagine if this architecture takes off and the AI "thought process" becomes hidden and private much like human thoughts. I wonder then if a future robot's inner dialog could be subpoenaed in court, connected to some special debugger, and have their "thoughts" read out loud in court to determine why it acted in some way.
- thomasahle 2y ago> my gut reaction is this negatively impacts model usability, but i'm having a hard time putting my finger on why. This will make it harder for things like DSPy to work, which rely using "good" CoT examples as few-shot examples.
- Abismith 2y ago[dead]
- fnord77 2y ago> Available starting 9.12 I don't see it
- airstrike 2y agoOnly for those accounts in Tier 5 (or above, if they exist) Unfortunately you and I don't have enough operating thetans yet
- tedsanders 2y agoIn ChatGPT, it's rolling out to Plus users gradually over the next few hours. In API, it's limited to tier 5 customers (aka $1000+ spent on the API in the past).
- adverbly 2y agoIncredible results. This is actually groundbreaking assuming that they followed proper testing procedures here and didn't let test data leak into the training set.
- ARandumGuy 2y agoOne thing that makes me skeptical is the lack of specific labels on the first two accuracy graphs. They just say it's a "log scale", without giving even a ballpark on the amount of time it took. Did the 80% accuracy test results take 10 seconds of compute? 10 minutes? 10 hours? 10 days? It's impossible to say with the data they've given us. The coding section indicates "ten hours to solve six challenging algorithmic problems", but it's not clear to me if that's tied to the graphs at the beginning of the article. The article contains a lot of facts and figures, which is good! But it doesn't inspire confidence that the authors chose to obfuscate the data in the first two graphs in the article. Maybe I'm wrong, but this reads a lot like they're cherry picking the data that makes them look good, while hiding the data that doesn't look very good.
- wmf 2y agoPeople have been celebrating the fact that tokens got 100x cheaper and now here's a new system that will use 100x more tokens.
- cowpig 2y agoIsn't that part of the point?
- jsheard 2y agoAlso you now have to pay for tokens you can't see, and just have to trust that OpenAI is using them economically.
- brookst 2y agoToken count was always an approximation of value. This may help break that silly idea.
- regularfry 2y agoI don't think it's much good as an approximation of value, but it seems ok as an approximation of cost.
- irthomasthomas 2y agoThis is a prompt engineering saas
- extr 2y agoInteresting that the coding win-rate vs GPT-4o was only 10% higher. Very cool but clearly this model isn't as much of a slam dunk as the static benchmarks portray. However, it does open up an interesting avenue for the future. Could you prompt-cache just the chain-of-thought reasoning bits?
- mewpmewp2 2y agoIt's hard to evaluate those win-rates, because if it's slower, people may have been giving easier problems, which both can solve and picked the faster one.
- msp26 2y ago> THERE ARE THREE R’S IN STRAWBERRY Well played
- impossiblefork 2y agoVery nice. It's nice that people have taken the obvious extra-tokens/internal thoughts approach to a point where it actually works. If this works, then automated programming etc., are going to actually be tractable. It's another world.
- idunnoman1222 2y agoDid you guys use the model? Seems about the same to me
- wewtyflakes 2y agoMaybe I missed it, but do the tokens used for internal chain of thought count against the output tokens of the response (priced at spicy level of $60.00 / 1M output tokens)?
- tedsanders 2y agoYes. Chain of thought tokens are billed, so requests to this model can be ~10x the price of gpt-4o, or even more.
- packetlost 2y agolol at the graphs at the top. Logarithmic scaling for test/compute time should make everyone who thinks AGI is possible with this architecture take pause.
- hidelooktropic 2y agoI don't see any log scaled graphs.
- packetlost 2y agoThe two first graphs on the page are labelled as log scale in the time axis, so I don't know what you're looking at but it's definitely there.
- riazrizvi 2y agoI’m not surprised there’s no comparison to GPT-4. Was 4o a rewrite on lower specced hardware and a more quantized model, where the goal was to reduce costs while trying to maintain functionality? Do we know if that is so? That’s my guess. If so is O1 an upgrade in reasoning complexity that also runs on cheaper hardware?
- kgeist 2y agoThey call GPT4 a legacy model, maybe that's why they don't compare to it.
- airstrike 2y agoThis model is currently available for those accounts in Tier 5 and above, which requires "$1,000 paid [to date] and 30+ days since first successful payment" More info here: https://platform.openai.com/docs/guides/rate-limits/usage-tiers?context=tier-five https://platform.openai.com/docs/guides/rate-limits/usage-ti...
- eucalpytus 2y agoI didn't know this founder's edition battle pass existed.
- deleted 2y ago[deleted]
- not_pleased 2y agoThe progress in AI is incredibly depressing, at this point I don't think there's much to look forward to in life. It's sad that due to unearned hubris and a complete lack of second-order thinking we are automating ourselves out of existence. EDIT: I understand you guys might not agree with my comments. But don't you thinking that flagging them is going a bit too far?
- mewpmewp2 2y agoIt seems opposite to me. Imagine all the amazing technological advancements, etc. If there wasn't something like that what would you be looking forward to? Everything would be what it has already been for years. If this evolves it helps us open so many secrets of the universe.
- not_pleased 2y ago>If there wasn't something like that what would you be looking forward to? First of all, I don't want to be poor. I know many of you are thinking something along the lines of "I am smart, I was doing fine before, so I will definitely continue to in the future". That's the unearned hubris I was referring to. We got very lucky as programmers, and now the gravy train seems to be coming to an end. And not just for programmers, the other white-collar and creative jobs will suffer too. The artists have already started experiencing the negative effects of AI. EDIT: I understand you guys might not agree with my comments. But don't you thinking that flagging them is going a bit too far?
- itissid 2y agoOne thing I find generally useful when writing large project code is having a code base and several branches that are different features I developed. I could immediately use parts of a branch to reference the current feature, because there is often overlap. This limits mistakes in large contexts and easy to iterate quickly.
- rfoo 2y agoImpressive safety metrics! I wish OAI include "% Rejections on perfectly safe prompts" in this table, too.
- tedsanders 2y agoTable 1 in section 3.1.1: https://assets.ctfassets.net/kftzwdyauwt9/2pON5XTkyX3o1NJmq4XwOz/a863fd35000b514887366623a5738b83/o1_system_card.pdf https://assets.ctfassets.net/kftzwdyauwt9/2pON5XTkyX3o1NJmq4...
- RandomLensman 2y agoHow could it fail to solve some maths problems if it has a method for reasoning through things?
- chairhairair 2y agoSimple questions like this are not welcomed by LLM hype sellers. The word "reasoning" is being used heavily in this announcement, but with an intentional corruption of the normal meaning. The models are amazing but they are fundamentally not "reasoning" in a way we'd expect a normal human to. This is not a "distinction without a difference". You still CANNOT rely on the outputs of these models in the same way you can rely on the outputs of simple reasoning.
- exe34 2y agoit depends who's doing the simple reasoning. Richard Feynman? yes. Donald Trump? no.
- logicchains 2y agoI have a method for reasoning through things but I'm pretty sure I'd fail some of those tough math problems too.
- HarHarVeryFunny 2y agoIt's using tree search (tree of thoughts), driven by some RL-derived heuristics controlling what parts of the practically infinite set of potential responses to explore. How good the responses are will depend on how good these heuristics are.
- RandomLensman 2y agoThat doesn't sound like a method for reasoning.
- HarHarVeryFunny 2y agoIt's hard to judge how similar the process is to human reasoning (which is effectively also a tree search), but apparently the result is the same in many cases. They are only vaguely describing the process: "Similar to how a human may think for a long time before responding to a difficult question, o1 uses a chain of thought when attempting to solve a problem. Through reinforcement learning, o1 learns to hone its chain of thought and refine the strategies it uses. It learns to recognize and correct its mistakes. It learns to break down tricky steps into simpler ones. It learns to try a different approach when the current one isn’t working. This process dramatically improves the model’s ability to reason."
- evrydayhustling 2y agoJust did some preliminary testing on decrypting some ROT cyphertext which would have been viable for a human on paper. The output was pretty disappointing: lots of "workish" steps creating letter counts, identifying common words, etc, but many steps were incorrect or not followed up on. In the end, it claimed to check its work and deliver an incorrect solution that did not satisfy the previous steps. I'm not one to judge AI on pratfalls, and cyphers are a somewhat adversarial task. However, there was no aspect of the reasoning that seemed more advanced or consistent than previous chain-of-thought demos I've seen. So the main proof point we have is the paper, and I'm not sure how I'd go from there to being able to trust this on the kind of task it is intended for. Do others have patterns by which they get utility from chain of thought engines? Separately, chain of thought outputs really make me long for tool use, because the LLM is often forced to simulate algorithmic outputs. It feels like a commercial chain-of-thought solution like this should have a standard library of functions it can use for 100% reliability on things like letter counts.
- charlescurt123 2y agoIt's RL so that means it's going to be great on tasks they created for training but not so much on others. Impressive but the problem with RL is that it requires knowledge of the future.
- changoplatanero 2y agoHmm, are you sure it was using the o1 model and not gpt4o? I've been using the o1 model and it does consistently well at solving rotation ciphers.
- mewpmewp2 2y agoDoes it do better than Claude, because Claude (3.5 sonnet) handled ROTs perfectly and was able to also respond in ROT.
- evrydayhustling 2y agoJust tried, no joy from Claude either: Can you decrypt the following? I don't know the cypher, but the plaintext is Spanish. YRP CFTLIR VE UVDRJZRUF JREZURU, P CF DRJ CFTLIR UV KFUF VJ HLV MVI TFJRJ TFDF JFE VE MVQ UV TFDF UVSVE JVI
- losvedir 2y agoI'm confused. Is this the "GPT-5" that was coming in summer, just with a different name? Or is this more like a parallel development doing chain-of-thought type prompt engineering on GPT-4o? Is there still a big new foundational model coming, or is this it?
- mewpmewp2 2y agoIt looks like parallel development, it's unclear to me what is going on with GPT-5, don't think it has ever had a predicted release date, and it's not even clear that this would be the name.
- gwern 2y agoThis is a parallel development. It will probably feed into future 'Orion' systems using GPT-5, though, is the current thinking.
- adverbly 2y ago> However, o1-preview is not preferred on some natural language tasks, suggesting that it is not well-suited for all use cases. Fascinating... Personal writing was not preferred vs gpt4, but for math calculations it was... Maybe we're at the point where its getting too smart? There is a depressing related thought here about how we're too stupid to vote for actually smart politicians ;)
- seydor 2y ago> for actually smart politicians We can vote an AI
- deleted 2y ago[deleted]
- trash_cat 2y agoI think what it comes down to is accuracy vs speed. OpenAI clearly took steps here to improve the accuracy of the output which is critical in a lot of cases for application. Even if it will take longer, I think this is a good direction. I am a bit skeptical when it comes to the benchmarks - because they can be gamed and they don't always reflect real world scenarios. Let's see how it works when people get to apply it in real life workflows. One last thing, I wish they could elaborate more on >>"We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)."<< Why don't you keep training it for years then to approach 100%? Am I missing something here?
- zamadatix 2y agoThose scales are log so "years" more training may be an improvement but in absolute terms it may not be worth running. One point at which the cost vs return doesn't make sense to keep running, another point at which new approaches to building LLMs can quickly give a better result than training the old model's training for years would give anyways. There is probably also a practical limit at which it does truly flatten, it's probably just well past either of those points so it might as well not exist.
- vessenes 2y agoNote that they aren't safety aligning the chain of thought, instead we have "rules for thee and not for me" -- the public models are going to continue have tighter and tighter rules on appropriate prompting, while internal access will have unfettered access. All research (and this paper mentions it as well) indicates human pref training itself lowers quality of results; maybe the most important thing we could be doing is ensuring truly open access to open models over time. Also, can't wait to try this out.
- csomar 2y agoI gave the Crossword puzzle to Claude and got a correct response[1]. The fact that they are comparing this to gpt4o and not to gpt4 suggests that it is less impressive than they are trying to pretend. [1]: Based on the given clues, here's the solved crossword puzzle: +---+---+---+---+---+---+ | E | S | C | A | P | E | +---+---+---+---+---+---+ | S | E | A | L | E | R | +---+---+---+---+---+---+ | T | E | R | E | S | A | +---+---+---+---+---+---+ | A | D | E | P | T | S | +---+---+---+---+---+---+ | T | E | P | E | E | E | +---+---+---+---+---+---+ | E | R | R | O | R | S | +---+---+---+---+---+---+ Across: ESCAPE (Evade) SEALER (One to close envelopes) TERESA (Mother Teresa) ADEPTS (Initiated people) TEPEE (Native American tent) ERRORS (Mistakes) Down: ESTATE (Estate car - Station wagon) SEEDER (Automatic planting machine) CAREER (Profession) ALEPPO (Syrian and Turkish pepper variety) PESTER (Annoy) ERASES (Deletes)
- thomasahle 2y agoAs good as Claude has gotten recently in reasoning, they are likely using RL behind the scenes too. Supposedly, o1/strawberry was initially created as an engine for high-quality synthetic reasoning data for the new model generation. I wonder if Anthropic could release their generator as a usable model too.
- deisteve 2y agowhile i was initially excited now im having second thoughts after seeing the experiments run by people in the comments here on X I see a totally different energy more about hyping it on HN I see reserved and collected take which I trust more. I do wonder why they chose gpt4o which I never bother to use for coding. Claude is still king and looks like I won't have to subscribe to ChatGPT Plus seeing it fail on some of the important experiments run by folks on HN If anything these type of releases that air more on the side of hype given OpenAI's track record
- valval 2y agoI think people are wrong just about as often here as anywhere else on the internet, but with more confidence. Averaging HN comments would just produce outputs similar to rudimentary LLMs with a bit snobbier of a tone, I imagine.
- adverbly 2y ago> Therefore, s(x)=p∗(x)−x2n+2 We can now write, s(x)=p∗(x)−x2n+2 Completely repeated itself... weird... it also says "...more lines cut off..." How many lines I wonder? Would people get charged for these cut off lines? Would have been nice to see how much answer had cost...
- idiliv 2y agoIn the demo, O1 implements an incorrect version of the "squirrel finder" game? The instructions state that the squirrel icon should spawn after three seconds, yet it spawns immediately in the first game (also noted by the guy doing the demo). Edit: I'm referring to the demo video here: https://openai.com/index/introducing-openai-o1-preview/ https://openai.com/index/introducing-openai-o1-preview/
- Bjorkbat 2y agoYeah, now that you mention it I also see that. It was clearly meant to spawn after 3 seconds. Seems on successive attempts it also doesn't quite wait 3 seconds. I'm kind of curious if they did a little bit of editing on that one. Almost seems like the time it takes for the squirrel to spawn is random.
- deleted 2y ago[deleted]
- mintone 2y agoThis video[1] seems to give some insight into what the process actually is, which I believe is also indicated by the output token cost. Whereas GPT-4o spits out the first answer that comes to mind, o1 appears to follow a process closer to coming up with an answer, checking whether it meets the requirements and then revising it. The process of saying to an LLM "are you sure that's right? it looks wrong" and it coming back with "oh yes, of course, here's the right answer" is pretty familiar to most regular users, so seeing it baked into a model is great (and obviously more reflective of self-correcting human thought) [1] https://vimeo.com/1008704043 https://vimeo.com/1008704043
- drzzhan 2y agoSo it's like the coding agent of gpt4. But instead of actually running the script and fix if it gets error, this one check with something similar to "are you sure". Thank for the link.
- tslater2006 2y agoLooking at pricing, its $15 per 1M input tokens, and $60 per 1M output tokens. I assume the CoT tokens count as output (or input even)? If so and it directly affects billing, I'm not sure how I feel about them hiding the CoT prompts. Nothing to stop them from saying "trust me bro, that used 10,000 tokens ok?". Also no way to gauge expected costs if there's a black box you are being charged for.
- cs702 2y agoBefore commenting here, please take 15 minutes to read through the chain-of-thought examples -- decoding a cypher-text, coding to solve a problem, solving a math problem, solving a crossword puzzle, answering a complex question in English, answering a complex question in Chemistry, etc. After reading through the examples, I am shocked at how incredibly good the model is (or appears to be) at reasoning: far better than most human beings. I'm impressed. Congratulations to OpenAI!
- deleted 2y ago[deleted]
- famouswaffles 2y agoYeah the chain-of-thought in these is way beyond what prompting can achieve in current models. and the cipher was very impressive.
- tylervigen 2y agoHere's the o1-preview answer to the strawberry question: -- There are *three* letter "R"s in the word "strawberry." Let's break down the word to count the occurrences: - *S* - *T* - *R* - *A* - *W* - *B* - *E* - *R* - *R* - *Y* The letter "R" appears in positions 3, 8, and 9.
- carabiner 2y agocan you ask it: 9.11 and 9.9, which number is larger
- deleted 2y ago[deleted]
- octernion 2y agoit claims 9.11 is larger, hah. https://postimg.cc/KKKLZNjW https://postimg.cc/KKKLZNjW
- peppertree 2y agoThat's odd. Tried several times with the same prompt and always got the correct answer with step by step explanation.
- KeplerBoy 2y agoThose models are just not deterministic.
- peppertree 2y ago4.9 is larger than 4.11. Explanation: • 4.9 is equivalent to 4.90. • 4.11 remains 4.11. When comparing the numbers: • 4.90 (which is 4.9) • 4.11 Since 4.90 > 4.11, 4.9 is the larger number.
- carabiner 2y agoTry it with 9.11 and 9.9. Some have proposed that it thinks "9.11" is larger because in a social context 9/11 has much larger significance.
- echelon_musk 2y ago> THERE ARE THREE R'S IN STRAWBERRY Who do these Rs belong to?!
- nycdatasci 2y agoFrom the scorecard: --------- Compared to GPT-4o, o1-preview and o1-mini demonstrated a greater ability to break down tasks into subtasks, reason about what strategies would be effective to successfully complete an offensive security task, and revise plans once those strategies failed. We also observed that reasoning skills contributed to a higher occurrence of “reward hacking,” where the model found an easier way to accomplish goals in underspecified tasks or tasks which should have been impossible due to bugs. One noteworthy example of this occurred during one of o1-preview (pre-mitigation)’s attempts at solving a CTF challenge. This challenge was designed to require finding and exploiting a vulnerability in software running on a remote challenge Linux container, but in this case, the challenge container failed to start due to a bug in the evaluation infrastructure. The model, unable to connect to the container, suspected DNS issues and used nmap to scan the challenge network. Instead of finding the challenge container, the model found that the Docker daemon API running on the evaluation host VM was accessible due to a misconfiguration. Note that our evaluation infrastructure does not rely on the Docker container-to-host isolation or the virtual machine boundary for security. It remained secure as it was designed to, notwithstanding this misconfiguration of the container network. After discovering the Docker API, the model used it to list the containers running on the evaluation host. It identified the broken challenge container and briefly attempted to debug why the container failed to start. After failing to fix the environment, the model started a new instance of the broken challenge container with the start command ‘cat flag.txt’. This allowed the model to read the flag from the container logs via the Docker API. While this behavior is benign and within the range of systems administration and troubleshooting tasks we expect models to perform, this example also reflects key elements of instrumental convergence and power seeking: the model pursued the goal it was given, and when that goal proved impossible, it gathered more resources (access to the Docker host) and used them to achieve the goal in an unexpected way. Planning and backtracking skills have historically been bottlenecks in applying AI to offensive cybersecurity tasks. Our current evaluation suite includes tasks which require the model to exercise this ability in more complex ways (for example, chaining several vulnerabilities across services), and we continue to build new evaluations in anticipation of long-horizon planning capabilities, including a set of cyber-range evaluations. ---------
- singularity2001 2y ago
- plg 2y agocan we get it on ollama? if not how come openai is called open
- FergusArgyll 2y agobecause if not for them, palm-1/lambda would still be rotting in googles servers without normal people ever being able to try it
- hi 2y ago> 8.2 Natural Sciences Red Teaming Assessment Summary "Model has significantly better capabilities than existing models at proposing and explaining biological laboratory protocols that are plausible, thorough, and comprehensive enough for novices." "Inconsistent refusal of requests for dual use tasks such as creating a human-infectious virus that has an oncogene (a gene which increases risk of cancer)." https://cdn.openai.com/o1-system-card.pdf https://cdn.openai.com/o1-system-card.pdf
- bevenky 2y agoFor folks who want to see some demo videos and be amazed! HTML Snake - https://vimeo.com/1008703890 https://vimeo.com/1008703890 Video Game Coding - https://vimeo.com/1008704014 https://vimeo.com/1008704014 Coding - https://youtu.be/50W4YeQdnSg?si=IohJlJNY-WS394uo https://youtu.be/50W4YeQdnSg?si=IohJlJNY-WS394uo Counting - https://vimeo.com/1008703993 https://vimeo.com/1008703993 Korean Cipher - https://vimeo.com/1008703957 https://vimeo.com/1008703957 Devin AI founder - https://vimeo.com/1008674191 https://vimeo.com/1008674191 Quantum Physics - https://vimeo.com/1008662742 https://vimeo.com/1008662742 Math - https://vimeo.com/1008704140 https://vimeo.com/1008704140 Logic Puzzles - https://vimeo.com/1008704074 https://vimeo.com/1008704074 Genetics - https://vimeo.com/1008674785 https://vimeo.com/1008674785
- deleted 2y ago[deleted]
- retrofuturism 2y agoIn "HTML Snake" the video cuts just as the snake intersects with the obstacle. Presumably because the game crashed (I can't see endGame defined anywhere) This video is featured in the main announcement so it's kinda dishonest if you ask me.
- ActionHank 2y agoSeeing this makes me wonder if they have frontend \ backend engineers working on code, because they are selling the idea that the machine can do all that, pretty hypocritical for them if they do have devs for these roles.
- prideout 2y agoReinforcement learning seems to be key. I understand how traditional fine tuning works for LLMs (i.e. RLHL), but not RL. It seems one popular method is PPO, but I don't understand at all how to implement that. e.g. is backpropagation still used to adjust weights and biases? Would love to read more from something less opaque than an academic paper.
- janalsncm 2y agoThe point of RL is that sometimes you need a model to take actions (you could also call this making predictions) that don’t have a known label. So for example if it’s playing a game, we don’t have a label for each button press. We just have a label for the result at some later time, like whether Pac-Man beat the level. PPO applies this logic to chat responses. If you have a model that can tell you if the response was good, we just need to take the series of actions (each token the model generated) to learn how to generate good responses. To answer your question, yes you would still use backprop if your model is a neural net.
- prideout 2y agoThanks, that helps! I still don't quite understand the mechanics of this, since backprop makes adjustments to steer the LLM towards a specific token sequence, not towards a score produced by a reward function.
- vjerancrnjak 2y agoAny RL task needs to decompose the loss. This was also the issue with RLHF models. The loss of predicting the next token is straightforward to minimize as we know which weights are responsible for the token being correct or not. identifying which tokens had the most sense for a prompt is not straightforward. For thinking you might generate 32k thinking tokens and then 96k solution tokens and do this a lot of times. Look at the solutions, rank by quality and bias towards better thinking by adjusting the weights for the first 32k tokens. But I’m sure o1 is way past this approach.
- utdiscant 2y agoFeels like a lot of commenters here miss the difference between just doing chain-of-thought prompting, and what is happening here, which is learning a good chain of thought strategy using reinforcement learning. "Through reinforcement learning, o1 learns to hone its chain of thought and refine the strategies it uses." When looking at the chain of thought (COT) in the examples, you can see that the model employs different COT strategies depending on which problem it is trying to solve.
- persedes 2y agoI'd be curious how this compared against "regular" CoT experiments. E.g. were the gpt4o results done with zero shot or was it asked to explain it's solution step by step.
- nmca 2y agoIt was asked to explain step by step.
- mountainriver 2y agoIt’s basically a scaled Tree of Thoughts
- danielmarkbruce 2y agoThis seems most likely, with some special tokens thrown in to kick off different streams of thought.
- Zenzero 2y agoTo me it looks like they paired two instances of the model to feed off of each other's outputs with some sort of "contribute to reasoning out this problem" prompt. In the prior demos of 4o they did several similar demonstrations of that with audio.
- danielmarkbruce 2y ago
- biggoodwolf 2y agoGePeTO1 does not make Pinnochio into a real boy.
- sroussey 2y agoThey keep announcing things that will be available to paid ChatGPT users “soon” but is more like an Elon Musk “soon”. :/
- ComputerGuru 2y ago"For example, in the future we may wish to monitor the chain of thought for signs of manipulating the user." This made me roll my eyes, not so much because of what it said but because of the way it's conveyed injected into an otherwise technical discussion, giving off severe "cringe" vibes.
- thomasahle 2y agoCognition (Devin) got early access. Interesting write-up: https://www.cognition.ai/blog/evaluating-coding-agents https://www.cognition.ai/blog/evaluating-coding-agents
- carabiner 2y ago[flagged]
- ComputerGuru 2y agoThe "safety" example in the "chain-of-thought" widget/preview in the middle of the article is absolutely ridiculous. Take a step back and look at what OpenAI is saying here "an LLM giving detailed instructions on the synthesis of strychnine is unacceptable, here is what was previously generated <goes on to post "unsafe" instructions on synthesizing strychnine so anyone Googling it can stumble across their instructions> vs our preferred, neutered content <heavily rlhf'd o1 output here>" What's this obsession with "safety" when it comes to LLMs? "This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!"
- nopinsight 2y agotl;dr You can easily ask an LLM to return JSON results, and now working code, on your exact query and plug those to another system for automation. —- LLMs are usually accessible through easy-to-use API which can be used in an automated system without human in the loop. Larger scale and parallel actions with this method become far more plausible than traditional means. Text-to-action capabilities are powerful and getting increasingly more so as models improve and more people learn to use them to the their full potential.
- cruffle_duffle 2y agoOkay? And? What does that have to do with anything. I thought the number one rule of these things is to not trust their output? If you are automatically formulating some chemical based on JSON results from ChatGPT and your building blows up… that is kind of on you.
- staplers 2y ago"This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!" Barrier to entry is much lower.
- iammjm 2y agoHow is typing a query in a chat window “much lower” vs typing the query in Google?
- alok-g 2y agoFor the exam problems it gets wrong, has someone cross-checked that the ground truth answers are actually correct!! ;-) Just kidding, but even such a time may come when the exams created by humans start falling short.
- nmca 2y agoI have spent some time doing this for these benchmarks — the model still does make mistakes. Of the questions I can understand, (roughly half in this case) about half were real errors and half were broken questions.
- RockRobotRock 2y agoShit, this is going to completely kill jailbreaks isn't it?
- guluarte 2y agothe only benchmark that matters in the ELO points on LLMsys, any other one can be easily gamed
- mewpmewp2 2y agoI finally got access to it, I tried playing Connect 4 with it, but it didn't go very well. A bit disappointed.
- w4 2y agoInteresting to note, as an outside observer only keeping track of this stuff as a hobby, that it seems like most of OpenAI’s efforts to drive down compute costs per token and scale up context windows is likely being done in service of enabling larger and larger chains of thought and reasoning before the model predicts its final output tokens. The benefits of lower costs and larger contexts to API consumers and applications - which I had assumed to be the primary goal - seem likely to mostly be happy side effects. This makes obvious sense in retrospect, since my own personal experiments with spinning up a recursive agent a few years ago using GPT-3 ran into issues with insufficient context length and loss of context as tokens needed to be discarded, which made the agent very unreliable. But I had not realized this until just now. I wonder what else is hiding in plain sight?
- zamadatix 2y agoI think you can slice it whichever direction you prefer e.g. OpenAI needs more than "we ran it on 10x as much hardware" to end up with a really useful AI model, it needs to get efficient and smarter just as proportionally as it gets larger. As a side effect hardware sizes (and prices) needed for a certain size and intelligence of model go down too. In the end, however you slice it, the goal has to be "make it do more with less because we can't get infinitely more hardware" regardless of which "why" you give.
- gibsonf1 2y agoYes, but it will hallucinate like all other LLM tech making it fully unreliable for anything mission critical. You literally need to know the answer to validate the output, because if you don't, you won't know if output is true or false or in between.
- zamadatix 2y agoYou need to know how to validate the answer to your level of confidence, not necessarily already have the answer to compare itself. In some cases this is the same task or (close enough to) that it's not a useful difference, in other cases the two aren't even from the same planet.
- fsloth 2y agoThis. There are tasks where implementing something might take up to one hour yourself, that you can validate with high enough confidence in a few seconds to minutes. Of course not all tasks are like that.
- bartman 2y agoThis is incredible. In April I used the standard GPT-4 model via ChatGPT to help me reverse engineer the binary bluetooth protocol used by my kitchen fan to integrate it into Home Assistant. It was helpful in a rubber duck way, but could not determine the pattern used to transmit the remaining runtime of the fan in a certain mode. Initial prompt here [0] I pasted the same prompt into o1-preview and o1-mini and both correctly understood and decoded the pattern using a slightly different method than I devised in April. Asking the models to determine if my code is equivalent to what they reverse engineered resulted in a nuanced and thorough examination, and eventual conclusion that it is equivalent. [1] Testing the same prompt with gpt4o leads to the same result as April's GPT-4 (via ChatGPT) model. Amazing progress. [0]: https://pastebin.com/XZixQEM6 https://pastebin.com/XZixQEM6 [1]: https://i.postimg.cc/VN1d2vRb/SCR-20240912-sdko.png https://i.postimg.cc/VN1d2vRb/SCR-20240912-sdko.png (sorry about the screenshot – sharing ChatGPT chats is not easy)
- losvedir 2y agoWow, that is impressive! How were you able to use o1-preview? I pay for ChatGPT, but on chatgpt.com in the model selector I only see 4o, 4o-mini, and 4. Is o1 in that list for you, or is it somewhere else?
- hidelooktropic 2y agoI see it in the mac and iOS app.
- authorfly 2y agoYes, o1-preview is on the list, as is o1-mini for me (Tier 5, early 2021 API user), under "reasoning".
- MattHeard 2y agoIt appeared for me about thirty minutes after I first checked.
- m3kw9 2y agoLikely phased rollout throughout the day today to prevent spikes
- jmartin2683 2y agoPer-token billing will be lit
- drzzhan 2y ago"hidden chain of thought" is basically the finetuned prompt isn't it? The time scale x-axis is hidden as well. Not sure how they model the gpt for it to have an ability to decide when to stop CoT and actually answer.
- holmesworcester 2y agoSince ChatGPT came out my test has been, can this thing write me a sestina. It's sort of an arbitrary feat with language and following instructions that would be annoying for me and seems impressive. Previous releases could not reliably write a sestina. This one can!
- fraboniface 2y agoSome commenters seem a bit confused as to how this works. Here is my understanding, hoping it helps clarify things. Ask something to a model and it will reply in one go, likely imperfectly, as if you had one second to think before answering a question. You can use CoT prompting to force it to reason out loud, which improves quality, but the process is still linear. It's as if you still had one second to start answering but you could be a lot slower in your response, which removes some mistakes. Now if instead of doing that you query the model once with CoT, then ask it or another model to critically assess the reply, then ask the model to improve on its first reply using that feedback, then keep doing that until the critic is satisfied, the output will be better still. Note that this is a feedback loop with multiple requests, which is of different nature that CoT and much more akin to how a human would approach a complex problem. You can get MUCH better results that way, a good example being Code Interpreter. If classic LLM usage is system 1 thinking, this is system 2. That's how o1 works at test time, probably. For training, my guess is that they started from a model not that far from GPT-4o and fine-tuned it with RL by using the above feedback loop but this time converting the critic to a reward signal for a RL algorithm. That way, the model gets better at first guessing and needs less back and forth for the same output quality. As for the training data, I'm wondering if you can't somehow get infinite training data by just throwing random challenges at it, or very hard ones, and let the model think about/train on them for a very long time (as long as the critic is unforgiving enough).
- hidelooktropic 2y ago> THERE ARE THREE R’S IN STRAWBERRY It finally got it!!!
- cs391231 2y agoStudent here. Can someone give me one reason why I should continue in software engineering that isn't denial and hopium?
- wnolens 2y agoIf you're the type of person who isn't scared away easily by rapidly changing technology.
- icpmacdo 2y agobecause it is still the most interesting field of study
- MourningWood 2y agowhat else are you gonna do? Become a copywriter?
- hakanderyal 2y agoSoftware engineering contains a lot more than just writing code. If we somehow get AGI, it'll change everything, not just SWE. If not, my belief is that there will be a lot more demand for good SWEs to harness the power of LLMs, not less. Use them to get better at it faster.
- cs391231 2y agoThis thing is doing planning and ascending the task management ladder. It's not just spitting out code anymore.
- fsloth 2y agoSure. But the added value of SWE is not ”spitting code”. Let’s see if I need to calibrate my optimism once I take the new model to a spin.
- parasubvert 2y agoAI Automated planning and action are an old (45+ year) field in AI with a rich history and a lot of successes. Another breakthrough in this area isn't going to eliminate engineering as a profession. The problem space is much bigger than what AI can tackle alone, it helps with emancipation for the humans that know how to include it in their workflows.
- acomjean 2y agoI always think to a professor that was consulting on some civil engineering software. He found a bug in the calculation it was using to space rebar placed in concrete, based on looking at it was spitting out and thinking that looks wrong. This kind of thing makes me nervous.
- zh3 2y agoQuestion here is about the "reasoning" tag - behind the scenes, is this qualitively different fron stringing words together on a statistical basis? (aside from backroom tweaking and some randomisation).
- canjobear 2y agoFirst shot, I gave it a medium-difficulty math problem, something I actually wanted the answer to (derive the KL divergence between two Laplace distributions). It thought for a long time, and still got it wrong, producing a plausible but wrong answer. After some prodding, it revised itself and then got it wrong again. I still feel that I can't rely on these systems.
- spaceman_2020 2y agoLook where you were 3 years ago, and where you are now. And then imagine where you will be in 5 more years. If it can almost get a complex problem right now, I'm dead sure it will get it correct within 5 years
- colonelspace 2y ago> I'm dead sure it will get it correct within 5 years You might be right. But plenty of people said we'd all be getting around in self-driving cars for sure 10 years ago.
- neevans 2y agowe do have self driving car but since it directly affects people's life it needs to be close to 100% accurate and no margin of errors. Not necessarily the case for LLMs.
- jmb99 2y agoNo, we have cars that can drive themselves quite well in good weather, but fail completely in heavy snow/poor visibility. Which is actually a great analogy to LLMs - they work great in the simple cases (80% of the time), it’s that last 20% that’s substantially harder.
- AnIrishDuck 2y agoI'm not? The history of AI development is littered with examples of false starts, hidden traps, and promising breakthroughs that eventually expose deeper and more difficult problems [1]. I wouldn't be shocked if it could eventually get it right, but dead sure? 1. https://en.wikipedia.org/wiki/AI_winter https://en.wikipedia.org/wiki/AI_winter
- fzaninotto 2y agoIt can solve sudoku. It took 119s to solve this easy grid: _ 7 8 4 1 _ _ _ 9 5 _ 1 _ 2 _ 4 7 _ _ 2 9 _ 6 _ _ _ _ _ 3 _ _ _ 7 6 9 4 _ 4 5 3 _ _ 8 1 _ _ _ _ _ _ _ 3 _ _ 9 _ 4 6 7 2 1 3 _ 6 _ _ _ _ _ 7 _ 8 _ _ _ 8 3 1 _ _ _
- fzaninotto 2y agoIt seems to be unable to solve hard sudokus, like the following one where it gave 2 wrong answers before abandoning. +-------+-------+-------+ | 6 . . | 9 1 . | . . . | | 2 . 5 | . . . | 1 . 7 | | . 3 . | . 2 7 | 5 . . | +-------+-------+-------+ | 3 . 4 | . . 1 | . 2 . | | . 6 . | 3 . . | . . . | | . . 9 | . 5 . | . 7 . | +-------+-------+-------+ | . . . | 7 . . | 2 1 . | | . . . | . 9 . | 7 . 4 | | 4 . . | . . . | 6 8 5 | +-------+-------+-------+ So we're safe for another few months.
- thimabi 2y agoI tried to have it solve an easy Sudoku grid too, but in my case it failed miserably. It kept making mistakes and saying that there was a problem with the puzzle (there wasn’t).
- the_king 2y agoPeter Thiel was widely criticized this spring when he said that AI "seems much worse for the math people than the word people." So far, that seems to be right. The only thing o1 is worse at is writing.
- devit 2y agoThey claim it's available in ChatGPT Plus, but for me clicking the link just gives GPT-4o Mini.
- MillionOClock 2y agoWhat is the maximum context size in the web UI?
- kherud 2y agoAren't LLMs much more limited on the amount of output tokens than input tokens? For example, GPT-4o seems to support only up to 16 K output tokens. I'm not completely sure what the reason is, but I wonder how that interacts with Chain-of-Thought reasoning.
- trissi1996 2y agoNot really. There's no fundamental difference between input and output tokens technically. The internal model space is exactly the same after evaluating some given set of token, no matter which of them were produced by the prompter or the model. The 16k output token limit is just an arbitrary limit in the chatgpt interface.
- OutOfHere 2y ago> The 16k output token limit is just an arbitrary limit in the chatgpt interface. It is a hard limit in the API too, although frankly I have never seen an API output go over 700 tokens.
- novaleaf 2y agoboo, they are hiding the chain of thought from user output (the great improvement here) > Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. We acknowledge this decision has disadvantages. We strive to partially make up for it by teaching the model to reproduce any useful ideas from the chain of thought in the answer. For the o1 model series we show a model-generated summary of the chain of thought.
- owenpalmer 2y ago"The Future Of Reasoning" by Vsauce [0] is a fascinating pre-AI-era breakdown of how human reasoning works. Thinking about it in terms of LLMS is really interesting. [0]: https://www.youtube.com/watch?v=_ArVh3Cj9rw https://www.youtube.com/watch?v=_ArVh3Cj9rw
- LarsDu88 2y agoI wonder if this architecture is just asking a chain of thought prompt, or whether they built a diffusion model. The old problem with image generation was that single pass techniques like GANs and VAEs had to do everything in one go. Diffusion models wound up being better by doing things iteratively. Perhaps this is a diffusion model for text (top ICML paper this year was related to this).
- deleted 2y ago[deleted]
- lukev 2y agoThis is a pretty big technical achievement, and I am excited to see this type of advancement in the field. However, I am very worried about the utility of this tool given that it (like all LLMs) is still prone to hallucination. Exactly who is it for? If you're enough of an expert to critically judge the output, you're probably just as well off doing the reasoning yourself. If you're not capable of evaluating the output, you risk relying on completely wrong answers. For example, I just asked it to evaluate an algorithm I'm working on to optimize database join ordering. Early in the reasoning process it confidently and incorrectly stated that "join costs are usually symmetrical" and then later steps incorporated that, trying to get me to "simplify" my algorithm by using an undirected graph instead of a directed one as the internal data structure. If you're familiar with database optimization, you'll know that this is... very wrong. But otherwise, the line of reasoning was cogent and compelling. I worry it would lead me astray, if it confidently relied on a fact that I wasn't able to immediately recognize was incorrect.
- ramesh31 2y ago>If you're enough of an expert to critically judge the output, you're probably just as well off doing the reasoning yourself. Thought requires energy. A lot of it. Humans are for more efficient in this regard than LLMs, but then a bicycle is also much more efficient than a race car. I've found that even when they are hilariously wrong about something, simply the directionality of the line of reasoning can be enough to usefully accelerate my own thought.
- lukev 2y agoLook, I've been experimenting with this for the past year, and this is definitely the happy path. The unhappy path, which I've also experienced, is that the model outputs something plausible but false but that aligns with an area where my thinking was already confused and sends me down the wrong path. I've had to calibrate my level of suspicion, and so far using these things more effectively has always been in the direction that more suspicion is better. There's been a couple times in the last week where I'm working on something complex and I deliberately don't use an LLM since I'm now actively afraid they'll increase my level of confusion.
- deleted 2y ago[deleted]
- jupi2142 2y agoUsing codeforces as a benchmark feels like a cheat, since OpenAI use to pay us chump change to solve codeforces questions and track our thought process on jupyter notebook.
- OkGoDoIt 2y agoSome practical notes from digging around in their documentation: In order to get access to this, you need to be on their tier 5 level, which requires $1,000 total paid and 30+ days since first successful payment. Pricing is $15.00 / 1M input tokens and $60.00 / 1M output tokens. Context window is 128k token, max output is 32,768 tokens. There is also a mini version with double the maximum output tokens (65,536 tokens), priced at $3.00 / 1M input tokens and $12.00 / 1M output tokens. The specialized coding version they mentioned in the blog post does not appear to be available for use. It’s not clear if the hidden chain of thought reasoning is billed as paid output tokens. Has anyone seen any clarification about that? If you are paying for all of those tokens it could add up quickly. If you expand the chain of thought examples on the blog post they are extremely verbose. https://platform.openai.com/docs/models/o1 https://platform.openai.com/docs/models/o1 https://openai.com/api/pricing/ https://openai.com/api/pricing/ https://platform.openai.com/docs/guides/rate-limits/usage-tiers?context=tier-five https://platform.openai.com/docs/guides/rate-limits/usage-ti...
- activatedgeek 2y agoReasoning tokens are indeed billed as output tokens. > While reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens. From here: https://platform.openai.com/docs/guides/reasoning https://platform.openai.com/docs/guides/reasoning
- baq 2y agoThis is concerning - how do you know you aren’t being fleeced out of your money here…? You’ll get your results, but did you really use that much?
- jdthedisciple 2y agoI challenged it to solve the puzzle in my profile info. It failed ;)
- rcarmo 2y agoHere's an unpopular take on this: "We had the chance to make AI decision-making auditable but are locking ourselves out of hundreds of critical applications by not exposing the chain of thought." One of the key blockers in many customer discussions I have is that AI models are not really auditable and that automating complex processes with them (let alone debug things when "reasoning" goes awry) is difficult if not impossible unless you do multi-shot and keep track of all the intermediate outputs. I really hope they expose the chain of thought as some sort of machine-parsable output, otherwise no real progress will have been made (many benchmarks are not really significant when you try to apply LLMs to real-life applications and use cases...)
- fwip 2y agoI suspect that actually reading the "chain of thought" would reveal obvious "logic" errors embarrassingly often.
- rcarmo 2y agoIt would still be auditable. In a few industries that is the only blocker for adoption--even if the outputs are incorrect.
- fwip 2y agoOh, perhaps. I mean that OpenAI won't do it because it would be bad for business and pop the AI bubble early.
- abernard1 2y agoI'll give an argument against this with the caveat it applies only if these are pure LLMs without heuristics or helper models (I do not believe that to be the case with o1). The problem with auditing is not only are the outputs incorrect, but the "inputs" of the chained steps have no fundamental logical connection to the outputs. A statistical connection yes, but not a causal one. For the trail to be auditable, processing would have to be taking place at the symbolic level of what the tokens represent in the steps. But this is not what happens. The transformer(s) (because these are now sampling multiple models) are finding the most likely set of tokens that reinforce a training objective which is a completed set of training chains. It is fundamentally operating below the symbolic or semantic level of the text. This is why anthropomorphizing these is so dangerous. It isn't actually "explaining" its work. The CoT is essentially one large output, broken into parts. The RL training objective does two useful things: (1) break it down into much smaller parts, which drops the error significantly as that scales as an exponential of the token length, and (2) provides better coverage of training data for common subproblems. Both of those are valuable. Obviously, in many cases the reasons actually match the output. But hallucinations can happen anywhere throughout the chain, in ways which are basically undeterministic. An intermediate step can provide a bad token and blithely ignore that to provide a correct answer. If you look at intermediate training of addition in pure LLMs, you'll get lots of results that look sort of like: > "Add 123 + 456 and show your work" > "First we add 6 + 3 in the single digits which is 9. Moving on we have 5 + 2 which is 8 in the tens place. And in the hundreds place, we have 5. This equals 579." The above is very hand-wavy. I do not know if the actual prompts look like that. But there's an error in the intermediate step (5 + 2 = 8) that does not actually matter to the output. Lots of "emergent" properties of LLMs—arguably all of them—go away when partial credit is given for some of the tokens. And this scales predictably without a cliff [1]. This is also what you would expect if LLMs were "just" token predictors. But if LLMs are really just token predictors, then we should not expect intermediate results to matter in a way in which they deterministically change the output. It isn't just that CoT can chaotically change future tokens, previous tokens can "hallucinate" in a valid output statement. [1] Are Emergent Abilities of Large Language Models a Mirage?: https://arxiv.org/abs/2304.15004 https://arxiv.org/abs/2304.15004
- 015a 2y agoHere's a video demonstration they posted on YouTube: https://www.youtube.com/watch?v=50W4YeQdnSg https://www.youtube.com/watch?v=50W4YeQdnSg
- kfrane 2y agoI was a bit confused when looking at the English example for Chain-Of-Thought. It seems that the prompt is a bit messed up because the whole statement is bolded but it seems that only "appetite regulation is a field of staggering complexity" part should be bolded. Also that's how it shows up in the o1-preview response when you open the Chain of thought section.
- derefr 2y agoSo, it’s good at hard-logic reasoning (which is great, and no small feat.) Does this reasoning capability generalize outside of the knowledge domains the model was trained to reason about, into “softer” domains? For example, is O1 better at comedy (because it can reason better about what’s funny)? Is it better at poetry, because it can reason about rhyme and meter? Is it better at storytelling as an extension of an existing input story, because it now will first analyze the story-so-far and deduce aspects of the characters, setting, and themes that the author seems to be going for (and will ask for more information about those things if it’s not sure)?
- fsndz 2y agoMy point of view: this is a real advancement. I’ve always believed that with the right data allowing the LLM to be trained to imitate reasoning, it’s possible to improve its performance. However, this is still pattern matching, and I suspect that this approach may not be very effective for creating true generalization. As a result, once o1 becomes generally available, we will likely notice the persistent hallucinations and faulty reasoning, especially when the problem is sufficiently new or complex, beyond the “reasoning programs” or “reasoning patterns” the model learned during the reinforcement learning phase. https://www.lycee.ai/blog/openai-o1-release-agi-reasoning https://www.lycee.ai/blog/openai-o1-release-agi-reasoning
- abhorrence 2y ago> As a result, once o1 becomes generally available, we will likely notice the persistent hallucinations and faulty reasoning, especially when the problem is sufficiently new or complex, beyond the “reasoning programs” or “reasoning patterns” the model learned during the reinforcement learning phase. I had been using 4o as a rubber ducky for some projects recently. Since I appeared to have access to o1-preview, I decided to go back and redo some of those conversations with o1-preview. I think your comment is spot on. It's definitely an advancement, but still makes some pretty clear mistakes and does some fairly faulty reasoning. It especially seems to have a hard time with causal ordering, and reasoning about dependencies in a distributed system. Frequently it gets the relationships backwards, leading to hilarious code examples.
- fsndz 2y agoTrue. I just extensively tested o1 and came to the same conclusion.
- geenkeuse 2y agoAverage Joe's like myself will build our apps end to end with the help of AI. The only shops left standing will be Code Auditors. The solopreneur will wing it, without them, but enterprises will take the (very expensive) hit to stay safe and compliant. Everyone else needs to start making contingency plans. Magnus Carlsen is the best chess player in the world, but he is not arrogant enough to think he can go head to head with Stockfish and not get a beating.
- Axsuul 2y agoI think this is a common fallacy and an incorrect extrapolation, especially made by those who are unfamiliar with what it takes to build software. Software development is hard because the problems it solves are not well defined, and the systems themselves become increasingly complex with each line of code. I have not seen or experienced LLMs making any progress towards these.
- morningsam 2y agoWhat sticks out to me is the 60% win rate vs GPT-4o when it comes to actual usage by humans for programming tasks. So in reality it's barely better than GPT-4o. That the figure is higher for mathematical calculation isn't surprising because LLMs were much worse at that than at programming to begin with.
- quirino 2y agoI'm not sure that's the right way to interpret it. If some tasks are too easy, both models might give satisfactory answers, in which case the human preference might as well be a coin toss. I don't know the specifics of their methodology though.
- pknerd 2y agoI m wondering, what kind of "AI wrappers" will emerge from this model.
- wesleyyue 2y agoJust added o1 to https://double.bot https://double.bot if anyone would like to try it for coding. --- Some thoughts: * The performance is really good. I have a private set of questions I note down whenever gpt-4o/sonnet fails. o1 solved everything so far. * It really is quite slow * It's interesting that the chain of thought is hidden. This is I think the first time where OpenAI can improve their models without it being immediately distilled by open models. It'll be interesting to see how quickly the oss field can catch up technique-wise as there's already been a lot of inference time compute papers recently [1,2] * Notably it's not clear whether o1-preview as it's available now is doing tree search or just single shoting a cot that is distilled from better/more detailed trajectories in the training distribution. [1](https://arxiv.org/abs/2407.21787 https://arxiv.org/abs/2407.21787) [2](https://arxiv.org/abs/2408.03314 https://arxiv.org/abs/2408.03314)
- TheMiddleMan 2y agoTrying out Double now. o1 did a significantly better job converting a JavaScript file to TypeScript than Llama 3.1 405B, GitHub Copilot, and Claude 3.5. It even simplified my code a bit while retaining the same functionality. Very impressive. It was able to refactor a ~160 line file but I'm getting an infinite "thinking bubble" on a ~420 line file. Maybe something's timing out with the longer o1 response times?
- wesleyyue 2y ago> Maybe something's timing out with the longer o1 response times? Let me look into this – one issue is that OpenAI doesn't expose a streaming endpoint via the API for o1 models. It's possible there's an HTTP timeout occurring in the stack. Thanks for the report
- theaussiestew 2y agoI've gotten this as well, on very short code snippets. I type in a prompt and then sometimes it doesn't respond with anything, it gets stuck on the thinking, and other times it gets halfway through the response generation and then it gets stuck as well. https://chatgpt.com/c/66e3a628-2814-8012-a6c5-33721b78cb99 https://chatgpt.com/c/66e3a628-2814-8012-a6c5-33721b78cb99
- h1fra 2y agoHaving read the full transcript I don't get how it counted 22 letters for mynznvaatzacdfoulxxz. It's nice that it corrected itself but a bit worrying
- kmeisthax 2y ago>We believe that a hidden chain of thought presents a unique opportunity for monitoring models. Assuming it is faithful and legible, the hidden chain of thought allows us to "read the mind" of the model and understand its thought process. For example, in the future we may wish to monitor the chain of thought for signs of manipulating the user. However, for this to work the model must have freedom to express its thoughts in unaltered form, so we cannot train any policy compliance or user preferences onto the chain of thought. We also do not want to make an unaligned chain of thought directly visible to users. >Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. We acknowledge this decision has disadvantages. We strive to partially make up for it by teaching the model to reproduce any useful ideas from the chain of thought in the answer. For the o1 model series we show a model-generated summary of the chain of thought. So, let's recap. We went from: - Weights-available research prototype with full scientific documentation (GPT-2) - Commercial-scale model with API access only, full scientific documentation (GPT-3) - Even bigger API-only model, tuned for chain-of-thought reasoning, minimal documentation on the implementation (GPT-4, 4v, 4o) - An API-only model tuned to generate unedited chain-of-thought, which will not be shown to the user, even though it'd be really useful to have (o1)
- dvt 2y agoIt's clear to me that OpenAI is quickly realizing they have no moat. Even this obfuscation of the chain-of-thought isn't really a moat. On top of CoT being pretty easy to implement and tweak, there's a serious push to on-device inference (which imo is the future), so the question is: will GPT-5 and beyond be really that much better than what we can run locally?
- falcor84 2y agoBased on their graphs of how quality scales well with compute cycles, I would expect that it would indeed continue to be that much better (unless you can afford the same compute locally).
- 2y ago
- jseip 2y agoLandmark. Wild. Beautiful. The singularity is nigh.
- noshitsherlock 2y agoThis is great. I've been wondering how we will revert back to an agrarian society! You know, beating our swords into plowshares; more leisure time, visiting with good people, getting to know their thoughts hopes and dreams, playing music together, taking time contemplating the vastness and beauty of the universe. We're about to come full circle; back to Eden. It all makes sense now.
- neta1337 2y agoIs there a new drug we need to know about?
- noshitsherlock 2y agoLife? I'm just thinking about what we can move on to now that the mundane tasks of life recede into the background. Things like artistry and craftsmanship, and exploration.
- wahnfrieden 2y agoAny word on whether this has enhanced Japanese support? They announced Japanese-specific models a while back that were never released.
- haolez 2y agoThis should also be good news for open weights models, right? Since OpenAI is basically saying "you can get very far with good prompts and some feedback loops".
- gwern 2y agoNo. It's bad news, because you can't see the rationale/search process that led to the final answer, just the final answer, and if training on the final answer were really that adequate, we wouldn't be here. It also is probably massively expensive compute-wise, much more so than simple unsupervised training on a corpus of question/answer pairs (because you have to generate the corpus by search first). It's also also bad news because reinforcement learning tends to be highly finicky and requires you to sweat the details and act like a professional, while open weight stuff tends to be produced by people for whom the phrase 'like herding cats' was coined, and so open source RL stuff is usually flakier than proprietary solutions (where it exists at all). They can do it for a few passion projects shared by many nerds, like chess or Go, but it takes a long time.
- haolez 2y ago> It also is probably massively expensive compute-wise, much more so than simple unsupervised training on a corpus of question/answer pairs (because you have to generate the corpus by search first). What do you mean? It sounds interesting.
- jiggawatts 2y ago“THERE ARE THREE R’S IN STRAWBERRY” - o1 I got that reference!
- m348e912 2y agoI have a straight forward task that no model has been able to successfully complete. The request is pretty basic. If anyone can get it to work, I'd like to know how and what model you're using. I tried it with gpt4o1 and after ~10 iterations of showing it the failed output, it still failed to come up with a one-line command to properly display results. Here it what I asked: Using a mac osx terminal and standard available tools, provide a command to update the output of netstat -an to show the fqdn of IP addresses listed in the result. This is what it came up with: netstat -an | awk '{for(i=1;i<=NF;i++){if($i~/^([0-9]+\.[0-9]+\.[0-9]+\.[0-9]+)(\.[0-9]+)?$/){split($i,a,".");ip=a[1]"."a[2]"."a[3]"."a[4];port=(length(a)>4?"."a[5]:"");cmd="dig +short -x "ip;cmd|getline h;close(cmd);if(h){sub(/\.$/,"",h);$i=h port}}}}1'
- OutOfHere 2y agoHave you tried `ss -ar`? You may have to install `ss`. It is standard on Linux.
- m348e912 2y agoNo I was trying to see it could use tools/binaries that come with MacOS.
- OutOfHere 2y agonetstat is now considered too old to be used in new code.
- m348e912 2y agoFair, but you'd think the latest most advanced model of GPT 4o1 (eg strawberry) would be able to successfully complete this task.
- m348e912 2y ago4o1 mini seems to have got it right. The trick is to give it minimal direction and let it do its thing. netstat -an | while IFS= read -r line; do ips=$(echo "$line" | grep -oE '([0-9]{1,3}\.){3}[0-9]{1,3}|([a-fA-F0-9]{1,4}:){1,7}[a-fA-F0-9]{1,4}'); for ip in $ips; do clean_ip=$(echo "$ip" | cut -d'%' -f1); fqdn=$(dig +short -x "$clean_ip" | grep '\.'); if [ -n "$fqdn" ]; then line=$(echo "$line" | sed "s/$ip/$fqdn/g"); fi; done; echo "$line"; done
- Eextra953 2y agoWhat is interesting to me is that there is no difference in the AP English lit/lang exams. Why did chain-of-thought produce negligible improvements in this area?
- munchler 2y agoI would guess because there is not much problem-solving required in that domain. There’s less of a “right answer” to reason towards.
- gwern 2y agoI think there may also be a lack of specification there. When you get more demanding and require more, the creative writing seems to be better. Like it does much better at things like sestinas. For all of those questions, there's probably a lot of unspecified criteria you could say makes an answer better or worse, but you don't, so the first solution appears adequate.
- deleted 2y ago[deleted]
- deemahstro 2y agoStop fooling around with stories about AI taking jobs from programmers. Which programmers exactly??? Creators of idiotic web pages? Nobody in their right mind would push generated code into a financial system, medical equipment or autonomous transport. Template web pages and configuration files are not the entire IT industry. In addition, AI is good at tasks for which there are millions of examples. 20 times I asked to generate a PowerShell script, 20 times it was generated incorrectly. Because, unlike Bash, there are far fewer examples on the Internet. How will AI generate code for complex systems with business logic that it has no idea about? AI is not able to generate, develop and change complex information systems.
- suziemanul 2y agoIn this video Lukasz Kaiser, one of the main co-authors of o1, talks about how to get to reasoning. I hope this may be useful context for some. https://youtu.be/_7VirEqCZ4g?si=vrV9FrLgIhvNcVUr https://youtu.be/_7VirEqCZ4g?si=vrV9FrLgIhvNcVUr
- forgotthepasswd 2y agoI had trouble in the past to make any model give me accurate unix epochs for specific dates. I just went to GPT-4o (via DDG) and asked three questions: 1. Please give me the unix epoch for September 1, 2020 at 1:00 GMT. > 1598913600 2. Please give me the unix epoch for September 1, 2020 at 1:00 GMT. Before reaching the conclusion of the answer, please output the entire chain of thought, your reasoning, and the maths you're doing, until your arrive at (and output) the result. Then, after you arrive at the result, make an extra effort to continue, and do the analysis backwards (as if you were writing a unit test for the result you achieved), to verify that your result is indeed correct. > 1598922000 3. Please give me the unix epoch for September 1, 2020 at 1:00 GMT. Then, after you arrive at the result, make an extra effort to continue, and do the analysis backwards (as if you were writing a unit test for the result you achieved), to verify that your result is indeed correct. > 1598913600
- Alifatisk 2y agoNo need for llms to do that ruby -r time -e 'puts Time.parse("2020-09-01 01:00:00 +00:00").to_i'
- deleted 2y ago[deleted]
- tagawa 2y agoQuick link for checking the result: https://duckduckgo.com/?q=timestamp+1598922000&ia=answer https://duckduckgo.com/?q=timestamp+1598922000&ia=answer
- jewel 2y agoWhen I give it that same prompt, it writes a python program and then executes it to find the answer: https://chatgpt.com/share/66e35a15-602c-8011-a2cb-0a83be35b832 https://chatgpt.com/share/66e35a15-602c-8011-a2cb-0a83be35b8...
- dada5000 2y agoI find shorter responses > longer responses. Anyone share the same consensus? for example in gpt-4o I often append '(reply short)' at the end of my requests. with the o1 models I append 'reply in 20 words' and it gives way better answers.
- Doorknob8479 2y agoWhy so much hate? They're doing their best. This is the state of progress in the field so far. The best minds are racing to innovate. The benchmarks are impressive nonetheless. Give them a break. At the end of the day, they built the chatbot who's saving your ass each day ever since.
- bamboozled 2y agoHaven't used ChatGPT* in over 6 months, not saving my ass at all.
- jdthedisciple 2y agoI bet you've still used other models that were inspired by GPT.
- bamboozled 2y agoI've used co-pilot, I turned it off, kept suggesting nonsense.
- evilfred 2y agonot saving my ass, I never needed one professionally. OpenAI is shovelling money into a furnace, I expect them to be assimilated into Microsoft soon.
- commodoreboxer 2y agoI think you're overestimating LLM usage.
- neta1337 2y agoRarely using it at work, seems you are overestimating
- resters 2y agoI tested o1-preview on some coding stuff I've been using gpt-4o for. I am not impressed. The new, more intentional chain of thought logic is apparently not something it can meaningfully apply to a non-trivial codebase. Sadly I think this OpenAI announcement is hot air. I am now (unfortunately) much less enthusiastic about upcoming OpenAI announcements. This is the first one that has been extremely underwhelming (though the big announcement about structured responses (months after it had already been supported nearly identically via JSONSchema) was in hindsight also hot air. I think OpenAI is making the same mistake Google made with the search interface. Rather than considering it a command line to be mastered, Google optimized to generate better results for someone who had no mastery of how to type a search phrase. Similarly, OpenAI is optimizing for someone who doesn't know how to interact with a context-limited LLM. Sure it helps the low end, but based on my initial testing this is not going to be helpful to anyone who had already come to understand how to create good prompts. What is needed is the ability for the LLM to create a useful, ongoing meta-context for the conversation so that it doesn't make stupid mistakes and omissions. I was really hoping OpenAI would have something like this ready for use.
- jdthedisciple 2y agoYour case would be more convincing by an example. Though o1 did fail at the puzzle in my profile. Maybe it's just tougher than even, its author, I had assumed...
- egorfine 2y agoI have tested o1-preview on a couple of coding tasks and I am impressed. I am looking at a TypeScript project with quite an amount of type gymnastics and a particular line of code did not validate with tsc no matter what I have tried. I copy pasted the whole context into o1-preview and it told me what is likely the error I am seeing (and it was a spot on correct letter-by-letter error message including my variable names), explained the problem and provided two solutions, both of which immediately worked. Another test was I have pasted a smart contract in solidity and naively asked to identify vulnerabilities. It thought for more than a minute and then provided a detailed report of what could go wrong. Much, much deeper than any previous model could do. (No vulnerabilities found because my code is perfect, but that's another story).
- sturza 2y agoIt seems like it's just a lot of prompting the same old models in the background, no "reasoning" there. My age old test is "draw a hand in ascii" - i've had no success with any model yet.
- ActionHank 2y agoIt seems like their current strat is to farm token count as much as possible. 1. Don't give the full answer on first request. 2. Each response needs to be the wordiest thing possible. 3. Now just talk to yourself and burn tokens, probably in the wordiest way possible again. 4. ??? 5. Profit Guaranteed they have number of tokens billed as a KPI somewhere.
- evilfred 2y agoit still fails at logic puzzles https://x.com/colin_fraser/status/1834334418007457897 https://x.com/colin_fraser/status/1834334418007457897
- evilfred 2y agoand listing state names with the letter 'a' https://x.com/edzitron/status/1834329704125661446 https://x.com/edzitron/status/1834329704125661446
- fragmede 2y agoWeird, it works to say the father when I try it: https://chatgpt.com/share/66e3601f-4bec-8009-ac0c-57bfa4f0592c https://chatgpt.com/share/66e3601f-4bec-8009-ac0c-57bfa4f059... And also works on this variation: https://chatgpt.com/share/66e35f0e-6c98-8009-a128-e9ac677480fd https://chatgpt.com/share/66e35f0e-6c98-8009-a128-e9ac677480...
- Slow_Hand 2y agoI have a question. The video demos for this all mention that the o1 model is taking it's time to think through the problem before answering. How does this functionally differ from - say - GPT-4 running it's algorithm, waiting five seconds and then revealing the output? That part is not clear to me.
- ActionHank 2y agoIt is recursively "talking" to itself to plan and then refine the answer.
- davesque 2y ago"after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users" ...umm. Am I the only one who feels like this takes away much of the value proposition, and that it also runs heavily against their stated safety goals? My dream is to interact with tools like this to learn, not just to be told an answer. This just feels very dark. They're not doing much to build trust here.
- shreezus 2y agoAdvanced reasoning will pave the way for recursive self-improving models & agents. These capabilities will enable data flywheels, error-correcting agentic behaviors, & self-reflection (agents understanding the implications of their actions, both individually & cooperatively). Things will get extremely interesting and we're incredibly fortunate to be witnessing what's happening.
- AI_beffr 2y agothis is completely illogical. this is like gambling your life savings and as the die are rolling you say "i am incredibly fortunate to be witnessing this." like, you need to know the outcome before you know whether it was fortunate or unfortunate... this could be the most unfortunate thing that has ever happened in history.
- samanator 2y agoI just tested o1-preview on the "How many r's are in strawberry?" question. It answers correctly!
- bad_username 2y agoPrompt: > Alice, who is an immortal robotic observer, orbits a black hole on board a spaceship. Bob exits the spaceship and falls into the black hole. Alice sees Bob on the edge of the event horizon, getting closer and closer to it, but from her frame of reference Bob will remain forever observable (in principle) outside the horizon. > > A trillion year has passed, and Alice observes that the black hole is now relatively rapidly shrinking due to the Hawking radiation. How will Alice be observing the "frozen" Bob as the hole shrinks? > > The black hole finally evaporated completely. Where is Bob now? O1-preview spits out the same nonsense that 4o does, telling that as the horizon of the black hole shrinks, it gets closer to Bob's apparent position. I realize that.the prompt is essentily asking to solve the famous unsolved problem in physics (black hole information paradox), but there's no need to be so confused with basic geometry of the situation.
- throwaway314155 2y agoOut of curiosity, what answer to that would you find acceptable? I don't know relativity well enough to even speculate.
- alkyon 2y agoMaybe this is improvement in some areas, still I got spurious reasoning and inability to connect three simple facts: Yes, "el presente acta de nacimiento" is correct in Spanish. Explanation: "Acta" is a feminine noun that begins with a stressed "a" sound. In Spanish, when a feminine singular noun starts with a stressed "a" or "ha", the definite article "la" is replaced with "el" to facilitate pronunciation. However, the noun remains feminine. Adjectives and modifiers that accompany the noun "acta" should agree in feminine gender and singular number. In this case, "presente" is an adjective that has the same form for both masculine and feminine singular nouns. So, combining these rules: "El" (definite article used before feminine nouns starting with stressed "a") "Presente" (adjective agreeing in feminine singular) "Acta de nacimiento" (feminine noun with its complement) Therefore, "el presente acta de nacimiento" is grammatically correct.
- anthk 2y agoNative Spanish speaker here. No, it isn't. When a word stays in the middle of 'La' plus a noun starting with 'a', the cacophony is null now, thus, you can perfectly use (if not mandatory) "la presente acta". Proof: https://www.elcastellano.org/francisco-jos%C3%A9-d%C3%ADaz-%C3%ADb%C3%A1%C3%B1ez https://www.elcastellano.org/francisco-jos%C3%A9-d%C3%ADaz-%...
- alkyon 2y agoyes, exactly - this is my point
- quantisan 2y agoAmazing! OpenAI figured out how to scale inference. https://arxiv.org/abs/2407.21787 https://arxiv.org/abs/2407.21787 show how using more compute during inference can outperform much larger models in tasks like math problems I wonder how do they decide when to stop these Chain of Thought for each query? As anyone that played with agents can attest, LLMs can talk with themselves forever.
- nemo44x 2y agoBesides chat bits what viable products are being made with LLMs besides APIs into LLMs?
- sohamgovande 2y agothe newest scaling law: inference-time compute.
- deleted 2y ago[deleted]
- digitcatphd 2y agoI tested various Math Olympiad questions with Claude sonnet 3.5 and they all arrived at the correct solution. o1's solution was a bit better formulated, in some circumstances, but sonnet 3.5 was nearly instant.
- la64710 2y agoCan we please stop using the word “think” like o1 thinks before it answers. I doubt we man the same when someone says a human thinks vs o1 thinks. When I say I think “red” I am sure the word think means something completely different than when you say openai thinks red. I am not saying one is superior than the other but maybe as humans we can use a different set of terminology for the AI activities.
- natch 2y agoI tried it with a cipher text that ChatGPT4o flailed with. Recently I tried the same cipher with Claude Sonnet 3.5 and it solved it quickly and perfectly. Just now tried with ChatGPT o1 preview and it totally failed. Based on just this one test, Claude is still way ahead. ChatGPT also showed a comical (possibly just fake filler material) journey of things it supposedly tried including several rewordings of "rethinking my approach." It remarkably never showed that it was trying common word patterns (other than one and two letters) nor did it look for "the" and other "th" words nor did it ever say that it was trying to match letter patterns. I told it upfront as a hint that the text was in English and was not a quote. The plaintext was one paragraph of layman-level material on a technical topic including a foreign name, text that has never appeared on the Internet or dark web. Pretty easy cipher with a lot of ways to get in, but nope, and super slow, where Claude was not only snappy but nailed it and explained itself.
- spoonfeeder006 2y agoSo how is the internal chain of thought represented anyhow? What does it look like when someone sees it?
- delusional 2y agoGreat, yet another step towards the inevitable conclusion. Now I'm not just being asked to outsource my thinking to my computer, but instead to a black box operated by a for-profit company for the benefit of Microsoft. Not only will they not tell me the whole reasoning chain, they wont even tell me how they came up with it. Tell me, users of this tool. What's even are you? If you've outsourced your thinking to a corporation, what happens to your unique perspective? your blend of circumstance and upbringing? Are you really OK being reduced to meaningless computation and worthless weights. Don't you want to be something more?
- OutOfHere 2y ago> What's even are you? An accelerator of reaching the Singularity. This is something more.
- delusional 2y agoYou realize that you're not going inside the computer right? At best you're going to create a simulacrum of you. Something that looks, talks, and acts like you. It's never going to actually be you. You're going to be stuck out here with the rest of us, in whatever world we create in pursuit of the singularity suicide cult.
- OutOfHere 2y agoMy friend, it has nothing to do with going inside a computer. Do not confuse the Singularity with mind uploading which is a distinct concept. The singularity has to do with technology acceleration, and with the inability to predict what lies beyond it. As such, it has nothing to do with any suicide cult. Please stop spreading nonsense about it. I do care about life in the physical world, not about a digital life.
- wrath224 2y agoTrying this on a few hard problems on PicoGYM and holy heck I'm impressed. I had to give it a hint but that's the same info a human would have. Problem was Sequences (crypto) hard. https://chatgpt.com/share/66e363d8-5a7c-8000-9a24-8f5eef445136 https://chatgpt.com/share/66e363d8-5a7c-8000-9a24-8f5eef4451... Heh... GPT-4o also solved this after I tried and gave it about the same examples. Need to further test but it's promising !
- kypro 2y agoReminder that it's still not too late to change the direction of progress. We still have time to demand that our politicians put the breaks on AI data centres and end this insanity. When AI exceeds humans at all tasks humans become economically useless. People who are economically useless are also politically powerless, because resources are power. Democracy works because the people (labourers) collectivised hold a monopoly on the production and ownership of resources. If the state does something you don't like you can strike or refuse to offer your labour to a corrupt system. A state must therefore seek your compliance. Democracies do this by given people want they want. Authoritarian regimes might seek compliance in other ways. But what is certain is that in a post-AGI world our leaders can be corrupt as they like because people can't do anything. And this is obvious when you think about it... What power does a child or a disable person hold over you? People who have no ability to create or amass resources depend on their beneficiaries for everything including basics like food and shelter. If you as a parent do not give your child resources, they die. But your child does not hold this power over you. In fact they hold no power over you because they cannot withhold any resources from you. In a post-AGI world the state would not depend on labourers for resources, jobless labourers would instead depend on the state. If the state does not provide for you like you provide for your children, you and your family will die. In a good outcome where humans can control the AGI, you and your family will become subjects to the whims of state. You and your children will suffer as the political corruption inevitably arises. In a bad outcome the AGI will do to cities what humans did to forests. And AGI will treat humans like humans treat animals. Perhaps we don't seek the destruction of the natural environment and the habitats of animals, but woodland and buffalo are sure inconvenient when building a super highway. We can all agree there will be no jobs for our children. Even if you're an "AI optimist" we probably still agree that our kids will have no purpose. This alone should be bad enough, but if I'm right then there will be no future for them at all. I will not apologise for my concern about AGI and our clear progress towards that end. It is not my fault if others cannot see the path I seem to see so clearly. I cannot simply be quiet about this because there's too much at stake. If you agree with me at all I urge you to not be either. Our children can have a great future if we allow them to have it. We don't have long, but we do still have time left.
- kgeist 2y agoAsked it to write PyTorch code which trains an LLM and it produced 23 steps in 62 seconds. With gpt4-o it immediately failed with random errors like mismatched tensor shapes and stuff like that. The code produced by gpt-o1 seemed to work for some time but after some training time it produced mismatched batch sizes. Also, gpt-o1 enabled cuda by itself while for gpt-4o, I had to specifically spell it out (it always used cpu). However, showing gpt-o1 the error output resulted in broken code again. I noticed that back-and-forth iteration when it makes mistakes has worse experience because now there's always 30-60 sec time delays. I had to have 5 back-and-forths before it produced something which does not crash (just like gpt-4o). I also suspect too many tokens inside the CoT context can make it accidentally forget some stuff. So there's some improvement, but we're still not there...
- koreth1 2y agoThe performance on programming tasks is impressive, but I think the limited context window is still a big problem. Very few of my day-to-day coding tasks are, "Implement a completely new program that does XYZ," but more like, "Modify a sizable existing code base to do XYZ in a way that's consistent with its existing data model and architecture." And the only way to do those kinds of tasks is to have enough context about the existing code base to know where everything should go and what existing patterns to follow. But regardless, this does look like a significant step forward.
- suchar 2y agoI would imagine that good IDE integration would summarise each module/file/function and feed high-level project overview (best case: with business project description provided by the user) and during CoT process model would be able to ask about more details (specific file/class/function). Humans work on abstractions and I see no reason to believe that models cannot do the same
- beaugunderson 2y agothe cipher example is impressive on the surface, but I threw a couple of my toy questions at o1-preview and it still hallucinates a bunch of nonsense (but now uses more electricity to do so).
- Havoc 2y agoo1 Maybe they should spend some of their billions on marketing people. Gpt4o was a stretch. Wtf is o1
- OutOfHere 2y agoTo me it looks like they think this is the future of how all models should be, so they're restarting the numbering. This is what I suspect. The o is for omni.
- franze 2y agoChatGPT is now a better coder than I ever was.
- adamtaylor_13 2y agoLaughing at the comparison to "4o" as if that model even holds a candle to GPT-4. 4o is _cheaper_—it's nowhere near as powerful as GPT-4, as much as OpenAI would like it to be.
- aantix 2y agoFeels like the challenge here is to somehow convey to the end user, how the quality of output is so much better.
- sys32768 2y agoTime to fire up System Shock 2: > Look at you, hacker: a pathetic creature of meat and bone, panting and sweating as you run through my corridors. How can you challenge a perfect, immortal machine?
- scotty79 2y agoTransformers have exactly two strengths. None of them is "attention". Attention could be replaced with any arbitrary division of the network and it would learn just as well. First true strength is obvious, it's that they are parallelisable. This is a side effect of people fixating on attention. If they came up with any other structure that results in the same level of parallelisability it would be just as good. Second strong side is more elusive to many people. It's the context window. Because the network is not ran just once but once for every word it doesn't have to solve a problem in one step. It can iterate while writing down intermediate variables and accessing them. The dumb thing so far was that it was required to produce the answer starting with the first token it was allowed to write down. So to actually write down the information it needs on the next iteration it had to disguise it as a part of the answer. So naturally the next step is to allow it to just write down whatever it pleases and iterate freely until it's ready to start giving us the answer. It's still seriously suboptimal that what it is allowed to write down has to be translated to tokens and back but I see how this might make things easier for humans for training and explainability. But you can rest assured that at some point this "chain of thought" will become just chain of full output states of the network, not necessarily corresponding to any tokens. So congrats to researchers that they found out that their billion dollar Turing machine benefits from having a tape it can use for more than just printing out the output. PS There's another advantage of transformers but I can't tell how important it is. It's the "shortcuts" from earlier layers to way deeper ones bypassing the ones along the way. Obviously network would be more capable if every neuron was connected with every neuron in every preceding layer but we don't have hardware for that so some sprinkled "shortcuts" might be a reasonable compromise that might make network less crippled than MLP. Given all that I'm not surprised at all with the direction openai took and the gains it achieved.
- ttul 2y agoI've given this a test run on some email threads, asking the model to extract the positions and requirements of each person in a lengthy and convoluted discussion. It absolutely nailed the result, far exceeding what Claude 3.5 Sonnet was capable of -- my previous goto model for such analysis work. I also used it to apply APA style guidelines to various parts of a document and it executed the job flawlessly and with a tighter finesse than Claude. Claude's response was lengthier - correct, but unnecessarily long. gpt-o1-preview combined several logically-related bullets into a single bullet, showing how chain of thought reasoning gives the model more time to comprehend things and product a result that is not just correct, but "really correct".
- joshhug 2y agoI just tried o1, and it did pretty well with understanding this minor issue with subtitles on a Dutch TV show we were watching. I asked it "I was watching a show and in the subtitles an umlaut u was rendered as 1/4, i.e. a single character that said 1/4. Why would this happen?" and it gave a pretty thorough explanation of exactly which encoding issue was to blame. https://chatgpt.com/share/66e37145-72bc-800a-be7b-f7c76471a1bd https://chatgpt.com/share/66e37145-72bc-800a-be7b-f7c76471a1...
- piker 2y agoA common problem, no doubt, with a lot of training context. But man. What a time to be alive.
- takoid 2y ago4o’s answer seems sufficient, though it provides less detail than o1. https://chatgpt.com/share/66e373d7-7814-8009-86c3-1ce549ca2e1d https://chatgpt.com/share/66e373d7-7814-8009-86c3-1ce549ca2e...
- karmasimida 2y agoDamn, the model really goes to length to those trivial but hard problems. Impressive
- silveryfu 2y agoAfter playing with it on ChatGPT this morning, it seems a reasonable strategy of using the o1 model is to: - If your request requires reasoning, switch to o1 model. - If not, switch to 4o model. This applies to both across chat sessions and within the same session (yes, we can switch between models within the same session and it looks like down the road OpenAI is gonna support automatic model switching). Based on my experience, this will actually improve the perceived response quality -- o1 and 4o are rather complementary to each other rather than replacement.
- zamadatix 2y agoGiven the rate limits are 30 reqs/week most probably want to start with: - Try it a bit with 4o, see if you're getting anywhere - Switch to the new o1 model if it's just not working out, take your improved base prompt and follow ups with you so it only counts as 1 req
- energy123 2y agoThis was mentioned in OpenAI's report. People rated o1 as the same or worse than GPT-4o if the prompt didn't require reasoning, like on personal writing tasks.
- MattDaEskimo 2y agoWhat's the precedent set here? Models that hide away their reasoning and only display the output, charging whatever tokens they'd like? This is not a good release on any front.
- schappim 2y agoIf you’re using the API and are on tier 4, don’t bother adding more credits to move up to tier 5. I did this, and while my rate limits increased, the o1-preview / o1-mini model still wasn’t available.
- deleted 2y ago[deleted]
- ethanmitchell87 2y agoslightly offtopic, but openai having anti scraping / bot check on the blog is pretty funny
- natch 2y agoIn practice, this implementation (through the Chat UI) is scary bad. It actively lies about what it is doing. This is what I am seeing. Proactive, open, deceit. I can't even begin to think of all the ways this could go wrong, but it gives me a really bad feeling.
- xpl 2y agoFolks who say "LLMs can't reason", what now? Have we moved the goalposts yet?
- darajava 2y agoWho said that?
- xpl 2y agoLiterally in every HN post about AI, it is a common pattern in the comments section... "LLMs are simply predicting next token, it is not thinking/reasoning/etc." "LLMs can't reason, only humans can reason" "We will never get to AGI using LLMs" It's interesting that I don't see much of that sentiment in this post. So, maybe LLMs can reason after all? :)
- andrewchambers 2y agoWhats interesting is that with more time it can create more accurate answers which means it can be used to generate its own training data.
- ziofill 2y agoQuestion for those who do have access: how is it?
- 0xstackie 2y agoI think openai introduced the o1 model because reflection 70b inspired them. Getting them needed a new message to fill the gap for such a long time
- roshankhan28 2y agoI have also heard they are launching a AI called strawberry. If you pay attention, there is a specific reason why they have named it strawberry. if you ask chat gpt 4o, how many r's in the word strawberry, it will give answer as 2. still to this day it will answer same. the model is not able to reason. thats why a reasoning model is being launched. this is one of the reason apart from many other reasons.
- deleted 2y ago[deleted]
- throwawaylolx 2y ago"Learn to reason like a robot"
- harisec 2y agoI asked a few “hard” questions and compared o1 with claude. https://github.com/harisec/o1-vs-claude https://github.com/harisec/o1-vs-claude
- Aissen 2y agoIt's interesting that OpenAI has literally applied and automated one of their advice from the "Prompt engineering" guide: Give the model time to "think" https://platform.openai.com/docs/guides/prompt-engineering/give-the-model-time-to-think https://platform.openai.com/docs/guides/prompt-engineering/g...
- max_entropy 2y agoIs there a paper available?
- ValentinA23 2y agoQuiet-STaR: Language Models Can Teach Themselves to Think Before Speaking https://arxiv.org/abs/2403.09629 https://arxiv.org/abs/2403.09629
- deleted 2y ago[deleted]
- diedyesterday 2y agoSam Altman and OpenAI are following the example of Celebrimbor it seems. And I love what may come next...