8 ms·
GPT o3 frequently fabricates actions, then elaborately justifies these actions
- YetAnotherNick 1y agoI wish there are benchmarks for these scenarios. Anyone who has used LLMs know that they are very different from human. And after certain context, it become irritating to talk to these LLMs. I don't want my LLM to excel in IMO or codeforces. I want it to understand my significantly easier but complex to state problem, think of solutions, understand its own issues and resolve it, rather than be passive agressive.
- otabdeveloper4 1y agoLLMs can't think. They are not rational actors, they can only generate plausible-looking texts.
- johnisgood 1y agoMaybe so, but they boost my coding productivity, so why not? (Not the mentioned LLMs here though.) I do the rational acting, and it does the rest.
- j_maffe 1y agoYou're being reductive. A system should be evaluated on how its measurable properties more than anything else.
- discreteevent 1y agoBeing "reductive" is how we got where we are today. We try to form hypotheses about things so that we can reduce them to their simplest model. This understanding then leads to massive gains. We've been doing this ever since we have observed things like the behavior of animals in order that we could hunt them more easily. In the same way it helps a lot to try to understand what the correct model of an AI is in order that we can use it more productively. Certainly based on it's 'measurable properties' it does not behave like a reasonable human being. Some of the time it does, some of the time it goes completely off the rails. So there must be some other model that is more useful. "They are not rational actors, they can only generate plausible-looking texts." - seems to be more useful to me. "They are rational actors" - would be more like magical thinking which is not what got us to where we are today.
- andrepd 1y ago"Benchmarks" in AI are hilarious. These tools can't even solve problems which are moderately more difficult than something that has a geeks4geeks page, but according to these benchmarks they are all IOI gold medallists. What gives?
- delusional 1y agoThe benchmarks are created by humans. So are the training sets. It turns out the sorts of problems that humans like to benchmark with are also the sorts of problems humans like to discuss wherever that training set was scraped. Well that and the whole field is filled with AI hypemen who "contribute" by asking ChatGPT about the quality and validity of some other GPT response.
- ramesh31 1y agoReasoning models are complete nonsense in the face of custom agents. I would love to be proven wrong here.
- latexr 1y ago> These behaviors are surprising. It seems that despite being incredibly powerful at solving math and coding tasks, o3 is not by default truthful about its capabilities. It is only surprising to those who refuse to understand how LLMs work and continue to anthropomorphise them. There is no being “truthful” here, the model has no concept of right or wrong, true or false. It’s not “lying” to you, it’s spitting out text. It just so happens that sometimes that non-deterministic text aligns with reality, but you don’t really know when and neither does the model.
- tezza 1y agoPrecisely. The tools often hallucinate: including in its instructions higher up even before your prompt portion. Also the behind the scenes stuff not show to the user during reasoning. You see binary failures all the time when doing function calls or JSON outputs. That is… “please call this function” … does not call function “calling JSON endpoint”… does not emit JSON so from the article the tool generates hallucinations that the tool has used external stuff: but that was entirely fictitious. it does not know that this tool usage was fictitious and then sticks by its guns. The workaround is to have verification steps, throw away “bad” answers. Instead of expecting one true output, expect a stream of results which have a yield (agriculture) of a certain amount. say 95% work, 5% garbage. never consider the results truly accurate, just “accurate enough”. Verify always
- atoav 1y agoAs an electrical engineer it is absolutely amazing how much LLMs suck at describing electrical circuits. It is somewhat ok with natural language, which works for the simplest circuits. For more complex stuff Chatgpt (regardless of model) seems to default to absolutely nonsensical ASCII circuit diagrams, you can ask it to list each part with each terminal and describe the connections to other parts and terminals and it will fail spectacularly with missing parts, missing terminals, parts no one ever heard of, short circuits, dangling nodes with no use.. If tou ask it to draw a schematic thigns somehow get even worse. But what it is good at is proposing ideas. So if you want to do a thing that could be solved by using a Gilbert cell, the chances it might mention a Gilbert Cell are realistically there. But I am already having students coming by with LLM slob circuits asking why the don't work..
- glial 1y agoI enjoy watching newer-generation models exhibit symptoms that echo features of human cognition. This particular one is reminiscent of the confabulation seen in split-brain patients, e.g. https://www.edge.org/response-detail/11513 https://www.edge.org/response-detail/11513
- SillyUsername 1y agoo3 has been the worst model of the new 3 for me. Ask it to create a Typescript server side hello world. It produces a JS example. Telling it that's incorrect (but no more detail) results in it iterating all sorts of mistakes. In 20 iterations it never once asked me what was incorrect. In contrast, o4-mini asked me after 5, o4-mini-high asked me after 1, but narrowed the question to "is it incorrect due to choice of runtime?" rather than "what's incorrect?" I told it to "ask the right question" based on my statement ("it is incorrect") and it correctly asked "what is wrong with it?" before I pointed out no Typescript types. This is the critical thinking we need not just reasoning (incorrectly).
- echoangle 1y ago> Ask it to create a Typescript server side hello world. It produces a JS example. Well TS is a strict superset of JS so it’s technically correct (which is the best kind of correct) to produce JS when asked for a TS version. So you’re the one that’s wrong.
- redox99 1y agoHe's not wrong. If the model doesn't give you what you want, it's a worthless model. If the model is like the genie from the lamp, and gives you a shitty but technically correct answer, it's really bad.
- echoangle 1y ago> If the model doesn't give you what you want, it's a worthless model. Yeah, if you’re into playing stupid mind games while not even being right. If you stick to just voicing your needs, it’s fine. And I don’t think the TS/JS story shows a lack of reasoning that would be relevant for other use cases.
- hatefulmoron 1y ago> Yeah, if you’re into playing stupid mind games while not even being right. If I ask questions outside of the things I already know about (probably pretty common, right?), it's not playing mind games. It's only a 'gotcha' question with the added context, otherwise it's just someone asking a question and getting back a Monkey's Paw answer: "aha! See, it's technically a subset of TS.." You might as well give it equal credit for code that doesn't compile correctly, since the author didn't explicitly ask.
- anshumankmr 1y agoIs it just me or it feels like bit of a dissapointment? I have been using it for some hours now, and its needlessly convoluting the code.
- jjani 1y agoIt feels similar to Llama4 - rushed. Sonnet had been king for at least 6 months, then Gemini 2.5 Pro recently raised the bar. They felt they had to respond. Ghibli memes are great, but not at the cost of losing the whole enterprise market. Currently for B2C, there's almost no lock in. Users can switch to a better app/model at very little cost. With B2B it's different, a product built on Sonnet generally isn't just going to switch to an OA model overnight unless there's huge benefits. OA will want a piece of that lock-in pie, which they'd been losing at a very rapid pace. Whether their new models solve that remains to be seen. To me, actually building products on top of these models, I still don't see much reason to use any of their models. From all testing I've been doing over the last 2 days, they don't seem particularly competitive. Potentially 4.1 or o4-mini for certain tasks, but whether they beat e.g. Deepseek v3 currently isn't clear-cut.
- anshumankmr 1y agoYeah. God knows. I was really surprised to see the Fchollet's benchmark being aced months ago, but whatever their internal QA was perhaps lacking. I was asking some fairly simple code, that too in Python, using Scikit learn for which I presume there must be a lot of training data, it for some reason, changed the casing of the columns, and didn't follow my instructions as I asked it, cause the function was being rewritten to reduce bloat, along with other random things I didn't ask it for.
- jjani 1y agoEveryone games the benchmarks, but a lot is pointing towards both Meta and OpenAI going to even further lengths than the others.
- 1y ago
- jjani 1y agoPower user here, working with these models (the whole gamut) side-by-side on a large range of tasks has been my daily work since they came out. I can vouch that this is extremely characteristic of o3-mini compared to competing models (Claude, Gemini) and previous OA models (3.5, 4o). Compared to those, o3-mini clearly has less of the "the user is always right" training. This is almost certainly intentional. At times, this can be useful - it's more willing to call you out when you're wrong, and less likely to agree with something just because you suggested it. But this excessive stubbornness is the great downside, and it's been so prevalent that I stopped using o3-mini. I haven't had enough time with o3 yet, but if it is indeed an evolution of o3-mini, it comes at no surprise it's very bad for this as well.
- dstick 1y agoSounds like we're getting closer and closer to an AI that acts like a human ;-)
- vitorgrs 1y agoYes! I always ask these models a simple question, that all models don't have the right answers. "List of mayors of my City X". All OF THEM, get it wrong. Hallucinate the names, wrong dates, etc. The list is on wikipedia, and for sure they trained on that data, but they are not able to answer properly. o3-mini? It just says it doesn't know lol
- jjani 1y agoYeah, that's the big upside for sure - it baseline hallucinates less. But when it does, it's very assertive in gaslighting you that it's hallucination is in fact the truth, it can't "fix" its own errors. I've found this tradeoff not to be worth it for general use.
- LZ_Khan 1y agoUm.. wasn't this what was mentioned was going to happen in AI 2027? "In a few rigged demos, it even lies in more serious ways, like hiding evidence that it failed on a task, in order to get better ratings."
- bjackman 1y agoI don't understand why the UIs don't make this obvious. When the model runs code, why can't the system just show us the code and its output, in a special UI widget that the model can't generate any other way? Then if it says "I ran this code and it says X" we can easily verify. This is a big part of the reason I want LLMs to run code. Weirdly I have seen Gemini write code and make claims about the output. I can see the code, the claims it makes about the output are correct. I do not think it could make these correct claims without running the code. But the UI doesn't show me this. To verify it, I have to run the code myself. This makes the whole feature way less valuable and I don't understand why!
- TobiWestside 1y agoI'm confused - the post says "o3 does not have access to a coding tool". However, OpenAI mentiones a Python tool multiple times in the system card [1], e.g.: "OpenAI o3 and OpenAI o4-mini combine state-of-the-art reasoning with full tool capabilities—web browsing, Python, [...]" "The models use tools in their chains of thought to augment their capabilities; for example, cropping or transforming images, searching the web, or using Python to analyze data during their thought process." I interpreted this to mean o3 does have access to a tool that enables it to run code. Is my understanding wrong? [1] https://openai.com/index/o3-o4-mini-system-card/ https://openai.com/index/o3-o4-mini-system-card/
- ddjohnson 1y agoOne of the blog post authors here! We evaluated o3 through the API, where the model does not have access to any specific built-in tools (although it does have the capability to use tools, and allows you to provide your own tools). This is different than when using o3 through the ChatGPT UI, where it does have a built-in tool to run code. (Interestingly, even in the ChatGPT UI the o3 model will sometimes state that it ran code on its personal MacBook Pro M2! https://x.com/TransluceAI/status/1912617941725847841 https://x.com/TransluceAI/status/1912617941725847841)
- TobiWestside 1y agoI see, thanks for the clarification!
- jlaternman 1y agoJust throwing this out there. Is it possible that in some way, it does have a MacBook Pro M2? For example, that the tools the ChatGPT UI have are exposed to it through access to one, which is can run whatever tools it wants through? That might actually be quite a sensible way to expose a “tools” UI to an LLM. If what it’s saying is technically accurate on the “flagship product” (ChatGPT), it could be the API version is simply confused about its differences (no access to tools).
- rsynnott 1y agoSo people keep claiming that these things are like junior engineers, and, increasingly, it seems as if they are instead like the worst possible _stereotype_ of junior engineers.