6 ms·
Is it me, or did it still get at least three placements of components (RAM and PCIe slots, plus it's DisplayPort and not HDMI) in the motherboard image[0] compl
by breakingcups 10mo ago
Is it me, or did it still get at least three placements of components (RAM and PCIe slots, plus it's DisplayPort and not HDMI) in the motherboard image[0] completely wrong? Why would they use that as a promotional image?
0: https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3ft1oP/2ee4e2a98c4725fab4e9eada8d38b6ad/image_8.png?w=1920&q=90&fm=webp https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3f...
- dolmen 10mo agoNot that bad compared to product images seen on AliExpress.
- timerol 10mo agoAlso a "stacked pair" of USB type-A ports, when there are clearly 4
- tedsanders 10mo agoYep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.
- d--b 10mo ago[flagged]
- wilg 10mo agoWhat did Sam Altman say? Or is this more of a vague impression thing?
- honeycrispy 10mo agoNot sure what you mean, Altman does that fake-humility thing all the time. It's a marketing trick; show honesty in areas that don't have much business impact so the public will trust you when you stretch the truth in areas that do (AGI cough).
- d--b 10mo agoI'm confident that GP is good faithed though. Maybe I am falling for it. Who knows? It doesn't really matter, I just wanted to be nice to the guy. It takes some balls posting as OpenAi employee here, and I wish we heard from them more often, as I am pretty sure all of them lurk around.
- rvnx 10mo agoIt's the only reasonable choice you can make. As an employee with stock options you do not want to get trashed on Hackernews because this affects your income directly if you try to conduct a secondary share sale or plan to hold until IPO. Once the IPO is done, and the lockup period is expired, then a lot of employees are planning to sell their shares. But until that, even if the product is behind competitors there is no way you can admit it without putting your money at risk.
- Esophagus4 10mo agoI know HN commenters like to see themselves as contrarians, as do I sometimes, but man… this seems like a serious stretch to assume such malicious intent that an employee of the world’s top AI name would astroturf a random HN thread about a picture on a blog. I’m fairly comfortable taking this OpenAI employee’s comment at face value. Frankly, I don’t think a HN thread will make a difference to his financial situation, anyway…
- rvnx 10mo agoMalicious ? No, and this is far from astroturfing, he even speaks as "we". It's just a logical move to defend your company when people claim your product is buggy. There is no other logical move, this is what I am saying, contrary to people above say this requires a lot of courage. It's not about courage, it's just normal and logic (and yes Hackernews matters a lot, this place is a very strong source of signal for investors). Not bad at all, just observing it.
- BoppreH 10mo agoThat would be a laudable goal, but I feel like it's contradicted by the text: > Even on a low-quality image, GPT‑5.2 identifies the main regions and places boxes that roughly match the true locations of each component I would not consider it to have "identified the main regions" or to have "roughly matched the true locations" when ~1/3 of the boxes have incorrect labels. The remark "even on a low-quality image" is not helping either. Edit: credit where credit is due, the recently-added disclaimer is nice: > Both models make clear mistakes, but GPT‑5.2 shows better comprehension of the image.
- hnuser123456 10mo agoYeah, what it's calling RAM slots is the CMOS battery. What it's calling the PCIE slot is the interior side of the DB-9 connector. RAM slots and PCIE slots are not even visible in the image.
- hexaga 10mo agoIt just overlaid a typical ATX pattern across the motherboard-like parts of the image, even if that's not really what the image is showing. I don't think it's worthwhile to consider this a 'local recognition failure', as if it just happened to mistake CMOS for RAM slots. Imagine it as a markdown response: # Why this is an ATX layout motherboard (Honest assessment, straight to the point, *NO* hallucinations) 1. *RAM* as you can clearly see, the RAM slots are to the right of the CPU, so it's obviously ATX 2. *PCIE* the clearly visible PCIE slots are right there at the bottom of the image, so this definitely cannot be anything except an ATX motherboard 3. ... etc more stuff that is supported only by force of preconception -- It's just meta signaling gone off the rails. Something in their post-training pipeline is obviously vulnerable given how absolutely saturated with it their model outputs are. Troubling that the behavior generalizes to image labeling, but not particularly surprising. This has been a visible problem at least since o1, and the lack of change tells me they do not have a real solution.
- deleted 10mo ago[deleted]
- furyofantares 10mo ago
- arscan 10mo agoI think you may have inadvertently misled readers in a different way. I feel misled after not catching the errors myself, assuming it was broadly correct, and then coming across this observation here. Might be worth mentioning this is better but still inaccurate. Just a bit of feedback, I appreciate you are willing to show non-cherry-picked examples and are engaging with this question here. Edit: As mentioned by @tedsanders below, the post was edited to include clarifying language such as: “Both models make clear mistakes, but GPT‑5.2 shows better comprehension of the image.”
- tedsanders 10mo agoThanks for the feedback - I agree our text doesn't make the models' mistakes clear enough. I'll make some small edits now, though it might take a few minutes to appear.
- g947o 10mo agoWhen I saw that it labeled DP ports as HDMI I immediately decided that I am not going to touch this until it is at least 5x better with 95% accuracy with basic things. I don't see any advantage in using the tool.
- jacquesm 10mo agoThat's a far more dangerous territory. A machine that is obviously broken will not get used. A machine that is subtly broken will propagate errors because it will have achieved a high enough trust level that it will actually get used. Think 'Therac-25', it worked in 99.5% of the time. In fact it worked so well that reports of malfunctions were routinely discarded.
- AdamN 10mo agoThere was a low-level Google internal service that worked so well that other teams took a hard dependency on it (against advice). So the internal team added a cron job to drop it every once in a while to get people to trust it less :-)
- iamdanieljohns 10mo agoIs Adaptive Reasoning gone from GPT-5.2? It was a big part of the release of 5.1 and Codex-Max. Really felt like the future.
- tedsanders 10mo agoYes, GPT-5.2 still has adaptive reasoning - we just didn't call it out by name this time. Like 5.1 and codex-max, it should do a better job at answering quickly on easy queries and taking its time on harder queries.
- iamdanieljohns 10mo agoWhy have "light" or "low" thinking then? I've mentioned this before in other places, but there should only be "none," "standard," "extended," and maybe "heavy." Extended and heavy are about raising the floor (~25% and ~45% or some other ratio respectively) not determining the ceiling.
- layer8 10mo agoYou know what would be great? If it had added some boxes with “might be X or Y, but not sure”.
- iwontberude 10mo agoBut it’s completely wrong.
- johnwheeler 10mo agoOh and you guys don't mislead people ever. Your management is just completely trustworthy, and I'm sure all you guys are too. Give me a break, man. If I were you, I would jump ship or you're going to be like a Theranos employee on LinkedIn.
- yard2010 10mo agoHey no need to personally attack anyone. A bad organization can still consist good people.
- johnwheeler 10mo agoI disagree. I think the whole organization is egregious and full of Sam Altman sycophants that are causing a real and serious harm to our society. Should we not personally attack the Nazis either? These people are literally pushing for a society where you're at a complete disadvantage. And they're betting on it. They're banking on it.
- deleted 10mo ago[deleted]
- whalesalad 10mo agoto be fair that image has the resolution of a flip phone from 2003
- malfist 10mo agoIf I ask you a question and you don't have enough information to answer, you don't confidently give me an answer, you say you don't know. I might not know exactly how many USB ports this motherboard has, but I wouldn't select a set of 4 and declare it to be a stacked pair.
- AstroBen 10mo agoNo-one should have the expectation LLMs are giving correct answers 100% of the time. It's inherent to the tech for them to be confidently wrong Code needs to be checked References need to be checked Any facts or claims need to be checked
- malfist 10mo agoAccording to the benchmarks here they're claiming up to 97% accuracy. That ought to be good enough to trust them right? Or maybe these benchmarks are all wrong
- AstroBen 10mo agoDoes code work if it's 97% correct? It's not okay if claims are totally made up 1/30 times Of course people aren't always correct either, but we're able to operate on levels of confidence. We're also able to weight others' statements as more or less likely to be correct based on what we know about them
- fooker 10mo ago> Does code work if it's 97% correct? Of course it does. The vast majority of software has bugs. Yes, even critical one like compilers and operating systems.
- jasonlotito 10mo agoFTA: Both models make clear mistakes, but GPT‑5.2 shows better comprehension of the image. You can find it right next to the image you are talking about.
- tedsanders 10mo agoTo be fair to OP, I just added this to our blog after their comment, in response to the correct criticisms that our text didn't make it clear how bad GPT-5.2's labels are. LLMs have always been very subhuman at vision, and GPT-5.2 continues in this tradition, but it's still a big step up over GPT-5.1. One way to get a sense of how bad LLMs are at vision is to watch them play Pokemon. E.g.,: https://www.lesswrong.com/posts/u6Lacc7wx4yYkBQ3r/insights-into-claude-opus-4-5-from-pokemon https://www.lesswrong.com/posts/u6Lacc7wx4yYkBQ3r/insights-i... They still very much struggle with basic vision tasks that adults, kids, and even animals can ace with little trouble.
- da_grift_shift 10mo ago'Commented after article was already edited in response to HN feedback' award
- an0malous 10mo agoBecause the whole culture of AI enthusiasts is to just generate slop and never check the results
- 8organicbits 10mo agoPromotional content for LLMs is really poor. I was looking at Claude Code and the example on their homepage implements a feature, ignoring a warning about a security issue, commits locally, does not open a PR and then tries to close the GitHub issue. Whatever code it wrote they clearly didn't use as the issue from the prompt is still open. Bizarre examples.
- fumeux_fume 10mo agoGeneral purpose LLMs aren't very good with generating bounding boxes, so with that context, this is actually seen as decent performance for certain use cases.
- tennisflyi 10mo agoYou seen the charts on their last release? They obviously don’t check - too rich
- az226 10mo agoAnd here is Gemini 3: https://media.licdn.com/dms/image/v2/D5610AQH7v9MtrZxxug/image-shrink_1280/B56ZsP9UUAIEAM-/0/1765499291160?e=1766131200&v=beta&t=AWL4EdNodgFtwjBEbKhVMFS_WyQsnX1zBdnGo3ckFMg https://media.licdn.com/dms/image/v2/D5610AQH7v9MtrZxxug/ima...
- saejox 10mo agoThis is very impressive. Google really is ahead
- pietz 10mo agoThey are definitely ahead in multi modality and I'd argue they have been for a long time. Their image understanding was already great, when their core LLM was still terrible.
- FinnKuhn 10mo agoThis is genuinly impressive. The OpenAI equivalent is less detailed AND less correct.
- Lionga 10mo agoWhen OpenAI Marketing Material is actually showing how far Gemini3 is ahead...