5 ms·
OpenAI’s o3 searches the web behind a curtain: you get a few source links and a fuzzy reasoning trace, but never the full chunk of text it actually pulled in. W
by patrickhogan1 1y ago
OpenAI’s o3 searches the web behind a curtain: you get a few source links and a fuzzy reasoning trace, but never the full chunk of text it actually pulled in. Without that raw context, it’s impossible to audit what really shaped the answer.
- simonw 1y agoYeah, I find that really frustrating. I understand why they do it though: if they presented the actual content that came back from search they would absolutely get in trouble for copyright-infringement. I suspect that's why so much of the Claude 4 system prompt for their search tool is the message "Always respect copyright by NEVER reproducing large 20+ word chunks of content from search results" repeated half a dozen times: https://simonwillison.net/2025/May/25/claude-4-system-prompt/#seriously-don-t-regurgitate-copyrighted-content https://simonwillison.net/2025/May/25/claude-4-system-prompt...
- Zopieux 1y agoThis is no secret or suspicion. It is definitely about avoiding (more accuratly, delaying until legislation destroys the business model) the warth of copyright holders with enough lawyers. I find this very hypocritical given that for all intents and purposes the infringement already happened at training time, since most content wasn't acquired with any form of retribution or attribution (otherwise this entire endeavor would not have been economically worth it). See also the "you're not allowed to plagiarize Disney" being done by all commercial text to image providers.
- NoraCodes 1y agoI don't understand how you can look at behavior like this from the companies selling these systems and conclude that it is ethical for them to do so, or for you to promote their products.
- simonw 1y agoWhat's happening here is Claude (and ChatGPT alike) have a tool-based search option. You ask them a question - like "who won the Superbowl in 1998" - they then run a search against a classic web search engine (Bing for ChatGPT, Brave for Claude) and fetch back cached results from that engine. They inject those results into their context and use them to answer the question. Using just a few words (the name of the team) feels OK to me, though you're welcome to argue otherwise. The Claude search system prompt is there to ensure that Claude doesn't spit out multiple paragraphs of text from the underlying website, in a way that would discourage you from clicking through to the original source. Personally I think this is an ethical way of designing that feature. (Note that the way this works is an entirely different issue from the fact that these models were training on unlicensed data.)
- deleted 1y ago[deleted]
- NoraCodes 1y agoI understand how it works. I think it does not do much to encourage clicking through, because the stated goal is to solve the user's problem without leaving the chat interface (most of the time.)
- simonw 1y agoYeah, I agree. I actually think an even worse offender here is Google themselves - their AI overview thing answers questions directly on the Google page itself, discouraging site visits. I think that's going to have a really nasty impact on site traffic.