3 ms·
I see that your prompt includes 'Do not use any tools. If you do, write "I USED A TOOL"' This is not a valid experiment, because GPT models always have access
by rybosworld 8mo ago
I see that your prompt includes 'Do not use any tools. If you do, write "I USED A TOOL"'
This is not a valid experiment, because GPT models always have access to certain tools and will use them even if you tell them not to. They will fib the chain of thought after the fact to make it look like they didn't use a tool.
https://www.anthropic.com/research/alignment-faking https://www.anthropic.com/research/alignment-faking
It's also well established that all the frontier models use python for math problems, not just GPT family of models.
- simianwords 8mo agoWould it convince you if we use the GPT Pro api and explicitly not allow tool access? Is that enough to falsify?
- jibal 8mo agoIt's not falsifiable because it's not false.
- simianwords 8mo agoThat’s not falsifiable means
- jibal 8mo agoI know what falsifiable means--you're misusing it and I simply adopted your misuse. A claim is falsifiable or not ... it can't be made falsifiable. The way you're using it is "Can we come up with a test to show that it's false"--no, we can't, because it's not false.
- simianwords 8mo agoHow do you know it’s not false? If one had to prove that it is false, what would you have to do?
- jibal 8mo ago[flagged]
- simianwords 8mo ago[flagged]
- dang 8mo agoPlease don't cross into posting like this, no matter how wrong someone else is or you feel they are. It's not what this site is for, and destroys what it is for. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- dang 8mo agoPlease don't cross into posting like this, no matter how wrong someone else is or you feel they are. It's not what this site is for, and destroys what it is for. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- rybosworld 8mo agoNo, it wouldn't be enough to falsify. This isn't an experiment a consumer of the models can actually run. If you have a chance to read the article I linked, it is difficult even for the model maintainers (openai, anthropic, etc.) to look into the model and see what it actually used in it's reasoning process. The models will purposefully hide information about how they reasoned. And they will ignore instructions without telling you. The problem really isn't that LLM's can't get math/arithmetic right sometimes. They certainly can. The problem is that there's a very high probability that they will get the math wrong. Python or similar tools was the answer to the inconsistency.
- simianwords 8mo agoWhat do you mean? You can explicitly restrict access to the tools. You are factually incorrect here.
- rybosworld 8mo agoI believe you're referring to the tools array? https://developers.openai.com/api/docs/guides/tools/ https://developers.openai.com/api/docs/guides/tools/ This is external tools that you are allowing the model to have access to. There is a suite of internal tools that the model has access to regardless. The external python tool is there so it can provide the user with python code that they can see. You can read a bit more about the distinction between the internal and external tool capabilities here: https://community.openai.com/t/fun-with-gpt-5-code-interpreter-and-why-it-likely-fails-to-deliver-files-in-many-instructed-cases/1356685 https://community.openai.com/t/fun-with-gpt-5-code-interpret... "I should explain that both the “python” and “python_user_visible” tools execute Python code and are stateful. The “python” tool is for internal calculations and won’t show outputs to the user, while “python_user_visible” is meant for code that users can see, like file generation and plots." But really the most important thing, is that we as end-users cannot with any certainty know if the model used python, or didn't. That's what the alignment faking article describes.
- simianwords 8mo ago
- chickenimprint 8mo agoAs far as I know, you can't disable the python interpreter. It's part of the reasoning mode. If you ask ChatGPT, it will confirm that it uses the python interpreter to do arithmetic on large numbers. To you, that should be convincing.