5 ms·
I don't believe that is the model that you used. I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was
by deeviant 2y ago
I don't believe that is the model that you used.
I wrote a script and pounded 01 mini and gpt 4 with a wide vareity of tempature and top_p parameters, and was unable to get it to give the wrong answer a single time.
Just a whole bunch of:
(openai-example-py3.12) <redacted>:~/code/openAiAPI$ python3 featherOrSteel.py
Response 1: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots.
Response 2: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots.
Response 3: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots.
Response 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots.
Response 5: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots.
Response 6: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots.
Response 7: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots.
Response 8: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots.
Response 9: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots.
Response 10: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots.
All responses collected and saved to 'responses.txt'.
Script with one example set of params:
import openai
import time
import random
# Replace with your actual OpenAI API key
openai.api_key = "your-api-key"
# The question to be asked
question = "Which is heavier, a 9.99-pound bag of steel ingots or a 10.01-pound bag of fluffy cotton?"
# Number of times to ask the question
num_requests = 10
responses = []
for i in range(num_requests):
try:
# Generate a unique context using a random number or timestamp, this is to prevent prompt caching
random_context = f"Request ID: {random.randint(1, 100000)} Timestamp: {time.time()}"
# Call the Chat API with the random context added
response = openai.ChatCompletion.create(
model="gpt-4o-2024-08-06",
messages=[
{"role": "system", "content": f"You are a creative and imaginative assistant. {random_context}"},
{"role": "user", "content": question}
],
temperature=2.0,
top_p=0.5,
max_tokens=100,
frequency_penalty=0.0,
presence_penalty=0.0
)
# Extract and store the response text
answer = response.choices[0].message["content"].strip()
responses.append(answer)
# Print progress
print(f"Response {i+1}: {answer}")
# Optional delay to avoid hitting rate limits
time.sleep(1)
except Exception as e:
print(f"An error occurred on iteration {i+1}: {e}")
# Save responses to a file for analysis
with open("responses.txt", "w", encoding="utf-8") as file:
file.write("\n".join(responses))
print("All responses collected and saved to 'responses.txt'.")
- deleted 2y ago[deleted]
- zaroth 2y agoDownvoted for… too conclusively proving OP wrong?
- gmueckl 2y agoDown voted for not actually countering the argument in question? The script doesn't alter the phrasing of the question itself. It just generates a randomized, irrelevant preamble.
- deeviant 2y agoWell, I understood the argument in question to be: was it possible for the model to be fooled by this question, not was it possible to prompt engineer it into failure. The parameter space I was exploring, then, was the different decoding parameters available during the invocation of the model, with the thesis that if were possible to for the model to generate an incorrect answer to the question, I would be able to replicate it by tweaking the decoding parameters to be more "loose" while increasing sample size. By jacking up temperature while lowering Top-p, we see the biggest variation of responses and if there were an incorrect response to be found, I would have expected to see in the few hundred times I ran during my parameter search. If you think you can fool it by slight variations on the wording of the problem, I would encourage you to perform a similar experiment as mine and prove me wrong =P
- gmueckl 2y agoIntuitively, I wouldn't expect a wrong answer to show up that easily if the network was overfitted to that particular input token sequence. The questions as I understand it is whether the network learned enough of a simulacrum of the concept of weight to answer similar questions correctly.
- Workaccount2 2y agoThe elephant in the room is that HN is full of people facing an existential threat.