4 ms·
At this point, we have to assume anything that becomes a published benchmark is specifically targeted during training. That's not something specific to LLMs or
by codeflo 2y ago
At this point, we have to assume anything that becomes a published benchmark is specifically targeted during training. That's not something specific to LLMs or OpenAI. Compiler companies have done the same thing for decades, specifically detecting common benchmark programs and inserting hand-crafted optimizations. Similarly, the shader compilers in GPU drivers have special cases for common games and benchmarks.
- darkerside 2y agoVW got in a lot of trouble for this
- TrueDuality 2y agoNot quite. VW got in trouble for running _different_ software in test vs prod. These optimizations are all going to "prod" but are only useful for specific targets (a specific game in this case).
- krisoft 2y ago> VW got in trouble for running _different_ software in test vs prod. Not quite. They programmed their "prod" software to recognise the circumstances of a laboratory test and behave differently. Namely during laboratory emissions testing they would activate emission control features they would not activate otherwise. The software was the same they flash on production cars. They were production cars. You could take a random car from a random dealership and it would have done the same trickery in the lab.
- TrueDuality 2y agoI disagree with your distinction on the environments but understand your argument. Production for VM to me is "on the road when a customer is using your product as intended". Using the same artifact for those different environments isn't the same as "running that in production".
- krisoft 2y ago“Test” environment is the domain of prototype cars driving at the proving ground. It is an internal affair, only for employees and contractors. The software is compiled on some engineer’s laptop and uploaded on the ECU by an engineer manually. No two cars are ever the same, everything is in flux. The number of cars are small. “Production” is a factory line producing cars. The software is uploaded on the ECUs by some factory machine automatically. Each car are exactly the same, with the exact same software version on thousands and thousands of cars. The cars are sold to customers. Some small number of these prodiction cars are sent for regulatory compliance checks to third parties. But those cars won’t become suddenly non-production cars just because someone sticks up a probe in their exhausts. The same way gmail’s production servers don’t suddenly turn into test environments just because a user opens the network tab in their browser’s dev tool to see what kind of requests fly on the wire.
- close04 2y agoOnly because what VW did is illegal, was super large scale, and could be linked to a lot of indirect deaths through the additional pollution. Benchmark optimizations are slightly embarrassing at worst, and an "optimization for a specific use case" at best. There's no regulation against optimizing for a particular task, everyone does it all the time, in some cases it's just not communicated transparently. Phone manufacturers were caught "optimizing" for benchmarks again and again, removing power limits to boost scores. Hard to name an example without searching the net because it's at most a faux pas.
- Swenrekcah 2y agoActually performing well on a task that is used as a benchmark is not comparable to decieving authorities about how much toxic gas you are releasing.
- ArnoVW 2y agoTrue. But they did not optimize for a specific case. They detected the test and then enabled a special regime, that was not used normally. It’s as if OpenAI detects the IP address from a benchmark organization, and then used a completely different model.
- K0balt 2y agoThis is the apples to apples version. Perhaps might be more accurate to say that when detecting a benchmark attempt the model tries the prompt 3 times with different seeds then picks the best answer, otherwise it just zero-shots the prompt in everyday use. I say this because the be test still uses the same hardware (model) but changed the way it behaved by running emissions friendly parameters ( a different execution framework) that wouldn’t have been used in everyday driving, where fuel efficiency and performance optimized parameters were used instead. What I’d like to know is if it actually was unethical or not. The overall carbon footprint of the lower fuel consumption setting, with fuel manufacturing and distribution factored in, might easily have been more impactful than the emissions model, which typically does not factor in fuel consumed.
- sigmoid10 2y agoApples and oranges. VW actually cheated on regulatory testing to bypass legal requirements. So to be comparable, the government would first need to pass laws where e.g. only compilers that pass a certain benchmark are allowed to be used for purchasable products and then the developers would need to manipulate behaviour during those benchmarks.
- 0xFF0123 2y agoThe only difference is the legality. From an integrity point of view it's basically the same
- boringg 2y agoHow so? VW intentionally changed the operation of the vehicle so that its emissions met the test requirements during the test and then went back to typical operation conditions afterwards.
- Thorrez 2y agoI think breaking a law is more unethical than not breaking a law. Also, legality isn't the only difference in the VW case. With VW, they had a "good emissions" mode. They enabled the good emissions mode during the test, but disabled it during regular driving. It would have worked during regular driving, but they disabled it during regular driving. With compilers, there's no "good performance" mode that would work during regular usage that they're disabling during regular usage.
- Lalabadie 2y ago> I think breaking a law is more unethical than not breaking a law. It sounds like a mismatch of definition, but I doubt you're ambivalent about a behavior right until the moment it becomes illegal, after which you think it unethical. Law is the codification and enforcement of a social contract, not the creation of it.
- Thorrez 2y ago
- tightbookkeeper 2y agoThis is 10 year old story. It’s very interesting which ones stay in the public consciousness.
- bluGill 2y agoMost of the time these days compiler writers are not cheating like VW did. In the 1980s compiler writers would insert code to recognize performance tests and then cheat - output values hard coded into the compiler instead of running the algorithm. Which is the type of thing that VW got in trouble for. These days most compilers are trying to make the general case of code fast and they rarely look for benchmarks. I won't say they never do this - just that it is much less common - if only because magazine reviews/benchmarks are not nearly as important as they used to be and so the incentive is gone.
- newerman 2y agoFunny response; you're not wrong.
- conradev 2y agoGPT-3.5 did not “cheat” on chess benchmarks, though, it was actually just better at chess?
- GolfPopper 2y agoI think the OP's point is that chat GPT-3.5 may have a chess-engine baked-in to its (closed and unavailable) code for PR purposes. So it "realizes" that "hey, I'm playing a game of chess" and then, rather than doing whatever it normally does, it just acts as a front-end for a quite good chess-engine.
- conradev 2y agoI see – my initial interpretation of OP’s “special case” was “Theory 2: GPT-3.5-instruct was trained on more chess games.” But I guess it’s also a possibility that they had a real chess engine hiding in there.
- gdiamos 2y agoIt’s approximately bad, like most of ML On one side: Would you expect a model trained on no Spanish data to do well on Spanish? On the other: Is it okay to train on the MMLU test set?
- deleted 2y ago[deleted]
- dang 2y agoWe detached this subthread from https://news.ycombinator.com/item?id=42144784 https://news.ycombinator.com/item?id=42144784. (Nothing wrong with it! It's just a bit more generic than the original topic.)