3 ms·
The Friedman functions are sufficiently well known to be in the training set, likely in the form of some csv file from somebody testing other methods on them.
by waldrews 2y ago
The Friedman functions are sufficiently well known to be in the training set, likely in the form of some csv file from somebody testing other methods on them. Though likely any data set would not share the same x-values. Still, it's surprising performance when we're used to the LLM's getting confused by straightforward math problems and with arithmetic often crippled by the tokenization system.
- nerdponx 2y agoIt's surprising, but in a weird way it makes sense if you think about it as "learning a pattern" instead of "doing math". The model doesn't have to understand logic as we know it (and this class of models doesn't seem to be able to do that in general), but it does have to be able to learn the pattern somehow, and we already know LLMs can learn patterns from the input context without any fine-tuning. I'd like to try it on a dataset like Titanic. That might be a more interesting experiment. Or to avoid training data bias, maybe a completely random dataset generated by some more-sophisticated means, like the scikit-learn data generator which can introduce clusters, irrelevant features, etc.
- rvacareanu 2y agoInitially I experimented with house prediction. But the models have probably seen this data, so I ended up experimenting with random functions. I tried with linear regression with irrelevant features (e.g., NI 1/2 -> 1 informative variable, 2 total variables). The models still perform reasonably well. All results can be seen in this heatmap: https://github.com/robertvacareanu/llm4regression/blob/main/heatmap_all.png https://github.com/robertvacareanu/llm4regression/blob/main/...
- ianand 2y agoThey did invent their own functions to test if the results were due to these functions being on the training date. See the section on data contamination in the paper. Agree it both kind of makes sense (regression is the best way to predict the next token in this context) and kind of ironic (LLMs can do high school regression but can’t do elementary school long digit arithmetic).
- choppaface 2y agoMissing from the paper: "We were able to verify that none of the in-context exemplars were in OpenAI's training set." I wonder if this paper will make it through peer review? Now that could be an interesting result.
- pama 2y agoThe authors explicitly generated new random data to address this concern.