5 ms·
Unfortunately, as every public benchmark, once it ends up in the training sets and/or the developers aware of it, it stops being effective, and I think we've st
by diggan 1y ago
Unfortunately, as every public benchmark, once it ends up in the training sets and/or the developers aware of it, it stops being effective, and I think we've started to reach that point.
The only thing I've found to give me some sort of quantitative idea of how good a new model is, is my own private benchmarks. It doesn't cover everything I want to use LLMs for, and only has 20-30 tests per "category", but at least I'm 99% sure it isn't in the training datasets.
- simonw 1y agoI have a few "SVG of an X riding a Y" tests that I don't publish online which I run occasionally to see if a model is suspiciously better at drawing a pelican riding a bicycle than some other creature on some other form of transport. I would be so entertained if I found out an AI lab had wasted their time cheating on my dumb benchmark!
- ajcp 1y ago-> I would be so entertained if I found out an AI lab had wasted their time cheating on my dumb benchmark! Que intro: "The gang wastes their time cheating on a dumb benchmark"
- mcny 1y agoA shower thought I just had: there must be some AI training company somewhere that has injested all It is always sunny in Philadelphia, not just the text but all the video from all episodes somehow...
- Imustaskforhelp 1y agoPlease do let us know through your blog post if you ever find AI labs to cheat on your benchmark. But now I am worried that since you have shared that you do SVG of an X riding a Y thing, maybe these models will try to cheat on the whole SVG of X riding Y thing instead of hyper focusing the pelican. So now I suppose you might need to come up with an entirely new thing though :)
- throwup238 1y agoThere are so many X and Y combinations that I find it hard to believe they could realistically train for a even a small fraction of them. Someone has to generate the graphics output for the training. A duck billed platypus riding a unicycle? A man o' war riding a pyrosome? A chicken riding a Quetzalcoatlus? A tardigrade riding a surf board?
- gnatolf 1y agoYou're assuming that given the collection of simonw's publicly available blog posts, the creativity of those combinations can't be narrowed down. Simply reverse engineer his brain this way and you'll get your Xs and Ys ;)
- throwup238 1y agoI feel like that would over fit on various snakes like pythons.
- fragmede 1y agoIf we accept ChatGPT telling me that there are approximately 200k common nouns in English, and then we square that, we get 40 billion combinations. At one second per, that's ~1200 years, but then if we parallelize it on a supercomputer that can do 100,000 per second that would only take 3 days. Given that ChatGPT was trained on all of the Internet and every book written, I'm not sure that still seems infeasible.
- throwup238 1y agoIt still can't satisfactorily draw a pelican on a bicycle because that's either not in the training data or the signal is too weak, so why would it be able to satisfactorily draw every random noun-riding-noun combination just because you threw a for loop at it? The point is that in order to cheat on @simonw's benchmark across any arbitrary combination, they'd have to come up with an absurd number of human crafted input-output training pairs with human produced drawings. You can't just ask ChatGPT to generate every combination because all it'll produce is garbage that gets a lot worse the further from a pelican riding a bicycle. It might work at first for the pelican and a few other animals/transport combination but what does it even mean for a man o' war riding a pyrosome? I asked every model I have access to generate an SVG for a "man o' war riding a pyrosome" and not a single one managed to draw anything resembling a pyrosome. Most couldn't even produce something resembling a man o' war except as a generic ellipsoid-shaped jellyfish with a few tenticles. Expand that to every weird noun-noun combination and it's just not practical to train even a tiny fraction of them.
- diggan 1y ago> I would be so entertained if I found out an AI lab had wasted their time cheating on my dumb benchmark! I don't think it's necessarily "cheating", it just happens as they're discovering and ingesting large ranges of content. A problem of public content, it's bound to be included sooner or later, directly or indirectly. Nice to hear you're doing some sort of contingency though, and looking forward to the inevitable blog post announcing the change to a different bird and vehicle :)
- deleted 1y ago[deleted]
- reissbaker 1y agoI doubt they'd cheat that obviously... But "SVG of X" has become common enough that I suspect most frontier labs train on it, especially since the models are multimodal now anyway. Not that I mind; I want models to be good at generating SVG! Makes icons much simpler.
- fragmede 1y agoBut how would you know it's from what you would consider cheating as opposed to pelicans on bicycles existing in the latest training data? Obviously your blog gets fed into the training set for GPT-6, as well as everyone else talking about your test, so how would the comparison to a secret X riding a Y tell you if an AI lab is cheating as opposed to merely there being more examples in the training data?
- simonw 1y agoMainly because if they train on the pelican on bicycle SVGs from my blog they are going to get some very weird looking pelicans riding some terrible looking bicycles.
- fragmede 1y agoIt's not that I claiming they're training on SVG pelicans on bicycles from your blog, it's that thanks to your popularity, there are simply now more pictures of pelicans on bicycles floating around on the Internet and thus ChatGPT's training data. Eg https://www.reddit.com/r/ColoredPencils/comments/1l9l4fq/pelican_riding_bicycle/ https://www.reddit.com/r/ColoredPencils/comments/1l9l4fq/pel... How would you determine that improvements to SVG pelicans on bicycles (and not your secret X on Ys) are from an OpenAI employee cheating your benchmark vs being an improvement on pelicans on bicycles thanks to that picture from Reddit and everywhere elsewhere in the training data?
- simonw 1y agoSee comment here: https://news.ycombinator.com/item?id=45454269 https://news.ycombinator.com/item?id=45454269
- jgalt212 1y agoYour benchmark may or may not be dumb, but it is definitely widely followed. So much so this is what Bing AI has to say on the matter. > Absolutely — the “pelican riding a bicycle” SVG test is a quirky but clever benchmark created by Simon Willison to evaluate how well different large language models (LLMs) can generate SVG (Scalable Vector Graphics) images from a prompt that’s both unusual and unlikely to be in their training data.
- ajcp 1y agoThat's the move right there.
- latemedium 1y agoWe need to know if big AI labs are explicitly training models to generate SVGs of pelicans on bicycles. I wouldn't put it past them. But it would be pretty wild in they did!
- londons_explore 1y agoAs soon as you use your private tests, all the AI companies vacuum up the input to use to train the next model. Obviously they're only getting the question and not a perfect answer, but with today's process of generating hundreds of potential answers and getting another model to choose the best/correct one for training, I don't think that matters.
- astrange 1y agoAre the models capable of judging a good SVG? They can't read ASCII art.
- londons_explore 1y agoIf you give the 'judge' models tool use, they could easily fire up a web browser to render an SVG and then use imagenet or something to see how 'pelican-y' the result is.
- Workaccount2 1y agoI honestly think people really blow out of proportion the effect of "being in the training set". The internet is ridden with examples of problem/solution posts that many models definitely trained on, but still get wrong. More important would be post training, where the labs specifically train on the exact question. But it doesn't seem like this is happening for most amateur benchmarks at least. All the models that are good at pelican bike have been good at whatever else you throw at them to SVG.