5 ms·
I believe that it may be misguided to focus on compute that much, and it would be more instructive to consider the effort that went into curating the training s
by fsh 1y ago
I believe that it may be misguided to focus on compute that much, and it would be more instructive to consider the effort that went into curating the training set. The easiest way of solving math problems with an LLM is to make sure that very similar problems are included in the training set. Many of the AI achievements would probably look a lot less miraculous if one could check the training data. The most crass example is OpenAI paying off the FrontierMath creators last year to get exclusive secret access to the problems before the evaluation [1]. Even without resorting to cheating, competition formats are vulnerable to this. It is extremely difficult to come up with truly original questions, so by spending significant resources on re-hashing all kinds of permutations of previous question, one will probably end up very close to the actual competition set. The first rule I learned about training neural networks is to make damn sure there is no overlap between the training and validation sets. It it interesting that this rule has gone completely out of the window in the age of LLMs.
[1] https://www.lesswrong.com/posts/8ZgLYwBmB3vLavjKE/some-lessons-from-the-openai-frontiermath-debacle https://www.lesswrong.com/posts/8ZgLYwBmB3vLavjKE/some-lesso...
- OtherShrezzing 1y ago> The easiest way of solving math problems with an LLM is to make sure that very similar problems are included in the training set. Many of the AI achievements would probably look a lot less miraculous if one could check the training data I'm fairly certain this phenomenon is responsible for LLM capabilities on GeoGuesser type games. They have unreasonably good performance. For example, being able to identify obscure locations from featureless/foggy pictures of a bench. GeoGuesser's entire dataset, including GPS metadata, is definitely included in all of the frontier model training datasets - so it should be unsurprising that they have excellent performance in that domain.
- YetAnotherNick 1y ago> GeoGuesser's entire dataset No, it is not included, however there must be quite a lot of pictures on internet for most cities.. Geoguesser data is same as Google's street view data and it probably contains billions of 360 degree photos.
- ivape 1y agoI just saw a video on Reddit where a woman still managed to take a selfie while being literally face to face with a black bear. There’s definitely way too much video training data out there for everything.
- lutusp 1y ago> I just saw a video on Reddit where a woman still managed to take a selfie while being literally face to face with a black bear. This is not uncommon. Bears aren't always tearing people apart, that's a movie trope with little connection to reality. Black bears in particular are smart and social enough to befriend their food sources. But a hungry bear, or a bear with cubs, that's a different story. Even then bears may surprise you. Once in Alaska, a mama bear got me to babysit her cubs while she went fishing -- link: https://arachnoid.com/alaska2018/bears.html https://arachnoid.com/alaska2018/bears.html .
- suddenlybananas 1y agoWhy do you say it's not included? Why wouldn't they include it.
- sebzim4500 1y agoIf every photo in streetview was included in the training data of a multimodal LLM it would be like 99.9999% of the training data/resource costs. It just isn't plausible that anyone has actually done that. I'm sure some people include a small sample of them, though.
- bluefirebrand 1y agoWhy would every photo in streetview be required in order to have Geoguessr's dataset in the training data?
- bee_rider 1y agoI’m pretty sure they are saying that Geoguessr's just pulls directly from Google Streetview. There isn’t a separate Geoguessr dataset, it just pulls from Google’s API (at least that’s what Wikipedia says).
- ACCount36 1y agoPeople tried VLMs on "closed set" GeoGuessr-type tasks - i.e. non-Street View photos in similar style, not published anywhere. They still kicked ass. It seems like those AIs just have an awful lot of location familiarity. They've seen enough tagged photos to be able to pick up on the patterns, and generalize that to kicking ass at GeoGuessr.
- astrange 1y ago> The easiest way of solving math problems with an LLM is to make sure that very similar problems are included in the training set. An irony here is that math blogs like Tao's might not be in LLM training data, for the same reason they aren't accessible to screen readers - they're full of math, and the math is rendered as images, so it's nonsense if you can't read the images. (The images on his blog do have alt text, but it's just the LaTeX code, which isn't much better.)
- prein 1y agoWhat would be a better alternative than LaTex for the alt text? I can't think of a solution that makes more sense, it provides an unambiguous representation of what's depicted. I wouldn't think an LLM would have issue with that at all. I can see how a screen reader might, but it seems like the same problem faced by a screen reader with any piece of code, not just LaTex.
- QuesnayJr 1y agoLLMs understand LaTeX extraordinarily well.
- MengerSponge 1y agoLLMs are decent with LaTeX! It's just markup code after all. I've heard from some colleagues that they can do decent image to code conversion for a picture of an equation or even some handwritten ones.
- alansammarone 1y agoAs others have pointed out, LLMs have no trouble with LaTeX. I can see why one might think they're not - in fact, I made the same assumption myself sometime ago. LLMs, via transformers, are exceptionally good any _any_ sequence or one-dimensional data. One very interesting (to me anyway) example is base64 - pick some not-huge sentence (say, 10 words), base64-encode it, and just paste it in any LLM you want, and it will be able to understand it. Same works with hex, ascii representation, or binary. Here's a sample if you want to try: aWYgYWxsIEEncyBhcmUgQidzLCBidXQgb25seSBzb21lIEIncyBhcmUgQydzLCBhcmUgYWxsIEEncyBDJ3M/IEFuc3dlciBpbiBiYXNlNjQu I remember running this experiment some time ago in a context where I was certain there was no possibility of tool use to encode/decode. Nowadays, it can be hard to certain whether there is any tool use or not, in some cases, such as Mistral, the response is quick enough to make it unlikely there's any tool use.
- disruptbro 1y agoLanguage modeling is compression, whittle down graph to reduce duplication and data with little relationship: https://arxiv.org/abs/2309.10668 https://arxiv.org/abs/2309.10668 Let’s say everyone agrees to refer to one hosted copy of a token “cat”, and instead generate a unique vector to represent their reference to “cat”. Blam. Endless unique vectors which are nice and precise for parsing. No endless copies of arbitrary text like “cat”. Now make that your globally distributed data base to bootstrap AI chips from. The data driven programming dream where other machines on the network feed new machines boot strap. American tech industry is IBM now. Stuck on recent success of web SaaS and way behind the plans of AI.
- eru 1y ago> It is extremely difficult to come up with truly original questions, [...] No, that's actually really easy. What's hard is coming up with original questions of a specific level of difficulty. And that's what you need for a competition. To elaborate: it's really easy to find lots and lots of elementary, unsolved questions. But it's not clear whether you can actually solve them or how hard solving them is, so it's hard to judge the performance of LLMs on them. > It it interesting that this rule has gone completely out of the window in the age of LLMs. No, it hasn't.