4 ms·
Came for: "A computer once beat me at chess, but it was no match for me at kick boxing." TFA was actually about leaps of intuition, sadly. One of the experime
by jvanderbot 2mo ago
Came for:
"A computer once beat me at chess, but it was no match for me at kick boxing."
TFA was actually about leaps of intuition, sadly.
One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
- Normal_gaussian 2mo agoThe curious case here is how much of a description do we give it of itself? That would almost certainly dominate success rates. My feeling is that a prompt would have to provide a vague description of a program that meaningfully passes something like a Turing test, an API to conform to, an expectation of novel construction (no 'ifs all the way down'), and then a requirement to search broadly and pursue promising ideas and not get hung up on the philosophy. Anything more precise feels like it would corrupt the test, but as it is that description feels doomed to loop before even trying the interesting parts.
- ModernMech 2mo agoI wonder if we could just tell it to invent itself without any description and see if it can I introspect enough through its own interface to figure out what it is.
- elar_verole 2mo agoI think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now
- jvanderbot 2mo agoEven simpler: Can GPT-2 anticipate and build Gwen/Deepseek? I think the answer is almost trivially "no", so I wonder what changed?
- ben_w 2mo agoLots of things changed, GPT-2 is small (1.5e9) and is also a base model, so it is only doing next-token/autocomplete rather than prompt-response like even the first ChatGPT-3.5 was doing.
- jebarker 2mo agoJust for the sake of clarity: all LLMs up to today are still only doing next-token/autocomplete. The training process got additional stages to shape the model weights, but standalone LLMs are still deployed essentially identically.
- ben_w 2mo agoIf you gave GPT-2 a question and ended with a "?", it might answer, but also it might write several more questions in a similar category. IMO, the mechanism isn't the important thing, the behaviour is. If you look at the step-by-step, we are also looking for the next word or motor action (and for whoever is about to suggest that we humans plan ahead, Transformer-based LLMs have been shown to also do this); as this is not a useful description of what it means to be a living brain, I'd say it's also not a useful description of what makes everything post-InstructGPT different from what came before.
- jebarker 2mo agoI agree completely - behaviorally the models have changed drastically due to RLHF, RLVR and now maybe even more so due to agentic harnesses. But the mechanism of prediction hasn’t changed, that was all I was clarifying.
- joefourier 2mo agoWhat about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.
- hackernudes 2mo agoMaybe we can synthesize large amounts of limited information. I thought that new training data is mostly synthetic anyway.
- inigyou 2mo agoWhy couldn't an LLM, if it was smart enough, generate and consume its own data? I know the answer: because it leads to model collapse. But why is that? Wouldn't a smart model not collapse? It's seeming like they keep getting smarter because we keep pouring more of our own knowledge into them, not because they are actually getting smarter. And yes, sometimes a dumb but persistent bruteforcer can make new discoveries.
- Garlef 2mo ago> if it was smart enough and i think this is exactly the crux; the really big models need really big datasets and current gen LLMs get a lot of training data beyond "all books + all of the internet" the objection is then that producing this additional data would already confound it with pre "virtual cutoff date" knowledge (since the training data probably implies mathematical and SWE concepts that were developed post "virtual cutoff date")
- smusamashah 2mo agoIf it is smart enough to generate data it can consume to train itself better, it is already smart enough to not need to do that.
- inigyou 2mo agoIf a human is smart enough to do the Michelson-Morley experiment, they are smart enough to not need to do that.
- Plasmoid 2mo agoIt's because LLMs are entropy generators. That's not a bad thing for what people are doing. But to prevent model collapse you need a way to pump down the entropy. Much like in thermo, it's an expensive and slow process.
- Kinrany 2mo agoThey are already trained on generated data I believe
- naasking 2mo agoTypical LLM pretraining is very inefficient with data. NanoGPT slowrun shows that data can be used much more efficiently.
- throwaway314155 2mo agoArticle was plenty interesting to me.
- Tade0 2mo agoIs that how chessboxing was invented? Genuinely asking.
- the_af 2mo agoNope. Chessboxing was the invention of comics book artist Enki Bilal (and he's credited with this in Wikipedia). I first saw it in his Nikopol trilogy. Because life is weird, it then became a real thing. It's unrelated to computers playing chess. It predates Kasparov's first defeat by Deep Blue. I don't remember any mention of computers being good at chess in the trilogy, either. Or any computers, for that matter.
- ambicapter 2mo agoFunny, I knew about chessboxing and Enki Bilal, but had no idea one birthed the other.
- ben_w 2mo ago> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself. Could be, but preventing leakage from more modern stuff can be challenging. This was attempted with Victorian public domain content: https://www.estragon.news/mr-chatterbox-or-the-modern-prometheus/ https://www.estragon.news/mr-chatterbox-or-the-modern-promet... I can't find the citation right now, but I think people found it was leaking anachronisms? So this probably wasn't as well filtered as the creator had hoped?
- morkalork 2mo agoProgress followed improvements in hardware, would you have to give access to modern hardware in the experiment for it to use? How much could it infer from it?
- ben_w 2mo ago> would you have to give access to modern hardware in the experiment for it to use? At a minimum, yes. IIRC, the sum total of all compute manufactured over history only reached the minimum needed to train an OK LMM in the mid 00s. > How much could it infer from it? Only way to find out is to try.
- bee_rider 2mo ago> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself. This is an interesting experiment but I wonder if it would be possible to prevent some sort of retrospective bias. For example, I’d expect the experiments that lead to relativity to be over-represented in our catalogue of scientific literature prior to 1900, just because in retrospect they were important, so the records about them were preserved.It would have to be a very intentionally constructed corpus, I think.
- unfitted2545 2mo agoIf would be interesting to see 5 billion LLM's working together, each with random mutations (temperature ig). Would we essentially be looking at a society through a petri dish? Ofc 5 billion is quite a lot of compute.
- handoflixue 2mo agoThey did that with "Talkie", a model trained on 1930 and before. It had the ability to assemble crude Python programs, but there were definitely some leaks in the training data (it knew about stuff like World War 2) so not 100% perfect. Still seems like a pretty reasonable "proof of concept" https://didof.dev/blog/talkie-1930-llm-reasoning/ https://didof.dev/blog/talkie-1930-llm-reasoning/ seems like a decent overview