9 ms·
I found the discussion around inference scaling with the 'Wait' hack so surreal. The fact such an ingeniously simple method can impact performance makes me wond
by mtrovo 2y ago
I found the discussion around inference scaling with the 'Wait' hack so surreal. The fact such an ingeniously simple method can impact performance makes me wonder how many low-hanging fruit we're still missing. So weird to think that improvements on a branch of computer science is boiling down to conjuring the right incantation words, how you even change your mindset to start thinking this way?
- nyoomboom 2y agoI think a skill here is learning a bias for experimentation and accepting the results one finds. Also the book "Why Greatness Cannot Be Planned" showcases the kind of open ended play that results in people discovering stuff like this.
- cubefox 2y agoNow imagine where we are in 12 months from now. This article from February 5 2025 will feel quaint by then. The acceleration keeps increasing. It seems likely we will soon have recursive self-improving AI -- reasoning models which do AI research. This will accelerate the rate of acceleration itself. It sounds stupid to say it, but yes, the singularity is near. Vastly superhuman AI now seems to arrive within the next few years. Terrifying.
- gom_jabbar 2y agoYes, and Accelerationism predicted this development back in the 1990s, perhaps most prominently in the opening lines of Nick Land's Meltdown (1994) text: [[ ]] The story goes like this: Earth is captured by a technocapital singularity as renaissance rationalization and oceanic navigation lock into commoditization take-off. Logistically accelerating techno-economic interactivity crumbles social order in auto-sophisticating machine runaway. As markets learn to manufacture intelligence, politics modernizes, upgrades paranoia, and tries to get a grip. > reasoning models which do AI research In the introduction to my research project on Accelerationism [0], I write: Faced with the acceleration of progress in Artificial Intelligence (AI) — with AI agents now automating AI research and development —, Accelerationism no longer seems like an abstract philosophy producing empty hyperstitional hype, but like a sober description of reality. The failed 2023 memorandum to stop AI development on systems more powerful than OpenAI's ChatGPT-4 perfectly illustrates the phenomenological aspects of Accelerationism: "To be rushed by the phenomenon, to the point of terminal institutional paralysis, is the phenomenon." [1] At the current rate of acceleration, if you don't write hyperstitionally, your texts are dead on arrival. [0] https://retrochronic.com/ https://retrochronic.com/ [1] Nick Land (2017). A Quick-and-Dirty Introduction to Accelerationism in Jacobite Magazine.
- versteegen 2y agoNice. Though I couldn't understand those "opening lines" until I read in your Introduction: > For Land, capitalism begins in Northern Italy around 1500 with "the emerging world of technologists and accountants", the spiral interexcitation of "oceanic navigation and place-value calculation", and zero-unlocked double-entry book-keeping Fibonacci, amongst many others, played a critical role that highly accelerative technology.
- pizza 2y agoHope we get the Nick Land the younger, and not Nick Land the elder, set of outcomes. Somewhere, sometime, along the way, it seems like everything from CCRU and Duginism leapt out of the page into the real. Maybe it's just the beginning of the Baudrilliardian millennium.
- zoogeny 2y agoThis is something I have been suppressing since I don't want to become chicken little. Anyone who isn't terrified by the last 3 months probably doesn't really understand what is happening. I went from accepting I wouldn't see a true AI in my lifetime, to thinking it is possible before I die, to thinking it is possible in in the next decade, to thinking it is probably in the next 3 years to wondering if we might see it this year. Just 6 months ago people were wondering if pre-training was stalling out and if we hit a wall. Then deepseek drops with RL'd inference time compute, China jumps from being 2 years behind in the AI race to being neck-and-neck and we're all wondering what will happen when we apply those techniques to the current full-sized behemoth models. It seems the models that are going to come out around summer time may be jumps in capability beyond our expectations. And the updated costs means that there may be several open source alternatives available. The intelligence that will be available to the average technically literate individual will be frightening.
- palmotea 2y ago> The intelligence that will be available to the average technically literate individual will be frightening. That's not the scary part. The scary part is the intelligence at scale that could be available to the average employer. Lots of us like to LARP that we're capitalists, but very few of us are. There's zero ideological or cultural framework in place to prioritize the well being of the general population over the profits of some capitalists. AI, especially accelerating AI, is bad news for anyone who needs to work for a living. It's not going to lead to a Star Trek fantasy. It means an eventual phase change for the economy that consigns us (and most consumer product companies) to wither and fade away.
- 101008 2y agoI agree with you and I am scared. My problem is: if most people can't work, who is going to pay for the product/services created with IA? I get a lot of "IA will allow us to create SaaS in a weekend" and "IA will take engineers jobs", which I think they both may be true. But a lot of SaaS surive because engineers pay for them -- if engineer don't exist anymore, a lot of SaaS won't either. If you eat your potential customers, creating quick SaaS doesn't make sense anymore (yeah, there are exceptions, etc., I know).
- koala_man 2y agoIt feels like we're back in 1900 when anyone's clever idea (and implementation) can give huge performance improvements, such as Ford's assembly line and Taylor's scientific management of optimizing shovel sizes for coal.
- andrewfromx 2y agoyes, it also feels like we are going to lose our just-in-time global shipments of anything to anywhere any day now. It will soon feel like 1900 in other ways.
- BobbyTables2 2y agoWe’ll have to raise our own chickens too…
- eru 2y agoHope we don't get 1914 again, too.
- xg15 2y agoI think the fact alone that distillation and quantization are techniques that can produce substantial improvements is a strong sign that we still have no real comprehensive understanding how the models work. If we had, there would be no reason to train a model with more parameters than are strictly necessary to represent the space's semantic structure. But then it should be impossible for distilled models with less parameters to come close to the performance of the original model. Yet this is what happens - the distilled or quantized models often come very close to the original model. So I think there are still many low-hanging fruits to pick.
- teruakohatu 2y ago> still have no real comprehensive understanding how the models work. We do understand how they work, we just have not optimised their usage. For example someone who has a good general understanding of how an ICE or EV car works. Even if the user interface is very unfamiliar, they can figure out how to drive any car within a couple of minutes. But that does not mean they can race a car, drift a car or drive a car on challenging terrain even if the car is physically capable of all these things.
- spiorf 2y agoWe know how the next token is selected, but not why doing that repeatedly brings all the capabilities it does. We really don't understand how the emergent behaviours emerge.
- ascorbic 2y agoI've noticed that R1 says "Wait," a lot in its reasoning. I wonder if there's something inherently special in that token.
- lionkor 2y agoSemantically, wait is a bit of a stop-and-breathe point. Consider the text: I think I'll go swimming today. Wait, ___ what comes next? Well, not something that would usually follow without the word "wait", probably something entirely orthogonal that impacts the earlier sentence in some fundamental way, like: Wait, I need to help my dad.
- ascorbic 2y agoYes, R1 seems to mostly use it like that. It's either to signal a problem with its previous reasoning, or if it's thought of a better approach. In coding it's often something like "this API won't work here" or "there's a simpler way to do this".
- fennecfoxy 2y agoI guess it goes to show how important reiteration is for general logic problems. And tbf when finding a solution to something myself I'll consider each part, and/or consider parts in relation to each other and/or consider all parts in relation to each other (on a higher level) before coming to a final solution. It's weird because I feel like we should've known that from work in general logic/problem solving studies, surely?
- katzenversteher 2y agoI bet a token like "sht!", "f*" or "damn!" would have the same or even stronger effect but the LLM creators would not like to have the users read them
- lodovic 2y agoI think you're onto something, however, as the training is done through on text and not actual thoughts, it may take some experimentation to find these stronger words.
- cyanydeez 2y agoits fascinating how certain political movements avoid that Wait moment...
- kevin009 2y agoThere are more than 10 different ways that I know for sure will improve LLMs just like `wait`. It is part if the CoT. I assume most researchers know this. CoT in old as 2019
- lostmsu 2y agoHm, I am surprised that people who are presumably knowledgeable with how attention works are surprised by this. The more tokens in the output, the more computation the model is able to do overall. Back in September, when I was testing my iOS hands-free voice AI prototype that was powered by 8B LLM, when I wanted it to give really thoughtful answers to philosophical questions, I would instruct it to output several hundred whitespace characters (because they are not read aloud) before the actual answer. What I am more surprised about is why models actually seem to have to produce "internal thoughts" instead of random tokens. Maybe during training having completely random tokens in thinking section derailed the model's thought process in a same way background noise can derail ours?
- deadbabe 2y agoI mean the “wait” thing is obvious if you’ve ever asked an LLM to look at its own response and ask if it’s really sure about its answer.
- rgovostes 2y ago> a branch of computer science It should be considered a distinct field. At some level there is overlap (information theory, Kolmogorov complexity, etc.), but prompt optimization and model distillation is far removed from computability, formal language theory, etc. The analytical methods, the techniques to create new architectures, etc. are very different beasts.
- BobbyTables2 2y agoAlmost seems more like computer engineering. Is it really that different than signal/image processing? I suspect CS departments don’t want to concede because they are now in the limelight…
- maginx 2y agoI agree - I don't know what field it formally is, but computer science it is not. It is also related to information retrieval aka "Google skills", problem presentation, 'theory of mind', even management and psychology. I'm saying the latter because people often ridicule AI responses for giving bad answers that are 'too AI'. But often it is simply because not enough context-specific information was given to allow the AI to giving a more personalized response. One should compare the response to "If I had asked a random person on the internet this query, what might I have gotten". If you write "The response should be written as a <insert characteristics, context, whatever you feel is relevant>" it will deliver a much less AI. This is just as much about how you pose a problem in general, as it is about computer science.
- BobbyTables2 2y agoMay sound like a conspiracy theory, but NVIDIA and a whole lot of AI startups have a strong vested interest to not seek+publish such findings. If I don’t need a huge model and GPU, then AI is little more than an open source program running on an idle PC. I feel like AI was NVIDIA’s lifeboat as GPU mining waned. Don’t see anything after that in the near future.
- philipswood 2y agoI think NVIDIAs future is pretty bright. We're getting to the run-your-capable-LLM on-prem or at-home territory. Without DeepSeek (and hopefully its successors) I wouldn't really have a usecase for something like NVIDIAs Project Digits. https://www.nvidia.com/en-us/project-digits/ https://www.nvidia.com/en-us/project-digits/
- Arn_Thor 2y agoExcept I can run R1 1.5b on a GPU-less and NPU-less Intel NUC from four-five years ago using half its cores and the reply speed is…functional. As the models have gotten more efficient and distillation better the minimum viable hardware for really cooking with LLMs has gone from a 4090 to suddenly something a lot of people already probably own. I definitely think a Digits box would be nice, but honestly I’m not sure I’ll need one.
- nickthegreek 2y agoR1 1.5b won’t do what most people want at all.
- Arn_Thor 2y agoNo, it won't. But that's not the point I was making
- fennecfoxy 2y agoYeah but what was R1 trained with? 50k GPUs as far as I've heard as well as distillation from OpenAI's models (basically leaning on their GPUs/GPU time). Besides the fact that consumers will still always want GPUs for gaming, rendering, science compute etc. No, I don't have any Nvidia stocks.
- tomaskafka 2y agoOne thing is to realize that we as humans have a thinking steps (internal monologue) before we output the texts. When LLMs produce text, we expect this thinking process to happen as well, but it does not - they are 'idiots that babble the first thing that comes to their minds'. The above 'hack' is one of many realizations of the above differences.
- codeulike 2y agoWait, so the trick is they reach into the context and basically switch '</think>' with 'wait' and that makes it carry on thinking?
- gield 2y agoYes, that's explicitly mentioned in the blog post: >In s1, when the LLM tries to stop thinking with "</think>", they force it to keep going by replacing it with "Wait".
- luc4sdreyer 2y agoYes, that's one of the tricks.
- danans 2y agoNot sure if your pun was intended, but 'wait' probably works so well because of the models being trained on text structured like your comment, where "wait" is followed by a deeper understanding.
- ozgune 2y agoAgreed. Here are three things that I find surreal about the s1 paper. (1) The abstract changed how I thought about this domain (advanced reasoning models). The only other paper that did that for me was the "Memory Resource Management in VMware ESX Server". And that paper got published 23 years ago. (2) The model, data, and code are open source at https://github.com/simplescaling/s1 https://github.com/simplescaling/s1. With this, you can start training your own advanced reasoning models. All you need is a thousand well-curated questions with reasoning steps. (3) More than half the references in the paper are from 2024 and Jan 2025. Just look at the paper's first page. https://arxiv.org/pdf/2501.19393 https://arxiv.org/pdf/2501.19393 In which other field do you see this?
- pradn 2y agoOmg, another fan of "Memory Resource Management in VMware ESX Server"!! It's one of my favorite papers ever - so clever.
- pradn 2y agoI mean is "wait" even the ideal "think more please" phrase? Would you get better results with other phrases like "wait, a second", or "let's double-check everything"? Or domain-dependent, specific instructions for how to do the checking? Or forcing tool-use?
- fennecfoxy 2y agoIn a way it's the same thing as finding that models got lazier closer to Christmas, ie the "Winter Break" hypothesis. Not sure what caused the above but In my opinion not only is the training affected by the date of training data (ie it refuses to answer properly because every year of the training data there was fewer or lower quality examples at the end of the year), or whether it's a cultural impression of humans talking about going on holiday/having a break etc in the training data at certain times and the model associating this with the meaning of "having a break". I still wonder if we're building models wrong by training them on a huge amount of data from the Internet, then fine tuning for instruct where the model learns to make certain logical associations inherent or similar to the training data (which seems to introduce a myriad of issues like the strawberry problem or is x less than y being incorrect). I feel like these models would have a lot more success if we trained a model to learn logic/problem solving separately without the core data set or to restrict the instruct fine tuning in some way so that we reduce the amount of "culture" it gleans from the data. There's so much that we don't know about this stuff yet and it's so interesting to see something new in this field every day. All because of a wee paper on attention.