31 ms·
S1: A $6 R1 competitor?
- sheepscreek 2y agoLLMs still feel so magical. It’s like quantum physics. “I get it” but I don’t. Not really. I don’t think I ever will. Perhaps a human mind can only comprehend so much.
- bberenberg 2y agoIn case you’re not sure what S1 is, here is the original paper: https://arxiv.org/html/2501.19393v1 https://arxiv.org/html/2501.19393v1
- mi_lk 2y agoit's also the first link in the article's first sentence
- bberenberg 2y agoGood call, I must have missed it. I read the whole blog then went searching for what S1 was.
- deleted 2y ago[deleted]
- addandsubtract 2y agoIt's linked in the blog post, too. In the first sentence, actually, but for some reason the author never bothered to attach the name to it. As if keeping track of o1, 4o, r1, r2d2, wasn't exhausting enough already.
- kgwgk 2y ago> for some reason the author never bothered to attach the name to it Respect for his readers’ intelligence, maybe.
- rahimnathwani 2y agoTo enforce a minimum, we suppress the generation of the end-of-thinking token delimiter and optionally append the string “Wait” to the model’s current reasoning trace to encourage the model to reflect on its current generation. Does this mean that the end-of-thinking delimiter is a single token? Presumably </think> or similar wasn't a single token for the base model. Did they just pick a pair of uncommon single-token symbols to use as delimiters? EDIT: Never mind, end of thinking is represented with <|im_start|> followed by the word 'answer', so the code dynamically adds/removes <|im_start|> from the list of stop tokens.
- dagurp 2y agoI don't know what R1 is either
- latexr 2y agoIt’s the DeepSeek reasoning model.
- ttyprintk 2y agohttps://huggingface.co/simplescaling https://huggingface.co/simplescaling
- anentropic 2y agoand: https://github.com/simplescaling/s1 https://github.com/simplescaling/s1
- mettamage 2y agoWhen you're only used to ollama, how do I go about using this model?
- davely 2y agoI think we need to wait for someone to convert it into a GGUF file format. However, once that happens, you can run it (and any GGUF model) from Hugging Face![0] [0] https://huggingface.co/docs/hub/en/ollama https://huggingface.co/docs/hub/en/ollama
- mettamage 2y agoSo this? https://huggingface.co/brittlewis12/s1-32B-GGUF https://huggingface.co/brittlewis12/s1-32B-GGUF
- withinboredom 2y agooh god, this is terrible! I just said "Hello!" and it went off the rails.
- delijati 2y agowhy how what? can you add a sample prompt with output ?
- 2y ago
- yapyap 2y ago> If you believe that AI development is a prime national security advantage, then you absolutely should want even more money poured into AI development, to make it go even faster. This, this is the problem for me with people deep in AI. They think it’s the end all be all for everything. They have the vision of the ‘AI’ they’ve seen in movies in mind, see the current ‘AI’ being used and to them it’s basically almost the same, their brain is mental bridging the concepts and saying it’s only a matter of time. To me, that’s stupid. I observe the more populist and socially appealing CEOs of these VC startups (Sam Altman being the biggest, of course.) just straight up lying to the masses, for financial gain, of course. Real AI, artificial intelligence, is a fever dream. This is machine learning except the machines are bigger than ever before. There is no intellect. and the enthusiasm of these people that are into it feeds into those who aren’t aware of it in the slightest, they see you can chat with a ‘robot’, they hear all this hype from their peers and they buy into it. We are social creatures after all. I think using any of this in a national security setting is stupid, wasteful and very, very insecure. Hell, if you really care about being ahead, pour 500 billion dollars into quantum computing so u can try to break current encryption. That’ll get you so much further than this nonsensical bs.
- mnky9800n 2y agoAlso the narrative that we are currently on the brink of Ai explosion and this random paper shows it has been the same tired old story handed out by ai hawks for years now. Like yes, I agree with the general idea that more compute means more progress for humans and perhaps having a more responsive user interface through some kind of ai type technology would be good. But I don’t see why that will turn into Data from Star Trek. But I also think all these ai hawks kind of narcissistically over value their own being. Like blink and their lives are over in the grand scheme of things. Maybe our “awareness” of the world around us is an illusion provided by evolution because we needed it to value self preservation whereas other animals don’t. There is an inherent belief in the specialness of humans that I suppose I mistrust.
- encipriano 2y agoI find the last part of the paragraph offputting and I agree
- GTP 2y agoSorry for being lazy, but I just don't have the time right now to read the paper. Is there in the paper or somewhere else a comparison based on benchmarks of S1 vs R1 (the full R1, not quantized or distilled)?
- pama 2y agoThe S1 paper is not meant to compete with R1. It simply shows that with 1k well curated examples for finetuning (26 minutes training on 16 GPU) and with a simple hack for controlling the length of the thinking process, one can dramatically increase the performance of a non-reasoning model and show a clear increase in benefit with increased test-time compute. It is worth a quick skim.
- swiftcoder 2y ago> having 10,000 H100s just means that you can do 625 times more experiments than s1 did I think the ball is very much in their court to demonstrate they actually are using their massive compute in such a productive fashion. My BigTech experience would tend to suggest that frugality went out the window the day the valuation took off, and they are in fact just burning compute for little gain, because why not...
- whizzter 2y agoMainly it points to a non-scientific "bigger is better" mentality, and the researchers probably didn't mind playing around with the power because "scale" is "cool". Remember that the Lisp AI-labs people were working on non-solved problems on absolute potatoes of computers back in the day, we have a semblance of progress solution but so much of it has been brute-force (even if there has been improvements in the field). The big question is if these insane spendings has pulled the rug on real progress if we head into another AI winter of disillusionment or if there is enough real progress just around the corner to show that there is hope for investors in a post-deepseek valuation hangover.
- wongarsu 2y agoWe are in a phase where costs are really coming down. We had this phase from GPT2 to about GPT4 where the key to building better models was just building bigger models and training them for longer. But since then a lot of work has gone into distillation and other techniques to make smaller models more capable. If there is another AI winter, it will be more like the dotcom bubble: lots of important work got done in the dotcom bubble, but many of the big tech companies started from the fruits of that labor in the decade after the bubble burst
- deleted 2y ago[deleted]
- svantana 2y agoBesides that, AI training (aka gradient descent) is not really an "embarrassingly parallel" problem. At some point, there are diminishing returns on adding more GPUs, even though a lot of effort is going into making it as parallel as possible.
- HenryBemis 2y ago> Going forward, it’ll be nearly impossible to prevent distealing (unauthorized distilling). One thousand examples is definitely within the range of what a single person might do in normal usage, no less ten or a hundred people. I doubt that OpenAI has a realistic path to preventing or even detecting distealing outside of simply not releasing models. (sorry for the long quote) I will say (naively perhaps) "oh but that is fairly simple". For any API request, add a counter of 5 seconds to the next for 'unverified' users. Make the "blue check" (a-la X/Twitter). For the 'big sales' have a third-party vetting process so that if US Corporation XYZ wants access, they prove themselves worthy/not Chinese competition and then you do give them the 1000/min deal. For everyone else, add the 5 second (or whatever other duration makes sense) timer/overhead and then see them drop from 1000 requests per minutes to 500 per day. Or just cap them at 500 per day and close that back-door. And if you get 'many cheap accounts' doing hand-overs (AccountA does 1-500, AccountB does 501-1000, AccountC does 1001-1500, and so on) then you mass block them.
- mark_l_watson 2y agoOff topic, but I just bookmarked Tim’s blog, great stuff. I dismissed the X references to S1 without reading them, big mistake. I have been working generally in AI for 40 hears and neural networks for 35 years and the exponential progress since the hacks that make deep learning possible has been breathtaking. Reduction in processing and memory requirements for running models is incredible. I have been personally struggling with creating my own LLM-based agents with weaker on-device models (my same experiments usually work with 4o-mini and above models) but either my skills will get better or I can wait for better on device models. I was experimenting with the iOS/iPadOS/macOS app On-Device AI last night and the person who wrote this app was successful in combining web search tool calling working with a very small model - something that I have been trying to perfect.
- cowsaymoo 2y agoThe part about taking control of a reasoning model's output length using <think></think> tags is interesting. > In s1, when the LLM tries to stop thinking with "</think>", they force it to keep going by replacing it with "Wait". I had found a few days ago that this let you 'inject' your own CoT and jailbreak it easier. Maybe these are related? https://pastebin.com/G8Zzn0Lw https://pastebin.com/G8Zzn0Lw https://news.ycombinator.com/item?id=42891042#42896498 https://news.ycombinator.com/item?id=42891042#42896498
- Havoc 2y agoThe point about agents to conceal access to the model is a good one. Hopefully we won’t lose all access to models in future
- cyp0633 2y agoQwen's QvQ-72B does much more "wait"s than other LLMs with CoT I tried, maybe they've somewhat used that trick already?
- theturtletalks 2y agoDeepseek R1 uses <think/> and wait and you can see it in the thinking tokens second guessing itself. How does the model know when to wait? These reasoning models are feeding more to OP's last point about NVidia and OpenAI data centers not being wasted since reason models require more tokens and faster tps.
- qwertox 2y agoProbably when it would expect a human to second guess himself, as shown in literature and maybe other sources.
- UncleEntity 2y agoFrom playing around they seem to 'wait' when there's a contradiction in their logic. And I think the second point is due to The Market thinking there is no need to spend ever increasing amounts of compute to get to the next level of AI overlordship. Of course Jevon's paradox is also all in the news these days..
- pona-a 2y agoIf chain of thought acts as a scratch buffer by providing the model more temporary "layers" to process the text, I wonder if making this buffer a separate context with its own separate FNN and attention would make sense; in essence, there's a macroprocess of "reasoning" that takes unbounded time to complete, and then there's a microprocess of describing this incomprehensible stream of embedding vectors in natural language, in a way returning to the encoder/decoder architecture but where both are autoregressive. Maybe this would give us a denser representation of said "thought", not constrained by imitating human text.
- bluechair 2y agoI had this exact same thought yesterday. I’d go so far as to add one more layer to monitor this one and stop adding layers. My thinking is that this meta awareness is all you need. No data to back my hypothesis up. So take it for what it’s worth.
- larodi 2y agoMy thought on the same guess being - all tokens live in same latent space or in many spaces and each logical units train separate of each other…?
- hadlock 2y agoThis is where I was headed but I think you said it better. Some kind of executive process monitoring the situation, the random stream of consciousness and the actual output. Looping back around to outdated psychology you have the ego which is the output (speech), the super ego is the executive process and the id is the <think>internal monologue</think>. This isn't the standard definition of those three but close enough.
- whimsicalism 2y ago> this incomprehensible stream of embedding vectors as natural language explanation, in a way returning to encoder/decoder architecture this is just standard decoding, the stream of vectors is called the k/v cache
- 2y ago
- sambull 2y agoThat sovereign wealth fund with tik tok might set a good precedent; when we have to 'pour money' into these companies we can do so with stake in them held in our sovereign wealth fund.
- TehCorwiz 2y agoExtra-legal financial instruments meant to suck money from other federal departments don't strike me as a good precedent in any sense. I don't disagree though that nationalizing the value of enormous public investments is something we should be considering, looking at you oil industry. But until congress appropriates the money under law it's a pipe dream or theft.
- ipnon 2y agoAll you need is attention and waiting. I feel like a zen monk.
- jebarker 2y agoS1 (and R1 tbh) has a bad smell to me or at least points towards an inefficiency. It's incredible that a tiny number of samples and some inserted <wait> tokens can have such a huge effect on model behavior. I bet that we'll see a way to have the network learn and "emerge" these capabilities during pre-training. We probably just need to look beyond the GPT objective.
- pas 2y agocan you please elaborate on the wait tokens? what's that? how do they work? is that also from the R1 paper?
- jebarker 2y agoThe same idea is in both the R1 and S1 papers (<think> tokens are used similarly). Basically they're using special tokens to mark in the prompt where the LLM should think more/revise the previous response. This can be repeated many times until some stop criteria occurs. S1 manually inserts these with heuristics, R1 learns the placement through RL I think.
- whimsicalism 2y ago? theyre not special tokens really
- jebarker 2y agoi'm not actually sure whether they're special tokens in the sense of being in the vocabulary
- whimsicalism 2y ago<think> might be i think "wait" is tokenized like any other in the pretraining
- throwaway314155 2y ago
- light_hue_1 2y agoS1 has no relationship to R1. It's a marketing campaign for an objectively terrible and unrelated paper. S1 is fully supervised by distilling Gemini. R1 works by reinforcement learning with a much weaker judge LLM. They don't follow the same scaling laws. They don't give you the same results. They don't have the same robustness. You can use R1 for your own problems. You can't use S1 unless Gemini works already. We know that distillation works and is very cheap. This has been true for a decade; there's nothing here. S1 is a rushed hack job (they didn't even run most of their evaluations with an excuse that the Gemini API is too hard to use!) that probably existed before R1 was released and then pivoted into this mess.
- danju 2y ago[dead]
- bloomingkales 2y agoThis thing that people are calling “reasoning” is more like rendering to me really, or multi pass rendering. We’re just refining the render, there’s no reasoning involved.
- frontalier 2y agosshhhh, let the money flow
- dleslie 2y agoThat was succinct and beautifully stated. Thank-you for the "Aha!" moment.
- bloomingkales 2y agoHah. You should check out my other comment on how I think we’re obviously in a simulation (remember, we just need to see a good enough render). LLMs are changing how I see reality.
- whimsicalism 2y ago[flagged]
- mistermann 2y ago"...there’s no reasoning involved...wait, could I just be succumbing to my heuristic intuitions of what is (seems to be) true....let's reconsider using System 2 thinking..."
- bloomingkales 2y agoOr there is no objective reality (well there isn’t, check out the study), and reality is just a rendering of the few state variables that keep track of your simple life. A little context about you: - person - has hands, reads HN These few state variables are enough to generate a believable enough frame in your rendering. If the rendering doesn’t look believable to you, you modify state variables to make the render more believable, eg: Context: - person - with hands - incredulous demeanor - reading HN Now I can render you more accurately based on your “reasoning”, but truly I never needed all that data to see you. Reasoning as we know it could just be a mechanism to fill in gaps in obviously sparse data (we absolutely do not have all the data to render reality accurately, you are seeing an illusion). Go reason about it all you want.
- whimsicalism 2y agothis isn't rlvr and so sorta uninteresting, they are just distilling the work already done
- bloomingkales 2y agoIf an LLM output is like a sculpture, then we have to sculpt it. I never did sculpting, but I do know they first get the clay spinning on a plate. Whatever you want to call this “reasoning” step, ultimately it really is just throwing the model into a game loop. We want to interact with it on each tick (spin the clay), and sculpt every second until it looks right. You will need to loop against an LLM to do just about anything and everything, forever - this is the default workflow. Those who think we will quell our thirst for compute have another thing coming, we’re going to be insatiable with how much LLM brute force looping we will do.
- MrLeap 2y agoThis is a fantastic insight and really has my gears spinning. We need to cluster the AI's insights on a spatial grid hash, give it a minimap with the ability to zoom in and out, and give it the agency to try and find its way to an answer and build up confidence and tests for that answer. coarse -> fine, refine, test, loop. Maybe a parallel model that handles the visualization stuff. I imagine its training would look more like computer vision. Mind palace generation. If you're stuck or your confidence is low, wander the palace and see what questions bubble up. Bringing my current context back through the web is how I think deeply about things. The context has the authority to reorder the web if it's "epiphany grade". I wonder if the final epiphany at the end of what we're creating is closer to "compassion for self and others" or "eat everything."
- zoogeny 2y agoI can't believe this hasn't been done yet, perhaps it is a cost issue. My literal first thought about AI was wondering why we couldn't just put it in a loop. Heck, one update per day, or one update per hour would even be a start. You have a running "context", the output is the next context (or a set of transformations on a context that is a bit larger than the output window). Then ramp that up ... one loop per minute, one per second, millisecond, microsecond.
- layer8 2y agoSame. And the next step is that it must feed back into training, to form long-term memory and to continually learn.
- incrudible 2y agoHmmm, 1 + 1 equals 3. Alternatively, 1 + 1 equals -3. Wait, actually 1 + 1 equals 1.
- falcor84 2y agoAs one with teaching experience, the idea of asking a student "are you sure about that?" is to get them to think more deeply rather than just blurting a response. It doesn't always work, but it generally does.
- latexr 2y agoIt works because the question itself is a hint born of knowledge. “Are you sure about that” is a polite way to say “that answer is wrong, try again”. Students know that, so instead of doubling down will redo their work with the assumption they made a mistake. It is much rarer to ask the question when the answer is correct, and in fact doing so is likely to upset the learner because they had to redo the work for no reason. If you want a true comparison, start asking that question every time and then compare. My hypothesis is students would start ignoring the prompt and answering “yes” every time to get on with it.
- ALittleLight 2y agoAt 6 dollars per run, I'm tempted to try to figure out how to replicate this. I'd like to try some alternatives to "wait" - e.g. "double checking..." Or write my own chains of thought.
- qup 2y agoLike the ones they tested?
- ALittleLight 2y agoYes, that is what "replicate" with my own ideas means.
- kittikitti 2y agoThank you for this, I really appreciate this article and I learned a bunch!
- Aperocky 2y agoFor all the hype about thinking models, this feels much like compression in terms of information theory instead of a "takeoff" scenario. There are a finite amount of information stored in any large model, the models are really good at presenting the correct information back, and adding thinking blocks made the models even better at doing that. But there is a cap to that. Just like how you can compress a file by a lot, there is a theoretical maximum to the amount of compression before it starts becoming lossy. There is also a theoretical maximum of relevant information from a model regardless of how long it is forced to think.
- psadri 2y agoI think an interesting avenue to explore is creating abstractions and analogies. If a model can take a novel situation and create an analogy to one that it is familiar with, it would expand its “reasoning” capabilities beyond its training data.
- zoogeny 2y agoI think this is probably accurate and what remains to be seen is how "compressible" the larger models are. The fact that we can compress a GPT-3 sized model into an o1 competitor is only the beginning. Maybe there is even more juice to squeeze there? But even more, how much performance will we get out of o3 sized models? That is what is exciting since they are already performing near Phd levels on most evals.
- jedbrooke 2y agomy thinking (hope?) is that the reasoning models will be more like how a calculator doesn’t have to “remember” all the possible combinations of addition, multiplication, etc for all the numbers, but can actually compute the results. As reasoning improves the models could start with a basic set of principles and build from there. Of course for facts grounded in reality RAG would still likely be the best, but maybe with enough “reasoning” a model could simulate an approximation of the universe well enough to get to an answer.
- hidelooktropic 2y ago> I doubt that OpenAI has a realistic path to preventing or even detecting distealing outside of simply not releasing models. Couldn't they just start hiding the thinking portion? It would be easy for them to do this. Currently, they already provide one sentence summaries for each step of the thinking I think users would be fine or at least stay if it were changed to provide only that.
- mtrovo 2y agoI found the discussion around inference scaling with the 'Wait' hack so surreal. The fact such an ingeniously simple method can impact performance makes me wonder how many low-hanging fruit we're still missing. So weird to think that improvements on a branch of computer science is boiling down to conjuring the right incantation words, how you even change your mindset to start thinking this way?
- nyoomboom 2y agoI think a skill here is learning a bias for experimentation and accepting the results one finds. Also the book "Why Greatness Cannot Be Planned" showcases the kind of open ended play that results in people discovering stuff like this.
- cubefox 2y agoNow imagine where we are in 12 months from now. This article from February 5 2025 will feel quaint by then. The acceleration keeps increasing. It seems likely we will soon have recursive self-improving AI -- reasoning models which do AI research. This will accelerate the rate of acceleration itself. It sounds stupid to say it, but yes, the singularity is near. Vastly superhuman AI now seems to arrive within the next few years. Terrifying.
- gom_jabbar 2y agoYes, and Accelerationism predicted this development back in the 1990s, perhaps most prominently in the opening lines of Nick Land's Meltdown (1994) text: [[ ]] The story goes like this: Earth is captured by a technocapital singularity as renaissance rationalization and oceanic navigation lock into commoditization take-off. Logistically accelerating techno-economic interactivity crumbles social order in auto-sophisticating machine runaway. As markets learn to manufacture intelligence, politics modernizes, upgrades paranoia, and tries to get a grip. > reasoning models which do AI research In the introduction to my research project on Accelerationism [0], I write: Faced with the acceleration of progress in Artificial Intelligence (AI) — with AI agents now automating AI research and development —, Accelerationism no longer seems like an abstract philosophy producing empty hyperstitional hype, but like a sober description of reality. The failed 2023 memorandum to stop AI development on systems more powerful than OpenAI's ChatGPT-4 perfectly illustrates the phenomenological aspects of Accelerationism: "To be rushed by the phenomenon, to the point of terminal institutional paralysis, is the phenomenon." [1] At the current rate of acceleration, if you don't write hyperstitionally, your texts are dead on arrival. [0] https://retrochronic.com/ https://retrochronic.com/ [1] Nick Land (2017). A Quick-and-Dirty Introduction to Accelerationism in Jacobite Magazine.
- maksimur 2y agoIt appears that someone has implemented a similar approach for DeepSeek-R1-Distill-Qwen-1.5B: https://reddit.com/r/LocalLLaMA/comments/1id2gox/improving_deepseek_r1_reasoning_trace/ https://reddit.com/r/LocalLLaMA/comments/1id2gox/improving_d... I hope it gets tested further.
- nullbyte 2y agoGreat article! I enjoyed reading it
- khazhoux 2y agoI have a bunch of questions, would love for anyone to explain these basics: * The $5M DeepSeek-R1 (and now this cheap $6 R1) are both based on very expensive oracles (if we believe DeepSeek-R1 queried OpenAI's model). If these are improvements on existing models, why is this being reported as decimating training costs? Isn't fine-tuning already a cheap way to optimize? (maybe not as effective, but still) * The R1 paper talks about improving one simple game - Countdown. But the original models are "magic" because they can solve a nearly uncountable number of problems and scenarios. How does the DeepSeek / R1 approach scale to the same gigantic scale? * Phrased another way, my understanding is that these techniques are using existing models as black-box oracles. If so, how many millions/billions/trillions of queries must be probed to replicate and improve the original dataset? * Is anything known about the training datasets used by DeepSeek? OpenAI used presumably every scraped dataset they could get their hands on. Did DS do the same?
- UncleEntity 2y ago> If these are improvements on existing models, why is this being reported as decimating training costs? Because that's what gets the clicks... Saying they spent a boatload of money on the initial training + iteration + final fine-tuning isn't as headline grabbing as "$5 million trained AI beats the pants off the 'mericans".
- torginus 2y agoIf what you say is true, and distilling LLMs is easy and cheap, and pushing the SOTA without a better model to rely on is dang hard and expensive, then that means the economics of LLM development might not be attractive to investors - spending billions to have your competitors come out with products that are 99% as good, and cost them pennies to train, does not sound like a good business strategy.
- khazhoux 2y agoWhat I still don’t understand is how one slurps out an entire model (closed source) though. Does the deepseek paper actually say what model it’s trained off of, or do they claim the entire thing is from scratch?
- janalsncm 2y ago> even the smartest people make hundreds of tiny experiments This is the most important point, and why DeepSeek’s cheaper training matters. And if you check the R1 paper, they have a section for “things that didn’t work”, each of which would normally be a paper of its own but because their training was so cheap and streamlined they could try a bunch of things.
- robrenaud 2y ago> "Note that this s1 dataset is distillation. Every example is a thought trace generated by another model, Qwen2.5" The traces are generated by Gemini Flash Thinking. 8 hours of H100 is probably more like $24 if you want any kind of reliability, rather than $6.
- zaptrem 2y ago"You can train a SOTA LLM for $0.50" (as long as you're distilling a model that cost $500m into another pretrained model that cost $5m)
- fizx 2y agoThat's absolutely fantastic, because if you have 1 good idea that's additive to the SOTA, you can test it for a dollar, not millions
- knutzui 2y agoThe original statement stands, if what you are suggesting in addition to it is true. If the initial one-time investment of $505m is enough to distill new SOTA models for $0.50 a piece, then the average cost for subsequent models will trend toward $0.50.
- nico 2y ago> Why did it cost only $6? Because they used a small model and hardly any data. > After sifting their dataset of 56K examples down to just the best 1K, they found that the core 1K is all that’s needed to achieve o1-preview performance on a 32B model. Adding data didn’t raise performance at all. > 32B is a small model, I can run that on my laptop. They used 16 NVIDIA H100s for 26 minutes per training run, that equates to around $6.
- nico 2y ago> In s1, when the LLM tries to stop thinking with "</think>", they force it to keep going by replacing it with "Wait". It’ll then begin to second guess and double check its answer. They do this to trim or extend thinking time (trimming is just abruptly inserting "</think>") I know some are really opposed to anthropomorphizing here, but this feels eerily similar to the way humans work, ie. if you just dedicate more time to analyzing and thinking about the task, you are more likely to find a better solution It also feels analogous to navigating a tree, the more time you have to explore the nodes, the bigger the space you'll have covered, hence higher chance of getting a more optimal solution At the same time, if you have "better intuition" (better training?), you might be able to find a good solution faster, without needing to think too much about it
- layer8 2y agoWhat’s missing in that analogy is that humans tend to have a good hunch about when they have to think more and when they are “done”. LLMs seem to be missing a mechanism for that kind of awareness.
- nico 2y agoGreat observation. Maybe an additional “routing model” could be trained to predict when it’s better to think more vs just using the current result
- sanxiyn 2y agoLLMs actually do have such hunch, they just don't utilize it. You can literally ask them "Would you do better if you started over?" and start over if answer is yes. This works. https://arxiv.org/abs/2410.02725 https://arxiv.org/abs/2410.02725
- janalsncm 2y agoI think a lot of people in the ML community were excited for Noam Brown to lead the O series at OpenAI because intuitively, a lot of reasoning problems are highly nonlinear i.e. they have a tree-like structure. So some kind of MCTS would work well. O1/O3 don’t seem to use this, and DeepSeek explicitly mentioned difficulties training such a model. However, I think this is coming. DeepSeek mentioned it was hard to learn a value model for MCTS from scratch, but this doesn’t mean we couldn’t seed it with some annotated data.
- insane-c0der 2y agoDo you have a reference for us to check? - "DeepSeek explicitly mentioned difficulties training such a model."
- janalsncm 2y agoSection 4.2: Unsuccessful attempts https://arxiv.org/pdf/2501.12948 https://arxiv.org/pdf/2501.12948
- talles 2y agoAnyone else wants more articles on how those benchmarks are created and how they work? Those models can be trained in way tailored to have good results on specific benchmarks, making them way less general than it seems. No accusation from me, but I'm skeptical on all the recent so called 'breakthroughs'.
- charlieyu1 2y ago> having 10,000 H100s just means that you can do 625 times more experiments than s1 did The larger the organisation, the less experiments you can afford to do. Employees are mostly incentivised by getting something done quick enough to not to be fired in this job market. They know that the higher-ups would get them off for temporary gains. Rush this deadline, ship that feature, produce something that looks OK enough.
- mmoustafa 2y agoLove the look under the hood! Specially discovering some AI hack I came up with is how the labs are doing things too. In this case, I was also forcing R1 to continue thinking by replacing </think> with “Okay,” after augmenting reasoning with web search results. https://x.com/0xmmo/status/1886296693995646989 https://x.com/0xmmo/status/1886296693995646989
- ConanRus 2y agoWait
- bxtt 2y agoCoT is widely known technique - what became fully novel was the level of training embedding CoT via RL with optimal reward trajectory. DeepSeek took it further due to their compute restriction to find memory, bandwidth, parallelism optimizations in every part (GRPO - reducing memory copies, DualPipe for data batch parallelism between memory & compute, kernel bypasses (PTX level optimization), etc.) - then even using MoE due to sparse activation and further distillation. They operated on the power scaling laws of parameters & tokens but high quality data circumvents this. I’m not surprised they utilized synthetic generation from OpenAI or copied the premise of CoT, but where they should get the most credit is their infra level & software level optimizations. With that being said, I don’t think the benchmarks we currently have are strong enough and the next frontier models are yet to come. I’m sure at this point U.S LLM research firms now understand their lack of infra/hardware optimizations (they just threw compute at the problem), they will begin paying closer attention. Now their RL-level and parent training will become even greater - whilst the newly freed resources to solve for sub-optimizations that have been traditionally avoided due to computational overhead
- cadamsdotcom 2y agoMaybe this is why OpenAI hides o1/o3 reasoning tokens - constraining output at inference time seems to be easy to implement for other models and others would immediately start their photocopiers. It also gave them a few months to recoup costs!
- mangoman 2y agoFrom the S1 paper: > Second, we develop budget forcing to control test-time compute by forcefully terminating the model's thinking process or lengthening it by appending "Wait" multiple times to the model's generation when it tries to end I'm feeling proud of myself that I had the crux of the same idea almost 6 months ago before reasoning models came out (and a bit disappointed that I didn't take this idea further!). Basically during inference time, you have to choose the next token to sample. Usually people just try to sample the distribution using the same sampling rules at each step.... but you don't have to! you can selectively insert words into the the LLM's mouth based on what it said previously or what it wants to say, and decide "nah, say this instead". I wrote a library so that you could sample an LLM using llama.cpp in swift and you could write rules to sample tokens and force tokens into the sequence depending on what was sampled. https://github.com/prashanthsadasivan/LlamaKit/blob/main/Tests/LlamaKitTests/LlamaKitTests.swift#L35-L52 https://github.com/prashanthsadasivan/LlamaKit/blob/main/Tes... Here, I wrote a test that asks Phi-3 instruct "how are you" and it if it tried to say "as an AI I don't have feelings" or "I'm doing " I forced it to say "I'm doing poorly" and refuse to help since it was always so dang positive. It sorta worked, though the instruction tuned models REALLY want to help. But at the time I just didn't have a great use case for it - I had thought about a more conditional extension to llama.cpp's grammar sampling (you could imagine changing the grammar based on previously sampled text), or even just making it go down certain paths, but I just lost steam because I couldn't describe a killer use case for it. This is that killer use case! forcing it to think more is such a great usecase for inserting ideas into the LLM's mouth, and I feel like there must be more to this idea to explore.
- jwrallie 2y agoSo what you mean is that if the current train of thought is going in a direction we find to be not optimal, we could just interrupt it and hint it into the right direction? That sounds very useful, albeit a bit different than how current "chat" implementations would work, as in you could control both ways of the conversation.
- latexr 2y ago> and a bit disappointed that I didn't take this idea further! Don’t be, that’s pretty common. https://en.wikipedia.org/wiki/Multiple_discovery https://en.wikipedia.org/wiki/Multiple_discovery
- Caitlynmeeks 2y agohttps://imgflip.com/i/9j833q https://imgflip.com/i/9j833q (ptheven)
- shaneofalltrad 2y agoWell dang, I am great at tinkering like this because I can’t remember things half the time. I wonder if the ADHD QA guy solved this for the devs?
- gorgoiler 2y agoThis feels just like telling a constraint satisfaction engine to backtrack and find a more optimal route through the graph. We saw this 25 years ago with engines like PROVERB doing directed backtracking, and with adversarial planning when automating competitive games. Why would you control the inference at the token level? Wouldn’t the more obvious (and technically superior) place to control repeat analysis of the optimal path through the search space be in the inference engine itself? Doing it by saying “Wait” feels like fixing dad’s laptop over a phone call. You’ll get there, but driving over and getting hands on is a more effective solution. Realistically, I know that getting “hands on” with the underlying inference architecture is way beyond my own technical ability. Maybe it’s not even feasible, like trying to fix a cold with brain surgery?
- Nurbek-F 2y agoTotally agreed this is not a solution we are looking for, in fact this is the only solution we have in our hands right now. It's a good step forward.
- code_biologist 2y agoWhat would a superior control approach be? It's not clear to me how to get an LLM to be an LLM if you're not doing stochastic next token prediction. Given that, the model itself is going to know best how to traverse its own concept space. The R1 chain of thought training encourages and develops exactly that capability. Still, you want that chain of thought to terminate and not navel gaze endlessly. So how to externally prod it to think more when it does terminate? Replacing thought termination with a linguistic signifier of continued reasoning plus novel realization seems like a charmingly simple, principled, and general approach to continue to traverse concept space.
- rayboy1995 2y agoThis is the difference between science and engineering. What they have done is engineering. If the result is 90% of the way there with barely any effort, its best to move on to something else that may be low hanging fruit than to spend time chasing that 10%.
- stefanoco 2y agoIs it me, or the affiliations are totally missing in the cited paper?? Looks like they come from a mix of UK / US institutions
- advael 2y agoI'm strictly speaking never going to think of model distillation as "stealing." It goes against the spirit of scientific research, and besides every tech company has lost my permission to define what I think of as theft forever
- surajrmal 2y agoMaybe but something has gotta pay the bills to justify the cutting edge. I guess it's a similar problem to researching medicine.
- ClumsyPilot 2y agoWell the artists and writers also want to pay their bills. We threw them under the bus, might as well throw openAI too and get an actual open AI that we can use
- advael 2y agoThe investment thrown at OpenAI seems deeply inflated for how much meaningful progress they're able to make with it I think it's clear that innovative breakthroughs in bleeding-edge research are not just a matter of blindly hurling more money at a company to build unprecedentedly expensive datacenters But also, even if that was a way to do it, I don't think we should be wielding the law to enable privately-held companies to be at the forefront of research, especially in such a grossly inconsistent manner
- eru 2y agoAt most it would be illicit copying. Though it's poetic justice that OpenAI is complaining about someone else playing fast and loose with copyright rules.
- tomrod 2y agoStochastic decompression. Dass-it.
- downrightmike 2y agoThe First Amendment is not just about free speech, but also the right to read, the only question is if AI has that right.
- svara 2y agoIt just occurred to me that if you squint a little (just a little!) the S1 paper just provided the scientific explanation for why Twitter's short tweets mess you up and books are good for you. Kidding, but not really. It's fascinating how we seem to be seeing a gradual convergence of machine learning and psychology.
- mig1 2y agoThis argument that the data centers and all the GPUs will be useful even in the context of Deepseek doesn't add up... basically they showed that it's diminishing returns after a certain amount. And so far it didn't make OpenAI or Anthropic go faster, did it?
- rayboy1995 2y agoWhat is the source for the diminishing returns? I would like to read about it as I have only seen papers referring to the scaling law still applying.
- adamc 2y agoI found it interesting but the "Wait" vs. "Hmm" bit just made me think we don't really understand our own models here. I mean, sure, it's great that they measured and found something better, but it's kind of disturbing that you have to guess.
- leopoldj 2y ago>it can run on my laptop Has anyone run it on a laptop (unquantized)? Disk size of the 32B model appears to be 80GB. Update: I'm using a 40GB A100 GPU. Loading the model took 30GB vRAM. I asked a simple question "How many r in raspberry". After 5 minutes nothing got generated beyond the prompt. I'm not sure how the author ran this on a laptop.
- coder543 2y ago32B models are easy to run on 24GB of RAM at a 4-bit quant. It sounds like you need to play with some of the existing 32B models with better documentation on how to run them if you're having trouble, but it is entirely plausible to run this on a laptop. I can run Qwen2.5-Instruct-32B-q4_K_M at 22 tokens per second on just an RTX 3090.
- leopoldj 2y agoMy question was about running it unquantized. The author of the article didn't say how he ran it. If he quantized it then saying he ran it on a laptop is not a news.
- kristianp 2y agoMaybe he has a 64GB laptop. Also he said he can run it, not that he actually tried it.
- coder543 2y agoI can't imagine why anyone would run it unquantized, but there are some laptops with the more than 70GB of RAM that would be required. It's not that it can't be done... it's just that quantizing to at least 8-bit seems to be standard practice these days, and DeepSeek has shown that it's even worth training at 8-bit resolution.
- mountainriver 2y ago> They used 16 NVIDIA H100s for 26 minutes per training run, that equates to around $6 Running where? H100s are usually over $2/hr, thats closer to $25
- kristianp 2y agoA couple of GGUF quantizations have been posted: https://huggingface.co/bartowski/simplescaling_s1-32B-GGUF https://huggingface.co/bartowski/simplescaling_s1-32B-GGUF https://huggingface.co/brittlewis12/s1-32B-GGUF https://huggingface.co/brittlewis12/s1-32B-GGUF
- vagab0nd 2y agoCool trick. But is this better than reinforcement learning, where the LLM decides for itself the optimal thinking time for each prompt?
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- _befireHack 2y agoI work at a mid-sized research firm, and there’s this one coworker who completely turned her performance around. A complete 180. A few months ago, she was one of the slowest on the team, now she’s always the first to get her work done. I was curious, so I asked her what changed. She just laughed and said she just used an AI tool that she randomly found on YouTube to do 90% of her work. We’ve been working on a project together, and every morning for the past two months, she’s sent me clean, perfectly organized FED data. I assumed she was just working late to get ahead. Turns out, she automated the whole thing. She even scheduled it to send automatically. Tasks that used to take hours. Gathering 1000s of rows of data, cleaning it, running a regression analysis, time series, hypothesis testing etc… she now completes almost instantly. Everything. Even random things like finding discounts for her Pilates class. She just needs to check and make sure everything is good. She’s not super technical so I was surprised she could do these complicated workflows but the craziest part is that she just prompted the whole thing. She just types something like “compile a list of X, format it into a CSV, and run X analysis” or “go to Y, see what people are saying, give me background of the people saying Z” And it just works. She’s even joking about connecting it to the office printer. I’m genuinely baffled. The barrier to effort is gone. Now we’ve got a big market report due next week, and she told me she’s planning to use DeepResearch to handle it while she takes the week off. It’s honestly wild. I don’t think most people realize how doomed knowledge work is.