4 ms·
Being able to see the thinking trace in R1 is so useful, as you can go back and see if it's getting stuck, making a wrong assumption, missing data, etc. To me t
by mechagodzilla 2y ago
Being able to see the thinking trace in R1 is so useful, as you can go back and see if it's getting stuck, making a wrong assumption, missing data, etc. To me that makes it materially more useful than the OpenAI reasoning models, which seem impressive, but are much harder to inspect/debug.
- thot_experiment 2y agoRunning it locally lets you INTERJECT IN IT'S THINKING IN REALTIME and I cannot stress enough how useful that is.
- amarcheschi 2y agoOh this is so cool
- Gooblebrai 2y agoYou mean it reacts to you writing something while it's thinking of that you can stop it while it's thinking?
- hmottestad 2y agoYou can stop it at any time, then modify what it's written so far...then press continue and let it continue thinking and answering.
- thot_experiment 2y agoFundamentally the UI is up to you, I have a "typing-pauses-inference-and-starts-gaslighting" feature in my homebrew frontend, but in OpenWebUI/Sillytavern you can just pause it and edit the chain of thought and then have it continue from the edit.
- Gracana 2y agoThat's a great idea. In your frontend, do you write in the same text entry field as the bot? I use oobabooga/text-generation-webui and I findit's a little awkward to edit the bot responses.
- thot_experiment 2y agoNo, but the chat divs are all contenteditable.
- Gracana 2y agoOh! That is an excellent solution. I wish it was that easy in every UI.
- thot_experiment 2y agoThanks, for what it's worth unless you particularly need to use exl2 ollama works great for local inference and you can prompt together a half decent chat UI for yourself in a matter of minutes these days which gives you full control over everything. I also lean a lot on https://www.npmjs.com/package/amallo https://www.npmjs.com/package/amallo which is a api wrapper i wrote for ollama which makes this sort of hacking very very easy. (not that the default lib is bad, i just didn't like the ergonomics)
- thenameless7741 2y agoInteresting.. In the official API [1], there's no way to prefill the reasoning_content: > Please note that if the reasoning_content field is included in the sequence of input messages, the API will return a 400 error. Therefore, you should remove the reasoning_content field from the API response before making the API request So the best I can do is pass the reasoning as part of the context (which means starting over from the beginning). [1] https://api-docs.deepseek.com/guides/reasoning_model https://api-docs.deepseek.com/guides/reasoning_model
- arresin 2y agoHow are you running it locally??
- thot_experiment 2y agoI am running a 4bit imatrix quant of the 70b distill with quantized context. It fits in the 43gb of vram I have.
- c-fe 2y agoI would actually love if it would just ask me simple questions (just yes/no) when its thinking about something i wasnt clear about and i could help it this way, its a bit sad seeing it write out the assumption and then take the wrong conclusion
- thot_experiment 2y agoYou can run it locally, pause it when it thinks wrong and correct it's chain of thought.
- c-fe 2y agoOh wow I did not know and dont have the hardware to run it locally unfortunately
- thot_experiment 2y agoYou probably have the hardware to run the smallest distill, it runs even on my ancient laptop. It's not very smart but it still does the CoT and you can have fun editing it.
- viraptor 2y agoYou can add that to the prompt. If you're running into those situation with vague assumption, ask it to provide either the answer or questions to provide any useful missing information.
- czk 2y agothe fact that openai hides the reasoning tokens from us to begin with shows that what they are doing behind the scenes isnt all that impressive, and likely easily cloned (r1) would be nice if they made them visible now
- orbital-decay 2y agoIt's almost like watching a stoned centipede having a panic attack about moving its legs. It also makes it obvious that these models (not just R1 I suppose) need to learn some kind of priority estimation to stop overthinking irrelevant issues and leave them to the normal token prediction, while focusing on the stuff that matters. Nevertheless, R1's reasoning chains are already shorter in tokens than o1's while having similar results, and apparently o3-mini's too.