4 ms·
> Based on the findings, we advocate for new metrics and tools that evaluate not just final outputs but the structure of the reasoning process itself. Maybe th
by WA 1y ago
> Based on the findings, we advocate for new metrics and tools that evaluate not just final outputs but the structure of the reasoning process itself.
Maybe the problem is to call them reasoning in the first place. All they do is expand the user prompt into way bigger prompts that seem to perform better. Instead of reasoning, we should call this prompt smoothing or context smoothing so that it’s clear that this is not actual reasoning, just optimizing the prompt and expanding the context.
- airstrike 1y agoThe process of choosing which of the available tools to use is, to me, the only part of AI that I'm comfortable referring to as "reasoning" today.
- neumann 1y agoRaining is essentially doing as a 'built in' feature for what users found earlier that requesting longer contextual responses tend to arrive at a more specific conclusion. Or put it inversely, asking for 'just the answer' with no hidden 'reasoning' gives answers far more brittle.
- leptons 1y agoIt seems to be something like brute forcing. Throw even more words at the LLM until something useful pops out.
- ACCount37 1y agoIf you go out of your way to avoid anthropomorphizing LLMs? You are making a mistake at least 8 times of 10. LLMs are crammed full of copied human behaviors - and yet, somehow, people keep insisting that under no circumstances should we ever call them that! Just make up any other terms - other that the ones that fit, but are Reserved For Humans Only (The Kind Made Of Flesh). Nah. You should anthropomorphize LLMs more. They love that shit.
- nurettin 1y agoSo we should invite them to dinner? Watch movies together? Would they enjoy shopping?
- anal_reactor 1y agoI talk to AI more than I talk to my family
- dpassens 1y agoThen perhaps you should seek help.
- elcritch 1y agoWell that is going to be a thing soon enough. LLMs running on humanoid robots as AI partners are gonna become a thing one day.
- ACCount37 1y agoWould you enjoy their company?
- a96 1y agoAnd would they enjoy yours?
- astrange 1y agoI read something today about a discord that has Claude join their movie nights.
- selfhoster11 1y agoIf it's a smart enough LLM with a bearable personality distinct from a ChatGPT customer service assistant, then why not?
- lcnPylGDnU4H9OF 1y ago> Nah. You should anthropomorphize LLMs more. They love that shit. I'm reminded of something I read in a comment, paraphrasing: it makes sense to anthropomorphize something that loudly anthropomorphizes itself when someone so much as picks it up.
- brrrrrm 1y agowhat about "test time scaling"?
- aurareturn 1y agoI agree. When I first heard of the term "reasoning" to describe these models, I thought, "wait, I thought normal models also reason pretty well".
- aytigra 1y agoI feel like "intuition" really fits to what LLM does. From the input LLM intuitively produces some tokens/text. And "thinking" LLM essentially again just uses intuition on previously generated tokens which produces another text which may(or may not) be a better version.