4 ms·
Did the model refuse to answer? Did it say that it doesn't know? If not, then it's a fair game in my opinion.
by Yajirobe 2mo ago
Did the model refuse to answer? Did it say that it doesn't know? If not, then it's a fair game in my opinion.
- dexterlagan 2mo agoLike most models it makes up things when it doesn't know. I had high hopes for Gemma4, which was said to having 'solved' this particular problem - it didn't. Gemma4 made less things up, did say it didn't know more, but it's still far from perfect. By comparison Qwen3.8 knows more, but still makes things up when it doesn't know. Coding abilities are very impressive however, and yeah these SVG tests do seem to 'scale' or generalize over its general reasoning+coding abilities - at least in JS and Rust. My next test will be to ask it to write some macros in Racket, just to see if it can balance parens. Most, if not all models cannot, no matter their size.
- croes 2mo agoModels don’t know that they don’t know.
- mdp2021 2mo ago> know that they don’t know And we are waiting for architectures that do - because it's duly.
- Zambyte 2mo agoHonestly it seems like a job for the harness, rather than the model. Sample the model with the same question, perhaps with varying temperature (?), and use that to establish a degree of confidence in the answer. If the model provides very different answers every time, respond that it doesn't know. If it responds with the same answer usually but a different answer sometimes, respond with moderate confidence. If the model always responds with the same answer, respond with certainty.
- jychang 2mo agothat could work sometimes but that's terribly hacky engineering
- Zambyte 2mo agoI... really don't agree. Being confident in using documentation to program in a variety of environmnets is not "terribly hacky engineering". Good engineering involves leveraging documentation well.
- mdp2021 2mo agoIf we want to implement Intelligence, and especially now that "the box is open" we must, we can play with the "intuitive" LLM architecture to understand it and squeeze it to its potential yeld, but at some stage we have to actually implement intelligence. That implies notions of confidence and a Foundational Theory of Knowledge (knowing why you know something), among the rest (one shot learning, update through reflection etc.).
- mbmbn 2mo ago[dead]
- ChucklsTheBeard 2mo agoIdeally it would run a web search to check questionable knowledge. Google-fu has been one of the most important SWE skills for a long time!