4 ms·
could you please tell me how it generates that certainty score?
by mylifeandtimes 5mo ago
could you please tell me how it generates that certainty score?
- colechristensen 5mo agoThe whole thing is a statistical model, that's just what it is. No, I cannot in a reasonable way dissect how an LLM works to a satisfactory level to a skeptic.
- skydhash 5mo agoIt's a statistical model for words and sentences, not knowledge. What does the LLM knows about having a pebble in your shoes, or drinking a nice cup of coffee?
- fc417fc802 5mo agoHe's not a skeptic, he's asking you to explicitly state your reasoning with the expectation that either the readers will learn something or (more likely) you will realize that your thought and speech pattern there was the equivalent of an LLM hallucinating. Yes you can prompt it as you suggested and yes you will generally receive a convincing answer but it is not doing what you seem to think it is doing ie the generated rating is complete bullshit that the model pulled out of its proverbial ass.
- colechristensen 5mo agoare you actually curious or do you just want to argue against it?
- clipsy 5mo ago"I can only explain my beliefs to people who promise they'll agree" is certainly a unique take.
- fc417fc802 5mo agoI think you're obviously wrong (based on my relatively detailed but certainly somewhat out of date and not expert level knowledge of LLM internals) but if you're willing to explain your reasoning I'm willing to reconsider my own position in light of any new information or novel observations you might provide.
- D-Machine 5mo agoGP is obviously wrong, and probably doesn't know about calibration and/or that it isn't even clear how to calibrate frontier models in the manner we need, given how complex and expensive the training is, and how tricky calibration becomes in e.g. mixture-of-experts and chain of thought approaches.
- mootothemax 5mo agoI suspect that introducing the calibration concept might be a case of too much too soon for some people. As far as I understand it, the various probability matrices boil down to: what token has the highest likelihood of coming next, given this set of input tokens. Which then all gets chucked away and rebuilt when the most likely token is appended to the input set. Objective assessment of internal state - again, to my non-expert eye - doesn’t appear to have any way to surface to me. Big-if my rough working understand is more or less correct - your calibration point makes a lot of sense to me. I’m not sure that it would make sense to someone who eg considers some form of active thinking process that is intellectualising about whether to output this or that token.
- adastra22 5mo agoVibes.