6 ms·
It is if the weights are sufficiently advanced.
by chpatrick 1y ago
It is if the weights are sufficiently advanced.
- blueflow 1y agoI find such statements frightening. Too many people can not tell the different between prevalence ("everybody does it") and factually correct.
- chpatrick 1y agoNothing to do with dice though.
- deleted 1y ago[deleted]
- blueflow 1y agoThe whole "stochastic means to find factual correctness" thing is an error of method, arguing about weights here is nonsense.
- chpatrick 1y agoIt isn't though, the most factually correct human expert is also stochastic. The only question is how the dice are weighted.
- blueflow 1y ago"human expert" as reference for "factually correct", oh just gently caress yourself. Appeal to authority (expert = social status) is as much bullshit as appeal to popularity.
- chpatrick 1y agoRight now the fully deterministic always correct oracle machine doesn't exist. The most authorative answer we can get on a subject is from a respected human in their field (who is still stochastic). It's unrealistic to hold LLMs to a higher standard than that.
- deleted 1y ago[deleted]
- Zigurd 1y agoThe weights, so to speak, come from the knowledge base. That means you can't get away from the quality of the knowledge base. That isn't uniform across all domains of knowledge. Then the problem becomes how do you make the training material uniformly high-quality in every knowledge domain? At best it becomes the meta problem of determining the quality of knowledge in some way that makes an LLM able to calibrate confidence to a knowledge domain. But more likely we're stuck with the dubious quality that comes from human bias and wishful thinking in supposedly authoritative material.
- chpatrick 1y agoSure, it's only as good as the training data. But human experts also output tokens with some statistical distribution. That doesn't mean anything.
- Zigurd 1y agoThat sounds plausible. But it doesn't explain why LLM's make laughably bad errors that even a biased and haphazard human researcher wouldn't make.
- chpatrick 1y agoI think that's been a lot less true over the last year or so. Gemini 2.5 Pro is the first LLM I actually find pretty damn reliable.
- Zigurd 1y agoGemini seems to have a user interface that, for the way most people encounter Gemini, is more closely linked to search results. This leads me to suspect that Google's approach to training could be uniquely informed by both current and historic web crawling.
- contagiousflow 1y agoIf you think talking to an LLM is the same experience as talking to a human you should probably talk to more humans