4 ms·
How do you notice hallucinations in a field you’re not familiar with? You may value focusing on different types of inputs or outputs than the model picker does
by hellisothers 1y ago
How do you notice hallucinations in a field you’re not familiar with? You may value focusing on different types of inputs or outputs than the model picker does and now you have no control.
We don’t know what we don’t know, we can’t always judge what is categorically right or wrong to make an informed decision. What we can do is decide who we want to ask a question based on competence.
- jstummbillig 1y agoWith 700 million(?) users, we have a lot of people familiar with every field. I am no biochemist, but if chatgpt starts spouting nonsense in that field biochemists will notice and speak up, and I will notice that they do. What's the idea? How does creeping, far reaching incompetence continually get past all of us?
- Topfi 1y agoEvery individual user would have to be consistently paying attention to discussions outside their expertise and interest. Considering prior stories of LLM usage among multiple legal professionals, wherein the model repeatedly output the potential for errors/“hallucinations”, I highly doubt that will happen. Heck, part of the outcry to reintroduce my personally least useful model 4o was grounded in a preference for subjective agreeableness in the output. The idea would/could be not intentional dissemination of missinformation, but purely financial. Models are expensive to run, hardware, rack space and power limited and making newer releases seem more robust subjectively can be a powerful incentive. With prior models we already have seen quantization post release and it’s been a personal pet peeve of mine that this should be communicated via a changelog, with the router there is one more quite powerful, potentially even less transparent way for providers to put their thumb on the scale. For now, GPT-5 does very impressively in my limited use cases and testing, especially considering pricing, but the concern that this may (and past experience tells me likely) change soon enough remains.
- jstummbillig 1y agoSide note, responding to AI written HN comments is something I will still have to get used to