4 ms·
> Peer review is not perfect, and may not be tuned to catch LLM’s style of errors This summarizes I think a lot of the challenges with validating LLM output. W
by appplication 2mo ago
> Peer review is not perfect, and may not be tuned to catch LLM’s style of errors
This summarizes I think a lot of the challenges with validating LLM output. We hear “humans make mistakes too”, but I would agree with you that our human detection of human-made mistakes and LLM-made mistakes is unlikely to have the same coverage.
- airstrike 2mo agoThe real problem is that human mistakes generally occur more frequently given the difficulty of the task, whereas LLM mistakes are somewhat random, like the carwash problem, because LLMs cannot truly reason. I'd rather have a human on my team for whom I can reasonably surmise what tasks they're good at than have a robot who randomly gets shit wrong.
- yndoendo 2mo agoI worded a questions similar to this with a person in the finance industry. He was very receptive of AI since it is one of the easiest ways to grow stock portfolio. Say you have two assistant employees. One that reads the reports and one that uses AI to summarize the content. You are sitting in a high value meeting. You only can have one assistant with you in the meeting. Which one would you pick, the persons that read the reports or the person who only used AI?
- aleph_minus_one 2mo ago> Say you have two assistant employees. One that reads the reports and one that uses AI to summarize the content. You are sitting in a high value meeting. You only can have one assistant with you in the meeting. Which one would you pick, the persons that read the reports or the person who only used AI? I would take the person who only used AI. Why? Your fictional problem has the silent assumption that if you are in a high value meeting, you shall act more conservatively (i.e. reduce risk for reliability, even if you loose chances/opportunities). But this is rather a statement that the respective person acts on a risk-averse profile. I rather have a risk-affine profile. Because it is likely that everybody else in the meeting has a risk-averse profile, it is very likely that if I act on the "non-conservative" choice and I am right, I will have a huge advantage. On the other hand, if I act on the conservative choice, I won't have any advantage over the other participants. Since I am risk-affine person, I am willing to risk it. :-)
- yndoendo 2mo agoThere was no time frame from reading the report and AI summary to having to apply the report is stated. It could be one hour to 30 days in between or more. There was no statement if the meeting was scheduled or impromptu. Client(s) might have saw you in public or at an event. Or at a trade show and there the report only applies to the single client versus to many clients. There was no statement of report ownership. All parties might have access to it. There is a chance that AI lied / hallucinated the summary. Only to know would be taking time to read the report and validate it. AI only consumption has shown to reduce retention of information compared to reading the core content. Ever have someone tell you why they want AI? I have, it was because they do not want to read the reports.
- gus_massa 2mo ago> whereas LLM mistakes are somewhat random, like the carwash problem At least in math problems, AI mistakes generally occur more frequently given the difficulty of the task. The carwash problem is very undespecified. The question should be something like "I am American. I live in a single family home in a suburb. Each member of my family has their own car. All the cars are parked during the night at home. My boss don't authorize me to go during working hour outside the office. I want to wash my car. The car wash is 50 meters away. Should I walk or drive?" Disclaimer: I am Argentinean. I live in an old apartment building without a parking lot. I don't have a car, but if I had one I'd probably move to a building with parking inside or nearby. The closest professional car washing facility is like 1000 meters away, but the closest parking lot is like 50 meters away and they guys would wash the car for a few bucks. I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
- airstrike 2mo agoYou're picking on the specifics of the carwash problem rather than acknowledging countless other similar problems exist. The original one was strawberry, but that got RL'd away
- gus_massa 2mo agoThe strawberry problem is a translation/encoding problem, not a context problem. To be more efficient and cheap, the implementations use tokens, that is very-slightly-similar to translate it to Japanese: イチゴには「r」がいくつありますか? (I don't speak Japanese, I used autotranslation, so it may be wrong.) There are no "r" in the token of strawberry.
- airstrike 2mo agoAnd yet, if models could indeed reason, that would not be an issue. You're arguing my point.