4 ms·
The streetlight effect: > A policeman sees a drunk man searching for something under a streetlight and asks what the drunk has lost. He says he lost his keys a
by post-it 4mo ago
The streetlight effect:
> A policeman sees a drunk man searching for something under a streetlight and asks what the drunk has lost. He says he lost his keys and they both look under the streetlight together. After a few minutes the policeman asks if he is sure he lost them here, and the drunk replies, no, and that he lost them in the park. The policeman asks why he is searching here, and the drunk replies, "this is where the light is"
All of your suggestions are better but they're hard, so someone casually evaluating an AI isn't going to do them.
- newaccountman2 4mo ago[flagged]
- blanched 4mo agoThis kind of hamfisted snark tends to make people take the actual and justified criticism of police less seriously.
- newaccountman2 4mo agoIf people were willing to take it seriously in the first place, then they wouldn't view it as "hamfisted snark"
- blanched 4mo agoI consider myself someone who takes it seriously, and have spent time and resources fighting for change. But it’s wholly unrelated to this particular thread, phenomenon, and story. So having a little “ha ha” moment accomplishes nothing towards the actual cause. It makes people uncomfortable, but not the useful kind of uncomfortable. That said, maybe we just disagree on how to drive change, and that’s fine. I’ll leave it.
- post-it 4mo agoIt could be a taxi driver if you like. Or an anarchist passing by on xir way to a protest.
- layer8 4mo ago…in the US.
- sanderjd 4mo agoSure, for casual evaluation, I agree. But are there serious analyses that are evaluating this kind of thing? I mean, these are the kinds of things I evaluate in my own work when a new model comes out, or when I'm evaluating a harness. But this is all very ad hoc and intuitional. I'd love to start bringing rigor to it, but I haven't found much prior art on this. In another thread someone said that's because it's probably impossible to do this rigorously because too much of it is subjective. And that does match my intuition. But I continue to suspect that intuition is wrong.
- jerf 4mo agoIt's hard to bring much rigor to it. I'm not saying impossible, but it's not like it's completely obvious how to do it and people are just too lazy. Intrinsically, if I'm going to test a back-and-forth with a model I have a human in the loop making frequent decisions. Did the model fail or succeed at whatever rate it did that because of the model or the human? Did the testing protocols capture the actual problem, e.g., maybe if the model was given some particular bit of information that a normal human would have given it it would have done much better or worse, but the testing protocol in the interests of "rigor" excluded the human in the loop from doing it. Is the human going to be willing to sit down and do the same task 25 times, refreshing the model from scratch each time for a "valid" test? Can you get the same human to analyze every model in the test? Is their 10th pass of the problem an invalid test because you can't as easily erase the human's knowledge of the previous 9 tests? What do you do with a model that succeeds wildly 75% of the time and spins off into a loop the other 25%? Is that loop real or, again, did your "rigorous" testing protocol prevent the human from saving the model from the loop like any developer would? And so on and so forth. Again, I'm not saying this is impossible but I am saying that if you tried to do it, and you got the money, and you built the test, and got the human subjects clearance, and you ignored that during the process of all that at least one more frontier model would come out, you can count on HN anklebiting your "rigorous" study even so, and probably being correct about a lot of the issues it could have because it would take several iterations of this to build a reasonable protocol... at which point it would quite possibly also be obsoleted by progress again.
- 4mo ago
- redsocksfan45 4mo ago[dead]
- echelon 4mo agoThe minute an open model breaks through and beats Claude Opus/Fable, it's over. There are far more opportunities that can be served when the world's intellectuals have the raw weights and can fine tune, splice, distill, and reapply. Imagine having raw unfettered access to Fable. It can be refit to structural biology. It can be fine tuned on the repo for smaller context requirements. It can be run cheaper and air gapped. The world wants this.
- barrenko 4mo agoAs crazy as this sounds, and as much I don't want to believe it myself, I think we're still underestimating LLMs, and we're gonna get to that point pretty soon.
- jupr 4mo agoThe world does want this. Opus capabilities, in a box, securely tunneled to my family and I utilizing the resources I already have available to me which is, energy + network.
- digitaltrees 4mo agoI don’t think we need them. I think the models we have are good enough. It’s the orchestration layer that makes the biggest difference at this point. The open source models we have are capable of calling tools and the work is getting them to be capable enough to know which tools to call and what to do in response. I think we are leaving the main frame era of AI and entering the PC era already. If there wasn’t a RAM shortage and we all had 2TB of ram and GPUs we would all have large local models or personal APIs serving our teams. That’s why all the labs are moving to the App layer and moving away from being the API for intelligence like they were originally.
- wahnfrieden 4mo agoThey are absolutely not good enough
- digitaltrees 4mo ago