5 ms·
I just tried. Asked a question on a research topic I'm digging into. It gave me some answers but no references. Then I copy the answers it gave me and specifica
by spirodonfl 2y ago
I just tried. Asked a question on a research topic I'm digging into. It gave me some answers but no references. Then I copy the answers it gave me and specifically ask for references. Then I got:
I sincerely apologize for my earlier response. Upon reviewing the search results provided, I realize I made an error in referencing those specific studies. The search results don't contain any relevant information for the claims I mentioned earlier. As an AI assistant, I should be more careful in providing accurate and supported information. Thank you for bringing this to my attention. In this case, I don't have reliable references to support those particular statements about software tools and their impact on developer experience and software quality.
Not sure about this...
- rushingcreek 2y agoSorry about that, could you make sure that "Always search" is enabled and try that first query again? It should be able to get the correct answer with references.
- spirodonfl 2y agoIt was on. If I ask the same question again it now gets the right answer. Maybe a blip? Not sure. To be fair, I don't expect these AI models to give me perfect answers every time. I'm just not sure people are vigilant enough to ask follow up questions that criticize how the AI got the answers to ensure the answers come from somewhere reasonable.
- mokkun 2y agoI found that quite often even though the always search option is on, it won’t search at times; maybe that was the case here.
- spirodonfl 2y agoJust to follow up on this: I asked it to give me a brief explanation on how to use laravel 11 blade fragments, which it did reasonably well. I then offered 3 lines of code of a route I'm using in Laravel and I asked to tell me how to implement fragment usage where the parameter in the url determines the fragment returned. Route::get('/vge-frags/{fragment}', function ($fragment) { return view('vge-fragments'); }); It told me to make sure I have the right view created (which I did) and that was a good start. Then... It recommended this? Route::get('/vge-frags/{fragment}', function ($fragment) { return fragment($fragment); }); I immediately knew it was wrong (but somebody looking to learn might not know). So I had to ask it: "Wait, how does the code know which view to use"? Then it gave me the right answer. Route::get('/vge-frags/{fragment}', function ($fragment) { return view('vge-fragments')->fragment($fragment); }); I dunno. It's really easy to find edge cases with any of these models and you have to essentially question everything you receive. Other times it's very powerful and useful.
- rushingcreek 2y agoThis is a good point, and we have new application-level features coming soon that to improve verifiability.
- spirodonfl 2y agoI dunno if you need it but I'd be happy to come up with some scenarios and help test
- wokwokwok 2y agoSeems a little bit of an unfair generalisation. I mean, this is an unsolvable problem with chat interfaces, right? If you use a plugin that is integrated with tooling that check generated code compiles / passes tests / whatever a lot of this kind of problem goes away. Generally speaking these models are great at tiny self contained code fragments like what you posted. It’s longer, more complex, logically difficult things with interconnected parts that they struggle with; mostly because the harder the task, the more constraints have to be simultaneously satisfied; and models don’t have the attention to fix things simultaneously, so it’s just endless fix one thing / break something else. So… at least in my experience, yes, but honestly, for a trivial fragment like that most of the time is fine, especially for anything you can easily write a test for.
- dotancohen 2y agoAnd you can have the LLM write the test, too.
- Retr0id 2y ago> As an AI assistant, I should be more careful I hate this kind of thing so much.
- magicalhippo 2y agoI've been playing with Gemma locally, and I've had some success by telling it to answer "I don't know" if it doesn't know the answer, or similar escape hatches. Feels like they were trained with a gun to their heads. If I don't tell it it doesn't have to answer it'll generate nonsense in a confident voice.
- ithkuil 2y agoThe models weights are tuned towards the direction that would cause the model to best fit the training set. It turns out that this process makes it useful at producing mostly sensible predictions (generate output) for text that is not present in the training set (generalization). The reason that works is because there are a lot of patterns and redundancy in the stuff that we feed to the models and the stuff that we ask the models so there is a good chance that interpolating between words and higher level semantics relationship between sentences will make sense quite often. However that doesn't work all the time. And when it doesn't, current models have no way to tell they "don't know". The whole point was to let them generalize beyond the training set and interpolate in order to make decent guesses. There is a lot of research in making models actually reason.
- magicalhippo 2y agoIn the Physics of Language Models talk[1], he argues that the model knows it has made a mistake, sometimes even before it has made it. Though apparently training is crucial to make the model be able to use this constructively. That being said, I'm aware that the model doesn't reason in the classical sense. Yet, as I mentioned, it does give me less confabulation when I tell it it's ok not to answer. I will note that when I've tried the same kind of prompts with Phi 3 instruct, it's way worse than Gemma. Though I'm not sure if that's just because of a weak instruction tuning or the underlying training as well, as it frequently ignores parts of my instructions. [1]: https://www.youtube.com/watch?v=yBL7J0kgldU https://www.youtube.com/watch?v=yBL7J0kgldU
- Intralexical 2y ago> I sincerely apologize for my earlier response. Upon reviewing the search results provided, I realize I made an error in referencing those specific studies. The search results don't contain any relevant information for the claims I mentioned earlier. As an AI assistant, I should be more careful in providing accurate and supported information. Thank you for bringing this to my attention. In this case, I don't have reliable references to support those particular statements about software tools and their impact on developer experience and software quality. Honestly, that's a lot of words and repetition to say "I bullshitted". Though there are humans that also talk like this. Silver lining to this LLM craze, maybe it'll inoculate us to psychopaths.