5 ms·
what evidence would be enough for you besides source code ? The thing is returning only one correct answer to a question that days before had 3 answers.
by disiplus 5y ago
what evidence would be enough for you besides source code ?
The thing is returning only one correct answer to a question that days before had 3 answers.
- wallfacer120 5y agoIt improved, the model improved. Because sometimes, when you do work on the model, it improves.
- ClumsyPilot 5y agoExemplary tautology, explains nothing but fills the reader with confidence.
- wallfacer120 5y agoNo, it's not a tautology. A tautology is a statement in a form that must always be true, regardless of its constituent parts. "Well, <blank> could be true, or it could be false" is an example of such a statement in natural language. My statement was an example of believing a simple/common explanation over the rarely seen and complex one.
- ben_w 5y agoHow does it respond to similar questions? If conversational AI could be implemented just by getting users to type stuff, getting humans to respond the first time, and merely caching the responses in a bunch of if-elses for future users, even home computers would have reached this standard no later than when “CD-ROM drive” started to become a selling point.
- mannykannot 5y ago> How does it respond to similar questions? Well, one of the more interesting examples in the article is where Garry Smith took a question that had received only a vague equivocating answer, and repeated it the next day, this time getting a straightforward and correct answer. When he followed up with a very similar question on the same topic, however, GPT-3 reverted to replying with the same sort of vague boilerplate it had served up the day before. One would have to be quite determined to not find out, I think, if one was not curious about how that came about.
- danielmarkbruce 5y agoGuess: Some human saw the q & a. They realized it wasn't good. They uploaded some example of a phrase and the meaning, which would fix it. They were kinda lazy and just fixed that specific scenario.
- viceroyalbean 5y agoDoesn't that still call into question the quality of GPT-3? Surely such a large model should be able to extrapolate to "which is better: a or b?" from "is a better than b?" when only provided with the latter.
- galaxyLogic 5y agoThere is no definitive answer to this question, as it depends on the situation.
- ben_w 5y agoIt doesn’t call it into question for me, but perhaps I just had lower expectations to start with. I forget where I saw this comparison so I can’t link to it, but the last few years in AI are like waking up and finding dogs can talk: while some complain they’re not the worlds greatest orators, I find it amazing they can string a few genuinely coherent sentences together and maintain a contextual thread over multiple responses even half the time.
- mannykannot 5y agoPerhaps you are thinking of Scott Aaronson's "AlphaCode as a dog speaking mediocre English"? [1] I agree with the sentiment, but to continue your analogy, if OpenAI is using people to improve the answers to specific questions, it is a bit like learning that Cicero, Lincoln and Churchill were merely reading the work of speechwriters. There is an argument that it does not matter how GPT-3 gets to its answers - after all, for a long time, the main approach to AI was for people to write a lot of bespoke rules in an attempt to endow a computer with common sense and knowledge, so GPT-3 + instructGPT might be described as a hybrid of machine learning and the old approach. If OpenAI wishes to pursue that path, it is fine by me (as if my opinion matters!) but, because the perception of GPT-3 depends very strongly on how its output looks to human readers, it is obviously misleading if some of the most impressive replies were largely the result of specific human intervention. The issue is transparency: I would just like to know, when I read a reply, if this was the case, and it would not help OpenAI for it to ignore the call, in this article, for it to be clear about this. There is another argument that says that, given how GPT-3 works, it is unreasonable to expect it to give good answers in these cases - but that's the point! It looks really impressive when GPT-3 apparently does so, but not if they were effectively hard-coded. [1] https://scottaaronson.blog/?p=6288 https://scottaaronson.blog/?p=6288
- bmicraft 5y agoI'll just drop the fact here, that 15% of all google searches in 2017 were new and unique. https://blog.google/products/search/our-latest-quality-improvements-search/ https://blog.google/products/search/our-latest-quality-impro...
- danielmarkbruce 5y agoI run a few small chat bots. I can correct specific answers to questions (like the example given) by mapping certain phrases to certain intents, probabilistically. A new model is trained, and deployed. I do it all the time. It takes minutes. No source code changes. Their model is certainly bigger than mine and while I'm not certain about their process or tech stack, but I'd be willing to bet at even money that their's works vaguely similarly and that they have people looking at usage to see bad responses, updating data, re-running the model.
- staticassertion 5y agoIDK, something a lot more compelling than "the answers change over time" ? For a model that learns over time?