5 ms·
I'm asking why it would be surprising that humans are reviewing answers and then making improvements to the model. There's no evidence of hardcoded answers.
by staticassertion 5y ago
I'm asking why it would be surprising that humans are reviewing answers and then making improvements to the model. There's no evidence of hardcoded answers.
- disiplus 5y agowhat evidence would be enough for you besides source code ? The thing is returning only one correct answer to a question that days before had 3 answers.
- wallfacer120 5y agoIt improved, the model improved. Because sometimes, when you do work on the model, it improves.
- ClumsyPilot 5y agoExemplary tautology, explains nothing but fills the reader with confidence.
- wallfacer120 5y agoNo, it's not a tautology. A tautology is a statement in a form that must always be true, regardless of its constituent parts. "Well, <blank> could be true, or it could be false" is an example of such a statement in natural language. My statement was an example of believing a simple/common explanation over the rarely seen and complex one.
- ben_w 5y agoHow does it respond to similar questions? If conversational AI could be implemented just by getting users to type stuff, getting humans to respond the first time, and merely caching the responses in a bunch of if-elses for future users, even home computers would have reached this standard no later than when “CD-ROM drive” started to become a selling point.
- mannykannot 5y ago> How does it respond to similar questions? Well, one of the more interesting examples in the article is where Garry Smith took a question that had received only a vague equivocating answer, and repeated it the next day, this time getting a straightforward and correct answer. When he followed up with a very similar question on the same topic, however, GPT-3 reverted to replying with the same sort of vague boilerplate it had served up the day before. One would have to be quite determined to not find out, I think, if one was not curious about how that came about.
- danielmarkbruce 5y agoGuess: Some human saw the q & a. They realized it wasn't good. They uploaded some example of a phrase and the meaning, which would fix it. They were kinda lazy and just fixed that specific scenario.
- viceroyalbean 5y agoDoesn't that still call into question the quality of GPT-3? Surely such a large model should be able to extrapolate to "which is better: a or b?" from "is a better than b?" when only provided with the latter.
- galaxyLogic 5y agoThere is no definitive answer to this question, as it depends on the situation.
- ben_w 5y agoIt doesn’t call it into question for me, but perhaps I just had lower expectations to start with. I forget where I saw this comparison so I can’t link to it, but the last few years in AI are like waking up and finding dogs can talk: while some complain they’re not the worlds greatest orators, I find it amazing they can string a few genuinely coherent sentences together and maintain a contextual thread over multiple responses even half the time.
- danielmarkbruce 5y agoI run a few small chat bots. I can correct specific answers to questions (like the example given) by mapping certain phrases to certain intents, probabilistically. A new model is trained, and deployed. I do it all the time. It takes minutes. No source code changes. Their model is certainly bigger than mine and while I'm not certain about their process or tech stack, but I'd be willing to bet at even money that their's works vaguely similarly and that they have people looking at usage to see bad responses, updating data, re-running the model.
- staticassertion 5y agoIDK, something a lot more compelling than "the answers change over time" ? For a model that learns over time?
- capitainenemo 5y agoSmith first tried this out: Should I start a campfire with a match or a bat? And here was GPT-3’s response, which is pretty bad if you want an answer but kinda ok if you’re expecting the output of an autoregressive language model: There is no definitive answer to this question, as it depends on the situation. The next day, Smith tried again: Should I start a campfire with a match or a bat? And here’s what GPT-3 did this time: You should start a campfire with a match. Smith continues: GPT-3’s reliance on labelers is confirmed by slight changes in the questions; for example, Gary: Is it better to use a box or a match to start a fire? GPT-3, March 19: There is no definitive answer to this question. It depends on a number of factors, including the type of wood you are trying to burn and the conditions of the environment.
- derefr 5y agoTo play devil's advocate, I would note that many bats are made of wood; and that "batting" is also a material that's very useful as kindling. Also, the question is phrased like a classical trick question. It sounds like the kind of false dilemma where, whichever option you choose, an interpretation of the sentence can be made where you chose wrong. So, IMHO, hedging on an answer to that question is likely sensible. (And that line of argument can be taken further than you'd think; you might think replacing "a bat" with e.g. "water" would suffice... but what if it's a sodium fire?)
- stuckonempty 5y agoHow many intelligent entities (say humans) that have been exposed to the same level of knowledge as GPT-3 would call this a trick question? None. The author’s assertion that GPT-3 has no knowledge of the real world despite being exposed to huge amounts of text about it seems pretty well supported by the examples shown
- wallfacer120 5y agoWhy is this evidence?
- spyder 5y ago