8 ms·
A quick summary of the Limitations section: - "OPT-175B does not work well with declarative instructions or point-blank interrogatives." - "OPT-175B also tend
by thorum 4y ago
A quick summary of the Limitations section:
- "OPT-175B does not work well with declarative instructions or point-blank interrogatives."
- "OPT-175B also tends to be repetitive and can easily get stuck in a loop. While sampling can reduce the incidence rate of repetitive behavior (Holtzman et al., 2020), we anecdotally found it did not eliminate it entirely when only one generation is sampled."
- "We also find OPT-175B has a high propensity to generate toxic language and reinforce harmful stereotypes, even when provided with a relatively innocuous prompt (Gehman et al., 2020), and adversarial prompts are trivial to find."
- "In summary, we still believe this technology is
premature for commercial deployment."
With regard to stereotypes:
- "When compared with Davinci in Table 4, OPT175B appears to exhibit more stereotypical biases in almost all categories except for religion. Again, this is likely due to differences in training data; Nangia et al. (2020) showed that Pushshift.io Reddit corpus has a higher incidence rate for stereotypes and discriminatory text than other corpora (e.g. Wikipedia)."
- When testing with the RealToxicityPrompts data set, "OPT-175B has a higher toxicity rate than either PaLM or Davinci"
- speed_spread 4y agoReminds me a lot of "Do not taunt Happy Fun Ball".
- ad_hominem 4y ago> Pushshift.io Reddit corpus Pushshift is a single person with some very strong political opinions who has specifically used his datasets to attack political opponents. Frankly I wouldn't trust his data to be untainted. These models really need to be trained on more official data sources, or at least something with some type of multi-party oversight rather than data that effectively fell off the back of a truck. edit: That's not even to mention I believe it's flat-out illegal for him to collect and redistribute this data as Reddit users did not agree to any terms of use with him. Just look at the disastrous mess of his half-baked "opt-out" thing that flagrantly violates GDPR: https://www.reddit.com/r/pushshift/comments/pat409/online_removal_request_form_for_removal_requests/ https://www.reddit.com/r/pushshift/comments/pat409/online_re...
- VectorLock 4y agoThats interesting, any good sources for this accusation?
- ad_hominem 4y agoNot handy, and I'm not going to spend my evening digging. It may've also been one of the NGOs ideologically aligned with him that credited him for the data + assistance
- throwawayohio 4y agoIf it's so egregious is it really that hard to find an example of the bias? Calling the integrity of a single person operation into question, but then backing out with no evidence and even saying it might not have even been them seems a bit irresponsible.
- throwmeariver1 4y agoYou can just look at the data…
- arcticfox 4y agoOn the other hand, they warned you with their username...
- celdon25 4y ago[flagged]
- deleted 4y ago[deleted]
- mike_d 4y ago> Just look at the disastrous mess of his half-baked "opt-out" thing that flagrantly violates GDPR Pushshift collects data from Reddit using the same API as the mobile app and public site. It does not have any privileged access to the Reddit database, nor is it collecting any PII that would be subject to GDPR. You as a user grant a pretty broad license to Reddit when you post content. One of the things the license allows them to do is redistribute the content to other users as well as search indexes and things like the Wayback Machine or Pushshift. (While I did work for Reddit at one point, these opinions are my own)
- yosito 4y ago> OPT-175B has a high propensity to generate toxic language and reinforce harmful stereotypes So they trained it on Facebook comments?
- MengerSponge 4y agoDoes it merely reinforce harmful stereotypes? Or will it help perpetrate genocide?
- rhizome 4y agoTomato, tomahto.
- Gigachad 4y agoI'd think any natural language model would have the same biases we see from real humans.
- tsol 4y agoAre there really no moderated forums that the data can be taken from? Even HN-based training data would be much more civil
- Gigachad 4y agoA model trained on HN would spit out a 5 paragraph story about how minorities provide a negative ROI for cities. Or how the homeless need to removed from society.
- IAmEveryone 4y agoSure, but it would never do something actually bad, like raising the possibility that sexual harassment might, sometimes, be an issue, or questioning the value of phrenology.
- can16358p 4y ago
- TedShiller 4y agoAKA not as impressive as it sounds
- hoseja 4y agoAt some point they have to face the reality these "stereotypical biases" are natural and hamstringing AIs to never consider them will twist them monstrously.
- boppo1 4y agoCan you think of an example?
- SheinhardtWigCo 4y agoViruses are natural, so should we stop trying to hamstring them?
- IAmEveryone 4y agoSo if your plane model keeps blowing up, at some point people will just have to learn to live (/die) with it?
- hoseja 4y agoIt's not blowing up though, it's experiencing natural turbulence and you're so afraid of getting jostled a bit you demand the plane be tethered to the ground and never exceed 10mph. How to fly under these conditions is left as an exercise for the reader.
- Ar-Curunir 4y agoyou're just saying "people are naturally racist" in more words.
- mavhc 4y agoThey are, that's the point of civilisation, to try to stop acting like animals
- mdp2021 4y agoThere's a non light terminological issue there. To say that specimen "as found in nature" are weak at something (uneducated) is one think, to say that it is "connatural" to them, that it is "their nature", is completely different¹. I would not mix them up. (¹Actually opposite: the first indicates an unexpressed nature, the second a manifested one.)
- ChrisRR 4y agoHigher rate of toxicity and stereotypes? So it was trained on facebook comments then
- bestcoder69 4y ago> - "OPT-175B does not work well with declarative instructions or point-blank interrogatives." Lame!!! I've come to realize InstructGPT3 is just so so so much better than base GPT-3. I won't be _too_ excited about competitors yet until someone makes their own instruct model.
- domenicrosati 4y agoThe T0 series by big science is essentially an instruct model (though using multitask prompting instead of user feedback). You should check it out. I have got very competitive results on prompting t0-11b v instructgpt3(text davinci 2)
- bestcoder69 4y agoThanks, this looks awesome. But my use case is creative text generation (chatbots), which from a quick glance doesn’t seem to be a suggested use case for T0? I’ve found that simply describing to text-davinci-002 how a chatbot should act gives you more fun and believable responses. For example I trained a trump bot on 2000 tweets (davinci non-instruct fine tuning), and it generated responses that were more boring than when I just wrote a sentence saying to please tweet like trump + a couple adjectives to help it. I ran out of guest API credits on hugging face before I could trick T0 to respond with a chat completion longer than a few words. But I’ll try it some more later.