5 ms·
When ChatGPT first came out, I was surprised at the depth of information it had about region-specific market sizing in our (relatively niche) industry. Turns o
by gomox 3y ago
When ChatGPT first came out, I was surprised at the depth of information it had about region-specific market sizing in our (relatively niche) industry.
Turns out the model had ingested a pay-to-read article that cost upwards of $2000 [0] and was quoting the figures in it and referencing it directly (i.e. attributing the info to the article in question).
I knew about the paper but had never purchased it. I was actually surprised I could access the data through ChatGPT.
A few days later the same information was gone. I assumed someone might have decided to keep it on the low regarding having ingested all of these sources of restricted access content. I now get a generic blurb about the same question.
The opacity of it all made me a little worried. What similar models exist, trained on actually private information, for more nefarious purposes?
[0] https://www.marketresearch.com/Infiniti-Research-Limited-v2680/ITSM-Latin-America-30367235/ https://www.marketresearch.com/Infiniti-Research-Limited-v26...
- jameshart 3y agoThe IT Services Market is not 'relatively niche'.
- gomox 3y agoImagine how your comment would come off if your guess was wrong. That's how it comes off.
- astrange 3y ago> A few days later the same information was gone. I assumed someone might have decided to keep it on the low regarding having ingested all of these sources of restricted access content. I now get a generic blurb about the same question. This didn't happen. There is no way to remove information from the model and retraining it costs millions of dollars.
- moneywoes 3y agoHow did they patch all the illegal stuff ChatGPT said at release?
- segmondy 3y agoYou can imagine a chain of GPTs, good chatGPT monitors your instruction and refuses to pass it to regular chatGPT. if your prompt does get passed in and generates an offensive content, monitor chatGPT will monitor to see if output said something bad and then refuse to return the result and apologize.
- astrange 3y agoSmaller side model that just prints the "umm sorry I can't do that" message. And they did retrain it at least once - that's why it got so much faster and why the model chooser has "default" and "legacy" GPT-3.5. Notice that if you hit regenerate it often actually answers the question though.
- nonethewiser 3y ago> How did they patch all the illegal stuff ChatGPT said at release? What did it say that was illegal? And where was it illegal? That seems non sensical. What text can a chat AI return that breaks the law?
- ruszki 3y agoI have no clear idea about the legality of AIs, but it definitely incited violence in several cases. If a person do it, it’s against the law even in the US.
- noobermin 3y agoNot retraining, likely traditional filtering before it is fed to chatgpt.
- htag 3y ago> The opacity of it all made me a little worried. What similar models exist, trained on actually private information, for more nefarious purposes? What do you mean by 'nefarious purposes'? If you have illicit access to private data, chatGPT won't reveal anything that isn't already in the data. If your access to private information is legitimate, anything you do with an LLM trained with the information will have the same moral/legal consequences as just using the information directly.
- michaelbuckbee 3y agoOne of my friends was trying to figure out "how" ChatGPT knew about tried something I hadn't thought out, just ask it "do you know paper xyz by author abc?" and in fact it had.
- soco 3y agoJust saying, if ChatGPT says it knows something it doesn't necessarily mean it knows something, or that said thing is even real.
- Mizza 3y agoI was hoping this article would be about that. It's not. It's a waste of time spent pearl clutching about how the authors find some small percentage of the internet/Common Crawl "troubling." I don't know what they expected.
- PaulHoule 3y agoIf you want a model to, say, never draw pornography or write a Hitler speech, you don't do that by excluding "bad" content from the foundation model but rather you tackle that in later stages of training, particularly in this phase https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback https://en.wikipedia.org/wiki/Reinforcement_learning_from_hu... Having some of that in the foundation model actually helps the model "understand" what it is it is not supposed to do. The whole point of the foundation model is that it generalizes from the examples you show it later to similar examples it saw during "pre-training" and a certain about of offensive content will help it learn what it is offensive much quickly later.
- Sunhold 3y agoUnless you actually checked the article, ChatGPT likely hallucinated the figures and picked a plausible-sounding source. Most likely, nothing changed when you tried a few days later. The responses are stochastic due to the temperature setting. If anything did change, it was probably an update to reduce hallucination.
- gomox 3y agoI never saw the actual article, but the quoted figures were not just plausible and reasonable but also internally consistent throughout several exchanges. It produced exact figures for key (i.e. largest) submarkets and growth estimates for smaller ones, in the same way that I would expect the original content to be structured. I would bet with a 95%+ confidence that it parroted the actual contents. $2000 for that extra 5% tho.
- aaron695 3y agoI've never seen anyone link a ChatGPT response to a source document[1] Yet you are saying you've linked it to a source, even though you haven't seen it. [1] GPT-2 was sometimes possible. Anyone for ChatGPT?
- gomox 3y agoI don't know what you mean. ChatGPT's reply said something along the lines of "according to study ABC done in 2020 by research firm XYZ...". It's not some theory I concocted.
- ipaddr 3y agoThat might have come from someone else making it freely available
- RC_ITR 3y agoYeah it was probably just a random pdf that commoncrawl got to. Plenty of people post that kind of stuff on accidentally indexed sites. The risk though is even if it remembers the right format, there’s no guarantee it remembers the right numbers. So caveat emptor.
- BiteCode_dev 3y agoImagine how powerful an ai would be if we would give it access to all the info behind firewalls.
- propogandist 3y agothis is what Microsoft is trying with Office 365 GPT approach, where they are siphoning training data from every company and offering graphs and summaries in return
- oriettaxx 3y agoand I suspect one day gmail content will be used to :)
- tsukikage 3y agoWhile we're at it, imagine how awesome an ordinary search engine could be if we would give it access to all the info behind firewalls.