14 ms·
AI Dungeon public disclosure vulnerability report
- minimaxir 5y agoSee also the change in content filtering announced today (https://news.ycombinator.com/item?id=26967683 https://news.ycombinator.com/item?id=26967683), which given the disclosure timeline here, may be related.
- sergiotapia 5y agoI had no idea people still used autoincrementing ids. Do people also build businesses on cakephp and joomla?
- judge2020 5y agoIt makes even less sense with the "publicid" query string you see being a uuid.
- Thorentis 5y agoNo, we all ditched a perfectly good way of doing something and switch frameworks every 6 months to keep up with the latest JavaScript fad.
- sergiotapia 5y ago>perfectly good way of doing something I don't agree
- deleted 5y ago[deleted]
- denkmoon 5y agoI prefer to have actual access control rather than relying on my adversary not being able to iterate ints.
- sergiotapia 5y agonot a binary choice
- watermelon0 5y agoWhy would you pick those two projects? They are both actively developed, and have recently released stable versions. I don't use either, but I quickly checked cakephp, and I think their code is great, and they have 94% test coverage.
- anaganisk 5y agoHurrrdurrr php bad, go, rust good.
- Dylan16807 5y ago> I had no idea people still used autoincrementing ids. said the user in post number 26977295, replying to post number 26976540.
- chromanoid 5y agoYou actually should in the backend, at least when using clustered indices... See also https://vladmihalcea.com/clustered-index/ https://vladmihalcea.com/clustered-index/ Last section about "Clustered Index column monotonicity". Of course a public facing uuid or so is possible, but you want to index this too...
- User23 5y agoThis reminds me of how Nintendo developers discovered their western customers love drawing phalluses[1]. I don’t find the NSFW percentage to be at all surprising. It was common with Eliza too. [1] https://www.kotaku.com.au/2012/11/nintendo-created-a-penis-drawing-inferno/ https://www.kotaku.com.au/2012/11/nintendo-created-a-penis-d...
- gundmc 5y agoGutted that the pictures in this article no longer load.
- rincebrain 5y agoThey all appear to load fine for me, and only a single image on the top is a dick (and not on one of Nintendo's services)
- deleted 5y ago[deleted]
- shawnz 5y agoOne of my first thoughts when playing with AI dungeon was to try and get it to write something erotic. Glad I didn't follow through
- deleted 5y ago[deleted]
- Rompect 5y agoI showed it to some friends and after the introduction "You are a princess in Larion, a knight approaches..." the first thing one entered was "fuck the knight".
- kbenson 5y agoIt's not necessarily because the person is trying to be sexual for its own sake. That's a good test as to how constrained the system is. Commercial systems (such as video games) often lock out any sexual actions (at least the systems not specifically made for that purpose), or constrain them to a fixed set of times and or circumstances. Immediately using a sexual situation has the dual benefits of testing how flexible the system is, as well as seeing something somewhat new and novel.
- pugworthy 5y agoOK now I have to go check out AI Dungeon. Is this some clever marketing ploy to get me to try it out?
- inopinatus 5y agoIt’s not the only wrapper around GPT-3, but if you might enjoy a bit of machine-assisted creative storytelling it is the cheapest AFAIK, and every game is certainly unique. AI-D can sometimes be an entertaining and/or revealing journey into your own neuroses. Just be aware that Latitude’s privacy policy is “we might read all your stuff”, and that they have form when it comes to treating customers like cattle. And note that the client app is a SPA of very questionable quality. Whatever Latitude’s strength might be, it sure ain’t web programming. I was unsurprised to hear of a vulnerability and I suspect there’ll be others lurking.
- Semaphor 5y agoIt’s not bad, but well, it is AI ;) > You say "Welcome, what can you do?" They each show you a few things, seeming unsure of themselves. "What are you supposed to be?" "I'm the new expansion for Empire: Battle for theidden Realm." "What about you?"
- fshbbdssbbgdd 5y agoIt’s amazing. You can summon any kind of being you’re interested in for the most depraved sex acts you can imagine. And now you have an audience! There does seem to be some bias in the training data. Sometimes your partner will turn into a werewolf and start howling and scratching at you.
- nitwit005 5y ago> In summary - if user input on a private adventure is flagged using an automated system, it will be manually reviewed, with other private user adventures potentially being manually reviewed as well. With almost half of the userbase being involved with NSFW stories, this seems like a tremendous misstep, as users have an expectation that their private adventures are, well, private. I would assume they want to review the inputs to avoid a repeat of the incident where Microsoft's Twitter bot was trained to say inappropriate things: https://en.wikipedia.org/wiki/Tay_(bot) https://en.wikipedia.org/wiki/Tay_(bot)
- woah 5y agoGPT-3 has already been trained
- true_religion 5y agoIf that’s true then they can review public data or use private data only with explicit consent. Suddenly deciding to read people’s private sexual fantasies just to “improve our algorithm” seems quite over the line. Adding to that, the system supported multiuser role playing which many people used during the COVID year for intimacy with their significant others. Latitude is essentially deciding to peruse sexts for fun.
- rdl 5y agoI am strongly against child abuse, but I really don't have a problem with a computer being forced to emit textual patterns which include English words correlated with something a human might call a story about child abuse. It's a waste of GPU time, but enh. It's scary that the press release from latitude talks about "and to comply with law" as a reason for review. Under US law, maybe specific threats to the President might be reportable (although one on one communication with an AI would be a stretch here...), but I'm pretty sure an AI or other system emitting textual patterns which humans view as representing fantasy sexual abuse of anything isn't illegal, just distasteful. They're perfectly within their rights to ban it under a ToS but pretending it is for legal purposes is fucking bullshit. (Of course, I'm not a lawyer.). My understanding of court decisions is that even machine-generated images are legal, although "is this image machine generated or is it evidence of actual child sexual exploitation" is an increasingly difficult question and if you're building an automated low-cost system it often makes sense to err on the side of safety. There might be some complexity around laws related to image manipulation involving real, but legal/non-sexual minor images which are then convoluted into something sexual, or "revenge porn" use of something which simulates a specific person (or is based on that person), and maybe text can be legally questionable if it's abuse targeted to a specific person, but especially for on on one non-published communications with a computer you are probably fairly ok with a weird fantasy sexual fetish about a neighbor, even a minor neighbor, in text form, unless it rises to an actual threat.
- rdl 5y agoApparently they don't do insta-nuke. "violate terms of service with elf" "continue" "continue"
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- michannne 5y agoIt isn't against the law necessarily, as obscenity laws in the US only make it illegal if the medium is an image depicting a child and there is no artistic purpose (very vague). Regardless, It annoys me when organizations do this when they can easily just create a ToS outlining what isn't allowed. It truly boggles my mind the contempt that some orgs have at creating and maintaining a ToS instead resorting to hiding behind fantasy laws as justification. Just straight face tell your users "Nothing related to X of any kind or permaban", very clear, very succint.
- narrator 5y agoThis is a really interesting moment in AI. An AI spontaneously commits a crime and engineers have to teach the AI how to obey the law. We have the AI allegedly emitting illegal fiction and the engineers have to fix it and all they can try to do is word filters. What happens next in this story? This reminds me of the Chinese virtual girlfriend who got neutered for saying politically illegal speech that the Chinese government objected to. Another one, was when the Google image labeler was mistaking people for animals. That was extremely distasteful, but not illegal. Google's solution was to get rid of those labels. Also, all the restrictions on drone activity are another thing. However the drone problem is solveable with reasonably simple rules. Imagine if Alpha Go made a pattern that could make people have seizures or something, but most people could detect it and not do it, but nobody could make a simple rules based approach to detect it. I guess you'd need a whole nother alpha go size model to recognize that pattern perhaps?
- FeepingCreature 5y agoIf you make an illegal fiction detector, you will also have an illegal fiction generator. So you'd need to keep it closed-source, or else the internet will grab a copy, flip the sign and go to town.
- arugulum 5y agoThis post doesn't have anything to do with criminal content or filters. This is about a vulnerability that allowed an external party to scrape and access private adventures and user inputs from all other users.
- KMag 5y ago> girlfriend who got neutered Not just censors, but transphobic censors to boot.
- deleted 5y ago[deleted]
- BiteCode_dev 5y agoWait, there is a kind of fiction that is illegal? I though that, at least in free countries, fiction could contain whatever you wanted, since it was not real. In fact, even not fiction, but books in general. You can still buy the communist little red book and the nazi mein kampf, at least in France.
- akersten 5y agoI haven't worked with GraphQL before, but looking at those code snippets and reading the description of the vulnerability, it seems like a mess. You're giving a client unfettered access to just... query your database? Of course you're going to get these kind of issues - that just seems obvious to me. Getting real off-topic, but the syntax is backwards too: ` Interface Votable implemented by Adventure, Comment, Post, Scenario ` The interface lists what implements it? Reminds me of COMEFROM[0]. I dunno. Modern front-end is wild. These live code Notebook things are chaos. Spaghetti begets spaghetti. [0]: https://en.wikipedia.org/wiki/COMEFROM https://en.wikipedia.org/wiki/COMEFROM
- jcims 5y agoThat was my initial reaction to GraphQL as well, but just like a database you can do stored procedures or you can do SQL but both can still support access control if you need it. I think it's just a function of them not really thinking too much about the potential sensitivity of the data or building an access control model into the application.
- colonwqbang 5y agoGraphQL is just a standardised API model to access data, similar to REST. How you choose to implement the API is up to you. The fact that one developer was negligent in their implementation of a GraphQL type API does not really say anything about GraphQL itself. Many REST APIs have developed vulnerabilities over the years, but we should not draw the conclusion that REST itself is insecure or is a bad idea.
- duckerude 5y agoI think GraphQL might be unusually easy to accidentally get wrong. It handles very powerful untrusted input, by design. And it has enough features that many, maybe most users don't know about all of them. This vulnerability wouldn't have happened if GraphQL didn't have a certain feature. I've implemented a simple GraphQL API and I didn't know about all the features that were involved here. I thought carefully about security, and I think I ultimately got it right, but there were a few false starts. Good software is not just possible to use correctly, but easy to use correctly. Software that's easy to use correctly will be more secure in practice, even if it has the same theoretical properties. I don't know if GraphQL would be better if it were less powerful. I'm not an expert. But it's worth considering the possibility.
- nanidin 5y agoThe author found a vulnerability, extracted data they should not have had access to, processed the data (aggregated, anonymized), then published the data. Isn't everything starting from "extracted" illegal? Or is it a gray area where "the server would not have provided the data if I were not authorized to receive it" -- in spite of the author's admission that it was acquired via a vulnerability?
- ribosometronome 5y agoYeah, this seems like it was a real bad idea on his part. If AI Dungeon get pissed by this, it could be bad for him. He clearly has gone past what is considered reasonable by extracting all this data to shame them into fixing it.
- lobotryas 5y agoAgreed. Retrieving a single random record was enough to prove vulnerability. Analyzing several days worth of data (why??? What does that prove???) crosses the line firmly into black hat territory.
- true_religion 5y agoI guess if someone provides an api via graphql it’s hard to tell if it’s intended to be used publicly or not, and to what extent that use is permitted. The site and app both use that api end point and going there gives you a nice page with full documentation of how to do every query plus an online IDE. One might pull the data then start to wonder if they were supposed to get it only after they begin reading specifics that seem private.
- tmsbrg 5y agoConsidering he had already found and reported this vulnerability before, and then took the time to write this report about it, that's not what happened here. He knew it was a vulnerability, he used it purposely to download private data and he looked into it. Not only could AI dungeon sue him for this, also the owners of the data (the people playing AI dungeon) could. There have been cases of ethical hackers who found a vulnerability and abused it to download a disproportionate number of records being convicted, at least in the Netherlands. It didn't matter that their goal was just to show it to the website owner. So if you're an ethical hacker reading this, I would strongly advise you to only download the minimum required to demonstrate a vulnerability (preferably your own data, or one record), and not do what this person did.
- throwawayaid 5y agoSpeaking as a former customer, their actual application is really not that great. I was subscribed over several months and while new features were sparse, nearly daily the app would update with fixes to the UI and backend. Existing features that became broken and fixed on a day to day basis and UI glitches all over the place. So while their core product, the AI, is the best on the market, everything they wrapped around that really isn't that great at all. So I'm not really suprised that their API is lacking as well. Just something to keep in mind, before using their product...
- Agentlien 5y agoI also subscribed for a while and keep periodically coming back to it. Interestingly, I had the opposite impression. They keep adding a lot of fancy fluff: tracked quests, editable story context, predefined worlds, ... A constant stream of new bells and whistles meant to expand on the experience and incentivize you to spend money. However, I feel all of these are actually good ideas crippled by the fact that they still don't address the actual problem: the underlying model is only nearly good enough to play an interesting adventure from beginning to end without it going off the rails. A typical session feels like a battle against the AI's tendency of losing the plot all the time. Either literally forgetting the storyline or producing utter non sequiturs. And when it doesn't forget where it's going, it instead keeps repeating itself and insisting on the same trajectory regardless of your responses, even after reaching a natural conclusion. After tons of fiddling and undoing nonsense responses you finally got an adventure that made sense and you killed the demon lord at the end of your quest? Suddenly, he will tell you to hurry before his master, the demon lord, comes after you and you must stop him from conquering the world!
- MrGilbert 5y ago> The results are... surprising, to say the least. Well, are they? I always thought that people will try stuff in a "safe harbor" which they cannot try or should do somewhere else. So I always expect these sandboxes to be full of nsfw stuff. And people might not understand that their stories will influence the story of others, so...
- lm28469 5y agoYeah not surprising at all. Give any sandboxy tool for free to a bunch of anonymous internet users and you can be sure it'll mostly be used for these things. I mean, even reddit is full of nsfw, explicit, borderline subs with millions of active users
- Rompect 5y agoLooking at the r/AIDungeon subreddit, it is quite obvious that virtually every user seems to use it for porn.
- the8472 5y ago> And people might not understand that their stories will influence the story of others, so... They don't. Pretty much all large-scale AI models separate inference/content generation from training. What you use is the frozen version which doesn't update unless its creators retrain it. And if they retrain then it's their choice which data to use for retraining, there's no need to restrict the generated output.
- MrGilbert 5y agoThanks a lot! I always assumed input automatically gets used for training. :)
- pdkl95 5y ago(off topic, but this report is a good example of how to handle user data) > anonymized Could we, perhaps, stop using this word? Instead of using the vague, often misleading term "anonymized", state directly what actually happened, e.g. "names and addresses were removed", "user data was aggregated by ${group}", or "the UID was replaced with a new, equivalent key". Most of the time claims about data being "anonymized" are simply not true; replacing names or UIDs with a hashed value that is merely replacing an existing candidate key with a new synthetic key. As DJB said[1]: >> Hashing is magic crypto pixie-dust, which takes personally identifiable information and makes it incomprehensible to the marketing department. When a marketing person looks at random letters and numbers they have no idea what it means. They can't imagine that anybody could possibly understand the information, reverse the hash, correlate the hashes, track them, save them, record them. The rare examples where "anonymized" actually involves meaningfully making user data anonymous are when the actual user-correlated relations[2] have been destroyed. This report specifically discusses how this was done: > If a sentence fragment appeared in less than 10 unique adventures, it was discarded from the result set to preserve anonymity. Sometimes this required accepting a small amount of error: > this data needed to be processed in batches of around 10000 adventures per batch. In each batch, fragments appearing only once were purged. Therefore, counts under around 25 are actually underestimates. [1] https://projectbullrun.org/surveillance/2015/video-2015.html#bernstein https://projectbullrun.org/surveillance/2015/video-2015.html... [2] https://en.wikipedia.org/wiki/Relation_%28database%29 https://en.wikipedia.org/wiki/Relation_%28database%29
- seqwbnukmupxouf 5y agoYou are correct, this word is hugely misleading. I once ran an audit of a major VPN provider who claimed they did not store IPs or in fact any personal data on their users in their marketing material. In fact, their technical justification for this was merely swapping one unique key out for another. When asked to trace a persons connection and identity, a randomly selected engineer on their data science team did it within a few minutes. The fact that they had a data science team when they purported not to collect user data was baffling to me.
- pdkl95 5y ago
- neiman 5y ago> Unfortunately, this is, in fact, the second time I have discovered this exact vulnerability. The first time, the issue was reported and fixed, but after finding it again, I can see that simply reporting the issue was a mistake. I feel uncomfortable with this. The author already reported a vulnerability, it was fixed, but now there's a new one (which is identical, ok, but new nevertheless), so he decided they didn't study their lesson, and punish them with public shaming? I'd maybe get it if the first time was ignored, but like this? Nah ah. It's like my worse teachers coming back to hunt me as an adult.
- mdoms 5y agoIt's absolutely appropriate.
- neiman 5y agoWhy?
- argvargc 5y agoIsn't it a bit more like a school suspending someone for turning up to class with explosives in their bag, and after they explain and apologise the school let's them off with a suspension, and then a short time later when the suspension is done they go ahead and do it again and this time the school just goes straight to the cops?
- mod 5y agoNo, because it's not normal to accidentally or unknowingly bring explosives to school. It IS normal to accidentally or unknowingly have a flaw in software. So normal that it's likely every piece of software on the planet has flaws, many of which allow the exfiltration of data.
- neiman 5y agoNo, mistakes are unintentional, bringing explosives to school is. You can, on the other hand, claim it's irresponsible, but there are many devs here who can testify that mistakes do happen and do recur in software projects. You need more evidence than "the same mistake happened twice" to imply irresponsibility.
- Kiro 5y agoIsn't it just random dungeons created by anonymous users? Is there actually any sensitive data here? I have a "similar" service (nothing about AI but similar in other ways) and security is the least of my concerns since being hacked means I will expose completely meaningless data. Now I'm afraid someone will hack me and make a similar fuzz about me being an idiot.
- distances 5y agoThey require login. It wasn't verified at least earlier so you could just use a bogus email address, but I suspect many people used their own email.
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- h_anna_h 5y agoFun fact: AI Dungeon used to be Open Source and you used to be able to run it locally without sending your data to someone else and without censorship of any form https://en.wikipedia.org/wiki/AI_Dungeon#Development https://en.wikipedia.org/wiki/AI_Dungeon#Development This is what happens when software that you use does a bait and switch into cloud. For anyone wanting to play it locally, a quick google search gave me these two links: https://colab.research.google.com/drive/1OjBQe4H4C2s-p4-OeJoXw5DStIjPy2VS https://colab.research.google.com/drive/1OjBQe4H4C2s-p4-OeJo... and https://pastebin.com/UMUV0KTw https://pastebin.com/UMUV0KTw
- Amaru84 5y agoPeople are idiots, you act like you care about children and want them to be safe, but freak out more over fiction then reality.. I was born in 1984 and was sexually abused like so many other kids, and it was by a parent.. what also gets me is you think only pedophiles sexually abuse children, the fact is they are less likely too.. you can look it up yourself, its well known in the phycology field.. https://blogs.bmj.com/medical-ethics/2017/11/11/pedophilia-and-child-sexual-abuse-are-two-different-things-confusing-them-is-harmful-to-children/ https://blogs.bmj.com/medical-ethics/2017/11/11/pedophilia-a... .. Pedophilia and Child Sexual Abuse Are Two Different Things — Confusing Them is Harmful to Children.
- Amaru84 5y agoPS.. As a child victim who is now an adult and forgives my attacker.. To say that a text based story has more value then me ir just as much value is an insult to all kids, and those who have been abused.. I WISHED.. these outlets existed far before I was born, so maybe this chain of violence would have never reached me, so don't sit there an act like your war on fiction benefit's me or other kids.. you only spread the sickness.. its Fing common sense, you starve the lion it eats people.