7 ms·
here‘s what i don‘t get about this whole discussion. AI companies already store all prompts and responses for future training. just make an API that returns th
by ch_sm 2mo ago
here‘s what i don‘t get about this whole discussion. AI companies already store all prompts and responses for future training.
just make an API that returns the string distance between a previously generated paragraph and the query?
that would sidestep this whole problem class.
regulators could even specify how that has to work.
what am i missing?
- dpoloncsak 2mo agoLocal models?
- dTal 2mo agoLocal models enable: - watermark-free generation - the stripping of watermarking from the output of SAAS models Any discussion of watermarking is dead in the water in a world where we are permitted to have these things. I fear for the future.
- thisoneworks 2mo agoChill my dude. This is just a sane default which will catch normies copy pasting stuff from claude and chatgpt. It's good enough.
- dTal 2mo agoDo you not want "normies" to have local AI? Do you believe that your ability to use esoteric software makes you special, makes you immune to law? The days where you can find refuge in running "weird hacker crap" instead of 100% goobermint approved software are numbered because news flash: AI means there are no normies anymore. Anyone can do anything. No, I will not chill. The war on general purpose computing is gonna get real hot real soon.
- bethekidyouwant 2mo agoPasting it where?
- queenkjuul 2mo agoHomework assignments, spam emails, LinkedIn
- deleted 2mo ago[deleted]
- bethekidyouwant 2mo agothe api to detect if text is AI generated will be for large bespoke companies, governments etc. not teachers or the general public,
- queenkjuul 2mo agoUniversities are functionally large companies
- bethekidyouwant 2mo agoOh so you already knew that? If they give it to every teacher it’s the same as giving it to everyone so they’re not going to do that.
- queenkjuul 2mo agoAnd giving it to employees of a large company is materially different in what way?
- dpoloncsak 2mo agoLocal models also enable: - not being locked into a provider - not being forced to have your prompts saved by a possible competitor - an alternative to the duolopy we quickly see forming - offline access I fear for the future without local models, much more than the future with them, and would rather everyone had access to a local model than be certain we catch everyone copy+pasting LLM responses. Watermarking would be cool, but it's not worth losing local for.
- dTal 2mo agoTo be clear, we agree. The problem is that unless local AIs become "normie friendly" real damn quick, we're gonna lose em, because they're damned inconvenient to power. That is what I fear.
- dpoloncsak 2mo agoMy apologies, I misunderstood your original comment to mean "We shouldn't have local models if it breaks watermarking" but yes, I think we're on the same page now
- bethekidyouwant 2mo agoI don’t get how you’re gonna lose them how is a guy from Brussels gonna come to your house and turn off your computer?
- dTal 2mo agoIn the short term, governments can suppress free distribution of open weight LLMs. For instance, they could "ask" Hugginface to require an account with a university affiliated email address to download anything. It won't stop you outright, but it will force it underground, slowing everything to a crawl. In the long run... well, general purpose computing is under threat anyway. Platforms with locked bootloaders and mandatory digital signatures outnumber those that don't. Even on "PCs", the openness is merely a cultural norm observed for a particular market segment[0] by Apple and Microsoft - for they are the only ones with true root keys. Even a brand new motherboard comes with Microsoft signing keys pre-flashed, and you are permitted to enroll your own only by grace. The technical infrastructure to flip a switch and lock it all down is now in place, ready to go at the stroke of a pen. Will you be able to get Llama.cpp from "the app store"? I hazard not. If you think this sounds hyperbolic, look at what is happening to Android. [0] They don't even segment it the same! Apple segments by "does it have a keyboard", and Microsoft segments it by "does it have a native x86 processor". Apple runs different software on the same basic hardware, Microsoft runs the same basic software on different hardware, but both have created artificially restricted second class computing categories.
- nradov 2mo agoNo need to fear, the future looks bright!
- graemep 2mo agoVery few people will use local models. In practical terms most people use a few big providers.
- dTal 2mo agoIs that a desirable state of affairs?
- mirashii 2mo ago> AI companies already store all prompts and responses for future training. They store some prompts and responses, not all, that's what you're missing.
- bonoboTP 2mo ago> AI companies already store all prompts and responses for future training. They claim to only do this when you agree to this in your personal settings. Though Google does say they will train on it, unless you disable history and only use ephemeral chats. Anthropic has a setting for it and claims not to train by default. Also it would be vulnerable to attacks and privacy problems. You could search for substrings about some suspected information, like "John Smith's medical records show advanced cancer" etc. Of course you'd have to guess the phrasing but still.
- ch_sm 2mo agoRight, my bad; I know they claim to not store _all_ prompts and responses, that was a bit cynical/hyperbolical of me. However, regulation could force them to do it — with all the downsides that come with that. Your privacy argument, on the other hand, makes total sense. Not even security measures like homomorphic encryption or locality-sensitive hashing could fix that basic issue.
- johnnyo 2mo ago1. AI company buys and trains on an author’s book when it gets published, it’s now part of the training data. 2. Attacker asks the LLM for the opening sentences of the book, it goes into the generated responses database. 3. Later, a malicious user shows that the first few sentences of the authors book are identical to a previously generated response.
- dpoloncsak 2mo agoIgnoring the other technical hurdles of the idea...this problem you outlaid is solved with a timestamp in the database, right? You can easily prove if the prompt was before/after publishing date?
- johnnyo 2mo agoIf it’s before publication date, does that prove the author used AI inappropriately in their writing? Not really, there are any number of reasons why they might do that. In my own personal writing, sometimes I run it through LLMs to give grammar or writing suggestions.
- echoangle 2mo agoWouldn't this become pretty difficult to do at scale over time? Is there a way to compare the similarity in a database of responses without doing a search over every entry and comparing them? Because that would probably become pretty slow if literally every LLM output is saved and has to be scanned.
- rsynnott 2mo agoWell, for a start, they’re probably on dodgy ground there with other European regulations. You’re not really supposed to hold onto hold onto potentially very sensitive data that you have no legitimate interest (a term of art; it doesn’t just mean “I want to keep this”) in indefinitely. As far as I know, most of these require user consent for retention?