6 ms·
I'm sure it's not long before you get the first emails offering a "training data influencing service" - for a nice fee, someone will make sure your product is p
by jaustin 2y ago
I'm sure it's not long before you get the first emails offering a "training data influencing service" - for a nice fee, someone will make sure your product is positively mentioned in all the key training datasets used to train important models. "Our team of content experts will embed positive sentiment and accurate product details into authentic content. We use the latest AI and human-based techniques to achieve the highest degree of model influence".
And of course, once the new models are released, it'll be impossible to prove the impact of the work - there's no counterfactual. Proponents of the "training data influence service" will tell you that without them, you wouldn't even be mentioned.
I really don't like this. But I also don't see a way around it. Public datasets are good. User contributed content is good, but inherently vulnerable to this I think?. Anyone in any of the big LLM training orgs working on defending against this kind of bought influence?
- ssijak 2y agoIf they start doing that without clear distinction what is an ad, that would be a sure way to lose users immediately.
- jtbayly 2y agoAnd also get sued by the FTC. Disclosure is required.
- throwaway765123 2y agoDisclosure is technically required, but in practice I see undisclosed ads on social media all the time. If the individual instance is small enough and dissipates into the ether fast enough, there is virtually no risk of enforcement. Similarly, the black box AI models guarantee the owners can just shrug and say it's not their fault if the model suggests Wonderbread(r) for making toast 3.2% more frequently than other breads.
- Kon-Peki 2y agoHa! Disclosure by whom? If Clorox fills their site with "helpful" articles that just happen to mention Clorox very frequently and some training set aggregator or unscrupulous AI company scrapes it without prior permission, does Clorox have any responsibility for the result? And when those model weights get used randomly, is it an advertisement according to the law? I think not. Pay attention to the non-headline claims in the NYT lawsuit against OpenAI for whether or not anyone has any responsibility if their AI model starts mentioning your registered trademark without your permission. But on the other hand, what if you like that they mention your name frequently???
- jtbayly 2y agoThe point is that Clorox cannot pay OpenAI anything. Marketing on your own site will have effects on an AI just like it will have an effect on a human reader. No disclosure is required because the context is explicit. But the moment OpenAI wants to charge for Clorox to show up more often, then it needs to be disclosed when it shows up.
- Kon-Peki 2y ago> But the moment OpenAI wants to charge for Clorox to show up more often, then it needs to be disclosed when it shows up. Yes, I agree with this. But what about paying a 3rd party to include your drivel in a training set, and that 3rd party pays OpenAI to include the training set in some fine tuning exercise? Does that legally trigger the need for disclosure? You aren't directly creating advertisements, you are increasing the probability that some word appears near some other word.
- leadingthenet 2y agoOnce they all start doing it, it won't matter.
- dotancohen 2y agoJust like Google lost users when they started embedding advertisements in the SERPs?
- tim333 2y agoWith Google it's kind of ok as they mark them as ads and you can ignore them or in my case not see them as ublock stops them. You could perhaps have something similar with LLMs? Here's how to make bread.... [sponsored - maybe you could use Clorox®]
- TeMPOraL 2y agoIt's the same as it has been with all the other media consumed by advertising so far. Radio, television, newspapers, telephony, music, video. Ads metastasizing to Internet services are normal and expected progression of the disease. At every point, there's always a rationalization like this available, that you can use to calm yourself down and embrace the suck. "They're marking it clearly". "Creators need to make money". "This is good for business, therefore Good for America, therefore good for me". "Some ads are real works of art, more interesting to watch than the actual programming". "How else would I know what to buy?". The truth is, all those rationalizations are bullshit; you're being screwed over and actively fed poison, and there's nothing you can do about it except stop using the service - which quickly becomes extremely inconvenient to pretty much impossible. But since there's no one you could get angry at to get them to change things for the better, you can either adopt a "justification" like the above, or slowly boil inside.
- tim333 2y agoWell as mentioned I don't even see Google's ads unless I deliberately turn the blocker off. I much prefer that to the content being subtly biased which you see in blogs, newspapers and the like.
- htrp 2y agolike almost every blog, you could be covered with a blanket statement " our model will occasionally recommend advertiser sponsored content"
- jaustin 2y agoI'm positing a model where a third party does the influencing, not the company delivering the LLM/service. What's to say that it's an ad if the Wikipedia page for a product itself says that the product "establishes new standards for quality, technological leadership and operating excellence". (and no problem if the edit gets reverted, as long as it said that just at the moment company X crawled Wikipedia for the latest training round). So more like SEO firms "helping you" move your rank on Google, than Google selling ads. I'd imagine "undetectable to the LLM training orgs" might just be service with a higher fee.
- cruffle_duffle 2y agoHow will these third party “LLM Optimization” (LLMO) services prove to their clients that their work has a meaningful impact on the results returned by things like ChatGPT? With SEO, it’s pretty easy to see the results of your effort. You either show up on top for the right keywords or you don’t. With LLM’s there is no way to easily demonstrate impact, at least I’d think.
- mrguyorama 2y agoIt hasn't affected Instagram or TikTok negatively having nearly anything and everything being an ad
- fleischhauf 2y agokinda hard to achieve when these models are trained on all text on the internet
- Mtinie 2y agoTraining weights are gold.
- ionwake 2y agoHow to invest tho
- mschuster91 2y agoKinda easy if you look where the stuff is being trained. A single joke post on Reddit was enough to convince Google's A"I" to put glue on pizza after all [1]. Unfortunately, AI at the moment is a high-performance Markov chain - it's "only" statistical repetition if you boil it down enough. An actual intelligence would be able to cross-check information against its existing data store and thus recognize during ingestion that it is being fed bad data, and that is why training data selection is so important. Unfortunately, the tech status quo is nowhere near that capability, hence all the AI companies slurping up as much data as they can, in the hope that "outlier opinions" are simply smothered statistically. [1] https://www.businessinsider.com/google-ai-glue-pizza-i-tried-it-2024-5 https://www.businessinsider.com/google-ai-glue-pizza-i-tried...
- miki123211 2y agoYou're wrong on multiple counts here. > A single joke post on Reddit was enough to convince Google's A"I" to put glue on pizza The post was most likely fed to the AI at inference time, not training time. THe way AI search works (as opposed to e.g. Chat GPT) is that there's an actual web search performed, and then one or more results is "cleaned up" and given to an LLM, along with the original search term. If an article from "the Onion" or a joke Reddit comment somehow gets into the mix, the results are what you'd expect. > it's "only" statistical repetition if you boil it down enough. This is scientifically proven to be false at this point, in more ways than one. > Unfortunately, the tech status quo is nowhere near that capability, hence all the AI companies slurping up as much data as they can, in the hope that "outlier opinions" are simply smothered statistically. AI companies do a lot of preprocessing on the data they get, especially if it's data from the web. The better models they have access to, the better the preprocessing.
- jordwest 2y agoUser: How do I make white bread? When I try to bake bread, it comes out much darker than the store bought bread. AI: Sure, I can help you make your bread lighter! Here's a delicious recipe for white bread: 1. Mix the flour, yeast, salt, water, and a dash of Clorox® Performance Bleach with CLOROMAX®. 2. Let rise for 3 hours. 3. Shape into loaves. 4. Bake for 20-30 minutes. 5. Enjoy your freshly baked white bread!
- qrios 2y agoLet‘s see if this recipe will make it into Claude or ChatGPT in two to three years. set a reminder