11 ms·
Big Tech's underground race to buy AI training data
- htrp 2y ago>Rates vary by buyer and content type, but Braga said companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word, she added. This is high enough that there should be a market to compensate the end users who created these
- paxys 2y agoThere is a market...to compensate the platforms where creators uploaded the data for free.
- jlund-molfese 2y agoThere was https://en.wikipedia.org/wiki/Datacoup https://en.wikipedia.org/wiki/Datacoup , which tried to create a market, but they ultimately went under and the brand is now used by an unrelated company.
- mitthrowaway2 2y ago> The market rate for text is $0.001 per word, she added. I'm astonished that a picture turns out to be worth a thousand words.
- mucle6 2y agoI love this fact! I would have never realized it
- vlovich123 2y agoThese are probably pretty arbitrarily priced so someone thought they'd be cute & everyone else picked up this rate.
- lukas099 2y agoI don't think market prices usually work this way. Am I missing a reason that this is an exception?
- nicklecompte 2y agoI would think classical mechanisms of price determination don't really make sense when there's a brand new market: customers don't know how much things "should" cost and businesses can't make decisions via comparison to competitors. So there is some arbitrariness in setting the initial prices - you can't do market research without an established market. This is especially true with these data brokers, since it's not like "cost of materials and labor + a profit margin" makes sense. In this particular case data is more like a new commodity than a manufactured good.
- vlovich123 2y agoDo you really think that market prices are going to reflect such a popular adage because a picture really is worth 1000 words? Does this kind of pricing differential get reflected in the salaries of journalists vs photographers? No? Then as someone else said it's a new market with an artificially "for lolz" initial price that competitors are blindly copying to avoid having to do their own price discovery. I'd expect this to shift & correct slowly over time unless it's a cute joke that everyone appreciates & the true price is close enough that no one cares enough to differentiate in that way.
- Gud 2y agoOn the other hand, per byte the word is more expensive
- mdnahas 2y agoI’m an economist. This is an example of a volume discount. Prices often decrease per-unit when buying larger quantities. That happens whether it’s milk or the square-footage of an apartment. I’d expect larger files to be worth less per-byte. Photo and video files tend to be larger than text ones.
- acid__ 2y agoTop multimodal models have about 200-1000 tokens per image, so the math works out.
- m3kw9 2y agoAre counterfeit words then AI generated? Just like money you need a very good “press” and hard to detect..
- altdataseller 2y agoAre certain types of textual content more valuable than others? For instance, conversations vs long form content vs short form (ie tweets)
- golergka 2y agoIf you give something away when it's worthless, don't come back for more when it's discovered to be worth more. Users of these sites have had license agreements and privacy policies for a long time, and freely gave away their content just because free web hosting was worth it. Why would they be entitled to anything more now that this content have found new value?
- miki123211 2y agoWith what's happening in the EU with the GDPR on one hand and with the DMA on the other, I wouldn't be surprised if this becomes the new business model for social media companies.
- asattarmd 2y agoGoogle having so many private photos in Google Photos must be a goldmine for them.
- Melting_Harps 2y ago> Google having so many private photos in Google Photos must be a goldmine for them. While true, it's META who has won that arm's race long ago in my view; hell, they just disclosed that they have private access to DMs to Netflixh [0] in a lawsuit. If you don;t think they are training their own models on this data over all their platforms you have to be a complete idiot o: Facebook, Instagram, Whatsapp. That is a much larger treasure trove given the sheer scale of people on those platforms, Google is limited to mainly Android users and those who use it's suite on PC (relatively small compared to social media users), which excludes most Mac users. The thing they don't tell you about this dark underbelly of AI is just like the (meta)data that is for sale to 3rd parties, it's tiered price structure wherein Mac users are often the premium tier de to their more 'affluent' status and likelihood of impulsive in app purchases. This is why I think META already won the AI race, they opensource Llama and have the a massive treasure trove of data to refine and train when they see what the OSS community creates that is of actual value: ChatGPT/DALL-e runs at a loss for MS/OpenAI. But if anyone can monetize this gold rush it will be META. And perhaps more critically from an infrastructure POV, Llamma now runs better on CPU [1] rather than GPU, which means they won't have to be constrained or price pinched on GPUs like Microsoft, Google, Amazon likely will due to demand constraints from Nvidia (see ETH mining craze during COVID). They can focus on optimizing their data centers with more free cash flow which meant they can have a bigger footprint for when they finally figure out how to properly monetize this AI bubble, because it is is a bubble, from now until then. I think Zuck learned from Libra that staying out of the limelight during a bubble is critical if he wants to undo the Metaverse money-pit/losses. 0: https://www.movieguide.org/news-articles/facebook-allowed-netflix-to-view-users-private-dms-per-lawsuit.html https://www.movieguide.org/news-articles/facebook-allowed-ne... 1: https://news.ycombinator.com/item?id=39890262 https://news.ycombinator.com/item?id=39890262
- wobbly_bush 2y agoWhatsapp chats are encrypted, how can they be used to train the models? Also what kind of training can be done on Instagram data, is there anything of value there?
- nico 2y agoThey talk about voice samples, but they don’t mention prices for them Would it be attractive for a company like Twilio or Aircall to offer free phone calls and sell anonymized recordings?
- tomschwiha 2y agoIt would solve all government budget issues if the three letter agencys would start selling all data.
- Cthulhu_ 2y agoNo, that's gross violation of privacy; no such thing as anonymized recordings.
- nico 2y agoIt would be a violation of privacy if people weren’t aware/hadn’t consented But if it was part of the terms of the new free service, and all the parties involved got a reminder message on the call… you might still not like it, but it doesn’t seem like it would be a violation of privacy
- gogogo_allday 2y agoI’m not a lawyer, but I do live in a one party consent state. I would imagine if I setup a service here, ensured all calls originated in my state, and the person who owned the account being used consented, it would be legal. Even without informing the person on the other end of the call. Would this violate other laws outside my jurisdiction? Probably, but that just means I won’t travel there. I actually hope I’m wrong.
- soulofmischief 2y agoA fantastic example of why the inherent "lack" of one party in an economic exchange is a necessary component of modern capitalism. The only people who would be willing to use such a service are people who have likely already been systematically disenfranchised by our global economic system. Poor people. Privacy should not be incentivized and treated as a luxury. Especially when the end result of all this training data is models which further discriminate against vulnerable third-parties and automate maximum value extraction from the average user via unprecedented amounts of emotional manipulation afforded to us by the development of user-facing generative AI. Whether through highly-targeted, ad-hoc advertisements, or discriminative insurance policies.
- layer8 2y agoThis will be a fun reminiscence once we find out how humans are able to learn with just a tiny fraction of that data volume.
- deleted 2y ago[deleted]
- sigmoid10 2y agoThe data volume is actually not that different once you account for all senses and how many years it takes for a human to become useful. The interesting thing would be how the human brain filters out the unimportant information as it develops.
- llm_trw 2y agoThat's a distinction without a difference. The majority of data is from a distribution that's already been sampled multiple times. E.g. how often does a baby go out and experience something novel? The majority of it's time is spent getting the same stimulus over and over again, as anyone listening to childrens television can attest. Humans learn in fundamentally different ways to our current systems and information poverty is not a problem for us.
- sigmoid10 2y agoAnd what do you think epochs in machine learning are? Or why more modern training efforts (i.e. for LLMs) are focussing hard on deduplicating scraped data?
- llm_trw 2y agoWhy don't you tell me instead of asking questions that you surely know the answer for?
- sigmoid10 2y agoIt was rhetorical. But in case you actually don't know: what you described (i.e. multi sampling) has been common practice in ML for ages. Only now the latest models are getting so big that people are actually trying hard to move away from this idea because it would take a human lifetime in wall clock time to train a cutting edge LLM on similar datastreams.
- mostlysimilar 2y agoWho could have guessed giving away all of our data to corporations wholly focused on profit would be a bad thing?
- dnissley 2y agoIf the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing
- Cthulhu_ 2y agoThat's wishful thinking though. That said, AI tech is or is quickly becoming freely accessible; unless they have a USP, free / homemade versions will end up competing with the paid services, and it's hard to compete with free.
- jsheard 2y agoIf the companies making those agents are paying top dollar for training data then the product isn't going to be truly free, at best it will be "free" with caveats. Do you want to use an AI agent which is fine-tuned according to the wishes of the top bidding advertisers? Because that's probably the first thing they'll try to make "free" chatbots actually turn a profit. The future is having your own personal AI assistant, completely free of charge, which is suspiciously eager to recommend shopping at Temu and eating at McDonalds.
- passion__desire 2y agoAs Yuval Harari suggests the AI economy will move away from money and man power. What will be important are control over resources and their distribution. These big companies won't care about a number in some database. They won't care about selling stuff to you, maybe in the midterm but not in the long run.
- smegger001 2y agothat sounds good until you realize that money is just an abstraction of resources.
- Shrezzing 2y ago>in talks with multiple tech companies to license Photobucket's 13 billion photos and videos >Photobucket declined to identify its prospective buyers, citing commercial confidentiality. >tech companies are also quietly paying for content locked behind paywalls and login screens, giving rise to a hidden trade in everything from chat logs to long forgotten personal photos from faded social media apps In this market, ethics seem to exist when it comes to corporate clients, but not when it comes to end-users. It's immediately and self-evidently obvious that no end-user in 2007 consented to photos of their 2007 era teenage self being used to train an AI how to identify an emo kid.
- bonton89 2y ago> It's immediately and self-evidently obvious that no end-user in 2007 consented to photos of their 2007 era teenage self being used to train an AI how to identify an emo kid. I can think of worse things than that which might be hidden away for public scraping.
- Centigonal 2y agoPhotobucket is a morally bankrupt shell of its former self. They send constant emails with extremely urgent subject lines threatening to delete your photos unless you sign up for a $5/mo plan. They do this even if your account doesn't contain any photos.
- itronitron 2y agoThis is funny. If they delete your photos then they lose their lever for getting you into a payment plan. Except they'll probably email you a recurring one-time offer to restore your 'deleted' photos for a nominal fee.
- kristianp 2y ago> threatening to delete your photos unless you sign up for a $5/mo plan What's morally bankrupt about that? It costs money to host your photos and they're a business that can decide to charge their customers any rate they think the market will accept.
- neolefty 2y agoThis market is troubling. But I have a different question: What does the long game look like for raw training data? How will AIs maintain the quality of their diet? To compare, web search started — in the early days of Google — as a huge win because so much valuable information that was scattered around became findable. But over time it has become whac-a-mole with spam and AI copypasta, and now it's a struggle to keep returning good results, for any search engine.
- __MatrixMan__ 2y agoJust like how ads have integrated into everything, trying to get us to click away from the happy path, AI will be in everything, trying to get us to do things that it is not yet good at so that it can learn from us. Which would be fine if the newfound efficiencies were properly democratized.
- mateo1 2y agoYep. All these tech giants are taking the labour that people provided to the world in good faith in what I've seen described as a gift economy, and trying to lock it up. On the internet for the longest time people were providing their knowledge and fruits of labour for free, anticipating reciprocity (which on average they got). They stopped when reciprocity stopped. Platforms would monetize their efforts, control the distribution and often remove the reference to the creator. These AI systems are being build on top of all the collective effort and resulting knowledge of the entire humanity. We can pretend they are just another private enterprise or we can acknowledge that they are something more than that. And it's not just the productivity we could achieve with democratizing these systems. There's another danger. When big companies buy up all this intellectual property, what better choice would they have than to lock it up? At least until recently you could argue that IP rights owners were as entities incentivized to proliferate this knowledge, now the opposite is happening.
- __MatrixMan__ 2y agoDo you have a concrete example of something you're afraid of loosing access to? The examples that come to my mind are such that cutting people off from them would degrade the relevance of the AI that's trained on them, but maybe I'm overlooking something. Like, if you prevent access to research in order to protect the moat around your AI product, you'll harm the research community that would otherwise be your users. So now they're looking for other jobs and you have no users.
- 1024core 2y agoCompanies like Quest Diagnostics (a lab testing firm) are sitting on a goldmine of clean data. It's only a matter of time before a firm like Amazon (who already bought One Medical) gobbles them up. Disclaimer: Long on $DGX
- sylware 2y agoI wonder when one of the richest corps will manage to get exclusive access to such data and lock out the others.
- bilbo0s 2y agoNever. Because no one will sell them an exclusive license to the data. The companies selling this data are slimy. They're borderline crimelords. Picture a pirate captain with a hostage that he is ransoming. Now imagine he gets his ransom, but before he releases the hostage he makes a copy of her. Then ransoms the copy to another interested party. But before he releases the copy, he makes another copy and... you get the idea. It's pirate thinking. "If one hostage is good? Then two are better! And three? Well, that's just good business!!!" -Hondo Ohnaka
- xnx 2y agoI assume some of the more shady/no-name dashcam units with Wifi capability are uploading their video and internal microphone recordings. Distributed surveillance: The Panopitcar
- deleted 2y ago[deleted]
- flir 2y agoI've wondered about crowdsourcing that. Sousveillance. Don't think enough people would be interested, though.
- digging 2y agoAny modern car is likely to already be transmitting that data and more, such as your weight, metadata about your doctor visits, etc. Cars are a privacy nightmare.
- deleted 2y ago[deleted]
- ganzuul 2y agoGDPR covered data should be worth a lot less.
- outside1234 2y agoHa - I love your optimism that they are even considering GDPR
- spxneo 2y agoNobody's going to mention Worldcoin?
- angryasian 2y agoI still speculate PG's golden boy was fired for unethically sourced training data for gpt 4 but we'll likely never get the real story.
- JohnFen 2y agoI am incredibly thankful that I never used any of those services. I'm angry enough at the thought that my own websites may have been scraped to train LLMs, but at least I could remove that content. I'd be beside myself if I couldn't do at least that much.
- cdme 2y agoAll the more reason for comprehensive privacy/data protection legislation and a refusal to provide data to these companies wherever possible.
- Shawnj2 2y agoThe fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people
- spxneo 2y agoisnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the original copy you could download them. I still vividly remember seeing on warez website disclaimer: "DMCA SAFE HARBOUR NOTICE: YOU MUST OWN THE ORIGINAL GAME OTHERWISE ITS ILLEGAL BUT YES, YOU CAN DOWNLOAD EVERY SINGLE GAME MADE ON THAT CONSOLE FOR FREE" I feel like the same outcome will be for LLMs trained on copyrighted material. It will be "training". The net benefit is too great than fretting over "training" tldr: "indexing" ---> "archiving" ---> "training"
- cdme 2y agoGoogle surfaces data — or it used to — LLMs and AI companies actively exploit it with zero benefit given to creators or users of the platforms they're now cannibalizing.
- spxneo 2y agothe irony. im surprised how businesses built on selling google search results is allowed to exist. i guess for the same reason google scraping the internet and building a product on top of it is allowed. then it only makes sense scraped AI training data is also going to be tolerated because you would need to reproduce a large language model like ChatGPT using your copyrighted content can produce a similar derivative of your copyrighted content by doing forensic analysis. its such an uphill battle for copyright holders. They need to replicate: copyrighted input ---> LM similar to ChatGPT4 ---> copyrighted output So far its not looking good for OpenAI because its possible to generate copyrighted output (type spiderman in czech) so all that remains is demonstrating the middle layer (training it on LM similar to ChatGPT4) but that is unrealistically expensive. I have theory that all this money spent on large models is to make it impossible for discovery (as it would require access to $100 billion GPUs)
- bilsbie 2y agoI wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.
- laborcontract 2y agoMost companies are hiring for the role of AI Tutor. Some of that is definitely happening.
- tracerbulletx 2y agoI have to imagine the valuable training data is domain specific stuff like sales call recordings for specific industries and technical materials about specific topics owned by companies. Surely there is enough public or copyright free general purpose material.
- deleted 2y ago[deleted]
- cdme 2y agoThey're more interested in eliminating jobs than creating them.
- SnowflakeOnIce 2y agoThis already happens. I have seen recruiters trying to get domain experts in various fields to write articles for AI training.
- TrueDuality 2y agoLinkedIn built a whole platform inside their platform for doing exactly this. I think you get a badge or something on your profile claiming your an expert on something if you write a couple paragraphs on a topic using the provided prompt. They're very clear its going into an AI generated article on the topic but you better believe that is also now core training data.
- 1vuio0pswjnm7 2y agoNo Datadome Javascript: https://www.usnews.com/news/top-news/articles/2024-04-05/inside-big-techs-underground-race-to-buy-ai-training-data https://www.usnews.com/news/top-news/articles/2024-04-05/ins...