15 ms·
AI groups spend to replace low-cost 'data labellers' with high-paid experts
- aspenmayer 1y agohttps://archive.is/dkZVy https://archive.is/dkZVy
- Melonololoti 1y agoYepp it continues the gathering of more and better data. Ai is not a hype. We have started to actually do something with all the data and this process will not stop soon. Aline the RL what is now happening through human feedback alone (thumbs up/down) is massive.
- KaiserPro 1y agoIt was always the case. We only managed to make a decent model once we created a decent dataset. This meant making a rich synthetic dataset first, to pre-train the model, before fine tuning on real, expensive data to get the best results. but this was always the case.
- noname120 1y agoRLHF wasn't needed for Deepseek, only gobbling up the whole internet — both good and bad stuff. See their paper
- rtrgrd 1y agoI thought human preferences was typically considered a noisy reward signal
- ACCount36 1y agoIf it was just "noisy", you could compensate with scale. It's worse than that. "Human preference" is incredibly fucking entangled, and we have no way to disentangle it and get rid of all the unwanted confounders. A lot of the recent "extreme LLM sycophancy" cases is downstream from that.
- smohare 1y ago[dead]
- TheAceOfHearts 1y agoIt would be great if some of these datasets were free and opened up for public use. Otherwise it seems like you end up duplicating a lot of busywork just for multiple companies to farm more money. Maybe some of the European initiatives related to AI will end up including the creation of more open datasets. Then again, maybe we're still operating from a framework where the dataset is part of your moat. It seems like such a way of thinking will severely limit the sources of innovation to just a few big labs.
- KaiserPro 1y ago> operating from a framework where the dataset is part of your moat Very much this. Its the dataset that shapes the model, the model is a product of the dataset, rather than the other way around (mind you, synthetic datasets are different...)
- andy_ppp 1y agoWhy would companies paying top dollar to refine and create high quality datasets give them away for free?
- flir 1y agoSame reason they give open source contributions away for free. Hardware companies attempting to commoditize their complement. I think the org best placed to get strategic advantage from releasing high quality data sets might be Nvidia.
- mh- 1y agoIndeed, and they do. https://huggingface.co/nvidia/datasets?sort=most_rows https://huggingface.co/nvidia/datasets?sort=most_rows
- charlieyu1 1y agoThere are some good datasets for free though, eg HLE. Although I’m sure if they are marketing gimmicks
- 1y ago
- panabee 1y agoThis is long overdue for biomedicine. Even Google DeepMind's relabeled MedQA dataset, created for MedGemini in 2024, has flaws. Many healthcare datasets/benchmarks contain dirty data because accuracy incentives are absent and few annotators are qualified. We had to pay Stanford MDs to annotate 900 new questions to evaluate frontier models and will release these as open source on Hugging Face for anyone to use. They cover VQA and specialties like neurology, pediatrics, and psychiatry. If labs want early access, please reach out. (Info in profile.) We are finalizing the dataset format. Unlike general LLMs, where noise is tolerable and sometimes even desirable, training on incorrect/outdated information may cause clinical errors, misfolded proteins, or drugs with off-target effects. Complicating matters, shifting medical facts may invalidate training data and model knowledge. What was true last year may be false today. For instance, in April 2024 the U.S. Preventive Services Task Force reversed its longstanding advice and now urges biennial mammograms starting at age 40 -- down from the previous benchmark of 50 -- for average-risk women, citing rising breast-cancer incidence in younger patients.
- matusp 1y agoThis is true for every subfield I have been working on for the past 10 years. The dirty secret of ML research is that Sturgeon's law apply to datasets as well - 90% of data out there is crap. I have seen NLP datasets with hundreds of citations that were obviously worthless as soon as you put the "effort" in and actually looked at the samples.
- panabee 1y ago100% agreed. I also advise you not to read many cancer papers, particularly ones investigating viruses and cancer. You would be horrified. (To clarify: this is not the fault of scientists. This is a byproduct of a severely broken system with the wrong incentives, which encourages publication of papers and not discovery of truth. Hug cancer researchers. They have accomplished an incredible amount while being handcuffed and tasked with decoding the most complex operating system ever designed.)
- skeezyboy 1y ago
- techterrier 1y agoThe latest in a long tradition, it used to be that you'd have to teach the offshore person how to do your job, so they could replace you for cheaper. Now we are just teaching the robots instead.
- verisimi 1y agoThis is it - this is the answer to the ai takeover. Get an ai to autogenerate lots of crap! Reddit, hn comments, false datasets, anything!
- Cthulhu_ 1y agoThat's just spam / more dead internet theory, and there will be or are companies that will curate data sets and filter out generated stuff / spam or hand-pick high quality data.
- vidarh 1y agoI've done review and annotation work for two providers in this space, and so regularly get approached by providers looking for specialists with MSc's or PhD's... "High-paid" is an exaggeration for many of these, but certainly a small subset of people will make decent money on it. At one provider I was as an exception paid 6x their going rate because they struggled to get people skilled enough at the high-end to accept their regular rate, mostly to audit and review work done by others. I have no illusion I was the only one paid above their stated range. I got paid well, but even at 6x their regular rate I only got paid well because they estimated the number of tasks per hour and I was able to exceed that estimate by a considerable margin - if their estimate had matched my actual speed I'd have just barely gotten to the low end of my regular rate. But it's clear there's a pyramid of work, and a sustained effort to create processes to allow the bulk of the work to be done by low-cost labellers, and then push smaller and smaller subsets of the data up more expensive to experts, as well as creating tooling to cut down the amount of time experts spend by e.g. starting with synthetic data (including model-generated reviews of model-generated responses). I don't think I was at the top of that pyramid - the provider I did work for didn't handle many prompts that required deep specialist knowledge (though I did get to exercise my long-dormant maths and physics knowledge that doesn't say too much). I think most of what we addressed would at most need people with MSc level skills in STEM subjects. And so I'm sure there are a few more layers on the pyramid handling PhD-level complexity data. But from what I'm seeing from hiring managers contacting me, I get the impression the pay scale for them isn't that much higher (with the obvious caveat given what I mentioned above that there almost certainly are people getting paid high multiples on the stated scale) Some of these pipelines of work are highly complex, often including multiple stages of reviews, sometimes with multiple "competing" annotators in parallel feeding into selection and review stages.
- ljlolel 1y ago[dupe]
- charlieyu1 1y agoI'll believe it when it happens. A major AI company got rid of an expert team last year because they think it is too expensive
- quantum_state 1y agoIt is expert system evolved …
- cryptokush 1y agowelcome to macrodata refinement
- joshdavham 1y agoI was literally just reached out to this morning about a contract job for one of these “high quality datasets”. They specifically wanted python programmers who’ve contributed to popular repos (I maintain one repository with approx. 300 stars). The rate they offered was between $50-90 per hour, so significantly higher than what I’d think low-cost data labellers are getting. Needless to say, I marked them as spam though. Harvesting emails through GitHub is dirty imo. Was also sad that the recruiter was acting on behalf of a yc company.
- apical_dendrite 1y agoThe latest offer I saw was $150-$210 an hour for 20hrs/week. I didn't pursue it so I don't know if that's what people actually make, but it's an interesting data point.
- antonvs 1y agoWhat kind of work was involved for that one? How specialist, I mean?
- Der_Einzige 1y agoDang - such a great opportunity for anyone which knows how to tweak their LLMs sampling settings. Whatever tricks they are using to try to detect folks who just answer with other LLM outputs will be defeated at high temperature with the weird settings and steering adapters I apply. If regular samplers still keep it getting detected, I can use https://github.com/sam-paech/antislop-sampler https://github.com/sam-paech/antislop-sampler to guarantee that it won't look AI generated. Send these my way. I'll be one of their most productive data labelers :)
- apical_dendrite 1y agoSure - where can I send you a referral link?
- deleted 1y ago[deleted]
- SoftTalker 1y agoIsn't this ignoring the "bitter lesson?" http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- TrackerFF 1y agoI don’t know if it is related, but I’ve noticed an uptick in cold calls / approaches for consulting gigs related to data labeling and data QA, in my field (work as an analyst). I never got requests like that 2++ years ago.
- the_brin92 1y agoI've been doing this for one of the major companies in the space for a few years now. It has been interesting to watch how much more complex the projects have gotten over the last few years, and how many issues the models still have. I have a humanities background which has actually served me well here as what constitutes a "better" AI model response is often so subjective. I can answer any questions people have about the experience (within code of conduct guidelines so I don't get in trouble...)
- merksittich 1y agoThank you, I'll bite. If within your code of conduct: - Are you providing reasoning traces, responses or both? - Are you evaluating reasoning traces, responses or both? - Has your work shifted towards multi-turn or long horizon tasks? - If you also work with chat logs of actual users, do you think that they are properly anonymized? Or do you believe that you could de-anonymize them without major efforts? - Do you have contact to other evaluators? - How do you (and your colleagues) feel about the work (e.g., moral qualms because "training your replacement" or proud because furthering civilization, or it's just about the money...)?
- deleted 1y ago[deleted]
- mNovak 1y agoCurious how one gets involved in this, and what fields they're seeking?
- kristianp 1y agoThey do advertise. E.g. outlier.ai : https://app.outlier.ai/en/expert/opportunities?location=All&type=All https://app.outlier.ai/en/expert/opportunities?location=All&...
- dbmikus 1y agoWhat kinds of data are you working on? Coding? Something else? I've been curious how much these AI models look for more niche coding language expertise, and what other knowledge frontiers they're focusing on (like law, medical, finance, etc.)
- rnxrx 1y agoIt's only a matter of time until private enterprises figure out they can monetize a lot of otherwise useless datasets by tagging them and selling (likely via a broker) to organizations building models. The implications for valuation of 'legacy' businesses are potentially significant.
- htrp 1y agoAlready happening.
- some_random 1y agoBad data has been such a huge problem in the industry for ages, honestly a huge portion of the worst bias (racism, sexism, etc) stems directly from low quality labelings.
- htrp 1y agoStarting a data labeling company is the least AI way to get into AI.
- glitchc 1y agoSome people sell shovels, others the grunts to use them.
- scotty79 1y agoTraining data should be open. Time to abolish copyright. Using any data for the purposes of training neural net and publishing data that was used for this purpose should be exempted from copyright protections. If you want beef with people consuming your content without license you should go after them individually or after people who sell them your content. But hands off the modern engine of progress. The entire reason for copyright was to promote progress. The moment it becomes obstacle it should go away. No one is entitled to their legacy business model.