4 ms·
Mechanical Turk is great for "open", public research. We used to use them a lot for machine learning tasks (data cleanup, model comparisons, label annotations),
by Radim 8y ago
Mechanical Turk is great for "open", public research. We used to use them a lot for machine learning tasks (data cleanup, model comparisons, label annotations), along with similar services like CrowdFlower / Figure Eight. We saw two primarily issues when applied to "non-open" (commercial) projects:
- business-related data too sensitive to share with strangers (contractual obligations, too much risk)
- some tasks required non-trivial subject matter expertise and context to annotate properly (quality control issues)
For this reason, we gradually moved to an in-house team of long-term annotators. It's not much more expensive (moms on maternity leave, students…), but infinitely more flexible and safer for our purposes. YMMV.
- therealmarv 8y agotry PYBOSSA, open source crowdsourcing platform, you host it yourself, you are in control of your data.
- k__ 8y agoI think the problem is more that you give the data to the turkers than to give it to Amazon.
- mrgordon 8y agoCrowdFlower / Figure Eight works with a lot of annotators who have signed NDAs and work in secure locations where they can't duplicate materials etc.
- mwexler 8y agoSeconded. Quality control is a nightmare with Turk, often requiring stimuli to be labeled multiple times and have a variety of judging approaches to "crown a winner". Companies like DefinedCrowd (https://www.definedcrowd.com/ https://www.definedcrowd.com/) have taken a quality-first approach which gives much better data, but of course, at a cost.
- ayw 8y agoAlex from Scale (www.scaleapi.com) here! We've taken an extremely quality-first approach and build out large workforces for datasets with high quality requirements and complexity. For example, we do a bunch of LIDAR / 3D labeling (https://www.scaleapi.com/sensor-fusion-annotation https://www.scaleapi.com/sensor-fusion-annotation) which is very complex and labor intensive, and provide extremely high quality that would not be possible otherwise.
- cwheatley 8y agoI work directly with several top-tier research universities use us for public opinion polling mostly, as we offer a representative sample at a great price point. Lucid is the creator of the world's largest programmatic survey sample marketplace. We have API integrations with hundreds of US and international panels that allow us to target very specific audiences across dozens of sources, ensuring excellent feasibility. Happy to discuss any needs the YC community has. - Cullen Wheatley Articles for more info: https://techcrunch.com/2017/04/10/lucid-60-million/ https://techcrunch.com/2017/04/10/lucid-60-million/ https://greenbookblog.org/2017/05/16/ceo-series-how-lucid-ceo-patrick-comer-plans-to-transform-insights/ https://greenbookblog.org/2017/05/16/ceo-series-how-lucid-ce...
- codydh 8y agoMoving to in-house annotators is probably the smart strategy. However, for tasks that are easily done by naive annotators, Prolific.ac might be a great source. I've had good luck getting quality data there, and they enforce a minimum hourly wage that, while not truly livable, is still heading in the right direction.
- deepakputhraya 8y agoDeepak from Playment(https://playment.io/ https://playment.io/). We help companies label data. We have dedicated managers for projects who take care of training annotators and ensuring quality output. We have a large number of annotators who are pre-trained on different annotations, and we train them for different specific business use-cases. Reach out to us if you are looking for fully managed quality annotated data. Links: https://playment.io/image-annotation/ https://playment.io/image-annotation/ https://playment.io/playment-vs-mechanical-turk/ https://playment.io/playment-vs-mechanical-turk/ https://playment.io/playment-vs-crowdflower/ https://playment.io/playment-vs-crowdflower/