3 ms·
About 8 years ago, we used mturk for reading data out of public PDFs generated by a huge range of producers. I am not sure that LLMs would have been able to con
by ealready_value 1mo ago
About 8 years ago, we used mturk for reading data out of public PDFs generated by a huge range of producers. I am not sure that LLMs would have been able to consistently extract this data until recently as a good number of these PDFs were scans, sometimes a scan of scan.
We had pretty good luck in the end, but getting there produced a somewhat large app on our side that would manage the whole process, including but not limited to asking for multiple responses, comparing them to each other, finding consensus, and determining which users would consistently produce bad responses and stop them from responding. We got to a confidence that about 85-95% of the data was correct, which was good enough for the company.
Through that process, I learned a couple things about managing mturk, primarily about how changes to the cost-per-task would change the process. Initially, we thought that price would be a quality knob, but quickly learned that price was a speed know. The higher the price the faster the tasks would be taken and completed. Quality did not change significantly as the price went up or down.
Overall, I still have fondness to mturk, but it was really bare-bones experience that needed a lot of work to get working effectively.