4 ms·
Show HN: An Agentic AI dataset for deepfake detection
Hi HN!
Really excited for our first dataset release. We're co-founders at Mendit.ai, a small startup in the PNW, we focus on customized and explainable AI models. One thing we noticed as we were building our first product was the need for a curated dataset with diverse resolutions, lighting conditions and scenarios resembling mundane, everyday depictions of characters and objects.
We tried a number of agentic ai frameworks and landed on our own lightweight implementation that first identifies the content properties that are relevant for the use case defined in the prompt and then randomly generates content based on those properties.
An interesting takeaway from our experiments is that the latest LLMs that include reasoning in their response tend to generate noisier outputs than instruction tuned LLMs that generate the output without an in depth plan and explanation.
The dataset can be found on huggingface at: https://huggingface.co/datasets/ninamoss/sleeetview_agentic_ai_dataset https://huggingface.co/datasets/ninamoss/sleeetview_agentic_...
Keep in mind this is only a first sample, we intend to add new content to the dataset based on community feedback on a regular basis.
Looking forward to any feedback, pointers or insights!
- maalber 2y agoThis looks interesting. Quick comment on the format of the dataset though: I would suggest placing both the image and the segmentation mask as the same row with meaningful column names. I think this would make it easier to use. I think it could also be a little more clear from the description how this can be used for deepfake detection. As a question, if I understood correctly, you automatically generate the segmentation masks. Did you do any sort of evaluation/validation of the quality of these?
- ninamoss 2y agoThanks for the valuable feedback, we really appreciate it! We'll make sure to update the format and change the description to explain how it's useful for deepfake detection. In terms of QA on segmentation masks, we did do a manual review of the output and removed ~15% of the original images that had invalid or empty segmentation masks, interestingly we found out this way that most segmentation models tend to generate incomplete or inaccurate masks when the aspect ratio of the input is not 1:1 . We're writing some code snippets that use multiple models to vote on the segments based on semantic class that is generated by the model, we plan to use this to speed up our assessment so we can release more image samples in the future, we could turn this into an open source library if that sounds useful. Happy to stay in touch!
- maalber 2y agoThank you for elaborating, and interesting insight regarding the aspect ratio! I would be interested to know a bit more about the manual review process, e.g. how the invalid masks were classified and how many hours were dedicated to this. At Rapidata, we have/are developing an API for easily creating and getting responses on small validation or annotation tasks from annotators around the world. On-demand and in near real-time. Maybe some form of collaboration could be interesting if you need similar validation in the future. Feel free to reach out, my email is in my profile. We also recently published a dataset to huggingface which we collected using our system: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback https://huggingface.co/datasets/Rapidata/text-2-image-Rich-H...