8 ms·
I'm one of the people building this. Hi HN, AMA :)
by 0xrisk 4y ago
I'm one of the people building this. Hi HN, AMA :)
- telotortium 4y agoWhat visual search engine are you using?
- simandl 4y agoThis is clip matching the text or image searched to the images in the Laion-5B dataset.
- tener 4y agoPrivacy policy for searched strings? "Rutkowski" returns a bunch of book covers, repeated a lot. Can you ensure images returned have diverse embeddings? I expected digital art, not detective stories. Do you use CLIP or just metadata? What is intended process that starts after you get artist's email (for either of purposes).
- simandl 4y agoIt's using clip to match the text to the image, so you can actually prompt it like you might an art generator. Here's "in the style of greg rutkowski": https://haveibeentrained.com/?search_text=in%20the%20style%20of%20greg%20rutkowski https://haveibeentrained.com/?search_text=in%20the%20style%2... In the next few weeks we'll be adding the ability to log in and flag or upload your works (if they aren't there). Those lists will have permissions assigned to them, starting with simple opt-in or opt-out.
- freefruit 4y agoHow do you have the rights to use any of the images? They are clearly the same as what a Google image search would result in. However google links to the source.
- simandl 4y agoIf that's a rights issue, we'll definitely add a link to the source. For now, you can right click -> open in new tab to see where it came from, but we'll look into this asap. The goal here is to give people the opportunity to remove images they don't want in this dataset or add images they do want in there.
- NicoleJO 4y agoThat's not how this works. Copyright owners have the right to control when, where, how and by whom their content may be used. Not you https://www.law.cornell.edu/uscode/text/17/106 https://www.law.cornell.edu/uscode/text/17/106
- deleted 4y ago[deleted]
- r-k-jo 4y agohi thanks for building this! Could you also enable a simple search by matching exactly words from the caption text? rather than semantic similarity?
- simandl 4y agoThanks! That is definitely on the list, but might be a few months away. We're focusing on using images to find other images so it will be easy for artists to flag all of their stuff quickly. But, once we have that in a good place, we will definitely be adding more to the text search side.
- kernelsanderz 4y agoThis sounds like a great initiative. I can't find anything on the site about a privacy policy, how the email addresses you're collecting will be used. If you're creating this for the AI community, then how will this data be made available to others?
- simandl 4y agoThese are important questions. We aren't storing any of the images used to search, and the email addresses are going to mailchimp lists for opt in and opt out. As we role out the next set of tools, which let users flag images and create lists from them, we'll email the mailchimp lists with more info. When we enable sign-in, we'll also add a privacy policy, because at that point, we will store some images, on request, to use them to make finding other works by the same artists easier. Opt-out image URL lists will be made available to the dataset owners for removal. Opt-in image lists will be public.
- moontear 4y agoAny way to search by URL / part of the URL?
- oth001 4y agoCool tool but don't you think the onus should be on Emad & team (to use a dataset with only public domain and licensed inages) rather than forcing every artist to opt-out?
- infinityio 4y agoObviously it might be a slightly futile task given the size of the dataset, but would there be value in adding community captions to some of the images? I'd be happy to spend 15 minutes captioning the worst-labelled (maybe cosine similarity between stabdiff im2txt and the prompt?) images in the set / reviewing other people's captions, and if you can get enough people on board you could probably get through a not-insignificant number of new captions. Equally a risk of anti-ai groups sabotaging this process though