Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
simandl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
1.
▲
by
simandl
4y ago
Last year's ICLR had a paper, "Data Poisoning Won't Save You From Facial Recognition" that included the Glaze team's previous project, Fawkes. This statement from that paper is quite damning. "This paper shows
2.
▲
by
simandl
4y ago
This assumes that the filter actually works in practice: https://www.reddit.com/r/StableDiffusion/comments/11v7sv9/ha...
3.
▲
by
simandl
4y ago
Thank you for the feedback! That's what we're hoping our opt-in tools will help artists do. For the problems you've posed here, you might actually want to both opt-in and opt-out. You'd be able to flag the imageurl-capti
4.
▲
by
simandl
4y ago
We think this is because the images are all links and the browser itself is pulling them in from across the web. With chrome it happened once during our testing, and we've had one user also experience it, but it's intermittent. Th
5.
▲
by
simandl
4y ago
These are important questions. We aren't storing any of the images used to search, and the email addresses are going to mailchimp lists for opt in and opt out. As we role out the next set of tools, which let users flag images and creat
6.
▲
by
simandl
4y ago
They have their own datasets and included Laion-400M, a subset of 5b that was released prior to 5b. You can see a short explanation in imagen's "Limitations and Societal Impact" section at: https://imagen.research.
7.
▲
by
simandl
4y ago
We (Spawning) did not create the dataset or train the models in question. We're working to make it easy for people to remove themselves from, or add themselves to, this dataset and future models.
8.
▲
by
simandl
4y ago
This article highlights several of the problems that we are working on. Here's another article, about our organization: https://www.inputmag.com/culture/mat-dryhurst-holly-herndon-...
9.
▲
by
simandl
4y ago
Reading through your comments, it appears you're under the impression that we (Spawning) trained models using this data. That is not the case. We're building tools to help people to remove themselves from, or add themselves to, da
10.
▲
by
simandl
4y ago
Our initial approach will be to validate the artists manually, and trust them to only flag or upload their own works. If we hit a scale where that starts to become an issue, we have plenty of ideas for mitigating this, but we're not op
11.
▲
by
simandl
4y ago
You might be surprised. We have almost as many opt-in requests as opt-outs since we announced this today. We don't see this as binary in the long term. Maybe artists want to release art from a prior period, for example, but withhold th
12.
▲
by
simandl
4y ago
We don't store any images used for searching. We are building an opt-in list, because a lot of people do want to be able to prompt AI with something like, "a cat in the style of me " or " me riding a dinosaur". Th
13.
▲
by
simandl
4y ago
Thanks! That is definitely on the list, but might be a few months away. We're focusing on using images to find other images so it will be easy for artists to flag all of their stuff quickly. But, once we have that in a good place, we w
14.
▲
by
simandl
4y ago
If that's a rights issue, we'll definitely add a link to the source. For now, you can right click -> open in new tab to see where it came from, but we'll look into this asap. The goal here is to give people the opportunity
15.
▲
by
simandl
4y ago
For sure! And as artists opt in, you'll be able to use it to see how they describe their work.
16.
▲
by
simandl
4y ago
Please do sign up and you'll be able to flag these images soon. We'll work to get them removed from this and future datasets built for AI training.
17.
▲
by
simandl
4y ago
It's using openai's clip ( https://openai.com/blog/clip/ ) to find the image similar to your query or image. Clip learned to match images to the captions that were paired with them from images on the web.
18.
▲
by
simandl
4y ago
Yes, that's the same dataset. This website has some additional tools coming so artists can flag and opt out, or upload to opt in, and we'll get those to the laion team to add or remove from the 5B (and future) datasets.
19.
▲
by
simandl
4y ago
I'm not 100% on this, but I think a big portion of the captions for the images come from the alt-text, and are probably auto-generated by the sites they were scraped from.
20.
▲
by
simandl
4y ago
This is Laion-5B, https://laion.ai/blog/laion-5b/ It's built off of common crawl, so it probably does have a pretty representative sample from whatever the big image searches use. Funny enough, the NSFW filte
21.
▲
by
simandl
4y ago
This is clip matching the text or image searched to the images in the Laion-5B dataset.
22.
▲
by
simandl
4y ago
It's using clip to match the text to the image, so you can actually prompt it like you might an art generator. Here's "in the style of greg rutkowski": https://haveibeentrained.com/?search_text=in%20the%2
23.
▲
by
simandl
4y ago
Stable Diffusion used an aesthetic filter to train on a subset of the English language images from this full 5.8 billion multi-language set. That probably got a lot of what you're finding.
24.
▲
by
simandl
4y ago
This is Laion-5B, you can read more about it here: https://laion.ai/blog/laion-5b/ Imagen and Stable-Diffusion both used subsets of this full 5.8B image set.
25.
▲
Time Series Visualisations: Kibana or Grafana?
(rittmanmead.com)
7 points
by
simandl
10y ago
|
0 comments
26.
▲
by
simandl
12y ago
This is a shiny + d3 side project I'm working on. It's still an MVP and I'd appreciate any feedback on improvements.
27.
▲
Optimize your liquor cabinet using data from three world-class bars
(rittmanmead.com)
4 points
by
simandl
12y ago
|
1 comments