7 ms·
I immediately found the results suspect, and think I have found what is actually going on. The dataset it was trained on was 2770 images, minus 982 of those use
by Aransentin 4y ago
I immediately found the results suspect, and think I have found what is actually going on. The dataset it was trained on was 2770 images, minus 982 of those used for validation. I posit that the system did not actually read any pictures from the brains, but simply overfitted all the training images into the network itself. For example, if one looks at a picture of a teddy bear, you'd get an overfitted picture of another teddy bear from the training dataset instead.
The best evidence for this is a picture(1) from page 6 of the paper. Look at the second row. The building generated by 'mind reading' subject 2 and 4 look strikingly similar, but not very similar to the ground truth! From manually combing through the training dataset, I found a picture of a building that does look like that, and by scaling it down and cropping it exactly in the middle, it overlays rather closely(2) on the output that was ostensibly generated for an unrelated image.
If so, at most they found that looking at similar subjects light up similar regions of the brain, putting Stable Diffusion on top of it serves no purpose. At worst it's entirely cherry-picked coincidences.
1. https://i.imgur.com/ILCD2Mu.png https://i.imgur.com/ILCD2Mu.png
2. https://i.imgur.com/ftMlGq8.png https://i.imgur.com/ftMlGq8.png
- kdma 4y agoGood find, when I read it I called bullshit but I got lost trying to understand the diagrams. Another gotcha is the semantic decoder, they are just looping the model on itself "A cozy teddy bear" + fMRI random input => A teddy bear!!!
- SubiculumCode 4y agoI feel like you might be moving the goal posts here a bit. Getting a reconstruction that is a bear, even if not the same bear, is impressive enough to be noteworthy.
- xmonkee 4y agoI think the point is that it's not a reconstruction. It's more like recognizing which letter of a thousand-letter alphabet is shown to the human after decoding their brain waves. Still impressive, but not really as impressive as visual reconstruction.
- groestl 4y agoTBH, I was not impressed up until now, but given the videos I have in mind from people trying to use brain computing interfaces to type a text, now I'm impressed.
- pedrosorio 4y agofMRI is useless for that purpose - latency is much higher than any BCI method you might’ve seen in those videos
- bawolff 4y agoEven if true, the result still seems very impressive to me as a layman.
- gfaure 4y agoThat’s the whole problem — that the reconstruction aspect of the contributions seems overstated given only a layperson’s understanding.
- Hakkin 4y agoI'm definitely not an expert in this subject, but even if the model is overfitted, doesn't the fact that it can pull out the similar images at all give credit to the idea that a larger, non-overfitted model could actually work as the paper describes? It means that there does exist some correlation between the shown subject, the captured fMRI data, and the resulting location in latent space.
- Double_a_92 4y agoThe output part is basically nonsense. It would be more honest if the output was a text. E.g. "Teddybear" instead of a bad image of a random teddybear.
- Hakkin 4y agoIn this specific case I agree, since the model may be overfitted, it seems like it's currently just a glorified object classifier based on what was in the training data, but the fact that it works at all may indicate that the underlying idea has merit. They would probably have to train a much larger network to see if it's able to separate features distinctly enough using the input fMRI data to be useful.
- gus_massa 4y agoThe problem is that it's impossible to know what is in the fMRI data and what is hallucinated by the reconstruction. In this case, the real bear has a blue ribbon and the "reconstructed" bear ha a red ribbon. Is the ribbon in the fMRI data and the computer choose the wrong color, or most of the images in the training set had ribbons and the computer just added one. Imagine this something like this is used in the future to get something like https://en.wikipedia.org/wiki/Facial_composite https://en.wikipedia.org/wiki/Facial_composite . People may give too much importance to the details and arrest someone only because the computer imagined some detail, like the logo in the baseball cap.
- deleted 4y ago[deleted]
- arnarbi 4y agoSubject 4 in the first line also looks very different from the ground truth, but clearly an airliner. I'm curious if there is also a closer match to that one in the set.
- 2-718-281-828 4y agothere is also no way that you could represent details as shown with such a small sample.
- ditchfieldcaleb 4y agoWhat are you talking about? They didn't train a model for this. That's why it's so impressive.
- Hakkin 4y agoQuoting from the paper, The only training required in our method is to con- struct linear models that map fMRI signals to each LDM component, and no training or fine-tuning of deep-learning models is needed. ... To construct models from fMRI to the components of LDM, we used L2-regularized linear regression, and all models were built on a per subject basis. Weights were estimated from training data, and regularization parame- ters were explored during the training using 5-fold cross- validation.
- ditchfieldcaleb 4y agoAh. This makes more sense. Thanks.
- brucethemoose2 4y agoIts still picking out the correct "overfitted" images, which is remarkable. Theoretically, the results would scale to more training images... we just need to fMRI all of LAION-5B. Easy peasy.
- mkagenius 4y agoThe only question is whether more images will confuse the model or not?
- razor_router 4y agoWhat evidence do you have that this technique is overfitting the training data rather than reading the brain?
- sillysaurusx 4y agoI don’t get the criticism here. Normally I’d be the first to err on the side of skepticism, but this work seems above board. I think the confusion is that this model is generating “teddy bear” internally, not a photo of a teddy bear. I.e. the diffusion part was added for flair, not to generate the details of the images that exist inside your mind. They could just as easily have run print(“teddy bear”), but they’re sending it to diffusion instead of printing it to console. The fact that it can correctly discern between a dozen different outputs is pretty remarkable. And that’s all that this is showing. But that’s enough. It’s not really a “gotcha” to say that it’s showing an image from the training set. They could have replaced diffusion with showing a static image of a teddy bear. It sounds like this is many readers’ first time confronting the fact that scientists need to do these kinds of projects to get funding. As long as they’re not being intentionally deceptive, it seems fine. There’s a line between this and that ridiculous “rat brain flies plane” myth, and this seems above it. Disclaimer: I should probably read the paper in detail before posting this, but the criticism of “the building looks like a training image” is mostly what I’m responding to. There are only so many topics one can think about, and having a machine draw a dog when I’m thinking about my dog Pip is some next-level sci-fi “we live in the future” stuff. Even if it doesn’t look like Pip, does it really matter? Besides, it’s a matter of time till they correlate which parts of the brain are more prone to activating for specific details of the image you’re thinking about. Getting pose and color right would go a long way. So this is a resolution problem; we need more accurate brain sampling techniques, i.e. Neuralink. Then I’m sure diffusion will get a lot more of those details correct.
- Aransentin 4y agoBecause pretty much everybody that reads the article will have taken away a grossly exaggerated idea of what the system is actually capable of. If Stable Diffusion was intentionally added "for flair" and really is unnecessary, then I would absolutely say that the researchers were being intentionally deceptive. Even if we do a massive goalpost-move and grant that the system is only identifying the label "dog" with a brain scan of a person looking at a dog, we would need to see actual statistics of its labelling accuracy before judging it in that way. If the images in the paper are cherry-picked(1), it could easily be only able to extract a handful of bits to no bits at all, and the entire thing could very well turn out the be replicable from random noise. (1) Note that the paper even states "We generated five images for each test image and selected the generated images with highest PSMs [perceptual similarity metrics].", so it even directly admits that the presented images are cherry-picked at least once.
- sampo 4y ago> The dataset it was trained on was 2770 images, minus 982 of those used for validation. I don't think you got that 2770 correct. Might be 9250 images, minus 982 (that one you got right). Then again, the paper is so badly written, I find it difficult to decipher what they did. From section 3.1: Briefly, NSD provides data acquired from a 7-Tesla fMRI scanner over 30–40 sessions during which each subject viewed three repetitions of 10,000 images. We analyzed data for four of the eight subjects who completed all imaging sessions (subj01, subj02, subj05, and subj07). We used 27,750 trials from NSD for each subject (2,250 trials out of the total 30,000 trials were not publicly released by NSD). For a subset of those trials (N=2,770 trials), 982 images were viewed by all four subjects. Those trials were used as the test dataset, while the remaining trials (N=24,980) were used as the training dataset. https://www.biorxiv.org/content/10.1101/2022.11.18.517004v2.full.pdf https://www.biorxiv.org/content/10.1101/2022.11.18.517004v2....