4 ms·
Although it's hard to tell from the images presented with the article, the face generation looks like it could be similar to the techniques used in Nishimoto et
by chch 13y ago
Although it's hard to tell from the images presented with the article, the face generation looks like it could be similar to the techniques used in Nishimoto et al., 2011, which used a similar library of learned brain responses, though for movie trailers:
http://www.youtube.com/watch?v=nsjDnYxJ0bo http://www.youtube.com/watch?v=nsjDnYxJ0bo
Their particular process is described in the YouTube caption:
The left clip is a segment of a Hollywood movie trailer that the subject viewed while in the magnet. The right clip shows the reconstruction of this segment from brain activity measured using fMRI. The procedure is as follows:
[1] Record brain activity while the subject watches several hours of movie trailers.
[2] Build dictionaries (i.e., regression models) that translate between the shapes, edges and motion in the movies and measured brain activity. A separate dictionary is constructed for each of several thousand points at which brain activity was measured.
(For experts: The real advance of this study was the construction of a movie-to-brain activity encoding model that accurately predicts brain activity evoked by arbitrary novel movies.)
[3] Record brain activity to a new set of movie trailers that will be used to test the quality of the dictionaries and reconstructions.
[4] Build a random library of ~18,000,000 seconds (5000 hours) of video downloaded at random from YouTube. (Note these videos have no overlap with the movies that subjects saw in the magnet). Put each of these clips through the dictionaries to generate predictions of brain activity. Select the 100 clips whose predicted activity is most similar to the observed brain activity. Average these clips together. This is the reconstruction.
With the actual paper here:
http://www.cell.com/current-biology/retrieve/pii/S0960982211009377 http://www.cell.com/current-biology/retrieve/pii/S0960982211...
- anigbrowl 13y agoForgot about this, thanks for re-posting.
- JackFr 13y agoThanks for posting that. I count myself as enormously skeptical of TFA research, but that paper appears to be quite good. I may need to re-evaluate my biases. On the other hand, this is presented as in the press release as mind reading, but the reality is more like trying design something similar to a cochlear implant.
- yeukhon 13y agoSelect the 100 clips whose predicted activity is most similar to the observed brain activity. Average these clips together. This is the reconstruction. Mind to explain to me why the right clip is inconsistent with the left clip? https://www.youtube.com/watch?v=nsjDnYxJ0bo https://www.youtube.com/watch?v=nsjDnYxJ0bo Starting at 20-second I feel like the right clip is out of touch with the left clip. between 20th and 22nd second I see at least three individuals rendered from the reconstruction. From 26th to the end of the clip I also see multiple individuals. The names also look different from one another... When you say find the closest is that an expected result?
- chch 13y agoI'll give the disclaimer that this paper isn't in my field, and I'm merely an observer. However, I'll do my best to explain, since it's a little unclear. Based on my perspective, there were three sets of videos: 1) The several hours of "training" video, that they used to learn how the test subject's brain acted based on different stimuli. (The paper (which I've only skimmed) says 7,200 seconds, which is two hours) 2) 18,000,000 individual seconds of YouTube video that the test subject has never seen. 3) The test video, aka the video on the left. So, the first step was to have the subject watch several hours of video (1), and watch how their brain responded. Then, using this data, they predicted a model of how they thought the brain would respond for eighteen million separate one second clips sampled randomly from YouTube (2). They didn't see these, but they were only predictions. As an interesting test of this model, they decided to show the test subject a new set of videos that was not contained in (1) or (2), the video you see in the link above, (3). They read the brain information from this viewing, then compared each one second clip of brain data to the predicted data in their database from (2). So, they took the first one second of the brain data, derived from looking at Steve Martin from (3), then sorted the entire database from (2) by how similar the (predicted) brain patterns were to that generated by looking at Steve Martin. They then took the top 100 of these 18M one second clips and mixed them together right on top of each other to make the general shape of what the person was seeing. Because this exact image of Steve Martin was nowhere in their database, this is their way to make an approximation of the image (as another example, maybe (2) didn't have any elephant footage, but mix 100 videos of vaguely elephant shaped things together and you can get close). They then did this for every second long clip. This is why the figure jumps around a bit and transforms into different people from seconds 20 to 22. For each of these individual seconds, it is exploring eighteen million second-long video clips, mixing together the top 100 most similar, then showing you that second long clip. Since each of these seconds has its "predicted video" predicted independently just from the test subject's brain data, the video is not exact, and the figures created don't necessarily 100% resemble each other. However, the figures are in the correct area of the screen, and definitely seem to have a human quality to them, which means that their technique for classifying the videos in (2) is much better than random, since they are able to generate approximations of novel video by only analyzing brain signal. Sorry, that was longer than I expected. :) Edit: Also, if you see the paper, Figure 4 has a picture of how they reconstructed some of the frames (including the one from 20-22 seconds), by showing you screenshots whence the composite was generated.