4 ms·
Very interesting watch. What sort of NN would be used to warp two compositionally similar images to match each other in structure? In the example the copy pain
by defaultname 5y ago
Very interesting watch.
What sort of NN would be used to warp two compositionally similar images to match each other in structure? In the example the copy painting had almost all of the same general components, laid out with fairly significant differences, and they warped it to match quite precisely.
- danbruc 5y agoWithout a neural network one would probably use optical flow [1] and I guess you could train a neural network to predict optical flow and then use that flow to align the two inputs. I have no answer to your specific question, what kind of neural network architecture would be best suited but given the task it seems reasonable to assume some form of convolutional neural network. [1] https://en.wikipedia.org/wiki/Optical_flow https://en.wikipedia.org/wiki/Optical_flow
- jvanderbot 5y agoI'm taking your question without context, so sorry if it doesn't make sense: Warping images to match based on structure / features is common in much of image processing and requires no NN since we know how to do it. (But may be enhanced by NN). The basic idea is: 1. Find features in images A and B (SURF / SIFT are common). However, maybe a NN could find interesting features here? 2. Match features so you know what feature from A is where in B. Again, a NN could potentially do this "better". 3. Compute a homography (translation between planes of features) or procrustes transform (or other types of transformations) 4. For each pixel (p) in the target frame (B), take the inverse of the transformation to find the pixel of the original frame (A) (which is usually a fractional pixel because it won't line up) 5. interpolate between nearby pixels in (A) to get the "true" pixel value at (p) in (B). Perhaps a NN could improve this interpolation? We've seen this in super resolution, for example. You can easily reproduce these steps with just a few function calls in OpenCV. Try it out. https://opencv.org/ https://opencv.org/ Up to step 3 is very helpful to answer questions like: "How much did the camera move between captures of image A and B. So now you can solve structure from motion, visual odometry, etc. This is typically how quadrotors navigate with cameras.
- danbruc 5y agoWould SIFT or SURF features actually work in this scenario, that was also one of my thoughts, but then I considered that the inputs are not photographs of the same objects from different angles but a painting and a manual reproduction of it. I would be really interested in knowing whether those features are robust enough under the distortions in such a scenario. And how far could one push this, could you still align the Mona Lisa with a good pencil drawing of it?
- jvanderbot 5y agoI have no idea! Careful feature selection is very important. If you only have edge information, then extracting edges in both images seems like a reasonable pre-processing step.