3 ms·
Another approach to neural-based data representation / encoding / decoding is Implicit Neural Representations (INRs). For INRs, you overfit a neural network, ty
by stefanpie 3y ago
Another approach to neural-based data representation / encoding / decoding is Implicit Neural Representations (INRs). For INRs, you overfit a neural network, typically a Multilayer Perceptron (MLP)[1], to a single data point. For instance, in the case of a single image (the data point), the inputs would be the xy pixel coordinates—usually scaled from [0, n_pixels]^2 to either [0, 1]^2 or [-1, 1]^2—and the outputs would be the RGB values of each pixel. Once trained on pixels from this singular image, the INR model essentially becomes a representation of that data point. The image can then be re-rendered or reconstructed by inputting all the pixel coordinates and receiving the corresponding RGB values. This approach even allows for sampling at intermediate coordinates or a subset of coordinates for partial, progressive, or super-resolution decoding.
This technique is generally effective as long as you can specify a reasonable coordinate system for your data, whether it's audio (n_samples, n_channels), 3D models (x, y, z)(x, y, z, θ, φ) [2], geospatial data (lat, lon, altitude), and so on. Interestingly, it can also be applied even if there isn't a well-defined or non-Euclidean coordinate system for your data [3].
Additionally, you can meta-learn an initial set of weights through meta-learning techniques tailored for a particular distribution of data points [4]. This enables you to quickly fit an INR model to a new data point that comes from the same distribution, such as microscope images of cells, MRI scans, or scenes for self-driving cars.
Most pertinent to the post, INRs can be compressed and quantized to achieve comparable or even superior performance to traditional data compression techniques, at various quality levels for different modalities [5]. By no means is it competitive for long standing approaches like for image compression but it works well nonetheless. This is an area that we have explored extensively in my lab, enabling end-to-end encoding, processing, and decoding of INRs in hardware for a variety of applications.
The most interesting topic to me is the fact that you can edit the weights of the INR to edit the underlying data [6][7], or use the weights of the INR to compute tasks on the data without ever having to decode it. For example, I can perform image classification on the weights of the INR model rather than on the actual images. This classification model is a hypernetwork that takes the INR parameters as inputs instead of the RGB image. These hypernetworks can also be designed in a way that preserves the invariance and equivariance properties of MLP structures, which is a very cool idea to me [8].
[1] https://arxiv.org/abs/2006.09661 https://arxiv.org/abs/2006.09661
[2] https://www.matthewtancik.com/nerf https://www.matthewtancik.com/nerf
[3] https://arxiv.org/abs/2205.15674 https://arxiv.org/abs/2205.15674
[4] https://www.matthewtancik.com/learnit https://www.matthewtancik.com/learnit
[5] https://arxiv.org/abs/2201.12904 https://arxiv.org/abs/2201.12904
[6] https://arxiv.org/abs/2201.12204 https://arxiv.org/abs/2201.12204
[7] https://arxiv.org/abs/2210.08772 https://arxiv.org/abs/2210.08772
[8] https://developer.nvidia.com/blog/designing-deep-networks-to-process-other-deep-networks/ https://developer.nvidia.com/blog/designing-deep-networks-to...