5 ms·
I actually work on this project! Feel free to ask any questions. Also, here’s an arxiv link to one of the papers if anyone is interested: https://arxiv.org/abs
by RaisinLoaf69 5y ago
I actually work on this project! Feel free to ask any questions.
Also, here’s an arxiv link to one of the papers if anyone is interested:
https://arxiv.org/abs/2108.00275 https://arxiv.org/abs/2108.00275
- mlajtos 5y agoQuite fascinating paper. How did you come up with this twin architecture? Can this effect be simulated in software?
- RaisinLoaf69 5y agoThank you! Definitely a group effort. Two big things that contributed to the choice: We knew we didn’t want the system to have to store memory on each edge so that eliminated a lot of our options. In our experimental learning rule we only had to compare voltages which is actually easier to do if we have two simultaneous networks.
- mlajtos 5y agoInteresting, I haven't seen this twin approach anywhere in ANNs. (I know about Barlow twins and Siamese nets, but this is different.) How did you decide on the topology of the network/graph?
- RaisinLoaf69 5y agoSince we have no global processor, each edge is changing using only local information, a lot of stuff that makes sense in our network doesn't really make sense in ANNs and vice versa. For example, in our newest paper (https://arxiv.org/abs/2201.04626 https://arxiv.org/abs/2201.04626) we desynchronize the updates of our edges. Instead of changing the entire system all at once, we change random parts of it each training step. This doesn't really make sense to do in an ANN where you require global information for every edge update. The shape of the network is actually inspired by jamming solids (we're a soft matter lab), but is completely arbitrary. We've done a ton of different shapes and sizes in simulation.
- aperrien 5y agoWhat type of component is an "adjustable resistor"? Looking the term up online only shows potentiometers, which I really don't think is what is being described here. Is this some sort of memristor?
- RaisinLoaf69 5y agoIt is in fact just a digital potentiometer. In the experiment shown in the paper I linked we’re using a 128 position one. In new work we’ve actually shifted to using transistors which are better for a number of reasons (smaller, faster, nonlinear, continuous).
- dekhn 5y agoThe Mark I perceptron used physical (analog) potentiometers that could rotate themselves.
- sitkack 5y agoFascinating! You can get digital pots in lots of different configurations. https://www.digikey.com/en/products/detail/microchip-technology/MCP4451-103E-ST/2601449 https://www.digikey.com/en/products/detail/microchip-technol... https://en.wikipedia.org/wiki/Perceptron https://en.wikipedia.org/wiki/Perceptron https://americanhistory.si.edu/collections/search/object/nmah_334414 https://americanhistory.si.edu/collections/search/object/nma...
- jareklupinski 5y agoI picked up a couple banks of memristors to replace the pots in your circuit :) Hope to create a small feedback circuit across each memristor, essentially letting it 'train itself'
- RaisinLoaf69 5y agoSweet! The basic learning comparison (Vc-Vf) can be implemented in tons of physical systems (memristors, springs, water pipes). So seeing it in other mediums would be pretty cool.
- detaro 5y ago> I picked up a couple banks of memristors Are those now something just available off the shelf?
- igorkraw 5y agoOn mobile so didn't have time skim the paper: is it a lienar layer trained with basically hebbian learning? If not, how do you handle backpropagation/credit assignment? How would you scale this to 100 million parameters if you had to?
- RaisinLoaf69 5y agoThis is strictly not a neural network, so there is no backpropagation. Credit assignment is done on each edge using a local rule (Eg. using only its current state and the state of touching edges). To scale the network you just have to add more edges (no limit on the amount). We have a design for a tiny version of this network using transistors that could have order 10^6 edges on the size of a microchip.
- igorkraw 5y agoI've looked over the paper now, unless I'm misunderstanding this seems very similar to the general trend of hebbian learning/STDP/local predictive coding/teacher forcing (for those unfamiliar, these are all distinct but the basic idea is always to have a signal adjust based on the difference with some local target. Hebbian learning is the basic "fire together wire together" principle, STDP is a specific instantiation that works with specific types of memristors, teacher forcing comes from RNN training and imposes the ground truth input on intermediate step, local predictive coding I can't recall the precise thing but but basically diffuses a local output error through a network similar to what is done on a single layer here [which can actually approximate backpropagation! It's very cool]). How would you differentiate yourself against this/what would you say is the core benefit of this approach?
- readingnews 5y agoOK, I looked over your paper. Could I actually build this from your paper and your single edge node circuit? I am not sure yet, I did not read it three times yet (typically, have to read it multiple times)... but my first impression is that I could not reproduce the papers conclusions. I _feel_ like something is missing. Are you holding off until your provisional patent is approved, or the like? I feel like there is some connecting device/circuitry/something left out... As you mentioned near the end, you do NOT overcome the bias in the AD5220s? You just accept the error floor??
- RaisinLoaf69 5y agoWe’re not holding anything back, you should be able to recreate our findings from the paper. More broadly, there are a number of ways to recreate the network using the relatively simple learning rule we provide. In theory, this learning rule will continue to decrease your error forever (in our simulations our error goes down to machine precision). However, with any physical learning network you’re always gonna hit an error floor based on the precision of your components. With our current variable resistors and network size that floor is around 10^-3. With more precise components (like we mention at the end of the paper), that floor will go down significantly.