9 ms·
Exploring Weight Agnostic Neural Networks
- phaedrus 7y agoI wrote a series of Markov chat simulators as a teenager. Often I used a simpler algorithm which ignored the probability weight (all out-links, once learned, given equal probability). These version performed subjectively as well as, if not better, than the versions which tracked the weight of links. I'm not surprised therefore that weight agnostic neural networks can work, too.
- meowface 7y agoI think it may not be a great comparison. N-grams (of words) of human speech/writing are way more deterministic than the kinds of things ML usually tries to tackle, I think. If you write the word "because", then "of", "the", or some pronoun are all extremely safe bets for the next word, regardless of their recorded probabilities. I imagine you could also totally randomize the probabilities and not see any issues. But I'm no expert and hardly even an amateur, so maybe it is a similar kind of thing here with ML. And I know randomized optimization is a big thing in ML, though I'm not sure to what extent that could be analogized with randomizing Markov model probabilities.
- s_Hogg 7y agoI'm pretty sure this was posted a while back (maybe a month or two)?
- mannykannot 7y agoYes, with some insightful and informative comments: https://news.ycombinator.com/item?id=20160693 https://news.ycombinator.com/item?id=20160693
- baylearn 7y agoPrevious discussion (about the actual research article at https://weightagnostic.github.io/ https://weightagnostic.github.io/ rather than the blog post): https://news.ycombinator.com/item?id=20160693 https://news.ycombinator.com/item?id=20160693
- DoctorOetker 7y agohow is this different from boring old evolutionary algorithms? In my opinion the big breakthrough that enabled optimization and machine learning was the discovery of reverse mode automatic differentiation, since the space or family of all possible decision-functions is high dimensional, while the goal (survival, reproduction) is low dimensional. Unless I see a mathematical proof that evolutionary algorithms are as efficient as RM AD, I see little future in it, and apparently neither did biology since it decided to create brains. It's not an ideological stance I take here (of nature vs nurture). For simplicity, lets pretend humans are single-cellular organisms, what does natural selection exert pressure on? our DNA code: both the actual protein codes and the promotor regions. I claim that variation on the proteins are risky (a modification in a proteinn coding region could render a protein useless) while a variation on the promotor regions is much less risky: altering a nucleotide there would slightly affect the affinity modulating transcription, so the cell would behave essentially the same but with different treshold concentrations, think of continuous parameters that describe our body (assuming same nurture, food, etc) some people are a bit taller, some people a bit stronger, etc... so how many of these continuous parameters do we have? On the order of the same number as the total number of promotor regions in DNA in the fertilized egg: both on human DNA and in one mitochondria (assuming there isn't a chemical signals addressing and reading and writing scheme for say 10 mitochondria)... EDIT: just adding that for a certain fixed environment, there are local (and a global) optimum of affinity values for each protein, so that near a local optimum the fitness is roughly shaped like -s(a-a_opt)^2 where s is spread and a_opt the local optimum affinity value. In other words, it is not so that "better affinity", means fitter, not at all, a collection of genomes from an identical environment will hover around an affinity sweet spot. According to wikipedia [0] that would result in about about 2x 20412 "floats" for just protein-coding genes about 2x 34000 "floats" when also including the pseudo-genes about 2x 62000 "floats" when also including long ncRNA, small ncRNA, miRNA, rRNA, snRNA, snoRNA these "floats" are the variables that allow a species to modulate the reaction constants in the gene regulatory network, since natural selection can not directly modulate the laws of physics and chemistry, and modulating the protein directly instead of the promotor region affinities / reaction rates risks disfunctional proteins... so my estimate of an upper limit of the number of "floats" in the genetic algorithm is ~120000 (and probably much less if not each of the above has a promotor region). thats not a lot of information, if we think about the number of synaptic weights in the brain, and many of these are shared in utilization by the other cell types besides neurons. I consider the possibility that: sperm cell, egg cell, or fertilized egg cell performs a kind of POST (power-on-self-test) that checks for some of the genes, although simply reaching the fertilized state may be enough of a selftest so no spontaneous abortion test may be needed (to save time and avoid resources spent on a probably malformed child). [0] https://en.wikipedia.org/wiki/Human_genome#Molecular_organization_and_gene_content https://en.wikipedia.org/wiki/Human_genome#Molecular_organiz... EDIT2: regarding: >This makes WANNs particularly well positioned to exploit the Baldwin effect, the evolutionary pressure that rewards individuals predisposed to learn useful behaviors, without being trapped in the computationally expensive trap of ‘learning to learn’. The computationally expensive trap of having to 'learn to learn' could end up being as mundane as a low number of hormones to which neurons in the brain globally or collectively respond, which enables learning by reward or punishment, and from then on anticipating reward or punishment, and our individual end goal stems from this anticipation, and anticipating the anticipation etc...
- elamje 7y agoUnpopular quote from my image and video processing professor - “The only problem with machine learning is that the machine does the learning and you don’t.” While I understand that is missing a lot of nuance, it has stuck with me over the past few years as I feel like I am missing out on the cool machine learning work going on out there. There is a ton of learning about calculus, probability, and statistics when doing machine learning, but I can’t shake the fact that at the end of the day, the output is basically a black box. As you start toying with AI you realize that the only way to learn from your architecture and results is by tuning parameters and trial and error. Of course there are many applications that only AI can solve, which is all good and well, but I’m curious to hear from some heavy machine learning practitioners - what is exciting to you about your work? This is a serious inquiry because I want to know if it’s worth exploring again. In the past university AI classes I took, I just got bored writing tiny programs that leveraged AI libraries to classify images, do some simple predictions etc.
- TaylorAlexander 7y agoI’m not a heavy practitioner but as a robotics engineer being able to use existing algorithms to perform previously difficult tasks is exciting. I’m also hopeful that the coming decades will bring a lot more to robotics software as machine learning research continues. One thing I’ve been learning is that the black box nature of machine learning algorithms is partially a myth. A lot of tools have been written to help explain models. However I’m a mere novice and student so that’s just something I’ve heard. Would love it if a skilled practitioner chimed in.
- meowface 7y agoI'd highly recommended watching this podcast interview between Lex Fridman and the creator of fast.ai, released yesterday: https://youtu.be/4CTDdxfSXF0 https://youtu.be/4CTDdxfSXF0 He covers a lot of relevant and interesting topics, including how he tries make it less of a black box when designing their courses, and also how it has the potential to confer increased rather than decreased insight into what's going on in a dataset. I don't personally know much about ML, but I think even though there will still be probably an opaque aspect in many cases for a while to come, immense value will still be continually gleaned, as long as people are aware of the limitations. If you accept something is a black box and don't oversell it, a black box is better than no box. All of our own brains are far more of a black box than any deep learning model, in many capacities. But we still use it daily for meat-machine learning, to great success, and can still tune the parameters a bit to improve outcomes, even if we very often don't really understand exactly what we're tuning or why it seems to cause certain effects for some brains (or why it doesn't have those effects for other brains). Consciousness may be the biggest black box of them all, but here we are all talking and making nearly non-stop use of it. I agree it's very important to try our hardest to reach a deeper understanding, but it's kind of like psychiatry vs. neuroscience, or experimental quantum physicists who "shut up and calculate" vs. theoretical quantum physicists who actually want to know what's really going on here at the most fundamental level beyond the useful black box of quantum behavior. While we're trying to solve the hard problems of deep understanding, we can make practical use of what we have in the meantime. We need both kinds of fields and people. If future AI architectures and ideas lead to some degree of convergence with a biological brain, I wonder if the black box problem might become amplified. Maybe one part of the solution is to not use the brain as the model to aspire to, and to eventually seek out alternative avenues to higher, and eventually general, intelligence? (Or maybe I'm completely talking out of my ass, because I'm not at all a researcher or practitioner. I'd appreciate any input from experts.)
- ilaksh 7y agoThis seems like it has the potential for massive efficiency gains and maybe could help with better generalization if the much simpler networks could more easily be reused or recursed or something.
- antpls 7y agoHow is it different than pruning a neural network? It seems you could train the weights of a state of the art NN, then quantizite it, then prune it. It will remove some weights of the NN, then all the remaining weights are set to the same value. Isn't training then pruning more efficient than using an architecture search algorithm ?
- drewm1980 7y agoAt the risk of broad oversimplification, pruning trains and then does an architecture search. This does an architecture search and then trains.
- deleted 7y ago[deleted]
- p1esk 7y agoNo, here the architecture search is the training.
- p1esk 7y agothen all the remaining weights are set to the same value During quantization, weights are set to different values, in the extreme case just 2 different values (binarization). In this case, they are using multiple activation functions to provide various paths for signals to be modified (effectively playing the same role as weights).
- ianamartin 7y agoNext thing you know, google will be telling us that the fastest websites are server side rendered from templates with minimal JavaScript. How could anyone possibly have known?
- scribu 7y agoDiscussion from 3 months ago: https://news.ycombinator.com/item?id=20160693 https://news.ycombinator.com/item?id=20160693
- jangid 7y agoThe analogy given in the article is interesting. Some organisms perform certain actions even before they start to learn. I myself have seen some animals start running immediately after birth. Less number of parameters (shared parameters) could also be thought of as less complexity and hence less processing power requirements; which implies faster training. Phew! too much similarity.
- nurettin 7y agohttps://github.com/google/brain-tokyo-workshop/tree/master/WANNRelease/prettyNEAT https://github.com/google/brain-tokyo-workshop/tree/master/W... to me, this is the really interesting part of the article. NEAT (neuro-evolution of augmenting topologies) is an algorithm for GANN. For those who are looking to implement the algorithm from scratch, see http://nn.cs.utexas.edu/downloads/papers/stanley.ec02.pdf http://nn.cs.utexas.edu/downloads/papers/stanley.ec02.pdf for hours of fun.
- kolar 7y agoHow is this different from genetic programming?
- zwaps 7y agoHow was early Machine Learning different from statistics? New names makes things exciting for people to oick up. Who wants to estimate multinomial regression when you can learn a shallow softmax activated neural network! Its all about creating hype.
- YorkshireSeason 7y agoHow was early Machine Learning different from statistics? I'd argue: in two ways. First: ML's algorithmic focus. Just about anything in modern AI/ML works because it uses compute at extreme scale. For example neural nets seem to work well only when trained with huge amounts of data. Statisticians lacked the background to make this happen. Second: most work in statistics assumed that data was generated by given stochastic data model. In contrast, ML has been using algorithmic models and the data given by an unknown mechanism. In most real-world situations, the mechanism is unknown. It's not just hype. Statistics was stuck in a local optimum, and it was ML's focus on algorithms, data structures, GPUs/TPUs, big data, ... together with the jump into 'weird' data (e.g. the proverbial cat photos), that propelled ML ahead of statistics.
- theferalrobot 7y agoThere are completely non statistical learning algorithms (many at that) which is part of why the distinction is needed. Stats are definitely crucial in parts of the ML domain but not everywhere. Another related reason is just one of focus, where ML doesn't care how to get to a result, it only concerns itself with getting the result. Stats can be a tool to get there along with many other things.
- YeGoblynQueenne 7y ago>> How was early Machine Learning different from statistics? Some of the very early work in machine learning, in the 1950's and '60s was not statistical. The first "artificial neuron", the Pitts & McCulloch neuron, from 1938 was a propositional logic circuit. Arthur Samuel's 1952 checkers-playing programs used a classical minimax search with alpha-beta pruning. Machine learning in the '70s and '80s was for the most part not statistical, but logic-based, in keeping with the then-current trend for logic-based AI. Early algorithms did not use gradient descent or other statistical methods and the models they learned were sets of logic rules, and not the parameters of continuous functions. For instance, a lot of work from that time focused on learning decision lists and decision trees, the latter of which are best remembered today. The focus on rules probably followed from the realisation of the problems with knowledge acquisition for expert systems, that were the first big success of AI. You can find examples of machine learning research from those times in the work of researchers like Ryszard Michalski, Ross Quinlan (know for C4.5 and IDR and the first-order inductive learner FOIL), (the) Stuart Russel, Tom Mitchell, and others.
- TekMol 7y agoIs each architecture given one set of random weights? Or is the architecture of the net tested against a bunch of random weights so that it performs well independently of the weights?
- reiinakano 7y agoNeither. The architecture performs well independently of weights BUT all weights must be the same e.g. Same network works well when all weights are 5.0 or when all weights are -3.0
- TekMol 7y agoHmm... that collides with what patresh said.
- reiinakano 7y agoNo, it doesn't. Shared weight means all weights are the same value, as I originally mentioned.
- patresh 7y agoEach architecture is tested multiple times against different samples of the shared weight From the paper : (1) An initial population of minimal neural network topologies is created (2) each network is evaluated over multiple rollouts, with a different shared weight value assigned at each rollout (3) networks are ranked according to their performance and complexity (4) a new population is created by varying the highest ranked network topologies, chosen probabilistically through tournament selection