13 ms·
DeepMind’s new AI with a memory outperforms algorithms 25 times its size
- amitport 5y agoDup https://news.ycombinator.com/item?id=29486607 https://news.ycombinator.com/item?id=29486607 (This is a different blogpost, but does not seem to add over the original) Edit: following derac's comment see https://news.ycombinator.com/item?id=29646112 https://news.ycombinator.com/item?id=29646112 for RETRO
- derac 5y agoThis article is about RETRO [0], not Gopher. [0] https://deepmind.com/research/publications/2021/improving-language-models-by-retrieving-from-trillions-of-tokens https://deepmind.com/research/publications/2021/improving-la...
- jari_mustonen 5y agoCould someone explain the article to layman engineer?
- visarga 5y agoIt's language modelling with search engine in-the-loop. Instead of training GPT-3 with 178B weights, you train a 25x smaller model and allow it to retrieve useful snippets from a large text index as additional information. This solves the problem of very large models and the problem of updating an already trained model, as you can swap the text corpus with a newer one. The model learns mostly syntax, burning less trivia in its weights than a regular LM as it can simply copy the relevant information from the index. This development was bound to happen as large LMs are expensive to use and it was an obvious idea. We've had these semantic search text indices for a few years already[1], they just weren't combined with text generation. [1] https://github.com/spotify/annoy https://github.com/spotify/annoy
- jamesblonde 5y agoYes, the key technology here is a scalable embedding store. The leading players here are the indexes - faiss and scann. The open source platforms are opensearch, elasticsearch, featureform, milvius. Then there are saas products like pinecone.
- alkonaut 5y agoSo the memory doesn't solve the context problem of e.g. "conversation context"? I.e. the storage isn't modified while the model is used? If I make an app that makes conversation using such a model model, then the storage isn't modified to insert knowledge about what the early parts of the conversation was about, and it's only bringing a database of fixed information into the conversation? (I have a friend who is just like that).
- TulliusCicero 5y agoI know AI Dungeon and Novel AI both factor in several recent text inputs when generating new text, and also have a memory section where you can add things you want the AI to never 'forget' about the current story.
- tveita 5y agoYou could update your storage as you go, the indexing doesn't appear to be that expensive. For many tasks it wouldn't be helpful because the input is small enough to be covered by the context already, and for summarizing and question answering tasks, you want it to repeat information from other documents, but not from earlier in its own output. It might be interesting for a long-context task like "given the first parts of this book, complete the next chapter".
- davefol 5y agoSeems odd to claim 25x reduction in size when the algo involves looking into a database of a trillion chunks of text.
- visarga 5y agoThe "algo" here refers to the neural net itself. The text index is considered an easy problem as you can do lookups in logarithmic time.
- ehsankia 5y agoThe word "Algo" here is definitely awkward. The point is though that what matters most here is the number of parameters, as those correlate quite closely with training and inference time. Storage space is pretty trivial, but TPU cycles are less so.
- davefol 5y agoThanks for this. Rereading + your comment and I think I have a better understanding of why this is progress.
- neom 5y agoA neural net with access to wikipedia is faster than than a neural net that contains Wikipedia? Seems odd to call it AI with a memory though... unless I'm misunderstanding. It's more like AI with a decent memory and an understanding of how to use an encyclopedia.
- robbedpeter 5y agoYeah, memory implies persisted state in the model, this is static lookups separate from the transformer. Still superb, though, there's no reason you can't use other gofai tools vs a static database, to trigger expert systems or formalized reasoning.
- neom 5y agoI got into a very long debate with an openai person 4/5 years ago about this + adversarial learning + access to a quantum computer (think just straight up world class abacus) was close to the primitives required for more generalized AI. They didn't agree with me, but that's ok! :)
- torbTurret 5y agoWhat’s your argument?
- neom 5y agoI hope I didn't come off like I think know anything about this field, because I don't. A friend who use to work for openai (Jack Clark) and I spend some time once discussing over beers some stuff around general purpose AI, and I proposed that I believe quantum is a dark horse on the road to general purpose AI, and he disagreed.
- stavros 5y agoSounds like a fun debate to me.
- 5y ago
- deleted 5y ago[deleted]
- ggm 5y agoIf we point it at the horrendously bad scots wiki (some kid in the US decided he'd translate Wikipedia into what he thought was lowland scots/Doric.. it's a disaster) we might get entertainingly bad outcomes.
- bawolff 5y agoNote, the stuff written by said kid has long since been deleted. However i have no idea what the quality of the rest of scowiki is.
- ggm 5y agoIt's not awful, but I feel it's still pretty meh. I say this as a person raised in Edinburgh in the sixties and seventies. How bad? Well.. in their backend meta pages they link to the DSL (Dictionary of the scots language/dictionars o' Scots Leid [0]) which says this: Written Scots In the written mode, Scots spelling remains variable. Attempts to make it more consistent, notably the Scots Style Sheet produced by the Makars’ Club in 1947 or the Recommendations for Writers in Scots published by the Scots Language Society in 1985, have had at best only limited success, competing with other systems that have been developed to represent more closely localized varieties of spoken Scots. When your reference text says the language isn't yet well captured in a single print, you better believe the wiki page is a hot mess. [0] https://dsl.ac.uk/about-scots/what-is-scots/ https://dsl.ac.uk/about-scots/what-is-scots/
- LudwigNagasena 5y agoWell, it is pretty hard to make something in a language when it is a dialect continuum and not a standardized variety that is forced onto the whole population through the education system and media.
- vkazanov 5y agoNot really relevant to the topic in question but... Isn't this how most languages begin? First you have a continuum of language dialects, then one of them dominates for political reasons, then it gets codified, then enforced onto everybody through centralised education. Dialects not under direct unified political control become related but separate languages... And so on.
- whazor 5y agoVery interesting. GPT-J is an opensource free alternative to GPT-3 and requires at least 12.1GB memory to run the model (which is reduced from original 48GB ram). But if the model stores some kind of index and does internet searches (or hard drive) instead, then it could scale much further as there is a limit on how much memory you can use in production.
- sbierwagen 5y ago>48GB ram 48GB VRAM? 48+ gigabytes of system ram is cheap, 48 gigabytes of ram on a GPU is still painfully expensive.
- charcircuit 5y agoYes, vram
- x3al 5y agoWould m1 max work?
- blovescoffee 5y agoRight now there's meh Deep Learning support on the m1 max. TensorFlow is ~supported~, PyTorch is not. Apple has their Neural Engine and CPU/GPU architectures sort of hidden and so it's hard to port or support them.
- schleck8 5y agoGPUs are the slowing factor in general if I'm not mistaken when it comes to Deep Learning progress.
- moffkalast 5y agoYes, in this GPU market that's essentially a new car's worth of cash.
- 5y ago
- Blikkentrekker 5y ago> Gebru, a widely respected leader in AI ethics research, is known for coauthoring a groundbreaking paper that showed facial recognition to be less accurate at identifying women and people of color, which means its use can end up discriminating against them. Surely this is a function of location? I understand the U.S.-English term “person o color” to be convoluted language for “not white”. One simple thing I notice is that if I search for, say, “child” on Google Image Search, the images indeed tend to look as what one would expect from the average inhabitant of an English-speaking nation, when I search “子供”, I indeed mostly see what I would expect from Japan. Similarly, if I search for “house”, what I find tends to look like a house most likely situated in the Netherlands; with “บ้าน”, it does resemble more so stereotypical Thai architecture. I would assume that a.i.'s made in, say, Japan would yield different results.
- xmprt 5y agoThe problem is that AI (and the English language to some extent) transcends borders. So even if it's an AI developed in the US, it can potentially impact people outside the US and it makes ethical sense to build something that doesn't exclude groups based on arbitrary conditions.
- Blikkentrekker 5y agoYes, but to offset that, many a.i. in English were also made outside of English-speaking regions, in what one assumes to be proportional degree. This is probably why there is more variance when searching for English terms as wel, as a Lingua Franca. If I search “house” I do see some styles of architecture not commonly found in Anglo-Saxon nations, whereas all occurrences of “huis” do seem to be situated in the Netherlands.
- sangnoir 5y ago> many a.i. in English were also made outside of English-speaking regions Different regions, yes - but where did the training and benchmark datasets come from? AI research is surprisingly monocultural (or use "standardized benchmarks" if you're feeling charitable). Not too long ago, there was a paper posted on HN that showed that a bunch of the datasets contain mislabeled data, which means a lot of "different" models are encoding similar biases.
- charcircuit 5y agoSo where can we download these models?
- rocgf 5y agoGiven that this is DeepMind and not some more open AI organization, I assume you cannot.
- deleted 5y ago[deleted]
- AJRF 5y agoI've not kept track of where large transformers like this have gotten to, GPT3 and the like - has GPT3 made any real difference to the world? Are people using it? Has it vastly improved any software?
- muzani 5y agoI don't know about world changing but it's saved me hundreds of hours. I use it to help read academic papers, put formatting on things like markdown and subtitles, and creative writing. A lot of the things that take it 15 seconds to do take me 2 minute and drain me mentally for about 15 mins. If anything, it's being used in force for social media marketing, where you're trying to say "buy this thing" in different ways every day.
- regularfry 5y agoForgive the ignorance, but how? What tools are you using on top of GPT3 to do those things?
- motoboi 5y agoThere is an API for GPT-3. GitHub copilot uses Codex, a descendent of GPT-3.
- muzani 5y agoIt's a funny question, like asking how to write a book with a computer, but perfectly valid. You can access GPT-3 directly now. There's no waitlist, but there still are restrictions. There's some examples here: https://beta.openai.com/examples https://beta.openai.com/examples You don't even need the API. Once you get access, it comes with access to the playground, which is enough to do anything you like. If you look at the examples, it's very "no code". You literally tell the AI what you're trying to do and it tries its best. Most of the work in prompt engineering is writing something that can't be misunderstood. But you just have to explain to it what you want like you would to a child.
- sdenton4 5y ago
- kkjjkgjjgg 5y agoSounds as if they stored all the correct answers in a database and call it "better". How do they even evaluate these models? Like they already have a billion preprepared correct answers in the database. How do they come up with new questions for the evaluation?
- charcircuit 5y agoIt's the equivalent of taking an a test where you can use the internet. Sure you know the information needed to answer the question exists, but it can be difficult to extract the answer and word it into at English sentence.
- blovescoffee 5y agoInstead of storing the correct answers in an encoded/embedded form in the weights of the neural net (certain neurons very loosely corresponding to certain "answers") the correct answers are stored elsewhere. That way we can scale down the model to the necessary "thinking" parts and we don't need to use excess neurons for the "memory" part. Kind of handwavey but hopefully that explains the general idea.
- kkjjkgjjgg 5y agoYou mean otherwise the whole words would be encoded in the net, and now you only need to encode the index in the database?
- tasty_freeze 5y ago> all the correct answers That is clearly not possible, so it can't be what they are doing. Rather than diffusely encoding that knowledge in a massive number of self-organized layers of weights, it is explicitly encoded. The remaining network can "focus" on mapping input to retrieve the relevant information stored in that database, and extracting/interpolating/extrapolating that information based on the current context to generate useful output.
- Billfiles 5y ago
- baalimago 5y agoOkay. Now make the small AI with memory 25 times bigger!
- yieldcurvinator 5y ago
- mik09 5y agooftentimes one can shrink a model down dramatically once one has a bigger, more robust model. but shrinking a huge model is still a great achievement.
- lebuffon 5y agoCould we say that they are re-inventing the human mind architecture by enhancing "fluid intelligence" with "crystallized intelligence". As humans age we apparently lose the former but compensate with the latter as best we can.
- meiji163 5y agoI'd be interested to see if these models are robust against algorithms like TextFooler [0]. I'm skeptical this trend of 10x'ing the parameters will solve the "clever hans" problem. [0]: https://github.com/jind11/TextFooler https://github.com/jind11/TextFooler
- rapjr9 5y agoThis seems like a very interesting approach to creating an AI that can continuously learn new things by just updating its database. Maybe a first step towards a general purpose AI? It would be interesting to create a personal assistant based on this whose database was fed the entire digital stream generated by a persons life. How would you protect such an AI from misuse? Add another AI with a database of information on ethics that acts as a gatekeeper? Could you somehow keep the gatekeeper from being turned off, perhaps by using cryptography in some fashion for access control?