4 ms·
A Primer in BERTology: What We Know About How Bert Works
- taneq 6y agoFor anyone who, like me, isn't a BERTologist, BERT is a neural network architecture.
- jeffrallen 6y agoHe's also Ernie's best friend.
- godelmachine 6y agoErnie who?
- cinntaile 6y agoI can't hear you Bert, I've got a banana in my ear.
- nullsense 6y agoOh hey Ernie! Oh hey Bert!
- kleiba 6y agoLet's not forget Elmo!
- mobilio 6y ago" BERT is a method of pre-training language representations, meaning that we train a general-purpose "language understanding" model on a large text corpus (like Wikipedia), and then use that model for downstream NLP tasks that we care about (like question answering). BERT outperforms previous methods because it is the first unsupervised, deeply bidirectional system for pre-training NLP." https://github.com/google-research/bert https://github.com/google-research/bert
- deleted 6y ago[deleted]
- niea_11 6y agoCan anyone please explain (in layman terms if it's possible) how did the researchers come up with the method in the first place if the process how the method finds the answers is not understood?
- JacobiX 6y agoThroughout human history, you can find many discoveries that were made before understanding why they work in the first place: we know exactly how the neural networks work, but we don’t know why they are so effective, maybe because of the lack of a theoretical understanding of deep learning and complex neural networks. For this particular case, BERT is the result of incremental enhancements of existing architectures and training procedures. Fundamentally, BERT is a sequence prediction algorithm, and historically, the sequence prediction models were based on complex recurrent or convolutional neural networks. Experiences showed that the best performing models were those having an attention mechanism (the concept of directing the focus on some words or sentences). Some researchers proposed a new simple network architecture, based solely on attention mechanisms, without complex recurrence or convolutions! and they showed that some of these models achieve state-of-the-art performances while being more parallelizable and requiring significantly less time to train.
- PaulHoule 6y agoCave men discovered fire before humans discovered phlogiston, which they discovered before they discovered there was no such thing as phlogiston but rather oxygen instead. BERT was motivated by the discovery that neural nets will try to learn functions you show them and some ideas of how to route information so it forms bottlenecks that force the network to learn general representations, but there was not a mature theory behind it and judging by that review paper there still isn't. That whole paper reads to me like an account of wandering in the dark. If I read a paper that long about liquid rocket fuels I'd learn that 99% of the things I might want to use as a rocket fuel won't work and I really have a choice of Hydrogen, Methane or RP-8 and Oxygen. If you were out to "build a better BERT", even a slightly better BERT, that paper doesn't give clear guidelines about what you should do. It's got all the trappings of a field which is preparadigmatic but could be mistaken for paradigmatic because of the sheer volume of researchers, conferences, papers, etc.
- hallqv 6y agoAny new information in the paper since the first version came out in Mars? Otherwise a 6 month old meta-study seems kind of dated given rate of progress in NLP atm.