4 ms·
I'm one of the authors. I may be able to answer some questions or concerns.
by phillypham 6y ago
I'm one of the authors. I may be able to answer some questions or concerns.
- lacker 6y agoIf you compare what BigBird is good at to what GPT-3 is good at, what are the relative strengths of each system?
- phillypham 6y agoRight now, BigBird is encoder only so it doesn't generate text. Causal attention with the global memory is a bit weird, but we could probably do it. GPT-3 is only using a sequence length of 2048. In most of our paper, we use 4096, but we can go much larger 16k+. Of course, we don't nearly have as many parameters as GPT-3, so our generalization may not be as good. BigBird is just an attention mechanism and could actually be complementary to GPT-3.
- teruakohatu 6y agoAny examples of turning complete interaction?
- phillypham 6y agoNot really. The proofs are more of a curiosity really. I think the strong performance on QA takes that require multihop reasoning give some evidence that the model is capable of complex reasoning.
- sytelus 6y agoI'm curious how long this work took. Did you encountered many negative results before you got good results? What inspired you for the main insight?
- phillypham 6y agoI think the insight is not incredibly original and follows naturally from OpenAI's Sparse Transformer. The idea is similar to Longformer. Two teams at Google had a similar insight, hence the high number of authors. The original implementation only took a couple of months and was primarily motivated by internal Google applications. Natural Questions was the first external benchmark we tried to validate on, which took a few months to find the right setup. All the other datasets, took a few weeks but the effort was done in parallel given the large team. There was quite a bit of frustration dealing with Tensorflow, TPUs, and the XLA compiler that maybe set us back a few months, too.
- phillypham 6y agoEdit: This is not meant to criticize the Tensorflow, TPU, or XLA team. They responded quickly to our bugs and made the work possible. I just meant there was some extra organizational overhead.
- visarga 6y agoThe main advantage of Big Bird is its linear complexity in sequence length. If it were to be trained on the same corpus as GPT-3 what would be the advantages/disadvantages? I am thinking maybe longer context window, faster training and less memory use, but what about performance, will it measure up?
- phillypham 6y agoThis is something we want to explore. BigBird just replaces the attention mechanism in BERT. We believe something like BigBird can be complementary to GPT-3. GPT-3 is still limited to 2048 tokens. We'd like to think that we could generate longer, more coherent stories by using more context.