4 ms·
How do they even train such huge models? Especially when you start talking abou multi-modal models. Is the underlying optimization process still relying on grad
by numbers_guy 4y ago
How do they even train such huge models? Especially when you start talking abou multi-modal models. Is the underlying optimization process still relying on gradient descent? I am just baffled how they manage to get it to train across such vastly different sets of data and manage to get it to work.
- deleted 4y ago[deleted]
- jah242 4y agoWhilst not the same I recommend you look at the DeepMind Gato paper to see surprisingly (relatively) simple multi modal can be - https://openreview.net/forum?id=1ikK0kHjvj https://openreview.net/forum?id=1ikK0kHjvj Essentially to merge lots of modalities they just go 'let's convert all modalities into integers in the same given range', e.g the word 'me' = 1001, up in Atari = 11002, joint torque of right motor of robot = 33000 and so on. From the paper: There are infinite possible ways to transform data into tokens, including directly using the raw underlying byte stream. Below we report the tokenization scheme we found to produce the best results for Gato at the current scale using contemporary hardware and model architectures. • Text is encoded via SentencePiece (Kudo & Richardson, 2018) with 32000 subwords into the integer range [0, 32000). • Images are first transformed into sequences of non-overlapping 16 × 16 patches in raster order, as done in ViT (Dosovitskiy et al., 2020). Each pixel in the image patches is then normalized between [−1, 1] and divided by the square-root of the patch size (i.e. The tokenized result is a sequence of integers within the range of [0, 1024). 16 = 4). • Discrete values, e.g. Atari button presses, are flattened into sequences of integers in row-major order. • Continuous values, e.g. proprioceptive inputs or joint torques, are first flattened into sequences of floating point values in row-major order. The values are mu-law encoded to the range [−1, 1] if not already there (see Figure 14 for details), then discretized to 1024 uniform bins. The discrete integers are then shifted to the range of [32000, 33024).
- jcims 4y agoThe interesting thing to me is that our brains probably do something similar, converting multi-modal sensory data into the same 'model' that we experience as our concsiousness.
- kelipso 4y agoUsually you stick a CNN (or whatever other vision encoding neural network) after the image inputs and pipe the output of the CNN as an additional token or tokens of the GPT transformer (into the layer after the text embedding layer). That way everything is differentiable. You can pre-train the CNN and Transformer independently and then train them together (usually easier to train this way), or just train them from scratch. There are lots of other ways to combine two networks also. So input is now one (or more) images and text, and output is text. There are ways to position the input image in a particular location within the input text as well. Training data can come from websites, etc.
- swatcoder 4y agoMore likely, this round of innovation is focused on integrating many engines into one product using techniques along the lines of langchain. Instead of grinding increasingly complicated models that do everything, you train the LLM input/output engine to delegate to other systems and synthesize the results. There’s a ton of headroom down this road now that the natural language interface of the LLM has become so capable. There are still surely active research tracks on integrating more data into more sophisticated models, but it looks like we’re at a maturity point where product engineering can start driving its own innovations.