4 ms·
I'll be interested to see how this performs if/when it is released. I'm not convinced that simply increasing the number of parameters improves the models, and
by techwizrd 3y ago
I'll be interested to see how this performs if/when it is released.
I'm not convinced that simply increasing the number of parameters improves the models, and I don't think we should be putting out press releases "spec-chasing" increases in the number of parameters. I've also found in my research that larger models often do not perform better, but this is becoming more difficult to explain to non-technical folks due to irresponsible marketing. If we aren't careful about the bold marketing claims, we'll disillusion many people and reduce the potential for future growth—we don't want another AI Winter.
There's a significant environmental cost to training these large models, and it's harder to understand, use, and control these very large models (even as a researcher).
- EGreg 3y agoWhat happened to “the scale hypothesis” LOL. Is scale really is all you need? Maybe this model can invent time travel and improve itself in the past..
- londons_explore 3y ago> I'm not convinced that simply increasing the number of parameters improves the models At every size people say this... And then someone comes out with a model 10x larger which outperforms the previous state of the art. Sure, you need to still take a little care with architecture choice and hyperparameter selection, but size really does seem to be king.
- weatherlite 3y ago> we don't want another AI Winter Speak for yourself...
- Der_Einzige 3y agoParameter count is strictly better IF the number of tokens (and ideally better quality tokens) trained on increases, and if the training is done for longer (most LLMs are way undertrained) Most of the huuuuuge models failed on most or all of these fronts and that's why they suck compared to Llama or Alpaca or Vicuna
- elcomet 3y agoThat's not true. For the same number of training tokens, bigger is better. And for the same size, more tokens is better. So obviously more tokens and bigger is better.
- rvz 3y ago> There's a significant environmental cost to training these large models, and it's harder to understand, use, and control these very large models (even as a researcher). Exactly. The worse part is there is NO viable efficient way of training, fine tuning etc and inference with these massive deep learning models for more than a decade since deep neural networks used GPUs for training, and it still requires a substantial amount of compute power and energy that is incinerating the planet and the result is untrustworthy large black-box models that cannot be trusted or transparently explain their decisions, especially with safety-critical or high risk tasks. Crypto at least has an alternative to the wasteful proof-of-work system, and Ethereum which was formerly PoW has shown it is possible to switch to a greener alternative consensus method. Deep Learning however, still has not shown such a viable switch and still needs to burn the planet with more data-centers of GPUs, ASICs and FPGAs to create hallucinating models that have shown to break on a single pixel or to confidently regurgitate nonsense as the truth with little understanding in reasoning and also answer with demonstrably false information. LLMs like this one is still essentially snake-oil BS generators hiding behind regurgitation and sophistry to pretend to show signs of 'intelligence'. EDIT: It is all true. [0] There is no amount of green-washing to hide the problem of deep learning systems wasting essential resources like thousands of running taps of water. Literally. [0] https://gizmodo.com/chatgpt-ai-water-185000-gallons-training-nuclear-1850324249 https://gizmodo.com/chatgpt-ai-water-185000-gallons-training...
- sebzim4500 3y ago>I'm not convinced that simply increasing the number of parameters improves the models How so? Llama 33B is clearly worse than llama-64B and that's 'only' a 2x increase. Are you aware of any example where a larger model fails to outperform a smaller one when all else is equal (tokens, architecture, data quality, etc.)? Obviously for a fixed amount of training compute more parameters can be bad, but there's a trade off where more parameters means you train on fewer tokens.