8 ms·
Bfloat16 support coming to Apple's Metal and PyTorch [video]
- eoskx 3y agoSomehow missed this from WWDC23, but it looks like Sonoma will add support for bfloat16 with Metal, and there's an active PR to add support with the PyTorch MPS back-end (PR #99272). Since M2 added bfloat16 support at the hardware level, I'm assuming this will only be supported on M2 Macs. That maxed out Mac Studio M2 w/ 192GB of memory now looks more appealing...
- kzrdude 3y agoVisible in the unofficial documentation for AMX instructions too - M2 only bf16 functionality - https://github.com/corsix/amx/blob/main/matfp.md https://github.com/corsix/amx/blob/main/matfp.md This matfp instruction computes an outer product and is a kernel for matrix multiplication.
- dlewis1788 3y agoI didn't even know about Apple's AMX instructions until I clicked on your link. Very interesting - thanks!
- my123 3y agobf16 in Metal on macOS 14 is supported on all Macs. Emulated in software transparently.
- LoganDark 3y agoYeah, Metal is pretty great because it runs the same on all Macs. Apple is really really good at this.
- minimaxir 3y agoI'm still confused by the proliferation of bf16. Although it certainly doesn't hurt compared to fp16, in my testing even with A100 GPUs optimized for it, both training speed and inference quality are the same between bf16 and fp16.
- YetAnotherNick 3y agobf16 is generally easier to train neural network than fp16 on due to no need for scaling. And most model training and inference performs the same with fp32 and bf16.
- gok 3y agoFp16 makes it easy to accidentally overflow, especially around summation operations.
- redox99 3y agoSometimes during training, fp16 will cause networks that would converge on fp32, to explode to Infs or NaNs with fp16, because of the limited range. bf16 generally speaking fixes that. It's true also that fp16 is often manageable with enough batch/layer norm and gradient clipping.
- voz_ 3y agoYea, I spent a few months comparing the two, and empirically i had a lot more issues with various normalized entropy problems (explosion, not converging, converging slower) with fp16 than with bf16. The transfer pipeline I wrote for fp32->fp16 also took a lot more work than fp32->bf16
- dlewis1788 3y agoMy understanding is for certain types of networks BF16 will train better than FP16, given the additional protection against exploding gradients and loss functions with the extended range of BF16 - at the loss of precision.
- bobbylarrybobby 3y ago(Not an ML guy.) bf16 and fp16 should be comparable if the weights are of the same magnitude, but what happens in a network where the weights are poorly regularized?
- eoskx 3y agoSomeone commented below that with enough batchnorm/layernorm/etc. and/or gradient clipping you can manage it, but BF16 just makes life easier if you can live without some precision.
- dlewis1788 3y agoConfirmed Apple M1 lacks bfloat16 support completely - M1: hw.optional.arm.FEAT_BF16: 0 vs M2: hw.optional.arm.FEAT_BF16: 1
- londons_explore 3y agoLuckily BF16 is just a truncated FP32. That means that the hardware can do BF16, just you don't get any performance benefit compared to FP32 (and depending on the hardware design, you might also have to space the data 4 bytes apart rather than 2), so you lose the memory bandwidth and RAM usage benefits too.
- pklausler 3y agoConversions from IEEE-32 to BF16 don't round?
- londons_explore 3y agoI don't believe the standard defines it. I believe implementations truncate (ie. round towards zero). Remember BF16 was invented specifically to be able to be backwards compatible with existing silicon - and pulling 2 bytes out of 4 is a far cheaper operation than any rounding.
- kelnos 3y agoJust to elaborate, as I was confused about this and had to look it up: BF16 is indeed designed to just be a truncated F32: you can grab the top 16 bits of a F32 value and it'll still "make sense": the sign bits are in the same place in both (unsurprisingly), and the exponent part of BF16 and F32 are both 8 bits. In the case of the mantissa, you end up grabbing the top 7 bits of the F32's 23-bit mantissa, so it all works out, as this will "round" the value toward zero.
- pclmulqdq 3y agoThere's no standardized definition of BF16.
- infogulch 3y agoI think posits are better. https://posithub.org/ https://posithub.org/
- pclmulqdq 3y agoPosits seem interesting, but they are fundamentally very different than floats and ints, and much harder to analyze.
- londons_explore 3y agoIt's a shame that large language models are mostly moving to 4 bit weights for inference, and a bunch of papers have shown promising techniques for training in 4 bit too... Remember that switching from 16 bit to 4 bit lets you have 4x as many weights, 4x as many weights loaded from RAM per second, and ~1/16 of the silicon area for the calculations (a multiplier scales with approximately the number of bits squared). That smaller silicon area will let you do more per $ too...
- MrBuddyCasino 3y agoI dimly remember reading that the mathematical compute-per-density optimum is around 3.x bits in a „brain like structure“, I don’t remember any details though or the precise context. Does this ring a bell with anyone?
- isoprophlex 3y agoWhat?! Can you also train with quantization? Incredible! I'd have thought the gradients were way too ugly for any convergence with 4 bits. Any particularly good papers you can recommend me on the topic?
- sbierwagen 3y agoA group at IBM has been working on minifloat training for a while. Here's a paper from 2020 on FP4 training: https://papers.nips.cc/paper/2020/file/13b919438259814cd5be8cb45877d577-Paper.pdf https://papers.nips.cc/paper/2020/file/13b919438259814cd5be8...
- londons_explore 3y agoTheir best performing 4-bit number format uses 1 sign bit, 3 exponent bits, and no mantissa bits! Ie. All weights, activations and gradients become powers of two! Which means all multiplications become simple bit shifts. That really changes mathematics and silicon design.
- 3y ago
- hospitalJail 3y agoMaybe someone can help me understand why people are investing into this. Inhousing typically means falling behind in technology but having lower operating costs. That makes the company win, not the users. If you hinge your career on Apple, they might make your technology obsolete on a dime. Its not the fastest, its not the best, its not the cheapest, its not some combination either. > 'compute per watt' With AI? The local LLM models are near useless already. There will be a time to cut down on power, but from what I've read, there is currently ~no value even with a 4090 with 512 RAM. I suggest avoiding Windows/M$, I am annoyed with Linux bugs, and google cannot be trusted. But all of that could be said about Apple as well. I just don't see a future with Apple hardware, it gives me some serious Nintendo vibes where they are going to be some quirky niche that is just enough for marketers to sell it. Compute per watt seems like a wiimote that no one asked for, but suddenly claim is ultra important. Maybe someone can change my view. I don't see who buys this when they are educated on the possible options.
- IOT_Apprentice 3y agoMy question to you is what are you currently using as an alternative for the COU/SOC in your personal & work environments? Intel? AMD Ryzen? Apple has taken their ARM approach and scaled it to all their platforms. Amazon now is on what, Gen 2 or 3 for their graviton platform in AWS. And what OS are you using if you don’t trust Microsoft, Linux or Apple?
- brucethemoose2 3y agoCPU arch isnt't even that critical here, as Apple is talking about Metal.
- brucethemoose2 3y ago> Maybe someone can help me understand why people are investing into this. Buying a Mac for running LLMs is kinda like buying a Mac for gaming. Its thoeretically interesting, but I don't think thats a serious driver of Mac sales. But: - Finetuned local LLMs are good for specific niches, like roleplaying, text games, and helper bots for your own pile of data. And they are getting better at other niches like code completion for specific languages, or summarization. - Remember that a huge selling point for Macs is iPhone/iPad development. The market for AI App Store apps is not small.This is also a reason to believe there will be some stability with the ML support.
- victor106 3y agoI think the trillion dollar question is: can Apple ever make Mac's / GPU's to compete with NVIDIA?