4 ms·
MLX doesn't use the neural engine still right? I still wish they would abandon that unit and just center everything around metal and tensor units on the GPU.
by fooblaster 1y ago
MLX doesn't use the neural engine still right? I still wish they would abandon that unit and just center everything around metal and tensor units on the GPU.
- zozbot234 1y agoWrt. language models/transformers, the neural engine/NPU is still potentially useful for the pre-processing step, which is generally compute-limited. For token generation you need memory bandwidth so GPU compute with neural/tensor accelerators is preferable.
- fooblaster 1y agoI think I'd still rather have the hardware area put into tensor cores for the GPU instead of this unit that's only programmable with onnx.
- hannesfur 1y agoOh, I overlooked that! You are right. Surprising… since Apple has shown that it’s possible through CoreML (https://github.com/apple/ml-ane-transformers https://github.com/apple/ml-ane-transformers) I would hope that the Foundation Models (https://developer.apple.com/documentation/foundationmodels https://developer.apple.com/documentation/foundationmodels) use the neural engine.
- hannesfur 1y agoEdit: Foundation Models use the Neural Engine. They are referring to a Neural Engine compatible K/V cache in this announcement: https://machinelearning.apple.com/research/introducing-apple-foundation-models https://machinelearning.apple.com/research/introducing-apple...
- fooblaster 1y agoThe neural engine not having a native programming model makes it effectively a dead end for external model development. It seems like a legacy unit that was designed for cnns with limited receptive fields, and just isn't programmable enough to be useful for the total set of models and their operators available today.
- hannesfur 1y agoThat's sadly true, over in x86 land things don't look much better in my opinion. The corresponding accelerators on modern Intel and AMD CPUs (the "Copilot PCs") are very difficult to program as well. I would love to read a blog post on someone trying though!
- fooblaster 1y agoI have a lot of the details there. Suffice to say it's a nightmare: https://www.google.com/url?sa=t&source=web&rct=j&opi=89978449&url=https://www.amd.com/content/dam/amd/en/documents/products/processors/ryzen/ai/iron-for-ryzen-ai-tutorial-micro-2024.pdf&ved=2ahUKEwiC8f-b1qaQAxWpMjQIHQ43HU8QFnoECBwQAQ&usg=AOvVaw0BYMxmtDG7G_fC1AAWyqUu https://www.google.com/url?sa=t&source=web&rct=j&opi=8997844... AMD is likely to back away from this IP relatively soon.
- llm_nerd 1y agoMLX is a training/research framework, and the work product is usually a CoreML model. A CoreML model will use any and all resources that are available to it, at least if the resource fits for the need. The ANE is for very low power, very specific inference tasks. There is no universe where Apple abandons it, and it's super weird how much anti-ANE rhetoric there is on this site, as if there can only be one tool for an infinite selection of needs. The ANE is how your iPhone extracts every bit of text from images and subject matter information from photos with little fanfare or heat, or without destroying your battery, among many other uses. It is extremely useful for what it does. >tensor units on the GPU The M5 / A19 Pro are the first chips with so-called tensor units. e.g. matmul on the GPU. The ANE used to be the only tensor-like thing on the system, albeit as mentioned designed to be super efficient and for very specific purposes. That doesn't mean Apple is going to abandon the ANE, and instead they made it faster and more capable again.
- almostgotcaught 1y ago> the work product is usually a CoreML model. What work product? Who is running models on Apple hardware in prod?
- llm_nerd 1y agoAn enormous number of people and products. I'm actually not sure if your comment is serious, because it seems to be of the "I don't, therefore no one does" variety.
- bigyabai 1y agoEnormous compared to what? Do you have any numbers, or are you going off what your X/Bluesky feed is telling you?
- llm_nerd 1y agoI'm super not interested in arguing with the peanut gallery (meaning people who don't know the platform but feel that they have absolute knowledge of it), but enough people have apps with CoreML models in them, running across a billion or so devices. Some of those models were developed or migrated with MLX. You don't have to believe this. I could not care less if you don't. Have a great day.