10 ms·
AMD Unveils Its First Small Language Model AMD-135M
- loufe 2y agoIt's always encouraging to see wider hardware platform competition for AI inference and training. Access to affordable and capable hardware for consumers will only benefit (I imagine) from increasing competition.
- diggan 2y ago> The training code, dataset and weights for this model are open sourced so that developers can reproduce the model and help train other SLMs and LLMs. Wow, an actual open source language model (first of its kind [from a larger company] maybe even?), includes all you need to be able to recreate it from scratch. Thanks AMD! Available under this funky GitHub organization it seems: https://github.com/AMD-AIG-AIMA/AMD-LLM https://github.com/AMD-AIG-AIMA/AMD-LLM
- bubaumba 2y agoNo, it's not open source till someone can actually reproduce it. That's the hardest part. For now it's open weights open dataset. Which is not the same.
- diggan 2y agoThat's... Not how open source works? The "binary" (model weights) is open source and the "software" (training scripts + data used for training) is open source, this release is a real open source release. Independent reproduction is not needed to call something open source. Can't believe it's the second time I end up with the very same argument about what open source is today on HN.
- dboreham 2y agoBut wouldn't failure to achieve independent reproduction falsify the open claim? Similar to you publish the source for Oracle (the database), but nobody can build a binary from it because it needs magic compliers or test suites that aren't open source? Heck when the browser was open-sourced, there was an explicit test where the source was given to some dude who didn't work for Netscape to verify that he could actually make a working binary. It's a scene in the movie "Code Rush".
- bubaumba 2y agoYou are missing key points here. "reproduce" means produce the same. Not just train similar model. I can simplify the task, can you convincingly explain how the same model can be produced from this dataset? We can start simple, how you can possibly get the same weights after the first single iteration? I.e. the same as original model got. Pay attention to randomness, data selection, initial model state. Ok, if you can't do that. Can you explain in believable way how to prove that given model was trained on give dataset? I'm not asking you for actually doing all these things, that could be expensive, only to explain how it can be done. Strict 'open source' includes not only open weights, open data. It also includes the word "reproducible". It's not "reproduced", only "reproducible". And even this is not the case here.
- worewood 2y agoHow often do people expect to compile open-source code and get _exactly_ the same binary as the distributed one? I've seen this kind of restriction only on decompilation projects e.g. the SM64 decompilation -- where they deliberately compare the hashes of original vs. compiled binaries, as a way to verify the decompilation is correct. It's an unreasonable request with ordinary code, even more with ML where very few ones have access to the necessary hardware, and where in practice, it is not deterministic.
- e12e 2y agoI expect that if I compile your 3d renderer, and feed it the same scene file you did - I get the same image?
- wrs 2y agoThe interesting part of the product we’re taking about (that is, the equivalent of the executable binary of an ordinary software product) is the weights. The “source” is not sufficient to “recompile” the product (i.e., recreate the weights). Therefore, while the source you got is open, you didn’t get all the source to the thing that was supposedly “open source”. It’s like if I said I open-sourced the Matrix trilogy and only gave you the DVD image and the source to the DVD decoder. (Edit: Sorry, I replied to the wrong comment. I’m talking primarily about the typical sort of release we see, not this one which is a lot closer to actually open.)
- littlestymaar 2y ago> The “source” is not sufficient to “recompile” the product (i.e., recreate the weights). Therefore, while the source you got is open, you didn’t get all the source to the thing that was supposedly “open source”. What's missing?
- wrs 2y agoWell, I’m not experienced in training full-sized LLMs, and it’s conceivable that in this particular case the training process is simple enough that nothing is missing. That would be a rarity, though. But see my edit above — I’m not actually reacting to this release when I say that.
- littlestymaar 2y agoOK, so you just like to be a contrarian…
- Jabrov 2y agoWhat’s the difference?
- avaldez_ 2y agoReproducibility? I mean what's the point of an open technology nobody knows if it works or not.
- frontalier 2y agothe goal posts moved?!
- jerrygenser 2y agoThis would be another example of open source. Not from such a large company but a good reference including code, data, weights, etc. https://allenai.org/olmo https://allenai.org/olmo
- brianjking 2y agoMolmo even more so! The 7b is wild.
- wrs 2y agoWe (developers and tech managers) really need to hold the line on this terminology. This is a full actual open source LLM. The usual “open inference” model is not.
- boulos 2y agoI assume by "open inference" you mostly mean "weights available"?
- wrs 2y agoUsually “open source” for an LLM means you get the weights and the inference code, which I’ve started calling “open inference”. It’s certainly good and useful, but it’s not the actual source of the model. I find people get into silly arguments about the terminology because they’re focused on whether the “source” is “open” and not on what the “source” is actually the source of. “Weights available” indicates even the weights aren’t “open” in the usual software meaning of the term, as they typically come with restrictive licenses (more restrictive than copyleft or attribution).
- throwawaymaths 2y agoBy "source" do you mean training data?
- nickpsecurity 2y agoI call them open weights or just freeware like when we got only the EXE’s on Windows.
- wmf 2y agoYou're not wrong, but if you come up with a definition that no one is willing to meet you're just making that definition irrelevant.
- deleted 2y ago[deleted]
- GeekyBear 2y ago> Wow, an actual open source language model (first of its kind Apple research has previously released another example of a model with open training code, data, and weights, but their model was sized for running inference workloads on mobile devices. However, Apple has a mobile device line of business and AMD has an enterprise AI accelerator line of business, so they are both doing work relevant to their bottom line.
- diggan 2y agoThanks, seems you're talking about the OpenELM family of models: https://github.com/apple/corenet/tree/main/projects/openelm https://github.com/apple/corenet/tree/main/projects/openelm
- kypro 2y agoSmart move from AMD. Helps develop an ecosystem around their tech and for their GPUs.
- jeff_carr 2y agoHas anyone tried it? I mean, I would, but as far as I can tell understand I need 4 boxes with 4 GPU's. Plus an interconnect. I mean, I could put in an order for my homelab but at around $80k per box and maybe $20k for the right switches and some other gear, my wife will probably frown at me ordering a $340,000 rig to try this code that I don't know what to do with it if it works. Maybe some other heavy hitter out there can explain what all this whatchamacallit newfangled synergy producing matrix algebra does after you have it running?
- Shadowmist 2y ago> that I don't know what to do with it if it works. After you get it up and running you can just ask it what to do with it.
- Rinzler89 2y agoHope it tells you to buy more AMD HW because that would be so funny.
- kypro 2y agoSeems kinda obvious this isn't targeted at hobbyists, but more towards SMEs and AI startups that want to build their own language models from scratch or experiment with the tech. Something like this would help small teams build an initial POC and do some experimentation. You have similar issues with robotics projects. It's very expensive for a hobbyist because of the hardware costs, but there's large number of small companies who benefit from open source tech to get their projects started.
- NitpickLawyer 2y ago> Wow, an actual open source language model I find it funny that the AI field has somehow normalised the goalpost moving from capabilities all the way to definitions about open source. And people seem really tribal about it... There absolutely are open source LLMs already. Phi3.5 (MIT), various Mistral models (Apache2.0), various Qwen2 models (Apache2.0) and so on. LLamas are not open source, nor are Gemmas. But to say this is "an actual open source model" is weird nitpicking for the sake of nitpicking, IMO. Requiring the methods and datasets that someone used to create some piece of IP is in no way a requirement for open sourcing said IP. It never has been! Imagine this analogy: A dev comes up with a way to generate source code that solves a real problem. This dev uses a secret seed, that only they know. The dev also uses thousands of hours of compute, and an algorithm that they created. At the end of the exercise they release the results on github, as follows: - here is a project that takes in a piece of text in english, and translates it into french. - the resulting source code is massive. 10 billions LOC. The lines of code are just if statements, all the way down, with some hardcoded integer values. - source code licensed under Apache 2.0, written in let's say python. - users can see the source code - users can run the source code - users can modify the source code and re-release the code Now, would anyone pre LLMs say "this isn't true open source" because it's too complicated? Because no one can reasonably understand the source code? Because it uses hard coded int values? Because it's 10b LOC? Because the dev never shared how they got those values? Of course not. The resulting code would have been open source because Apache 2.0 is open source. It's the same with model weights. Just because they're not source code, and just because you don't know how they were created, it does not mean the weights are not open source. You can see the weights. You can change the weights. You can re-distribute the weights. It's open source. The definition of something being open source does not cover you understanding why the weights are like they are. Nor do they require you having access to the methods of creating those weights. Or datasets. Or whatever the devs had for breakfast.
- diggan 2y ago> that the AI field has somehow normalised the goalpost moving from capabilities all the way to definitions about open source The problem is that Facebook and others are trying to move the goalpost, while others like me would like the goalpost to remain where it is, namely we call projects "Open source" when the required parts to build it on our own machines, is sufficiently accessible. As I probably wouldn't be a developer in the first place if it wasn't for FOSS, and I spend literally all day long contributing to others FOSS projects and working on my own, it's kind of scary seeing these large companies trying to change what FOSS means. I think you're forgetting about the intent and purpose of open source. The goal is that people can run software for whatever purpose they want, and they can modify it for whatever purpose. This is the intent behind the licenses we use when we "create FOSS". This means, in practice, that the source code has to be accessible somehow, so the compiler I have on my computer, can build a similar binary to the one the project itself offers (if it does). The source code has to be accessible so I can build the project, but also modify it for myself. Taking this idea that mostly only applied to software before (FOSS) but applying it to ML instead, it's clear to see what we need in order to 1) be able to use it as we want and 2) be able to modify it as we want. > You can see the weights. You can change the weights. You can re-distribute the weights. It's open source. Right. If I upload a binary to some website, you can see the binary, you can change the binary and you can re-distribute it. Would you say the binary is open source? The weights are the binary in ML contexts. It's OK for projects to publish those weights, but it's not OK to suddenly change the definition and meaning of open source because companies want to look like they're doing FOSS, when in reality they're publishing binaries without any ways of building those binaries with your own changes. Imagine if the Linux kernel was just a big binary blob. Yes, you can change it, re-distribute and what not, but only in a binary-blob shape. You'd be kind of out there if you insist on calling this binary-blob kernel FOSS. I'm sure you'd be able to convince some Facebook engineers about it, seems they're rolling with that idea already, but the rest of us who exist in the FOSS ecosystem? We'd still have the same goalpost in the exact same spot it's been for at least two decades I've been involved.
- benterix 2y agoI'm happy to see a truly open source model. Actually, AMD has excellent reasons to make this kind of development and I hope they continue.
- craftkiller 2y agoI see multiple mentions of NPU on this page, but its still not clear to me: is this something that can finally use the NPU on my processor?
- lhl 2y agoThere's actually seems to be a bunch of stuff now: * https://github.com/amd/RyzenAI-SW https://github.com/amd/RyzenAI-SW - has a list of demos and how to use it directly (including apparently w/ PyTorch and LLMs) * https://github.com/huggingface/optimum-amd https://github.com/huggingface/optimum-amd - can use RyzenAI to use the NPU for HF transformers There's now a Linux driver even https://github.com/amd/xdna-driver https://github.com/amd/xdna-driver although it looks like a sufficiently PITA that I haven't even bothered to try it (my 7940HS only has like 10 TOPS anyway, so not much point even if it worked perfectly).
- n_ary 2y agoNow this here is the beginning on real innovation of AI. With AMD coming in(albeit late and slowly), meta with LLama improving, we will soon see some real adaptation and development in next few thousand days. At this moment, I see OAI as the yahoo of the pre-Google era.
- imjonse 2y ago"next few thousand days" can we stick to years as a unit of measure and not spread Sam Altman's phrase :)
- washadjeffmad 2y agoTwenty two thousand days Twenty two thousand days It's not a lot, it's all we got Twenty two thousand days - Sam Altman?
- raffraffraff 2y agoI was just listening to that song in the last hour
- rsolva 2y agoCan this model run on ollama?
- deleted 2y ago[deleted]
- luyu_wu 2y agoThe section on speculative execution is interesting. "This approach allows each forward pass to generate multiple tokens without compromising performance, thereby significantly reducing memory access consumption, and enabling several orders of magnitude speed improvements." Does anyone know if the "several orders of magnitude speed improvement" is accurate? I'm doubtful. Very interesting though! I'll be playing around with this on the weekend!
- lhl 2y agoOrders of magnitude seems a bit ambitious. The implementation from the DeepMind paper achieved a 2-2.5X https://arxiv.org/pdf/2302.01318 https://arxiv.org/pdf/2302.01318 and most of the tests I've seen [1][2] have been similar, but there are different variations (Medusa, Ouroboros, etc) that can do better/be combined. Recently Together.ai published SpecExec, a SD variant which did claim to get a 10-18X speedups: https://www.together.ai/blog/specexec https://www.together.ai/blog/specexec [1] https://www.reddit.com/r/LocalLLaMA/comments/17h4rqz/speculative_decoding_in_exllama_v2_and_llamacpp/ https://www.reddit.com/r/LocalLLaMA/comments/17h4rqz/specula... [2] https://arxiv.org/pdf/2402.01528v3 https://arxiv.org/pdf/2402.01528v3
- lhl 2y agoBTW, I got a chance to read through the model card and there's a section that shows their SD gains: https://huggingface.co/amd/AMD-Llama-135m#speculative-decoding https://huggingface.co/amd/AMD-Llama-135m#speculative-decodi... - 1.75x-2.80x on MI250 - 2.83x-2.98x on NPU - 3.57x-3.88x on CPU Note they were testing on AMD-Llama-135m-code as draft model for CodeLlama-7b, both of which do similarly badly on Humaneval Pass@1 (~30%), so it's likely if they were using a similarly trained 135m to SD for say, Qwen2.5-Coder (88.4% on HumanEval), the perf gains would probably be much worse.
- deleted 2y ago[deleted]
- highfrequency 2y agoLooks like they are using sixteen $13k GPUs [1] (around $210k hardware) for 6 days of training. Anyone know the recommended cloud provider and equivalent rental price? [1] https://www.wiredzone.com/shop/product/10025451-supermicro-gpu-amdmi250-oam-0029h-graphics-processing-unit-gpu-instinct-mi250-128gb-hbm2e-amd-100-300000029h-10725?srsltid=AfmBOoqud_fxtvuGVSjxigEx4DSMbozywAE5-dI9GfBsBWNDzLE9-wN2 https://www.wiredzone.com/shop/product/10025451-supermicro-g...
- wmf 2y agoHot Aisle seems to the (only?) place to rent AMD. (Ryan, please don't spam this thread. It's not a good look.)
- lhl 2y agoRunpod.io rents the next-gen MI300X's for $4/hr, although since they also rent H100's for $3/hr (that are easier to work with/faster for training) it might be more of a novelty.
- highfrequency 2y agoI thought the whole selling point of AMD GPUs was that they were a lot cheaper than Nvidia GPUs?
- deleted 2y ago[deleted]
- dagmx 2y agoCheaper for the cloud company. But that doesn’t always translate to cheaper for the end user. Maybe they cost more to run or maybe there’s fewer of them so they’re more expensive to book?
- knotimpressed 2y agoAt least a couple years ago, a big advantage of Nvidea cards was how much cheaper they were to run power wise-often the dies that made it into cloud level cards would be binned consumer dies. Not sure if that’s still the case, but I’d say it’s plausible.
- Decabytes 2y agoSince most people can’t run these LLMs locally, I wonder what a model would look like where we have hyper tuned models for specific purposes, IE a model for code, a model for prose, etc. you have a director model that interprets what downstream model should be used and then it runs that. That way you can run the model locally, without needing beefy GPUs. It’s a trade off of using more disk space vs needing more vram
- wmf 2y agoThe whole point of this model is that it's so tiny that even a weak RPi could run it. Apple has also done some interesting work with a common <4B base model that is customized with different LoRAs for different purposes.
- kristianp 2y agoAbout 3b if you're referring to this: https://machinelearning.apple.com/research/introducing-apple-foundation-models https://machinelearning.apple.com/research/introducing-apple...
- Philpax 2y agoYou're essentially describing Apple Intelligence :-) https://machinelearning.apple.com/research/introducing-apple-foundation-models https://machinelearning.apple.com/research/introducing-apple... (see Model Adaptation)
- fennecbutt 2y agoA rip off of LLMs and loras. Wrapping it in a shiny sounding name for the normies doesn't mean they contributed anything to the space.
- Philpax 2y agoThey're not hiding anything; they've very clearly described what they've done and how they've done it. They've branded their specific architecture and integration, which allows me to easily refer to it as an example. I understand that it's easy to be cynical about Apple's approach to product development, but it seems unwarranted in this case.
- bjt12345 2y ago> [1] The training code for AMD-135M is based on TinyLlama, utilizing multi-node distributed training with PyTorch FSDP. I thought PyTorch didn't work well with AMD architecture, and read of many people using JAX instead?