5 ms·
Magma: A foundation model for multimodal AI agents
- erikig 2y agoThe multimodal capabilities especially on next action prediction are quite impressive; watching the github to see if & when they'll open source this: https://github.com/microsoft/Magma https://github.com/microsoft/Magma Also, I wonder why they named it Magma?
- lanternfish 2y ago`M(ultimodal) Ag(ent) [ma]` maybe
- jwyang 2y agoGood catch! A minor correction: Magma - M(ultimodal) Ag(entic) M(odel) at M(icrosoft) (Rese)A(rch), the last part is similar to how the name Llama came out, :)
- throw310822 2y agoHow many 'M's in "Magma"? ;)
- gostsamo 2y agoI know that AWS have an AI product for foundational models called Bedrock so MS might've decided to go even deeper.
- pyinstallwoes 2y agoWow
- jauntywundrkind 2y agoFrom the news section of that github README: > [2025.02.19] We will be releasing our code, model and UI navigation demo at MSR Forum on 02.25 next Tuesday!
- deleted 2y ago[deleted]
- Lockal 2y agoA bit sad that they reused name of https://icl.utk.edu/magma/ https://icl.utk.edu/magma/ (Matrix Algebra on GPU and Multi-core Architectures). This library is already heavily used in machine learning, for example, it is included in every pytorch-based project.
- fullstackchris 2y agolooking at the paper some other agentic models they compared to were named LLaVA... maybe it's just a play on words
- digitaltrees 2y agoThey need to build an epistemology and theory of mind engine into models. We take it for granted when dealing with other humans that they can infer deep meaning, motivations, expectations of truth vs fiction. But these agents don’t do that and so will be awful collaborators until those behaviors are present
- szundi 2y ago[dead]
- kvirani 2y agoWe're in the 56k modem era of generative AI, so I wouldn't be surprised if we had that in the next few years, or weeks.
- kolinko 2y agoDid you read any research on theory on mind and models? Since gpt4 they were tested using similar metrics to humans and it seems the bigger models “have” it
- MattGaiser 2y agoAnd it causes a ton of chaos that we do take that for granted between humans. The annoying collaborator is the person who takes information for granted.
- energy123 2y agoTheory of mind should naturally emerge when the models are partly trained in an adversarial simulation environment, like the Cicero model, although that's a narrow AI example.
- jwyang 2y agoThanks for your great interests on our Magma work, everyone! We will gradually roll out the inference/training/evaluation/data preprocessing code on our codebase: https://github.com/microsoft/Magma https://github.com/microsoft/Magma, and this will be finished by next Tuesday. Stay tunned!
- dr_dshiv 2y agoHow far are we from making peanut butter sandwiches? Is that a valid benchmark to look towards, in this space?
- yurimo 2y agoMultimodal agents notoriously fail at long horizon tasks, how does Magma perform on it?
- jwyang 2y agovery good question, now we are mainly focusing on building the foundtion for multimodal perception and atomic action taking. Of course, integrating the trace-of-mark prediction for robotics and human video data enhances the model's medium length reasoning but this is not sufficient for sure. The current Magma model will serve as the basis for our next step, i.e., longer horizong reasoning and planning! We are exactly looking at this part for our next version of Magma!
- ygouzerh 2y agoThe rate of progress on multimodal agents is impressive. OpenVLA was released in June 2024 and was state of the art at that time... 8 months later, on tasks like "Pick Place Hotdog Sausage" the success rate is passing from 2/10 to 6/10
- paulluuk 2y ago"Pick Place Hotdog Sausage" is such a bizarre name, though. Is it meant to be human readable? AI-readable? Just a label for the researchers? Same with "Put Mushroom Place Pot". As far as I can see both labels are only used in this Magma paper, nowhere else that Google can find.
- ekidd 2y ago"Pick & place" is a term for a kind of robot that can pick up scattered items from a conveyor belt and arrange them in a regular fashion. The really fast multi-arm versions can be hypnotic to watch. You can see an example at 1:00 in this video: https://youtu.be/aPTd8XDZOEk https://youtu.be/aPTd8XDZOEk The limitation of industrial pick & place robots is that they're configured for a single task, and reconfiguring them for a new product is notoriously expensive. Magma's "pick & place" demo is much slower and shakier than a specialized industrial robot. But Magma can apparently be adapted to a new task by providing plain English instructions.
- fewdaysto2025 2y ago[dead]
- sorz 2y agoIn the mug-scrubbing video, the person clearly pretends to wash the cup but does not seem to want to get their hands wet anyway. I'm curious as to when models can figure out that subtle thing.
- regularfry 2y agoYou want that to still work so that the human can demonstrate an action without putting themselves in the path of a danger to squishy human bits that the robot is safe from.
- funnyAI 2y agoIt's all probabilistic, my guess. I.e. model produces probabilities for a set of actions from the same video. Even pretended action may look more like it than anything else. Thus getting higher probability.
- funnyAI 2y agoJust wondering if there is any research done in incremental training? That could be used in robots as alternative to RAG.
- Oras 2y agoLooking at industrial robots they don't mimic how humans do things, and hence, they are efficient. That's why I don't understand how these propsals to teach robots how humans do things will make any sense. To have robots at homes, they will need their tools to be efficient. It will not be the same washing machine, oven, or dishwasher that we use now, there will be new ones made for robots.
- 4ndrewl 2y agoThis. The paucity of imagination in the AI space is mind-numbing.
- ben_w 2y agoFashions in AI. Everyone is piling on Transformers and Diffusion (and in robotic, humanoids) today; but for most of the history of AI, we've been making things so simple they can only mono-task, and the only way to make commercial sense of that is to be much more efficient (on one of the many axies) than humans. Now we have models that seem (at least at first glance) to cover the full breadth of what humans can do, so the question has become: can we make them perform at a decent skill level, rather than like someone who is book-smart enough to pass the tests but has almost no real experience of anything.
- whatever1 2y agoBut Humans generalize very well across tasks. You can have an employee driving a forklift, then stop pick-up a pallet that blocks his way and continue.
- Oras 2y agoAnd robots will not do that either, what if the employee used hearing to determine if there is a hazard (another moving vehicle around) before jumping to pick a pallet? How would the robot know by just “looking”? How to prioritise visuals, audio, sense … etc?
- 2y ago
- Mizza 2y agoHave any multimodal models been reasoning-trained yet?
- coder543 2y agohttps://platform.openai.com/docs/models/#o1 https://platform.openai.com/docs/models/#o1 > The latest o1 model supports both text and image inputs
- macrolime 2y agoBut not multimodal reasoning, the intermediate and output tokens are text only, at least in the released version, they probably have actual multimodal reasoning that's not been shown yet, as they already showed gpt-4o can output image tokens,but that's not been released yet either.
- coder543 2y agoThat wasn’t the question… they asked if any multimodal models had been reasoning trained. o1 fits that criteria precisely, and it can reason about the image input. They didn’t ask about a model that can create images while thinking. That’s an entirely unrelated topic.
- bosky101 2y agoSpent 10 mins on the website, all the examples are single agent examples. There is 0 value add for yet another wrapper on an openai call, parading as an agent. The whole point of agents is knowing what to do among potentially 100's of intents and actions. Disappointing.
- lelag 2y agoReally interesting model, I'm looking forward to play with it. But what I want is a multimodal agent model capable of generating embeddings for a humanoid control model like Meta motivo[0] rather than directly outputting coordinates. Meta motivo is still a toy model, trained on the SMPL skeleton, which lacks fingers which limits its capabilities beside having some fun with it. They could have used a more advanced based model, SMPL-X, which includes fingers, but there isn’t enough open motion data with precise finger motion to train a robust manipulation model anyway. Most existing motion datasets come from academic motion capture setups, which are complex and not focused on manipulation tasks (and also pretty old). I believe this gap will be filled by improvements in 3D HPE from 2D video. With access to thousands of hours of video, we can build large-scale motion datasets covering a wide range of real-world interactions. This will enable training the two components needed for dexterous humanoid robots: the agentic model that decides what actions to take and generates embeddings that can be read by a control model that accurately models hand and finger joint movement. Given the rapid progress in the capabilities of SoTA 3D HPE from 2D video, and the vast amount of videos online (Youtube), I expect we will see humanoid robots with good manipulation capabilities it the not so distant future. [0]: https://github.com/facebookresearch/metamotivo https://github.com/facebookresearch/metamotivo
- michaelbuckbee 2y agoTrying to wrap my head around this - are you saying that those models are trained around the concept of fingers (some kind of physical manipulators with set dimensions)?
- lelag 2y agoThe SMPL-x body model, a standard in this academic field does model fingers https://smpl-x.is.tue.mpg.de/ https://smpl-x.is.tue.mpg.de/ The issue is that there are much less dataset available for it than for the simplier SMPL model. Regarding fingers, you already have "dumb" models like https://github.com/google-deepmind/mujoco_mpc https://github.com/google-deepmind/mujoco_mpc which can control finger mouvement to achieve specific task. Look at this video to see it action: https://www.youtube.com/watch?v=2xVN-qY78P4&t=387s https://www.youtube.com/watch?v=2xVN-qY78P4&t=387s Pretty cool stuff.
- bilsbie 2y agoWhy do no multimodels fluidly create images. It seems like they pass off to another model to generate images? They’re not really aware what’s in the images they make and the can edit images in place.
- dartos 2y agoWhat do you mean by fluidly?
- kittikitti 2y agoThese benchmarks are not really representative of what agents are capable of. The slow process of checking the weather through UI elements is not a good use case which is non-peer reviewed paper showcases.
- ok123456 2y ago[flagged]
- deleted 2y ago[deleted]
- bob_theslob646 2y agoAm I the only one that read that title in Dr.Evil's voice? All kidding aside. This looks promising