6 ms·
MiniMax-M1 open-weight, large-scale hybrid-attention reasoning model
- deleted 1y ago[deleted]
- swyx 1y ago1. this is apparently MiniMax's "launch week" - they did M1 on Monday and Hailuo 2 on Tuesday (https://news.smol.ai/issues/25-06-16-chinese-models https://news.smol.ai/issues/25-06-16-chinese-models). remains to be seen if they can keep up the pace of model releases for the rest of this week - these 2 were big ones, they aren't yet known for much else beyond llm and video models. just watch https://x.com/MiniMax__AI https://x.com/MiniMax__AI for announcements. 2. minimax m1's tech report is worthwhile: https://github.com/MiniMax-AI/MiniMax-M1/blob/main/MiniMax_M1_tech_report.pdf https://github.com/MiniMax-AI/MiniMax-M1/blob/main/MiniMax_M... while they may not be the SOTA open weights model, they do make some very big/notable claims on lightning attention and their GRPO variant (CISPO). (im unaffiliated, just sharing what ive learned so far since no comments have been made here yet
- behnamoh 1y ago> they did M1 on Monday and Hailuo 2 on Tuesday It would've been fun to see them name their models like Apple chips: M1, M1 Pro, M1 Ultra.
- plufz 1y agoYeah MiniMax M1 certainly directed my thoughts to Mac mini M1. :)
- drewbitt 1y agoThey are also known for audio models, having the best TTS on some leaderboards (and my personal favorite) https://artificialanalysis.ai/text-to-speech/arena?tab=leaderboard https://artificialanalysis.ai/text-to-speech/arena?tab=leade...
- noelwelsh 1y agoA few thoughts: * A Singapore based company, according to LinkedIn. There doesn't seem to be much of a barrier to entry to building a very good LLM. * Open weight models + the development of Strix Halo / Ryzen AI Max makes me optimistic that running great LLMs locally will be relatively cheap in a few years.
- rfoo 1y ago> A Singapore based company, according to LinkedIn Nah, this is a Shanghai-based company.
- diggan 1y ago[flagged]
- noelwelsh 1y agohttps://en.wikipedia.org/wiki/MiniMax_(company) https://en.wikipedia.org/wiki/MiniMax_(company)
- diggan 1y agoWikipedia in itself is no source, and after reading parents message I went there to check to and surprise surprise, neither of the statements have sources attached to it. None of the linked articles have any information about where their headquarters is either. If someone knows of a trustworthy article that states it outright, please feel free to share.
- noelwelsh 1y agoI'm the OP who claimed it was Singaporean, after checking LinkedIn. I then found the Wikipedia page, which I posted above. Amongst the comments here there is also a link to a Bloomberg article about a potential IPO. I don't have a dog in the race. Just passing on what I found.
- vintermann 1y ago"We publicly release MiniMax-M1 at this https url" in the arxiv paper, and it isn't a link to an empty repo! I like these people already.
- npteljes 1y agoThis is stated nowhere on the official pages, but it's a Chinese company. https://en.wikipedia.org/wiki/MiniMax_(company) https://en.wikipedia.org/wiki/MiniMax_(company)
- iLoveOncall 1y agoWhy would you expect them to mention that on their project's page?
- noelwelsh 1y ago1. It's conventional to do so. 2. It's a legal requirement in some jurisdictions (e.g. https://www.gov.uk/running-a-limited-company/signs-stationery-and-promotional-material https://www.gov.uk/running-a-limited-company/signs-stationer...) 3. It's useful for people who may be interested in applying for jobs
- iLoveOncall 1y ago1. No it's not. Top GitHub repository from Google as an example: https://github.com/google/material-design-icons https://github.com/google/material-design-icons I think you'd actually be hard pressed to find a single repository where the company that owns it lists where they are registered. 2. This is a requirement for companies registered in the UK. You should also read your own link, it doesn't say anything about the company's presence on 3rd party websites. 3. This is such a remote reason it's laughable, there are plenty more things that are more relevant to potential job applications, such as whether they are hiring at all or not. You just want them to mention it because it's a Chinese company. If they were American, Mexican, German or Zimbabwean you wouldn't give the slightest fuck.
- noelwelsh 1y agoOP said "official pages", which I took to mean the company website: https://www.minimax.io/ https://www.minimax.io/ Also, thanks for putting words in my mouth. If they were Mexican or Zimbabwean I would find it very interesting to see a roughly SOtA model coming from that country.
- htrp 1y agoThey apparently building buzz for an IPO https://www.bloomberg.com/news/articles/2025-06-18/alibaba-backed-ai-dragon-minimax-is-said-to-plan-hong-kong-ipo https://www.bloomberg.com/news/articles/2025-06-18/alibaba-b...
- markkitti 1y agoPlease come up with better names for these models. This sounds like the processor in my Mac Studio.
- chvid 1y agohttps://en.wikipedia.org/wiki/Minimax https://en.wikipedia.org/wiki/Minimax They named themselves after a classic ai algorithm.
- JoeDaDude 1y agoAs best I can tell from a gloss-over read, it doesn't use anything like the Minimax algorithm. Astute readers are aware that one of the first applications of Minimax was in an AI chess program designed by Claude Shannon. https://en.wikipedia.org/wiki/Claude_Shannon#Shannon's_computer_chess_program https://en.wikipedia.org/wiki/Claude_Shannon#Shannon's_compu...
- bjord 1y agoit's the name of the company
- npteljes 1y agoThe company supplies contemporary AI solutions, like LLM and video generation. The name is just a reference, like in the case of Tesla, or like how there is a kaliapparat in the American Chemical Society logo.
- thawab 1y agoDoes facebook use llamas in their model? it's a name, it doesn't have to be 100% true to its meaning.
- badc0ffee 1y agoBut then there's the "M1" part.
- diggan 1y agoAlso sounds like my long lost dog whose name was Max but he was tiny. Absolutely horrible name, borderline criminal I say.
- reedlaw 1y agoIn case you're wondering what it takes to run it, the answer is 8x H200 141GB [1] which costs $250k [2]. 1. https://github.com/MiniMax-AI/MiniMax-M1/issues/2#issuecomment-2982368797 https://github.com/MiniMax-AI/MiniMax-M1/issues/2#issuecomme... 2. https://www.ebay.com/itm/335830302628 https://www.ebay.com/itm/335830302628
- incomingpain 1y agoThat's full quantization. If you run Q4 or Q8 you can run this on <$10,000 equipment.
- cma 1y agoAnd if you add in heavy sparsification it should fit and run on a raspberry pi.
- rvz 1y agoSo in around 6 months, we will see that the person who bought this H200 in the listing just got scammed for $250k and will realize that you just needed specific quantizations to the model and a few optimizations to run locally. Unless they want to train their own model, buying this for inference for $250k is unnecessary and still isn't enough for a full production deployment.
- vFunct 1y agoIt's already sparsified from the 150T parameter model..
- yorwba 1y agoIt took me several hours to realize that 150 trillion parameters is a reference to the number of synapses in a human brain.
- deadbabe 1y agoNo point in running anything but full quantization.
- insider123 1y ago[dead]
- deleted 1y ago[deleted]
- b0a04gl 1y agoif they trained this scale without western cloud infra, i'd want to know what their token throughput setup looks like
- econ 1y agoSneakernet
- jaggs 1y agoThey trained on 512 H800 GPUs for three weeks, equivalent to around half a million dollars. https://xcancel.com/MiniMax__AI https://xcancel.com/MiniMax__AI
- yorwba 1y agoThat is for the reinforcement learning part. The base model was likely trained on more GPUs for significantly longer.
- killerstorm 1y ago> "In our attention design, a transformer block with softmax attention follows every seven transnormer blocks (Qin et al., 2022a) with lightning attention." Alright, so it's 87.5% linear attention + 12.5% full attention. TBH I find the terminology around "linear attention" rather confusing. "Softmax attention" is an information routing mechanism: when token `k` is being computed, it can receive information from tokens 1..k, but it has to be crammed through a channel of a fixed size. "Linear attention", on the other hand, is just a 'register bank' of a fixed size available to each layer. It's not real attention, it's attention only in the sense it's compatible with layer-at-once computation.