6 ms·
MAI-Thinking-1
https://microsoft.ai/wp-content/uploads/2026/06/main_20260602_2.pdf https://microsoft.ai/wp-content/uploads/2026/06/main_2026060...
Launching seven new MAI models: https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/ https://microsoft.ai/news/building-a-hillclimbing-machine-la...
- simjnd 4mo agoAbsolutely disgusting scroll jacking, even when "Accessibility mode" is turned on
- dang 4mo agoI'm sure most of us agree, but: "Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- simjnd 4mo agoForgot about this, my bad!
- wmf 4mo agoAt least there shouldn't be any complaints about benchmaxing this time.
- i_have_an_idea 4mo agoJust because it is performing rather poorly by comparison, it doesn’t mean it isn’t benchmaxxed. It can still be worse than it appears.
- wasabi991011 4mo agoIt isn't benchmaxxed because they are using human preference as an evaluation.
- pixeldash928 4mo agoLooks like the OAI divergence is finally taking place. Seems like the comparisons are mainly with Opus 4.6 and GPT 5.4 though. Still, exciting to see a new frontier player.
- i_have_an_idea 4mo agoIs it a frontier player though, or perhaps a new benchmaxxed model? People were saying similar things about Grok but it ultimately amounted to little.
- wasabi991011 4mo ago"preferred by humans over Sonnet 4.6" makes it pretty clearly not benchmaxxed though. At least when you define benchmaxxed as "good in benchmarks but not human preference".
- dude250711 4mo agoPost 4.6 Anthropic models do not exactly have a stellar reputation, so that choice is smart.
- keeda 4mo ago> Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Shots fired? It would be interesting to see how far "clean data" can go on the scaling laws.
- onlyrealcuzzo 4mo agoI'm interested how much "Clean Data" is synthetic data from "unclean" models...
- xavriley 4mo ago“ We trained it from the ground up on enterprise grade, clean and commercially licensed data, without distillation from third-party models.”
- azinman2 4mo agoaka all of GitHub OSS
- ChicagoDave 4mo agoYeah this is exactly what I was thinking.
- rurban 4mo agoNot OSS only, likely also the enterprise private repos, with a lot of business secrets.
- ertgbnm 4mo ago> with AI-generated content excluded from pre-training. > without distillation from third-party models sounds like zero unless they are lying.
- bossyTeacher 4mo ago7 modes launched. 5 models in the dropdown. Only 4 actually usable :( About time Microsoft joined the fray. After the OpenAI divorce, it really looked like Microsoft was going to become another Uber.
- giancarlostoro 4mo agoThey still own 27% of OpenAI, this IPO will feed them a lot of easy cash.
- lordmauve 4mo agoWe need to see DeepSWE scores. SWE Bench Pro is junk.
- kstenerud 4mo agoThey've hijacked scrolling. They've hijacked the spacebar. It flickers like crazy when I try to move through the article. Trying to get through it is an exercise in madness.
- AirMax98 4mo agoI normally don't comment on matters of taste like this, but wow this is brutal. It's like someone threw the site in a vat of molasses.
- t-sauer 4mo agoI do not understand how scroll hijacking is still a thing. Who thinks this is a better experience?
- maelito 4mo agoDesigners.
- bensyverson 4mo agoAs a designer, let me tell you: scroll jacking is not good design
- aniceperson 4mo agothere is also a gap between the header and the top of the page... they should ask the ai to make it better a few more times...
- blisstonia 4mo agoI gave up after the first scroll.
- grassfedgeek 4mo agoEven without flicker it is very distracting. Why do people think this is a good idea?
- deleted 4mo ago[deleted]
- BeetleB 4mo agoBased on the first table, why would I pick this over GLM?
- missedthecue 4mo agoBecause your employer might make you exclusively use enterprise copilot.
- BeetleB 4mo agoAs long as my employer is footing the bill, fine. For personal stuff this release is not noteworthy.
- hartator 4mo agoI like it so much when a website hijacks the way my scroll works. This is truly innovative.
- campital 4mo agoYeah, you might get disoriented and throw up if they didn't smooth it out.
- vcryan 4mo agoIt really looks like they used Claude to design this webpage. I guess the color taupe it the marker of good AI today.
- Handy-Man 4mo agoInflection AI
- Alifatisk 4mo ago> MAI-Thinking-1 is built with enterprise readiness in mind. It supports long context with a 256k token window Isn’t 1M becoming the norm?
- stingraycharles 4mo agoYes it is, but I can imagine that they want to start out a bit smaller to see how well things scale, and/or did not yet have the time to work on optimizing for the large context windows.
- droidjj 4mo agoI struggle to get quality results from the frontier models at contexts > 256k anyway.
- stingraycharles 4mo agoYup, same experience, it’s because the attention basically has exponential complexity. So at large context windows, they need to compress the attention (eg group multiple tokens together), when then leads to loss in accuracy. It’s almost always better to keep your context windows small.
- vb-8448 4mo ago1M it's only marketing, in my experience above 150k quality noticeable drops. Claude code will suggest you to start a new session or compact if you go above 100k.
- Bolwin 4mo agoIn my experience above 60k quality noticeably drops. 30k for open source models
- __natty__ 4mo agoIt's good there is a new player on the market, I take benchmark tables with a grain of salt, however. Speaking about model presentation it's funny to see how clearly their website is inspired by other AI company blogs with extra innovation of hijacked scrollbar.
- deleted 4mo ago[deleted]
- Centigonal 4mo ago> MAI-Thinking-1 is a 35B-active, ~1T-total parameters, sparse Mixture of Experts model, a smaller inference footprint than much larger models. This seemingly nonsensical sentence (of course this will have a smaller inference footprint than larger models) suggests this model's competitors have larger inference footprints and total parameter sizes.
- dr_kiszonka 4mo agoWhen would a larger model have a smaller inference footprint? If the larger was MoE and the smaller was dense?
- Centigonal 4mo agoyes, MoE reduces the inference compute requirements (inference memory reqs remain the same)
- rajveerb 4mo agoAs someone who has spent quite a lot of time on inference, I would a add a small note: Deployment looks very different for MoE than dense style models so I would say that it is more nuanced than "inference memory reqs remain the same". Memory can be very different for MoE style models.
- kaicianflone 4mo agoIs that a pretext zoom effect when changing screen dimensions? Very cool.
- gigatexal 4mo agoAnyone believing those benchmark numbers from a 35B model?
- jeffdn 4mo agoIt says right at the top, 35B active, 1T total.
- jampekka 4mo agoThe benchmarks are a bit of a disaster? It's at about DeepSeek V3.2 level, but with about 50% more parameters. Loses handily to the also smaller GLM-5.1, and even worse to the similarly sized Kimi K2.6.
- sailingparrot 4mo agoYes and no. Yes from a user PoV, I don't really see a great reason to use this other than for enterprises that care about using a model not trained on copyrighted data (not sure what the market really is for this anymore, feels like this concern has been forgotten by most customers). From a strategic PoV for MS, all the models you cited are distilling GPT/Claude/Gemini and wouldn't be anywhere as good as they are without this distillation, which in turn means you are dependent on OAI/Anthropic/G first shipping a good model to generate data for your training. This MAI model is trained from scratch with no synthetic data or distillation. So in term of benchmark its obviously much harder to get strong score and thus not a disaster if they can keep on improving.
- usef- 4mo agoThey claim to not be training to the benchmarks at all. It'll be interesting to see how it stacks up in actual use.
- nojito 4mo agoNo distillation. Comparing it to DeepSeek or GLM doesn't make much sense.
- andai 4mo ago[dead]
- dang 4mo agoRelated ongoing thread: MAI-Code-1-Flash - https://news.ycombinator.com/item?id=48374466 https://news.ycombinator.com/item?id=48374466 - June 2026 (131 comments)
- adt 4mo agohttps://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- euphetar 4mo agoHonestly, a lame release of mediocre models. I was most excited about the "frontier tuning." Like, it will actually watch you do stuff and learn to do it for you? That would be actually interesting. But no, it's just a data labelling interface: https://learn.microsoft.com/en-us/microsoft-365/copilot/copilot-tuning-overview https://learn.microsoft.com/en-us/microsoft-365/copilot/copi.... You have to provide the instruction and give feedback and there is a whole UI with hour-lonf wait between steps. So basically they want you to do the labelling to train a model, or at least that's how it looks from the outside Also the mission statement of Humanist AI is the most boring, but tries to sound way too grand. Like "all the cool labs have a mission statement, so we should also have one" vibes
- throwawayffffas 4mo agoMeh, 1T parameters no weights? I am running a better model right now on 40GB of VRAM.
- basilikum 4mo agoWhy is microsoft.ai hosted on an ASN called WPEngine and not by Microsoft themselves?
- aesthesia 4mo agoWhat's interesting is that although they don't seem to be releasing the model weights, they have published a technical report (https://microsoft.ai/wp-content/uploads/2026/06/main_20260602_2.pdf https://microsoft.ai/wp-content/uploads/2026/06/main_2026060...) that's more extensive than the typical open-weights model gets.
- deflator 4mo agoDoes this mean that work created with it can be copyrighted? Since the courts ruled that the inclusion of pilfered IP was the reason other model's work cannot be copyrighted, I would think so! In that case, this is a completely different beast. It can maybe be used for things that need a durable copyright.