8 ms·
Inkling: Our Open-Weights Model
- bobkb 3mo agoHappy to see an open weight model ! This has all the right ingredients for success.
- kancha 3mo agoNot compared against Gemma 4? That is a big omission.
- ls_stats 3mo agoAmerica needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice. Thinking Machines might be it.
- verdverm 3mo agoIts not as good as GLM 5.2 for agentic workflows while also being bigger. Competition is going to be ruthless because the super low cost to switching. There is also AllenAi in the US, but they have yet to produce a model at this scale. Thankfully, new contenders can come out of nowhere and do well, as long as they can produce a competitive model.
- InsideOutSanta 3mo ago> Its not as good as GLM 5.2 for agentic workflows while also being bigger GLM 5.2 underwent extensive post-training and iteration since its original release to reach its current state. This seems like an extremely strong model for a first release, with a lot of potential for improvement, just like DS4. Sometimes I wish Meta had stuck with Llama 4 a bit longer to see how much further it could be pushed.
- verdverm 3mo agoThis is a great point
- hirako2000 3mo agoLlama 4 wasn't deemed a success, and Meta pivoted away as its now former head of AI couldn't demonstrate, nor even showed interest in, business profit. They overspent on llama 3 anyway so money ran dry, LeCun is good at running research, but budgets didn't stretch. Meta isn't investing in frontier big models anymore.
- nl 3mo ago> Meta isn't investing in frontier big models anymore Yes they are. Meta Muse is their attempt. It's below frontier performance at the moment but they are spending on getting there.
- nl 3mo agoLlama 4 was a bad architecture. Meta Spark is moderately promising but of course closed source.
- deleted 3mo ago[deleted]
- gkapur 3mo agoIt could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc. That said, the fine-tuning API + open weight model at least is a semblance of a viable business that could work so I will be curious about it. I'm not sure the synergy is fully there (why is someone with an open weights model privelaged to fine-tune it better if it's just QLora or Lora) but let's see!
- andriy_koval 3mo ago> It could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc. my bet is that Chinese government fund Chinese models way more compared to what those companies receive (except llama, which is outdated but was strong foundation at its time)
- deleted 3mo ago[deleted]
- bostonvaulter2 3mo agoWhat is the business model for an open weight model?
- deleted 3mo ago[deleted]
- matsur 3mo agoThinky has a potential answer in Tinker — give away the weights and charge for the SFT (and maybe RL down the line) to make the model more capable for specific tasks.
- andriy_koval 3mo agoSFT/RL can be done without parent company.
- alightsoul 3mo agoBut not conveniently. This is why you outsource to vendors. Not everyone will do it
- ergocoder 3mo agoThe same business model that Deepseek is using. Open-source models + services. This is more attractive because it doesn't lock in the vendors. If I grow larger, I can decide to deploy the open-source models.
- tyre 3mo agoSo they're constantly hemorrhaging their most valuable clients? Tech history is littered with the corpses of "open source but we sell hosting" services. Models are so expensive to train, you can't be losing the big clients once they get super profitable.
- MikeTheGreat 3mo ago
- joshmarlow 3mo agoI don't hear about them a lot but it looks like arcee.ai is aiming to be just that. Here are some of their current open weight offerings: https://www.arcee.ai/open-source-catalog https://www.arcee.ai/open-source-catalog
- wgd 3mo agoYou don't hear about them much because their models aren't really competitive. I really wanted to try Trinity Large as a daily-driver in the MiniMax M2 sort of niche but I couldn't make it through a single day. The models need another couple point releases worth of post-training to make useful agents and if memory serves they weren't any less slopped in writing style and those are really the only two things people look for in models.
- tonic_note 3mo agoisn't that what Reflection is trying to be?
- UncleOxidant 3mo agoHopefully they'll release some smaller models (<100B) that we can run on home hardware at faster than 10tok/s.
- fastball 3mo agoWhat about Meta?
- insane_dreamer 3mo agoIt’s what Meta was supposed to do but Llama fell of the wagon. There’s also Prism
- jauntywundrkind 3mo agoAlso the fact that China is building solar power like crazy: that makes it fantastically more well spirited an endeavor to wish well.
- codemog 3mo agoI’m trying to be charitable but your comment reads as “China bad” propaganda to me. Who cares that DeepSeek and Z.ai are Chinese companies?
- nodja 3mo agoIt doesn't matter until it does. If the chinese government decides that open weight model releases are no longer allowed, that's a lot of companies that can't release new models. Same with the US government, etc. Having diversity is important.
- jas- 3mo agoIt's a similar problem the human DNA solved by telling our teenage selves that our parents are dumb and we needed to move to a new tribe. Genetic diversity, but a digital equivalent.
- sschueller 3mo agoHowever unlike the US models, China banning the release of new models would not break existing ones. Betting on US models only can get you locked out in just a few hours.
- Systemerror7A69 2mo agoNo, Open weights US models would not break as well - this isn't related to China or USA, it's about Open Weights and the fact that you can download the models.
- MattDamonSpace 3mo agoChina’s got absolute control over its outputs. For America to have any guarantees around long-term availability of OW models, they need domestic production. FWIW this is the same logic for China’s need for their own OW models
- HaloZero 3mo agoI think practically every government will want to put restrictions on private companies building models. Frankly the EU and the US will practically be less involved and have more pushback from the public in this than China. I think that’s less “China bad” than recognizing that China is a more authoritarian state and has far more proclivity to interfere than western states. Maybe I’m wrong? What does deep seek say about Tiananmen square in 1989?
- soundworlds 3mo agoAllenAI is also one to keep your eye on. Founded by Paul Allen of Microsoft, they are one of the best teams working towards truly transparent / open AI (including training data)
- maxloh 3mo agoI love Allen AI. I find it wonderful that, as a non-profit, they are only one to two years behind SOTA models that cost billions of dollars to build, if not more.
- FrankBooth 3mo agoI love the tasteful thickness of Paul Allen’s model.
- nl 3mo agoAllenAI is great, but they don't have the budget or remit to build large models.
- timmg 3mo agoI wonder if the recent sale of the Seahawks will change that. IIRC, ~$10B and all is supposed to go to charity. Not sure how much of that will go to AllenAI, though. (If any.)
- nl 3mo agoNo reason to think it will. Paul Allen died after donating the money to create AllenAI and I don't think there are any links. Hopefully it somehow works out though!
- ReptileMan 3mo agoI will wait for Modernist AI by Myhrvold.
- mstank 3mo agoDo you think American companies will secretly distill frontier models to build open weight ones?
- icase 3mo agoi refuse to root for our enemies, but otherwise you are correct.
- upmind 3mo agounlikely I think, they're likely doing this to garner some interest in their company but they seem pretty interested in revenue (judging by the companies they're working with)
- xnx 2mo ago> root for open chinese models to win What does "winning" mean to you?
- alansaber 3mo agoI never thought i'd see the day they released a model, rather than a blog post. The Figure 3 demo being a screencap of chrome in localhost made me feel better about myself. Jokes aside, best western open weights model- very cool.
- pr337h4m 3mo agoThey are one of the few labs (perhaps even the only one at this level) that are doing something both unique and useful, rather than simply imitating what the others are doing: https://thinkingmachines.ai/blog/interaction-models/ https://thinkingmachines.ai/blog/interaction-models/
- verdverm 3mo agoIf it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model? Maybe for the multi modal?
- gkapur 3mo agoIf they have a really seamless fine-tuning experience and maybe can help you extract the data you need to FT (which is one of the big challenges in actually getting fine-tuning democratized), maybe you would use it because "Tinker" defaults to it. The model could also be more flexible for non-coding use-cases (they show the results for reasoning being strong) so maybe the argument is to use it for non-coding use-cases to drive relatively deterministic conclusions for non-coding agents (they have also done some determinism work on kernels, which could be useful in pulling on that thread of deterministic models that are fine-tuned for everything that is not writing code.) That said, I'm not sure how much all the work they have done actually synergizes or if the market size (at least in the short to medium term) is big enough for a huge outcome from the company's current valuation with those bets as the enterprise agent estate is taking a while to evolve. Hence companies like Anthropic and OpenAI are throwing tons of consulting money at the problem.
- Aurornis 3mo ago> If it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model? The benchmarks never tell the full story. Some of the open weights models have been benchmaxxed for a while. Their utility on real work can be different than the benchmark number. The multimodal input is also a big deal. Having vision input is really helpful for a lot of tasks.
- Reubend 3mo agoSeems like this is particularly good at instruction following, but not as strong at coding as others. It's always great to get more diversity of open weight models though! I'll need to test this out to see what its "personality" is like.
- jakswa 3mo agoseems pretty dang snappy and I like it's tone/personality so far. > look at today's hackernews frontpage and generate me a daily briefing report (create an artifact) to read later for today's nerd news https://chat.home.jake.town/artifacts/019f679d-99e5-7000-b02f-fb64e97169d3 https://chat.home.jake.town/artifacts/019f679d-99e5-7000-b02...
- arrowleaf 3mo agoThis is the best voice/tone I've seen from any model so far. It's using filler words and phrases in places that normal people would put them, rather than sounding like a corporate customer support agent!
- ianbutler 3mo agoIt's nice to see a strong long context open weights model that is multi-modal. There are many applications that will benefit from the strength in audio here and until z.ai and co work in visual this could be very strong for general agentic applications, though I see there's a bit of weakness in the benches for areas that might make that less true. Like all models need to slap it in your harness and do proper evals on the tasks you care about.
- 0xbadcafebee 3mo agoMiniMax M3 and DeepSeek v4-Pro are highly capable long context open weight multi-modal models. But long-context is a trap, because performance still falls dramatically after 150k-200k context.
- InsideOutSanta 3mo ago> But long-context is a trap, because performance still falls dramatically after 150k-200k context. I'm not sure exactly what causes the difference, but this heavily depends on the model. In my experience with Opus 4.8, I can go well over 500k and still get extremely good results. A drastically different example was GLM-5.1, which worked great until about 100k and then turned insane almost immediately. They did fix that with 5.2, though.
- gunalx 3mo ago5.1 going insane was probably also a inference quirk. Because it sometimes remained coherent the entire 200k context length.
- subscribed 2mo agoI'm an amateur so it influences my setup a lot, but Opus 4.8 above 250k context in my experience with planning and implementing its own plans gets much dumber than fresh Sonnet 5, to the point of forgetting / ignoring things in the last prompt, forgetting half of the convention for the (very small and simple) codebase, etc.
- 3mo ago
- MaxPock 3mo ago[flagged]
- CurbStomper 3mo ago[dead]
- amarble 3mo agoThey also indicate they have a 276B A12B version, but it doesn't seem the weights are available. This might actually be able to fit in 128GB when quantized to 2 bits or so which makes it interesting.
- Flux159 3mo agoThey mention in the announcement link https://thinkingmachines.ai/news/introducing-inkling/ https://thinkingmachines.ai/news/introducing-inkling/ that they are still testing Inkling-Small and it will also still be multimodal. This makes it super interesting as a Deepseek V4 Flash replacement (and would be interesting with DwarfStar / ds4 if it gets supported).
- pants2 3mo agoThe Artifical Analysis has a link on their homepage but it 404's :/ https://artificialanalysis.ai/models/inkling https://artificialanalysis.ai/models/inkling
- raverbashing 3mo agoCool, now we just need the GPU that supports it
- janalsncm 3mo agoFor the most part it’s better than Nemotron, worse than GLM. This makes it the best American open weights model from what I can tell?
- deleted 3mo ago[deleted]
- nickludlam 3mo agoIt's nearly double the size of Nemotron 3 Ultra, so I'd expect it to be considerably better, although the active parameter count seems to be a touch lower at 41B vs 55B
- vcryan 3mo agoI'm surprised that Nemotron gets mentioned at all. In my experiments with it for coding tasks it performed extremely poorly, essentially unusable.
- nostrebored 3mo agoit is pretty good at instruction following and has extremely fast decode.
- jameshush 2mo agoI focus on realtime voice AI uses cases and nemotron's time to first token is INSANELY fast. It's become a legit option for voice use cases
- solomatov 3mo agoIt looks like HuggingFace shows Apache-2.0 but they have AUP. How does it work together?
- deleted 3mo ago[deleted]
- firasd 3mo agoLooks like it can be tried at https://tinker.thinkingmachines.ai/playground https://tinker.thinkingmachines.ai/playground
- bbstats 3mo agotoo bad we'll never know how good it is, since they used a radar plot to show its benchmark scores!
- InsideOutSanta 3mo agoHow does the radar plot prevent you from looking at just one of its axes?
- inkvi 3mo agoDo they have an api to try the model in real envs?
- jakswa 3mo agoThey've got an openai + anthropic compatible endpoints. I got far enough to run some tests on the openai endpoint, albeit with some finagling (their /models list is empty, my tool auto-configures using that, was an initial stumble).
- inkvi 3mo agoThanks! I found the OpenAI-compatible endpoint and got it working. I ran Inkling on a couple of my own evals. It looks promising, but on my cases it still fell short of GPT-5.4 and GPT-5.6 Luna.
- dr_dshiv 3mo agoWhat are the different business models for open-weight AI companies?
- firasd 3mo agoJust serving the model over API seems like a natural fit and is what many of them are doing. So simply being the cloud provider for your own open weight model can be a source of revenue
- dyauspitr 3mo agoBut so can everyone else. What’s the moat for spending all those billions. I understand the Chinese angle, they need to undermine American models as a matter of statecraft, but what is the business model here? It just seems like VC charity.
- 3848488459 3mo ago[flagged]
- fragmede 3mo agoMira Murati's success isn't because she's a woman.
- 3848488459 3mo agoshes not gonna date you
- kingleopold 3mo agouse open models to gain marketing/users/attention and then go closed? maybe
- wyre 3mo agoThere are no moats. LLM's are a commodity. The point in spending all of the billions is to have strong domestic open-weight models. One of the worst case scenarios regarding LLM's is monopoly control, so these billionaires know they need to invest in competition.
- MaxPock 3mo agoRaised 2 billion dollars at a 12 billion valuation and debuts at 41 on the Artificial Analysis Intelligence Index, while KIMI and DeepSeek will release Fable-class models this week. What a joke.
- KronisLV 3mo ago> ...while KIMI and DeepSeek will release Fable-class models this week. What new model is DeepSeek releasing? Their current V4 Pro at Max reasoning is consistently worse than GLM 5.2 at Max reasoning, though the latter is close to Opus 4.8 at Extra/Max reasoning, albeit a little bit worse in my experience (though if they gave comparable amounts of tokens to Anthropic 5x Max subscription I could see myself moving over, currently they give you less though even with their ZCode discount). In practical agentic development, none of those seem to be that close to Fable to me. Spent 181 million tokens with GLM 5.2 with ZCode in the past month, 142 million with DeepSeek V4 Pro with ZCode and OpenCode and about 3.45 billion across all Anthropic models with Claude Code, though understandably with my workload between 95-99% of them are cached (very docs/plan/tooling/read heavy work to limit slop, albeit with sub-agents and workflows).
- segmondy 3mo agoDeepSeekV4 was a preview model, read the papers. It's not the final model. They released it to demonstrate architectural capabilities. They are still training and the model release is planned within the next month.
- minraws 3mo agoFor a first model, and given it's open, I am gaining some faith in American Open research labs again... I couldn't test it since it's not on openrouter or something, but even if it's only as good as GLM5.1 that's more than good enough first attempt, I think. Perhaps a lot more labs will catch up to ballpark frontier esque level soon, I am all for more competition in any field.
- cpt100 3mo agoNVIDIA is building Nemotron
- minraws 3mo agoI don't want to say this but Nemotron is not worth running on any sillicon, given Nvidia has been doing it for 3+ years, if Nvidia instead gave away GLM or KIMI API for free no one would use Nemotron the reason it's so wildly used is because Nvidia offers a Free API...
- ljlolel 3mo agoit's on TrustedRouter https://trustedrouter.com/models/thinkingmachines/inkling-1m https://trustedrouter.com/models/thinkingmachines/inkling-1m
- deleted 2mo ago[deleted]
- mhluongo 3mo agoInterested in the implied strategy - that training a bespoke model for what you need will make economic sense over using a mass-trained model. I wonder if that's true?
- thatoneengineer 3mo agoSame. Gutsy bet to make in the face of Fable / Mythos, but the multimodal quality is at least a promising technical/ product story to tell. Everyone knows throwing Opus at everything is wasteful and domain expertise should live in the weights eventually; the question is whether foundation model scaling will slow down enough soon enough for that to matter. Or maybe this is just a warm-up / stopgap and Thinking Machines is betting on finding the next architectural breakthrough that lets it compete with the big foundation models?
- androiddrew 3mo agoGive me a good 180B param model that fits snuggly on an single DGX spark and I will sing your praises.
- ggcr 3mo agoMy personal bet is that this model should really shine in Autoresearch NanoGPT-style speedruns because its first-class integration with Tinker
- GodelNumbering 3mo agoInterestingly, when opening this page, the first thought I had was not that the benchmarks should be high, but 'I really hope they did not benchmaxx'. I think a model with modest benchmark scores can have much better real world utility as opposed to the current frontiers that are RL'd into being robotic and rigid.
- segmondy 3mo agoVery nice, multi modal, largest open weight model that supports audio. Would be interesting to see how good the audio capability is. If you want to run locally, checkout https://github.com/danielhanchen/llama.cpp/tree/add-inkling https://github.com/danielhanchen/llama.cpp/tree/add-inkling https://unsloth.ai/docs/models/inkling https://unsloth.ai/docs/models/inkling https://huggingface.co/unsloth/inkling-GGUF https://huggingface.co/unsloth/inkling-GGUF https://huggingface.co/unsloth/inkling-NVFP4 https://huggingface.co/unsloth/inkling-NVFP4 This supposedly is better than KimiK2.7, as much hype as GLM5.2 gets, I find myself using KimiK2.7 half of the time, so if the benchmark is true, then this can definitely go in the mix. My hope is that it might have strengths in some areas to beat all other open weight models.
- danielhanchen 3mo agoOh thanks for sharing! The llama.cpp PRs should generally be fine for now - I'm fixing a few small edge cases as well!
- dimgl 3mo agoWhat harness do you use for Kimi?
- rpdillon 2mo agoNot OP, but I tend to use Kimi 2.6/7 more than other models, works great with omp.
- paxys 3mo agoNot to mention - it is American. This is the first competitive non-Chinese open weights model since what, Llama 3?
- whimsicalism 3mo agonemotron
- segmondy 3mo agoIt's not the better model since Llama3. Trinity Large is American and quite decent, unfortunately tons of crazy good models have been out and it's harder to run locally at 400B. I think Arcee did a terrible job of promoting their model. https://www.arcee.ai/blog/trinity-large https://www.arcee.ai/blog/trinity-large
- trilogic 3mo agoYou certainly cooking smth, Good Luck Mira.
- RohoSwagger 3mo agowhy is this website ai slop
- eisbaw 3mo agodo you want them to focus on the website or their model? do you buy a device because of its unboxing experience?
- hs86 3mo agoIs it really that bad? I always get the impression that their blog posts look especially beautiful with their font choices and overall design. They are typographically pleasing, and if I could, I would use this as the distraction-free reading mode for every web page. It feels like I’m reading a newspaper, but oddly, without them resorting to any skeuomorphic tricks.
- SarahNickler 3mo ago[flagged]
- jijji 3mo agocompetition in this space is great, especially with open models/weights. I think the answer is not closed source models. Similar to the Unix versus Linux situation in the 1990's, open source wins out. Yesterdays story about how OpenAI has now began encrypting traffic between model and agent [0], this story brings a breath of fresh air. There is nothing "Open" about hiding the communication between model and agent, especially with software that is running within a trusted environment/network. It needs to be more transparent, not less. [0] https://www.theregister.com/ai-and-ml/2026/07/15/openai-hides-codex-agent-instructions-behind-encryption-leaving-developers-in-the-dark/5271484 https://www.theregister.com/ai-and-ml/2026/07/15/openai-hide...
- AlexCoventry 3mo agoOpen Source won out because the cost of compute fell through the floor. I'm not sure whether we're going to see a similar dynamic play out this time, although I would greatly prefer it to.
- jijji 2mo agocost of compute, also the less likelyhood of underhanded tactics or vendor lock-in
- luciana1u 3mo ago[flagged]
- wxw 3mo ago> Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning. Open base models that can be fine tuned on Tinker is a great business model IMO. You (i.e. an enterprise) can own your own model & have it perform frontier-or-better at your task at potentially much lower cost and Thinking Machines gets to be your essential infra/service provider in this world. Also, > Inkling-Small matches or exceeds its larger sibling on many benchmarks — the result of improvements we made to the pre-training data and recipe for the smaller model. Very cool! Excited to see the next generations of Thinky models.
- JumpCrisscross 3mo ago> that can be fine tuned on Tinker Good source to understand why this is valuable?
- ford 3mo agoFrontier models need to do everything for everyone. It's expected (though not often done) that smaller models fine-tuned on specific tasks can approach frontier performance on a specific area. [0] Post-training/fine-tuning is not trivial and having it as a service might make it more accessible. [0] https://surgehq.ai/blog/training-on-complexconstraints https://surgehq.ai/blog/training-on-complexconstraints
- ralusek 3mo agoIf you want an LLM to have knowledge about something, the knowledge has to either exist in its weights or be provided to it in its context. Because context is expensive and limited, and models tend to get dumber the more their context is filled, there is usually more that you'd need to put into context than can reasonably fit in it in order for the model to answer questions about your data. So your options are basically 1.) stuff it into context 2.) figure out a way to determine what to put into context based off of what is being asked of the model (RAG) 3.) change the weights of the model to have knowledge of your data baked into it (fine tuning)
- potwinkle 3mo agoVery impressive model, exciting to see an American open-source lab with such competitive results.
- logicprog 3mo agoThis seems like a really really great debut model for a new lab. I'm happy
- nickandbro 3mo agoLol slither.io is the new benchmark now? I guess my game slitherworld.com is now something that can be vibecoded too
- luciana1u 3mo ago[flagged]
- hahahaa 3mo agoHow much mortgage equity would I need to do that 27min fine tune demo on local :) Self fine tuning like that though seems like a whole new set of possibilities unlocked.
- Topfi 3mo agoVery preliminary testing so far, but there is something here, far beyond what the benchmarks suggest. Only ever saw such outperformance of public evals vs my private ones with Anthropic models and while it is far to early to make any judgement at this stage, this model will take up a lot of mine time in the coming weeks by the look of things. Only ever viewed Moonshot AIs models as something I'd be able to live with open-weight-wise (Z.AIs output simply does not perform as well in my task set), but this has the potential to be the second. If Mistral came out with something like this, I suspect every Europhile (me included) would never stop talking about it.
- Topfi 3mo agoQuick and still very early update, the model has (with web search disabled which was verified via the reasoning traces) accurately answered a number of questions focused on very niche details (engine specific maintenance in certain newtimers, very niche bag construction and material details) that I have only ever seen Gemini 3 and 3.1 Pro get correct. Neither Fable 5, nor GPT-5.6 Sol or any other model by any other lab has ever provided accurate information without web access for these specific questions for which an objectively correct answer absolutely exists and is general knowledge if one is versed in the specifics. Being ahead of Fable 5 in any task, that is not included in public benchmarks and thus could be overfitted for, is impressive to say the least. Last time a model exceeded the expectations I had based on the release notes to such an extent was Haiku 4.5, which I still wish we got a solid replacement for.
- nullbio 2mo ago[dead]
- aabhay 3mo agoWhat strikes me the most is just how many different tasks are involved in modern model design. It used to be the case that you come up with a new loss function, slight architecture changes, etc., run your train and eval loop, and publish the artifacts. Now, there’s so much work to do just to keep up. It’s the ultimate red queen race. All of the 500 steps involved, each of which is its own little optimization loop, is sort of awe inspiring. But obviously this inverts the previous rules that small teams run faster than big teams. AI requires a big team. It’s only once the team pushes past the 1000s that organizational inertia seems to become an issue. Because until then, there’s way too many pieces for even a dozen super stars.
- hham 2mo agohttps://news.ycombinator.com/item?id=48892559 https://news.ycombinator.com/item?id=48892559
- hibijibies 2mo agoIt's probably not that many people necessary for creating a decent model. Soofi was small team: https://news.ycombinator.com/item?id=48870978 https://news.ycombinator.com/item?id=48870978
- insane_dreamer 3mo agoI think we’re going to start seeing more OSS models that perform especially well on certain tasks instead of trying to be generalists like the frontier models. That’s a winning formula because if you’re building an app on a model it often has a specific set of use cases
- veber-alex 3mo agoYour first mission should be providing a working dark mode site. Holy flashbang.
- 2001zhaozhao 3mo agoI really respect the epistemtics work here. It might become an accurate, inexpensive open-weight workhorse for high-level prioritization and decision-making work. (Finance bros will also love this)
- christinetyip 3mo agoExcited to try out its capability, especially audio and video. It's nice that it has a long context window, but in practice, I find I always have to clear context btw 150k-400k context even if the context window is 1M on paper.
- figomore 3mo agoSplatoon LLM
- slim 3mo agoThe MoE design largely follows DeepSeek-V3 why is the model never compared to deepseek in their blog post ?
- dominotw 3mo agoeveryone and their grandma shipping top tier models now. anthropic and openai trying to capture the app layer with their shitty 'super app'
- yellowlimetea 3mo agoWe always have been, Big Tech has been extremely slow to catch up to the indies. Nobody is making money lmao. I would not bet on OpenAI creating any good products, they never have. They are like Meta in all this - never innovated anything themselves, can only acquire others to stay relevant. They'll never do an incredible consumer experience on the level of a PlayStation or Blizzard or even Google.
- paulsutter 3mo agoThe most important observation is that open source matches their business model, which is to provide fine tuning services for enterprises, etc. That what makes this a (potentially) safer model to build on top of
- k__ 3mo agoI tried Hy3 today and liked it. It's a (small) step up from DSV4P. Something on that level but multi-modal would be quite nice!
- simonw 3mo agoHere's a pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F8117ac4376371dd3fc2b5dbce27e0855 https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- calny 3mo agowell it's good they didn't train on the test!
- dozerly 3mo agoI’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now.
- tyre 3mo agoIf they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.
- tstrimple 3mo agoWait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?!
- argee 2mo agoThat’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
- 0-_-0 2mo agoIt's not a benchmark, it's a meme
- OrangeMusic 2mo agoTo be fair, the "they're trained on this benchmark" response is also a meme.
- greenlimetea 3mo ago"Our model has TRILLIONS of parameters!" "But it's worse than Mistral 7b" cape
- thatxliner 3mo agoaudio??? can it listen to music
- yellowlimetea 3mo ago[flagged]
- deleted 3mo ago[deleted]
- metmac 3mo agoYou may be new here. This is Simon’s de facto benchmark for models. I happen to find it a really good one. Small aside: It’s crazy to me that while it’s improved over time it does seem like most of the models haven’t been trained specifically to defeat this one.
- TurdF3rguson 3mo agoThe pelican thing is definitely getting deeper into model weights over time just by getting fed threads like this.
- tomhow 3mo agoWe detached this comment from https://news.ycombinator.com/item?id=48929037 https://news.ycombinator.com/item?id=48929037 and marked it off topic.
- SilverSlash 3mo agoDo you think the barbarians are at the gates of OpenAI and Anthropic? If cheaper, open weights models can seriously take revenue away from those two labs for (frontier - 1) model use cases (which are the models most enterprises will choose) then OpenAI and Anthropic are left only with users using their latest and greatest model AND who will keep upgrading to the newer ones?
- fwipsy 3mo agoBull case: iteration becomes so quick that frontier-1 won't cut it. OpenAI and Anthropic are both betting on the singularity, I suppose.
- ford 3mo agoWithout the singularity, I think Frontier labs will offer intelligent model blends. They'll have their own versions of "cheap" models and be expert at using the appropriate amount of compute for a task.
- sigbottle 3mo agoThe "singularity" as stated requires AI to either make a technological advancement strong enough to be deadly to humans (besides just intelligence), or spreading deep institutional support for itself among society. The idea that "we will get superintelligence first, then... ???" is kind of a weird notion. I mean, it's pretty arguable that we do have at least some form of superintelligence. The AI itself needs to actually do something with it though. Either that, or more likely, someone needs to do something bad with the superintelligence. That could be both re-assuring or not. Because under that view, given how AI is being integrated so quickly into society, it's not going to take this fantasy view of superintelligence to reach the singularity. If you have broad institutional support and crowd out the thing we call humanity over time (the two ways to 'solve' a problem: solve it, or declare it meaningless), that is another way to reach the singularity.
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- helloplanets 3mo agoThe actual part on fine-tuning seems very short in the article. Did I miss a page where they have examples of fine-tuning it for different niche use cases? Optimizing models to be fine-tuned is an amazing direction, but just makes me wonder how much better this actually is at being fine-tuned compared to other models. As none of the modern models are great at being fine-tuned afaik. Basically looking for some sort of benchmark showing that it's resistant to overfitting / catastrophic forgetting, etc. Would be very interesting to see concrete demonstrations of different fine-tunes of the model. I'd imagine they've done hundreds of those internally.
- malshe 3mo agoThey have good docs on finetuning in general here: https://thinkingmachines.ai/tinker/ https://thinkingmachines.ai/tinker/ I used it last week for an application using a small model just as an experiment. It all went very well. The model did not turn out to be good though because my training data was of bad quality. I plan to work on it more this weekend.
- luciana1u 3mo ago[flagged]
- nikcub 3mo agoThis is a winner IMO. Lots of cost pressure on token spend atm within enterprises and tasks that don't require Opus / Codex class models. These companies have hopefully captured all of their traces and now have enough to fine-tune an open model and host themselves. Inkling feels like the right base - not obsessed with benchmaxxing on coding but rather being adaptable to the task required For tasks like GTM, support, content writing etc. seeing 80%+ savings
- hdemirev 2mo agoOpen weight models as a category might be a winner due to cost pressure, but I don't think Inkling is a top performer in that respect. Pricing on a token basis is 6-9x higher depending on the provider. Can chime in on the support use case specifically: GPT OSS performs really well here and has been somewhat of a benchmark with our customers, limited testing [0] against Inkling reveals basically identical performance, but with a significant cost increase at scale. I'd say that for real-world tasks that aren't coding most companies don't see value by being on the latest and greatest model. [0] https://valiopt.com/blog/inkling-model-customer-support-review https://valiopt.com/blog/inkling-model-customer-support-revi...
- semiinfinitely 3mo ago45 trillion tokens is a lot
- hdjdjdjdjdjdjd 3mo ago[dead]
- kamranjon 3mo ago“Alongside Inkling we are sharing a preview of Inkling-Small, a 276B-parameter Mixture-of-Experts model (12B active, vs. 41B for Inkling) with a different performance/latency trade-off.” Buried at the end there is the details I was most interested in - a possible competitor for DeepSeek V4 Flash? Excitedly awaiting the release of the weights for this one.
- 13639366668 3mo ago[flagged]
- maxignol 3mo agoGiving the accuracy-token graph and not the accuracy-cost graph, thus we cannot easily compare costs with other models, is not a way to gain my trust
- throwaw12 3mo agoIs this the same model they are using for interaction models? https://thinkingmachines.ai/blog/interaction-models/ https://thinkingmachines.ai/blog/interaction-models/
- rasmus1610 2mo agoyes. At least they want to use it for that, I am not sure they already do. (look under "multimodality" in the blog post: https://thinkingmachines.ai/news/introducing-inkling/ https://thinkingmachines.ai/news/introducing-inkling/)
- small_model 3mo agoHow does this compare with Grok 4.5, Fable, GPT 5.6 etc I can use them for a few bucks a month, whats the benefit of using this model? is it more intelligent, faster, cheaper (and no I don't want to spin-up my own mini datacenter to run it 'in house') I want to install a harness/app/visit web page auth and start as 99% of people using AI want to do.
- mtxeat 3mo ago[flagged]
- hham 2mo agoSmart that she says "not the strongest overall model, open or closed". This is a rare for an AI lab to say out loud. They basically decided to compete on customizability, and not on topping the temporary leaderboard. Also corroborates what we recently wrote: any Lab's capability lead cant hold for long anyway, it's a "red queen race" that never settles: https://news.ycombinator.com/item?id=48892559 https://news.ycombinator.com/item?id=48892559
- arisAlexis 2mo agoartificial analysis puts it quite low below previous gen open weights models from China. Why are people so ecstatic?
- KomorKomor 2mo ago[flagged]
- dipakk 2mo agois there something in the space of "taskifying" enterprises data for them? inkling on its own looks high-quality, but expecting companies to spend $ figuring it out before spending more $ on the actual fine-tuning job seems ... hard, especially if making the model especially customizable is the goal? or do the unit economics just work out with a small number of fine-tuners training a ~1T model on Tinker?
- neo_core 2mo ago[flagged]
- finecode 2mo ago[dead]
- zftnb666 2mo agoOpen weights are great, but open payment matters too. DeepSeek V4 is open-weight but closed-payment. api-hub.cc makes it accessible with just a credit card.