14 ms·
Qwen 3.8
https://www.qwencloud.com/pricing/token-plan https://www.qwencloud.com/pricing/token-plan
- revolvingthrow 3mo agoThe few tests I ran were by no means comprehensive, but while kimi felt like the real deal qwen seems a bit of a benchmark princess.
- dannyw 3mo agoQwen3.6 is still the best agentic open weight LLM around 30b params (Gemma isn’t very good at agentic execution). I also find the model is a lot more predictable and less “glitchy” when made to think in Chinese. You can do this in the system prompt.
- akazantsev 3mo ago> Gemma isn’t very good at agentic execution I had no issues with it for C++ development with https://pi.dev https://pi.dev. I'm yet to try it with Zed Editor. I don't rely on agents too much. However, I used it on Chromium's codebase to research some functionalities, let's say for searching. Requests like: check my last commit and do the same for SetterA and SetterB; it also ran without any errors.
- anana_ 3mo agoApparently agentic performance in Gemma was improved recently: https://x.com/googlegemma/status/2077449152062247219 https://x.com/googlegemma/status/2077449152062247219 Too little too late imo
- wajahatraja 3mo ago[dead]
- adnane4 3mo ago[dead]
- molmos 3mo ago[flagged]
- lebovic 3mo agoI'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page. Looks like they're previewing the model only on their subscription plan.
- LaurensBER 3mo ago> With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5. That's a massive model! The shift from "value" models to "intelligent, huge and slow" models coming from China is an interesting change in strategy. My main issue with GLM 5.2 and Kimi 3 is that they're extremely token hungry and thus feel slow(er) to use.
- charcircuit 3mo agoThe shift isn't new. Kimi K2, a 1T model came out July last year. I am happy that more labs are following the trend as its important for competitive open models to exist.
- chronogram 3mo agoAnd DeepSeek has been making huge progress on efficiency, and publishing about, so they came with a 1.6T model that is both fast and cheap to run.
- selcuka 3mo agoAlso DeepSeek R1 was announced 1.5 years ago with ~0.7T parameters, which was a huge model back then.
- hodgehog11 3mo agoValue models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companies are still hanging in there.
- adrian_b 3mo agoI assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to better compete with Moonshot AI. In any case, from this competition in LLMs, we win.
- gardnr 3mo agoIt's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
- barrenko 3mo agoHumanity is a bit of a stretch, and to be seen over time, not that I'm saying it won't happen; let's get some hubris here.
- hodgehog11 3mo agoI think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here to stay in one form or another.
- DanielHB 3mo agoThe fear is not about the models open weights it is the erosion of training capability in other countries. Why train models when they do it for free? Until they don't of course, or they start doing what the US is doing right now by locking out some models to government only or internal market only. What people should be afraid is the rug pull.
- comandillos 3mo agoJust imagine Anthropic making Opus open-weights now for the sake of trolling everyone. Wouldn't surprise me at this point xD
- hodgehog11 3mo agoThat would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, because it is all too dangerous in the hands of anyone else.
- anon373839 3mo agoI’ve seen no evidence that he believes in anything. He comes off as just another slimy would-be monopolist to me.
- cyanydeez 3mo agoI see no evidence that any ceo retains anything but the desire to capitalize on their marketplace of ideas for their own benefit. Like wolves inn sheep clothing, they'll put on any skin suit to convince people to keep giving them money and power. And it has nothing to do with the individual, from what I can tell, 70% of the population placed in their position would become the same type of uberpath.
- cindyllm 3mo ago[dead]
- embedding-shape 3mo agoThat OpenAI releases Sol as downloadable weights feels way more likely than Anthropic releasing even the tiniest of models for download.
- deleted 3mo ago
- brunooliv 3mo agoQwen is so much better than GLM or Kimi that this makes me genuinely excited!!
- Pesto 3mo agoThe bigger models are usually worse though, hopefully they nail it this time.
- nwhnwh 3mo agoOpen what?
- khurs 3mo agoGo China, screw America* *within the scope of open models only
- dannyw 3mo agoI like my Apache 2.0 licensed Gemma, and NVIDIA’s Nemotrons are decent bases for finetuning or continued pretraining, esp thanks to good documentation and tooling. Oh, and Mira’s thinking machines lab dropped Inkling, a ~1T open weight model too. This isn’t US vs China. This is open vs closed.
- khurs 3mo agoIt's Sunday morning so I'm allowed to be facetious!
- jimbob45 3mo agoThis isn’t open though. Promises to be open later aren’t worth anything, given what we’ve seen and heard from AI execs making promises in this industry.
- ycui7 3mo agoit takes extra effort to open source a model even if you had it running internally. even traditional software takes extra effort to get released as open source.
- jimbob45 3mo agoI’m sympathetic but AI claims are to be disbelieved until proven anymore. That’s a well poisoned thoroughly by not just American companies.
- hodgehog11 3mo agoThe "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.
- neevans 3mo agotbh even if its better model due to lot of restrictions its not that useful than opus.
- hodgehog11 3mo agoTell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.
- NitpickLawyer 3mo ago100% this. There's currently this [1] submission that hasn't gained much attention, but is really important. In this [2] incident report from HuggingFace, they talk about detecting an attack and not being able to analyse the logs / IoC with API models because of guardrails. If not even highly regarded reputable companies can't sort out access to SotA models for blue team use, the raw capabilities don't matter. They're useless paperweights (hah!), and nothing else. Having to resort to open models is insane! [1] - https://news.ycombinator.com/item?id=48965243 https://news.ycombinator.com/item?id=48965243 [2] - https://huggingface.co/blog/security-incident-july-2026 https://huggingface.co/blog/security-incident-july-2026
- throwa356262 3mo agoKey part from [2]: "When we started the log analysis, we first used frontier models behind commercial APIs. This did not work [...] We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. [...] The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout [...]"
- nsbk 3mo agoBring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data
- psychoslave 3mo agoWhat hardware do you have?
- nsbk 3mo agoI run a 2x 3090 rig, but a single 3090 already provides a great experience at a reasonable quant and context size. On a single card I used to run Qwen_Qwen3.6-27B-Q4_K_M or similarly quantized 35B MoE at 65536 context size
- xiconfjs 3mo agoyou should look into https://github.com/noonghunna/club-3090 https://github.com/noonghunna/club-3090
- mft_ 3mo agoI think everyone is hoping this! It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.
- nsbk 3mo agoIndeed! That would be the sweet spot for my 2x3090 rig
- embedding-shape 3mo ago> the 122B version of 3.5 Yeah, this is what I'm holding out for, the NVFP4 variant of 3.5 122B is blazing fast with reasonable quality and even with max context fits perfectly within 96GB.
- jane_hilly 3mo ago[dead]
- SwellJoe 3mo agoQwen is the most censored of the Chinese models in my testing, which makes me wonder in what other ways it is compromised. Open weights doesn't really reveal what's in there. And, in my tests, existing Qwen models are not at the pareto frontier of any metric; DeepSeek V4 Pro is better, faster, and much cheaper than Qwen 3.7 Max. (DeepSeek is also among the least censored of the Chinese models.) I guess we'll see if the "second only to Fable" hype pans out. In my limited experience with Kimi K3 (I signed up for a month of the $19 plan) it's slower and chews a lot more, so ends up being pretty expensive; one little feature burned through almost the entirety of my five hour limit. The $20 GPT plan is a lot more useful and includes 5.6 Sol, which is fast and token-efficient enough to be quite usable even with the small plan.
- dannyw 3mo agoWhat censorship? ;) https://github.com/p-e-w/heretic https://github.com/p-e-w/heretic
- SwellJoe 3mo agoYou still don't know what's going on in there.
- dannyw 3mo agoI can explore and find out _something_. LLM interpretability has come a long way, even if we don't have all the answers, the weights and activation do tell a lot; and when analyzed collectively, each weight isn't a random number anymore. I can use techniques from the simple logit-lens at different layers, to J-Space analysis, to more advanced techniques for identifying deliberate misalignment. I can create and inject steering vectors, whether it's to align a model's CoT (which can be deceptively trained to misinform) closer towards what its underlying activations suggest, or just to probe or steer it. I can also statistically analyse and understand _if_ steering vectors have been applied; and if so, from the vectors themselves it's very possible to translate those vectors back to the intent. Think of it as analysing the complete, heavily obfuscated source code of something that is self-contained. It's not 100% the same, but weights are incredibly illuminating.
- rcarmo 3mo agoI do hope they provide optimized A3B quants--that's been the sweet spot for usable local inference for me, at least.
- docheinestages 3mo agoQwen has set an excellent track record for architecting and releasing open-weight models that consumer-grade devices can run. What is needed the most right now is something similar to Bonsai 27B, with a modest memory footprint, but faster and more capable. On-device models can make up for intelligence by being faster, thinking longer, or doing more quick iteration rounds.
- embedding-shape 3mo ago> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable Yeah, that'd be neat, but that's not what this announcement is about at all: > With a massive 2.4T parameters
- docheinestages 3mo agoTrue. It was more of an open letter, with hopes that the Qwen team sees the comments in this thread.
- cyanydeez 3mo agodont we all deem the ability to improve large models as the defacto capability to produce small ones?
- embedding-shape 3mo agoI don't think so, they have different constraints and require different optimizations, being able to produce one of them doesn't mean you'll automagically be good at the other.
- cyanydeez 3mo agobut that's denying the singularity boostrap theory. Which I don't agree with, but if you can't harness a large model to make a small model, then we're going to have problems brining about the singularity. If you do think there's some magical singularity, how do you comport?
- eurekin 3mo agoWith 3.6 27b, I just stopped changing local models and started tinkering with things on top (like mem0). Feels genuinely useful and more than a toy
- androiddrew 3mo agoI have only been using 3.6 27B for coding. Is mem0 for agents like Openclaw or Hermes? How are you using it?
- eurekin 3mo agoIt's a mcp, so connects quite easily to agents. With mcpo, I also connected it to open-webui (which has better support for OpenAI style tools/functions). Used it in claude code with that mcp plugin set-up too. Only ever used it for managing homelab information, but it met initial expectations. 27b is a great model, if grounded. The query about physical hosts and routing... I haven't found a single hallucination (altough Codex 5.6 as a reviewer mentioned something was wrong with some parts, and those were exactly the never properly documented ones. Codex/gpt had extra knowledge, because it was the conversation I used to set it up).
- baist0 3mo agocan i get "code instruct" version of this? i want 7B and 14B to launch on my hardware.
- vitorgrs 3mo agoDeepseek 4 "final" version is imminent as well. Will probably be at Opus 4.8 level, and I find it pretty big deal because of Deepseek price...
- drob518 3mo agoYea the performance/price ratio for Deepseek is off the charts. I’ve been using V4 Flash a lot lately and it’s quite good.
- mark_l_watson 3mo agoI like that v4 flash is so fast! I run it on both FireWorks.ai in the US and bought some tokens directly from DeepSeek as an experiment. I only work on Open Source projects, so I don’t have to worry about my work being used to train models - I welcome AI’s being trained on my open content books and code (but not my conventionally published books: I am a party to the copyright suit against Anthropic).
- drob518 3mo agoYea, Flash is quite fast, though looking at model data on Open Router some of the other models are quite fast (Muse Spark, Grok, etc). I’m sure all these models have been trained on my conventionally published books as well, but I don’t care.
- XCSme 3mo agoDeepSeek V4 pricing is insane, 10x-30x cheaper to use than most other models, and it usually is good enough for most tasks.
- onlyrealcuzzo 3mo agoWho do you buy DeepSeek from? I bought it through OpenRouter and used it with Pi agent. The model was good, but there appeared to be a pricing glitch or something, because it burned through $50 in under an hour on pretty trivial stuff. Pi agent claimed it only used like $1. OpenRouter claimed differently and said I used all $50.
- gxs 3mo agoThe open weights vs frontier models is reminding me more and more of the Linux vs Windows I grew up with (slashdot randomly popped into my head saying that) I have a feeling this is the next…frontier of that fight One can only hope it eventually does as well as Linux
- deleted 3mo ago[deleted]
- Alifatisk 3mo ago> You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out. You can also try it out on Qwen chat, Its free.
- 5701652400 3mo agoin my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far. and it is super expensive compare to Deepseek. cannot delegate anything to it, cannot use it real-time low-level tasks either. totally unusable.
- ph4rsikal 3mo ago> Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. D Anthropic should not have bugged their knowledge distillation attacks.
- chewz 3mo ago> Anthropic should not have bugged their knowledge distillation attacks. It is like one of Pizzaro's men crying that someone have stolen his precious golden dublons As Lenin have said - "Loot the looters" (Russian: Грабь награбленное)
- RazorBucksICO 3mo agoAppealing to the Belsheviks for moral authority is, well I will just say an interesting approach. I do not have that much sympathy for Anthropic, but I do not have much sympathy for publishing companies either whose rights to a revenue stream they violated either. Are Chinese AI companies the Robin Hood in this story? Would they be so magnanimous if they had the upper hand? I don’t think so.
- trollbridge 3mo agoConsidering the results from Kimi K3, it appears most the accusations of them “stealing” via distillation are unfounded accusations.
- zobzu 3mo agohow many hn posts do you believe arent propaganda these days? its billions, trillions were talking about. imo hn should display posters origin, such as country, bon, datacenter registered ips, and the discourse will change dramatically.
- ernsheong 3mo agoSo are locally-runnable models frozen at Qwen 3.6 now :/
- worldsavior 3mo agoEveryone wanted open models that would challenge Opus and Codex, here, you got it.
- ernsheong 3mo agoWe need better coding models that can run on local hardware, i.e. 128GB VRAM or less
- seanmcdirmid 3mo agoQueen has that already, although they seem to be moving away from local models unfortunately.
- zozbot234 3mo agoYou can run larger models by offloading to SSD (for weights), it's just slow so people don't do it all that much. But you can get back at least some of that performance by using either MTP (at least for dense models; not effective for sparse MoE models unless you're batching them already and have VASTLY more parallel compute than you'd know what to do with) or batching multiple requests in parallel (note, this hurts throughput for your single sessions but running more sessions in parallel still boosts your total amount of inference. This requires careful management of memory requirements for your context/KV cache, and Qwen models tend to be KV-cache heavy). Broadly speaking, this ultimately pushes local inference towards a challenging world where you use SSD offload for weights as a matter of course; then smaller requests (or requests sharing the bulk of their context, e.g. subagent swarms) can be batched together and run quickly in aggregate, but running very large contexts will actually limit you to single-session inference and require swapping out even the KV cache itself to some external scratch SSD, further hurting your performance. Then feel free to add wide use of MTP in a probably futile effort to go back to tolerable tok/s numbers.
- 3mo ago
- try-working 3mo agoin case someone wants to speculate in why chinese labs open source their models: https://try.works/why-chinese-ai-labs-went-open-and-will-remain-open https://try.works/why-chinese-ai-labs-went-open-and-will-rem...
- nullbio 3mo agoI predict that no one will use this and everyone will use Kimi K3.
- embedding-shape 3mo agoI've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time.
- sunaookami 3mo agoSame problem with every chinese model currently, they overthink way too much and take too much tokens and time.
- EgregiousCube 3mo agoA consequence of aggressive distillation?
- embedding-shape 3mo agoMore or less, yeah. I've found mild success with deepseek-v4-flash though, and also Qwen3.5-122B-A10B-NVFP4 running locally, especially in terms of "doesn't overthink every single prompt" and somewhat reasonable quality. Really wishing for a 3.8 update of the 122B variant, that'd be really competitive (for local usage) :)
- szundi 3mo ago[dead]
- jadbox 3mo agoWhat's the price difference?
- rubslopes 3mo agoWhy? Price? If the reason is performance, I've been using non-frontier models for cheap, and they run great for my needs (GLM 5.2, DeepSeek v4 Pro).
- dluan 3mo agowaic go brr
- hermes_scanner 3mo ago[flagged]
- antiloper 3mo agoDoes anyone have the privacy policy of their token plan available? Want to check if they retain/train on inputs/outputs.
- moffkalast 3mo agoLol, lmao even. Of course they train on literally everything they get their hands on, like everyone else. If you need privacy, that's what local models are for.
- adamtaylor_13 3mo agoIt's a flippant answer to a real question. Anthropic, OpenAI, and even Grok have "Don't train on my data" knobs. Whether you trust them is different, but there ARE knobs on other hosted AI companies.
- moffkalast 3mo agoThose knobs don't do anything, don't be silly. It's just optics.
- villish 3mo agoDo you have proof of that?
- moffkalast 3mo agoDo you have proof it's not? There are no laws they have to follow in regards to it, and it's practically impossible for them to go against their core self interest. More data is literally a direct component to better models and more revenue. If someone proves it's fake worst they get is a month of bad PR and then people will move on, as they always do.
- villish 3mo ago
- smnplk 3mo agoDo this giant open-weight models have less active params and could be run on consumer hardware or no ?
- adrian_b 3mo agoAny open-weights model that has ever been published can be run on consumer hardware, even on a mini-PC. The right question is which is the speed that can be achieved on a given hardware and whether it is high enough for the model to be useful. Until now, the speeds reported for running big LLMs with the weights stored on SSDs have ranged from as low as a token every 10 seconds or so, to as high as a few tokens per second. With open weights models that you host yourself, you are not constrained to use any single model, because that is the one for which you pay a subscription. You can use many models, each for whatever it is more suitable. You can use frequently a small model with a high inference speed, but for some tasks you may actually save time with a better model, even if it is much slower. In my opinion, even at 1 token per second a big model may be useful for some tasks.
- apitman 3mo agoI can only imagine what that does to the poor SSD
- Alpha3031 3mo agoPaging in experts is mostly reading so on the first order effects it would be fine. These might be second order effects from e.g. write caches needing to be flushed more often (and maybe swapping other applications, if you use a swap file or partition) but it probably wouldn't be too much of an issue.
- adrian_b 3mo agoPretty much nothing, if you only read the weights from SSDs, i.e. if you keep any writable caches in your actual DRAM.
- netdur 3mo agoThe only problem I had with Qwen, fine tuning on Colab, it takes 31 t/s while Gemma 4 is around 9 t/s, otherwise, one of best local LLM
- vitorgrs 3mo agoSVG's pelican https://gist.github.com/vitordelucca/521c2d63c9b852c622e76489df72b5ce https://gist.github.com/vitordelucca/521c2d63c9b852c622e7648... Made on the website, so not sure if on the API there's more thinking options...
- esrauch 3mo agoI feel like the pelican test can't be relevant anymore; the whole point was to to something that wouldn't be in the training set at all and now it is?
- LatencyKills 3mo agoAgree. It was interesting/fun for a bit though.
- joegibbs 3mo agoWhat about an armadillo playing a piano? There are so many potential combinations It would say something if the pelican looked great but the armadillo looked terrible
- rhdunn 3mo agoA parachuting flamingo? An aardvark driving a bus? It should be easy to randomize the animal and the mode of transport (or vary it with animal playing a sport) to create images not in the training data.
- onlyrealcuzzo 3mo agoHow about an animated SVG of a pelican doing the Macarena, profile view, spinning to face the camera on the last beats?
- rvz 3mo ago> I feel like the pelican test can't be relevant anymore; It never was. The point of this "pelican test" was for performative reasons, or just for attention of the joke. It is like trying to test whether if an adult elephant could actually climb up a tree and reporting that some elephants are slightly better at doing that than others while also reporting at the same time that they are all bad at tree climbing anyway. This is an example of testing for the sake of testing. The "pelican test" tests for nothing.
- Alifatisk 3mo agoI remember when they released Qwen 3.7 Plus and Max. These models behaved way different from all prior models, it became too verbose. It wrote multiple paragraphs just to answer my prompt instead of the usual concise and direct way responding to me. I didn't like that at all, and I know Gemini also had this behaviour with with the Flash series until I managed to reduce it a bit with personal instructions (in the settings on Gemini website). I haven't tried Qwen 3.8 Max yet, looking forward to it. My hope is that its way less verbose. Another thing I experience with the Qwen models is that I do not trust their benchmark scores at all. Have anyone played with Qwen 3.8 Max and can share their experience? Which model it come close to? Sonnet 5? GLm-5? DS V4 Pro? Flash? Gemini 3.5 Flash?
- cakbeslik 3mo agoUsing QWEN models since 2.5. I never used the chat properly but as an API I can say they're quite good, especially when you compare with OpenAI models. Cheaper and almost same level. I will try this now also.
- sieste 3mo agoWhat is a "credit" and how does it translate to tokens for the different models?
- xyzsparetimexyz 3mo agoIts the currency of the future.
- whynotmaybe 3mo agoThey use it in Babylon5 (supposedly) in 2260!
- ahartmetz 3mo agoSince Euro and Dollar values are reasonably close, you can call both of them credits, maybe
- tclancy 3mo agoYes, but credits are money you don’t actually own. So much more convenient, wave of the future and all that.
- sbinnee 3mo agoIf it offers more than opencode go, the entry plan looks enticing
- jane_hilly 3mo ago[dead]
- rhdunn 3mo agoDoes anyone know if they intend on releasing open source/weights variants for 3.8 or whether 3.6 was the last model they are/were doing that for?
- rolls-reus 3mo agowill be releasing weights per their tweet announcing the model https://xcancel.com/Alibaba_Qwen/status/2078759124914098291 https://xcancel.com/Alibaba_Qwen/status/2078759124914098291
- deleted 3mo ago[deleted]
- bdxn 3mo ago[dead]
- Archit3ch 3mo agoObligatory "Does it answer security questions?".
- beefsack 3mo agoFor those trying to get it to work in OpenCode with a Qwen Cloud Token Plan, this is what worked for me. Note that I've just matched Qwen 3.7 Max for the limits as I don't know exactly what they are. "provider": { "alibaba-token-plan": { "models": { "qwen3.8-max-preview": { "limit": { "context": 1048576, "output": 65536 }, "modalities": { "input": [ "text" ], "output": [ "text" ] }, "name": "Qwen3.8 Max Preview" } } } }
- 5701652400 3mo agoalso, be very careful which API endpoint and API Token you use. make sure you use right one (obseve your quota is used up. if you hit right endpoint quota used almost immediately). so that you do not accidentally burn API endpoint tokens (they are expensive, can easily hit 200 USD / 3 days which do not count towards your membership "Credits", if you say purchased it with 200 USD signup bonus in Alibaba Cloud)
- cadlernox 3mo agoNice
- hunmernop 3mo ago[dead]
- corv 3mo agoWho is behind this site? Is this another frontend to Alibaba or a reseller in Singapore?
- jxmorris12 3mo agoWhy did Qwen stop producing open models? They've gone from building the best open models ~1 year ago to producing like the 10th-best closed models. I don't understand this pivot at all. Edit: I saw online they do in fact plan to release this openly at some point – x.com/Alibaba_Qwen/status/2078759124914098291
- fragmede 3mo agoIt's not a pivot, giving away the weights was a marketing strategy that they don't need to keep up with.
- InsideOutSanta 3mo agoThey've announced that they're releasing the weights for a 2.4T model soon: https://xcancel.com/Alibaba_Qwen/status/2078759124914098291 https://xcancel.com/Alibaba_Qwen/status/2078759124914098291
- mikae1 3mo agohttps://news.ycombinator.com/item?id=48966120 https://news.ycombinator.com/item?id=48966120
- sinuhe69 3mo agoThe title is misleading. The link led to a pricing page/token plan and not about the new QWen 3.8 model.
- Schiendelman 3mo agoThe submission link is to the twitter announcement. The body just has a different link to pricing.
- MichaelNolan 3mo agoIf 3.8 max goes open weight, what are the odds they retroactively open weight the earlier releases?
- Gecko4072 3mo agoWhat would be the point? Kimi is kind of forcing their hand.
- kennywinker 3mo agoWell, if nothing else, posterity.
- notnullorvoid 3mo agoAlways nice to see more open-weights in the heavy model class. I can only hope this trend continues, causing OpenAI and Anthropic to crash and burn.
- blfr 3mo agoAs much as I dislike 'em, this sounds mean spirited. And Alibaba admits in this very tweet that Fable is next level (it is).
- Hamuko 3mo agoI want as much misfortune as possible to befall OpenAI and Sam Altman after what they did to the memory market.
- brap 3mo agoHow dare they buy things
- notnullorvoid 3mo agoI dislike their practices, but the main motivation for hoping they'll crash is that I think their immense overvaluation posses too much economic risk.
- ludydev 3mo ago[flagged]
- Umair_khan2324 3mo agonice
- adnane4 3mo ago[dead]
- adnane4 3mo ago[dead]
- sidcool 3mo agoThese models are great, but what's the potential use? No small entity can run them.
- alex43578 3mo agoChina encourages/prioritizes their release because it directly competes with American companies closed models.
- raised_hand 3mo agointeresting, when will this race end?
- monster_truck 3mo agoI really like Qwen, even the Q2KP quants of 3.6 27B have genuinely impressive local performance on a 24GB card. It has been good enough that I am happily giving them $60 right now to try this instead of waiting to try a slightly lesser version locally. Was there ever an explanation for why we never got the weights of 3.7? I would like sourced quotes and not weird/cringe accusative speculation about distillation, or your take on The Big D.
- jared0x90 3mo agodo you mind sharing your settings? i just picked up an r9700 to start playing with local qwen3.6 27b and your setup sounds promising and efficient on 24gb.
- kadoban 3mo agoDifferent quants, but I've been having success with: https://unsloth.ai/docs/models/qwen3.6 https://unsloth.ai/docs/models/qwen3.6 I do llama.cpp
- monster_truck 3mo agoWhich settings exactly do you want? As far as the model settings go I just follow what's on the card. I'm using HauhauCS's models, they seem to do a slightly better job with their "P" quants. Especially wrt patching them to eliminate the "doom loops" that will time out the GPU (esp if you have not already given it the extra power budget, set fans to max, and lowered your max clock by about 8% to save yourself a crash/reboot). Without getting into the weeds, unsloth covers a lot more ground so, when they're good they're great, but I've also had the most problems with them. Basically, don't be afraid to shop around and fuck with sliders. These days it should almost always be enough to open up LMStudio, set context to max, K/V quant to 16 or 8, and off you go. I'm using a 7900XTX, with 128GB of memory for the cases where things don't fit. The default settings should be fine otherwise. I don't know anything about the r9700, seems neat. Looks like it might play nicer with the vulkan backend than rocm, and you might have to `set GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1`. Perf should be at least as good as what I'm getting (180+ tk/s prompt, >50tk/s response), which imho is just fast enough that I can read it as it reasons and responds. You can probably get more aggressive with KV quant (any of the 4's) and bump up the batch sizes (4096/1024 vs 2048/512, etc), it's quite situational. These days there generally should not be any serious loss of accuracy from the former. Keeping the rig cool will likely matter more, so turn your music up to hide the fans and set your rig on top of the AC exhaust lol. Oh and I guess make sure ReBAR is enabled and working, adrenaline should let you know if it's not enabled.
- dartharva 3mo agoI very much appreciate the existence of these free models, but in my experience Qwen has too high of a tendency to confidently give the wrong answer as compared to other frontier models.
- mannanj 3mo agoIt's an interesting time to be alive when your local models are supposedly the pinnacle of what a free nation is capable of, yet the ethicality of the companies is disliked and their models restrict and limit you so much you root for the models from a socialist/communist state. If it wasn't for the effectiveness of propaganda, tribalism and psyops in this scenario my words wouldn't even be controversial and would just be seen as a truthful observation.
- Elzair 3mo agoHas there been any news on open weighting Qwen Image 2.0 and WAN? We are spoiled in the LLM segment, but I would love to see an open source competitor to Flux.2, etc.
- mchusma 3mo agoI counted the other day and there were at least 12 different providers with "better than Opus 4.5 performance" on Artificial Analysis, Opus 4.5 being Anthropic's December release that many say kicked off the latest acceleration. Which is totally insane competition, particularly given how low switching costs. I personally think that Opus 4.5 level performance is sufficient for most apps and usecases, as they get deployed.
- margorczynski 3mo ago> Which is totally insane competition, particularly given how low switching costs Which is why OAI and Anthropic will most probably push for more governmental control and bans. Without it their whole income model is cooked.
- culi 3mo agoI'm not sure if corporations will be willing to allow LLMs to go the way of EVs. Then again, all of Chinese models are open. And DeepSeek even publishes research papers alongside their models that go in depth into the methodology. I guess there's not much stopping USian companies from copying
- wolttam 3mo agoDeepSeek isn't the only one publishing; Kimi published about both their Attention Residuals and KDA / sparse attention.
- wolttam 3mo agoThe U.S. has ~350 million people. Anthropic and OAI can piss and moan all they want - limiting the U.S. to only their models would hurt the U.S. economy in myriad more ways than the failure of a couple of companies that scaled too quickly. If they get that outcome, the rest of the world would simply keep moving forward with access to open models and tokens at pennies on the dollar.
- sampton 3mo agoThis is reminiscent of the operating system wars and browser wars. In the end there will only be 2 models that can survive. 1: give it away for free or 2: locked in with top notch hardware.
- maxrumpf 3mo ago> "compatible" instead of "comparable." It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.
- segmondy 3mo agolet's see your grammatical correct tweet in chinese.
- rrhjm53270 3mo agoWell, I think Chinese doesn't really have any grammar most of the time.
- aloknnikhil 3mo agoIt's OK because I'm not prompting the person who tweeted for my usecases.
- pvorb 3mo agoNow you can be sure they write their announcements by hand. Doesn't really matter, does it?
- lardosaurusrex 3mo agoat the risk of upsetting a lot of people mentioning this but like are you really surprised? if youre going to use ai for everything youre gonna start losing your edge as you focus less and less on what youre doing and this isnt me just talking out my ass, like... the front page here is peppered with study after study and blogpost after blogpost about how its overuse can come to the detriment of one's own abilities and skills. coca cola had the ad with the magical truck that changed its design, shape and amount of tires it had and if nobody noticed that before releasing it then im not sure why anyone might think that the people peddling the LLMs would somehow be immune to this phenomenon
- halJordan 3mo agoThis is the sort of unmitigated pedantry thats the real problem. Everyone makes typos and errors of this nature. Feymann and Hemmingway both did. You deliberately chose this level of error multiple times when you chose to not put an apostrophe in youre and failed to capitalize proper nouns Come off that high horse
- souravsspace 3mo ago[flagged]
- cloudengineer94 3mo agoThings are heating up in China. Looking forward to see what Antrophic and OpenAI does next.
- softwaredoug 3mo agoIt feels like an inflection point of lost US leadership in technology? A year plus ago you would say while China led in green energy and manufacturing, at least the US was ahead in software - as demonstrated by the state of US AI models. We could point at a lot of factors on the US side. From political paralysis / head-in-the-sand attitudes towards emerging tech like green energy. To something of disdain for workers that will be impacted by AI (creating a backlash). To education that continues to lag. Add to this so many other self-inflicted economic wounds from the current administration. I don't know if its nearly as terminal, as say the UK after WW2. The US is still large, wealthy, and resource rich. Yet at a minimum the triumphalism about US leadership after Trump was elected by the tech elite feels silly in retrospect. Something I also think about is how much stronger The West overall would be if instead of antagonizing allies, there was a single ecosystem working closer together.
- pessimizer 3mo ago> We could point at a lot of factors on the US side. It's just lack of antitrust enforcement. China pours money into tons of different businesses in the same industry and lets them fight it out. The only US business model left now is to shut down (or collude with) competitors and raise prices while cutting costs. All they have to do is cut Congress (and regulators, and individual judges) in. We've financialized everything for the sake of scammers, rather that finance being used for the sake of getting cash to the most productive organizations. We've optimized for corruption. If we hadn't let the stupidest people in the world buy up everything, and made doing nothing with it the most profitable option, China would have never have blown past us. The US Supreme Court has explicitly legalized "tipping" politicians. That's the biggest sign of degeneracy that a government could possibly achieve. https://en.wikipedia.org/wiki/Snyder_v._United_States https://en.wikipedia.org/wiki/Snyder_v._United_States
- 2001zhaozhao 3mo agoNow there are not one, but two incredibly powerful open LLMs. I think this level of capability makes general prioritization / high level decision making doable with the right harness, and now everyone has hard-to-interrupt access to them (since these are open weights and someone in the world is going to run them). This world is going to get really weird soon, in both good and bad ways...
- theplumber 3mo agoLet’s better download them fast before Dario is making a scene again!
- kingo55 3mo agoIf he bans them, they'll just show up as torrents. Or we'll download them from Chinese hugging face.
- overgard 3mo agoI've been using Qwen 3.6 27B with LMStudio, and I was pleasantly surprised with it, although it was a little slow. I found mtplx last night, and it really wasn't an exaggeration to say that it ran the model 2-3x faster which was super impressive. I'm trying to move to local models as much as I can, and I'm finding that it's becoming more and more practical. Admittedly this is on a $6000 dollar laptop (M5 Max Macbook with the specs maxxed out), so the hardware is still a bit out of reach for most people (the AI industry isn't exactly helping here..), but I'm getting the impression that the future is going to be smaller models with more focused training running locally. The danger of giving all your data to these cloud providers just seems too big to me, and I think they're going to start charging insane amounts when they need to show a profit.
- Aurornis 3mo ago> Admittedly this is on a $6000 dollar laptop (M5 Max Macbook with the specs maxxed out) Sadly that laptop is likely $8000 now, after the recent price increase.
- overgard 3mo agoWoof, I thought you might be exaggerating but I priced it out and you're right, it's actually way over $8000 now. I hate how the AI industry is making everything more expensive for everyone.
- bertili 3mo agoWait.. the Qwen Max models have never been open-weight. But it sure sound like that's what they intend now? "Qwen3.8 is launching and going open-weight soon! With a massive 2.4T parameters..."
- simonw 3mo ago(I can't draw a pelican for this one because Alibaba Cloud have flagged my email address and won't let me pay them for access. So I'm waiting for the open weights release, or for the new model to show up on OpenRouter.)
- samxli 3mo agolol the pelican benchmark is basically the only review process I trust at this point. kinda wild that alibaba of all companies is making it hard to give them money tho, you'd think they'd want prominent devs testing their stuff. openrouter usually picks these up pretty fast, hopefully it shows up there soon. You can also just download the qwen app and do this in the chat interface using their MCP tools for local dev.
- timClicks 3mo agoIt's an interesting example of the metric becoming a target. Simon started drawing pelicans because it wasn't something that existed before.
- nextaccountic 3mo ago> kinda wild that alibaba of all companies is making it hard to give them money tho Okay so this is entirely off topic but, are there any options to subscribe to these companies without having a credit card? (I am Brazilian if this makes a difference) It's kind of wild that a Chinese AI lab will require something like.. a credit card with visa or mastercard label, and very little else.
- hizyyo 3mo ago[flagged]
- blmayer 3mo agoLooking forward for a 20ish billion parameter version
- kumanday 3mo agoI did some comparisons to Kimi K3... and found the best is combining the two! Model diversity is powerful. https://trilogyai.substack.com/p/qwen-38-max-benchmark-how-it-compares https://trilogyai.substack.com/p/qwen-38-max-benchmark-how-i...
- cnhwl 3mo agoI’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source competitions—it’s got nothing to do with politics; you’ve simply lost sight of your original aspirations.
- mannycalavera42 3mo ago> instead starts going on about politics, human rights and all that rubbish look, it's just a matter of priorities ¯ \ _ (ツ) _ / ¯
- retox 3mo ago[dead]
- adithyassekhar 3mo agoUnfortunately this is a US centric site, for a US based investor, who themselves and probably most commenters as well who are highly paid silicon valley people who have stakes in ai companies on their side of the pond. Not to mention the political unrest and fear of losing the technological superiority which they once earned but now is trying so hard to hold on to through legally grey monopolistic practices. You dare not outcompete the US in any tech they can sanction. At some point people who made genuine innovation and wanted to make the world a better and equal place were replaced with capitalists. Wait for the downvotes.
- cnhwl 3mo agoRankings have never been something I’ve cared about. Just as with the smear campaigns against China, there is far too much noise in the world. We simply hope that the geeks who are genuinely committed to making the world a better place can unite and focus on getting things done, rather than being brainwashed by ideology and social media.
- colortiles 3mo ago[dead]
- winterscott 3mo ago[dead]
- scottwen 3mo ago[dead]
- scottwen 3mo ago[dead]
- The_resa 3mo agoOpen Source Era is comming
- polterguy-hyper 3mo ago[dead]
- phs318u 3mo agoOff topic, but comment threads like this, where the early comments rabbit hole into something I didn’t come to read, really make me wish HN had collapsible comment controls. My phone screen nearly melted from the insane amount of swiping/scrolling to get to the part of the discussion about the actual LLM.
- boutell 3mo agoI'm making good use of the next button on original replies, but it's true that walking back through the parent button is a long walk sometimes.
- Schlagbohrer 3mo agoI have been using the docling software for PDF and XLSX reading/viewing/comprehension by my local qwen3.6. Claude was telling me that docling includes within it a very small LLM model just to assist with what is basically "super OCR". We are definitely in this era of ultra-mega-huge frontier models and super-tiny-micro models, all finding their uses. On an aside I tested docling by giving it the Ronin TTRPG rulebook PDF and it did an astounding job of converting it all into .txt and .md. Given the wild graphic design of the book that's pretty astounding. Next we'll see if it can read the other Borg books like Mörk Borg and Cy Borg, which are also famous for their messy information.
- jchoksi 3mo agoAn alternative to Docling is Xberg: https://github.com/xberg-io/xberg https://github.com/xberg-io/xberg "A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 97+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server." Amongst other things, you can use it with Open Webui: https://docs.xberg.io/integrations/openwebui/#choosing-an-engine-mode https://docs.xberg.io/integrations/openwebui/#choosing-an-en... fyi, they are releasing their v1.0.0 release shortly and until then you will need to a release candidate docker image tag: https://github.com/xberg-io/xberg/issues/1192 https://github.com/xberg-io/xberg/issues/1192 As a layman, I would say Xberg is better than Docling because its one engine for everything while the former feels like a bunch of scripts and glue logic.
- jingpostmedia 3mo ago[flagged]
- EchooAI 3mo agoWithout a doubt, the Chinese models are going to overtake the American ones. In my opinion, this will happen this year!
- HerrlichDigital 3mo agoI used Qwen before, and I'm pretty curious what the new version is all about.
- niko323 3mo agoKeep up the great work Qwen. You give us options.
- cellardoor32 3mo ago[flagged]
- Helldez 3mo agoI tried it and it's incredible
- elaz48 3mo agothe pricing is this competitive
- zftnb666 3mo agoQwen's models keep getting better. Same problem as the other Chinese models though — international payment is painful. If anyone wants to try Qwen or DeepSeek V4 without the payment headache, api-hub.cc supports both with credit cards.
- nerdalytics 3mo agoI wrote up my experience here: https://dly.to/erfRBMyJhSC https://dly.to/erfRBMyJhSC TL;DR: Qwen3.8-Max-Preview is already good enough that I cancelled my Anthropic subscription. There are other factors too, but Qwen is on a great path. As a European, this is choosing between the plague and cholera. Trusting US companies is as troubling as trusting Chinese companies. But I get more bang for the buck from the qwencloud.com Token Plan than from any US AI lab.
- nerdalytics 3mo agohttps://dly.to/lQTrxZ2ZJYr https://dly.to/lQTrxZ2ZJYr
- amiino 3mo agoThe Yankee Empire is falling apart, they are to blindsided aboiut what is really happening outside US. China both in robotics and in AI is currently leapfrogging US innovation, not they are not distilling US models, its an argument only the loosing side adopts to have an excuse, reality is the chinese engineers are very freaking good. I am from Europe, Denmark to be more precise, and I for sure notice the shift that has heppened the last few years in AI from the Eastern part of the world. You can try to deny it all you want, but China is leading the AI race currently, and what a freaking brilliant move to OpenSource it, it shatters the business model of team Yankee completely.