3 ms·
Qwen 3.8 Omni Flash
- tolugenius 9d agoCurious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.
- _ache_ 9d agoI don't think Qwen3.8-Omni-X will ever be released. The last one was: Qwen3-Omni-30B-A3B https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.
- imrehg 9d agoOut of curiosity, what's makes the 125B unsuable? (performance of running it, the quality of that version of the model, or something else?)
- bitexploder 9d agoNo, it is actually very good. Qwen Flash 3.8 Next is fine. But you need ~128GB of RAM to get it going and not a lot of people have that or can serve it very quickly. I have been running it on an old gaming system around 25 t/s to do overnight work and it is very strong, even at 3 bit quant.
- diddid 8d agoQwen 3.8 Flash Next is amazing, i did hundreds of turns and billions of prefill and it may not be as smart as sota but then again it does what i tell it to and it does it well.
- mdp2021 8d ago> definitively slowing down Surely it was meant to be 'definitely' - the "good news" at this stage are that given the speed of history and important levels of uncertainty, it is difficult to label trends with "definitively" ;) Some would not have bet that the change of management at Qwen would have kept similar good results, but there we are, presumably satisfied. Other changes will happen, there or elsewhere - the situation is still very open. And when the "40Watts Intelligence" (which we know possible) will be implemented... It will be a testimony that the current was only a middle-way, temporary, dynamic stage.
- Iolaum 8d agoWhy mostly unusable 125b? I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.
- _ache_ 8d agoVery capable yes but very slow. 27B is relatively easy to run, but the 125b one need around 128Gb of RAM (DDR4 isn't enough, you need DDR5 to be quick enough, that's $3000 alone, you also need a graphic card). DDR4 is caped @20tps. So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people. With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow. Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).
- vinzenzu 8d agoQwen 3.8 Max was open-weight released [1], as was Qwen 3.8 Flash Next [2]. I still agree that they aren't as aggressively releasing the open-weights models as before, but there hasn't been a major release they haven't published the weights for yet afaik. [1] https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B [2] https://huggingface.co/Qwen/Qwen3.8-Flash-Next https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- npodbielski 8d agoI am using Flash Next for few weeks and it is very capable model. I just wish there would a way to have faster prefill because reloading longer sections of session sometimes can take even 2h. I stopped using Qwen 3.8 27B completely on my Strix Halo.
- hgoel 8d agoThe entire point of Qwen3.8-Next-Flash is to allow inference engines to implement support for the Qwen4 architecture, so they're ready by the time qwen4 is ready for release.
- conception 9d ago3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
- rubslopes 9d ago> RL it to oblivion. What would that mean in this context?
- cleaning 9d agoSee 5.6, Astra, and Opus 4.8 for examples
- smallerfish 9d agoWhat are they examples of? Opus 4.8 was much better than the infamous 5, and I find Astra generally competent.
- pennomi 9d agoTuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness. I swear I spend more time telling Claude not to do things than telling it what to do.
- mdp2021 8d ago> aggressively useful ... in the name of helpfulness But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?
- girvo 8d agoIt’s absolutely down to their post-training RL, yeah. It’s where most of its strongest behaviour comes from, with regards to this kind of agentic behaviour
- syntaxing 9d ago> audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash. They also made a new harness but github link seems to 404.
- testaburger 8d agoprobably their distill target
- JLO64 8d ago[dead]
- _ache_ 9d agoIf the performances are comparable, and there is no evidence it's not. in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47 That is a massive cost reduction. Refs: https://www.alibabacloud.com/help/en/model-studio/model-pricing#china-beijing-h4 https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni https://runware.ai/gemini-omni
- killingtime74 8d agoYou can't just look at the per token cost, but how many tokens it takes on average to do a task. The difference can be massive.
- vntok 8d agoTrue, but it would have to be more than massive (order(s) of magnitude) to offset that gap.
- kaliqt 8d agoWe notice with frontier models like Astra and Fable that one might use a lot less tokens than the other to complete the task thereby being the better deal in spite of the far higher token cost.
- LeBit 8d agoWhat is Astra $/task? Even if OpenAI end up using 1 token for per task, if the token costs 1M$ , some people will find it expensive.
- nater5000 8d agohttps://artificialanalysis.ai/#intelligence-comparison-tabs https://artificialanalysis.ai/#intelligence-comparison-tabs Astra on xhigh has a cost per task of $2.31 with an intelligence index of 53. Qwen3.8 Max has a cost per task of $5.41 with an intelligence index of 45. Pricing for GPT-6 Astra (xhigh) is $10.00 per 1M input tokens and $50.00 per 1M output tokens. Pricing for Qwen3.8 Max (0902) is $2.00 per 1M input tokens and $6.00 per 1M output tokens. Obviously this is just one measure of all of this (and Qwen 3.8 Omni Flash isn't yet available), but I think this illustrates the point well. These relative task costs are pretty consistent across different analysts. Cost per token is arguably a useless measure at this point in most circumstances.
- lxe 8d agoLooks like the harness repo is already removed?
- andy_ppp 8d agoWhy can the Chinese build models and Europe cannot? The algorithms behind this stuff are not that complicated, are they? Is it the cost of energy? The illegality and difficulty of obtaining all the data in the world? Lack of capital to start moonshot labs? Lack of optimism? The Chinese just seem to have an ability to get it done without anywhere near the GPUs of the US and Europe can buy these GPUs. I think relying on the US and China for AI is probably not ideal? For example I think Qwen have not released the Omni models as open weights in the past, it’d be good to know if they’re doing this here?
- apexalpha 8d agoThe Chinese were directed to it from their government. We have no such government with a mandate to do that.
- EagnaIonat 8d agoWhile there is no official mandate, many countries of the EU are working on models that are government funded. The main one (IIRC) is Spains ALIA. https://alia.gob.es/eng https://alia.gob.es/eng There is also OpenEuroLLM and EuroLLM. https://www.openeurollm.eu https://www.openeurollm.eu
- jambutters 8d agoNah, they wouldn't have released deepseek open source if they were the case
- Zambyte 8d agoHow are you sure about that? Undercutting your competition (at cost even) to minimize their power is a common strategy.
- podocarp 8d agoSource?
- rw2 8d agoGreat engineering schools and good tech companies to train them post graduation. Europe has none; The best tech university in Europe when compared to Chinese/US equivalents won't even rank in the top 10.
- podocarp 8d agoPlease what is flash pro ultra and all these, can they just use semver or something
- spacebanana7 8d agoI believe in general flash models prioritise speed, ultra/pro/omni do more slow reasoning to the effect of sometimes better intelligence, and lite models prioritise cost.
- drbscl 8d agoOmni usually means multimodality (in terms of input and/or output type, text, images, audio, etc)
- xutopia 8d agoThe different names have different meanings and help make decisions on which to use. Omni means you can use multiple types of input and have multiple types of outputs like audio, video, images and text. Flash means that it is built for speed and smaller than the more complete ones.
- mavamaarten 8d agoI'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed. I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try. E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.
- apimade 8d agoFor TTS I launched this like this week based on a Reddit thread of recommendations, added a new one to it.. I want to say on Wednesday, but this week has been a blur. Problem I’ve found with similar sites is I can’t run a lot of the models, or the results are beyond stale. But this is only stuff I can run locally, or it’s a cloud model. So.. This is good for right now! https://apimade.com/audio-compare.html https://apimade.com/audio-compare.html
- Systemerror7A69 8d agoMy approach to this problem is to just...not try them all. As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe. So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't have to change. And even if it can't do something, the first thing I change is see if I can give it more tools or better context (useful even if I switch models later) or trying a different approach to the problem. If google/gemma-4-31b works, you don't need to overthink it.
- bmordue 8d ago
- esquire_900 8d agoThe blog post itself is technically unimpressive, and feels like the slop future. Amongst many things - Scrolling in firefox is a nightmare (only uBlock origin) and the console is full or warnings and debug data. - Videos are all flashy but just fail to communicate anything beyond what can be said in a small paragraph (and with horrible stock music). - Figure 1, the headpiece; too small to read, can't zoom in
- zepearl 8d agoFunny, that website ( https://qwen.ai/blog?id=qwen3.8-omni-flash https://qwen.ai/blog?id=qwen3.8-omni-flash ) downloads automatically 441 MiB of mp4 files... (using Firefox on Linux PC - noticed it because my network chart in Gkrellm spiked for several seconds).