11 ms·
Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
- dvh 10mo agoI asked: "How many resistors are used in fuzzhugger phantom octave guitar pedal?". It replied 29 resistors and provided a long list. Answer is 2 resistors: https://tagboardeffects.blogspot.com/2013/04/fuzzhugger-phantom-octave.html https://tagboardeffects.blogspot.com/2013/04/fuzzhugger-phan...
- iFire 10mo ago> How many resistors are used in fuzzhugger phantom octave guitar pedal? Weird, as someone not having a database of the web, I wouldn't be able to calculate either result.
- iFire 10mo agoI tend to pick things where I think the answer is in the introduction material like exams that test what was taught.
- dvh 10mo ago"I don't know" would be perfectly reasonable answer
- MaxikCZ 10mo agoI feel like theres a time in near future where LLMs will be too cautious to answer any questions they arent sure about, and most of the human effort will go into pleading the LLM to at least try to give an answer, which will almost always be correct anyways.
- plufz 10mo agoThat would be a great if you could have a setting like temperature 0.0-1.0 (Only answer if you are 100% to guess as much as you like).
- littlestymaar 10mo agoIt's not going to happen as the user would just leave the platform. It would be better for most API usage though, as for business doing just a fraction of the job with 100% accuracy is often much preferable than claiming to do 100% but 20% is garbage.
- kaoD 10mo ago> as someone not having a database of the web, I wouldn't be able to calculate either result And that's how I know you're not an LLM!
- esafak 10mo agoThis is just trivia. I would not use it to test computers -- or humans.
- parineum 10mo agoEverything is just trivia until you have a use for the answer. OP provided a we link with the answer, aren't these models supposed to be trained on all of that data?
- esafak 10mo agoThere is nothing useful you can do with this information. You might as well memorize the phone book. The model has a certain capacity -- quite limited in this case -- so there is an opportunity cost in learning one thing over another. That's why it is important to train on quality data; things you can build on top of.
- DennisP 10mo agoJust because it's in the training data doesn't mean the model can remember it. The parameters total 60 gigabytes, there's only so much trivia that can fit in there so it has to do lossy compression.
- littlestymaar 10mo agoIt's good way to assess the model with respect to hallucinations though. I don't think a model should know the answer, but it must be able to know that it doesn't know if you want to use it reliably.
- brookst 10mo agoWhere did you try it? I don’t see this model listed in the linked Qwen chat.
- strangattractor 10mo agoMaybe it thinks some of those 29 are in series:)
- bongodongobob 10mo agoLol I asked it how many rooms I have in my house and it got that wrong. Llms are useless amirite
- cindyllm 10mo ago[dead]
- mettamage 10mo agoI wonder if with that music analysis mode, you can also make your own synths
- sosodev 10mo agoDoes Qwen3-Omni support real-time conversation like GPT-4o? Looking at their documentation it doesn't seem like it does. Are there any open weight models that do? Not talking about speech to text -> LLM -> text to speech btw I mean a real voice <-> language model. edit: It does support real-time conversation! Has anybody here gotten that to work on local hardware? I'm particularly curious if anybody has run it with a non-nvidia setup.
- dsrtslnd23 10mo agoit seems to be able to do native speech-speech
- sosodev 10mo agoIt does for sure. I did some more digging and it does real-time too. That's fascinating.
- red2awn 10mo agoNone of inference frameworks (vLLM/SGLang) supports the full model, let alone non-nvidia.
- sosodev 10mo agoThat's unfortunate but not too surprising. This type of model is very new to the local hosting space.
- AndreSlavescu 10mo agoWe actually deployed working speech to speech inference that builds on top of vLLM as the backbone. The main thing was to support the "Talker" module, which is currently not supported on the qwen3-omni branch for vLLM. Check it out here: https://models.hathora.dev/model/qwen3-omni https://models.hathora.dev/model/qwen3-omni
- red2awn 10mo agoNice work. Are you working on streaming input/output?
- binsquare 10mo agoDoes anyone else find that there's hard to pin down reason of life-lessness in the speech of these voice models? Especially in the fruit pricing portion of the video for this model. Sounds completely normal but I can immediately tell it is ai. Maybe it's intonation or the overly stable rate of speech?
- colechristensen 10mo agoI'm perfectly ok with and would prefer an AI "accent".
- esafak 10mo ago> Sounds completely normal but I can immediately tell it is ai. Maybe that's a good thing?
- sosodev 10mo agoI think it's because they've crammed vision, audio, multiple voices, prosody control, multiple languages, etc into just 30 billion parameters. I think ChatGPT has the most lifelike speech with their voice models. They seem to have invested heavily in that area while other labs focused elsewhere.
- Lapel2742 10mo agoIMHO it's not lifeless. It's just not overly emotional. I definitely prefer it that way. I do not want the AI to be excited. It feels so contrived. On the video itself: Interesting, but "ideal" was pronounced wrong in German. For a promotional video, they should have checked that with native speakers. On the other hand its at least honest.
- nunodonato 10mo agoI hate with a passion the over-americanized "accent" of chatgpt voices. Give me a bland one any day of the week
- wkat4242 10mo ago
- rarisma 10mo agoGPT4o in the charts is crazy.
- BoorishBears 10mo agoWhy? gpt-realtime is finalized gpt-4o. Gemini Live is still 2.5. Not their fault frontier labs are letting their speech to speech offerings languish.
- banjoe 10mo agoWow, crushing 2.5 Flash on every benchmark is huge. Time to move all of my LLM workloads to a local GPU rig.
- embedding-shape 10mo agoJust remember to benchmark it yourself first with you private task collection, so you can actually measure them against each other. Pretty much any public benchmark is unreliable at this moment, and making model choices based on other's benchmarks is bound to leave you disappointed.
- MaxikCZ 10mo agoThis. Last benchmarks of DSv3.2spe hinted at beating basically everything, yet in my testing even sonnet is miles ahead both in terms of speed and accuracy
- red2awn 10mo agoWhy would you use an Omni model for text only workload... There is Qwen3-30B-A3B.
- skrunch 10mo agoExcept the image benchmarks are compared against 2.0, which seems suspicious that they would casually drop to an older model for those.
- gardnr 10mo agoThis is a 30B parameter MoE with 3B active parameters and is the successor to their previous 7B omni model. [1] You can expect this model to have similar performance to the non-omni version. [2] There aren't many open-weights omni models so I consider this a big deal. I would use this model to replace the keyboard and monitor in an application while doing the heavy lifting with other tech behind the scenes. There is also a reasoning version, which might be a bit amusing in an interactive voice chat if it pronounces the thinking tokens while working through to a final answer. 1. https://huggingface.co/Qwen/Qwen2.5-Omni-7B https://huggingface.co/Qwen/Qwen2.5-Omni-7B 2. https://artificialanalysis.ai/models/qwen3-30b-a3b-instruct https://artificialanalysis.ai/models/qwen3-30b-a3b-instruct
- gardnr 10mo agoI can't find the weights for this new version anywhere. I checked modelscope and huggingface. It looks like they may have extended the context window to 200K+ tokens but I can't find the actual weights.
- pythux 10mo agoThey link to: https://huggingface.co/collections/Qwen/qwen3-omni-68d100a86cd0906843ceccbe?spm=a2ty_o06.30285417.0.0.24bac921KI4bG2 https://huggingface.co/collections/Qwen/qwen3-omni-68d100a86... from the blog post but it does seem like this redirects to their main space on HF so maybe they didn't yet make the model public?
- olafura 10mo agoLooks like it's not open source: https://www.alibabacloud.com/help/en/model-studio/qwen-omni#2d8d6c9ca5e1c https://www.alibabacloud.com/help/en/model-studio/qwen-omni#...
- coder543 10mo agoNo... that website is not helpful. If you take it at face value, it is claiming that the previous Qwen3-Omni-Flash wasn't open either, but that seems wrong? It is very common for these blog posts to get published before the model weights are uploaded.
- deleted 10mo ago[deleted]
- Aissen 10mo agoIs this a new proprietary model?
- deleted 10mo ago[deleted]
- sim04ful 10mo agoThe main issue I'm facing with realtime responses (speech output) is how to separate non-diegetic outputs (e.g thinking, structured outputs) from outputs meant to be heard by the end user. I'm curious how anyone has solved this
- artur44 10mo agoA simple way is to split the model’s output stream before TTS. Reasoning/structured tokens go into one bucket, actual user-facing text into another. Only the second bucket is synthesized. Most thinking out loud issues come from feeding the whole stream directly into audio.
- pugio 10mo agoThere is no TTS here. It's a native audio output model which outputs audio tokens directly. (At least, that's how the other real-time models work. Maybe I've misunderstood the Qwen-Omni architecture.)
- artur44 10mo agoTrue, but even with native audio-token models you still need to split the model’s output channels. Reasoning/internal tokens shouldn't go into the audio stream only user-facing content should be emitted as audio. The principle is the same, whether the last step is TTS or audio token generation.
- regularfry 10mo agoThere's an assumption there that the audio stream contains an equivalent of the <think>/</think> tokens. Every reason to think it should, but without seeing the tokeniser config it's a bit of a guess.
- stevenhuang 10mo agoWayback for those that can't reach https://web.archive.org/web/20251210164048/https://qwen.ai/blog?id=qwen3-omni-flash-20251201 https://web.archive.org/web/20251210164048/https://qwen.ai/b...
- terhechte 10mo agoIs there a way to run these Omni models on a Macbook quantized via GGUF or MLX? I know I can run it in LMStudio or Llama.cpp but they don't have streaming microphone support or streaming webcam support. Qwen usually provides example code in Python that requires Cuda and a non-quantized model. I wonder if there is by now a good open source project to support this use case?
- mobilio 10mo agoYes - there is a way: https://github.com/ggml-org/whisper.cpp https://github.com/ggml-org/whisper.cpp
- novaray 10mo agoWhisper and Qwen Omni models have completely different architectures as far as I know
- tgtweak 10mo agoYou can probably follow the vLLM instructions for omni here, then use the included voice demo html to interface with it: https://github.com/QwenLM/Qwen3-Omni#vllm-usage https://github.com/QwenLM/Qwen3-Omni#vllm-usage https://github.com/QwenLM/Qwen3-Omni?tab=readme-ov-file#launch-local-web-ui-demo https://github.com/QwenLM/Qwen3-Omni?tab=readme-ov-file#laun...
- aschobel 10mo agoLooks to be API only. Bummer.
- deleted 10mo ago[deleted]
- readyplayeremma 10mo agoThe models are right here, one of the first links in the post: https://huggingface.co/collections/Qwen/qwen3-omni https://huggingface.co/collections/Qwen/qwen3-omni edit: Nevermind, in spite of them linking it at the top, they are the old models. Also, the HF demo is calling their API and not using HF for compute.
- aschobel 10mo agoIt is super confusing. I also thought this initially was open weights.
- Alifatisk 10mo agoIt seems to be available on Qwen chat? https://chat.qwen.ai/settings/model?id=qwen3-omni-flash-2025-12-01 https://chat.qwen.ai/settings/model?id=qwen3-omni-flash-2025...
- devinprater 10mo agoWow, just 32B? This could almost run on a good device with 64 GB RAM. Once it gets to Ollama I'll have to see just what I can get out of this.
- plipt 10mo agoI see that their HuggingFace link goes to some Qwen3-Omni-30B-A3B models that show a last updated date of September The benchmark table in their article shows Qwen3-Omni-Flash-2025-12-01 (and the previous Flash) as beating Qwen3-235B-A22B. How is that possible if this is only a 30B-A3B model? Also confusing how that comparison column starts out with one model but changes them as you descend down the table. I don't see any FLASH variant listed on their Hugginface. Am i just missing it or are these specifying a model only used for their API service and there are no open weights to download?
- apexalpha 10mo agoI run these on a 48gb Mac because of the universal ram.
- vessenes 10mo agoInteresting - when I asked the omni model at qwen.com what version it was, I got a testy "I don't have a version" and then was told my chat was blocked for inappropriate content. A second try asking for knowledge cutoff got me the more equivocal "2024, but I know stuff after that date, too". No idea how to check if this is actually deployed on qwen.com right now.
- zamadatix 10mo ago> No idea how to check if this is actually deployed on qwen.com right now. Assuming you mean qwen.ai, when you run a query it should take you to chat.qwen.ai with the list of models in the top left. None of the options appear to be the -Omni variant (at least when anonymously accessing it).
- vessenes 10mo agoThanks - yes - I did. The blog post suggests clicking the 'voice' icon on the bottom right - that's what I did.
- mh- 10mo agoFor what it's worth, that's not a reliable way to check what model you're interacting with.
- vessenes 10mo agoIt’s a good positive signal, but not a good negative one. It would be convincing if it said “I’m qwen-2025-12-whatever”. I agree it’s not dispositive if it refuses or claims to be llama 3 say. Generally most models I talk to do not hallucinate future versions of themselves, in fact it can be quite difficult to get them to use recent model designations; they will often autocorrect to older models silently.
- forgingahead 10mo agoI truly enjoy how the naming conventions seem to follow how I did homework assignments back in the day: finalpaper-1-dec2nd, finalpaper-2-dec4th, etc etc.
- mohsen1 10mo agoHaving lots of success with Gemini Flash Live 2.5. I am hoping 3.0 to come out soon. Benchmarks here claim better results that Gemini Live but have to test it. In past I've always been disappointed with Qwen Omni models in my English-first case...
- andy_ppp 10mo agoQwen seem to be deliberately confusing about if they are releasing models open weight or not. I think largely not any more and you can go on quite a wild goose chase looking for different things that are implied they are released but are actually only available via API.