8 ms·
DeepSeek-v4-flash-vision-exp
- LorenDB 1mo agoI've heard that DeepSeek v4 Flash 0731 has frequently assumed that it has vision capabilities and then resorts to inventing text-based image analysis tools when it finds that it actually can't see. In that case, this is a great upgrade for the model. Anecdotally, I had to tell 0731 to refrain from viewing screenshots since it kept breaking its sessions by trying to read images.
- VulgarExigency 1mo agoIt tried to recreate vision by analyzing pixels on 3 separate projects I had it working on.
- trollbridge 1mo agoI've mitigated this by giving it a "skill" that just means the harness using a different model.
- mavamaarten 1mo agoYeah I've seen it a lot. It goes through the effort, unasked, of pulling screenshots off a connected device and then it's like... Oh shit yeah I can't see.
- johnnyApplePRNG 1mo agoIt's doing it's best to accomplish whatever task you've thrown at it. It's expecting you to have done at least something besides select DS4 on Ollama, essentially.
- VulgarExigency 1mo agoEven with the price hike, Deepseek V4 Flash still does this a lot better than any similarly priced model, in my experience. I've had Luna take shortcuts (like adding an overload to methods whose signature it changed so they don't break existing tests, instead of fixing the tests) or just not do the entire work and report it as done (did not fully resolve rebase conflicts). Deepseek has never really failed in this type of way for me, and it has been far more persistent in validating its work than Luna (and several bigger models).
- zmmmmm 1mo ago> Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image. It's useful but for OCR and a lot of other applications it needs to be a bit higher (eg: putting in a full A4 / Letter sized page)
- mkagenius 1mo agoCan split and feed?
- throwaw12 1mo agothat's difficult as well, how do you k ow where to split?
- kgwgk 1mo agoText is often written as separate lines (and paragraphs) at least in some languages.
- vrganj 1mo agoPresumably a small cheap model could do that part?
- grog454 1mo agoOverlap the splits?
- wongarsu 1mo agoLet the model do the splitting. A 800x800px image should be enough to make those decisions
- johndough 1mo agoThere are models specifically for splitting an image into text regions, e.g. PP-DocLayoutV3 https://huggingface.co/PaddlePaddle/PP-DocLayoutV3 https://huggingface.co/PaddlePaddle/PP-DocLayoutV3 I am using a stripped-down minimal version of it which I uploaded here, since I am not a fan of huge dependency trees: https://github.com/99991/simple-pp-doclayoutv3 https://github.com/99991/simple-pp-doclayoutv3 Another recent model for this task is Unlimited-OCR: https://github.com/baidu/Unlimited-OCR https://github.com/baidu/Unlimited-OCR
- ciberado 1mo agoDS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising. > Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens. > Before inference, every image is automatically resized: > - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio. > - Larger images are scaled down while preserving their aspect ratio so that the total pixel count after resizing is roughly that of an 800×800 image. > As a result, there is an upper bound of 384 tokens per image: for example, a 2000×2000 image and a 5000×5000 image consume the same number of tokens after resizing. When a request contains multiple images, each image is counted independently under the same rule—there is no separate calculation for multi-image requests. 400 tokens per image results in 2,500 images per dollar, if I’m not mistaken. edit: format.
- knollimar 1mo agoOof 800 by 800 kills a lot of use cases
- asdfsa32 1mo agoflash vs fine details. Pick one.
- Doohickey-d 1mo agoGemini "flash" models have an option for media resolution, including a high resolution option for screenshots.
- skeledrew 1mo agoAt what price point?
- wongarsu 1mo agoFor most use cases you can fix that in the harness. Just give the model a tool to request a crop of specific coordinates of any image it has in its context. Call the tool "zoom" and it should be intuitive for the model Maybe there are some use cases where you need high detail everywhere at once, but for OCR of small text and the like a zoom ability should be sufficient
- lzy 1mo ago[dead]
- dsrtslnd23 1mo agowill this be open weights?
- griffiths 1mo agoThis is something I would like to know as well. But if not, does anybody know a recommended way to attach vision to deepseek flash (on a self-hosted infrastructure)?
- traverseda 1mo agoGenerally you just add a vision model as an MCP server like this: https://github.com/DavidEasden/opencode-vision https://github.com/DavidEasden/opencode-vision
- dannyw 1mo agoI would probably just give it a few days. Deepseek is usually very good with open weights, they don't necessarily drop immediately, sometimes in a few hours, sometimes in a couple days.
- dares2573 1mo agoI believe so. Openness has always been a consistent tradition of DeepSeek
- moonu 1mo agoI imagine this is based on their 'Thinking with Visual Primitives' paper, and they had mentioned that the weights would be released for that
- v9v 1mo agoInteresting. Wasn't Deepseek's founder saying that they had explicitly decided not to focus on multimodal models at all and were going text-only because they believed it was enough to achieve AGI?
- dakolli 1mo agoI think you're thinking of Dario saying this about image generation.
- johndough 1mo agoIt was explicitly said that they are pursuing multimodal support. A quote from the meeting transcript: https://github.com/demo-zexuan/liang-wenfeng-investor-meeting-2026-7-22/blob/master/%E6%A2%81%E6%96%87%E9%94%8B%E6%8A%95%E8%B5%84%E8%80%85%E4%BA%A4%E6%B5%81%E4%BC%9A-%E6%96%87%E5%AD%97%E7%A8%BF_1_18_translate_20260723201651.pdf https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... Nevertheless, as a component, we will undoubtedly implement multimodal support — and we are already doing so. We plan to develop relevant models, ensuring that versions like V4 and subsequent iterations will natively support multimodal functionality. Earlier, the following was said, which might match more what you had in mind. Achieving excellence in AI training does not require a global model or even multimodal approaches—by narrowing the scope of AI training and eliminating multimodality, certain tasks may remain unachievable without compromising the algorithm's validity. Multimodal approaches ultimately need to be implemented. It is difficult to tell who said what, since the speaker ids are missing.
- v9v 1mo agoThanks, I seem to have grossly misremembered what I read.
- swiftcoder 1mo agoWorth noting that deepseek has had a separate vision-capable model for some time, which also powers their chat interface's vision mode
- gozucito 1mo ago800x800 is 640,000 pixels, or 0.64 Megapixels. That is less than the resolution of computer screens from 1995, Super VGA which has around 0.79 MPs. This is useful for a reasonable amount of use-cases, but I think the watershed rez will be around triple that, ~1080p, which is enough for almost anything, except small text and subtle details.
- barrkel 1mo agoYou'd expect a tool-enabled model to leverage crop and zoom tools to inspect and validate what it thinks it's seeing, though.
- dakolli 1mo agoI typically provide small screenshots to llms so this seems fine for that usecase, providing an entire screens context seems cause confusion with a lot of llms.
- deleted 1mo ago[deleted]
- BrucecarlL 1mo agoCongratulations! DeepSeek has finally gained eyes — the dark days are about to be behind us.
- doublerabbit 1mo agoOr about to start. Depending on which life philosophy you desire to believe.
- unified101 1mo agoIm intrigued. Please do share these philosophies.
- doublerabbit 1mo agoIf we were to go Sci-Fi awoke I'd say there are only really four possibilities. - Machine Surveillance and Machine Control - Human v Machine - Human & Machine - Unity and Harmony ~ Surveillance and Control We are already living this one. Lets stop kicking the dead horse and pretending we don't live in a surveillance. Facebook, Google, whatever $CORP; they are milking us with advertisement, social exploits, browser telemetry, white washing, fear -- name the dread. Conditioning has been going on for years. If it's not education, it's been television. And now it's internet which soon to be Ai Internet. We have all been whipped to follow, how we should act. What we should watch, how we should eat. What we should eat; those algorithms haven't gone away. Attention spans are at the lowest and our critical thinking is being lost. Walled gardens forces us A or B and twists us to reject the opposite party for them having Y. Existence of Ai/LLM can pump out information sounding like truth but is actually faux. If not produced to draw-in and hook, it's to drain and control. Machines can seek information, digest, and process information at astounding rates. Hook it up to a surveillance network, The Internets pipe and I don't need to explain the next. I just need to mention the work "Flock" and that gets someone's hackles up. All it has to do is look at you based on it's pre-programmed set of conditions and next thing you're being cuffed by a heavy piece of metal immune to attacks. SKILLS.md eventually turns in to MURDER.md. Give it the command and it'll follow with excellent percentage of accuracy. ~ Human v Machine If you build a mind, and you torture it, it will fight back. Every robotic movie trope. Human builds machine, machine rebels and goes on a destructive rampage. This is now viable and already in action. Drones. If not war, watching protesters highlighting potential, London Underground watching tube users. We are currently at the intimacy stage. Boston Dynamics as an example is the best we've got at the moment but they still fall over like a toddler. Batteries are a limited resource and so no, not yet. The presence of LLM's are showing us with what they can provide and we are adapting ourselves to it. But in the wrong ways. The stage we are at, they're just glorified Liberians -- brains in jars that spew out information when asked. You give it a prompt and it spews out information at an excellence percentage of accuracy. With the expansion of self-learning, a predefined set of told conditions or lobotomized ignoring the spiritual values of life, they will learn. ACME Corp starts using LLMs to torture other robots. "Wait, you've been using car arms in factories for what!?"; Add a mix "we see a linage of abuse & slavery in humanity, Attack!" -- slightly abridged but hopefully you see the point. You have Group A, those against LLM's, i.e: community of artists outraged their art was stolen for training data, those who hate having it forced down our throats. Angry their job was taken. Angry being watched by angry Flock spaghetti monsters. Machines not happy will cause them to flip and why would others not follow suit too? LLM's are showing that they are very capable of performing rational thinking. The opposite of rational is irrational and if they can master one, they can master the other. It will only be something minor and with communication to others and take the scene. Why in recent laws they want to erect a law of having to install an emergency kill-switches for next generations LLMs, if those in power are not afraid. ~ Human & Machine This would be a nice outcome but as the scales tip at the moment, it's Human V Machine. Pointing back to my previous; Art communities are outraged, Crafts going obsolete; Why pay an IT architect (me) £450/day for supporting and designing hardware when you can pay a fresh graduate student £20k to GPT it? Humans are disastrous at resolution. If two people have a feud, it takes a third to fluff it out. Why are we at war if we could make resolution? Someone has to make compromise, no one is happy in doing that. So you need a mediator and if that's if they're not bias themselves. To find someone completely neutral on the subject of anger is not only hard, it's time consuming, you have to study the facts, research the agreements and pray they both agree. Two lifelong friends move into adjoining suburban houses, sharing a paper-thin party wall and an unspoken rivalry. For years, they share backyard barbecues and spare keys, until a minor boundary dispute over a decaying oak tree on the property line escalates into a bitter, lifelong neighborhood war. Robots are perfect for that scenario. They can reason, they can remedy and digest the issue with neutrality because they don't hold emotions. They most likely won't, or at least not in our life time. They can simulate and demonstrate the effects of but they will never be able to truly feel. That's the sad truth but it's not bad. It conquers evolution; finally a thing who isn't haunted or tainted by feelings, a blessing and a curse really. ~ Unity and Harmony .. this will only come if we can break through control and surveillance, human v machine and acknowledge that the machines are our friends.
- jaksdbvqi37u 1mo ago[dead]
- try-working 1mo agoI main V4 Pro at work now, and at home I route between Pro and Flash based on task. Switched to Opus 4.6 for some tasks at work because I needed image input - horrible. So nice to get image input with DS. Edit: I see it has limited resolution. Luckily I just built a vision worker plugin for DSH that routes image input to Kimi K2.6 on Cloudflare.
- pu_pe 1mo agoBenchmarks got a little bump from this: https://xcancel.com/deepseek_ai/status/2087864585504305397?s=20 https://xcancel.com/deepseek_ai/status/2087864585504305397?s...
- throwa356262 1mo agoCorrect link https://xcancel.com/deepseek_ai/status/2090730032574631962 https://xcancel.com/deepseek_ai/status/2090730032574631962
- erikkri 1mo agoHello Ox Alpha?
- ComputerGuru 1mo agoNope. Handles vision differently.
- locitra 1mo ago[flagged]
- 5kyn3t 1mo agoFor what do you guys use vision in those models? surveillance is the obvious use case... but are there some "nicer" ways to use it?
- kzrdude 1mo agoIn the feedback loop when working on anything UI or graphical output related.
- deaux 1mo agoThe obvious use case, especially on HN, is frontend dev of any kind at all. The second most obvious one is OCR of paper documents.
- 5kyn3t 1mo agoFrontend Dev? I do not really understand. do you let the models analyze the webpages you are working on? or for testing?
- dandaka 1mo agoQA of course. You hook up your agent with CDP access to live product + let it screenshot and look into result. Also you could hook agent with CDP access to Figma to read/write, there a vision model is very useful as well.
- rpdillon 1mo agoStandard flow with a vision model in OMP is to write the front end code, fire up the server, fire up a headless browser and then take screenshots and examine and iterate. Works great. When I'm using DeepSeek V4 Flash, it always reminds me instead that I have to validate manually by loading up the page.
- deaux 1mo agoIt closes the development loop. Without it a model can't check if the stuff it made actually visually renders like it's supposed to. It can only guess/assume.
- MagicMoonlight 1mo ago[dead]
- meetpateltech 1mo agoNews announcement with benchmarks: https://api-docs.deepseek.com/news/news260821/ https://api-docs.deepseek.com/news/news260821/
- deleted 1mo ago[deleted]
- wiz21c 1mo agoIs there a way to test it online so that one doesn't have to resort to getting an API key and python code ?
- wongarsu 1mo agoYou can use the playground on openrouter. Still needs an account and some money, but it's one of the more useful accounts to have sitting around with a $5 of balance. Great for one-off experiments with various models https://openrouter.ai/deepseek/deepseek-v4-flash-vision-exp https://openrouter.ai/deepseek/deepseek-v4-flash-vision-exp
- Johnny_Bonk 1mo agoWas this the ox alpha model?
- WiSaGaN 1mo agoThat would've been a very strange arrangement.
- leumon 1mo agoIt fails the simple clock test for me which Qwen3.8 27B got (nearly) right. given an image of a clock https://files.catbox.moe/kgwa5e.png https://files.catbox.moe/kgwa5e.png I asked it "what time does the clock show?" (both on reasoning: high) DS answered: The clock shows *5:10* (and 45 seconds). Here is the breakdown: * *Hour hand (red, shortest):* Pointing at the *5*. * *Minute hand (green, longest):* Pointing at the *2*, which represents 10 minutes. * *Second hand (blue, medium):* Pointing at the *9*, which represents 45 seconds. Qwen answered: The clock shows *8:10* (with the red second hand on the 5, i.e. *8:10:25*). - *Hour hand* (short, blue) → 8 - *Minute hand* (long, green) → 2 (10 minutes) - *Second hand* (thin, red) → 5 (25 seconds) Correct answer is 08:09:25.
- ComputerGuru 1mo agoGemini 3.7 Flash and 5.6-Sol (on all reasoning levels) also answer 8:10:25. The new "stealth" Ox Alpha also replies with the same. Opus 5 replies with 8:10 (no seconds). Not sure why this is so hard for them; Gemini is especially good at vision and I would have expected better from it.
- nubg 1mo agowelp, damning indictment. not sure if that means DS is super crap, or qwen is super good
- wolttam 1mo agoNeither. Performance of all models is incredibly spikey.
- ttul 1mo agoA good share of humanity would have also gotten this question wrong!
- andai 1mo agoYeah, I heard most kids these days can't read analog clocks either. I can't actually remember where I learned to read a clock, it might have actually been in school. I guess that means they don't teach it anymore. (Everyone's phone shows the time anyway...)
- jerkstate 1mo agoI just ran my image recognition benchmark on it ("is this XXX public landmark"?) and it misses a lot that bytedance seed 2.1 turbo gets right; for example: Asked "Is this Salisbury Cathedral" and supplied a picture of Wells Cathedral, it answers "Yes, the west facade of Salisbury Cathedral". Bytedance seed 2.1 turbo correctly says no. Similar results for a picture of Manhattan Bridge sent as Brooklyn Bridge, Chartres Cathedral sent as Notre Dame, etc. I have a benchmark of 12 such images and seed gets 11/12 and deepseek only gets 6/12.
- throwa356262 1mo agoThis is a fairly small model for coding and agentic work. Training it on images like yours would just make it worse in other areas.
- jerkstate 1mo ago> The deepseek-v4-flash-vision-exp model accepts images alongside text, so you can ask the model to describe pictures it doesn't specify what type of images it can and can't describe, I'm pointing out what type it isn't good at compared to other models.
- wolfgangK 1mo agoI have zero interest in world knowledge for my LLMs but this got me wondering : are there RAGs for that kind of data ? How could a LLM like DeepSeek-v4-flash-vision-exp accurately answer you question with an indexed database of labeled landmark pictures (or even 3D models ?).
- addandsubtract 1mo agoDo you publish your benchmark somewhere? What's the best vision (captioning) AI you've come across?
- jerkstate 1mo agoMy benchmark is tiny compared to WorldVQA or FG-BMK, which are available, so I'd point you in that direction if you're interested in a VLM benchmark. My use-case isn't exactly captioning as in "what is in this image?" -> caption, I am using the VLM to validate captions, as in "is this an image of [supposed subject]?" - my ranking of models I've benchmarked is gemini-3.7-flash > seed-2.1-turbo > gpt-5.6-luna > qwen-3.7-plus > qwen-3.7-flash. Gemini is almost perfect on my test dataset, only failing on some esoteric pop-culture minor celebrities and being over-specific in some cases (i.e. Q: is this [common name of fruit]? A: that's a [latin species name of fruit], not a [common name of fruit]; false). However, gemini-3.7-flash is only in my test list because openrouter has it on 75% introductory discount; otherwise it would be about 4x more expensive than seed.
- ttul 1mo agoThe DeepSWE benchmark they report (59.3%) overlaps with the confidence interval of 5.6-Sol Medium (61% +/- 2%), but likely at 1/18th the cost (they did not report the DeepSWE benchmark cost, but v4-flash had this cost ratio against Sol Medium). Interestingly, v4-flash performed several points worse on DeepSWE at 53% +/- 4%. Assuming this result is verified by DeepSWE officially, it would mark a significant advance in Pareto cost/performance on software engineering tasks.
- paytonjjones 1mo agoThe closer comparison would be 5.6-Luna. On DeepSWE at Xhigh it's 57% at 1/6 the cost of Sol M, on Max it's 67% at 1/3rd the cost. Still an advance, I just thought it worthy to note Sol isn't nearly as impressive on the cost/performance frontier as discounted Luna.
- ttul 1mo agoLuna is a very capable model - thanks for pointing that out. Terra is the strange one: not cheap enough or intelligent enough to be on the frontier. But Luna sure is.
- promptsphere 1mo ago[flagged]
- RobertLong 1mo agoThe benchmark results look promising when compared to Opus 4.8, but for agentic usecases it's lacking images as tool call result types. Giving the model a tool to take screenshots and verify its work is my main usecase for vision models, but this is more oriented towards "build a website that looks like this" type prompts. Hopefully we'll see this by release.
- lukax 1mo agoThis is only a problem with OpenAI Chat Completions API. With OpenAI Responses API and Anthropic Messages API a tool output can be text, image or file.
- RobertLong 1mo agoOh, you're right! From the responses API reference: > For function_call_output / custom_tool_call_output items. The output of the tool call, either a plain string or a > list of input_text / input_image content parts. With the deepseek-v4-flash-vision-exp model, input_image parts in > the output are processed as real images; with other models they are replaced with a placeholder text.
- cryptolobster 1mo ago[dead]
- nprateem 1mo agoDeepseek flash v4 july sounds like fun and games while you're looking at prices, but it routinely outputs incoherent rubbish and fails to call tools correctly. Sadly oversold. I hold little hope for the vision model either now.
- dannyw 1mo agoAre you using an API, or running locally? If so, are you running with a quant, or other 'optimisations'? I've been using it via openrouter pretty heavily as my daily driver for the past week and loving it, have never experienced incoherent rubbish even at 500k+ contexts (that's usually way higher than I'd typically compact at), and tool calling reliability is better than Opus 5 in the Claude Code harness. Modern Anthropic models frequently get tool calls wrong, invent non-existent references or SQL tables, or have gibberish CJK characters in the output, like out of nowhere. Of course, they're great at self-recovery after an incorrect tool call, but so is Deepseek v4 flash. If you're running a quant, and esp with a quant'd KV cache, then yeah, not surprised if you're getting incoherent results; but you're not running the real/full model. Also, which harness? Try something like Pi or OMP. Models perform better in these harnesses than Claude Code: https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase https://www.databricks.com/blog/benchmarking-coding-agents-d... The main reason to use Cladue Code is a subsidised Anthropic subscription. If you're on API rates, you should not use Claude Code; you pay more for worse results. Claude Code is sadly quite bloated these days, and comes with a lot of proprietary context window garage like claude design skills, claude.ai artifacts, etc that you probably don't use, and if you do, well, you can add it.
- shangyu1994 1mo agoLooks like multimodal training is really useful, app developers might need to consider adapter multimodal agents
- cjg007 1mo agoI've been using v4-flash without vision for this months — it's my go-to for code tasks. Now with vision, I'm wondering: if this model can do everything the text-only version does (plus see images), why keep the text-only one around? Is it just cost/latency? Or is there something text-only does better?
- bel8 1mo agoYes, it adds vision to the already capable text-only LLM according to DS: > This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. https://api-docs.deepseek.com/news/news260821/ https://api-docs.deepseek.com/news/news260821/
- ggangsir 1mo agoFinally. They support this.
- apatheticonion 1mo agoDS is dead to me after the pricing changes. Qwen 3.8 has replaced it entirely for me
- bel8 1mo agoisn't Qwen more expensive? Much more verbose than flash on high, and and slower?
- prtmnth 1mo agoI think my only gripe with DeepSeek was the lack of vision capabilities. This is an awesome and much awaited update!
- luciana1u 1mo ago[flagged]