2 ms·
I've been using v4-flash without vision for this months — it's my go-to for code tasks. Now with vision, I'm wondering: if this model can do everything the text
by cjg007 1mo ago
I've been using v4-flash without vision for this months — it's my go-to for code tasks. Now with vision, I'm wondering: if this model can do everything the text-only version does (plus see images), why keep the text-only one around?
Is it just cost/latency? Or is there something text-only does better?
- bel8 1mo agoYes, it adds vision to the already capable text-only LLM according to DS: > This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. https://api-docs.deepseek.com/news/news260821/ https://api-docs.deepseek.com/news/news260821/