12 ms·
DeepSeek OCR
- farseer 1y agoHow good is this compared to most commercial OCR software?
- ozim 1y agoAny vision model is better than commercial OCR software.
- szundi 1y ago[dead]
- Etheryte 1y agoI'm not really sure if that's an accurate summary of the state of the art, [0] is a better overview. In short, SOTA multi-modal LLMs are the best option for handwriting, nearly anything is good at printed text, for printed media, specialty models from hyperscalers are slightly better than multi-modal LLMs. [0] https://research.aimultiple.com/ocr-accuracy/ https://research.aimultiple.com/ocr-accuracy/
- dragonwriter 1y agoSince “commercial OCR software” includes VLM-based commercial offerings, that's clearly not correct.
- empressplay 1y agoThis could be great for extracting text from old magazines; traditional OCR gives you a bit of a mess you have to clean up, but this looks like it can properly identify columns and track the flow accurately (and extract images!) It appears it can convert magazine layouts to markdown too
- piker 1y agoThis looks really cool for prototyping and playing around. It seems to me though if one is building a modern application that needs to get image segmentation and/or text recognition right there are better APIs available than natural language? It seems like a lot of effort to make a production-scale CV application to weigh it down with all of an LLM’s shortcomings. Not a field I’m familiar with but I would assume that this doesn’t produce state of the art results—that would change the analysis.
- randomNumber7 1y agoImagine you build an image segmentation model for a e.g. specific industrial application. With this LLM approach you can at least create your training data from the raw images with natural language.
- piker 1y agoThat does make sense
- CheeseFromLidl 1y agoAs a hobby photographer, I organise everything for speedy retrieval but this would be amazing to search my collection.
- krackers 1y agoThe paper is more interesting than just another VLM for OCR, they start talking about compression and stuff. E.g. there is this quote >Our work represents an initial exploration into the boundaries of vision-text compression, investigating how many vision tokens are required to decode text tokens. The preliminary results are encouraging: DeepSeek-OCR achieves near-lossless OCR compression at approximately 10× ratios, while 20× compression still retains 60% accuracy. (I guess you could say a picture token is worth 10 textual tokens...) Could someone explain to a noob what the information-theoretic intuition is here? Why does this work, is it that text tokens are still too "granular"/repetitive and don't come close to the ideal entropy coding? Or is switching to vision tokens escaping the limitation of working "one word-ish at a time", allowing you to get closer to entropy (similar to the way that arithmetic encoding does compared to huffman codes)? And then they start talking about handling long-context by literally(?) downscaling images, forming a correspondence between information loss in the textual domain and the image domain.
- looobay 1y agoLLMs are compute heavy with quadratic scaling (in compute) per tokens. They are trying to compress text tokens into visual tokens with their VLM. Maybe they would render texts to an image before tokenizing to reduce the compute cost.
- krackers 1y agoBut naively wouldn't you expect the representation of a piece of text in terms of vision tokens to be roughly the same number of bits (or more) than the representation as textual token? You're changing representation sure, but that by itself doesn't give you any compute advantages unless there is some sparsity/compressability you can take advantage of in the domain you transform to right? So I guess my question is where is the juice being squeezed from, why does the vision token representation end up being more efficient than text tokens.
- looobay 1y agoVision tokens are a good compression medium because with one vision token you have one vector of N elements, but with textual tokens you have M vectors of N elements, because one vision token represent multiple pixels (and possibly multiple words). This is why its a good compression medium for compute. It will never be as precise as textual tokens but it can be really good as they show in the paper.
- bugglebeetle 1y agoLooks great, but looking at the benchmark, can’t help but think about how crazy good dots-ocr is as a model. Too bad they’re not as open as the Deepseek team because its so crazy good and would love to know how it was trained.
- rfoo 1y agoIf you look you'd notice that it's the same Haoran Wei behind DeepSeek-OCR and GOT-OCR2.0 :p
- bugglebeetle 1y agoOh you’re right! Good catch!
- bethekind 1y agoDid we read the same graph? DeepSeek Gundam 200 dpi appeared to get similar perf as dots-ocr, but with less tokens needed. The x axis is inverted, descending with distance from the origin.
- k_sze 1y agoIt's interesting how they use "Gundam" in their variant names. I gather that Gundam-M and Gundam are their most powerful ones.
- daemonologist 1y agoI think maybe to distinguish their dynamic resolution approach from the t-shirt sizes, which have a fixed input. (Although I don't know why "Gundam")
- brightUiso 1y agoPlease a bit of education, what does it do?
- yoran 1y agoHow does an LLM approach to OCR compare to say Azure AI Document Intelligence (https://learn.microsoft.com/en-us/azure/ai-services/document-intelligence/overview?view=doc-intel-4.0.0 https://learn.microsoft.com/en-us/azure/ai-services/document...) or Google's Vision API (https://cloud.google.com/vision?hl=en https://cloud.google.com/vision?hl=en)?
- sandblast 1y agoNot sure why you're being downvoted, I'm also curious.
- ozgune 1y agoOmniAI has a benchmark that companies LLMs to cloud OCR services. https://getomni.ai/blog/ocr-benchmark https://getomni.ai/blog/ocr-benchmark (Feb 2025) Please note that LLMs progressed at a rapid pace since Feb. We see much better results with the Qwen3-VL family, particularly Qwen3-VL-235B-A22B-Instruct for our use-case.
- CaptainOfCoit 1y agoMagistral-Small-2509 is pretty neat as well for its size, has reasoning + multimodality, which helps in some cases where context isn't immediately clear, or there are few missing spots.
- cheema33 1y agoOmni OCR team says that according to their own benchmark, the best OCR is the Omni OCR. I am quite surprised.
- numpad0 1y agoClassical OCR still probably make undesirable su6stıtutìons in CJK from there being far too many of similar ones, even some absurd ones that are only distinguishable under microscope or by looking at binary representations. LLMs are better constrained to valid sequences of characters, and so they would be more accurate. Or at least that kind of thing would motivate them to re-implement OCR with LLM.
- CloseChoice 1y agoIt's deepseek so one can expect an open-source license but for anyone (like me) who wants to see that explicitly, since it's not obvious in the GitHub repo: https://huggingface.co/deepseek-ai/DeepSeek-OCR/blob/main/LICENSE https://huggingface.co/deepseek-ai/DeepSeek-OCR/blob/main/LI... TLDR: It's MIT licensed
- AndroTux 1y ago> since it's not obvious in the GitHub repo Literally says MIT license on the right sidebar and in the readme tab and in the file called LICENSE
- maxloh 1y agoModel weights are MIT too: https://huggingface.co/deepseek-ai/DeepSeek-OCR/blob/main/LICENSE https://huggingface.co/deepseek-ai/DeepSeek-OCR/blob/main/LI...
- x______________ 1y ago>先天下之忧而忧 How is this an example of a prompt? Google translated this to "Worry about the world first" while Bing says "Worry before the worries of the world." Can anyone shed some light on this saying or why it's in the article?
- fspeech 1y agoGoogle is closer. This is from a famous essay expressing tbe author's desire to bear the burden for the world. Essay is 岳阳楼记 by 范仲淹 in year 1046 https://zh.wikisource.org/zh-hans/%E5%B2%B3%E9%99%BD%E6%A8%93%E8%A8%98 https://zh.wikisource.org/zh-hans/%E5%B2%B3%E9%99%BD%E6%A8%9...
- SequoiaHope 1y agoAsk a language model - ChatGPT says it’s a line from a famous poem “Memorial to Yueyang Tower” which expresses the Confucian ideal of selfless concern for people and society.
- raincole 1y agoIt's a very famous (classical) Chinese phrase. Both translations don't catch the meaning well though. It means: "worry before the rest of the world (notice that they have something to) worry." The next part is 後天下之樂而樂("be happy only after the rest of the world is happy.") I don't know why it's a prompt example.
- jdthedisciple 1y agoSibling comment has the second part as 后天下之乐而乐 which one is correct?
- ellisd 1y agoThe paper makes no mention of Anna’s Archive. I wouldn’t be surprised if DeepSeek took advantage of Anna’s offer granting OCR researchers access to their 7.5 million (350 TB) Chinese non-fiction collection ... which is bigger than Library Genesis. https://annas-archive.org/blog/duxiu-exclusive.html https://annas-archive.org/blog/duxiu-exclusive.html
- throawayonthe 1y agohahaha also immediately thought of this, wonder when the ocr'd dataset would be getting released
- singularfutur 1y agoYes it means they will never release their dataset :(
- _vqpz 1y agoWhy do they need to grant access for people to use copies of books they don’t own?
- JohnLocke4 1y agoNot to rationalize it, but it appears that they're gatekeeping the dataset to get access to the OCR-scans from the people they choose to share it with. This is to improve their existing service by making the content of books (and not just their title/tags) searchable. As per the blog post: >What does Anna’s Archive get out of it? Full-text search of the books for its users.
- _vqpz 1y agoFair enough, it just seems like they're painting an even bigger target on their backs by restricting access to copyrighted material they don't own the rights to
- est 1y ago> The books from Duxiu have long been pirated on the Chinese internet. Usually they are being sold for less than a dollar by resellers. They are typically distributed using the Chinese equivalent of Google Drive, which has often been hacked to allow for more storage space Ownership laundering.
- mrasong 1y agoKinda reminds me of PaddleOCR. Would be awesome if DeepSeek OCR could be integrated into a mobile app someday. That’d make OCR way more convenient!
- pzo 1y agoiOS already have on device both text detector and document scanner in apple Vision API. Hard to say how good are they compared to LLM based solutions. Similarly google had MLKit with OCR working on devices also for many years.
- pietz 1y agoMy impression is that OCR is basically solved at this point. The OmniAI benchmark that's also referenced here wasn't updated with new models since February 2025. I assume that's because general purpose LLMs have gotten better at OCR than their own OCR product. I've been able to solve a broad range of OCR tasks by simply sending each page as an image to Gemini 2.5 Flash Lite and asking it nicely to extract the content in Markdown under some additional formatting instructions. That will cost you around $0.20 for 1000 pages in batch mode and the results have been great. I'd be interested to hear where OCR still struggles today.
- kbumsik 1y ago> My impression is that OCR is basically solved at this point. Not really in practice to me. Especially they still struggle with Table format detection.
- coulix 1y agoThis. Any complex parent table span cell relationship still has low accuracy. Try the reverse, take a complex picture table and ask Chatgpt5, claude Opus 3.1, Gemini Pro 2.5 to produce a HTML table. They will fail.
- bobsmooth 1y agoMaybe I misunderstood the assignment but it seems to work for me. https://chatgpt.com/share/68f5f9ba-d448-8005-86d2-c3fbae028bcf https://chatgpt.com/share/68f5f9ba-d448-8005-86d2-c3fbae028b... Edit: Just caught a mistake, transcribed one of the prices incorrectly.
- kbumsik 1y agoRight, I wouldn't use full table detection to VLM model because they tend to mistake with numbers in table...
- pietz 1y agoMaybe my imagination is limited or our documents aren't complex enough, but are we talking about realistic written documents? I'm sure you can take a screenshot of a very complex spreadsheet and it fails, but in that case you already have the data in structured form anyway, no?
- singularity2001 1y agoInstead of downloading a specific OCR model how would one fare just downloading the currently best multi-modal foundation model? And what would that be at less than 30 GB?
- prats226 1y agoThen you can just download finetuned version of same multi-modal foundation model that's trained on documents?
- saltserv 1y ago[dead]
- ammar_x 1y agoLanguage support is not mentioned in the repo. But from the paper, it offers extensive multilingual support (nearly 100 languages) which is good, but I need to test it to see how it compares to Gemini and Mistral OCR.
- zacmps 1y agoI suspect the number of langauges it can do with reasonable accuracy is actually much smaller, probably <15.
- 2big2fail_47 1y agoI find it interesting that there's all these independent AI-OCR Projects but still no commercial offering. Is it still too inaccurate, too complex or simply too expensive?
- Annatar01 1y agoI dont know, but maybe existing commercial OCR is still on top, and also using ML. Recently tried a free trial for OCR/reading Sütterlin and it was a weird feeling being so outclassed in reading.
- Eisenstein 1y agoIt is because the AI is not actually doing OCR. It is giving an interpretation of what the text in an image is by ingesting vision tokens and mapping them onto text tokens. So you either have to be fine with a lot of uncertainty as to the accuracy of that interpretation or you have to wait for an LLM that can do it in a completely reproducible way every time.
- rsolva 1y agoMistral offers their OCR commercially through their API and in their Chat services, at least. https://mistral.ai/news/mistral-ocr https://mistral.ai/news/mistral-ocr
- simlevesque 1y agohttps://cloud.google.com/document-ai https://cloud.google.com/document-ai
- daemonologist 1y agoThere are commercial OCR offerings from the big cloud providers (plus, like, Adobe). In my experience they generally outperform anything open-weights, although there's been a lot of improvement in VLMs in the past year or two.
- aleinin 1y agoOne that I’ve seen recently is https://reducto.ai https://reducto.ai It appears to be an OCR wrapper.
- tinyhouse 1y agoOCR is not a great name for these models. While they can do traditional OCR such as digitize and scanned PDF for example, they do so much more.
- intalentive 1y ago>they do so much more I'm not familiar. What else are they good for?
- tinyhouse 1y agoThey can take something like an image of a graph and provide a description of it. From my understanding, these are multimodal models with reasoning capabilities.
- breadislove 1y agoFor everyone wondering how good this and other benchmarks are: - the OmniAI benchmark is bad - Instead check OmniDocBench[1] out - Mistral OCR is far far behind most Open Source OCR models and even further behind then Gemini - End to End OCR is still extremely tricky - composed pipelines work better (layout detection -> reading order -> OCR every element) - complex table parsing is still extremely difficult [1]: https://github.com/opendatalab/OmniDocBench https://github.com/opendatalab/OmniDocBench
- hakunin 1y agoWish someone benchmarked Apple Vision Framework against these others. It's built into most Apple devices, but people don't know you can actually harness it to do fast, good quality OCR for you (and go a few extra steps to produce searchable pdfs, which is my typical use case). I'm very curious where it would fall in the benchmarks.
- wahnfrieden 1y agoIt is unusable trash for languages with any vertical writing such as Japanese. It simply doesn’t work.
- thekid314 1y agoYeah, and fails quickly at anything handwritten.
- hakunin 1y agoI mostly OCR English, so Japanese (as mentioned by parent) wouldn't be an issue for me, but I do care about handwriting. See, these insights are super helpful. If only there was, say, a benchmark to show these. My main question really is: what are practical OCR tools that I can string together on my MacBook Pro M1 Max w/ 64GB Ram to maximize OCR quality for lots of mail and schoolwork coming into my house, all mostly in English. I use ScanSnap Manager with its built in OCR tools, but that's probably super outdated by now. Apple Vision does way better job than that. I heard people say also that Apple Vision is better than Tesseract. But is there something better still that's also practical to run in a scripted environment on my machine?
- loaderchips 1y agoGreat work guys, how about we replace the global encoder with a Mamba (state-space) vision backbone to eliminate the O(n²) attention bottleneck, enabling linear-complexity encoding of high-resolution documents. Pair this with a non-autoregressive (Non-AR) decoder—such as Mask-Predict or iterative refinement—that generates all output tokens in parallel instead of sequentially. Together, this creates a fully parallelizable vision-to-text pipeline, The combination addresses both major bottlenecks in DeepSeek-OCR.
- loaderchips 1y agonot sure why i m getting downvoted. Would love to have a technical discussion on the validity of my suggestions.
- neves 1y agoI see that the project uses conda for development. Is it still a good tool now that pip also install binaries?
- modeless 1y agoNo. Everyone should be using uv instead.
- foofoo12 1y agoHow does it compare to Tesseract? https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract I use ocrmypdf (which uses Tesseract). Runs locally and is absolutely fantastic. https://ocrmypdf.readthedocs.io/en/latest/ https://ocrmypdf.readthedocs.io/en/latest/
- utopiah 1y agoIndeed, seems the default benchmark is LLM/VLM based alternatives as if they somehow "solved" the problem but IMHO even if it goes from (totally made up numbers) 80% with tesseract to 95% with this or Qwen or whatever but it takes 100x harddisk with containers or a CUDA stack, dedicated hardware, e.g. GPU with 16GB or VRAM, etc then it's such a trade of it should be considered.
- tidbeck 1y agoHow does this compare to https://huggingface.co/ibm-granite/granite-docling-258M https://huggingface.co/ibm-granite/granite-docling-258M in performance and how they work?
- bugglebeetle 1y agoThe granite Dockling models are unfortunately quite far below SOTA. dots-ocr and PaddleOCR were best here.
- hank2000 1y agoHave yall seen tensorlake? I’m curious how this compares to a model custom built for the problem. My guess is it can be as good. But can it be as efficient? disclaimer: I do not work for tensorlake—but i know the folks behind it.
- modeless 1y agoHmm, at first I was thinking "why OCR?", but maybe the reason is to ingest more types of training data for LLM improvement, e.g. scanned academic papers? I imagine all the frontier labs have a solution for this due to the value of academic papers as a data source. Edit: Oh I see the paper abstract says this explicitly: "In production, DeepSeek-OCR can generate training data for LLMs/VLMs at a scale of 200k+ pages per day (a single A100-40G)". This is just part of the training data ingestion pipeline for their real models. Explains why the architecture is not using all of their latest tricks: it's already good enough for their use case and it's not the main focus.
- polytely 1y agoIf we get ocr working it makes it possible to store all human knowledge now stored in PDF's with way less resources https://annas-archive.org/blog/critical-window.html https://annas-archive.org/blog/critical-window.html
- rsp1984 1y agoCan someone ELI5 to me (someone who doesn't have the time to keep up with all the latest research) what this is and why it's a big deal? It's very hard to guess from the github and paper. For example, there is OCR in the title but the abstract and readme.md talk about context compression for LLMs, which I find confusing. Somebody care to explain the link and provide some high-level context?
- intalentive 1y agoSuppose you have an image with 1000 words in it, and suppose for simplicity that every word is 1 token. Then the image is “worth” 1000 tokens. But under the hood, the image will have to be transformed into features / embeddings before it can be decoded into text. Suppose that the image gets processed into 100 “image tokens”, which are subsequently decoded into 1000 “text tokens”. Now forget that we are even talking about images or OCR. If you look at just the decoding process, you find that we were able to compress the output into a 10x smaller representation. The implication for LLMs is that we don’t need 1000 tokens and 1000 token embeddings to produce the 1001st token, if we can figure out how to compress them into a 10x smaller latent representation first.
- rsp1984 1y agoExcellent, thanks. So basically this is saying: "our pixels-to-token encoding is so efficient (information density in a set of "image tokens" is much higher as compared to a set of text tokens), why even bother representing text as text?" Correct?
- intalentive 1y agoBasically. Some people are even saying, hey, if you encode text as an image then you don’t need tokenizers any more, and you get more expressivity from the graphic styling. Another takeaway is that you don’t need to pass a tensor of shape (batch_size, sequence_length, d_model) through your transformer. Not every token needs its own dedicated latent embedding. You can presumably get away with dividing sequence_length by a constant. This isn’t super ground breaking but it does reinforce the validity of a middle ground between recurrent models, where context is compressed into a single “memory token”, and transformers, where context is uncompressed. 1 < n/k < n
- dumpsterkid 1y agoI haven't fired this up yet to try but I've been evaluating & working with quite a few different VLMs from the small granite, qwen etc models up to the larger VLMs available to see if we can fully replace traditional OCR in our system but I've been disappointed so far - our system takes documents from customers and supplies them back normalized documents (i.e rasterized multi-page bitmaps) marked up as they've requested - however in our use case we need accurate coordinates of data down to the letter/word level and from my experience the positional outputs from these VLMs are either wildly inconsistent, completely hallucinated, or so vague that it doesn't allow us to target anything with any kind of accuracy or granularity. our solution so far has been to stick to using tesseract with good clean-up routines and then augmenting/fixing-up the output using the VLM OCR text where we don't have structured source document data available it could be that we just have a very niche use-case and it doesn't matter to most people, I'm sure if you just want a text dump or restructured markdown/html representation of documents these VLMs work well but the number of articles & comments I've seen claiming that these models have 'solved' OCR just seems counter to our experiences
- sampton 1y agoYou can train a cnn to find bounding boxes of text first. Then run VLM on each box.
- kamranjon 1y agoHave you tried moondream yet[1]? The moondream 3 preview model[2], according to the blogpost[3] appears to outperform many frontier models on VLM tasks and does so with a relatively small footprint. [1] https://moondream.ai/ https://moondream.ai/ [2] https://huggingface.co/moondream/moondream3-preview https://huggingface.co/moondream/moondream3-preview [3] https://moondream.ai/blog/moondream-3-preview https://moondream.ai/blog/moondream-3-preview
- jmpeax 1y agoYour customers don't have any handwritten text?
- shepardrtc 1y agoHow does this fair with the Vidore benchmark? https://huggingface.co/spaces/vidore/vidore-leaderboard https://huggingface.co/spaces/vidore/vidore-leaderboard
- schopra909 1y agoIt’s not clear to me what the bottleneck for OCR to “100%” work with LLMS is. In my work we do a lot of stuff with image understanding and captioning (not OCR). There object identification and description works great, since all the models are using a CLIP like visual backbone. But it falls apart when you ask about nuances like left/right or counting (reasoning kind of improves the latter but it’s too expensive to matter IMO). For our tasks, it’s clear that there’s more fundamental research that needs to be done on vision understanding to push past CLIP. That would really improve LLMs for our usecases. Curious if there’s something similar going on for OCR in the vision encoder that’s fundamentally holding it back.
- edtechdev 1y agoI tried this out on huggingface, and it has the same issue as every other multimodal AI OCR option (including MinerU, olmOCR, Gemini, ChatGPT, ...). It ignores pictures, charts, and other visual elements in a document, even though the models are pretty good at describing images and charts by themselves. What this means is that you can't use these tools yet to create fully accessible alternatives to PDFs.
- mediaman 1y agoI have a lot of success asking models such as Gemini to OCR the text, and then to describe any images on the document, including charts. I have it format the sections with XML-ish tags. This also works for tables.
- allanren 1y agoIt says the conversation can reduce size with large compression, which basic make the image blur but still contaim import information. This is indeed amazing. It's actually how human try to understand and remember things. BY VISUAL! And when memory fade out, the image are getting blurred. Not sure if those close source multimodal models are already using this method.
- dlowe24 1y agoThe only model that I found so far that extract table data with OCR is dots.ocr. Models that came after it have not done a good job. Interesting on testing this new model.
- joshstrange 1y ago> [2025/x/x] We release DeepSeek-OCR, a model to investigate the role of vision encoders from an LLM-centric viewpoint. So close but it should be 2025/X/XX as "X" = 10 in Roman Numerals /s Jokes aside, this is really neat and I'm looking forward to getting this running. For most OCR-type stuff I just use AWS Textract since I need it so rarely and that service does a decent job. I really like how well this model seems to extract images/figures as well from the original document.
- simonw 1y agoI figured out how to get this running on the NVIDIA Spark (ARM64, which makes PyTorch a little bit trickier than usual) by running Claude Code as root in a new Docker container and having it figure it out. Notes here: https://simonwillison.net/2025/Oct/20/deepseek-ocr-claude-code/ https://simonwillison.net/2025/Oct/20/deepseek-ocr-claude-co... Here's a result I got https://github.com/simonw/research/blob/main/deepseek-ocr-nvidia-spark/output_text/free_ocr/result.md https://github.com/simonw/research/blob/main/deepseek-ocr-nv... - against this image: https://static.simonwillison.net/static/2025/ft.jpeg https://static.simonwillison.net/static/2025/ft.jpeg
- jjcm 1y agoLooks like this did really solid, with the exception of the paragraph directly below the quote. It hallucinated some filler there and bridge it with the next column. Thanks for running the test quickly!
- djmips 1y agoBy my eye it just bridge. I didn't see any filler. It went from "Code is a language" - above the quote and then to "in a garden by name." which was the top of the next column but missing the chicken subject.
- throwaway314155 1y ago> by running Claude Code as root in a new Docker container How do you get the "as root" part of that to work? (sorry if it's explained in your article)
- simonw 1y agoRun it on a root account and do: IS_SANDBOX=1 claude --dangerously-skip-permissions
- throwaway314155 1y agoThanks!!
- prats226 1y agoTop 3 models on huggingface are all OCR models. Most automation projects involve documents where you need a model finetuned to understand all elements inside documents and provide grounding and confidence scores etc which is why these subset of models are gaining popularity
- dcl 1y agoIs there any 'small' OCR models around? Say I only care about reading serial numbers from photos in a manufacturing process, not whole document parsing. Using a 3B param model to do this seems like a bit of overkill...
- Thma 1y agoIs this only at the level of visual compression? For example, are there any applications in terms of understanding (being able to represent the actual meaning it stands for) and reasoning? Technically, it seems to have no connection with current reinforcement learning and other techniques. The model is quite small, yet there appears to be no explanation regarding its understanding capabilities. If it is merely for compression, what impact will it have on the current large models?
- h14h 1y agoThis feels like another one of those stairstep ML/AI advances that makes computers behave eerily more like humans. We tend to think in images rather than plaintext, and here we are discovering it's more efficient for a computer to do so as well.
- vladpowerman 1y agoThe compression framing is super interesting. It makes me wonder if there’s an equivalent notion for source code - like how much “information” or entropy a commit contains vs. boilerplate churn. I’ve been exploring Git activity analysis recently and ran into similar trade-offs: how do you tokenize real-world code and avoid counting noise?
- Karen667 11mo agoDeepSeek-OCR revolutionizes document processing by converting text into high-resolution images, achieving up to 20× compression while maintaining impressive accuracy. At a 10× compression ratio, it retains approximately 97% accuracy, and even at 20×, it maintains around 60% accuracy Tom's Hardware . This approach reduces token usage significantly, making it particularly beneficial for industries like finance, healthcare, and legal sectors. For a comprehensive guide on implementing and utilizing DeepSeek-OCR, you can check https://deepseeksguides.com/deepseek-ocr-guide/ https://deepseeksguides.com/deepseek-ocr-guide/
- giardini 11mo agoThis model (DeepSeek-OCR) aligns particularly well with what we know about written language and the human act of reading. The Visual Word Form Area (VWFA) on the left side of the brain is where the visual representation of words is transformed to something more meaningful to the organism. https://en.wikipedia.org/wiki/Visual_word_form_area https://en.wikipedia.org/wiki/Visual_word_form_area The DeepSeek-OCR encoding (rather than simple text encoding) appears analogous to what occurs in the VWFA. This model may not only be more powerful than text-based LLMs but may open the curtain of ignorance that has stymied our understanding of how language works and ergo how we think, what intelligence is precisely, etc. Kudos to the authors: Haoran Wei, Yaofeng Sun, Yukun Li. You may have tripped over the Rosetta Stone of intelligence itself! Bravo!
- sherlockxu 11mo agoSeems there’s still some confusion around what DeepSeek-OCR really does. Learn about the model, Contexts Optical Compression, and its impact on LLMs here: https://www.bentoml.com/blog/deepseek-ocr-contexts-optical-compression-explained https://www.bentoml.com/blog/deepseek-ocr-contexts-optical-c...