10 ms·
Karpathy on DeepSeek-OCR paper: Are pixels better inputs to LLMs than text?
https://xcancel.com/karpathy/status/1980397031542989305 https://xcancel.com/karpathy/status/1980397031542989305
- yunwal 1y ago> The more interesting part for me (esp as a computer vision at heart who is temporarily masquerading as a natural language person) is whether pixels are better inputs to LLMs than text. Whether text tokens are wasteful and just terrible, at the input. > Maybe it makes more sense that all inputs to LLMs should only ever be images. So, what, every time I want to ask an LLM a question I paint a picture? I mean at that point why not just say "all input to LLMs should be embeddings"?
- smegma2 1y agoNo? He’s talking about rendered text
- rhdunn 1y agoFrom the post he's referring to text input as well: > Maybe it makes more sense that all inputs to LLMs should only ever be images. Even if you happen to have pure text input, maybe you'd prefer to render it and then feed that in: Italicized emphasis mine. So he's suggesting that/wondering if the vision model should be the only input to the LLM and have that read the text. So there would be a rasterization step on the text input to generate the image. Thus, you don't need to draw a picture but generate a raster of the text to feed it to the vision model.
- fspeech 1y agoIf you can read your input on your screen your computer apparently knows how to convert your texts to images.
- CuriouslyC 1y agoAll inputs being embeddings can work if you have embedding like Matryoshka, the hard part is adaptively selecting the embedding size for a given datum.
- awesome_dude 1y agoI mean, text is, after all, highly stylised images It's trivial for text to be pasted in, and converted to pixels (that's what my, and every computer on the planet, does when showing me text)
- dang 1y agoRecent and related: Getting DeepSeek-OCR working on an Nvidia Spark via brute force with Claude Code - https://news.ycombinator.com/item?id=45646559 https://news.ycombinator.com/item?id=45646559 - Oct 2025 (43 comments) DeepSeek OCR - https://news.ycombinator.com/item?id=45640594 https://news.ycombinator.com/item?id=45640594 - Oct 2025 (238 comments)
- sabareesh 1y agoIt might be that our current tokenization is inefficient compared to how well image pipeline does. Language already does lot of compression but there might be even better way to represent it in latent space
- ACCount37 1y agoPeople in the industry know that tokenizers suck and there's room to do better. But actually doing it better? At scale? Now that's hard.
- typpilol 1y agoIt will require like 20x the compute
- _lyxd 1y agoWhy do you suppose this is a compute limited problem?
- ACCount37 1y agoIt's kind of a shortcut answer by now. Especially for anything that touches pretraining. "Why aren't we doing X?", where X is a thing that sounds sensible, seems like it would help, and does indeed help, and there's even a paper here proving that it helps. The answer is: check the paper, it says there on page 12 in a throwaway line that they used 3 times the compute for the new method than for the controls. And the gain was +4%. A lot of promising things are resource hogs, and there are too many better things to burn the GPU-hours on.
- typpilol 1y agoThanks. Also, saying it needs 20x compute is exactly that. It's something we could do eventually but not now
- ACCount37 1y ago
- hbarka 1y agoChinese writing is logographic. Could this be giving Chinese developers a better intuition for pixels as input rather than text?
- anabis 1y agoYeah, mapping chinese characters to linear UTF-8 space is throwing a lot of information away. Each language brings some ideas for text processing. sentencepiece inventor is Japanese, which doesn't have explicit word delimiters, for example.
- ComputerGuru 1y agoIt's not throwing any information away because it can be faithfully reconstructed (via an admittedly arduous process), therefore no entropy has been lost (if you consider the sum of both "input bytes" and "knowledge of utf-8 encoding/decoding").
- hobofan 1y agoYeah, that sounds quite interesting. I'm wondering whether there is a bigger gap in performance (= quality) between text-only<->vision OCR in Chinese language than in English. There is indeed a lot of semantic information contained in the signs that should help an LLM. E.g. there is a clear visual connection between 木 (wood/tree) and 林 (forest), while an LLM that purely has to draw a connection between "tree" and "forest" would have a much harder time seeing that connection independent of whether it's fed that as text or vision tokens.
- est 1y agoChinese text == Method of loci Many Chinese student have good memory to recall a particular paragraph, understand the meaning, but no idea how those words were pronouced.
- yandie 1y agoI can read Kanji (Japanese) and sometimes I will understand the sentence but can't pronounce it (Japanese Kanji rules are quite arbitrary). Your brain definitely handles information differently with Chinese letters
- varispeed 1y agoText is linear, whereas image is parallel. I mean when people often read they don't scan text from left to right (or different direction, depending on language), but rather read the text all at once or non-linearly. Like first lock on keywords and then read adjacent words to get meaning, often even skipping some filler sentences unconsciously. Sequential reading of text is very inefficient.
- sosodev 1y agoLLMs don't "read" text sequentially, right?
- olliepro 1y agoThe causal masking means future tokens don’t affect previous tokens embeddings as they evolve throughout the model, but all tokens a processed in parallel… so, yes and no. See this previous HN post (https://news.ycombinator.com/item?id=45644328 https://news.ycombinator.com/item?id=45644328) about how bidirectional encoders are similar to diffusion’s non-linear way of generating text. Vision transformers use bidirectional encoding b/c of the non-causal nature of image pixels.
- Merik 1y agoDidn’t anthropic show that the models engage in a form of planning such that it is predicting a possible future subsequent tokens that then affects prediction of the next token: https://transformer-circuits.pub/2025/attribution-graphs/biology.html#dives-poems https://transformer-circuits.pub/2025/attribution-graphs/bio...
- ACCount37 1y agoSure, an LLM can start "preparing" for token N+4 at token N. But that doesn't change that the token N can't "see" N+1. Causality is enforced in LLMs - past tokens can affect future tokens, but not the other way around.
- 1y ago
- cnxhk 1y agoThe paper is quite interesting but efficiency on OCR tasks does not mean it could be plugged into a general llm directly without performance loss. If you train a tokenizer only on OCR text you might be able to get better compression already.
- ianbutler 1y agohttps://arxiv.org/abs/2510.17800 https://arxiv.org/abs/2510.17800 (Glyph: Scaling Context Windows via Visual-Text Compression) You can also see this paper from the GLM team where they explicitly test this assumption to some pretty good results.
- scotty79 1y agoI couldn't imagine how rendering text tokens to images could bring any savings, but then I remembered esch token is converted into hundreds of floating point numbers before feeding it to neural network. So in a way it's already rendered into a multidimensional pixel (or hundreds of arbitrary 2-dimensional pixels). This papers shows that you don't need that many numbers to keep the accuracy and that using numbers that represent the text visually (which is pretty chaotic) is just as good as the way we currently do it.
- viraptor 1y agohttps://xcancel.com/karpathy/status/1980397031542989305 https://xcancel.com/karpathy/status/1980397031542989305
- kirubakaran 1y agoThanks. There are also these: - https://addons.mozilla.org/en-US/firefox/addon/toxcancel/ https://addons.mozilla.org/en-US/firefox/addon/toxcancel/ - https://chromewebstore.google.com/detail/xcancelcom-redirector/pgbegepjkmpcolalkpbakcbjlkodakie?hl=en https://chromewebstore.google.com/detail/xcancelcom-redirect...
- dang 1y agoThanks! Added to toptext also.
- dgfitz 1y ago[flagged]
- scotty79 1y agoIt's kind of beautiful that they can actually do that.
- nl 1y agoKapathy's points are correct (of course). One thing I like about text tokens though is that it learns some understanding of the text input method (particularly the QWERTY keyboard). "Hello" and "Hwllo" are closer in semantic space than you'd think because "w" and "e" are next to each other. This is much easier to see in hand coded spelling models, where you can get better results by including a "keybaord distance" metric along with a string distance metric.
- swyx 1y agoim particularly sympathetic to typo learning, which i think gets lost in the synthetic data discussion (mine here https://www.youtube.com/watch?v=yXPPcBlcF8U https://www.youtube.com/watch?v=yXPPcBlcF8U ) but i think in this case you can still generate typos in images and it'd be learnable. not a hard issue relevant to the OP
- harperlee 1y agoBut assuming that pixel input gets us to an AI capable of reading, they would presumably also be able to detect HWLLO as semantically close to HELLO (similarly to H3LL0, or badly handwritten text - although there would be some graphical structure in these latter examples to help). At the end of the day we are capable of identifying that... Might require some more training effort but the result would be more general.
- tcdent 1y ago"Kill the tokenizer" is such a wild proposition but is also founded in fundamentals. Tokenizing text is such a hack even though it works pretty well. The state-of-the-art comes out of the gate with an approximation for quantifying language that's wrong on so many levels. It's difficult to wrap my head around pixels being a more powerful representation of information, but someone's gotta come up with something other than tokenizer.
- dgently7 1y agoI consume all text as images when I read as a vision capable person so it kinda passes the evolution does it that way test and maybe we shouldn’t be that surprised that vision is a great input method? Actually thinking more about that I consume “text” as images and also as sounds… I kinda wonder if instead of render and ocr like this suggests we did tts and just encoded like the mp3 sample of the vocalization of the word if that would be less bytes than the rendered pixels version… probably depends on the resolution / sample rate.
- visarga 1y agoFunny, I habitually read while engaging TTS on same text. I have even made a Chrome extension for web reading, it highlights text and reads it, while keeping the current position in the viewport. I find using 2 modalities at the same time improves my concentration. TTS is sped up to 1.5x to match reading speed. Maybe it is just because I want to reduce visual strain. Since I consume a lot of text every day, it can be tiring.
- hiddencost 1y agoBack before transformers, or even LSTMs, we used to joke that image recognition was so far ahead of language modeling that we should just convert our text to PDF and run the pixels through a CNN.
- jimdavid 1y agoDid anyone check the token feature dimension? If we're talking about compression, "token length" is just one of the dimensions.
- shikon7 1y agoSeems we're now at a point of time when OCR is doing so well, that printing text out and letting computers literally read it is suggested to be superior to processing the endoded text directly.
- programmarchy 1y agoPDF is arguably a confusing format for LLMs to read.
- Legend2440 1y agoNeural networks have essentially solved perception. It doesn't matter what format your data comes in, as long as you have enough of it to learn the patterns.
- Sharlin 1y agoThe information density of a bitmap representation of text is just silly low compared to normal textual encodings, even compressed.
- orliesaurus 1y agoone of the MOST interesting aspects of the recent discussion on this topic is how it underscores our reliance on lossy abstractions when representing language for machines. Tokenization is one such abstraction, but it's not the only one.... using raw pixels or speech signals is a different kind of approximation. what excites me about experiments like this is not so much that we'll all be handing images to language models tomorrow, but that researchers are pressure testing the design assumptions of current architectures. Approaches that learn to align multiple modalities might reveal better latent structures or training regimes, and that could trickle back into more efficient text encoders without throwing away a century of orthography. BUT there’s also a rich vein to mine in scripts and languages that don’t segment neatly into words: alternative encodings might help models handle those better.
- bni 1y agoOf course PowerPoint is the best input to LLMs. They will come to that eventually.
- cat5e 1y agoYeah, I’ve seen great results with this approach.
- jtwaleson 1y agoIt's slides all the way down. Once models support this natively, it's a major threat to slides ai / gamma and the careers of product managers.
- brokencode 1y agoI'd actually prefer to communicate to ChatGPT via Microsoft Paint. Much more efficient than typing.
- saaaaaam 1y agoLeading scientists claim interpretative dance is the AI breakthrough the world has been waiting for!
- falcor84 1y agoIn all seriousness, I found those sorting dance videos to be really educationally effective (when coupled with going over the pseudocode) - e.g. https://youtu.be/3San3uKKHgg?si=09EQYJNIkRqvQgWG https://youtu.be/3San3uKKHgg?si=09EQYJNIkRqvQgWG
- amelius 1y agoClippy knew this all along.
- seydor 1y agowe re going to get closer and closer to removing all hand-engineered features of neural network architecture, and letting a giant all-to-all fully connected network collapse on its own to the appropriate architecture for the data, a true black box.
- justlikereddit 1y agoWhich is the Logical conclusion. If the neural network can distill a model out of complex input data. Especially when many model are frequently trained through data augmentation practices that actively degrade input to achieve generalisation abilities. Then why are we stuck wearing silk glove tokenizers?
- alexchamberlain 1y agoI'm probably one of the least educated software engineers on LLMs, so apologies if this is a very naive question. Has anyone done any research into just using words as the tokens rather than (if I understand it correctly) 2-3 characters? I understand there would be limitations with this approach, but maybe the models would be smaller overall?
- murkt 1y agoYou will need dictionaries with millions of tokens, which will make models much larger. Also, any word that has too low frequency to appear in the dictionary is now completely unknown to your model.
- mhuffman 1y agoAlong with the other commenter, the reason the dictionary would start getting so big is that words with a stem would have all its variations being different tokens (cat, cats, sit, sitting, etc). Also any out-of-dictionary words or combo words, eg. "cat bed" would not be able to be addressed.
- plaguuuuuu 1y agopresumably anyone tokenizing chinese characters, which are basically entire words.
- lyu07282 1y agoThe way modern tokenizers are constructed is by iteratively doing frequency analysis of arbitrary length sequences using a large corpus. So what you suggested is already the norm, tokens aren't n-grams. Words and any sequence really that is common enough will already be one token only, the less frequent a sequence is the more tokens it needs. That's the Byte-pair encoding algorithm: https://en.wikipedia.org/wiki/Byte-pair_encoding https://en.wikipedia.org/wiki/Byte-pair_encoding It's also not lossy compression at all, it's lossless compression if anything, unlike what some people have claimed here. Shocking comments here, what happened to HN? People are so clueless it reads like reddit wtf
- alexchamberlain 1y ago
- hunglee2 1y agoReally interesting analysis on the latest DeepSeek innovation. I’m tempted to connect it to the information density of logographic script, which DeepSeek engineers would all be natively fluent.
- pcwelder 1y agoThere are many unicode characters that look alike. There are also those zero width characters.
- foundonechar 1y ago[dead]
- js8 1y agoNot pixels, but percels. Pixels are points in the image, while a "percel" is unit of perceptual information. It might be a pixel with an associated sound, in a given moment of time. In case of humans, percels include other senses as well, and they can also be annotated with your own thoughts (i.e. percels can also include tokens or embeddings). Of course, NNs like LLM never process a percel in isolation, but always as a group of neighboring percels (aka context), with an initial focus on one of the percels.
- falcor84 1y agoI love this idea, but can't find anything about it. Is this a neologism you just coined? If so, is there any particular paper or work that led you to think about in those terms?
- js8 1y agoYes, I just coined the neologism. It was supposed to be partly sarcastic (why stay at pixels, why not just go fully multimodal and treat the missing channels as missing information?), I am kind of surprised why it got so upvoted. (IME, often my comments which I think are deep get ignored but silly things, where I was thinking "this is too much trolling or obvious", get upvoted; but don't take it the wrong way, I am flattered you like it.)
- throwaway-aws9 1y agoShould future attributions in white papers go to js8 from HN?
- SJMG 1y agoDeep things often, not always, take more attention to appreciate than the superficial. It's a precious resource people are seldom disposed to allocate a lot of when headline-surfing HN.
- jaredhansen 1y agoI think there's a decent chance you may have just created the ideal name for what will become one of the most important concepts ever. Bravo!
- a_bonobo 1y agoSomewhat related: There's this older paper from Lex Flagel and others where they transform DNA-based text, stuff we'd normally analyse via text files, into images and then train CNNs on the images. They managed to get the CNNs to re-predict population genetics measurements we normally get from the text-based DNA alignments. https://academic.oup.com/mbe/article/36/2/220/5229930 https://academic.oup.com/mbe/article/36/2/220/5229930
- deleted 1y ago[deleted]
- antirez 1y agoThis should be "pixels are (maybe) a better representation than the current representation of tokens". Which is very different. Text is surely more information dense than the image containing the same text, so the problem is finding the best representation of text. If each word is expanded to a very large embedding and you see pixels doing better, than the problem is in the representation and not in the text vs image.
- redbell 1y agoFor reference, here's the paper: https://github.com/deepseek-ai/DeepSeek-OCR/blob/main/DeepSeek_OCR_paper.pdf https://github.com/deepseek-ai/DeepSeek-OCR/blob/main/DeepSe...
- koushikn 1y agoIs it feasible that if we have a tokeniser that works on ELF (or PE/COFF) binaries, then we could have LLMs trained on existing binaries and have them generate binary code directly, skipping the need for programming languages?
- anon291 1y agoPossible but not precise depending on your use case. LLM compilers would suffer from the same sort of propensity to bugs as humans.
- trollbridge 1y agoI can attest that existing LLMs work surprisingly well for disassembly.
- kkukshtel 1y agoI've thought about this a lot, and it comes down ultimately to context size. Programming languages themselves are sort of a "compression technique" for assembly code. Current models even at the high end (1M context windows) do not have near enough workable context to be effective at writing even trivial programs in binary or assembly. For simple instructions sure, but for now the compression of languages (or DSLs) is a context efficiency.
- koushikn 1y agoWouldn't all binaries be in the training data, rather than the context? And output context could be in pieces, with something concatenating the pieces into a working binary? ChatGPT claims its possible, but not allowed due to OpenAI safety rules: https://chatgpt.com/share/68fb0a76-6bf8-800c-82f7-605ff9ca22e6 https://chatgpt.com/share/68fb0a76-6bf8-800c-82f7-605ff9ca22...
- bahmboo 1y agoNot criticizing per se but I just watched this recent (and great!) interview where he extols how special written language is. That was my take away at least. Still trying to wrap my head around this vision encoder approach. He’s way smarter than me! https://youtu.be/lXUZvyajciY https://youtu.be/lXUZvyajciY
- bob1029 1y agoI think the DCT is a compelling way to interact with spatial information when the channel is constrained. What works for jpeg can likely work elsewhere. The energy compaction properties of the DCT mean you get most of the important information in a few coefficients. A quantizer can zero out everything else. Zig zag scanned + RLE byte sequences could be a reasonable way to generate useful "tokens" from transformed image blocks. Take everything from jpeg encoder except for perhaps the entropy coding step. At some level you do need something approximating a token. BPE is very compelling for UTF8 sequences. It might be nearly the most ideal way to transform (compress) that kind of data. For images, audio and video, we need some kind of grain like that. Something to reorganize the problem and dramatically reduce the information rate to a point where it can be managed. Compression and entropy is at the heart of all of this. I think BPE is doing more heavy lifting than we are giving it credit for. I'd extend this thinking to techniques like MPEG for video. All frame types also use something like the DCT too. The P and B frames are basically the same ideas as the I frame (jpeg), the difference is they take the DCT of the residual between adjacent frames. This is where the compression gets to be insane with video. It's block transforms all the way down. An 8x8 DCT block for a channel of SDR content is 512 bits of raw information. After quantization and RLE (for typical quality settings), we can get this down to 50-100 bits of information. I feel like this is an extremely reasonable grain to work with.
- jacquesm 1y agoI can listen to music in my head. I don't think this is an extraordinary property but it is kind of neat. That hints at the fact that I somehow must have encoded this music. I can't imagine I'm storing the equivalent of a MIDI file, but I also can't imagine that I'm storing raw audio samples because there is just too much of it. It seems to work for vocals as well, not just short samples but entire works. Of course that's what I think, but there is a pretty good chance they're not 'entire', but it's enough that it isn't just excerpts and if I was a good enough musician I could replicate what I remember. Is there anybody that has a handle on how we store auditory content in our memories? Is it a higher level encoding or a lower level one? This capability is probably key in language development so it is not surprising that we should have the capability to encode (and replay) audio content, I'm just curious about how it works, what kind of accuracy is normally expected and how much of such storage we have. Another interesting thing is that it is possible to search through it fairly rapidly to match a fragment heard to one that I've heard and stored before.
- nottorp 1y agoThe text should be printed and a photo of the printed paper on a wooden table should be passed as input into the LLM.
- taneq 1y agoAll questions must now be posed to the Oracle through interpretive dance.
- anon291 1y agoI made exactly this point at the inaugural Portland AI tinkerers meetup. I had been messing with large document understanding. Converting PDF to text and then sending to gpt was too expensive. It was cheaper to just upload the image and ask it questions directly. And about as accurate. https://portland.aitinkerers.org/talks/rsvp_fGAlJQAvWUA https://portland.aitinkerers.org/talks/rsvp_fGAlJQAvWUA
- InkCanon 1y agoCould someone explain to me the difference? They both get turned to tensors of floats.
- 0x264 1y agoJavaScript code and Haskell code ultimately both get turned into instructions for a microprocessor, so there really isn't much of a difference between both.
- sd9 1y ago> more information compression (see paper) => shorter context windows, more efficiency It seems crazy to me that image inputs (of text) are smaller and more information dense than text - is that really true? Can somebody help my intuition?
- spiderfarmer 1y agoI absolutely think that it can, but it depends on what mean you associate with each pixel.
- vjerancrnjak 1y agoIt must be the tokenizer. Figuring out words from an image is harder (edges, shapes, letters, words, ...), yet internal representations are more efficient. I always found it strange that tokens can't just be symbols but instead there's an alphabet of 500k tokens, completely removing low level information from language (rhythm, syllables, etc.), side-effect being a simple edge case of 2 rs in strawberry, or no way to generate predefined rhyming patterns (without constrained sampling). There's an understandable reason for these big token dictionaries, but feels like a hack.
- krackers 1y agoSee this thread https://news.ycombinator.com/item?id=45640720 https://news.ycombinator.com/item?id=45640720 As I understood the responses, the benefit comes from making better use of the embedding space. BPE tokenization is basically like a fixed lookup table, whereas when you form "image tokens" you just throw each 16x16 patch into a neural-net and (handwave) out comes your embedding. From that, it should be fairly intuitive that since current text tokenization embedding vectors won't even form a subspace (it can only just be ~$VOCAB_SIZE points), image tokens have the capacity to be more information dense. And you might hope that the neural network can somehow make use of that extra capacity, as you're not encoding one subword at a time.
- rustyconover 1y agoYet again Hollywood is prescient. This post reminds me of the language of the aliens in Arrival. It seems like the OP would see that as a reasonable input to an LLM.
- cschmidt 1y agoThere is other research that works with pixels of text, such as this recent paper I saw at COLM 2025 https://arxiv.org/abs/2504.02122 https://arxiv.org/abs/2504.02122.
- deleted 1y ago[deleted]
- yalogin 1y agoI don’t quite follow. The way I see it, I hat the llm “reads” depends on the input modality. If the input is a human it will be in text form, has to be. If the input is through a camera then yes, even text will be camera frames and pixels, and that is how I expect the llms to process. So I would a vision llm would already be doing this.
- danans 1y ago> if the input is a human it will be in text form, has to be. Why can't it be a sequence of audio waveforms from human speech?
- bonoboTP 1y agoSometimes you want to be Unicode-precise, such as when checking if domain names are legit.
- daxfohl 1y agoI wouldn't think it would be good for coding assistants, or things where character precision is important. OTOH maybe the information implied by syntax coloring could make syntax patterns easier to recognize and internalize? Once internalized, perhaps it'd retain and use that syntax understanding on plaintext too if you fine tune it by gradually removing the color coding. Similar approaches have worked for improving their innate (no "thinking", no tool use) arithmetic accuracy.
- bigyikes 1y agoIt might be helpful for intuiting the structure of a program. Imagine if you had to read code all on a single line, with newlines represented with \n. I can get the feel of a piece of code just by looking at it. Even if you blurred the image, just the shape of the lines of code conveys a lot of information.
- daxfohl 1y agoTrue, but LLMs are already really good at that kind of thing. Even back in 2015, before transformers, here's a karpathy blog post showing how you could find specific neurons that tracked things like indent position, approx column location, long quotes, etc. https://karpathy.github.io/2015/05/21/rnn-effectiveness/ https://karpathy.github.io/2015/05/21/rnn-effectiveness/ That said, I do think algorithms and system designs are very visual. It's way harder to explain heaps and merge sorts and such from just text and code. Granted, it's 2025 now and modern LLMs seem to have internalized those types of concepts ~perfectly for a while now, so IDK if there's much to gain by changing approaches at that level anymore.
- CamperBob2 1y agoAnother example might be the way people used to show off their Wordle scores on Twitter when the game first came out. Just posting the gray, green and yellow squares by themselves, sans text, communicates a surprising amount of information about the player's guesses.
- daxfohl 1y agoIt seems like we're still pretty far away from that being viable, if chatgpt is any indication. Whenever it suggests "should I generate an image of that <class design, timeline, data model, etc>, it really helps visualize it!", the result is full of hallucinations.
- valine 1y agoImage generation and image input are two totally different things. This is about feeding text into LLMs as images, it has nothing to do with image generation.
- daxfohl 1y agoYeah but IIUC they're both just representations of embeddings in a latent space, translated from one format to another. So if the image interpretation of a text embedding is full of hallucinations, it's unlikely that the other direction works well either (again, IIUC). That said, I'll be interested to see what the DeepSeek model can do once they've trained it in the other direction. It'd be great to have it output architecture diagrams that actually correspond to what it says in the chat.
- ninetyninenine 1y agoeh, some part of the model will be translating those pixels into tokens. We're just moving the extra step into the blackbox.
- superconduct123 1y ago> more information compression (see paper) => shorter context windows, more efficiency I'll ask the dumb question here How is that possible? Wouldn't different sizes of text/fonts/rendering/spacing end up with way worse compression?
- qarl 1y agoHm. When I think to myself, I hear words stream across my inner mind. It's not pages of text. It's words.
- teleforce 1y agoPlease check this new promising GenAI architecture namely Discrete Distribution Networks or DDN that's recently posted here at HN [1]. The proposed DDN can even work as LLM based on binary strings instead of text [3]. [1] Show HN: I invented a new generative model and got accepted to ICLR (91 comments): https://news.ycombinator.com/item?id=45536694 https://news.ycombinator.com/item?id=45536694 [2] DDN+GPT for LLM: https://github.com/Discrete-Distribution-Networks/Discrete-Distribution-Networks.github.io/issues/1 https://github.com/Discrete-Distribution-Networks/Discrete-D...
- giardini 1y agoThis model (DeepSeek-OCR) ties particularly well with what we know about written language and the human act of reading. The Visual Word Form Area (VWFA) on the left side of the brain is where the visual representation of words is transformed to something more meaningful to the organism. https://en.wikipedia.org/wiki/Visual_word_form_area https://en.wikipedia.org/wiki/Visual_word_form_area The DeepSeek-OCR encoding (rather than simple text encoding) appears analogous to what occurs in the VWFA. This model may not only be more powerful than text-based LLMs but may open the curtain of ignorance that has stymied our understanding of how language works and ergo how we think, what intelligence is precisely, etc. Kudos to the authors: Haoran Wei, Yaofeng Sun, Yukun Li. You may have tripped over the Rosetta Stone of intelligence itself! Bravo!