3 ms·
DeepSeek interpreting screenshots and images I send it at fractions of what I pay Claude and ChatGPT, for me, is of far higher priority than supporting dictatio
by testbjjl 3mo ago
DeepSeek interpreting screenshots and images I send it at fractions of what I pay Claude and ChatGPT, for me, is of far higher priority than supporting dictation. There are workarounds for dictation but not image processing.
- corimaith 3mo agoOr you could just use a CNN...
- Jabrov 3mo agoTransformers are superior
- nullstyle 3mo agoWhich?
- bigmadshoe 3mo agoCNNs are not SoTA anymore when it comes to large models, and also are not used to provide interpretations of images as text, but rather to classify, do semantic segmentation, etc.
- tehjoker 3mo agoCan you say more about that? I haven't kept up.
- crypto420 3mo agoCNNs excel in vision tasks where you have limited compute, limited memory, limited data, and want something that works super well and quick. People usually don't hook CNNs up to a transformer to get language understanding either, you have to train bespoke CNNs for specific tasks ViTs excel where you're unbounded in compute + data and also want text understanding or have a conversation about an image
- bonoboTP 3mo agoThese are vibes. ViT has been shown to work fine on small data with proper hyperparam and most of what you mention is actually doable just fine with the other architecture as well.
- bonoboTP 3mo agoCNNs are fine when trained with a good recipe. There are very few good studies comparing them with proper hyperparam search and all the training tricks applied consistently. Transformers are good but ViT vs CNN is not some settled issue. Transformers are more hyped and more popular with the tech enthusiasts who just read forums and news, but if you need stuff done, CNNs are still great.
- bigmadshoe 3mo agoI agree, but since we're talking about imagine understanding with text output, clearly a CNN is unsuitable. My previous comment was overly reductive and CNNs can still be SoTA depending on your performance metrics. I spent the earlier part of my career training CNNs, and they are very pleasant to work with.
- bonoboTP 3mo agoYou can run a CNN and use the downsampled feature map the same way as patch tokens.
- famouswaffles 3mo ago>Transformers are more hyped and more popular with the tech enthusiasts who just read forums and news, but if you need stuff done, CNNs are still great. Vits are straight up more popular for ML research now, it's not just 'tech enthusiasts'.
- bonoboTP 3mo agoThere's a dearth of research properly comparing them.
- famouswaffles 3mo agoI'm talking about research pushing state of the art in computer vision. Vits have 100% become more popular than CNNs in most CV research.
- anthonypasq 3mo agojust use one of the various cheap gemini models
- carterschonwald 3mo agogemini models are also fantastic at understanding non spoken sounds
- jauntywundrkind 3mo agoI don't know what runs on my phone's Google Translate app, but whatever it is, they are doing an insult to their models by it being so bad. It's amazing at picking up sound if spoken directly into the unit, but if trying to hold any kind of conversation or listen to anything even a little bit far away, it falls completely apart, is good for basically nothing. This is obviously different than the models most people are discussing here, which are much bigger. But it's damaging the Gemini brand in general, by association, if nothing else.
- Royce-CMR 3mo agoI’ve long wondered if this was deliberate - only conversations where the participants are overtly using the translator get parsed.
- freedomben 3mo agoIndeed, Gemini really is incredible at image analysis. Yesterday I pointed it at some sloppy handwritten notes and asked it to add up the numbers in the right column, and it did it no problem. I've also used it to find out what TV show or actor is on screen, and various other things. It's quite impressive.
- winstonp 3mo agoGemini pretty clearly has the best underlying model, and the worst RL and post-training of the lot.
- 3mo ago
- segmondy 3mo agoYou can do that with smaller models at home. Gemma-4-E4B will run on a 12gb GPU, and supports audio, image, video input
- NooneAtAll3 3mo ago12GB GPU is a lot