4 ms·
CV tools have been better than most people at reading bad cursive since at least 2016. Is there a specific benchmark comparison between a GPT and any OCR app/a
by passwordoops 3y ago
CV tools have been better than most people at reading bad cursive since at least 2016.
Is there a specific benchmark comparison between a GPT and any OCR app/algo? Any cases where you've seen a free OCR app fail spectacularly where a GPT succeeded without question?
Language models are just that - language models. They are good at predicting what word should come next. Because it's language we impart capabilities way above and beyond their reality. Sure, they can incorporate CV data (among other inputs) that feed them the resulting text. Maybe when we develop something with a functional world model, we might be able to tackle the real physical world. Until then, problems like this, self driving, folding clothes, etc will keep failing at the edge cases, rendering them more trouble than they are worth.
And "just walk out" is a fun, interesting problem only in that it might give rise to interesting solutions and approaches for more pressing or relevant problems
- CamperBob2 3y agoTell me that the 'language model' behind this won't be capable of solving the just-walk-out problem before long: https://shot.3e.org/ss-20240310_145736.png https://shot.3e.org/ss-20240310_145736.png Go ahead, tell me. Life is short of opportunities for good sensible chuckles.
- xboxnolifes 3y agoA single good result does not prove accuracy.
- CamperBob2 3y agoYou could say the same for the army of 1000 mechanical Turks they were using before.
- passwordoops 2y agoAbsolutely. And what does that tell you? It tells me this is a solution without a problem.
- passwordoops 2y agoGo ahead, tell me this image and scenario wasn't already available in its training data (it was [1]). Tell me that if it got it wrong in the first pass (which it most likely did), the developers didn't explicitly tell it what the scenario was. I'm not saying the feat of LLMs is not impressive, they certainly are. Just don't tell me they have developed a "world model" and display understanding because they have not. They will always suffer the same issues that have plagued self-driving cars and autonomous robotics: they can only process what's in their training data, therefore they need to be trained on all scenarios that will ever exist to function outside of well-curated, well-defined closed systems. I would love a good chuckle too, unfortunately the total lack of critical thinking and understanding when it comes to these stochastic correlative black boxes leaves me greatly disappointed [1] https://www.reddit.com/r/Wellthatsucks/comments/j67atm/1_second_before/ https://www.reddit.com/r/Wellthatsucks/comments/j67atm/1_sec...