5 ms·
Partially like being an ass for fun. My first question is what does a GPT have to do with solving this problem? Honest question, because I don't see how a lang
by passwordoops 3y ago
Partially like being an ass for fun.
My first question is what does a GPT have to do with solving this problem? Honest question, because I don't see how a language model helps solve this, so if your do have a clear explanation it will be appreciated. I get they parse image data but have they demonstrated better accuracy for this use case (grocery items/food) than what was available when Amazon started their campaign?
There's also the fact that, yes more data gives better results... To a point. At this stage, I'm inclined to think we've hit that point when it comes to deep learning models' ability for object recognition and we're in a plateau until a rational new approach (i.e. not deep learning) allows for another leap.
My last is more socialogically rooted. Namely, it's pretty clear that since sama said he'd need $7 Trillion, AI peddlers have jumped the shark and we're seeing a slow motion deflation of the major claims, even concerning LLMs (don't get me wrong, very impressive tech, but no where near the heights and capabilities of the hucksters). For that reason I'm very inclined to lean towards skepticism when people claim they've found a solution by incremental improvements to a problem that's proven to be very, very difficult for a computer to handle.
I think we're nearing a general plateau in deep learning and need the next foundational approach or what we could find ourselves in another deep AI winter.
Any reason you think my skepticism is unfounded and why 5 years?
- CamperBob2 3y agoMy first question is what does a GPT have to do with solving this problem? Have you seen how effective transformer-based networks are at image recognition? They are literally better at reading bad cursive handwriting than I am, at this point. So yes, they can be pressed into service recognizing products being taken off of shelves and placed into carts or bags. "Just walk out" is a difficult problem but very much worth solving. We now have tools to attack it that nobody was even dreaming of when Amazon designed their current approach. Tools with near science-fiction levels of capability.
- passwordoops 3y agoCV tools have been better than most people at reading bad cursive since at least 2016. Is there a specific benchmark comparison between a GPT and any OCR app/algo? Any cases where you've seen a free OCR app fail spectacularly where a GPT succeeded without question? Language models are just that - language models. They are good at predicting what word should come next. Because it's language we impart capabilities way above and beyond their reality. Sure, they can incorporate CV data (among other inputs) that feed them the resulting text. Maybe when we develop something with a functional world model, we might be able to tackle the real physical world. Until then, problems like this, self driving, folding clothes, etc will keep failing at the edge cases, rendering them more trouble than they are worth. And "just walk out" is a fun, interesting problem only in that it might give rise to interesting solutions and approaches for more pressing or relevant problems
- CamperBob2 3y agoTell me that the 'language model' behind this won't be capable of solving the just-walk-out problem before long: https://shot.3e.org/ss-20240310_145736.png https://shot.3e.org/ss-20240310_145736.png Go ahead, tell me. Life is short of opportunities for good sensible chuckles.
- xboxnolifes 3y agoA single good result does not prove accuracy.
- CamperBob2 3y agoYou could say the same for the army of 1000 mechanical Turks they were using before.
- passwordoops 2y agoAbsolutely. And what does that tell you? It tells me this is a solution without a problem.
- passwordoops 2y agoGo ahead, tell me this image and scenario wasn't already available in its training data (it was [1]). Tell me that if it got it wrong in the first pass (which it most likely did), the developers didn't explicitly tell it what the scenario was. I'm not saying the feat of LLMs is not impressive, they certainly are. Just don't tell me they have developed a "world model" and display understanding because they have not. They will always suffer the same issues that have plagued self-driving cars and autonomous robotics: they can only process what's in their training data, therefore they need to be trained on all scenarios that will ever exist to function outside of well-curated, well-defined closed systems. I would love a good chuckle too, unfortunately the total lack of critical thinking and understanding when it comes to these stochastic correlative black boxes leaves me greatly disappointed [1] https://www.reddit.com/r/Wellthatsucks/comments/j67atm/1_second_before/ https://www.reddit.com/r/Wellthatsucks/comments/j67atm/1_sec...