3 ms·
> Just because we rely on vision to interface with computer software doesn't mean it's optimal for AI models. It's optimal for beings that have general purpose
by raducu 1y ago
> Just because we rely on vision to interface with computer software doesn't mean it's optimal for AI models.
It's optimal for beings that have general purpose inteligence.
> would severely limit your capabilities as well because you'd have to dedicate more effort towards handling basic editing tasks
Yes, but humans will eventually get used to it and internalize the keyboard, the domain language, idioms and so on and their context gets pushed to long term knowledge overnight and thei short term context gets cleaned up and they get bettet and better at the job, day by day.
AI starts very strong but stays at that level forever.
When faced with a really hard problem, day after day the human will remember what he tried yesterday and parts of that problem will become easier and easier for the human, not so for the AI, if it can't solve a problem today, running it for days and days produces diminishing returns.
That's the General part of human intelligence -- over time it can aquire new skills it did not have yesterday, LLMs can't do that -- there is no byproduct of them getting better/aquiring new skills as a result of their practicing a problem.
- daxfohl 1y agoRight, and also the ability to know when it's stuck. It should be able to take a problem, work on it for a few hours, and if it decides it's not making progress it should be able to ping back asynchronously, "Hey I've broken the problem down into A, B, C, and D, and I finished A and B, but C seems like it's going to take a while and I wanted to make sure this is the right approach. Do you have time to chat?" Or similarly, I should be able to ask for a status update and get this answer back.
- ctoth 1y ago> It's optimal for beings that have general purpose inteligence [Sic]. Hi. I'm blind. I would like to think I have general-purpose intelligence thanks. And I can state that interfacing with vision would, in fact, be suboptimal for me. The visual cortex is literally unformed. Yet somehow I can perform symbolic manipulations. Converse with people. Write code. Get frustrated with strangers on the Internet. Perhaps there are other "optimal" ways that "intelligent" systems can use to interface with computers? I don't know, maybe the accessibility APIs we have built? Maybe MCP? Maybe any number of things? Data structures specifically optimized for the purpose and exchanged directly between vastly-more-complex intelligences than ourselves? Do you really think that clicking buttons through a GUI is the one true optimal way to use a computer?
- Jensson 1y ago> Do you really think that clicking buttons through a GUI is the one true optimal way to use a computer? There are some tasks you can't do without vision, but I agree it is dumb to say general intelligence requires vision, vision is just an information source it isn't about intelligence. Blind people can be excellent software engineers etc they can do most white collar work just as well as anyone else since most tasks doesn't require visual processing, text processing works well enough.
- jermaustin1 1y ago> There are some tasks you can't do without vision... I can't think of anything where you require vision that having a tool (a sighted person) you protocol with (speak) wouldn't suffice. So why aren't we giving AI the same "benefit" of using any tool/protocol it needs to complete something.
- ctoth 1y ago> I can't think of anything where you require vision that having a tool (a sighted person) you protocol with (speak) wouldn't suffice. Okay, are you volunteering to be the guide passenger while I drive?
- jermaustin1 1y agoThank you for making my point: We have created a tool called "full self driving" cars already. This is a tool that humans can use, just like we have MCPs a tool for AI to use. All I'm trying to say, is AGIs should be allowed to use tools that fit their intelligence the same way that we do. I'm not saying AIs are AGIs, I'm just saying that the requirement that they use a mouse and keyboard is a very weird requirement like saying People who can't use a mouse and keyboard (amputees, etc.) aren't "Generally" intelligent. Or people who can't see the computer screen.
- daxfohl 1y agoOf course not. The visual part is window dressing on the argument. The real point is, before declaring AGI, I think the way we interact with these agents needs to be more like human to human interaction. Right now, agents generally accept a command, figure out which from a small number of MCPs that have been precoded for it to use, do that thing you wanted right or wrong, the end. If it does the right thing, huge confirmation bias that it's AGI. Maybe the MCP did most of the real work. If it doesn't, well, blame the prompt or maybe blame the MCPs are lacking good descriptions or something. To get a solid read on AGI, we need to be grading them in comparison to a remote coworker. That they necessarily see a GUI is not required. But what is required is that they have access to all the things a human would, and don't require any special tools that limit their search space to a level below what a human coworker would have. If it's possible for a human coworker to do their whole job via console access, sure, that's fine too. I only say GUI because I think it'd actually be the easiest option, and fairly straightforward for these agents. Image processing is largely solved, whereas figuring out how to do everything your job requires via console is likely a mess. And like I said, "using the computer", whether via GUI or screen reader or whatever else, isn't going to be the hard part. The hard part is, now that they have this very abstract capability and astronomically larger search space, it changes the way we interact with them. We send them email. We ping them on Slack. We don't build special baby mittens MCPs and such for them and they have to enter the human world and prove that they can handle it as a human would. Then I would say we're getting closer to AGI. But as long as we're building special tools and limiting their search space to that limited scope, to me it feels like we're still a long way off.