13 ms·
Using GPT-4 Vision with Vimium to browse the web
- bnchrch 3y agoPersonally. This is what Im really excited about chatgpt for. Data has just become alot more free to access.
- burcs 3y agoThis is amazing, I feel like these vision models are going to make everything so much more accessible. Between the Be My Eyes app integration and now this, I'm really excited for how this transforms the web.
- ctoth 3y agoI agree, and I think we're a year or two away from a full end-to-end trained screen reader. The ground truth from existing systems would provide great training material. As a technical blind person, my only concern is the inherent loss of privacy while sharing stuff with the big models.
- supriyo-biswas 3y agoThere are open source models such as https://github.com/THUDM/CogVLM https://github.com/THUDM/CogVLM and https://github.com/haotian-liu/LLaVA https://github.com/haotian-liu/LLaVA.
- jackconsidine 3y agoLooks extremely cool. Trying to run it though, I get stuck at "Getting actions for the given objective..." (using the example on the repo)
- ishan0102 3y agoHuh weird, I'm getting that too. OpenAI has been having periodic outages today, think that might be why since it was working fine earlier.
- jechamt 3y agohttps://www.bleepingcomputer.com/news/security/openai-confirms-ddos-attacks-behind-ongoing-chatgpt-outages/ https://www.bleepingcomputer.com/news/security/openai-confir... News reports and their https://status.openai.com/incidents/21vl32gvx3hb https://status.openai.com/incidents/21vl32gvx3hb incident reports indicate they are mitigating / fighting off attacks recently
- ishan0102 3y agoHey! Creator here, thanks for sharing! Let me know if anyone has questions and feel free to contribute, I've left some potential next steps in the README.
- squeegmeister 3y agoHow does this differ from how ChatGPT currently browses the web?
- jgalentine007 3y agoVery cool use for Vimium, I like the approach!
- ishan0102 3y agoThank you!
- celeste_lan 3y agoOmg I also just released something pretty similar earlier today https://github.com/Jiayi-Pan/GPT-V-on-Web https://github.com/Jiayi-Pan/GPT-V-on-Web. But it received little attention.
- ishan0102 3y agoWoah looks great, not surprised that multiple people thought of this! Your prompt looks much better than mine, I'm not really taking advantage of any of the default Vimium shortcuts.
- jimmySixDOF 3y agoNice. I know Open Interpreter are trying to get Selenium automated to natural language control and quite a few other projects are also popping up on HN lately. The vimium approach is a lot lighter so looks promising. One way or another the as-published world wide web is turning into its own dynamic API overlay server. Ingest all the Sources!
- roland35 3y ago
- transistorfan 3y agoAt my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective
- Garlef 3y ago"Chinese Room Automation"
- deleted 3y ago[deleted]
- morkalork 3y agoKinda sci-fi, we're so close to a future where when/if original source code is lost, a mainframe runs in an emulator and the human operating it is also emulated.
- haswell 3y agoThe industry buzzword is "Robotic Process Automation", which as a category of products has been focused on using various forms of ML/AI to glue these things together in a common/structured way (in addition to good old fashioned screen scraping). Up this this point, these products have been quite brittle. The recent explosion of AI tech seems like quite a boon for this space.
- leovander 3y agoIn the OP's specific instance when would you reach out for a traditional ETL tool vs an RPA solution?
- teaearlgraycold 3y agoRPA is for data sources and destinations that are meant for human consumption and entry. So you’d use RPA to take an image of a table and enter every row into a web form.
- comment_ran 3y agoIt's so cool. I was wondering if we can make crawler tool much easier and better. It's more similar to the "human" way to interact with a website.
- imranq 3y agoIs the vision model directly reading the screen and therefore also reading the Vimeo tags? It might be more effective to export the DOM tags and the associated elements as a Json object that is fed into chatGPT without using the vision component
- dymk 3y ago> Currently the Vision API doesn't support JSON mode or function calling, so we have to rely on more primitive prompting methods.
- thekid314 3y agoI'm curious to see what it does when it sees a captcha.
- ishan0102 3y agoFrom OpenAI docs[1]: "For safety reasons, we have implemented a system to block the submission of CAPTCHAs." [1] https://platform.openai.com/docs/guides/vision https://platform.openai.com/docs/guides/vision
- circuit10 3y agoThere was an exploit that let you access the GPT-4 vision model months before release (and this restriction) and it could do this: https://media.discordapp.net/attachments/1020661972322230272/1086257844053098506/image.png https://media.discordapp.net/attachments/1020661972322230272...
- xur17 3y agoYeah, I've been feeding screenshots from selenium to the vision API, and when I trigger bot detection on a website, chatgpt refuses to process the image.
- NorwegianDude 3y agoIt does solve, or at least try to solve, captchas for me. It gets like half the characters correct, it's very bad at it.
- circuit10 3y agoMaybe try asking it to list what's in each square before giving the final answer
- snake_doc 3y agoAh, very similar to Adept’s[1] concept? Though, their product seems not yet ready. [1] https://www.adept.ai/ https://www.adept.ai/
- ishan0102 3y agoYep, took inspiration from them and a couple other startups
- QkPrsMizkYvt 3y agoWhat other startups did you use for inspiration?
- karmasimida 3y agoThis is precisely the demo I am thinking.
- jatins 3y agoIt's also a little insane to me that what Adept has been supposedly building for years with 300+ mil in funding can now be built in a day with Open AI APIs? I think Adept pivoted along the way but original concept was very similar to this.
- sunshadow 3y agoBut its too expensive to become practical with the OpenAI API. Also, demo is cool until you see the real-world webpages, then you'll realize that this only works less than %50 of webpages.
- famouswaffles 3y agoGPT-4V may be surprisingly robust here. Set of mark prompting(which is accomplished here with Vim) improves grounding by a silly high amount. https://som-gpt4v.github.io/ https://som-gpt4v.github.io/
- abrichr 3y ago
- snthpy 3y agoLooks cool. Unfortunately I expected this to enhance my Vimium experience but it looks like this is using Vimium to enhance GPT4, right?
- maccam912 3y agoI've been playing with a similar idea of screenshots and actions from GPT-4 Vision for browsing, but after trying and failing to overlay info in the screenshot, I ended up just getting the accessibility tree from playwright and sending that along as text so the model would know what options it had for interaction. In my case it seemed to work better, I see the creator is here and has a list of future ideas, maybe add this to the list if you think its a good idea?
- ishan0102 3y agoCool that’s a solid idea, I was trying to only use visual data but this could make the agent a lot more powerful, I’ll try this really soon
- manmal 3y agoProbably better to capture all the content and not just what fits on one screen. Most pages should fit as text (or HTML?) in the new extended token window.
- arbuge 3y agoBetter watch token costs. The per token costs are lower now but even so a full context load still costs almost $4.
- karmasimida 3y agoWe can create an autopilot for browser. It is going to incredibly difficult moving forward to distinguish bot traffic, if this is deployed at scale. The problem I see is this isn't going to be cheap or even affordable in short term.
- ishan0102 3y agoI think costs can come down if you finetune open source models like llava or cogvlm. This demo also cost about 6 cents so it's not insanely expensive either, especially with clever prompting.
- owenpalmer 3y agoThis will be fantastic for accessibility
- reqo 3y agoHow will tools like this affect web tracking or generally advertisements on the internet? Imagine you could have an agent browse the web for you and fetch exactly what you are seraching for without you seeing any ads/pop ups or being tracked along the way! Could be a great ”ad blocker”! Could it perhaps also make SEO useless and thus improve the quality of internet? But I wonder if it also could have negative effects such as the ads being “interweaved” into the fetch content somehow!
- famouswaffles 3y agoSince this is sending screenshots of pages to GPT, won't it see the ads as well?
- braindead_in 3y agoWhy not build a new browser with GPT baked in?
- reustle 3y agoCurious, how would that differ? Assuming it is just grabbing the rendered HTML DOM after each action, isn’t it nearly the same?
- lachlan_gray 3y agoI think vim is unintentionally a great “embodiment” for chatgpt. There’s nothing that can’t be done with a stream of text, and the internet is full of vimscript already I started a similar experiment if anyone else is thinking along the same lines :) https://github.com/LachlanGray/vim-agent https://github.com/LachlanGray/vim-agent
- gsuuon 3y agoThis is a neat idea!
- gvv 3y agoNice job! The horrors GPT-4 must endure to watch ads, truly inhumane
- FooBarWidget 3y agoMany Dutch companies pay salaries by 1. receiving payslips from the accountant, and then 2. manually initiating bank transfers to each employee for the amount in the corresponding payslip, and then 3. manually initiating a bank transfer to the tax authority to pay the withholded salary taxes. This is completely useless manual labor. There should be no reason for this to be a manual procedure. And yet it's almost impossible to automate this. The accountant portal either has no API, or it has an API but lets you download the data as PDF, and/or the API costs good money. The bank either has no API, or it requires you to sign up for a developer account as if you're going to publish a public app, when you're just looking to automate some internal procedures. So the easiest way to pay salaries and taxes is still to hire a person to do it manually. Hopefully one day that won't be necessary anymore. I wouldn't trust an AI to actually initiate the bank transfers, but maybe they can just prepare the transactions and then a person has to approve the submission.
- martinald 3y agoI don't think this really has much to do with AI. In the UK there are solutions like Pento now which do all this, including automating payments via open banking to the user and the tax authority and automatically filing tax filings: https://www.pento.io/la/payroll-software https://www.pento.io/la/payroll-software
- is_true 3y agoIn my country it's similar but for some data you have to upload to the government agency's site, I think it was earlier this year that they released a statement saying that people using software to perform actions on the website could get banned.
- nvm0n2 3y agoThat's just a bank problem. Certainly this isn't how payroll works for large companies. Banks usually let you upload XML files that define a set of SWIFT payments, this is how I do payroll even for a small company. The accountants supply the XML file too, presumably they have an app that generates it.
- abrichr 3y ago
- rizpaki 3y ago[flagged]
- rizpaki 3y ago[flagged]
- ranulo 3y agoThis could enable human language test automation scripts and could either improve my life as a QA engineer a lot or completely destroy it. Not sure yet.
- sunshadow 3y agoYou're good until this is cheaper than your salary.
- mackross 3y agoBeen playing with this through the ChatGPT interface for the past few weeks. Couple of tips. Update the css to get rid of the gradients and rounded corners. I found red with bold white text to be most consistent. Increase the font size. If two labels overlap, push them apart and add an arrow to the element. Send both images to the API, a version with the annotations added and a version without.
- bilekas 3y agoThis is actually pretty interesting.. I am thinking maybe it would be faster than writing up selenium tests themselves if we could just give a few instructions. I'm still going through the source, but really nice idea and great example of enriching the GPT with tools like vimium.
- startages 3y agoThere is just so much you can do with GPT-4 vision, I just hope it's more affordable.
- jonathanlb 3y agoHmm interesting. I'm curious what this means for accessibility and screen readers.
- e12e 3y agoIt's insane that this is now possible: https://github.com/ishan0102/vimGPT/blob/682b5e539541cd6d710e6723ef891f70506f64e9/vision.py#L35 https://github.com/ishan0102/vimGPT/blob/682b5e539541cd6d710... > "You need to choose which action to take to help a user do this task: {objective}. Your options are navigate, type, click, and done. Navigate should take you to the specified URL. Type and click take strings where if you want to click on an object, return the string with the yellow character sequence you want to click on, and to type just a string with the message you want to type. For clicks, please only respond with the 1-2 letter sequence in the yellow box, and if there are multiple valid options choose the one you think a user would select. For typing, please return a click to click on the box along with a type with the message to write. When the page seems satisfactory, return done as a key with no value. You must respond in JSON only with no other fluff or bad things will happen. The JSON keys must ONLY be one of navigate, type, or click. Do not return the JSON inside a code block."
- Maxion 3y agoThe speed at which this is moving at is mind boggling. This may become crazier than the dot.com boom.
- pms 3y agoUntil you realize that it doesn't work well with less popular videos (any items really), because "Large Language Models Struggle to Learn Long-Tail Knowledge" [1]. [1] https://proceedings.mlr.press/v202/kandpal23a.html https://proceedings.mlr.press/v202/kandpal23a.html
- heroprotagonist 3y agoExcept in this case, the knowledge is 'how to search the web for X" instead of 'an understanding or familiarity with X'.
- DalasNoin 3y agoI tried to use it, but unfortunately it often did not add the little annotations for the different options to the screen and it got stuck in a loop. This bot works by adding a two letter combination to each clickable option, but sometimes they don't show up. It managed to sign in to twitter ones, but really quickly I burned through the 100 images api limit. Maybe for a future version it only uses vision for difficult situations in which it gets stuck and otherwise uses the text based browser?
- nostrowski 3y agoThis will be in a future history book under a chapter titled "the beginning of the end"
- dangerwill 3y agoHow is this making your browsing experience any better? You still have to know what you want to do, and it is just faster to type Rick roll into youtube directly and click the links directly instead of having to type k, or vh, or whatever. You are just adding a useless chatgpt middleman between you and the browser that you likely spend all day in anyway and should be adept at navigating
- circuit10 3y agoIt's a proof of concept for how it could do more complicated tasks
- ternaus 3y agoLove the idea. It also shows that GPT-4V created a new angle in web scraping. I guess, this or similar code would be leveraged in many projects like: 1. Scrape XXX websites, say LinkedIn or Twitter use all types of methods in the DOM to prevent it, but fighting working well GPT-4V + OCR would be ultra hard. 2. Give me an analysis of what these XXX companies are doing. And this could be done for competitors, to understand the landscape of some industry, or even plainly to get news. Large-scale scrapping, not depending on the source code of the pages is a powerful infrastructural change.
- sebastiennight 3y agoIt took me a while to get what you meant, because... I'm not sure "XXX websites" usually means what you intended to convey here :)
- ternaus 3y agoI feel very innocent now, as it did not even cross my mind ;)
- doctorM 3y agoi think this is actively dangerous. well not yet. but getting there. i know - ai isn't meant to be sentient. but if it looks like a duck and quacks like a duck... how do i know that the comments here aren't done by dedicated hacker news ai bots? the potential danger could come from lack of supervision down the road. i didn't get much sleep last night so this is less coherent than it could be.
- mediumsmart 3y agothis is awesome and great news, nevermind that the AI found the wrong video in the demo https://www.youtube.com/watch?v=jRyX1tC2OS0 https://www.youtube.com/watch?v=jRyX1tC2OS0
- rpigab 3y agoThis is amazing that it's possible and works, but I wonder if the electricity cost is sustainable in the long run. For handicapped people who depend on tools like this for accessibility, it's justified, but I wouldn't use it myself if it uses too much power. I'm sure OpenAI and friends love operating at a loss until everyone uses their products, then enshittify or raise prices, like Netflix, Microsoft, Google, etc., but CO2 emissions can't be easily reversed. I'd be glad to listen to other points of view though, maybe everything we do on computers is already bad for the environment anyway and comparing which one pollutes more is vain, idk.
- silentguy 3y agoI think this can be extended to desktop as well. There are programs that act like vimium for your desktop (win-vind, etc.). I don't have the openai API key to try it but I wish someone gave it a try (in obviously an isolated environment).
- silentguy 3y agoUsually there are a lot of comments about how text is the best interface and it's making a comeback in the LLMs but in this case picture is the better medium since parsing the webpage js would prove too difficult. I think a screenshot of a webpage has a smaller footprint than the raw payloads (js, assets, etc.).