6 ms·
You've raised valid points about the cost and efficiency of our approach, which aims to make the LLM function as closely as possible to a human user. We chose t
by keremyilmaz 3y ago
You've raised valid points about the cost and efficiency of our approach, which aims to make the LLM function as closely as possible to a human user. We chose this approach primarily for its compatibility with various websites, as it aligns closely with a website's intended audience, which is typically human.
Addressing complex website interactions is a key advantage of this approach. For instance, in the process of generating an auto insurance quote, the sequence of questions and their specifics can vary greatly depending on prior responses. A simple example is the choice of a foreign versus a California driver's license. Selecting a foreign license triggers additional queries about the country of issuance and expiry date, illustrating the complexity and branching nature of such web interactions.
However, we recognize the concerns about cost and are actively working on strategies to reduce it:
- Optimizing the context provided to the LLM
- Implementing caching mechanisms for certain repeated actions and only use LLMs when there's a problem
- Anticipating advancements in LLM efficiency and cost-effectiveness, with the hope of eventually finetuning our own models for greater efficiency
- dtnewman 3y agoI like this approach. Just as an example, if I'm getting a car insurance quote, I'd rather pay $1 to have the tool fill out the forms for me and be 90% that it filled them out correctly rather than pay $0.01 and only be 70% sure it did it correctly. And there are plenty of use cases like that.
- amne 3y agoisn't that crazy rabbit thingy supposed to do just that? I hope you pre-ordered. I hear they're in great demand.
- daniel_iversen 3y agohttps://www.rabbit.tech/research https://www.rabbit.tech/research
- Kerbonut 3y agoYou would still be willing to pay $1 if it got it wrong 10% of the time, or if it got 10% of the information wrong every time?
- dtnewman 3y agoIt really depends on the use case.
- dinobones 3y agoThere are two things here: 1) Using the LLM to find elements/selectors in HTML 2) Use LLMs to fill out logical/likely/meaningful answers to things I highly recommend you decouple these 2 efforts. While you gave a good example of "insurance quote step by step webapp", the vast majority of web scraping efforts are much more mundane. Additionally, even in this instance, the selector brain/intelligence brain don't need to be coupled. For example: Selector brain: "Find/click the button for foreign drivers license." Selector brain: "Find the country of origin field." Selector brain: "Find the expiry date field." LLM-intelligence brain: "Use values from prompt to fill out the country of origin and expiry date fields." Not-LLM intelligence brain: Inputs values from a JSON object of documentSelector=>value.
- suchintan 3y agoInteresting. We've decoupled navigation and extraction for specifically this reason, but I suppose decoupling selector with input could let us use cheaper smaller LLMs to "select" and answer We've been approaching it a little bit differently. We think larger more capable models would actually immediately improve the performance of Skyvern. For example, if you run it with LLaVa, the performance significantly degrades, likely because of the coupling But since we use GPT-4V, and it's rumoured to be a MoE model, I wonder if there's implicit decoupling going on. I'm gonna spend some more time thinking about this
- bravura 3y agoI still think you're missing the point. The idea is that you should use vision APIs and LLMs to build traditional browser automation using a DSL or Python. I don't want to use vision and LLMs for every page. I just want to use vision and LLMs to figure out what elements need to be clicked once. Or maybe every time the site changes the frontend.
- suchintan 3y agoThis is a great point. This is something already on our roadmap. We call it "prompt caching", but I realize writing this that it's a terrible name. Will update! (https://github.com/Skyvern-AI/Skyvern?tab=readme-ov-file#feature-roadmap https://github.com/Skyvern-AI/Skyvern?tab=readme-ov-file#fea...) Thank you for this feedback