Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
suchintan
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
suchintan
1y ago
That would be amazing!!
32.
▲
by
suchintan
1y ago
Very cool. The benchmark can be found here if you want to take a look at it: https://github.com/Halluminate/WebBench
33.
▲
by
suchintan
1y ago
We definitely plan to expand it. I want to get to ~10,000 for a reasonable benchmark. 15 blew my mind -- it's too easy to overfit that dataset
34.
▲
Web Bench: a new way to compare AI browser agents
(blog.skyvern.com)
34 points
by
suchintan
1y ago
|
9 comments
35.
▲
by
suchintan
2y ago
Definitely. It gives you a lot more control on which models you want powering these browser agents, which is the important part
36.
▲
Show HN: MCP Server to let agents control the browser
(github.com)
14 points
by
suchintan
2y ago
|
3 comments
37.
▲
by
suchintan
2y ago
Hey HN, we were playing around with MCPs over the weekend and thought it would be cool to build an MCP that lets Claude / Cursor / Windsurf control your browser: https://github.com/Skyvern-AI/skyvern/tree
38.
▲
by
suchintan
2y ago
I've seen people use semantic versioning to version APIs and SDKs https://semver.org/ We at Skyvern are still doing patch versions only
39.
▲
Raw Dump of Skyvern's navigations and decisions while evaluating WebVoyager
(eval.skyvern.com)
2 points
by
suchintan
2y ago
|
0 comments
40.
▲
by
suchintan
2y ago
I wonder if more companies should open source their eval model outputs alongside the eval results We tried doing that here at Skyvern (eval.skyvern.com)
41.
▲
by
suchintan
2y ago
Definitely need a newer benchmark. I couldn't find where browser-use published their run results (expected to see it here https://github.com/browser-use/eval ) We went ahead and published our full run at https:&#x
42.
▲
by
suchintan
2y ago
I'd love to chat to see how we can help! Here's my email: suchintan@skyvern.com We're working on 2 major improvements that will get cost down at scale: 1. We're building a code generation layer under the hood that will s
43.
▲
by
suchintan
2y ago
This is a great point -- the example we chose was meant to be a consumer example that we could relate with.. however a similar example exists for the enterprise which may be more interesting Let's say that you are a parts procurement s
44.
▲
Skyvern Browser Agent 2.0: How We Reached State of the Art in Evals
(blog.skyvern.com)
49 points
by
suchintan
2y ago
|
31 comments
45.
▲
by
suchintan
2y ago
It creates a feedback loop for the LLM to help it correct hallucinations or misunderstandings about how it's planned actions actually played out!
46.
▲
Show HN: Skyvern 2.0 – open-source AI Browser Agent scoring 85.8% on WebVoyager
(eval.skyvern.com)
9 points
by
suchintan
2y ago
|
3 comments
47.
▲
by
suchintan
2y ago
This is very interesting, thank you for sharing it What are your opinions on lists like this? https://github.com/mmccaff/PlacesToPostYourStartup
48.
▲
by
suchintan
2y ago
Yep. Makes a lot of sense. Your side project is super noble -- thank you for that
49.
▲
by
suchintan
2y ago
This is really cool. Do you have any interest in helping people auto apply to them? We can help you set it up with a really simple API call via Skyvern ( https://github.com/Skyvern-AI/skyvern )
50.
▲
A practical example of YC's advice to launch early and often
(blog.skyvern.com)
2 points
by
suchintan
2y ago
|
1 comments
51.
▲
Using Cedana to cut time to first token by almost 50%
(docs.cedana.ai)
1 points
by
suchintan
2y ago
|
0 comments
52.
▲
by
suchintan
2y ago
we started this last year. Back then, autogpt was more of a prototypical framework -- I'm sure it's improved dramatically since then. We ran into too many issues while developing with it (how do you run tasks effectively? How do y
53.
▲
by
suchintan
2y ago
Yep. This is totally fair feedback -- we're still a super early product and haven't had a chance to optimize the phone experience.. largely because it's tough to see the magic from the phone We'll improve it soon!
54.
▲
by
suchintan
2y ago
Yes -- in theory. You'd need to use our workflows feature to get that set up and chain a few tasks together to collect that information!
55.
▲
by
suchintan
2y ago
Great to see you on here Anton! Just curious, how do you differentiate from other open source competitors like n8n and Active pieces?
56.
▲
by
suchintan
2y ago
You're making some really good points here 1/ the current prompt + payload structure is definitely on the complicated end of the spectrum, but we've found that we can use an LLM to help generate this payload for our users The
57.
▲
by
suchintan
2y ago
1. Yes absolutely. But the issue is a little bit more nuanced than that. Websites without APIs don't have them for one of two reasons: (1) They want to protect their data (LinkedIn) or (2) can't be bothered to make an API (boutiqu
58.
▲
by
suchintan
2y ago
Depends on the scope of the changes. What did you have in mind?
59.
▲
by
suchintan
2y ago
We have an open issue for this right now -- we would LOVE some contributions here. The biggest problem until Llama 3.2 came out was that most (good) open source llms were text-only, and Skyvern needs vision to perform well This isn't t
60.
▲
by
suchintan
2y ago
Give it a try! It's very capable of doing simple tasks like logging in and clicking around. You'll need to prompt assertions like "Complete if..." and "Terminate if..."
More ›