Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anerli
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
anerli
1y ago
Hey, awesome to hear! We are definitely open to contributions :) We plan to (very soon) enable mixing standard Playwright or other code in between Magnitude steps, which should enable doing exact assertions or anything else you want to do.
32.
▲
by
anerli
1y ago
Well yeah it's kind of ambiguous, it's just our way of saying that we're trying to use AI to make testing easier!
33.
▲
by
anerli
1y ago
You can run it against any URL, not just node projects! You'll still need a skeleton node project for the actual Magnitude tests, but you could configure some other public or staging URL as the target site.
34.
▲
by
anerli
1y ago
The planner can plan out multiple web actions at once, which Moondream can then execute in sequence on its own. So Moondream is never deciding how to execute more than one web action in a single prompt. What this really means for developers
35.
▲
by
anerli
1y ago
Originally we were actually thinking about doing exactly this and building agents for usability testing. However, we think that LLMs are much better suited for tackling well defined tasks rather than trying to emulate human nuance, so we pi
36.
▲
by
anerli
1y ago
Yeah good criticism for sure. We definitely want to keep this in mind as we continue to build. Some kind of accessibility tests which run in parallel with each visual test that are only allowed to use the accessibility tree could make it mu
37.
▲
by
anerli
1y ago
So this is a path that we definitely considered. However we think its a half-measure to generate actual Playwright code and just run that. Because if you do that, you still have a brittle test at the end of the day, and once it breaks you w
38.
▲
by
anerli
1y ago
So the architecture is built with determinism in mind. The plan-caching system is still a work in progress, but especially once fully implemented it should be very consistent. As long as your interface doesn't change (or changes in tri
39.
▲
by
anerli
1y ago
So the prompts that are sent to the planner vs executor are completely distinct. We allow complete customization of the planner LLM with all major providers (Anthropic, OpenAI, Google AI Studio, Google Vertex AI, AWS Bedrock, OpenAI compati
40.
▲
by
anerli
1y ago
Oh this is interesting. In our case we are being very specific about which types of prompts go where, so the planner essentially creates prompts that will be executed by Moondream, instead of trying to route prompts generally to the appropr
41.
▲
by
anerli
1y ago
So it's key to still have a big model that is devising the overall strategy for executing the test case. Moondream on its own is pretty limited and can't handle complex queries. The planner gives very specific instructions to Moon
42.
▲
Show HN: Magnitude – open-source, AI-native test framework for web apps
(github.com)
179 points
by
anerli
1y ago
|
44 comments
43.
▲
Can LLMs accurately evaluate their own confidence?
(github.com)
2 points
by
anerli
2y ago
|
2 comments
44.
▲
by
anerli
2y ago
I ran a simple experiment to try and understand whether self-rated answer confidence reflects the actual probability of the LLM generating that answer. I've always been skeptical of prompting techniques that ask the LLM to output a sco
45.
▲
Will AI Ruin Your Codebase?
(magnitude.run)
4 points
by
anerli
2y ago
|
0 comments
46.
▲
Why LLM Agents Today Don't Work
(langur.ai)
2 points
by
anerli
2y ago
|
0 comments
47.
▲
Show HN: Langur – consistent, observable LLM agents
(github.com)
1 points
by
anerli
2y ago
|
0 comments