Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
edunteman
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
edunteman
1y ago
you mean like https://www.anthropic.com/news/prompt-caching or just saving LLM chat message history? If the latter, saving chat history is useless without some snapshot of the environment in which it was performed. Mus
32.
▲
by
edunteman
1y ago
Thank you!
33.
▲
by
edunteman
1y ago
This is very similar to what Voyager did https://arxiv.org/abs/2305.16291 Their implementation uses actual code, JS scripts in their case, as the stored trajectories, which has the neat feature of parameterization buil
34.
▲
by
edunteman
1y ago
another currently unanswered question in Muscle Mem is how to more cleanly express the targeting of named entities. Currently, a user would have to explicitly have their @engine.tool call take an element ID as an argument to a click_element
35.
▲
by
edunteman
1y ago
Our experience working with A11y apis like above is that data is frequently missing, and the APIs can be shockingly slow to read from. The highest performing agents in WindowsArena use a mixture of A11y and yolo-like grounding models such a
36.
▲
by
edunteman
1y ago
Correct. In the RPA world, if there's an API available, or even a sqlite server you can tap into, you absolutely should go directly to the source. Emulating human mouse and keyboard is an absolute last resort for getting data across th
37.
▲
by
edunteman
1y ago
I struggled with this one for a while in the design, and didn't want to be hasty in making any decisions that lock us into a direction. I definitely want to support sub-trajectories. In fact, I believe an absolutely killer feature for
38.
▲
by
edunteman
1y ago
Thanks for sharing your experience! I'd love to chat about what you did to make this work, if I may use it to inform the design of this system. I'm at erik [at] pig.dev To clarify, the use of CLIP embeddings in the CUA example is
39.
▲
by
edunteman
1y ago
I agree, Cache Validation is the singular concern of Muscle Mem. If you boil it down, for a generic enough task and environment, the engine is just a database of previous environments and a user-provided filter function for cache validation
40.
▲
by
edunteman
1y ago
In a way, a Muscle Mem trajectory is just a new "meta tool" that combines sequential use of other tools, with parameters that flow through it all. One form factor I toyed with was the idea of a dynamically evolving list of tool sp
41.
▲
by
edunteman
1y ago
I believe explicit trajectories for learned behavior are significantly easier for humans to grok and debug, in contrast to reinforcement learning methods like deep Q-learning, so avoiding the use of models is ideal, but I imagine they'
42.
▲
Show HN: Muscle-Mem, a behavior cache for AI agents
(github.com)
226 points
by
edunteman
1y ago
|
51 comments
43.
▲
by
edunteman
2y ago
This was one of my favorite demos I've seen live at YC - congrats on the launch!
44.
▲
by
edunteman
3y ago
we have no plans for image or other modalities. I'd like to keep it just text-to-text so it can be as sharp of a tool as possible
45.
▲
by
edunteman
3y ago
interesting idea! yeah next big challenge for us is hallucination in null fields (IE: if you ask for a "name" from text that doesn't have a name you usually get "John") so need to add more sampling heuristics to dou
46.
▲
by
edunteman
3y ago
openai json mode will ensure JSON, but not strictly the right JSON schema to match whatever object you're expecting to return. Keys could be missing, fields could be improperly typed. We found gpt-4 reasoning capabilities could, most o
47.
▲
by
edunteman
3y ago
hey, a vouch is a vouch!
48.
▲
Show HN: Anything To JSON – a language model for structured extraction
(anythingtojson.com)
18 points
by
edunteman
3y ago
|
8 comments
49.
▲
by
edunteman
3y ago
Thanks everyone for a great Show HN! This turned out much larger than expected, thanks for all the comments and github stars. We had a fun time with it. Lots of takeaways, blogs to read, things to implement, issues to address. On it!
50.
▲
by
edunteman
3y ago
Beautiful, this does the trick!! Thanks for the tip. @ai() def stub() -> int: """docstring""" ... # (use ... instead of "pass" in the function body)
51.
▲
by
edunteman
3y ago
valid json, yes, but not a specific json schema (yet, who knows, maybe they ship schema support, I'm surprised they haven't)
52.
▲
by
edunteman
3y ago
it seems this is in the context of "extraction" where all of the data is already present in the input text, and all that's needed is the reformatting. This is something we've been wrestling with (even today): is the role
53.
▲
by
edunteman
3y ago
the king has arrived! Instructor is a clever API on this, clean
54.
▲
by
edunteman
3y ago
feel free to PR the readme if you feel it was misleading
55.
▲
by
edunteman
3y ago
absolutely! My DMs are open at @erikdunteman on twitter, or erik at banana dot dev
56.
▲
by
edunteman
3y ago
Yeah the pyright doesn't like the annotated return type not being honored by the empty stub function. I wonder if there's a way to trick it. For your suggestion, the decorator would still be required to overload the function execu
57.
▲
by
edunteman
3y ago
ah shoot, yes I meant vLLM, sorry for the confusion, lots of comments to reply to :)
58.
▲
by
edunteman
3y ago
one idea we're cooking is to offer a proxy with a hosted reformatting model on-board, to rewrite payloads on their way back in the case of type parse failure. fructose, the clientside sdk, would be optional
59.
▲
by
edunteman
3y ago
Thanks for asking, and I'd agree. I'd give the same answer as the folks asking about instructor: we built this in a week and are sharing it early, this package API happens to have landed on what Marvin is doing, we're likely
60.
▲
by
edunteman
3y ago
Thanks!
More ›