3 ms·
Wanted to introduce NOVA - agent with dynamic management of prompts, functions and document access. Key features - Create and manage multiple prompts, enabling
by samueltates 3y ago
Wanted to introduce NOVA - agent with dynamic management of prompts, functions and document access.
Key features
- Create and manage multiple prompts, enabling, disabling, removing,
- Persistent memory, essentially rolling summary windows plus traversable index
- Turn on 'auto-gpt like' functionality like write, read, query
- Embed pdfs and google docs using the amazing LlamaIndex
- Create new channels / configurations, and share those with other people
Intro video here : https://www.youtube.com/watch?v=FB_g_8ofSlE https://www.youtube.com/watch?v=FB_g_8ofSlE
You can essentially craft a set of prompts and document connections, interact with that agent over time, with persistent memory, and shared docs. You can make that public, so being able to share embeddings was the thinking, but also shared memory.
The stack
It's a react front end, to a python cloud server. Hope is to make that server client agnostic, delivered via API. So people could compose agents, behaviours, permissions in NOVA, and deliver agent behaviours into their own flows, thinking AMS (agent management system).
Having it on a server puts all logs, convos, documents cloud access. My actual goal is a local app, and able to add / connect your own endpoints, or connect your account, but potentially sync with the cloud.
Right now it is openAI using my key, and I've got a credit system, it works out to be about double the cost, but I've included $10USD. I'll add ability to connect your own API soon, as well as sliding scale for credits.
Using it
- Jumping on, you'll see a 'public' space I've made, and nova is configured for guests
- Sign in with single sign on and it'll generate a 'new user' space designed to do goal setting
- Both spaces were created 'with the tools' so an example of its utility
- You can then either clear all those prompts, or make a new space and start fresh, creating instructions for an agent you might find useful (or using it to test different agents for your own work)
I made a walkthrough of the basics here : https://www.youtube.com/watch?v=iQpt0B5LzNI https://www.youtube.com/watch?v=iQpt0B5LzNI
More advanced stuff
- So the toggles on the side are 'cartridges' (my term for json blob of prompts, commands, settings, all injecting at runtime)
- If you add an 'index' cartridge, it directly adds an index using LlamaIndex - so you can add a pdf, or google doc (the auth flow there isn't approved so you'll get warnings) - you can query directly in cartridge, or agent can query if commands are on
- If you add a command cartridge, it turns on 'auto-gpt-lite' - basically switches to json returns, switches on command parsing, you'll see the commands are cofigurable, but I'm going to rethink all that
- If you add a settings cartridge, it then adds settings, main ones are 'give context' injects your name, date, number of convos (based on user account)
- Most importantly you can switch between gpt-3.5-turbo and gpt-4, any typos will cook it, you can see here the sketch of configurable agents, different api sources etc
Longer video talking through these features and ideas :
https://www.youtube.com/watch?v=MM9pSd8ADuQ https://www.youtube.com/watch?v=MM9pSd8ADuQ
This has been a pretty big mission, but I'm happy with the first offer and excited to keep working on it, adding multi-step behaviours, other librarys for image rec etc. Next steps are improve command behaviour. Most importantly get user feedback!
But as it is today, configuring the agent in the way you like, and giving it an easy memory overview, and playing with different 'types' of agent, is pretty great.
New users get $10 worth of credt, can top up or donate if you find it useful. Any issues or thoughts catch me via me@samueltate.com or @samueltates on twitter.
Mostly I just hope you can try it out, I've had some great feedback about the prompt composition and document embedding. But also I just like Nova - that bit of persistence, context and agency makes a pretty cool pal.
thanks!
- ssd532 3y agoWhat does it mean that it has persistent memory?
- Der_Einzige 3y agoVector DB storage
- samueltates 3y agoSo there's a few aspects to the persistent memory, Der_Einzige is correct, vector storage is one part. You can upload and embed documents, which then get indexed by LlamaIndex (just featured and one of my favourite AI tools and actually a big driver behind me making this project). There's another aspect that is custom, that is the summary system. Basically coming from initial idea using api with ChatGPT launch last year. Takes past convos, summarises them, brings into current context. The issue is however that even that summary list gets too big, so you summarise that ad infinitum. So that was a version I had, which was fine, but there was what I'm calling lossy temporal compression, so further back things got squished, and the 'detail' of the summary was variable depending on whether it'd 'filled up/ got squished'. So I made this system that basically has rolling windows of detail, that when they get filled up, get summarised, which then puts them into the next level of summary (calling them epoches but kinda confusing). So each level of summary has a sort of 'open face' of unsummarised chunks (so latest unsummarised from each epoch), creating an exposed face of latest summaries for what essentially becomes each time period. Its kinda hard to explain i had to go into a sort of jazz trance to make it but imagine a pyramid being built from the side, but the side is staying still and the pyramid is moving backwards. But on top of that, as the summaries are happening, theyre also pulling out keywords, notes and meta data, so bubbling that up the to top, so then that memory is traversable via the 'time based' pointers (top level summaries) and keywords or notes. That way you have a 'temporally biased' view (highest detail lowest level of summary is latest), but also a flat searchable structure on topic. It is one of those OCD things where I could probably just be summarising the pyramid 'straight up' but I don't want summaries from one level mixed with another, and I don't want there to be too much variability with how many of each summary (at least for level one) there is. But what this means is that the agent has in its context an overview (pulled from next part) - so its like 'hey sam did you do the thing, are we working on the R&D report today hows your mum), but then pointers to the 'exposed face' of summaries, so latest level 1, 2, 3, so it can see 'level 1 (direct summaries of convos) -R&D report finally finished, here are details)' up to 'level 6 - september - march - sam and nova start on conversation logging system', and basically choose to 'open' those pointers, or use keywords. All of this is designed to try and keep like 500 tokens in context, so it can sort of traverse through it, (like you would skim through notes). The traversal itself I need to finish my looping system, where it can 'flick through' the notes itself (thats another story). So right now past a certain point I just flick the summaries to GPTINDEX to query (which is almost like it calling in another bot as an assistant). Anyways long story short, I was OCD on how you would manage summaries in context and this is what I came up with, I'm pretty happy with the results, but really want to improve the recall and traversal, but goal is that Nova has the right info at the right time when you're talking, like you'd expect from a pretty organised person.