7 ms·
Multi AI agent systems using OpenAI's assistants API
- deleted 2y ago[deleted]
- spiritplumber 2y agoOoh, shiny!
- xrendan 2y agoI'd be interested in knowing if anyone is seriously using the assistants API, it feels like such a lock in to OpenAIs platform when your can alternatively just use completions that are much more easily interchanged.
- phh 2y agoI've indeed refused to work with some providers giving only a chat interface and not a completion interface because it made the communication "less natural" to the model (like adding new system messages in between for function calling on models which don't officially does it, or adding other categories than system/user/assistant)
- metaskills 2y agoGreat points. Dont even get me started about how function calling in other LLMs costs me tokens. Something OpenAI provides OOTB. I'm also not a big fan of OpenAI's lock in. Right now I'm on a huge Claude 3 Haiku kick. That said, OpenAI does seem to get the APIs right and my hunch is the new Assistants API is going to potentially disrupt things again. Time will tell.
- heggy 2y agoI would love to be using Claude, but you can't get API access (beyond an initial trial period) in the EU without providing a European VAT number. They don't want personal users or people to even learn and experiment I guess.
- bjterry 2y agoYou can use the Claude APIs via OpenRouter with a pre-paid account.
- heggy 2y agoThanks, this did the job!
- metaskills 2y agoInteresting, would Amazon Bedrock be an alternative? That's how I use Claude.
- Jimmc414 2y agoI'd guess it's more likely about the additional programming needed to meet GDPR compliance requirements.
- msp26 2y ago> Dont even get me started about how function calling in other LLMs costs me tokens. Something OpenAI provides out of the box. Not sure what you mean by this.
- metaskills 2y agoI have some assumptions/guesses on how billing works. Gonna do a post on this on my unremarkable.ai blog, please do signup for posts there, no spam. I could be right or wrong but need to do some experiments and publish later.
- benreesman 2y agoOpus is really cool. I’ve found it to have a few persistent bugs in what I initially assumed is tokenization but now wonder if might be more fundamental, but modulo a few typographical-level errors, I personally think it’s the most useful of the API-mediated models for involved sessions. And there are some serious people at Anthropic, they’ll get the typo thing if they haven’t already (been a busy week and change, they easily could have shipped a fix and I overlooked it).
- BoorishBears 2y agoI'm not sure you're talking about the same thing: OpenAI specifically has a "Assistants API" that manages long term memory and tool usage for the consumer: https://platform.openai.com/assistants https://platform.openai.com/assistants I'd guestimate 99% of people using LLMs are using instruct-based message interfaces that have a variation of system/user/assistant. The top models mostly only come as a completion models, and even Anthropic has switched to a message based API
- Nedomas 2y agoI do and built Assistants API compat layer for Groq and Anthropic: https://github.com/supercorp-ai/supercompat https://github.com/supercorp-ai/supercompat I’d argue that Assistants API DX > manual completions API.
- tomrod 2y agoAye, but your FinOps will be comolaining even with simple use.
- Nedomas 2y agoAssistants API use in prod used to suck because it would send full convo on each message. But last month they added an option to send truncted history so its no longer 2$ a pop thankfully. Also Grok, Haiku and Mistral is cheap
- brianjking 2y agoAre you using Assistants API v2 with streaming?
- Nedomas 2y agoYeah, I do both in prod and in the lib. In the lib I even ported Anthropics streaming API to be OpenAI compatible. Will write the docs over the coming days if interested.
- oddthink 2y agoI know at least one team is at work is using the Assistants API, and I'm talking with another team that is leaning pretty heavily towards using it over building a custom RAG solution themselves, or even over other in-house frameworks.
- j45 2y agoI've used it and in some cases it's taking days and weeks of development away to get to testing the market. In some cases the lock in is what it is for now because a particular model in reality is so far ahead, or staying ahead. It doesn't mean other options won't become available, but it does matter to relate your need to your actions. Getting something working consistently for example might be the first goal, and then learning to implement it with multiple models might be secondary. The chances of that increase the later other models are explored in some cases. It should be possible to tell pretty quickly if something works in a particular model that's the leader, how others compare to it and how to track the rate of change between them.
- stavros 2y agoI use it mostly exclusively (I've even developed a Python library for it, https://github.com/skorokithakis/ez-openai https://github.com/skorokithakis/ez-openai), because it does RAG and function calling out of the box. It's pretty convenient, even if OpenAI's APIs are generally a trash fire.
- beoberha 2y agoFrom the website linked in the readme: “A lot of research has been doing in this are and we can expect a lot more in 2024 in this space. I promise to share some clarity around where I think this industry is headed. In personal talks I have warned that multi-agent systems are complex and hard to get right. I've seen little evidence of real-world use cases too” These assistant systems fascinate me, but I just don’t have the time and energy to set something up. I was going to ask if anyone had a good experience with it, but the above makes it sound like there’s not much hope at the moment. Curious what other people’s experience are.
- m3kw9 2y agoBy the time you do get around to it OpenAi would have built a full interface for this. This is the type of stuff that’s gonna get steamrolled.
- metaskills 2y agoThanks @beoberha, I am too. I like one take I heard on Twitter. The sentiment was something like these types of systems are useful under the AI-Powered Productivity industry which has incremental gains, no big bangs. Said another way, if your job was to help a TON of your employees be more productive individually, it is worth it because companies measure those efforts broadly and the payoff is there. But again, not big. My advice for folks to stay lower level and hook AI automation up with simple, closed loop, LLM patterns that feel more like basic API calls in a choreographed manner. OMG, hope all that made sense
- csouzaf 2y agoWhat's the use cases people are using Multi AI Agents to solve problems that deliver real value? Someone has something with your hands on right now?
- coffeebeqn 2y agoI tried the last crop. Interesting idea but the success rate of any real multi step task always approached 0% the longer it went
- ww520 2y agoI imagine having an agent set up with specific RAG context to solve a specific problem and having another with a different RAG context to solve a different problem can be useful.
- csouzaf 2y agoI see customer support as a very talked subject to solve this. But these system really manage to solve the issue removing the human feedback dramatically?
- LASR 2y agoWe’ve tried. A lot. Custom frameworks and all. There is really no way to make the ensemble behave with an acceptable level of consistency. Where we ended up is now having a frontier model generate a whole tree of possible execution plans, and then have the user select one of those path, and then we just run whatever the user chose in a plain sequence until the next decision point that needs user approval.
- avereveard 2y agoI've encountered two viable cases: instructions are too complex, too many tools, or wildly different processing steps, in which case it semplify a lot the processing to have a few well defined steps each doing their thing, and a coordinator on top, either sequential, or intelligent, that is only focuesed on next step routing. the other is memory for conversational retrieval. ai memory is still quite limited, especially if there needs to be a lot of token in context, and context too long impede the ability of llm of focus on the task itself, especially if the context is itself a conversation or a request, so spreading the context along a few agents, and propagating the user request among agent, and having those produce answer fragment for another llm to formulate an answer allows to not lose the conversational context without swamping the llm with noise. the problem tho remains latency as son as you nest them latency explodes as you can only stream the last layer of llm output
- obiefernandez 2y agoMy main conversation “loop” at https://olympia.chat https://olympia.chat has tool functions connected to “helper AIs” for things such as integrating with email. It lets me minimize functions on the main loop and actually works really well.
- bongodongobob 2y agoI'm sorry but that is absolutely hilarious.
- Terretta 2y agoSid Kapoor, Content Specialist, forgot to include himself in Growth or Pro plans. Guess he is Basic!
- deleted 2y ago[deleted]
- moltar 2y agoBare JS. What is this 2001?
- metaskills 2y agoLMAO. Yes, I love ESM modules. So maybe more like 2012 or 2015. Would you like to see TypeScript?
- alluro2 2y agoThank you for using vanilla JS!
- taf2 2y agoYes this is great so much easier to work with
- metaskills 2y agoY'all just made my day!
- fy20 2y agoNot OP, but I use TypeScript because it adds a layer of safety to the codebase. It's like having good test coverage - you can make large changes and if the tests pass (the code compiles), you can be fairly confident that you didn't mess anything up. I've written Ruby for years, so I'm used to dynamically typed languages. But JavaScript is it's own level of special, and there's so many ways you can accidentally mess things up. Having tests cover every single path (especially failure paths) can be very time consuming, and often hard or messy to setup (how would you mock the OpenAI module returning an error when adding metadata to a thread?), where as using something like TypeScript can make sure your code handles all paths somewhat correctly (at least as well as the types you defined). Your code looks clean, and you appear to have good test coverage, so you do you though :-)
- deleted 2y ago[deleted]
- 2y ago
- __loam 2y agoI've not seen any of these "agentic" systems be all that useful in practice. Complicated chain of software where a lot can wrong at any step, and the probability of failure explodes when you have many steps.
- behnamoh 2y agoI stay away from such frameworks because: - Writing what I want in Python/other-lingo gives me much more customizability than these frameworks offer. - No worries about the future plans of the repo and having to deal with abandonware. - No vendor lock in. Currently most repos like this focus on OpenAI's models, but I prefer to work with local models of all kinds and any abstraction above llama.cpp or llama-cpp-python is a no-no for me. The last point means I refuse to build on top of ollama's API as it's yet another wrapper around llama.cpp.
- rcarmo 2y agoNot using the ollama API means you have to keep track of context yourself, and run all your stuff in the same box. Hardly ideal.
- yatz 2y agoAssistants API is promising, but earlier versions have many issues, especially with how it calculates the costs. As per OpenAI docs, you pay for data storage, a fixed price per API call, + token usage. It sounds straightforward until you start using it. Here is how it works. When you upload attachments, in my case a very large PDF, it chunks that PDF into small parts and stores them in a vector database. It seems like the chunking part is not that great, as every time you make a call, the system loads a large chunk or many chunks and sends them to the model along with your prompt, which inflates your per request costs to 10 times more than the prompt + response tokens combined. So, be mindful of the hidden costs and monitor your usage.
- deleted 2y ago[deleted]
- iamflimflam1 2y agoThere isn’t really any other way for this to work. The only way for the model to answer questions on your pdf is for the information to be somewhere in the prompt.
- benreesman 2y agoThat might be true of specific models or specific APIs for accessing them, but I’d argue isn’t even remotely true of neural networks generally or generatively-pretrained decoder-only attention-inspired language models in particular. Ideally if you want a model’s weights to include a credible representation of non-trivial data you want it somewhere in the training pipeline (usually earlier is better for important stuff but that’s a hubristic at best), but there’s transfer learning of various kinds, and joint losses of countless kinds (CLIP in SD-style diffusors come to mind), and fine tunes (if that doesn’t just count as transfer learning), and dimensionality reduction that is often remarkably effective, and multi-tower models like what evolved into DLRM, and I’m forgetting/omitting easily 100x the approaches I mentioned. It’s possible I misunderstand you, so please elaborate if so?
- barfbagginus 2y agoThe way they vectorized the PDF could be less efficient than simply extracting the text and dropping it into context as text. If it's a 100 MB PDF then it's probably a scanned PDF, and OpenAI is probably using an OCR model to vectorize each page directly. It seems an opaque process with room to be inefficient. So I would be interested to know if we could save on token/vector fees by preprocessing the PDF to text with our own OCR.
- tr14 2y agoAI botnet?
- mentos 2y agoAnyone recommend the best way to use AI to search all of my documents for a project. I've got specifications, blueprints, emails, forms, etc. Would be great to be able to ask it, 'have we completed the X process with contractor Y yet?'
- valiant-comma 2y agoTry h2ogpt: https://github.com/h2oai/h2ogpt https://github.com/h2oai/h2ogpt
- squirrel 2y agoZenfetch
- ec109685 2y agoI don’t understand the comment about server send events not being async friendly. What is unfriendly about this? import OpenAI from 'openai'; const openai = new OpenAI(); async function main() { const stream = await openai.chat.completions.create({ model: 'gpt-4', messages: [{ role: 'user', content: 'Say this is a test' }], stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content || ''); } } main(); It’s easy to collect the streaming output and return it all when the llm’s response is done.
- deleted 2y ago[deleted]
- wokwokwok 2y agoThey’re referring to the: > assistant.on("textDelta”, () => … Callbacks, which are not async and can’t be streamed that way directly without wrapping it in some helper function. (Which does seem obvious; I’m also not sure why they called it out specifically as not being async friendly? I guess most callback style functions these days have async equivalents in popular libraries and these ones don’t)
- arresin 2y ago> const stream = await… Is this right? Aren’t you prematurely unwrapping the promise here?
- ec109685 2y agoI believe that is what gets the call started so awaiting there is okay. There isn’t anything to stream at that point.
- david_shi 2y agoA bit off topic, but has anyone seen any agent systems focused on improving the agents capabilities with more usage?
- jackbravo 2y agoFrom their linked main page: > In my opinion, exploration of multi-agent systems is going to require a broader audience of engineers. For AI to become a true commodity, it needs to move out of the Python origins and into more popular languages like JavaScript , a major fact on why I wrote Experts.js. I wholeheartedly agree