13 ms·
RasaGPT: First headless LLM chatbot built on top of Rasa, Langchain and FastAPI
- riter 3y agoUnfortunately there were not a whole lot of end-to-end examples of integrating Rasa with OpenAI nor functional boilerplates on github so I put a working prototype together in a few days and thus RasaGPT was bron. RasaGPT is a python-based boilerplate and reference implementation of Rasa and Telegram utilizing an LLM library like Langchain for indexing, retrieval and context injection. FastAPI end-points are made available for you to build your application on top of. Features include: - Automated hand-off to human if queries are out of bounds - "Training" pipeline done via API - Multi-tenant support - Generate category labels from questions - Works right out of the box with docker-compose - Ngrok reverse tunnel and dummy data included - Multiple use cases and a great starting point Hope you like it, more @ rasagpt.dev
- contravariant 3y agoI haven't worked with Rasa so I was wondering if I understood things correctly. Are you using a language model to look up the correct reply to a particular response inside Rasa? Where Rasa presumably connects to some kind of backend to retrieve information or 'do stuff'?
- riter 3y agothanks for asking. this implementation leverages Rasa and stands up a FastAPI server where it receives the user response webhook first and gets processed by (or bypasses) Rasa. The LLM queries a set of documents indexed by Langchain. Dummy data has been included (Pepe Corp.) Rasa has support for a "fallback" mechanism whereby if a user's response scores low on your pre-configured Rasa intents (like Greet) you can have it route directly to the LLM as well. But for now RasaGPT capture and routes the Telegram response to the FastAPI webhook endpoint. the LLM itself and prompts I configured provides a boolean on whether the response should be escalated to a human or not, based on LLM+Langchain not knowing the answer to the user's query from the indexed documents. I hope that answers your question, if not happy to follow-up!
- 0x008 3y agoCan somebody ELI5 Rasa for me? I read through the README and I still don't get what it does.
- zwaps 3y agoI guess it connects your langchain bot to some API like e.g. slack?
- riter 3y agoTotally. Rasa (https://github.com/RasaHQ/rasa https://github.com/RasaHQ/rasa) is an open source chatbot platform. It allows you to setup "Input Channels" e.g. slack telegram, and has an intents and response pipeline. It leverages pre-LLM NLU models (NLTK, BERT, etc.) to score intents and based on that intent it will automate a pre-configured response. My implementation allows you directly route (or fallback to) to GPT-3 or GPT-4 via Langchain document retrieval. So essentially this is an example of a knowledgebase customer support bot. I hope that makes sense, let me know if not!
- janmo 3y agoA bit off topic but you better change the name and remove the GPT. OpenAI is claiming AI products that are using GPT in their name are causing confusion and is sending legal threats now. One of many examples: https://twitter.com/pbteja1998/status/1654095756200931328 https://twitter.com/pbteja1998/status/1654095756200931328
- riter 3y agoI appreciate the feedback. I didn't realize they were acting on it. Would Rasa-LLM sound as compelling?
- KaoruAoiShiho 3y agoYes it sounds better and less confusing.
- danjc 3y agoIs it a GPT?
- riter 3y agoIt itself is not a GPT. It is a a framework of a framework project built on top of Rasa (https://github.com/RasaHQ/rasa https://github.com/RasaHQ/rasa) and Langchain which by default uses gpt3.5-turbo (change it in the .env file) or any foundation model you wish.
- mirekrusin 3y agoOther good alternatives may include: * KnowsItAllKaren * GeniusJack * GuruGary * BotBecky * ChatterBoxChantelle * SmartypantsSam
- cehrlich 3y agoI'd suggest to not put a bunch of 4chan memes in your product demos.
- riter 3y agowhy is that exactly? is it offensive, if so I'm unaware and appreciate the feedback.
- bheadmaster 3y ago[flagged]
- hereonout2 3y agoLots of references to that weird cartoon frog in the JSON output. Using something that's been quite controversial in the past does seem at least a little naive ... https://en.m.wikipedia.org/wiki/Pepe_the_Frog https://en.m.wikipedia.org/wiki/Pepe_the_Frog
- Beaver117 3y agoevery meme started as a 4chan meme
- becquerel 3y agothis is a great disrespect to the memes that were born on SA
- Der_Einzige 3y agoIt worked just fine for the stable diffusion community, where automatic1111 puts a ton of credit to 4chan for the development of stable diffusion tooling
- data_maan 3y agoI'm not sure what the advantage the use of a somewhat comprehensive framework like Langchain gives you for this use case? It starts to feel as AI tech is slowly turning into web tech with a million tools and frameworks, so I'm just wondering whether all of these are needed and if it isn't easier to code your own than learning a foreign framework...
- riter 3y agoNot off-topic at all. After struggling with LangChain's hyper-opinionated implementation of classes I agree. In fact, this is better off leveraging Llamaindex. This is a proof-of-concept and ultimately leveraging a library / framework helps afford the following: - easy implementation of chunking strategies when you're unsure - OpenAI helper functions - embeddings and vector store management Again, even with the above I struggled and had to implement PGVector myself. Going into production once I have my document retrieval strategy and prompt-tuning optimized, I would never use Langchain in production simply bc of the bloat and inflexible implementation of things like the PGVector class. Also the footprint is massive and the LLM part can be done in 5% of the footprint in Golang and 5% of the cloud costs. So I actually agree with you :)
- yawnxyz 3y agoSomeone needs to create a “Langchain, but less complicated” framework
- riter 3y agolololol. i think this opportunity gets bigger post $10m seed round. they'll likely double down and expand footprint vs the inverse. check out llama-index. its purpose-built for document indexing and retrieval and less agents and "everything else"
- data_maan 3y agoWhat do you by post 10m seed round? Do you mean if LlamaIndex starts collecting VC? I'm not sure, are they for-profit?...
- aantti 3y agoAlso, with Haystack and a smaller Transformer model to address the long-tail of answers https://github.com/deepset-ai/rasa-haystack https://github.com/deepset-ai/rasa-haystack (and https://www.deepset.ai/blog/build-smart-conversational-agents-with-chatbots-qa https://www.deepset.ai/blog/build-smart-conversational-agent...)
- depr 3y agoCan you actually build a reliable customer-facing chatbot on top of LLM's? With the "jailbreaking" and not knowing if it's actually using the data you're supplying it or other data it was trained on and so on.
- riter 3y agoyes. there are a few approaches which i intend to take and some helpful resources: You could implement a Dual LLM Pattern Model https://simonwillison.net/2023/Apr/25/dual-llm-pattern/ https://simonwillison.net/2023/Apr/25/dual-llm-pattern/ You could also leverage a concept like Kor which is a kind of pydantic for LLMs: https://github.com/eyurtsev/kor https://github.com/eyurtsev/kor in short and as mentioned in the README.md this is absolutely vulnerable to prompt injection. I think this is not a fully solved issue but some interesting community research has been done to help address these things in production
- depr 3y agoThanks, I hadn't seen those. I did find https://github.com/NVIDIA/NeMo-Guardrails https://github.com/NVIDIA/NeMo-Guardrails earlier but haven't looked into it yet. I'm not sure it solves the problem of restricting the information it uses though. For example, as a proof of concept for a customer, I tried providing information from a vector database as context, but GPT would still answer questions that were not provided in that context. It would base its answers on information that was already crawled from the customer website and in the model. That is concerning because the website might get updated but you can't update the model yourself (among other reasons).
- bravura 3y agoCurious if people want to suggest alternatives to Rasa for writing stateful chatbots. Or share feedback about using Rasa.
- aantti 3y agoThis was an interesting read :) https://www.pinecone.io/learn/javascript-chatbot/ https://www.pinecone.io/learn/javascript-chatbot/
- riter 3y agothe next best platform I could find for my friend I was helping was google's dialog flow. again, it was managed, closed-source opinionated and not as flexible. and most importantly design considerations were for a pre-LLM world. i personally think there is an acute opportunity for creating a bare bones rasa built with LLMs in mind. the core concepts behind rasa are useful (domains, intents, actions, etc.) but the underlying NLU technology and assumptions around the platform are obsolete so 70% of the footprint is unnecessary. just my humble Ξ0.02
- dinerodiva 3y agoWhat would you use if you were building a consumer facing chat bot, if not Rasa?
- lachlan_gray 3y agoI’m making a similar “chatbot backend” kind of thing right now, meant to be configurable and versatile, while avoiding dependencies. It’ll be compatible with stuff like llama index and pinecone but doesn’t impose it. I’ll have a minimal foundation ready this week, I can send it your way when it’s up if you’d like
- sergioprada 3y agoI've been working on Motorhead! https://github.com/getmetal/motorhead https://github.com/getmetal/motorhead
- xchip 3y agowhat problem is this solving?
- riter 3y agoit solves how to integrate LLMs (Langchain) an application API pipeline with Rasa... of which I could not find an out-of-the-box public example on github. and so here we are :) TL;DR: i solved a friend's headache (at the time)
- MichaelMoser123 3y agoonly google bard could tell me (chatgtp3.5 has a too old knowledge cutoff date). They are trying to get the capabilities of ChatGTP plugins (apparently) >what is Langchain? LangChain is a framework for developing applications powered by large language models (LLMs). It provides a standard interface for connecting LLMs to other sources of data, as well as for persisting state between calls of a chain/agent. LangChain also provides a collection of memory implementations, and examples of chains/agents that use memory. Here are some of the features of LangChain: Data-aware: LangChain allows you to connect your LLM to other sources of data, such as databases, APIs, and file systems. This makes it possible to build applications that are more powerful and differentiated than those that simply call out to an LLM via an API. Agentic: LangChain allows you to build agents that use LLMs to make decisions and take actions. This makes it possible to build applications that can interact with their environment in a more sophisticated way. Chains: LangChain allows you to chain together multiple calls to an LLM, or to other utilities. This makes it possible to build more complex applications that can perform multiple tasks. LangChain is a powerful tool for building applications powered by LLMs. It is easy to use and provides a wide range of features. If you are interested in building an application that uses an LLM, I encourage you to check out LangChain. >is langchain trying to extend the LLM with plugins, similar to what is possible with ChatGtp? Yes, LangChain is trying to extend the LLM with plugins, similar to what is possible with ChatGtp. This will allow users to add new functionality to their LLMs without having to modify the underlying code. For example, a user could create a plugin that adds the ability to translate text, or to generate different creative text formats. This would allow the user to use their LLM for a wider range of tasks. LangChain is still under development, but the team is working on making it easy for users to create and deploy plugins. They are also working on adding new features to the plugin ecosystem, such as the ability to share plugins and to rate plugins. I believe that LangChain has the potential to be a powerful tool for developers who want to build applications powered by LLMs. The ability to extend the LLM with plugins will make it even more powerful and versatile.
- BaculumMeumEst 3y agoSorry for the off topic question, but does anyone know how to buy consumer hardware optimal for running emerging open source chat models with the largest parameter chat models possible? Would it be more cost effective to try to buy an absurd amount of ram and run on the cpu? Or buy an Nvidia card with the biggest capacity available? Or maybe buy a Mac with the most memory you can get?
- moffkalast 3y agoWell for running the average model as-is without spending a few days figuring out why you're getting strange errors and can't get it working you more or less need CUDA support. As much VRAM as you can get is probably also a good idea. For reference I can seemingly run Vicuna-7B (I think the 4 bit version) on my 6G 1660 Ti at roughly 1.5 tokens per second. Way too slow for anything useful, so you can imagine what CPU inference would look like.
- londons_explore 3y agoCPU inference is only a little slower. GPU's aren't good for a batch size of 1 and everything quantised.
- execveat 3y agoI get 3 tokens per second on M1 Max running 30B models compared to 1 token per second on a GPU (P40), both quantized to 4bit. So, in my opinion CPUs are better for inference (at least fast CPUs with DDR 5 versus cheapest GPUs). The reason why GPUs seem to be the standard de facto is that they scale better, are more power efficient and are better supported by pytorch & co. Also, academia cares more about getting the best quality for their benchmarks, than about the performance and accessibility.
- londons_explore 3y agoGPU's win for training... And those who write papers and publish code tend to do lots of training and only a little inference.
- deet 3y agoTangentially, it's interesting seeing an open source project like this actually spin up a domain name, contact email, and some branding (the image in the Readme), for a project the author said was created in just a few days. I wonder what the objective is for that extra polish. If it's optimizing star count growth, how much do these touches help?
- riter 3y agoOP here. that's a somewhat cynical interpretation. what if i just care about aesthetics and want to raise the bar. my primary motivation was to get users of Rasa out of a directional hole bc that's where i was. of course i like stars. it's a video game and i like winning. it was actually created in a few days all by me. no ulterior motive, literally indexing a solution to my problem from ~a week ago. my bg is eng + product so i do these things as reflex and have a love for good UX. nothing more. nothing less.
- deet 3y agoSorry, I didn't mean to imply any nefarious ulterior motives here. I'm more just intellectually curious about the dynamics of Github and marketing on it these days, whether it's for attracting contributors to non-commercial OSS projects or more commercial objectives where rapid growth leads to userbase, funding, etc. The project looks quite interesting and I agree we need a way to bridge the gap between traditional bot creation frameworks and the more LLM-centric approaches of late.
- riter 3y agoah! in that case glad you asked. my objective falls into neither bucket. i want rasa users to find it so i optimize for search (GH tags, clear description), ease of use (video, addt'l MD files) and perception (logos) but i'll be honest, for my intention it has a diminishing rate of return. at minimum i find canonical README sections like quick start, installation, how it works is necessary if you want to be helpful. helpfulness is difficult to measure outside of inbound emails thanking you / forks w/ actual commits. hope that gives some kind of insight. just make everything awesome :)
- darepublic 3y agoEverybody racing into the AI space to plant their flag and say "First!". But first isn't going to be correlated with the winner much, I'd wager
- riter 3y agoOP here. i agree. perhaps you're confused on the intent. the only flag being planted is for folks using rasa looking for a reference implementation just like i was a week ago. not sure if you're being intentionally cynical but trying is good thing. why? bc most ppl don't try. you make 0 of the shots you never take. and of course, if you're not intentionally being cynical -- gucci. if you are i encourage you to make your next comment substantial or encouraging :)