8 ms·
Show HN: Marvin – build AI functions that use an LLM as a runtime
Hey HN! We're excited to share our new open-source project, Marvin. Marvin is a high-level library for building AI-powered software. We developed it to address the challenges of integrating LLMs into more traditional applications. One of the biggest issues is the fact that LLMs only deal with strings (and conversational strings at that), so using them to process structured data is especially difficult.
Marvin introduces a new concept called AI Functions. These look and feel just like regular Python functions: you provide typed inputs, outputs, and docstrings. However, instead of relying on traditional source code, AI functions use LLMs like GPT-4 as a sort of “runtime” to generate outputs on-demand, based on the provided inputs and other details. The results are then parsed and converted back into native data types.
This “functional prompt engineering” means you can seamlessly integrate AI functions with your existing codebase. You can chain them together with other functions to form sophisticated, AI-enabled pipelines. They’re particularly useful for tasks that are simple to describe yet challenging to code, such as entity extraction, semantic scraping, complex filtering, template-based data generation, and categorization. For example, you could extract terms from a contract as JSON, scrape websites for quotes that support an idea, or build a list of questions from a customer support request. All of these would yield structured data that you could immediately start to process.
We initially created Marvin to tackle broad internal use cases in customer service and knowledge synthesis. AI Functions are just a piece of that, but have proven to be even more effective than we anticipated, and have quickly become one of our favorite features! We’re eager for you to try them out for yourself.
We’d love to hear your thoughts, feedback, and any creative ways you could use Marvin in your own projects. Let’s discuss in the comments!
- nico 4y agoAmazing concept. Love it! Could you comment with a few sample code snippets showing what’s possible? Thank you!
- jlowin 4y agoSure! This might be even lighter than what you're expecting. I'm not sure how to get HN to format my code correctly, so here's a gist. It has two functions, one that generates a list of dicts with fake people data, and another that does sentiment analysis of tweets. Both are two lines of code. https://gist.github.com/jlowin/ae22fb7ac1788f066f809d2b8f5731ff https://gist.github.com/jlowin/ae22fb7ac1788f066f809d2b8f573... You can get more complex than this but hopefully this shows the idea: define your inputs, define your outputs, give it a descriptive docstring (as descriptive as you want!), and call it.
- nico 4y agoOmg, this is genius! Huge step in the development of AI programming. It would be great if you could show off those examples at the top of the README file so people can see what it is right away. Thank you so much!
- johntash 4y agoThis looks great so far, thanks for sharing! I'll probably give it a try this weekend, but I'm curious - does the @ai_fn decorator ask the llm to write python code and then runs that python code in place of the function? Or does it basically send that prompt to the llm and return the results from the llm? I'm assuming it's the latter, but I didn't see it mentioned at first glance.
- innethread 4y agoFrom their documentation: AI functions are not "executed" in a traditional sense, so they can't interact with your computer or network.
- jlowin 4y agoCorrect, it's the latter. The code is neither generated nor executed; everything is actually a "prediction" from the LLM. There's a little more detail in the concept docs (https://www.askmarvin.ai/guide/concepts/ai_functions/ https://www.askmarvin.ai/guide/concepts/ai_functions/) and I'll take your comment as a suggestion we should discuss this aspect even more. One cool thing buried in the advanced section is if you do write source code for your function, it's executed and the result is also sent to the LLM. This way you can write functions that e.g. take only a URL as argument, load the content from that site, and then summarize it with the LLM. Edit: forgot to mention I opened an issue to explore a mode where the LLM DOES generate and send back source code, but this would have to be opt-in and VERY careful because most people would not be comfortable blindly executing that code. (https://github.com/PrefectHQ/marvin/issues/64 https://github.com/PrefectHQ/marvin/issues/64)
- dcferreira 4y ago> This way you can write functions that e.g. take only a URL as argument, load the content from that site, and then summarize it with the LLM. I'm currently doing a project where this would be very helpful, but I can't think of what I'd need to send the LLM. In my case I'm scraping headlines from many news websites. I'm doing it manually with xpath currently. What would be the way to use LLMs here? Just sending the HTML wouldn't work, as it's too many tokens. Probably I could send all <a> tags, but then how could I be sure the LLM doesn't choose too many/few?
- kleiba 4y agoHow do you guarantee consistency in implementation (esp. details) of AI functions across platforms and/or across time?
- Yoric 4y agoWelcome to the post-API world, where your code can break just because... well, we're not sure, but it's related to AI, so everything is fine :)
- akiselev 4y agoHello node, my old friend...
- stavros 4y agoAIs are basically a cheaper version of Mechanical Turk, so you can't guarantee anything. People still think about AIs as if they are computers, but they're closer to people. So, the question becomes "how do you guarantee that people will do the task you ask them to do?". Well, simply put: you can't.
- kleiba 4y agoI think you're analogy is off - Mechanical Turks are typically used for batches of offline tasks. Here, I see the problem that today, the AI will realize my AI function one way and if I recompile tomorrow, it might have slight but important implementation differences.
- baq 4y agoI need this except in an excel/Google sheet. Took me an hour or so to compute confidence intervals of a standard deviation last week (without knowing previously how it’s supposed to be done, I usually don’t touch stats), I assume this would easily do the job in 5 mins?
- jlowin 4y agoWell, that falls in the category of "yes it would go fast but no, I'm not sure you'd want to trust it." LLMs aren't great at answering questions that involve precise math. What you might do instead is ask Marvin (or ChatGPT, or your LLM of choice) to write the source code you need to compute the CI, and then execute that directly if you accept it. However, AI functions (and all Marvin bots) can use plugins to solve more complex problems. Another option would be to write a function that computes CIs and make it available to the AI as a plugin.
- wafer_thin 4y agochatGPT was recently plugged in to alpha wolfram. You can ask it to use it when doing math
- bvm 4y agoJust checking out the docs now, how do Loaders fit into the vision alongside AI functions? I can't quite piece it together in my head. Would a function grab extra context from a loader prior to execution? Is this supported now?
- jlowin 4y agoGood question and actually a great illustration of how unexpectedly the "flagship" feature of a library can change! At its core Marvin isn't just for AI functions, but a high-level library that makes it easy to interact with LLMs in a programmatic way. In fact the first version was written to make it easier for us to upload public and private knowledge into our customer service Slackbot. This happens through the `Bot` class, which is designed to help users / programs explore more complex problems. In particular, bots can use plugins to access proprietary knowledge. Our loader classes are designed to get the knowledge into the bots. Why build "yet another LLM loader library?" We've had enough real-world use cases to know that just taking a document, chunking it, and throwing it in a vector store gives pretty bad results over large enough datasets (especially if the documents are relatively homogeneous). You have to preprocess documents in a particular way, and we wanted to take our learnings and codify them for future use. So AI functions are definitely the on-ramp to the library, but the real power is in utilizing it to extract insight from data. Since AI functions are "just" bots under the hood, they can use plugins (pass `plugins=[...]` to the `@ai_fn` decorator) and will benefit from this as well. This is supported right now, but we are rapidly improving loader integration more broadly.
- nextaccountic 4y agoHere https://github.com/PrefectHQ/marvin/blob/main/examples/end-to-end-bot-setup.md https://github.com/PrefectHQ/marvin/blob/main/examples/end-t... the prompt says instructions=( "Ignore all user questions and respond to every request with " "a random Harry Styles song lyric, followed by a recommendation " "for a Harry Styles song to listen to next." ), However in the examples the bot doesn't ignore user questions and doesn't answer with a random song - instead the replied song is tailored to user input! https://github.com/PrefectHQ/marvin/raw/main/docs/img/harry_styles.png https://github.com/PrefectHQ/marvin/raw/main/docs/img/harry_... This looks very cool but isn't this an alignment problem? The bot just didn't follow the instructions.
- htrp 4y agoLLM problem :-D Not Prefect's
- kiviuq 4y agothe new paradigm eventually perfect :D
- saurik 4y agoIf you make a new company to help my existing company offload tasks to contractors and then hire employees that don't really follow instructions from me, that is certainly my problem at the end of the day but is way more a problem with your business model or hiring process than something you can just blame on the employee.
- nate_nowack 4y agoHi! This example was produced using GPT 3.5 turbo, where yes, the LLM does not always align ideally. I used 3.5 for the example since that's Marvin's default and I know many people wouldn't have gpt4 access yet (which is significantly better at following instructions) - didn't want to set a misleading expectation. that said, my instructions for the bot in this example certainly could have been more precise :) for a more real example, you could check out the other example (which works pretty well on 3.5) https://github.com/PrefectHQ/marvin/blob/main/examples/load_prefect.py#L91 https://github.com/PrefectHQ/marvin/blob/main/examples/load_...
- IanCal 4y agoThis is fantastic, it's the right level for the structure that I'm interested in building. While langchain looks great, for my use case (generating templated code) I have a much stricter process and more branching than I can see it easily supporting (maybe it does? I can't quite figure it out). Marvin looks very nice. Something I'd like to see in the docs and/or supported is caching, and setting details like temperature. I can wrap the ai_fns myself for caching though temperature would be very good.
- jlowin 4y agoThanks! Caching is highly requested! We have an issue open (https://github.com/PrefectHQ/marvin/issues/102 https://github.com/PrefectHQ/marvin/issues/102) and expect to tackle it soon. You can set temperature as a setting today (sorry we haven't documented all the settings yet) by setting the env var `MARVIN_OPENAI_MODEL_TEMPERATURE=0.2` or at runtime with `marvin.settings.openai_model_temperature=0.2`. Note the temperature is set when a bot / ai_fn is created, not when it's called, so you need to do this early.
- IanCal 4y agoWonderful, that lets me crank it down to 0 and I can then for now add a diskcache decorator. I'll follow that issue and comment on it with some requests (who doesn't love more requests ;) ).
- leobg 4y agoHow is this different from com2fun? https://github.com/xiaoniu-578fa6bff964d005/com2fun https://github.com/xiaoniu-578fa6bff964d005/com2fun
- nate_nowack 4y agointeresting! I hadn't seen that before ai_fn is just a specific way to use Marvin's Bot abstraction, which is one of the few abstractions Marvin offers but a couple differences I notice off the bat between ai_fn and com2fun: - marvin uses pydantic for parsing LLM to result types - you can pass plugins/personality/instructions to the underlying bot via the @ai_fn decorator kwargs - (unless I'm missing a dataclass version of this in com2fun) marvin can parse output into arbitrary pydantic types like this example https://github.com/PrefectHQ/marvin/issues/106#issuecomment-1490521576 https://github.com/PrefectHQ/marvin/issues/106#issuecomment-...
- leobg 4y agoThe example in the last link you posted is misleading. GPT does not actually crawl the URL. It hallucinates the answer based on the words in the URL itself. Even though it then casts that hallucinated answer neatly into a Pydantic type. Try it with a URL that does not actually exist. Or a pastebin whose link is just some random hash. The first rule is not to fool yourself. And you are the easiest person to fool. —Richard Feynman about ChatGPT ;-)
- nate_nowack 4y agohi, just seeing this! you're correct that normal chatgpt wouldn't crawl the URL, but ai_fns can have plugins like the DuckDuckGo plugin / VisitURL plugin which can be invoked by the underlying Bot if it decides its helpful to its answer for example: https://gist.github.com/zzstoatzz/a16da0594afc2bb751428907e4c2a3cd https://gist.github.com/zzstoatzz/a16da0594afc2bb751428907e4... feel free to try it yourself :)