4 ms·
You are absolutely right. I have a built a handful of complex interconnected LLM calls, and the thing I am dying for is a simple way to inspect what is happenin
by pstorm 3y ago
You are absolutely right. I have a built a handful of complex interconnected LLM calls, and the thing I am dying for is a simple way to inspect what is happening at each step. For instance, one chain has 5+ steps where data is fetched, transformed, sent to an LLM, and transformed again. I have a homegrown solution to seeing what's going on, but it is missing a lot of what I would want to see.
I've seen some attempts at providing RAG tracing as a service, but for my specific use case, where in some instances I am making 100+ LLM calls per chain, it isn't a fit quite yet.
- phillipcarter 3y agoWhat we do, admittedly not with chaining, is use OTel tracing end-to-end. Our RAG pipeline is (currently) 42 spans, and the end-to-end of RAG + LLM call + parse/validate is 48 spans. That's just a part of the whole trace, though. Our feature is a natural language querying system built onto a querying engine. And so when someone has their input typed out and clicks the button (or hits enter), we trace that user interaction, RAG, LLM call, parse/validate, and queuing up to run against the querying engine. Getting that "whole experience" visibility really helped us iterate in production. All of it was done just by manually instrumenting things with opentelemetry tracing and having our own conventions for naming key:value pairs that we're looking to upstream. It's more work to set up if you're not familiar with opentelemetry, but it's worth it because you can then start plugging in user IDs or really any other interesting data that helps correlate with LLM behavior so you get a full picture of what's going on.