4 ms·
We did something similar for our reporting service which is based duckdb. Overall it works great, though we've ran into a few things: * Even with low temperatu
by qiller 3y ago
We did something similar for our reporting service which is based duckdb. Overall it works great, though we've ran into a few things:
* Even with low temperature, GPT-4 sometimes deviates from examples or schema. For example, sometimes it forgets to check one or another field...
* Our service hosts generic data, but customers ask to generate reports using their domain language (give me top 10 colors... what's a color?). So we need to teach customers to nudge the report generator a bit towards generic terms
* Debugging LLM prompts is just tricky... Customers can confuse the model pretty easily. We ended up exposing the "explained" generated query back to give some visibility of what's been used for the report
- ignoramous 3y agoCurious, as we're looking to build / use a similar setup. > Debugging LLM prompts is just tricky... Customers can confuse the model pretty easily. Would a RAG like how Vanna.ai uses, help? > For example, sometimes it forgets to check one or another field Do prompting techniques like CoT improve the outcome? > So we need to teach customers to nudge the report generator a bit towards generic terms. Did you folks experiment with building an Agent-like interface that asks more questions before the LLM finally answers?
- qiller 3y agoOur primary issue is that our DB is a dynamic Entity-Attribute-Value schema, even quite a bit denormalized at that. The model has to remember to do subqueries to retrieve "attributes" based on what's needed for the query and then combine them correctly. NLQ is a somewhat new feature for us, so we don't have a great library to pull from for RAG. Experimenting, I found that having a few-shot examples with some CoT (showing examples of chaining attributes retrieval) sprinkled around did help a lot. Even still, some queries come out quite ugly, but still functional. I'm thankful that DuckDB is a beast when tackling those :D > Did you folks experiment with building an Agent-like interface that asks more questions before the LLM finally answers? That's something I want to figure out next: 1) try to check if a generated query would work but would generate absolutely junk results (cause the model forgot to check something) and ask to rephrase 2) or show results (which may look "real" enough), but give an ability to tweak the prompt. A good example is something like "top 5 products on Cyber Monday" <- which returns 0 products, cause 2024 didn't happen yet, and should trigger a follow up.
- totalhack 3y agoMaybe you could utilize views to make your EAV schema more friendly for the LLM? Whether that's realistic depends on the specifics of your situation of course.
- atomicvibe 3y ago>Did you folks experiment with building an Agent-like interface that asks more questions before the LLM finally answers? We built something similar to query DB. Created two versions, one of which was agent based that had a maker-checker style of generation. Basically one generates and the other checks its correctness and if objective has been acheived. The accuracy improves in the agent driven framework, at the cost of latency.