3 ms·
I think the biggest area for growth for LLM based tools for data analysis is around helping users _understand what edits they actually made_. I'm a co-founder
by narush 3y ago
I think the biggest area for growth for LLM based tools for data analysis is around helping users _understand what edits they actually made_.
I'm a co-founder of a non-AI data code-gen tool for data analysis -- but we also have a basic version of an LLM integration. The problem we see with tooling like Pandas AI (in practice! with real users at enterprises!) is that users make an edit like "remove NaN values" and then get a new dataframe -- but they have no way of checking if the edited dataframe is actually what they want. Maybe the LLM removed NaN values. Maybe it just deleted some random rows!
The key here: how can users build an understanding of how their data changed, and confirm that the changes made by the LLM are the changes they wanted. In other words, recon!
We've been experimenting more with this recon step in the AI flow (you can see the final PR here: https://github.com/mito-ds/monorepo/pull/751 https://github.com/mito-ds/monorepo/pull/751). It takes a similar approach to the top comment (passing a subset of the data to the LLM), and then really focuses in the UI around "what changes were made." There's a lot of opportunity for growth here, I think!
Any/all feedback appreciated :)
EDIT: Also, shout out to JupyterAI (https://github.com/jupyterlab/jupyter-ai https://github.com/jupyterlab/jupyter-ai) -- it's an official Jupyter project with some really awesome LLM support directly in JupyterLab. I saw it debuted at JupyterCon last week :)
- gventuri 3y agoHey narush, I'm the author of PandasAI. The way PandasAI works is slightly different. It's not the LLM that edits data, but a python script. This is in my opinion the best approach, as it gives a predictable result. There are still some potential issues, including hallucinations, that we are taking care of, but all in all we hope that this could represent an advancement in the way we interact with data. Any questions, feel free to ask :)