5 ms·
I like the concept, but I'm not sure I like the implementation. The demo they show has an excel sheet source being transformed into another sheet by grouping b
by mpeg 2y ago
I like the concept, but I'm not sure I like the implementation.
The demo they show has an excel sheet source being transformed into another sheet by grouping by a specific column. I don't like how it does a transformation step, sends a new transformed file and then switches the display to this new file. I worry that this loses the history of what exactly has happened to the file, so you're letting a black box run an unknown transformation and trusting the result to be correct.
A better implementation would require a custom UI for this, similar to existing data wrangling tools where every action is logged so that it's clear what has happened to the file and the steps can be checked transparently and rolled back if needed. As it exists now I would find it hard to trust it
- ramoz 2y agoFile lineage doesn’t seem like a complex feature addition here.
- mpeg 2y agoYes and no, apart from the UI for it it really depends on how they've implemented the transformations. For example, if it's writing and running python behind the scenes which I suspect is what is happening here as that's what the old ChatGPT data analysis did that would not work very well without having to read the code which defeats the point of it being no-code... You'd want deterministic steps, like it would need to have a limited list of functions to choose from that do specific things eg "group_by()" "filter()" etc so that the code it runs is always the same for it.
- StarlaAtNight 2y agoMight make sense for them to anchor to ibis for the code part (since it compiles down to SQL or Pandas) - and, being inspired by tidyverse, could easily translate to R
- bhl 2y agoTo collaborate with LLMs with text editors and spreadsheets, yeah I do think we will need deterministic declarative primitives that both users and LLMs can use. Getting those primitives right would also get us generalized autocomplete for apps.
- zarathustreal 2y agoV0 by Vercel does essentially this using the RadixUI components as the declarative primitives to generate components
- bhl 2y agoFor a generative and collaborative web app, you would have to hook up those declarative components to function calls which are also declaratively written and perhaps use OT/CRDTs to mutate state.
- andy99 2y agoGoing to turn out like every other no code / low code tools, works great for a demo of something you'd never do, customizing to a real application is just as much work as coding it
- hermitcrab 2y agoNo code / low code tools are widely used in industry for ETL, data wrangling and data analysis. It's a multi billion dollar industry. Of course, they aren't the best tool for every job.
- Hansenq 2y agoThe way ChatGPT Data Analysis works is that ChatGPT generates python code that does the transformation step. Because python code is static and deterministic, you can always re-run the code to get the exact output over and over again. If you want it rolled back, just re-run the code, but run fewer lines of it. In fact, this is probably more accurate than custom UIs that log actions you take, since to build that UI, engineers need to make sure each action is logged (and it's easy to forget to log one!).
- anotherpaulg 2y agoThis is why I like to use LLMs to write the code to manipulate my data, draw my graphs, etc. That way I end up with the data/graph/etc artifact, but I also have the code that created it tucked away safely in my repo. So I can tweak and improve that code over time, either on my own or with the help of AI. Here's an example where I recently used aider and GPT-4o to plot a graph. https://aider.chat/2024/05/13/models-over-time.html https://aider.chat/2024/05/13/models-over-time.html The graph itself is kind of interesting. It shows how LLM code editing skill has been changing over time as new models have been released by OpenAI, Anthropic and others.
- krainboltgreene 2y ago> This is why I like to use LLMs to write the code to manipulate my data One of the nice things about human programmers is that you can derive intent and create responsibility. Sometimes we have to encode that in a commit message or ticket, but it's there. When you find out that subtly the program was piping all data into devnull without realizing it you at least have a human to figure out how you got here. Another example of this is the xz debacle. What are you supposed to do when that's a thousand layers deep? What about when your next generation of programmer's ability to do things is exceptionally stunted by this effort?
- weitendorf 2y agoI'm explicitly working on this in my startup's product (a GenAI for code product). The obvious answers: record the human's intent in the form of their prompt, and record the LLM's raw output (if you use a conversational LLM out of the box, it almost always includes this even if you explicitly prompt it not to, lol). Of course, depending on your UX this may or may not work. For autocomplete there is no obvious user intent. There are additional approaches which I'm exploring that require more intentional engineering, but essentially involve forcing "structure" so that more intent gets explicitly specified.
- krainboltgreene 2y agoTo be clear you're not working on what I pointed out, you're just doing the same thing. The prompt "may" encode intent, but that has no bearing on what code gets written or stored or changed. Think about this as well: You're creating processes that have no reasonability. I am 100% responsible for all code I write, even if I wrote it wrong, but if the code your tool generates is wrong it's not my fault, I didn't write it. Multiply this by the hundreds of thousands of times this will happen in a given year and by each employee. Frankly, you should re-evaluate if you even want your product in the world. What kind of future hellscape are you enabling?
- mritchie712 2y ago(note: I'm founder of Definite) I think we're closer to what you're describing. All transformations in Definite[0] are done with SQL which you can easily inspect. The transformations are stored in your data model so everyone else can reuse them. The setup we normally see at a company is one (or more) data people control the model which can be used by the rest of the company thru our AI assistant. Quick demo here: https://youtu.be/p6BoqX0cYnU?si=hg2V_KM8ScUzwrDx&t=147 https://youtu.be/p6BoqX0cYnU?si=hg2V_KM8ScUzwrDx&t=147 0 - https://www.definite.app/ https://www.definite.app/