3 ms·
I'm not sure about this approach. From what I have seen, most researchers have no idea how to get their data in a format which can be efficiently analysed. Onc
by davidktr 3y ago
I'm not sure about this approach. From what I have seen, most researchers have no idea how to get their data in a format which can be efficiently analysed.
Once you have that, it's trivial to do any kind of statistical analysis. In R, a regression is simply lm(y ~ x1 + x2 + ... + xn).
You can always look up how an API works, but thinking about data in terms of structures is what hinders effective analysis in most cases.
- cl42 3y agoTotally appreciate the feedback and I agree with you that a well structured data set can be trivially analyzed. Heck, at that point you can use drag and drop stats packages too. The data set I used for the demo has strings for income categories and a mix of categorical variables that the LLM had to transform, which is incredibly promising. The insights that Claude generated also imply that it can do follow-up analysis. This is less of a “hey write my regression code for me” and more of a “suggest the analysis, do it, find insights, and run follow up analyses”. That’s way more powerful and interesting.
- davidktr 3y agoGood points, and I really appreciate your work. You are addressing a real problem. I'm simply skeptical that someone who can make use of Claude (or any LLM) for data analysis is in need of making use of it. Let's hope I'm too pessimistic here.
- cl42 3y agoOooh, if that’s your concern then give me a few months to launch a product. :-)
- reacharavindh 3y agoAnecdotally, my wife - a researcher in management accounting who does a lot of analysis of corporate data was very excited about this tool because it allows her to explore the dataset in almost natural language and have a starter Python code base to tinker with. I have seen her use Python. She uses it like a research notebook. Sequential pipeline like analysis steps. Any little change to a step, and she will run the whole thing :-)