3 ms·
Are you planning on submitting this model to be evaluated against the Spider holdout set? Also, wondering if anyone has found research on the inverse of this a
by philipodonnell 3y ago
Are you planning on submitting this model to be evaluated against the Spider holdout set?
Also, wondering if anyone has found research on the inverse of this approach to the problem, i.e., instead of training the model to understand the data, you improve the data to be more understandable to a model? This seems more promising when you are looking at enterprise use cases without much training data. Spider seems like quite a simple dataset compared to the ones I encounter on the job, and LLMs struggle even with those.
- MrezaPourreza 3y agoYes, we have already submitted the model for evaluation on the Spider holdout test set. While your suggestion is certainly intriguing, implementing a universal solution could be quite challenging, as it would heavily depend on the specifics of the dataset.
- philipodonnell 3y agoI don’t think it’s necessarily about a “universal” solution, just “better”. Make the column names more verbose, changing numeric enums to text ones, disambiguating column names, etc. One of the spider datasets is a stadium table and one of the column names is “average”, which means average capacity, but it’s super ambiguous. If you asked an LLM to “make these table columns more verbose” I bet it would call that “average_capacity” and all of the sudden some NLQ queries that confused the function and the column name would start to work.
- aazo11 3y agoIn theory one could create domain specific (or industry specific) templates for data. However coming up with a universal structure might be challenging since data is so varied. Since the issue is often the context, plugging in data dictionaries (and passing those to the LLM) can help