Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aazo11
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
aazo11
2y ago
Currently does not but looking to add support. Would love to connect and learn more about your use case.
32.
▲
by
aazo11
2y ago
Sure will reach you out. Currently Dataherald blocks DML or DDL commands from being generated/executed.
33.
▲
by
aazo11
2y ago
This is not a text to semantic layer but it does far more than just inject schema into the prompt: - the engine keeps an updated catalog of the data (low cardinality columns, their values etc) - taps into query history and finetunes the mod
34.
▲
by
aazo11
2y ago
The agent is LLM agnostic and you can use it with OpenAI or self-hosted LLMs. For self hosted LLM we have benchmarked performance with Mixtral for tool selection and CodeLlama for code generation.
35.
▲
by
aazo11
2y ago
The agent currently executed the generated SQL (limited to 10 rows) and recovers from errors.
36.
▲
by
aazo11
2y ago
Yes. There is a schema linking step which identifies relevant columns and tables including any foreign key relationships if they exist. The agent also can be finetuned on sample NL <> SQL pairs or they can be used in few shot promptin
37.
▲
by
aazo11
2y ago
The entirety of the codebase is now open source.
38.
▲
by
aazo11
2y ago
Hi -- the license is Apache 2.0
39.
▲
Show HN: We open sourced our entire text-to-SQL product
(github.com)
464 points
by
aazo11
2y ago
|
144 comments
40.
▲
by
aazo11
2y ago
We use Stainless. Great team and product.
41.
▲
Did the "modern data stack" not deliver?
(benn.substack.com)
8 points
by
aazo11
3y ago
|
1 comments
42.
▲
by
aazo11
3y ago
Companies have invested a lot on data tools and data stack. It seems the ROI is either not there or hard to measure and budgets are being cut, or being repurposed to collecting data for LLM training/fine-tuning.
43.
▲
Text-to-SQL that asks the LLM to predict the result set
(arxiv.org)
2 points
by
aazo11
3y ago
|
1 comments
44.
▲
by
aazo11
3y ago
I work in the field and have read a good number of papers on the topic. All existing approaches send DDL commands, semantic layer data, few shot samples etc. This one actually sends a version of the DB. First time I have seen this approach.
45.
▲
by
aazo11
3y ago
While the agent does execute the query to recover from errors, the SQLalchemy call execution is limited to a few rows only so if the server is locked there is probably something else going on ;)
46.
▲
by
aazo11
3y ago
We do store golden sql, query history and schemas. Of course everything is encrypted at rest and in transit.
47.
▲
by
aazo11
3y ago
Glad you have been liking it. Feel free to reach out at amir (at) dataherald.com if you need any additional help setting up.
48.
▲
by
aazo11
3y ago
All DML commands are blocked by the engine. You can wrap the returned SQL in a CTE only passing the rows the customer is allowed to access.
49.
▲
by
aazo11
3y ago
Hi -- we do use Fine-tuning together with RAG. To get best in class performance for NL-to-SQL you definitely need to combine both. The good folk at OpenAI dove into this during the last dev-day https://youtu.be/ahnGLM-RC1Y?s
50.
▲
by
aazo11
3y ago
The largest we have successfully deployed is on the OSQuery schema https://osquery.io/ which is 277 tables and lots of business context (malwares, vulnerabilities, Windows registry keys, etc).
51.
▲
by
aazo11
3y ago
For large databases, LLMs do not perform well if you pass the entire schema (either run into context window issues or confuse the LLM with too much info). There is a schema linking step that identifies the relevant schema and only passes th
52.
▲
by
aazo11
3y ago
How does this compare with LlamaIndex + Langchain SQL agents? It seems these tools still are not ready for production use cases, especially since the data gets sent to OpenAI.
53.
▲
by
aazo11
3y ago
Open source LLMs are catching up with OpenAI. They are also considerably cheaper in the long run. We took a look at how they performed in terms of accuracy vs. OpenAI in an NL-to-SQL workload. tldr -> For zero-shot prompting, the closed
54.
▲
by
aazo11
3y ago
Is langsmith still useful if your app is not built with Langchain? Given how early these technologies wouldn't it be better to use a framework agnostic tool to not be tied in to a single framework?
55.
▲
by
aazo11
3y ago
We have been using Langchain since the beginning. We recently deployed Langsmith and it has helped us a lot with monitoring our app in production. Specifically it helped uncover a very nasty bug which was driving up our LLM costs.
56.
▲
Natural Language to SQL on real estate data
(share.streamlit.io)
5 points
by
aazo11
3y ago
|
1 comments
57.
▲
by
aazo11
3y ago
Question answering with an LLM on relational data is hard because the LLM does not know the business and data context. Current approaches are to use few-shot prompting and instructions. This is a sample set up showing how to set up with US
58.
▲
by
aazo11
3y ago
The mathematical representation of algorithms is now pretty standard undergrad stuff. Why is relational algebra not better known?
59.
▲
Ask HN: Why Use an OLAP Database?
3 points
by
aazo11
3y ago
|
0 comments
60.
▲
Ask HN: Pointers to resources for an internal security review
3 points
by
aazo11
3y ago
|
1 comments
More ›