Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aazo11
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
61.
▲
by
aazo11
3y ago
Same here. I hit "run" but get an error
62.
▲
by
aazo11
3y ago
Most RAG approaches use a vectorDB and embeddings for schema linking. In this case the fine-tuning is handling schema linking and there is no vectorDB.
63.
▲
by
aazo11
3y ago
In theory one could create domain specific (or industry specific) templates for data. However coming up with a universal structure might be challenging since data is so varied. Since the issue is often the context, plugging in data dictiona
64.
▲
by
aazo11
3y ago
A tutorial on how to fine-tune a GPT3.5 model for Natural Language to SQL tasks and a comparison of its performance vs Retrieval Augmented Generation. Based on the results, fine-tuning can match and outperform RAG (the approach matches the
65.
▲
by
aazo11
3y ago
Hi -- the /question endpoint does return the generated SQL under the sql_query field of the JSON response: https://dataherald.readthedocs.io/en/latest/api.question.htm... You can even see the entire chain of
66.
▲
by
aazo11
3y ago
Hi -- thanks for the kind words. While right now the engine only works with OpenAI, swapping out other LLMs simply from the envars is on the roadmap. Since we use Langchain for LLM calls this logic is already abstracted away so it will be a
67.
▲
by
aazo11
3y ago
Steve from Hegel-AI had showed us Vanna. Looks super cool. Let us know if you think there is room to collaborate.
68.
▲
by
aazo11
3y ago
Will shoot you an email to connect and chat further.
69.
▲
by
aazo11
3y ago
Thanks for the feedback regarding the Readme. We have some additional detail in the docs here https://dataherald.readthedocs.io/en/latest/index.html (including an architecture doc), but will definitely be incorpor
70.
▲
by
aazo11
3y ago
Thanks. Let us know if we can help out in any way. We just set up a Discord server you can join as well https://discord.gg/A59Uxyy2k9
71.
▲
by
aazo11
3y ago
Depends what you mean by 'work.' You can definitely use them, but GPT-4 is the only model that generates good SQL without lots of training + fine-tuning (and it is very difficult/impossible to come up with Nl <> SQL tra
72.
▲
by
aazo11
3y ago
Sort of. Having perfect data engineering is a requirement if you want to connect an LLM straight to your data warehouse. For real world scenarios, you need a way to add context over time (including examples of how to answer questions from m
73.
▲
by
aazo11
3y ago
SQL is definitely an old technology. However it is still the main language for interfacing with structured data, even for newer tools like Clickhouse. Down the road it is conceivable that it will be replaced with something else, but the cos
74.
▲
by
aazo11
3y ago
Amazing! Feel free to shoot me a line at amir at dataherald
75.
▲
by
aazo11
3y ago
We are looking to partner with other open source projects. Let's connect. Shoot me a line at amir at dataherald.com
76.
▲
by
aazo11
3y ago
Interesting. So far we have been focused on allowing non-technical users to self-serve from enterprise data warehouses. It might be possible to enable this scenario with a simple addition to the NL-2-SQL engine to translate the generated SQ
77.
▲
by
aazo11
3y ago
Totally agree. In the hosted version built on top of the engine we block the answer from going to the question asker until an authorized user 'verifies' the answer. However these 'verified' answers are then stored in the
78.
▲
by
aazo11
3y ago
TLDR -- We currently do not automatically load stored procedures as context for in context learning/prompting but it is on the roadmap. Additional color: Once you configure a connection to a DB, you can trigger a 'scan' of th
79.
▲
by
aazo11
3y ago
Hi -- would love to connect and discuss more. Shoot me an email amir at dataherald.com.
80.
▲
by
aazo11
3y ago
Hi Dylan -- new GPT-4 class LLMs have gotten good at writing correct SQL, but while the SQL they generate almost always executes correctly they often write SQL that generates incorrect answers to the question. Some reasons for this can be t
81.
▲
Show HN: Dataherald AI – Natural Language to SQL Engine
(github.com)
215 points
by
aazo11
3y ago
|
103 comments
82.
▲
by
aazo11
3y ago
This is great. Thanks for the contributions on the LangChain discord as well.
83.
▲
by
aazo11
3y ago
At Dataherald our mission is to make it easy to create data-heavy content. While LLMs like ChatGPT have revolutionized content creation, they still cannot create factual and visually-appealing data-intensive content that many industries suc
84.
▲
Deploying thousands of SEO-optimized Zillow-style landing pages
(pseo.dataherald.com)
2 points
by
aazo11
4y ago
|
1 comments
85.
▲
by
aazo11
4y ago
Programmatic SEO is the practice of creating landing pages at scale to capture long tail search traffic. The solution works especially well if your business deals with lots of data, since you can create unique pages that also update. At Dat
86.
▲
Show HN: No-Code access to all FRED and BLS data
(app.dataherald.com)
7 points
by
aazo11
4y ago
|
0 comments
87.
▲
by
aazo11
5y ago
Amir here. We allow users to download the cleaned data and load it into any visualization tool they want. That being said, we have found a lot of users (especially those who tell stories with data) have a lot of use for having the visualiza
88.
▲
Ask HN: What datasets (public or proprietary) do you use on a regular basis?
7 points
by
aazo11
5y ago
|
1 comments
89.
▲
by
aazo11
5y ago
Question for the HN community: Is there a dataset (free or proprietary) which you regularly use for as part of your work?
90.
▲
by
aazo11
6y ago
Our eventual goal is to remove the need for anyone to do redundant data engineering work on public datasets. So while Superset/Metabase are tools to explore you enterprise data, we are focused on building a growing library of feeds fro
More ›