Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
adrianbr
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
1.
▲
Show HN: An agent that fixes broken data pipelines from their logs (3min video))
(youtube.com)
4 points
by
adrianbr
3mo ago
|
0 comments
2.
▲
Show HN: We built a LLM-native workflow with 3500 scaffolds and a debug app
(dlthub.com)
1 points
by
adrianbr
1y ago
|
0 comments
3.
▲
Show HN: SQL for APIs. Iceberg under the hood, no warehouse on the bill
(colab.research.google.com)
2 points
by
adrianbr
2y ago
|
0 comments
4.
▲
by
adrianbr
2y ago
The whole idea is pretty nuts! I can imagine it being used as the dev env for teams that use SQLMesh and can thus port the sql. Might be worth investigating with them
5.
▲
Show HN: Automated data extraction from REST APIs in Python
(colab.research.google.com)
2 points
by
adrianbr
2y ago
|
0 comments
6.
▲
by
adrianbr
2y ago
That's really cool, did you already saw the dlt library? That one's done for very easy to use EL in python. It's similarly modular and built by senior data engineers for the data team, and the sources are generators which you
7.
▲
by
adrianbr
2y ago
congrats on the hard work and this launch!
8.
▲
by
adrianbr
3y ago
ahh good old manual fine tuning and maintenance. We are adding data contracts for things like event ingeston where schema needs to be strict or cases where you know ahead of time what to expect. Our experience comes from startups that usual
9.
▲
by
adrianbr
3y ago
Thank you! That's the example we looked at for our dlt-airflow integration :) the dlt dag becomes an airflow dag.
10.
▲
by
adrianbr
3y ago
Duckdb is analytical and gained popularity with the analytics crowd. it has multiple features that make it play well with use cases in that ecosystem such as aggregation speed, parquet support, etc
11.
▲
by
adrianbr
3y ago
for now :) Thanks for pointing it out - and it looks like we should add an aws lambda guide too :) If you want to deploy to lambda, try asking in the slack community, some folks there do it. Or if you wanna try yourself, here is a similar g
12.
▲
by
adrianbr
3y ago
Thank you for the heads up! it is unfortunate, and with 3 letter acronyms this will happen. An easy way to remember is that we are the one you can pip install and plays well in the ecosystem. Databricks has interesting choice in marketing n
13.
▲
by
adrianbr
3y ago
This is amazing! to figure out the website apis has always been a huge pita. With our dlt library project we can turn the openapi spec into pipelines and have the data pushed somewhere https://www.loom.com/share/2806b87
14.
▲
by
adrianbr
3y ago
There are multiple ways to run together - we will show a few in a demo coming out soon. We also consider a tighter integration like with Airflow described here as a possible next step https://dlthub.com/docs/walkthrough
15.
▲
by
adrianbr
3y ago
you can request a source or a feature by opening an issue on sources/dlt repo https://github.com/dlt-hub
16.
▲
by
adrianbr
3y ago
Since dlt generates a schema, and tracks evolution etc, contains lineage, and follows data vault standard it can easily provide metdata or lineage info to the other tools. At the same time, dlt is a pipeline building tool first - so if peop
17.
▲
by
adrianbr
3y ago
Thank you for the feedback! I can see now how it could be confusing. The reason we used chatgpt is because it's an easy starting point - why read through examples when you can get the one you want in seconds? Because dlt is a library,
18.
▲
by
adrianbr
3y ago
dlt is a python library that you can probably plug into the OpenRefine java application to enable moving the data somewhere easily and into different formats, making OpenRefine more useful in a connected environment. I would not say they ar