Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jtagliabuetooso
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
jtagliabuetooso
5mo ago
How is the Git semantics of merge, rebase and diff defined in the system?
2.
▲
by
jtagliabuetooso
5mo ago
“Why it’s distinctive” is misleading (perhaps LLM-generated)? Imo, we cite other work because it puts our work in context for experts and beginners alike; because it makes clear that we all stand on someone else’s shoulders (progress is,
3.
▲
by
jtagliabuetooso
5mo ago
Cool release. IMO, "Why it's distinctive" is a bit misleading on a few points: certainly the dbt and DX folks can add their POV, but even considering stuff I know / authored ;-), https://arxiv.org/pdf
4.
▲
by
jtagliabuetooso
5mo ago
Data contracts as types and compile time checks (even across languages) are not new - this is a recent paper exposing the idea of correctness-by-design pipeline, which is a super set of this particular issue obviously (disclaimer: I'm
5.
▲
We solved trust for AI Agents in 1973 (we just forgot)
(bauplanlabs.com)
5 points
by
jtagliabuetooso
9mo ago
|
0 comments
6.
▲
by
jtagliabuetooso
11mo ago
Author here: happy to answer questions on our experience with duckDB / DF as part of a larger system running OLAP workloads.
7.
▲
by
jtagliabuetooso
1y ago
Marimo and astral are great and we use them both, they are not however infrastructure companies, so the parallel is a bit imperfect. Wouldn't you use AWS because it's closed source? And BigQuery? Or Motherduck? There is no "o
8.
▲
by
jtagliabuetooso
1y ago
We are not proposing or advocating for any approach to development (I personally almost never use notebooks these days and run Bauplan with preview). The blog together with our marimo friends is to showcase that you can have notebook develo
9.
▲
by
jtagliabuetooso
1y ago
Hey Ben, thanks for your message. We have people building stuff featured here ( https://www.bauplanlabs.com/build-with-bauplan ) as well as online (e.g. https://blog.det.life/bauplan-the-serverless-data-lakeho
10.
▲
by
jtagliabuetooso
1y ago
Thanks for the comment: your frustration is the default in the industry, and it's part of the reasons why Bauplan was built. "but it always takes forever and I'm never even able to articulate why." -> there are way mo
11.
▲
by
jtagliabuetooso
1y ago
Importantly, kedro does not run things for you, resulting in a suboptimal experience because the runtime and dsl are separated: in particular, it does not solve the problem of having K different systems with scattered logs and not easy to
12.
▲
by
jtagliabuetooso
1y ago
Thanks for your comment. As stated elsewhere, we understand the need for people to know how the system works, and have contributed back our ideas (and quite a bit of open source code) to the community: if you want to check our blogs and &#x
13.
▲
by
jtagliabuetooso
1y ago
Spark is technically not Python, even if we support PySpark with the relevant decorator but it's a very niche use case for us. As for all the other Python packages, including proprietary ones, the FaaS model is such that you can declar
14.
▲
by
jtagliabuetooso
1y ago
Thanks for checking out bauplan (which also supports BYOC, so I guess it is indeed hostable by you in a sense!). We've done quite a lot of open source in our life, at Bauplan (you can check our github), and before (you can check me ;-)
15.
▲
by
jtagliabuetooso
1y ago
Yeah, terms are confusing sometimes! "Data lakehouse" is weirdly enough a "technical term". The canonical reference is from CIDR https://www.cidrdb.org/cidr2021/papers/cidr2021_paper17.pdf , but
16.
▲
by
jtagliabuetooso
1y ago
Mhmm, it doesn't resolve to empty but the full SecureFrame monitoring: https://security.bauplanlabs.com/#resources-b2152df0-4179-48... - if you wait a second, this is the entire report: https://www.loom.com&
17.
▲
by
jtagliabuetooso
1y ago
You mean on the data side? Data access in the example (and in real-world) is mediated by production-grade Iceberg compatible catalog, sandboxed changes, and full auditability trail ( https://docs.bauplanlabs.com/en/lates
18.
▲
by
jtagliabuetooso
1y ago
Thanks for the feedback. Bauplan actually features a few innovative points in this area, and full Pythonic at that: Git for Data ( https://docs.bauplanlabs.com/en/latest/concepts/git_for_data... ) to sandbox an
19.
▲
by
jtagliabuetooso
1y ago
Hey, founder of Bauplan here. Happy to field any questions or thoughts. Yes, marimo is great, and it's the only way to work within a real Python ecosystem for production use cases shipping proper code.
20.
▲
by
jtagliabuetooso
1y ago
Awesome, you can write me anytime to geek out (jacopo.tagliabue@bauplanlabs.com) or follow us for more community sharing (papers, deep tech blog posts: https://www.linkedin.com/in/jacopotagliabue/ ). The full audit
21.
▲
by
jtagliabuetooso
1y ago
Sorry for the confusing example. So, the AWS lambda in the data product example is a bit of a red herring, and it's used as the outer process to create branches and launch bauplan pipelines through the Python client ( https:/&#x
22.
▲
by
jtagliabuetooso
1y ago
Thanks for your interest! Aside from the demo video in the home page, our quick start takes <3 minutes, which is way less! Just ask for an invite to the free sandbox in our website. If you love videos and would like to understand the dec
23.
▲
by
jtagliabuetooso
1y ago
Glad it resonates! Happy to help answer any question you may have - the sandbox in our home page is free to try: just ask for an invite!
24.
▲
by
jtagliabuetooso
1y ago
"In the notebook I'll typically try to replicate (as close as possible) the state of the data inside some intermediate step, and will then manually mutate the pipeline between the original and branch versions to determine how the
25.
▲
by
jtagliabuetooso
1y ago
Glad to see it resonates, especially the Python part <3
26.
▲
by
jtagliabuetooso
1y ago
Glad to see it resonates! Happy to help if needed!
27.
▲
by
jtagliabuetooso
1y ago
The pricing is a bit bespoke at the moment as we work closely with our customers - you can reach out at anytime to any of us for a chat (jacopo.tagliabue@bauplanlabs.com). The general driver is just compute capacity: how much resources you
28.
▲
by
jtagliabuetooso
1y ago
Thanks for the question! On the data side of things, DVC is more about versioning static datasets / local files, while Bauplan manages your entire lakehouse, potentially hundreds of tables with point in time versioning (time travel) an
29.
▲
by
jtagliabuetooso
1y ago
Correct. Unlike warehouses or SQL lakehouses, we also any Python code, including from your private AWS repositories for example, through a simple decorator, while giving you transactional pipelines, fully versioned and revertible, like it&#
30.
▲
by
jtagliabuetooso
1y ago
You technically just need storage (files in a bucket you own and control forever). We bring you the compute as ephemeral functions, vertically integrated with your S3: table management, containerization, read / write optimizations, per
More ›