Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dangoldin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
dangoldin
2y ago
FWIW - Redpanda open sources their core product - https://github.com/redpanda-data/ while WarpStream keeps their core product proprietary - https://github.com/warpstreamlabs
32.
▲
Building a data pipeline using Snowflake primitives
(blog.twingdata.com)
1 points
by
dangoldin
2y ago
|
0 comments
33.
▲
by
dangoldin
2y ago
You don’t need to. dbt/sqlmesh are competitive. I just like the model of sqlmesh over dbt but dbt is much more dominant.
34.
▲
by
dangoldin
2y ago
From egress + storage cost standpoint absolutely which ends up being a big factor for these large scale data systems. There’s a prior discussion on HN about that post: https://news.ycombinator.com/item?id=38118577 And full
35.
▲
by
dangoldin
2y ago
Author here. Basic idea is you want some way of defining metrics. So something like “revenue = sum(sales) - sum(discount)” or “retention = whatever” which need to be generated via SQL at query time vs built in to a table. Then you can have
36.
▲
by
dangoldin
2y ago
Yes but most data-heavy tasks are parallelizable. SQL itself is naturally parallelizable. There's a reason Apache RAPIDs, Voltron, Kinetica, Sqream, etc exist. Full transparency I don't have huge amount of experience at working on
37.
▲
by
dangoldin
2y ago
Author here and there's nuance here but as a rule of thumb data size is a decent enough proxy. Audience here isn't everyone and the goal was to give less experienced data engineers and folk a sense of modern data tools and a possi
38.
▲
Building an open data pipeline in 2024
(blog.twingdata.com)
103 points
by
dangoldin
2y ago
|
32 comments
39.
▲
Predicting Solar Eclipses with Python
(erikbern.com)
2 points
by
dangoldin
2y ago
|
0 comments
40.
▲
Show HN: Identify unused columns in your database
(github.com)
1 points
by
dangoldin
3y ago
|
0 comments
41.
▲
Identify unused columns in Snowflake and other data warehouses
(blog.twingdata.com)
1 points
by
dangoldin
3y ago
|
0 comments
42.
▲
by
dangoldin
3y ago
That's when your batch jobs are running. Partially kidding but companies will wait until there's lower demand to take advantage of spot pricing.
43.
▲
To CTE or Not to CTE: The Case for Subqueries
(blog.twingdata.com)
1 points
by
dangoldin
3y ago
|
0 comments
44.
▲
The Evolution of a Data Stack
(blog.twingdata.com)
1 points
by
dangoldin
3y ago
|
0 comments
45.
▲
by
dangoldin
3y ago
Not to change your direction but something I've been toying around is being able to support Algebraic types when defining tables. That way you can offload a lot of the error checking to the database engine's type system and keep a
46.
▲
The Curious Case of a Snowflake CTE
(blog.twingdata.com)
3 points
by
dangoldin
3y ago
|
0 comments
47.
▲
Instead of pinching pennies on Snowflake, make dollars
(dansdatathoughts.substack.com)
1 points
by
dangoldin
3y ago
|
0 comments
48.
▲
AI in BI: Rhyming doesn't make it so
(dansdatathoughts.substack.com)
2 points
by
dangoldin
3y ago
|
1 comments
49.
▲
SQL Scoping Is Surprisingly Subtle and Semantic
(buttondown.email)
18 points
by
dangoldin
3y ago
|
1 comments
50.
▲
by
dangoldin
3y ago
Author here but some ideas I was thinking about: - An open source data pipeline built on top of R2. A way of keeping data on R2/S3 but then having execution handled in Workers/Lambda. Inspired by what https://www.boilin
51.
▲
by
dangoldin
3y ago
Author here - have you tried using R2? As others mentioned there's also Sippy ( https://developers.cloudflare.com/r2/data-migration/sippy/ ) which makes this easy to try.
52.
▲
by
dangoldin
3y ago
Author here and it is true that costs within a region are free and if you do design your system appropriately you can take advantage of it but I've seen accidental cases where someone will try to access in another region and it's
53.
▲
by
dangoldin
3y ago
Author here and really cool link to Sippy. I love the idea here since you're really migrating data as needed so the cost you incur is really a function of the workload. It's basically acting as a caching layer.
54.
▲
From S3 to R2: An economic opportunity
(dansdatathoughts.substack.com)
274 points
by
dangoldin
3y ago
|
173 comments
55.
▲
Automating Dead Code Cleanup
(engineering.fb.com)
111 points
by
dangoldin
3y ago
|
29 comments
56.
▲
McKinsey Developer Productivity Review
(dannorth.net)
1 points
by
dangoldin
3y ago
|
0 comments
57.
▲
Product Model at Spotify
(svpg.com)
3 points
by
dangoldin
3y ago
|
0 comments
58.
▲
by
dangoldin
3y ago
Yes - he gives it credit at the bottom of the page.
59.
▲
by
dangoldin
3y ago
Probably a function of what looks to be an AWS Lambda outage.
60.
▲
by
dangoldin
4y ago
I read an interview a while back with a game developer who was asked why video games have historically had so much fighting and he responded with "it's easy to write." Take away is that as AI improves we will move to a world
More ›