Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mildbyte
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
18 ms
·
61.
▲
by
mildbyte
6y ago
Currently we control and set up all the FDWs (well, an orchestration layer does it on the fly as the query comes in and routes the query to the correct schema with foreign tables). You can also run a Splitgraph engine locally and add your o
62.
▲
by
mildbyte
6y ago
We use pglast [0]: it's basically a Python wrapper around Postgres's query parsing code. [0] https://github.com/lelit/pglast
63.
▲
by
mildbyte
6y ago
The Splitgraph core code on GitHub [0], around which we've built the DDN, is all about managing "data images" which are basically snapshots of PostgreSQL schemata. You can build them with a format similar to Dockerfiles as we
64.
▲
by
mildbyte
6y ago
Absolutely, having ability to download the actual data and keep it is always going to be important. We want to facilitate access to data and think it should be available from the source. But, there will inevitably be fragmentation, so it&#x
65.
▲
by
mildbyte
6y ago
We currently limit all queries to 30s of execution and 10000 rows returned (by adding a `LIMIT` clause to queries that don't have it). We also have some mechanisms like query result caching and rate limiting for better QoS. One of our
66.
▲
by
mildbyte
6y ago
Co-founder here. The 63-char limit still applies (we didn't recompile Postgres!) but we have some code in front, embedded in a layer of PgBouncers, that intercepts the query, parses it and rewrites it into a shorter dataset ID hash tha
67.
▲
The Splitgraph Data Delivery Network – query over 40k public datasets
(splitgraph.com)
297 points
by
mildbyte
6y ago
|
95 comments
68.
▲
by
mildbyte
6y ago
We hit it in real time for most datasets. A lot of government open data portals are powered by Socrata [0] and we wrote a foreign data wrapper that translates the query into their proprietary query language. If it's a dataset that we h
69.
▲
by
mildbyte
6y ago
Do you mean to run or to buy? To run: this is very lean (the whole stack, including our public website, our REST API etc is currently running on a ~60EUR/pcm Scaleway instance). This is because for most datasets we proxy queries to ups
70.
▲
by
mildbyte
6y ago
Thanks for the report! Our error reporting could definitely be more informative here. I've just looked at the logs and in this case the problem is that the upstream government data portal ( https://data.brla.gov/ ) is te
71.
▲
Show HN: Splitgraph DDN – Public PostgreSQL proxy to 40k+ datasets
(splitgraph.com)
30 points
by
mildbyte
6y ago
|
11 comments
72.
▲
Using Docker Compose in Production
(splitgraph.com)
2 points
by
mildbyte
6y ago
|
0 comments
73.
▲
Supercharging dbt with Splitgraph: versioning, sharing, cross-DB joins
(splitgraph.com)
1 points
by
mildbyte
6y ago
|
0 comments
74.
▲
Querying 40k Datasets with SQL
(splitgraph.com)
4 points
by
mildbyte
6y ago
|
0 comments
75.
▲
Towards a Data Delivery Network
(splitgraph.com)
3 points
by
mildbyte
6y ago
|
0 comments
76.
▲
by
mildbyte
6y ago
It varies depending on how the user chooses to structure storage (we're flexible with that) and what mode of querying they use. We have a more in-depth explanation and some benchmarks in an IPython notebook at [1]. We store Splitgraph
77.
▲
by
mildbyte
6y ago
Shameless plug (I'm a co-founder) but this is basically what we've built with Splitgraph[0]: we can add change tracking to tables using PostgreSQL's audit triggers and let the user switch between different versions of the tab
78.
▲
Write tests. Not too many. Mostly integration
(splitgraph.com)
2 points
by
mildbyte
6y ago
|
0 comments
79.
▲
Treat your datasets like cattle, not pets
(splitgraph.com)
1 points
by
mildbyte
6y ago
|
0 comments
80.
▲
by
mildbyte
6y ago
Splitgraph co-founder (and post author) here. Most PostgreSQL clients don't treat foreign tables any differently than real ones, but note that things like FK constraints or triggers will have to be resolved on the remote server (your a
81.
▲
by
mildbyte
6y ago
There's some query planner tweaks you can use to speed up JOINs with FDWs [0]. In layered querying [1], we had an issue with the planner choosing nested loop joins (which essentially run as multiple small single-row fetches) which tank
82.
▲
by
mildbyte
6y ago
Splitgraph co-founder here. cstore_fdw is great (we even use it in Splitgraph to store data [0])! It's not as fast as purpose-built columnar stores like MonetDB. However, it plugs seamlessly into PostgreSQL and supports all types, even
83.
▲
by
mildbyte
6y ago
You always have the latency/bandwidth overhead from moving queries/data between instances, but FDW performance can be surprisingly fast. There's a performance-FDW complexity spectrum and you can choose a point on it that'
84.
▲
by
mildbyte
6y ago
Splitgraph co-founder (and post author!) here. By far the biggest problem with foreign data wrappers is that you're still forced into PostgreSQL's format of treating and returning each tuple separately. There's some research
85.
▲
Using Make to build multiple Docker images efficiently
(splitgraph.com)
1 points
by
mildbyte
6y ago
|
0 comments
86.
▲
Foreign data wrappers: PostgreSQL's secret weapon?
(splitgraph.com)
4 points
by
mildbyte
6y ago
|
0 comments
87.
▲
It took 10 minutes to add support for DataGrip to Splitgraph
(splitgraph.com)
1 points
by
mildbyte
6y ago
|
0 comments
88.
▲
by
mildbyte
6y ago
That does sound very ambitious for now! I discussed below that we're focused on the OLAP use case (manage actual data, not DDL around it), so triggers, indexes and functions that you create won't be stored in the Splitgraph image
89.
▲
by
mildbyte
6y ago
There aren't many places where Splitgraph intersects with Dolt. Dolt aims to build a database from the ground up to have Git semantics and a real commit graph, whereas Splitgraph works on top of an existing RDBMS (PostgreSQL) and perfo
90.
▲
by
mildbyte
6y ago
Thanks, glad you like it! Theoretically, yes, you can use Splitgraph as a PostgreSQL replication client and then occasionally run a commit to produce a new delta. But we're currently focused on the OLAP use case and so Splitgraph image
More ›