Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
akarve
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
akarve
9y ago
We welcome your contributions to the R interface. If you email me, aneesh at quiltdata dot io, I can add you to our Slach channel where our engineers can support your efforts. Several users have asked about R and if we combine them together
62.
▲
by
akarve
9y ago
Those are both interesting projects. Quilt has a rather different emphasis, though. The difference from CVMFS is that Quilt is a full set of services around data (build, push, install) whereas CVMFS is just the file system. In the big data
63.
▲
by
akarve
9y ago
We've talked about this. Though streams (e.g. Apache beam) are closer to where we think realtime data is going. It would be possible to wire Quilt to something like firebase to get realtime behavior... Happy to brainstorm other solutio
64.
▲
by
akarve
9y ago
Parquet is stationary data on-disk. Arrow is focused on in-memory analytics and serializes out to a variety of formats (including Feather and Parquet).
65.
▲
by
akarve
9y ago
Java has the right idea. Namespace by [reverse] domain name. It's sadly a tad verbose though. There's high contention for short names which is at odds with uniqueuness.
66.
▲
by
akarve
9y ago
It's not about people not being able to load their data, but about accelerating the loading with serialization and about whether or not people want to focus on data cleaning, or have the cleaning done once and then available for poster
67.
▲
by
akarve
9y ago
There are some important differences from git: https://news.ycombinator.com/item?id=14772036
68.
▲
by
akarve
9y ago
We are rolling out on-prem where the customer runs the package registry on their own infrastructure. The set up process takes a bit of customization depending on your environment. If you have Docker containers it's easier. If you email
69.
▲
by
akarve
9y ago
I know only a bit about Synapse. It seems like they have some valuable data. I think the biggest differences are in the culture and user experience. Quilt gets its DNA from the open/community-driven cultures of GitHub and npm. We want
70.
▲
by
akarve
9y ago
Gigabytes should be no problem. We have yet to implement multi-part uploads though, so if your network gets interrupted you will need to re-upload the package from scratch. By design when the upload is in progress your client is talking dir
71.
▲
by
akarve
9y ago
I agree wholeheartedly with this statement: "application layer should be responsible for storage management and performance tuning". I would take it one step further and say that storage should be virtualized in a high-performance
72.
▲
by
akarve
9y ago
Done. You should see a more obvious search bar on the next push. We'll also make search case insensitive :)
73.
▲
by
akarve
9y ago
Is there an example in the docs that just shows `examples.sales`? If so please let me know and I'll fix it. I searched and couldn't find such an example (though maybe, through the "magic" of Chrome's service-worker,
74.
▲
by
akarve
9y ago
You can absolutely edit in place and that will go into `quilt log` for the package--as long as you are the package owner. Our docs were a bit confusing on this point. I just updated them: https://docs.quiltdata.com/edit-a-pa
75.
▲
by
akarve
9y ago
We charge business and on-prem users in TB-sized blocks. So that part is variable cost, not flat. And we sell user seats in blocks of 10. What else should we be thinking about? We want to be fair and also price in a way that encourages shar
76.
▲
by
akarve
9y ago
We really like pachyderm and know the founders. Quilt is zero-config focused on storage and versioning, pachyderm is more focused on on-prem and compute (as a Hadoop replacement).
77.
▲
by
akarve
9y ago
Hi. R is really important to us and we want to add support (probably through SparklyR). If you email me I can add you to our Slack channel and we can talk through extending Quilt (aneesh at quiltdata dot io). There is a sliver in our docs b
78.
▲
by
akarve
9y ago
We support arbitrary data formats in that Quilt falls back to a raw copy if it can't parse the file. On the columnar side (things we convert to Parquet) we support XLS, CSV, TSV, and actually anything that `pandas.read_csv` can parse.
79.
▲
by
akarve
9y ago
Understood. We can do something very close to that. Right now we give the registry source to our on prem users as part of the license. Email sales at quiltdata dot io and we can discuss.
80.
▲
by
akarve
9y ago
Data cleaning is so necessary. `build.yml` already supports a limited set of feature (through pandas). In addition to custom data transformations, any "out of the box" cleaning functions you'd like? In the spirit of dplyr? We
81.
▲
by
akarve
9y ago
The package name + hash is an implicit DOI. What if we added web support for it so that users could https://quiltdata.com/packages/USER/PKG?doi=SOME_HASH ?
82.
▲
by
akarve
9y ago
The client is fully open source. You can indeed run your own and we are just starting to roll that out. I can get you started: feedback at quiltdata dot io. We are deliberating open sourcing the registry as well (making everything open sour
83.
▲
by
akarve
9y ago
Hmm. We were able to get the pip handle so didn't see major conflicts in our target space. Are there places in code/cli where we could name conflict with the patch sets manager?
84.
▲
by
akarve
9y ago
As to your slicing question, yes. Data is lazily loaded. With Parquet as our data store we can do even more (but haven't yet): load only the columns referenced.
85.
▲
by
akarve
9y ago
Four things: serialization, virtualization, querying, big data (even Git LFS isn't super performant for large files). Quilt actually transforms data into Parquet and wraps it in a virtualization layer so that the data can be injected d
86.
▲
by
akarve
9y ago
Thanks. Where do you see it going in 2-5 years? Sometimes outsiders see things about the future that we haven't thought about :)
87.
▲
by
akarve
9y ago
To start with, we can add stars (pay with prestige). Getting more into science fiction--but very possible science fiction--we can put data on the blockchain and let people transact. The data owner would get the lion's share of the tran
88.
▲
by
akarve
9y ago
Dat is a distributed transport layer for raw data. Quilt is a centralized (your infrastructure or ours) transport and consumption layer for virtualized data. As such we'll be able to, for example, run efficient queries across all of Qu
89.
▲
by
akarve
9y ago
As Kevin mentioned we can extend support to frictionless (and are acceptign PRs on GitHub :). The thing we didn't love about frictionless is that it requires the user to fully specify the schema. We take a slightly more automated appro
90.
▲
by
akarve
9y ago
Yes :) Happy to discuss in detail if you have specific types of abuse in mind. Our first line of defense is to get help from the community through downvoting useless content.
More ›