6 ms·
Hey HN, We're happy to introduce Dropbase 2.0! It's a tool that helps you bring offline files, such as CSV, Excel, and JSON files, into Postgres database. You
by ayazhan 6y ago
Hey HN,
We're happy to introduce Dropbase 2.0! It's a tool that helps you bring offline files, such as CSV, Excel, and JSON files, into Postgres database. You can also process your data before uploading it using a spreadsheet-like interface or by writing a custom Python script. Once your data is in the database, you can query it using any third party tool (credentials will be provided). You can also access your data via REST API (powered by PostgREST) or create custom endpoints to serve a more specific use case.
A bit about the tech:
Currently, we support .csv, .json, .xls, .xlsx files. For data processing, we use Pandas, so if you are comfortable using Python, you can write your own custom functions to process the data. We also give you a free shared Postgres database to test the tool with (your data is isolated and hidden from others). Each one of these databases come with an instance of PostgREST preinstalled, so you can query your database using REST API ([http://postgrest.org/en/v7.0.0/](http://postgrest.org/en/v7.0.0/) http://postgrest.org/en/v7.0.0/](http://postgrest.org/en/v7....). You can also generate an access token with an expiry date to share your data with others.
There are many more features that we baked into the product. Come check it out, it's open for HN community.
- mtVessel 6y agoQuestion about the TOS: "However, by posting Content using Service you grant us the right and license to use, modify, publicly perform, publicly display, reproduce, and distribute such Content on and through Service. You agree that this license includes the right for us to make your Content available to other users of Service, who may also use your Content subject to these Terms" Does this mean I should have no expectation of privacy or control over anything I upload?
- jimmyechan 6y agoYour data is private and you own all of your data. We do not and will not share your data with anybody else unless you share it yourself through the sharing of projects, pipelines, endpoints, or exports. We do however store and process your data. We also let you generate endpoints so we need some wording to cover these cases. We'll double check our terms to make this point clearer, but we added this because you can generate live endpoints and you can share those.
- rapnie 6y agoUnfortunately IANAL and the formulation of the ToS/PP in your, and that of most other online service providers, always give me that naggy feeling that the legalese leaves so many loopholes, texts open to different interpretation, that effectively - even though it may seem so - I have no privacy guarantees whatsoever. That might be entirely unwarranted of me, but the feeling is there. Unease.
- gwd 6y ago> even though it may seem so - I have no privacy guarantees whatsoever. I mean, fundamentally, really consistent security is hard; and the best you can reasonably expect from someone you're not paying is "best effort". For them to make real promises about security opens them up to being sued if they fail; it's not really reasonable to ask someone to do that unless you're paying them a reasonable chunk of cash to offset that risk.
- rapnie 6y agoSorry, but this almost feels like a GPT-3 response to me. I don't see what security, paid vs. free or best-effort has got to do with my argument, which is that the loopholes in legalese are so hard to spot for anyone but a lawyer, that effectively my data might still be used in any way and possibly against my wishes or expectations (but which becomes legal when I consent to the PP and ToC).
- Oranj 6y agoWe have also stumbled upon this point in the T&C. A real show-stopper for us as we are planning to work with sensitive client data.
- jimmyechan 6y agoWe're working on an on-prem / self-hosted version Dropbase that will better address use cases where sensitive data needs to stay within the company. Would you mind describing the kind of data you work with and which constraints you're subject to?
- 2mol 6y agoLooks very cool so far, congratulations! I have two questions: - How do you handle incremental loads from files (or even google sheets)? Am I able to only load the diff, load full snapshots and get bi-temporality, etc? - Are you supporting PostgREST as sponsors in any way? It's one of the most solid tools I've ever used, and I love to see that companies build great products on top of it!
- ayazhan 6y ago- we use pandas to process data and load it to Postgres using .to_sql (https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to_sql.html https://pandas.pydata.org/pandas-docs/stable/reference/api/p...). for incremental loads, we set "if_exists" to "append". we're working to add more flexibility to load function, so you can specify how to load your data, handle conflicts, and so on. we're open for suggestions - PostgREST is great. We are not sponsors at the moment, but looking to do this when we can!
- hans_castorp 6y agoCan this be installed on-premise?
- reportgunner 6y agoIf not I really don't like how it says "offline files". I can install a database on-prem and upload files to it and then I can too query it as if it was in a database. From what I've seen on the this seems like a fancy cloud database client.
- jimmyechan 6y agoGood point. We can probably work on clarifying our value prop. You could do something like what you described with a local database. We could add that by building some desktop code. At the moment we are only focused on the cloud part. That way we can get data, let you process it, and also easily share as APIs or endpoints.
- jimmyechan 6y agoNot at the moment, but we're actively developing a teams/enterprise version of this that allows for self-hosted/on-prem.
- mikorym 6y agoI've used Pandas recently, not sure if this will ever help you, but dictionaries are much faster if you continually add rows. pandas.concat and similar functions for appending to a Pandas dataframe can be quite slow. Just mentioning this in case you ever encounter this issue; maybe you don't need to. In my case I changed to dictionaries for important parts and 2 mins changed to 5 seconds execution time. However, in my case, I have to change logical row structure and not just read in rows as is.
- ayazhan 6y agothank you for advice! we'll consider it
- ayazhan 6y agoPostgREST links are broken. Here is the correct link: http://postgrest.org/en/v7.0.0/ http://postgrest.org/en/v7.0.0/