7 ms·
What Happens When You Put a Database in the Browser?
- gregw2 2y agoI am not 100% clear how this works... If you query a Parquet file from your lake via DuckDB-in-browser, does DuckDB run in WASM on the web client and pull the compressed parquet to your browser where it is decompressed? Or are you connecting some DuckDB on the web client to some DuckDB component on a server somewhere? I presume yes to the first and no to the second but just checking I have my mental model correct.
- 4b11b4 2y agoThe data would be sent to the browser
- throwaway-blaze 2y agoUm, my data lake is measured in Pb, not Tb. How is that going to work exactly?
- undefuser 2y agoDuckDB has certain optimizations which allows it to read only parts of a parquet file. It can also read remote file in streaming fashion so it does not have to wait for the entire file to be downloaded or to store a large amount of data in memory. Relevant documentation: https://duckdb.org/2021/06/25/querying-parquet.html https://duckdb.org/2021/06/25/querying-parquet.html
- 1egg0myegg0 2y agoHowdy! I work at MotherDuck and DuckDB Labs (part time as a blogger). At MotherDuck, we have both client side and server side compute! So the initial reduction from PB/TB to GB/MB can happen server side, and the results can be sliced and diced at top speed in your browser!
- code_biologist 2y agoPlease spend a sentence or two explaining the server side filtering mechanism and linking to documentation! I would like to know the conditions required for streaming queries! From the sibling comment and a search of the docs it seems like this is a Parquet only feature, which seems pretty important to note!
- justincormack 2y agoNot specific to Ducks but S3 select https://docs.aws.amazon.com/AmazonS3/latest/userguide/selecting-content-from-objects.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/select... can filter Parquet server side on S3 and is supported by some other object stores.
- FridgeSeal 2y agoParquet is designed with predicate push-down in mind. Partitions are laid out on disk, and then blocks within files are laid out so that consumers can very, very easily narrow in on which files they need to read, before doing anymore IO than a list, or a small metadata read. Once you know what you are reading, many parquet/arrow libraries will support streaming reads/aggregations, so the client doesn’t need to load the whole working set in memory.
- LunaSea 2y agoThis only covers very simple min / max / sum cases. For all others you'll need to download all columns you are filtering or selecting.
- victor106 2y agoDoes duckdb work with delta files?
- simonw 2y agoDuckDB can use the HTTP range header trick agains parquet, which means a lot of analytical questions against multi-GB files can be answered by fetching only small portions of the overall data. Here's a post about applying this same trick to SQLite from few years ago: https://phiresky.github.io/blog/2021/hosting-sqlite-databases-on-github-pages/ https://phiresky.github.io/blog/2021/hosting-sqlite-database...
- zX41ZdbW 2y agoI have an example of a database hosted on GitHub: https://github.com/ClickHouse/web-tables-demo https://github.com/ClickHouse/web-tables-demo Which you can query anywhere using a ClickHouse local database engine. When you do a query, it downloads only the required ranges of required columns from the server and does computations on your machine. In the same way you can query any data sources with ClickHouse, either local or remote, which is a strict superset in power and usability than duckdb.
- voidUpdate 2y agoSo when I want to browse your website on my phone with a limited data plan, I have to download an entire database client and database, as well as any of your other huge JS libraries?
- wruza 2y agoI'd prefer that to page reloads, lost focus, scroll position and any input between clicking on a checkbox and receiving a response. Bonus points if I only have to download a db-diff next time, compared to downloading 5x db worth through POST /get-products?<filters>&page=<n> in every e-shopping session. Rarely a client-facing business database exceeds your average youtube cat video in size (assuming pics stored as urls).
- threeseed 2y agoWhat year was this comment written ? Most web apps these days are single page applications which don't require page reloads for every UI interaction.
- wruza 2y agoSadly the current one. Most web apps that have a backend database do all of the aforementioned. See e.g. https://amazon.com https://amazon.com Also, https://www.ebay.com/b/Digital-Cameras/31388/bn_779 https://www.ebay.com/b/Digital-Cameras/31388/bn_779 -- try choosing a brand https://www.bestbuy.com/site/video-games/video-games-accessories/abcat0715000.c?id=abcat0715000 https://www.bestbuy.com/site/video-games/video-games-accesso... -- only loses scroll, probably a winner https://www.newegg.com/p/pl?N=50001157%20100007671%20601393085%20601351801%20601298157%20601304866%20600565702%20600095610%20600005864%20600005862%20600005860%20600005851%20600005846%20600005857 https://www.newegg.com/p/pl?N=50001157%20100007671%206013930... -- uses "Apply" button, out of competition, but still better https://www.levi.com/US/en_US/sale/mens-sale/c/levi_clothing_men_sale_us https://www.levi.com/US/en_US/sale/mens-sale/c/levi_clothing... -- a winner of the "wtf is going on after I click my size" category
- vundercind 2y ago
- jeroenhd 2y agoWhy run a database in WASM when IndexedDB exists? Browsers already have a database built in, I don't see the need to download another one.
- isodev 2y agoIt feels very impractical indeed. Also the size of the binary to load the compiled wasm. All this would be much better done on the server and if really needed, users may be given a way to download results (ideally with their own preferred tool for fetching files)
- jampekka 2y agoA server sounds quite inpractical if you could otherwise serve the application statically. Or offline. Also having user data on server causes problems with privacy etc.
- isodev 2y agoWell the data has to come from somewhere right? If the goal is to facilitate client-side (bring your own data) scenarios, I'd make a proper native desktop app and take full advantage of the system. A hybrid something running in the browser feels like a compromise between both solutions.
- jampekka 2y agoWhy would you make a "proper" native desktop app (or more specifically apps for every platform you want to support) if you can do it with a PWA (which you can do for vast majority of apps).
- isodev 2y agoBecause PWAs are limited in more ways than practical to list here, just to name a few - the browser they run in, restricted by the availability and quality of internet connection, they're not sustainable as build tools don't even bother with backwards compatibility given the pace of evolution, practically no user control over a pwa "app" running in the browser. Remember, PWAs exist as a work-around for gatekeepy OS vendors making it hard to create cross-platform apps. PWAs don't resolve anything - PWAs move the problem to the browser space, where (today at least) we only have closed, proprietary, very-much revenue-driven browser implementations. The related web standards have also largely been influenced by FAANGS as the likes of Google wanting to turn "the web as their webstore".
- Zambyte 2y agoAnother interesting option is PouchDB[0], which is a Javascript implementation of the CouchDB[1] synchronization API. It allows you to acheive eventual consistency between a client with intermittent connectivity, and a backend database. [0] https://pouchdb.com/ https://pouchdb.com/ [1] https://couchdb.apache.org/ https://couchdb.apache.org/
- sandwitches 2y ago[dead]
- xnorswap 2y agoI don't understand these "DB in browser" products. If the data "belongs" to the server, why not send the query to the server and run it there? If the data "belongs" on the client, why have it in database form, particularly a "data-lake" structured db, at all? A lot of the benefits of such databases are their ability to optimise queries for improving performance in a context where the data can't fit in memory (and possibly not even on single disks/machines), as well as additional durability and atomicity improvements. If the data is small enough to be reasonable to send to a client, then it's small enough to fit in memory, which means it'll be fast to query no matter how you go about it. The page says one advantage is "Ad-hoc queries on data lakes", but isn't that possible with the most basic form that simply sends a query to the database? What am I failing to understand about this category of products?
- tlarkworthy 2y agoThat analytics in the browser is about 10000x times more performant, and doesn't contest on a shared resource.
- lmeyerov 2y agoRight - so not a database, but a columnar analytics compute engine. Half of why we wrote arrow js was for pumping server arrow data to webgl, and the other half for powering arrow columnar compute libraries & packages like this. For sub-100ms smooth interactivity, for data in a certain size range sweet spot, can be very nice!
- jimberlage 2y agoThere is an increasing subset of people (think those that used to work in MS Access, Excel power users) who learned SQL in a business IT course or on the job. They don’t have data sized to fit in a DB, but they do want to do analyses that use window functions and things that SQL makes more natural than Excel vtables or functions. They may have to give reports to an equally technical boss who would like to play with the report and explore assumptions using SQL. The data size is typically not large; SQL and integrating with cloud-based spreadsheets is the selling point.
- threeseed 2y agoThey didn't mention the lifecycle of the database. Because if it's anything longer lived than a week then it could be used by marketers to evade Apple's ITT for retargeting. Which would be a huge win for advertisers and a loss for privacy.
- paulgb 2y agoAs far as I can tell they're not storing anything locally, they're pulling the data from Google Cloud Storage as you access it. That said, a Wasm module doesn't have access to any storage facilities that the browser doesn't already expose (IndexedDB, OPFS, cookies, etc.) so even if it could be used by marketers, they would gain nothing by using DuckDB for that over just using the same underlying browser storage API.
- zX41ZdbW 2y agoI tried https://shell.duckdb.org/ https://shell.duckdb.org/, but it was a very rough experience. The "delete" button does not work. The "home" button inserts a whitespace. Pasting with "Ctrl+v" also does not work. Every keypress results in blinking, and there is a notable input lag. When I tried a query duckdb> SELECT * FROM 'https://clickhouse-public-datasets.s3.amazonaws.com/github_events/partitioned_json/*.gz' ...> ; Catalog Error: Table with name https://clickhouse-public-datasets.s3.amazonaws.com/github_events/partitioned_json/*.gz does not exist! Did you mean "sqlite_master"? LINE 1: SELECT * FROM 'https://clickhouse-public-datasets.s3.... Suggesting the "sqlite_master" database is also misleading.
- niutech 2y agoHow would DuckDB know which *.gz are in that folder? Is there a directory listing?
- cryptonector 2y agoSupercookies?