6 ms·
Shocked to find out that this is very much a static website, it's "merely" downloading a small portion of the 43GB SQLite file in question with HTTP Range reque
by arianon 5y ago
Shocked to find out that this is very much a static website, it's "merely" downloading a small portion of the 43GB SQLite file in question with HTTP Range requests, and then it uses a WASM-compiled copy of SQLite to query the requested data. Very impressive.
- simonw 5y agoIt's the same trick that was described here - it's absolutely brilliant: https://phiresky.github.io/blog/2021/hosting-sqlite-databases-on-github-pages/ https://phiresky.github.io/blog/2021/hosting-sqlite-database...
- JoeyBananas 5y agowouldn't this be vulnerable to DOS attacks? I can make the database run arbitrarily long and complicated queries
- CodesInChaos 5y agoThe database runs on the client. It can't do anything you couldn't do through any other http client.
- simonw 5y agoSince the work all happens in your browser, the only victim of a long complicated query would be you own browse and the S3 bandwidth bill of the person hosting the database (since you'd end up sucking down a lot of data). But if you want to jack up someone's S3 bandwidth bill there are already easier ways to do it!
- floatingatoll 5y agoIt's a static file, CDN it.
- simlevesque 5y agoMany CDNs cannot cache 43GB files. Cloudflare's limit is 10GB, 20GB for Cloudfront, 150GB for Akamai and Azure.
- phiresky 5y agoYou can chunk the file into e.g. 2MB chunks. The CDN can then cache all or the most commonly used ones. That's what I did in the original blog post to be able to host it on GitHub Pages.
- dmos62 5y agoI would look into caching range requests. Simpler than pre-chunking or caching the whole database.
- sltkr 5y agoIt should be trivial to split it up into 1GB chunks or whatever. In fact, if you only request single pages, you could split the database up into pagesize-sized files. This is a lot of files, but at least it avoids the need to do range requests at all. You probably want to increase the pagesize a bit though (e.g. maybe 256 KiB instead of the default 4 KiB?)
- floatingatoll 5y agoRange requests are the star, so I would vote keep them? I mean, technically, the CDN could represent each 4kb block of the file as an individual URL, so that you do range requests by filename instead of by .. range request .. but at some point I definitely think RRs are more sane.
- rzzzt 5y agoProbably both, compute file name and relative offset from the original offset. (A minor complication is requests crossing file boundaries, in which case you have to do two requests instead of one.)
- jlg23 5y agoIf I get the essence of what you are saying... "the only victims would be both ends of the communication"? ;)
- quickthrower2 5y agoWhat is said is this site isn’t any more susceptible to DOS than any other site. Of all the ways of architecting a site this is probably the least vulnerable.
- chrisseaton 5y ago> I can make the database run arbitrarily long and complicated queries Do you understand that the database runs on your computer? You can only DOS yourself.
- carlhjerpe 5y agoIn a smart webserver this could mean not a whole lot more than a memcpy, and between you and the CDN it probably won't be too high latency. One would have to version the SQLite DB by version to avoid corruption by cache invalidation.
- rkeene2 5y agoMy static file HTTP server called "filed" [0] will satisfy a request in as few as 1 system call (no memcpy involved -- the kernel reads the file and sends the buffer to the NIC), by using sendfile(2). Most other webservers do a bit more work, like open the file for every request. [0] http://filed.rkeene.org/ http://filed.rkeene.org/
- scrame 5y agoWow! The end where they use sql to manipulate the DOM I think just blew a fuse in my brain. "Brilliant" feels like an understatement.
- MuffinFlavored 5y agoOn insert from Netlify: > [error: RuntimeError: abort(Error: server uses gzip or doesn't have length). Build with -s ASSERTIONS=1 for more info.] On insert from Github Pages: > [error: Error: SQLite: ReferenceError: SharedArrayBuffer is not defined] Your browser might either be too old to support SharedArrayBuffer, or too new and have some Spectre protections enabled that don't work on GitHub Pages since they don't allow setting the necessary isolation headers. Try going to the Netlify mirror of this blog for the DOM demos.
- k__ 5y agoDidn't know you could pump 43GB of data into a browser.
- wmf 5y agoThe trick is that the browser is not downloading all 43 GB, just the parts it needs.