8 ms·
Searchcode.com’s SQLite database is probably 6 terabytes bigger than yours
- leighleighleigh 2y agoI've been looking for a service just like searchcode, to try and track down obscure source code. All the best, hope it can be sustainable for you.
- antithesis-nl 2y agoYup, they win. My biggest SQLite database is 1.7TB with, as of just now 2314851188 records (all JSON documents with a few keyword indexes via json_extract). Works like a charm, as in: the web app consuming the API linked to it returns paginated results for any relevant search term within a second or so, for a handful of concurrent users.
- k_bx 2y agoI think FS-level compression would be a perfect match. Has anyone tried it successfully on large SQLite DBs? (I tried but btrfs failed to do so, and I didn't get to the bottom of why).
- zimpenfish 2y ago> I think FS-level compression would be a perfect match. Has anyone tried it successfully on large SQLite DBs? I've had decent success with `sqlite-zstd`[0] which is row-level compression but only on small (~10GB) databases. No reason why it couldn't work for bigger DBs though. [0] https://github.com/phiresky/sqlite-zstd https://github.com/phiresky/sqlite-zstd
- piterrro 2y agoI did a small benchmark for compression with VFS and column compression a while ago: https://logdy.dev/blog/post/part-3-log-file-compression-with-zstandard-vfs-in-sqlite-benchmark https://logdy.dev/blog/post/part-3-log-file-compression-with... https://logdy.dev/blog/post/part-4-log-file-compression-with-column-zstandard-in-sqlite-benchmark https://logdy.dev/blog/post/part-4-log-file-compression-with... It all depends on the use cases and read/write patterns. Imo if well designed could yield added value
- pinoy420 2y agoYou should use mongodb. It’s web scale
- Traubenfuchs 2y ago> My biggest SQLite database is 1.7TB with What do you run this on? Just some aws vpc with a huge disk attached?
- antithesis-nl 2y agoA Windows Server VM on a self-hosted Hyper-V box, which has a whole bunch of 8TB NVMe drives; this VM has a 4TB virtual volume on one of those (plus a much smaller OS volume on another).
- immibis 2y agoI can see that you're a user of AWS. Check some prices on dedicated servers one day. They're an order of magnitude cheaper than similar AWS instances, and more powerful because all compute and storage resources are local and unshared. They do have a higher price floor, though. There are no $5/month dedicated servers anywhere - the cheapest is more like $40. There are $5/month virtual servers outside of AWS which are cheaper and more powerful than $5/month AWS instances.
- cm2187 2y agoHow do you backup a file like that?
- antithesis-nl 2y agoUsing the SQLite backup API, which pretty much corresponds to the .backup CLI command. It doesn't block any reads or writes, so the performance impact is minimal, even if you do it directly to slow-ish storage.
- hu3 2y ago> It doesn't block any reads or writes. That's neat! I bet it keeps growing a WAL file while the backup is ongoing right?
- rco8786 2y agoHard to imagine doing it any other way, which is probably fine up until you hit some larger files sizes.
- ncruces 2y agoThat copies the entire file each time (not just deltas). You may find sqlite_rsync better.
- leosanchez 2y agosqlite_rsync is new tool created by sqlite team. It might be useful.
- homebrewer 2y agoI use zfs snapshots, they work in diffs so they're very cheap to store, create, and replicate.
- blitzar 2y agoCtrl-c, Ctrl-v
- 2y ago
- 1f60c 2y agosearchcode doesn't seem to work for me. All queries (even the ones recommended by the site) unfortunately return zero results. Maybe it got hugged? https://searchcode.com/?q=re.compile+lang%3Apython https://searchcode.com/?q=re.compile+lang%3Apython
- adulion 2y agoEven taking of the language filter you only get 5 results for a very common function!
- RandomRandy 2y agoI had to scroll down the search page and select the sources and languages to get a result
- mojosam 2y agoYeah, I just searched for “driver_register”, a call that would show upin a large number of Linux drivers in the open source Linux kernel, not to mention other public-facing repos, and it only returned two results, neither from the mainline Linux kernel repo.
- lenkite 2y agoIt doesn't list Go types from several k8s projects on github that I contribute to. Feel something is buggy about the filtering as well. I guess he will take some time to iron out all issues - suspect not all his data got migrated into the new db and the DB size should be far greater than 6TB. That feels rather low for github. But I liked his tip about SQLite driver scalability to avoid that stupid locked error that I too have faced regularly. numCPUS for readers and single writer - will try that out.
- StilesCrisis 2y agoI searched for several Skia classes and it never found the actual Skia repo, just forks and references from unrelated repos. It also failed to find several classes entirely. Skia exists in GitHub as well as in Chromium CodeSearch so it should have come up at least twice. As a sanity check, "fwrite" only has 8 references in the entire database. Yeah, agreed, I think the migration didn't actually work.
- bborud 2y agoI've been using RWMutex'es around SQLite calls as a precaution since I couldn't quite figure out if it was safe for concurrent use. This is perhaps overkill? Since I do a lot of cross-compiling I have been using https://modernc.org/sqlite https://modernc.org/sqlite for SQLite. Does anyone have some knowledge/experience/observations to share on concurrency and SQLite in Go?
- KingOfCoders 2y agoWould be very interested too.
- adenta 2y agoDo you really need concurrency? My understanding is read are like… picoseconds, because everything happens in memory. You don’t have a separate server to call. This guy is great: https://fractaledmind.github.io/2024/04/15/sqlite-on-rails-the-how-and-why-of-optimal-performance/ https://fractaledmind.github.io/2024/04/15/sqlite-on-rails-t...
- adenta 2y agohttps://fractaledmind.github.io/images/railsworld-2024/111.png https://fractaledmind.github.io/images/railsworld-2024/111.p... https://fractaledmind.github.io/2024/10/16/sqlite-supercharges-rails/ https://fractaledmind.github.io/2024/10/16/sqlite-supercharg...
- tiagod 2y agoYou're off by two orders of magnitude.
- lokimedes 2y agoI’ll bet you some CERN PhD student has a forgotten 100 TB detector calibration database in sqlite somewhere in the dead caverns of collaboration effort.
- deleted 2y ago[deleted]
- rvnx 2y agoIt could really happen. It's an organization with an unpredictable return on investment, in practice, they don't really have any negative consequences if they waste public money, or if it was actually useless (unless too obvious to external people). It's somewhat part of investing into experimental science.
- lokimedes 2y agoI've been there (when it was 100GBs scale, 15 years ago), trust me, it can and do happen :)
- golergka 2y agoSabine Hossenfelder has recently released a video about it: https://youtu.be/shFUDPqVmTg?si=xZAQZ725UEcO8lf_ https://youtu.be/shFUDPqVmTg?si=xZAQZ725UEcO8lf_
- lokimedes 2y agoSabine really represents the “apathetic disillusioned” - no hope for fundamental physics. I left after we discovered the Higgs based on similar observations, there was too much invested in group think. If you have to raise $50B to ask “what happens if…” then the “if” has to be damned likely, or you have to share the risk with a large group. My own conclusion, back then, was that the collider paradigm had run its course. Without economical tools, and no theory to guide experiments, the field was stuck. I’m not apathetic though, and believe there are ways for both theories and experiments to break the stall. Theory could dig into the backlog of shortcuts and dirty tricks, that underlines the machinery of QFT, are there other ways of probing the quantum fields? Experimentalists can at the least get behind novel acceleration schemes like laser plasma wake fields, to reduce the massive capital risk of conducting model-free searches for new signatures. Or as the theorists, hunt for alternative ways to “excite the quantum fields”. This may Not be recipes, but as a community Particle physics has been way to focused on chasing resonances with ever larger machines.
- Alifatisk 2y agoIs the site like grep.app?
- feverzsj 2y agoI'd consider no relational db scales reads vertically better than SQLite. For writes, you can batch them or distribute them to attached dbs. But, either way, you may lose some transaction guarantee.
- taurknaut 2y agoReading without write contention is not a terribly difficult problem. You could use any database and it'd work fine. It's the mutations that distinguish db engines. Sqlite is indeed close to ideal but the comparison to other databases (at scale no less) is without substance.
- hobs 2y agoGenerally to scale reads appropriately (to >10k readers) you need things like connection pooling and effective replication of some time readable secondaries (which can maintain transaction guarantees if you do synchronous writes, which... maybe not great) - I dont think SQLite has either of those.
- aerioux 2y agonit: this code snippet has a typo (dbWrite wasn't defined) ``` dbRead, _ := connectSqliteDb("dbname.db") defer dbRead.Close() dbRead.SetMaxOpenConns(runtime.NumCPU()) dbRead, _ := connectSqliteDb("dbname.db") defer dbWrite.Close() dbWrite.SetMaxOpenConns(1) ```
- franciscop 2y agoYeah, the second `dbRead` should prob be `dbWrite`: dbWrite, _ := connectSqliteDb("dbname.db") I was a bit confused by the code, then assumed it was a typo
- rednafi 2y agoFascinating read. Those suggesting Mongo are missing the point. The author clearly mentioned that they want to work in the relational space. Choosing Mongo would require refactoring a significant part of the codebase. Also, databases like MySQL, Postgres, and SQLite have neat indexing features that Mongo still lacks. The author wanted to move away from a client-server database. And finally, while wanting a 6.4TB single binary is wild, I presume that’s exactly what’s happening here. You couldn’t do that with Mongo.
- christophilus 2y agoThe binary would only contain the database engine; not the database. It’s probably a few MB, if I had to guess.
- taurknaut 2y agoPresumably the database needs to be distributed to servers, too. The engine needs to access something. This is a necessity whether or not it's referred to as a binary.
- taurknaut 2y agoThis comment appears to be the only place between article and thread where mongo getting recommended is even mentioned. Maybe stop giving them free press. I certainly haven't heard anyone seriously suggest them for about about ten years now. Edit: ok the other guy mentioning mongo is clearly being sarcastic
- sohzm 2y agothe person suggesting mongodb in the replies isnt serious, its just a meme reference
- deleted 2y ago[deleted]
- throwaway984393 2y ago[dead]
- aliasav 2y agoI have been contemplating postgres hosted by AWS vs locally using SQLite with my Django app. I have low volume in my app, low concurrent traffic. I may have joins in future since the data is highly relational. I still chose postgres managed by AWS mainly to reduce the operational overhead, I keep thinking if I should have just gone with sqlite3 though
- pietz 2y agoI host my stuff on railway and while I love SQLite, I usually go with their Postgres offering instead. It's actually less work for me and although I haven't tested this scientifically, it's hard too imagine the SQlite way would save me serious money.
- morphle 2y agoWithout reading the article, I always quickly estimate what a DRAM based in-memory database would cost. You build a mesh network of FPGA PCBs with 12 x 12.5 Gbps links interconnecting them. Each $100 FPGA has 25 GB/s DDR3 or DDR2 controllers with up to 256 GB. DDR3 DIM have been less than $1 per GB for many years. 6.4x($100+256x4x$1)=$7,193.6 worst case pricing So this database would cost less than $8000 in DRAM chips hardware and can be searched/indexed in parallel in much less than a second. Now you can buy old DDR2 chips harvested from ewaste at less than $1 per GB, so this database could cost as low as $1000 including the labour for harvesting. You can make a wafer scale integration with SRAM that holds around 1 terabyte SRAM for $30K including mask set costs. This database would cost you $210.000. Note that SRAM is much faster than DDR DRAM. You would need 15000/7 customers of this size database to cover the manufacturing of the wafers. I'm sure there are more than 2143 customers that would buy such a database. Please invest in my company, I need 450,000,000 up front to manufacture these 1 terabyte wafers in a single batch. We will sprinkle in the million core small reconfigurable processors[1] in between the terabyte SRAM for free. For the observant: we have around 50 trillion transistors[2] to play with on a 300mm wafer. Our secret sauce is a 2 transistor design making SRAM 5 times denser that Apple Silicon SRAM densities at 3nm on their M4. Actually we would not need reconfigurable processors, we would intersperse Morphle Logic in between the SRAM. Now you could reprogram the search logic to do video compression/decompression or encryption by reconfiguring the logic pattern[3]. Total search of each byte would be a few hundred nanoseconds. The IRAM paper from 1997 mentions 180nm. These wafers cost $750 today (including mask set and profit margin) and now you would need less than a million investment up front to start mass manufactoring. You just would need 3600 times the wafer amount compared to the latest 3nm transistor density on a wafer. [1] Intelligent RAM (IRAM): chips that remember and compute (1997) https://sci-hub.ru/10.1109/ISSCC.1997.585348 https://sci-hub.ru/10.1109/ISSCC.1997.585348 [2] Merik Voswinkel - Smalltalk and Self Hardware https://www.youtube.com/watch?v=vbqKClBwFwI&t=262s https://www.youtube.com/watch?v=vbqKClBwFwI&t=262s [3] Morphle Logic https://github.com/fiberhood/MorphleLogic/blob/main/README_MORPHLE_LOGIC.md https://github.com/fiberhood/MorphleLogic/blob/main/README_M...
- dist1ll 2y ago> 64000x($100+256x$4)=$7,193,600 worst case pricing Not sure how you arrived at this calculation. 256x$4 already accounts for 256GB. The database in the OP is 6400GB large. So shouldn't it be 25x($100+256x$4) = $28100? FWIW in practice this number should be much, MUCH lower. You can get 1TB of ECC DDR4 for ~$2k, probably lower if you buy wholesale.
- deleted 2y ago[deleted]
- lutusp 2y agoWait ... dbRead, _ := connectSqliteDb("dbname.db") defer dbRead.Close() dbRead.SetMaxOpenConns(runtime.NumCPU()) dbRead, _ := connectSqliteDb("dbname.db") defer dbWrite.Close() dbWrite.SetMaxOpenConns(1) Is dbWrite ever declared? I know it's just an example, but still ... Now I have to ask myself whether this error resulted from relying on a human, or not relying on one.
- qingcharles 2y agoTypo
- gedw99 2y agoI use marmot which gives me a multi master SQLite. It uses a simple crdt table structure, and allows me to have many live SQLite instances in all data enters. A Nat Jetstream server is used as a core with all SQLite DBS connected to it. It operates off the WAL and is simple to install. It also replicates all blobs to S3 , with the directory structure in the SQLite db. With a Cloudflare domain , the users request is automatically sent to the nearest Db. So it replaces cloudscapes D1 system for free . Just a hetzner 4 euro cos is enough. Mine is holding about 200 gb of data. https://github.com/maxpert/marmot https://github.com/maxpert/marmot