6 ms·
Well, first doing `find > .my-index` and then measuring `cat .my-index` would give you even better results... I don't find it noteworthy that reading from an in
by filmor 5y ago
Well, first doing `find > .my-index` and then measuring `cat .my-index` would give you even better results... I don't find it noteworthy that reading from an index is faster than actually recursively walking the filesystem.
- yakubin 5y agoYes. I once wrote a tool which spends a lot of time traversing the filesystem and to my surprise in one scenario most of its time is spent on-CPU (!) in kernel-mode in the readdir_r(2) syscall implementation. I still haven't dived into what it's doing on the CPU, but it sure is interesting.
- cseleborg 5y agoYeah, that's not really the interesting part. It's clever because if you're a developer editing source files, the Git index is relatively up-to-date. So, as the author says, why not take advantage of it?
- rakoo 5y agoNo, it's not surprising, so why do we still not use indexes for this ? NTFS maintains a journal of all files modification (https://en.wikipedia.org/wiki/USN_Journal https://en.wikipedia.org/wiki/USN_Journal). This is used by Everything (https://www.voidtools.com/support/everything/ https://www.voidtools.com/support/everything/) to quickly and efficiently index _all_ files and folders. Thanks to that, searching for a file is instantaneous because it's "just an index read". The feature is common: listing files. We know that indexes help solve the issue. But we still use `find` because there's no efficient files indexing strategy. Yes, there's updatedb/mlocate, but updating the db requires a full traversal of the filesytem because there's no efficient listing of changes in linux, only filesystems-specific stuff. So we will still have articles like this until the situation changes, and it will still be relevant to the discussion because it's not "solved"
- roywiggins 5y agoWizTree also uses NTFS metadata to perform shockingly fast reports on space usage. Just much, much faster than anything else I've seen, and I'm not sure there's any equivalents for other filesystems.
- rakoo 5y agoWizTree is so fast it makes me yearn for the day we have the same funcitonality on Linux. Especially not because I'm particularly interested in visualizing free disk space on the regular, but because I want to backup changed files as soon as possible and the only reliable way to do it is to have some kind of journal.
- ww520 5y agoUpdating an index adds additional time, adding a random seek to the index file location. Also requires some transaction to group the update to the file metadata and the index. It just adds more complexity with little benefit.
- ziml77 5y agoAs someone who uses Everything many times per day, I can say that there is a significant amount of benefit. I don't think I'd be able to function at work without being able to instantly search all my files. The lack of a similar solution on Linux is one of the big barriers to me using it. The best options I've seen there all refresh the index on a schedule.
- ReidZB 5y agoI suppose eBPF could be used to implement real-time index updates, which could be neat. Maybe I'll investigate doing that as a little side project. Though from a quick look, doing it efficiently may require some modification to mlocate's updatedb program (depending on how -U works exactly...).
- Groxx 5y agoAs always, it depends. Doing a hefty build can make half a million files on my machine - even minor additional file-creation latency can add up VERY quickly in that scenario. And frankly I do far more builds per day than I do `find`, though the build system likely does a fair number of shallow ones. In user-visible-oriented folders though, oh heck yes it should all be indexed.
- rakoo 5y agoI have to use Windows for work and I never see Everything take any CPU ever, even when I have to compile (maybe not a million files though). It's all asynchronous so builds aren't slowed down, indexing happens in the background but never takes long and when you need it it's all there because as you said you don't need to search for files every second. So, in practice, it works.
- ryantriangles 5y agoEverything's approach has important tradeoffs: the indexing service must be run as administrator/root, and by reading the USN journal it gives users a listing of _all_ files and mtimes on the filesystem regardless of directory permissions. That means that any user who can run a file search can also see every other users' files, which includes their partial web history (since most browsers, including Firefox and Chrome, cache data in files named for the website they're from, and the mtime will often be the last visit time). If you want to avoid this, you can only switch to Everything's traditional directory-traversal-based index, which has the same performance as updatedb/mlocate, Baloo, or fsearch. I think these tradeoffs are why there hasn't been as much interest in replicating it, combined with the fact that mlocate/fd/find/fsearch are already a lot faster in most circumstances than the default Windows search and good enough for the most common use cases (although there are certainly usecases where they're not).
- charcircuit 5y ago>That means that any user who can run a file search can also see every other users' files, which includes their partial web history (since most browsers, including Firefox and Chrome, cache data in files named for the website they're from, and the mtime will often be the last visit time). Most people only have a single user. Sharing a computer with someone is not that common of a thing to do anymore.
- rakoo 5y agoThat's an extremely important tradeoff but it's not an issue in the process: it is possible to split indexing and searching in a root process and querying in a user process that asks the first one, and the first one filters with the different rights. Nothing we've never heard of.
- sillysaurusx 5y agoIt feels like the index update should be a part of the filesystem layer itself. Not a separate process like you're saying.
- pabs3 5y agoIn Linux there is fanotify for monitoring for filesystem-wide events.
- rakoo 5y agoThe problem with *notify is that it's a push-based system: the receiver (ie the process that is interested in changes) needs to be running to receive a change. Because there is no ACK, if it's not running, you miss changes and there is no way to get those events. Also, even if you received the change but have some failure to process it, you must find a way to not lose the change. What the USN Journal does is implement a pull-based system: the sender stores everything, and the receiver queries the sender when it wants/can, on its own rhythm, starting from an offset it manages. In a generic pull-based system the sender can optimistically send a notification for the receiver to be informed as soon as possible. I have my own personal views, but a push-based system with no acknowledgments only makes sense if missing events is ok, typically because you know you'll receive another event about the same "thing" in a short time; this system is not viable for file changes. A push-based system with acks requires the sender to register each receiver, which is a bit heavy. A pull-based system is just the simplest to implement and solves all problems.
- dufferzafar 5y agoHey! Are you aware of any non-ntfs filesystems that also maintain a USN style journal? upon which tools like Everything could be created? I wonder why common linux filesystems like ext2/ext4 don't support this. After having used locate and all its friends (rlocate, plocate, lolcate-rs) - tools like Everything & WizTree on Windows feel like a breath of fresh air!
- rakoo 5y agoI'm absolutely not an expert, but I feel like log-structured filesystems (https://en.wikipedia.org/wiki/Log-structured_file_system https://en.wikipedia.org/wiki/Log-structured_file_system) are a natural fit for this kind of things: an index "just" has to read the latest written entries. But if we're talking about the future, we're probably talking about btrfs and zfs, both of which have the internal machinery to give you a feed of "recently changed files" up to the beginning of the filesystem. While writing this answer I stumbled upon https://github.com/rflament/loggedfs https://github.com/rflament/loggedfs which is probably a very nice solution to this problem.
- c_joly 5y ago> I don't find it noteworthy that reading from an index is faster than actually recursively walking the filesystem I agree, what I found noteworthy though is that git ls-files uses this index you already have for. (author here, “proof”: https://cj.rs/contact/ https://cj.rs/contact/ & https://keybase.io/leowzukw https://keybase.io/leowzukw)