3 ms·
More technically, here’s what we have: A “baseline” filesystem watcher which uses only the standard library. It has been made to beat kqueue. And it does. A p
by e-dant 4y ago
More technically, here’s what we have:
A “baseline” filesystem watcher which uses only the standard library. It has been made to beat kqueue. And it does.
A platform filesystem watcher for Darwin is used, but certain event properties are handled by the standard library. Namely, the event time and the path type.
A platform filesystem watcher is schedule for Windows. Work hasn’t been started.
A platform filesystem watcher for Linux (> 2.4 or so) was toyed with but ultimately rejected out of accuracy concerns. It was far more efficient than the cross-platform implementation “warthog”, no doubt, but it lacked accuracy. Work is being done to get most of the benefits from both worlds.
There are problems with the “baseline” watcher (which I’ve named “warthog” because it’s sturdy and reliable). But those are potential efficiency losses when watcher more than a few million paths. They are, thankfully, not accuracy or safety problems.
Maybe you can see the solution emerging here?
Here’s where we’re going next:
The most efficient kernel watchers can be used on most platforms, but checked for their accuracy periodically by the “warthog” watcher.
- diffxx 4y agoWhat do you mean by beat kqueue? Is it faster than kqueue? Does it use less memory than kqueue? How does the baseline filesystem watcher work? If it doesn't use kqueue, does it poll the filesystem periodically and diff against an in memory representation? If yes, see my other comments. If not, I am genuinely curious what you are doing because you know something that I do not.
- e-dant 4y agoWhen I began this project, I started with kqueue. The performance was wanting and there were bugs with very large file trees. I moved to a minimal std::filesystem-based watcher and optimized it from there. There hasn’t been a formal head-to-head test between the two. That should be about halfway down my todo list. It’s worth revisiting more formally. My response to this question should help here: https://news.ycombinator.com/item?id=33247155#33251437 https://news.ycombinator.com/item?id=33247155#33251437 In short, there’s no secret sauce. There’s an efficiency spread in (what I consider) edge-cases. Every potential gain over other naive watchers implemented with kqueue is likely algorithmic. I store events in a historical map, compare differences to the current state of the file tree, prune them, and send events when they change. That’s the whole implementation: scan paths, record their attributes, check for differences in the map, and send events when they happen. I haven’t given much thought to exactly why it beats kqueue, nor are there any good tests showing by how much. (Again, this is worth doing.)
- diffxx 4y agoMakes sense. I have only used kqueue on macos to monitor a small number of files and I find it quite painful to use and the semantics were confusing, not sure if it is different on say freebsd. Just as a heads up, one of the strange fsevents issues is that it fails if you register two directories where one directory is a prefix of the other. So say that you want to monitor directories $ROOT/foo and $ROOT/fo and you register an event stream first with $ROOT/foo and then $ROOT/fo, you will only receive events for paths in $ROOT/fo and no events for paths in $ROOT/foo (I just double checked that this is still the case in Monterey at least). I never bothered to report this to apple but worked around it by just registering a stream with $ROOT if I detected that one path name was a substring of another.
- marwis 4y agoHave you tried using auditpipe?