12 ms·
Show HN: FSQL – Search through your file system with SQL-esque queries
- m0skit0 9y agoI think SQL is too verbose for use on the terminal. find + grep does the trick with way less verbose syntax (but also probably less readable). With that said, it is quite cool.
- deleted 9y ago[deleted]
- misterdata 9y agoFor those of us on Windows: Everything [1] does the job quite nicely with much less verbose syntax. [1] http://voidtools.com http://voidtools.com
- nirav72 9y agoAgent Ransack is another good one.
- saganus 9y agoDoes it keep everything local? I couldn't find anything on the FAQ and I remember a similar tool posted here on HN and people complained that it called home for some search functionality.
- degenerate 9y agoI've been using this for years and it is LIGHTNING fast. No need to "index" all the files because it reads directly from the MFT. If you create a new file matching the search pattern it's already sitting in the results window by the time you alt-tab. Also, the "directory size" equivalent of Everything is WizTree [1] ... much faster than WinDirStat, which I see recommended way too often. [1] http://antibody-software.com/web/software/software/wiztree-finds-the-files-and-folders-using-the-most-disk-space-on-your-hard-drive/ http://antibody-software.com/web/software/software/wiztree-f...
- kevindqc 9y agoAre you sure it doesn't index anything? There is even a section in the settings called Indexes. Also when you launch it first, it's going to be empty and says it's scanning your folders and it takes a little bit until you see something. I think it is still indexing (maybe using the MFT instead of recursively listing files and directories), it's just a lot better than Windows search indexing. And it might use this [1] to keep up to date? It's mentioned in the settings [1] https://en.wikipedia.org/wiki/USN_Journal https://en.wikipedia.org/wiki/USN_Journal
- degenerate 9y agoNo I'm not sure. Maybe it builds a rudimentary index... but give it a shot yourself and see. 15 seconds after installing, you can search your entire system instantly. It's crazy whatever it is doing. And yes I do believe it uses the USN Journal to stay up to date.
- kevindqc 9y agoThere's actually a wikipedia page, and it explains how it works. As we suspected, it uses the MFT for initial indexing and change journal for updates. https://en.wikipedia.org/wiki/Everything_(software) https://en.wikipedia.org/wiki/Everything_(software)
- joshschreuder 9y agoThanks for WizTree, you're right that it's an order of magnitude faster than WinDirStat. Only thing it's missing is that graphical block view, but for my usecase it will be perfect.
- jazoom 9y agoI'd really like to know why Microsoft can't do something like this in Windows itself. Everything is very useful.
- jonas21 9y agoIsn't that what they tried to do with WinFS? https://en.wikipedia.org/wiki/WinFS https://en.wikipedia.org/wiki/WinFS
- jazoom 9y agoI don't know their reasons for WinFS, but I do know that Everything's authors figured out a long time ago how to get great search with Windows' current FS.
- sixothree 9y agoMy guess is that MS needs an excuse to index the contents of your files.
- chinhodado 9y agoEverything does not index the content of your files, only the name and some attributes. Also MS already index file contents by default, it just sucks at it. There have been several occasions where I use the find file by name syntax and Windows can't even find the file in the current folder.
- sixothree 9y agoThis is exactly my point. The key to finding files is to use name:*mylostfile*.txt But Everything excels at this so pretend you never saw it.
- hokkos 9y agoThis is my secret weapon.
- drumttocs8 9y agoEverything is awesome
- kileywm 9y agoThis is pretty amazing. I've been using a Windows program called Agent Ransack [0] for finding files with regex, but this Everything program is so much faster. Incredibly so. Thanks for the tip! [0] https://www.mythicsoft.com/agentransack https://www.mythicsoft.com/agentransack
- Adverblessly 9y agoEverything is pretty awesome, one of the things I really miss on Linux. When I last looked for a Linux replacement for it, none had the real time updates or the instant search ui, and even those that claimed to index the file system for a quick search were very slow. I actually ended up writing my own hacky and rudimentary GUI over locate to achieve something that fits my needs (and of course doesn't support real time updates). Maybe things changed since then? Any chance that the HN crowd knows of a good Linux replacement for Everything?
- enedil 9y agoI don't know what features exactly Everything has, but maybe fsearch is a good alternative? http://www.fsearch.org/ http://www.fsearch.org/
- kuschku 9y agoTry out KDE's search. It does realtime fulltext indexing and just works. Also integrated with the KRunner framework.
- super_mario 9y agoHave you tried FZF (the fuzzy finder)? https://github.com/junegunn/fzf https://github.com/junegunn/fzf It integrates well with UNIX shells, editors (VIM integration is excellent), git log search etc. It's really versatile. Of course, there is also q: http://harelba.github.io/q/ http://harelba.github.io/q/ for something that is more like the software described by OP.
- netheril96 9y agoI heard that Everything is dependent on the exact format of NTFS. The same algorithm isn't applicable to ext4 or HFS+.
- yread 9y agoI recently learned that you can also do advanced search (for e.g. date of modification) in Everything 1.4 (beta): filename dm:>=16/05/2017 size:5mb..9mb parents:>=3
- xurukefi 9y agoEverything is one of the first tools I install on every windows box. For me personally it is a must have. I don't think I know of any other tool which altered my workflow that heavily.
- samsk 9y agoWe have implemented smth. like this with sqlite extension, pretty powerful, with all the goodies sqlite (and its extensions) provides...
- kshvmdn 9y agoSounds interesting, would love to take a look if it's available anywhere.
- samsk 9y agoUnfortunately, it was commercial development, so it's not released. But implementation is relatively easy - it was a sqlite virtual table that (as much as I remember) looked in where condition for dir field, and listed that directory (= returned stat() data). Whole thing was quite interesting, because almost every component was somehow hooked into sqlite (either vith function or virtual table), so one could do pretty interesting things only with SQL.
- zitterbewegung 9y agoThis is really neat. MacOS has a easy to use smart folder which I use to find recent files and large files. An interface like this is an advantage because it's easy to understand what it's doing and it's cross platform . Other people make the claim it may be verbose (but being verbose makes the operation clearer) and SQL is so familiar to programmers that are power users.
- jstimpfle 9y agoFROM dir1, dir2 doesn't mean the same as in SQL. In SQL that's a join of dir1 and dir2, but here it's a union.
- kshvmdn 9y agoHence the esque. :) Good point though, I'll make a note of that in the README.
- devnonymous 9y agoNice project, wish you the best ! Although tbh, I personally won't use this simply because I know enough of find(1) to not see the cognitive overhead of switching to sql to do filesystem /queries/. Any examples where this would be better than using find (with the occasional filter thrown in) ?
- _jal 9y agoSubselects would be a pretty awesome feature. "select name from foo where name not in (select name from ../bar where date < ...)" I'm usually fine with `find`, but when doing things more interesting than just "find files in this directory that are not in that directory", while uncommon, tend to make me think about my pipeline a bit.
- taeric 9y agoOut of curiosity, I'm interested in how folks would do the "in this not that" folder query. At a gut shot, I'd assume that diff would be used. I'm about to dig through the find man page to see if it has something directly to help.
- hoytech 9y agoAssuming no dups in file1, this outputs lines in file1 that aren't in file2: sort file1 file2 file2 | uniq -u (double file2 is not a typo :)
- isometry 9y agoOr (a little faster): comm -23 <(sort file1) <(sort file2)
- keithlfrost 9y agoOne tool for this is the 'comm' utility: given two files containing sorted lines, it can output one or more of (1) lines only in file 1, (2) lines only in file 2, and (3) lines common to both files.
- rekwah 9y agoReminds me of osquery [0]. [0] - https://github.com/facebook/osquery https://github.com/facebook/osquery
- kshvmdn 9y agoHaha, you're not alone -- https://github.com/kshvmdn/fsql/issues/2 https://github.com/kshvmdn/fsql/issues/2.
- bodhibyte 9y agoThanks for sharing! I've installed this locally and I'm really impressed by how easy to use and powerful this is! I must have missed the previous mentions on HN [0],[1]. [0] https://news.ycombinator.com/item?id=8528460 https://news.ycombinator.com/item?id=8528460 [1] https://news.ycombinator.com/item?id=12600790 https://news.ycombinator.com/item?id=12600790
- MrBuddyCasino 9y agoDoesn't Windows have something like this built-in? WMI or something?
- ziikutv 9y agofind/grep/awk/ag get me a long way to be honest. However, I think this is a cool project because it makes filtering of file attributes (such as size) so much easier. No need for splitting strings and using regex. Cheers.
- anuragbiyani 9y agoNot to take anything away from this project, but you can filter on size, permission, etc easily and robustly using just `find` (try -perm, -size, -{c,m}time, etc flags): https://linux.die.net/man/1/find https://linux.die.net/man/1/find If you are splitting strings (from output of `ls -l` presumably) for such tasks, then definitely take a look at find.
- m00s3 9y agoSeems if the query is always going to start with SELECT, that maybe it should be assumed? I would never use this though, ack or find seem sufficient to me.
- dTal 9y agoIt's a shame Bash used 'select' as an elaborate menu built-in - it'd be quite neat to name the binary that (and drop the quotes). The you could just type the query right into your prompt!
- koolba 9y agoYou could use an alias. They're case sensitive so SELECT could be mapped without impacting the built in "select".
- labster 9y agoJust use the fish shell instead, and you can avoid the years of shell cruft of bash, or the endless customization of zsh. Bash is a good environment for shell scripting, but not really the best for user interaction. Although perl is probably the best environment for shell scripting.
- fnj 9y agoXonsh is far better both interactively and for scripting than bash, fish, zsh, or perl.
- lorenzhs 9y ago"alias select=command select" lets you override it.
- emmelaich 9y agoYeah, I like the idea of using sql but it is painful to write. It would be nice to omit the select and the quoting; my suggestion would be that fsql "select * from ..." could be written as select all from ... and all could be assumed if ommitted; so you could write select from or just from ... But unfortunately `select` is a sh(1) reserved word and `from` is an existing command! (shows who who your mail is from) So maybe select and from could be shortened to sel and frm.
- bpchaps 9y agoNice! I'm actually working on a similar project to push lsof and files from /proc into some postgres tables. Lets me do cool things like query log files across a ~6000 server infrastructure similar to: SELECT distinct(l.name) FROM lsof l, lsofer_runs r WHERE l.lsofer_id = r.id AND fd_type = 'REG' AND l.fd ~ '[0-9][uw]' AND l.name like '%log' GROUP BY l.name, r.hostname ORDER BY name Best of luck!
- matthewaveryusa 9y agoso you're rewriting osquery? https://osquery.io/ https://osquery.io/
- bpchaps 9y agoHah, apparently.
- nthcolumn 9y agoI need this immediately.
- bpchaps 9y agoI'll see about getting it into a public repo soon and let you know.
- tyingq 9y agoHis description sounded like it would do joins across different hosts. Osquery looks to be single host at a time only.
- bpchaps 9y agoYep. I'm specifically writing it to find any log file that isn't being pushed into our third party logging service. It's a surprisingly difficult problem, especially considering the amount of tech sprawl that's accumulated. Since it's also a relatively low latency environment, it has to be written in a way that doesn't add too much load (without core isolation..).
- aardvark291 9y agoDidn't BeOS have some awesome database-like file system indexing and query system?
- Koshkin 9y agoSure; in general, a file system endowed with extended/extensible attributes can be naturally seen as a relational database (in which the files themselves are BLOBs).
- nayuki 9y agoThis video talks in detail about the extended attributes in BeOS: https://systemswe.love/archive/minneapolis-2017/ivan-richwalski https://systemswe.love/archive/minneapolis-2017/ivan-richwal... - "Metadata Indexes & Queries in the BeOS Filesystem" https://player.vimeo.com/video/209021697 https://player.vimeo.com/video/209021697
- amasad 9y agoDid someone come up with a generalized rule about putting SQL on top of every possible system that contains queryable information? Here is first-pass: >eventually every system that contains information that can be queried will have a sql interface
- koolba 9y agoThat's just a lemma in the wider theorem that every sufficiently complicated application evolves to have ad-hoc SQL reports exported to Excel.
- Koshkin 9y agoSQL is the standard language for querying relational data, so why not.
- Retra 9y agoIt's not a great standard. Practically keywords the whole English language...
- Koshkin 9y agoYes, it is a "Structured-English Query Language", formerly abbreviated as SEQUEL.
- devnonymous 9y agoConsidering its longevity as compared with other 'standards' that came around 40 years ago, I'd say it is in fact, a great standard.
- atemerev 9y agoYes, it is called "Spark driver" these days.
- drinchev 9y agoReally cool idea, but I'm really missing the WHY section in the Readme file. Thinking about a use case is quite hard. Anyone?
- CaseFlatline 9y agoDon't see it offhand so asking: 1) How/where are you storing the index 2) Have you tried this on large (30+ TB filesystems)?
- agumonkey 9y agoeven without an index, having a way to project declaratively instead of relying on cut/sed is giving me hot flashes.
- bangonkeyboard 9y agoOn macOS, there is a query syntax [0] that's usable in Spotlight and the mdfind(1) command. Richer searchable attributes [1], but the results may have to be piped through other tools for formatting or other output. [0]: https://developer.apple.com/library/content/documentation/Carbon/Conceptual/SpotlightQuery/Concepts/QueryFormat.html https://developer.apple.com/library/content/documentation/Ca... [1]: https://developer.apple.com/library/content/documentation/CoreServices/Reference/MetadataAttributesRef/Reference/CommonAttrs.html https://developer.apple.com/library/content/documentation/Co...
- jklehm 9y agoReminds me of WSSQL [0] on Windows. [0] https://msdn.microsoft.com/en-us/library/windows/desktop/bb231255(v=vs.85).aspx https://msdn.microsoft.com/en-us/library/windows/desktop/bb2...
- atemerev 9y agoI wanted to write the exact opposite: a Mysql/Postgres client as a FUSE filesystem driver. Namespaces -> folders, tables -> (editable) CSV files, stored procedures and settings accessible as (editable) plain text files.
- howderek 9y agoIf someone put data in a column that wasn't valid, like a string in a bigint column, would the table be altered or would the FUSE driver refuse to make the change?
- bbcbasic 9y agoSounds dangerous!
- atemerev 9y agoNo more than DELETE FROM ... WHERE ... wait, where is WHERE?
- philsnow 9y agoEven this can be made safe(r) if you only only only connect to your database nthrough a proxy that sanitizes queries. IIRC vitess adds an implicit LIMIT 10 to queries that don't have a limit.
- cdmckay 9y agoA simple solution is to not allow UPDATE or DELETE statements without a WHERE clause.
- vidarh 9y agoThere's been some attempts. Here's one: https://github.com/BMDan/DFuse https://github.com/BMDan/DFuse
- crivabene 9y agoDoes anyone remember WinFS (1)? Bill Gates described it as his biggest product regret (2). I remember I thought it was brilliant. Too bad it was probably a little bit too futuristic for its time, as for a few other things they launched when it just was not the right time... the clunky Tablet PCs (3) were for sure another example. (1) https://en.m.wikipedia.org/wiki/WinFS https://en.m.wikipedia.org/wiki/WinFS (2) http://www.zdnet.com/article/bill-gates-biggest-microsoft-product-regret-winfs/ http://www.zdnet.com/article/bill-gates-biggest-microsoft-pr... (3) https://en.m.wikipedia.org/wiki/Microsoft_Tablet_PC https://en.m.wikipedia.org/wiki/Microsoft_Tablet_PC
- UnfairIsaac 9y agoThat was my first thought, too. I'd love a file system that was more like a relational database. Also see the Pick operating system.
- LgWoodenBadger 9y agoTransactional ACID updates to file systems would be pretty fantastic.
- mpfundstein 9y agoyes. thinking about this for a long time already
- GlobalServices 9y agoAlso MUMPS. See https://en.wikipedia.org/wiki/Pick_operating_system https://en.wikipedia.org/wiki/Pick_operating_system for info. In some ways, Microsoft is on the way there with the way PowerShell works, and the ability to script things through OS functions that return objects which can be queried. If we ever see a WinFS, it would be very powerful combined with PowerShell
- Jweb_Guru 9y agoBTRFS has had a bunch of problems trying to actually compete with traditional filesystems, though. In the distributed world, CalvinFS seems pretty promising to me.
- vondur 9y agoI think the BeOS had a file system that was set up like a database that could be queried.
- danans 9y agoIt sure did. Alas at the time I had a BeBox in college (mid-late 90s) I didn't know SQL yet ;) I think it just searched over file metadata, not contents, though I might be mistaken.
- rs86 9y agoReally cool idea
- rodorgas 9y agoI never know when I have to use find vs grep. And linux grep is different from macOS grep so I google about it every day lol. I just never figure it out. I think I'll be a heavy user of FSQL.
- ScalaNovice 9y agoIt probably would be a good idea to mention powershell here. Very powerful and non-kludgy syntax. I do all my FS operations using that now. I'll give an example, I recently wanted to back up my arch packages but only the latest version for each package. Here's how I did it after maybe 5-10 minutes of trial and error: `$k = (Get-ChildItem | Group-Object{ $ns = $_.name.split( "-" ); $filtered= ($ns | %{ if( $_ -match $regex ) {""} else {$_} } ); $joined = ($filtered -join "" ); $joined } )` `($k | %{ ($_.Group | Sort-Object -Descending LastWriteTime)[0] } ).name | %{ cp $_ /mnt/Data/SuperShare/pkg/ }` Contrast this with bash where I would still be correcting the spaces after and before '[' Quick PS crash course for those who care: `Get-Help <command>` e.g. `Get-Help New-Alias` `New-Alias gci Get-ChildItem` Get-ChildItem is like ls. `Get-ChildItem -Recurse` has, to my knowledge, no equiv in bash and is equiv to `ls /*` in zsh. I guess I can write more if any of you guys show interest.
- adamhepner 9y agoThis is exactly why I love PowerShell.
- deleted 9y ago[deleted]
- Jaruzel 9y agoThis is nice, but what I'm actually looking for is a lightweight clone of SharePoint Search[1]. Something that has a self-hosted Web Interface, and an engine that I can point at some file servers, and let it index the files to it's hearts content. All I then have to do is search the index 'google style' for my files. Any suggestions? -- [1] https://i-technet.sec.s-msft.com/dynimg/IC423463.jpg https://i-technet.sec.s-msft.com/dynimg/IC423463.jpg
- legulere 9y agoReminds me a bit of SPARQL, that is used on the linux desktop e.g. by Gnome Music to find your music collection through Gnome Tracker
- gigatexal 9y agoKudos to the authors for taking an idea I've had for a while and actually doing it. Very, very cool.