10 ms·
Hacking ls -l
- cheald 14y agoWhile I appreciate the story, what's wrong with `ls -lh`?
- ghostfish 14y agoMy thoughts exactly. This is just complexity for complexity's sake. Useful as an exercise, but the -h flag already does this is an even more readable manner.
- bajsejohannes 14y agoThis works with sorting, and it's easier to pick big files out at a glance. I would use this as often as -h
- cheald 14y ago`ls -lhS` works just fine to sort with human-readable sizes, at least on Fedora.
- runejuhl 14y agoYou could also use 'ls -lh | sort -h'.
- sliverstorm 14y agoI believe that's only supported in newer versions of sort.
- pepve 14y agoI guess people can be very different in this respect. I really need the full number (or at least all of the numbers in the same unit) to be displayed. I don't find the 'human' format helpful at all when looking at ls output.
- sliverstorm 14y agoIt just depends what info you desire. I typically use "ls -lh" to quickly find large files to clean up to recover disk space.
- tsahyt 14y agoI use "du <directory> -h -d 1|sort -h" for this because large files may be nested in directories and ls doesn't display the size of content directories (as far as I know). The output is sorted. Note that the "-d" flag for "du" doesn't work with all versions, but there were similar flags on all systems.
- sliverstorm 14y agoI don't usually have access to a version of "sort" that supports the "-h" flag.
- paddyoloughlin 14y agoFor me, -h makes it more difficult to quickly compare the sizes of files in one list by glance. This is something that I have to do often enough that it has prevented me from adding -h to my ls alias. I'd have to use it for a bit to be sure, but the post's suggestion seems a pretty good 'best of both worlds' solution to me.
- dredmorbius 14y agoThis. I run into the situation more often with 'du' (when trying to find which subdirectory tree has excess junk in it), to the point that, while 'du -h' is human readable, it's not particularly sortable so: du -hs $( du -s * | sort -k1nr,1 -k2 | head ) .... which will return the human-readable output, based on numerically sorting the full numeric output. Eyeball comparisons are easier as you're aware that results are already sorted by size.
- pwg 14y agoA reasonably recent sort from GNU coreutils has a -h (and equivalent long option --human-numeric-sort) which properly sorts the output of du -h, meaning you can do: du -h | sort -h And get properly size-sorted output.
- dredmorbius 14y agoTIL! Thanks.
- zapman449 14y agoOr, you could use 'ls -h'... (that said, I do see the utility, since it gives a more obvious visual queue as to the order of size differences... but if you're doing anything with the sizes programatically, you have to remove the commas afterwards... Short version: if you're going to do this, make it a unique flag, or a new flag modifier to the -l flag... don't overload the -l flag without recourse...)
- Evbn 14y agoIdeally a parser should respect locale, and use a sane format (not commas) for multiple numbers in a list. Even better if the _ separator used by programming languages were a supported locale LC=C_FOR_HUMANS :-)
- guylhem 14y agoIs this front page worth materiel? I mean, it's good you took the time to add a ', but doing the same with sed would have been faster. Alternatively, do you know about ls -lhrS? It will print size in human formats and reverse sort the files by size - ie the bigger will be at the end of the list
- drp4929 14y agoLast sentence in the post says it all. Hacking is more than just coding and the steps he followed are illustrative in general.
- memset 14y agoYes, it is front-page worthy, because the details give us insight into how one might make any sort of change to a codebase like this. "If you want to be a hacker, these are specific steps I took to hack something up." In a domain that most of us are not very familiar with!
- guylhem 14y agoEditing ls/print.c sourcecode and recompiling is not hacking in my definition. I usually consider the portability of the solution. I have linux i386 and x64 machines, my arm n900, an osx latop, etc. Recompiling ls (or, heavens forbid, cross compiling!) for each machine may be a bright idea. Adding a line to your profile that will take advantage of the existing tools like sed is closer to hacking in my definition, because it tries to think about the bigger problem - but still I wouldn't dare calling the following "hacking": echo "alias lll=\"ls -l | sed -e :a -e 's/\(.*[0-9]\)\([0-9]\{3\}\)/\1,\2/;ta'\"">> ~/.bashrc
- dalke 14y agoIt's definition under my definition of 'hacking.' Consider that the GNU tools already implement non-portable extensions; this would be yet another one, were it added. Also, portability is not an essential requirement for "hacking." I do small programming exercises sometimes in order to understand a facet of how things work. Consider this as an exercise in how to use locale-dependent specifiers. As I pointed out elsewhere, your sed code doesn't work because it changes too many numbers - including filenames and dates - in the output. Also, if the exercise is to understand localization then your sed code isn't appropriate because it hard-codes "," when some locales use a "." as the thousands separator. And for completeness, your alias can't then mix and match other flags, like "ls -lRt". It's a single command with strange side-effects if you use it incorrectly: % lll -art sed: illegal option -- r usage: sed script [-Ealn] [-i extension] [file ...] sed [-Ealn] [-i extension] [-e script] ... [-f script_file] ... [file ...]
- thaumasiotes 14y agoSurely the appropriate option character for this new, human-readable output is "-h". Makes you wonder whether anyone ever considered the problem before...
- chris_wot 14y agoHuman readable form in bytes? I thought it you had a file greater than a megabyte hen it shows in MB? What if you want to read it in bytes with the thousands separator?
- memset 14y agoGuys! The point of this article is not to prescribe the only method of displaying human-readable file sizes. Obviously one could use `ls -lh`; the author clearly demonstrates that he is willing and able to read man pages to find answers. Rather, this is a pretty interesting look into what it actually entails to make what ought to be a very simple and straightforward change. It turns out that these simple changes are hard! Not just in identifying the piece of code to modify, but that man pages are often incomplete or unclear. It also illustrates the complexities behind making software portable - in this case, using the nation-neutral place separator. It also reminds us that solving what is on the surface a simple problem lets one uncover all sorts of interesting and messy details underneath - including more problems to solve! These are steps that he'd have to take no matter what the code or feature. This article is not "complexity for complexity's sake", it's illustrating the complexity of making changes to any piece of code - and that it is surprisingly difficult for something that one would think is very easy!
- guylhem 14y agoIt is not very easy because it is an unusual request. But I still wonder if this is easier than ls -l | sed -e :a -e 's/\(.*[0-9]\)\([0-9]\{3\}\)/\1,\2/;ta' ?? The investigations would be interesting it they were more complete, i.e. if the actual result was a change in the locale which could be appliable to other tools printing numbers besides ls (in the author TODO). I mean, will it work with bc? At the moment it's not better than an shell script alias giving the output to sed, but it is more complex - you have to recompile a binary for every OS you use.
- dalke 14y agoWhen I use your sed transformation on one of my directories, I see entries like: drwxr-xr-x 1,653 dalke admin 56,202 Mar 5 2,012 pubchem -rw-r--r-- 1 dalke staff 59,252 Nov 16 2,011 pubchem_10,000.fps.0.9.cluster I don't want to see the year written "2,012", and the file name is 'pubchem_10000" not "pubchem_10,000".
- guylhem 14y ago
- 3amOpsGuy 14y agoMountain, meet molehill.
- secure 14y agoIn case anyone wonders, I recently looked into how locales work with respect to LANG, LC_ALL and LC_*: http://c.i3wm.org/6799926 http://c.i3wm.org/6799926 By the way, by looking at http://www.lemis.com/grog/index.php http://www.lemis.com/grog/index.php you can see that the author uses FreeBSD, just in case you were wondering about /usr/src
- pixelbeat 14y agoI enable this for GNU ls like: alias ls="BLOCK_SIZE=\'1 ls --color=auto" The above is a bit hacky and not very UNIXy as it's lumping more logic into ls, rather than splitting out into functional units. Number formatting being a very common requirement, I've proposed a design for a new numfmt GNU coreutil http://lists.gnu.org/archive/html/coreutils/2012-02/msg00085.html http://lists.gnu.org/archive/html/coreutils/2012-02/msg00085... which would be used like: ls -l | numfmt --field=5 --format=%'d
- Aardwolf 14y agoI really don't understand why block size is 512 by default. It should really be 1 by default. Except for someone with an ancient hard disk who thinks in blocks instead of (mega, giga, etc...)bytes, who ever needs or wants that?
- alayne 14y ago512 byte blocks for ls is specified by POSIX. That would probably be a painful change now.
- fafner 14y agoWhy not alias ls="ls --block-size=\'1 --color=auto"
- sophiabatka464 14y agonice movie http://freemoviesite24.com/2012/09/get-the-gringo-2012-dvdrip-free-download http://freemoviesite24.com/2012/09/get-the-gringo-2012-dvdri...
- osteele 14y agoMr. Lehey managed to improve the system in such a way that it will subsequent changes for him and others easier, independently of whether the specific change to `ls` is never adopted. It's “five whys” applied to “why is this hard” and “how can I make it easier”. It's more effort, with a greater chance that much of it will survive the current context and requirements. Some people improve the area they travel through, others leave debris, and many are noops who make no difference to those who come after. If there's not enough entropy fighters like Mr. Lehey working a system, it turns to kipple.
- jrockway 14y agoMost annoying is that gcc warns about perfectly valid and logical code. That causes people to ignore warnings, and before you know it, you have a piece of software that has more warnings than lines of code. Alternatively, when you cleverly figure out how to work around the warning, like the author does, you now prevent that rule from triggering even when it's right. Clearly a better unit test is needed.
- Evbn 14y agoThat warning seems like a nasty hack anyway, if the compiler can't inline a local when running safety checks. It is super scary that the compiler appears to be using a different constant from printf for its format checker, that shows it probably isn't using a pattern supplied by printf.
- jrockway 14y agoYes to both points. I haven't read the source code, but this feels like a case of code-oriented-programming instead of data-oriented-programming. In other words, they write printf twice: once in the C library, and once in the warning system. A more careful programmer might write both in terms of a ruleset that's declared a single time. (Or they did that and it's just a bug somewhere.)
- EdiX 14y agoCompiler and standard library are two separate codebases, in fact gcc gets routinely used with standard libraries other than GNU's. Reimplementing printf parsing probably is the cleanest solution.
- jrockway 14y agoNo reason that glibc can't include a validate_format_string routine to be run at compile-time by gcc. There are already so many conditional compilation sections in both codebases that another #ifdef GNU in gcc isn't going to hurt anyone :)
- 14y ago
- righyeah 14y agoImagine if all UNIX users had taken this approach - making modifications to suit their own tastes, and writing a public diary entry, instead of making their own changes _and_ then trying to get them into the source tree so every UNIX user has them by default, regardless of whether they need/want them? Oh well. Too late now. (There are still people trying to add their own personalized features into UNIX source distributions... watch out for them.) There is something very appealing about having these utilities be simple enough that you can quickly hack them to add some feature that you might need. They are arguably much more useful as basic building blocks than as complex programs that purport to handle any task. When they start to become complex, with many features (cat -v?), that simple quick and dirty hacking, adding a little ad hoc feature, becomes more involved. Who knows, by the time you finish, you may want it added to the base system for everyone, to justify the time you spent! (I'm guessing here. I honestly don't know why people try to push their added features into the base distributions of UNIX utilities that everyone must use.)
- dsr_ 14y agoNow let's consider software lifecycle in a large context: longevity of forks. If he doesn't send the changes off to upstream, and make a case good enough for them to be approved, then all this dooms him to maintaining his fork on all the platforms where he wants it until he gets sick of it or convinces someone else to do it for him.
- cjg_ 14y agoFYI, Greg Lehey is a longtime FreeBSD committer.
- Aardwolf 14y agoMan, such fragile stuff. Why not code a function yourself that turns a number into a string representing it decimally with the commas every three digits. I normally like and use good library functions and standards, but if they're that fragile and depend on your environment then no thanks.
- michael_h 14y agonot everyone uses commas to separate their number groupings. his solution will work for any locale.
- Aardwolf 14y agoYes, in my country we use spaces to separate thousands and "," is used for what the decimal point does in the US. In my handwriting I use that notation. However, from a computer, I (and I'm certain I'm not the only one) actually expect it to output the US notation. I'd much rather have a computer always output the same format (and that happens to be the US format), than try to be smart with locales, when the end result is that some things will do this, others that. Makes stuff harder to use, and when programming, harder to parse. I've once had to touch Excel on a Windows machine configured for a non-US language, and it refused to import a CSV file that had commas, even though CSV means comma separated values. It required semicolons due to the locale settings of Windows. This stuff should not happen. A CSV is meant for computers, and to be interchangeable, not to use different types of commas and refuse to work with other types depending on user locale settings... Of course, when publishing or printing, that's a whole different matter, and there it better get the locale of your country perfect. But this here was about output in the console, which is often meant as input of other scripts etc...
- barrkel 14y agoBecause that would not be correct in Germany or other locales which use . for digit group separation and , for the decimal separator.
- codegeek 14y agoi always use one hack for ls. alias lsd="ls -ltrF | grep ^d" This way, I quickly run lsd to only look for directories.
- bartv331 14y agols -alh You guys have to much time.
- lucian303 14y ago-h
- chris_wot 14y agoNot for files greater than a MB.
- joeyh 14y agoIncidentially, I completed ls's set of -a-z options recently. http://joeyh.name/~joey/blog/entry/ls:_the_missing_options/ http://joeyh.name/~joey/blog/entry/ls:_the_missing_options/ (Well, actually, I never got around to writing -z, but it's clear what it should do, and any ls hackers are encouraged to finish that up.)
- haldean 14y agoYour link is broken, FYI
- alexpeattie 14y agoI think it should be: http://joeyh.name/blog/entry/ls:_the_missing_options/ http://joeyh.name/blog/entry/ls:_the_missing_options/
- al1x 14y agogobble.wa@gmail.com made a similar post to the freebsd-questions mailing list a month ago. In his case the question was how to print an md5sum along with the file names in a given directory. I saved it because I thought it was a clever hack. http://lists.freebsd.org/pipermail/freebsd-questions/2012-September/244933.html http://lists.freebsd.org/pipermail/freebsd-questions/2012-Se... A lot of times I catch myself in the mindset of taking a step back and saying "here are the set of tools I have at hand to accomplish a task" without realizing that I should simultaneously be taking a step "in"--so to speak--and acknowledging that the tools I have to work with are not immutable tools cast of iron; they are malleable and can be re-tooled to suit my purposes.. and that sometimes going that route can be the simplest--and in fact "best"--solution.
- deleted 14y ago[deleted]
- jtgeibel 14y agoFor such large numbers, would it make more sense to use groups of 6 instead of 3? This would allow you to easily identify the megabyte position with the next separator at the terabyte position.
- rcthompson 14y agoHere's a wrapper I wrote for ls a while ago that allows you to spell "--color" as "--colour": http://ubuntuforums.org/showthread.php?t=684239 http://ubuntuforums.org/showthread.php?t=684239
- njharman 14y agoMan, this sounds like every change I try to make to "legacy" code. There's so much debt and smell. I find it very, very hard to leave alone.
- meyering 14y agoFYI, there is no need to change GNU ls to get that behavior. You can make it use your locale's separator with either the --block-size="'1" option or by setting the LS_BLOCK_SIZE envvar to that same string: $ LC_ALL=en_US.UTF8 ls -og --block-size="'1" . -rw-------. 1 5,145,416 Oct 5 16:44 A -rw-------. 1 5,137,692 Oct 4 14:37 B -rw-------. 1 5,147,168 Oct 8 07:52 C This feature is documented in the "Block size" section of the coreutils manual: i.e., you can type this to see it: info coreutils 'block size'