18 ms·
When I say “alphabetical order”, I mean “alphabetical order”
- Someone 1y agohttps://www.unicode.org/reports/tr10/#Contextual_Sensitivity https://www.unicode.org/reports/tr10/#Contextual_Sensitivity: “There are additional complications in certain languages, where the comparison is context sensitive and depends on more than just single characters compared directly against one another, […] Numbers. A customization may be desired to allow sorting numbers in numeric order. If strings including numbers are merely sorted alphabetically, the string “A-10” comes before the string “A-2”, which is often not desired. This behavior can be customized, but it is complicated by ambiguities in recognizing numbers within strings (because they may be formatted according to different language conventions). Once each number is recognized, it can be preprocessed to convert it into a format that allows for correct numeric sorting, such as a textual version of the IEEE numeric format.”* I think those file browsers made the right choice, even given that they don’t (as in this example) always do the right thing.
- afrisch 1y agoBut -10 is smaller than -2, right?
- JoshTriplett 1y agoFilenames rarely have negative numbers in them, and it'd usually be ambiguous whether they were negative or dash-separated positive.
- pwdisswordfishz 1y agoThat's a hyphen, not a minus sign, silly.
- ZoomZoomZoom 1y agoI know you jest, but this just further demonstrates why Natural Sorting is complicated and might not be the best default choice. my_photos_at_-3c my_photos_at_-10c Do users want smaller numbers first, or do they want them in counting order, away from zero?
- JoshTriplett 1y agoI almost always want the version-sorting that's being presented in this article, rather than an "alphabetical" sort. But on the other hand, it absolutely seems like a valid bug that this is presented as an "alphabetical" sort, rather than something like "alphabetic/numeric" or similar. In other words, a problem of labeling rather than one of sorting.
- lisper 1y agoYeah, exactly. The behavior described is actually very useful. The problem is imposing it on the user with no warning or option to turn it off.
- sebtron 1y agoAuthor here - I Agree both with you and with the parent's comment. Having two options in the "sort by menu" - like "Name (natural)" and "Name (strict)" or something - would have solved everything.
- parineum 1y ago> The problem is imposing it on the user with no warning or option to turn it off. You can say that about every single design decision made about every product. The gripe about this particular feature seems misplaced because almost all users will want the sort that's offered and the actual alphabetical sort is likely the desire of a more advanced user who, in fact, is offered a choice through registry editing and/or using a more advanced cli option for the occasion they might need an alternative sort. This is a sensible default.
- jmclnx 1y agoIf I understand the article, the author wants magic :) I take it to mean they want the system to know file_9.txt is less then file_10.txt. I never saw that happen in any OS, so I do not know what he is referring to. Maybe whatever that old system was, it sorted by create time as opposed to file name. So, the author can try and create "aisort" that will look at all file names and add leading zeros to the file numeric portion, sort, then remove the zeros added. That will probably as slow as s***t and use gobs pf memory, depending on the number of files.
- anopenben 1y agoNo the author is saying the opposite. They expect file9.txt to be after file10.txt, but it many modern operating systems, it isn’t!
- jmclnx 1y agoReally, I do not know how I missed that :) I read it a couple of times to and I still thought he wanted it the other way. So my original comment kind of stands but in a opposite way. I have never see file_9.txt sorted before file_10.txt, I just tested it on OpenBSD and I got this, which I have always seen: $ ls|sort file_1.txt file_10.txt file_12.txt file_2.txt file_20.txt file_3.txt file_9.txt
- sebtron 1y agoAuthor here - My surprise stems exactly from the fact that for the last few years I have exclusively managed my files via a the UNIX shell, which behaves in the classical way.
- zahlman 1y agoWhen I started using Linux as my daily driver after many years of Windows (but with familiarity with UNIX systems going way back), I knew it would be like that in the terminal, but it still took some adjustment. But actually, Nemo does the same "natural sort" thing, and also sorts case-insensitively.
- 1970-01-01 1y agoDigits and any other characters that are not A-Z and a-z should not get sorted. That's the true result of doing what you asked and not what you meant. Pedantic, but that's why we are here.
- orphea 1y ago> Digits and any other characters that are not A-Z and a-z should not get sorted. Do you suggest that sorting in any language other than English should be broken?
- 1970-01-01 1y agoNo.
- pbw 1y ago"I created the Alphanum Algorithm to solve this problem. The Alphanum Algorithm sorts strings containing a mix of letters and numbers. Given strings of mixed characters and numbers, it sorts the numbers in value order, while sorting the non-numbers in ASCII order. The end result is a natural sorting order." https://web.archive.org/web/20210207124255/http://www.davekoelle.com/alphanum.html https://web.archive.org/web/20210207124255/http://www.daveko...
- JoshTriplett 1y agoThere are many older instances of that, such as "versionsort" from various Linux tools and libraries. I think this has likely been independently recreated several times, with various subtle differences.
- HeavyStorm 1y agoWhen you get a bunch of files (let's say 1000+) without leading zeroes, this is a blessing. But I get the author's frustration, the expected behavior is not there, instead, he gets magical sorting that is wrong for his use case. I'm not sure what the ux should be, and maybe the algorithm here could be smarter, but it's a trade-off.
- nebezb 1y ago> the expected behavior is not there The expected behaviour is ambiguous (and thus subjective). Older versions of windows shipped with alpha sort. New versions ship with natural alpha sort. According to the UX designers over at Microsoft (and surely the user feedback), natural sort _is_ the expected behaviour. I certainly agree with natural sort being the expected behaviour too.
- dfxm12 1y agoI get it, but if all these major operating systems are handling this same ambiguous [0] situation in the same way, perhaps one needs to reevaluate their mental model or expectations. Am I out of touch? No, it's the operating systems who are wrong 0 - numbers are not part of the alphabet.
- nenenejej 1y agoSort by the time the photo was a taken in the metadata?
- vachina 1y agoThat’s why I don’t even bother with the file name for photos. 1. Sync all equipments to the same clock. 2. Sort by Date Taken, if unavailable, sort by Date Created.
- alain94040 1y agoYes, sounds to me like the user really wanted to sort by time created. And got used to sorting alphabetically as a poor proxy for that.
- meindnoch 1y agoI thought this was pretty well known. E.g. the macOS Foundation library even exposes NSString.localizedStandardCompare() [1] which implements the sorting algorithm used by Finder, and should be used by any well-behaved macOS application. Windows uses StrCompareLogical [2]. [1] https://developer.apple.com/documentation/foundation/nsstring/localizedstandardcompare(_ https://developer.apple.com/documentation/foundation/nsstrin...:) [2] https://learn.microsoft.com/en-us/windows/win32/api/shlwapi/nf-shlwapi-strcmplogicalw https://learn.microsoft.com/en-us/windows/win32/api/shlwapi/...
- nielsbot 1y agoI tried it just for kicks. The Finder sorts these as: IMG_20250820_095716_607.jpg IMG_20250820_103857_991.jpg IMG_20250820_103903_811.jpg IMG_20250820_055436307.jpg IMG_20250820_092016029_HDR.jpg IMG_20250820_092440966_HDR.jpg IMG_20250820_092832138_HDR.jpg Whereas `ls -l` gives me IMG_20250820_055436307.jpg IMG_20250820_092016029_HDR.jpg IMG_20250820_092440966_HDR.jpg IMG_20250820_092832138_HDR.jpg IMG_20250820_095716_607.jpg IMG_20250820_103857_991.jpg IMG_20250820_103903_811.jpg
- freetime2 1y agoI would have assumed it worked the same as ls, so I found the article interesting. But now that I know, I think this way is better. I can’t think of any case where I would need purely alphabetical sort. In most photo browsing apps, photos will be sorted by timestamp rather than filename. If I really needed it to sort properly in file explorer, I would try sorting on created date. And failing that I would probably just normalize the file names.
- furyofantares 1y agoNumbers aren't in the alphabet. So no, you don't mean alphabetical order.
- furyofantares 1y agoI felt a little bad about this snark but actually, author barely understands their own use case (says they want alphabetical order but they actually want something more) and barely understands the UI they're using (says they asked for alphabetical order but none of the file managers they used says it has any such setting) and then they go on to claim this is to satisfy dumb users: > Well, apparently all these operating systems have decided that no, users are too dumb and they cannot possibly understand what alphabetical order means.
- crazygringo 1y agoIt really is world-class irony. Impressive.
- croes 1y agoThere isn't the alphabet
- xerox13ster 1y agothis is an ID-10T PEBKAC ERR. Not this keyboard not this chair, but the problem is with idiots between keyboards and chairs. The author is not the ID10T it’s the other general users. The author is intelligent enough to recognize that this is not alphabetical sort, but the term that they are looking for to describe the sort that they see in dolphin windows, google etc. is *lexical* sort, not alphabetical. The engineering problem is ID10Tic not technical. How do you educate an illiterate public on what the difference between alphabetical and lexical sort is in practice? You can’t, so you engineer around it and call lexical sort alphabetical.
- armchairhacker 1y agoI agree with Microsoft/Google/KDE's order. The author's situation is extremely rare, and the situation where someone wants "10" to be before "9" is far more common. Moreover, desktops don't label this sorting "alphabetical" (E: and it would really be "lexicographic"*), they label it "by name" (an informal criteria), so technically they're not lying. > I miss the time when computers did what you told them to, instead of trying to read your mind. You may be looking at that time through rose-tinted glasses. I don't like when computers lie to me either, but "mind-reading" is really helpful in ways we take for granted, like autosave. Desktops can have an option to sort files truly alphabetically, but the more common case should always be the default; that's the definition of "intuitive". * https://news.ycombinator.com/item?id=45404022#45405279 https://news.ycombinator.com/item?id=45404022#45405279
- zweifuss 1y agoYou mean file9 before file10? I have some beef with microsoft, that you can only change this at the Computer level, not per user (see registry key below). Also they call it natural sorting for users, but logical sorting internaly. Unify your termini! [HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows\CurrentVersion\Policies\Explorer] "NoStrCmpLogical"=dword:00000001
- nerdile 1y agoTo change it per user, set it in the user's hive instead of in the local machine hive (e.g. HKEY_CURRENT_USER instead of HKEY_LOCAL_MACHINE)
- yegle 1y agoTIL they are called "hives". Windows Registry is an interesting thing. Even casual users have to interactive with it once or twice w/o fully understand it. https://learn.microsoft.com/en-us/windows/win32/sysinfo/registry-hives https://learn.microsoft.com/en-us/windows/win32/sysinfo/regi...
- 1y ago
- ineedasername 1y agoWell, lots of interfaces don’t say “alphabetical” anymore, they say “name” or some variant, and then they can define it however they want, regardless or because of the frustration it causes users but not some other users which will now be inverted for long term-frustration averaged user experience.
- nebezb 1y ago> But 1 is smaller than 9, so file-10.txt should be first in alphabetical order. Everyone understands that, and soon people learn to put enough leading zeros if they want their files to stay sorted the way they like. No. Not “everyone understands that”. Natural sort happens in real life and everyone understands that. Only those who understand ASCII — not the average user of graphical file managers — will deduce the reason for your definition of “alphabetical order”. > Now that I know what the issue is, I can solve it by renaming the files with a consistent scheme. Intensely ironic given the previous suggestion.
- zahlman 1y ago> Well, apparently all these operating systems have decided that no, users are too dumb and they cannot possibly understand what alphabetical order means. So when you ask them to sort your files alphabetically, they don’t. Instead, they decide that if some piece of the file name is a number, the real numerical value must be used. Well, no. You don't actually ask them to sort in alphabetical order. You ask them to sort "by name", and that is up to their interpretation. And they choose the interpretation that (per their reasoning, and possibly some actual data) seems most likely to correspond to what the user wants. Maybe future versions of those OSes will add a rule that says that if any of the number groups have leading zeros then it reverts back to actual alphabetic order. Or maybe they'll give you configurable options. (Maybe some of them already do.)
- sebtron 1y ago> And they choose the interpretation that (per their reasoning, and possibly some actual data) seems most likely to correspond to what the user wants. Yes, that make sense, but the problem is that this interpretation changed in the last 10 (15? 20?) years. It used to be that "by name" meant "by name, il alphabetical / lexicographical order" in pretty much every file manager.
- deleted 1y ago[deleted]
- KuSpa 1y agoReminds me of https://xkcd.com/1172/ https://xkcd.com/1172/
- johanyc 1y agonice. first time seeing this
- pseudalopex 1y agoMicrosoft and Apple changed to natural order in 2001.
- 1y ago
- shawnz 1y ago> Of course, the user who named those files probably wants file-9.txt to come before file-10.txt. But 1 is smaller than 9, so file-10.txt should be first in alphabetical order. Everyone understands that, and soon people learn to put enough leading zeros if they want their files to stay sorted the way they like. Well, apparently all these operating systems have decided that no, users are too dumb and they cannot possibly understand what alphabetical order means. So when you ask them to sort your files alphabetically, they don’t. Instead, they decide that if some piece of the file name is a number, the real numerical value must be used. I think there are many things wrong with your assessment of the situation. First, where does it say in these file managers that they're sorting by alphabetical order? I see that you've specified that you want the files sorted by name, but I don't see that you've specified you want them sorted by name alphabetically. And what does "alphabetical sort" even mean when you're sorting characters which are not letters? What you mean is probably "lexicographical sort". Second, you admit yourself that users probably want natural sort. Why would you expect these products to do the thing which they know users usually don't want by default? That just seems like bad design to me. They know users usually want natural sort, and you know users usually want natural sort, so why would you expect the default behaviour to be a lexicographical sort? Third, just like how you've learned to work around the lack of natural sort in poorly designed products of years past by adding leading zeroes, you can just add trailing zeroes to get the lexicographical ordering that you want. Why do you seem to be implying that the latter is more user-hostile than the former? It doesn't make sense to me. A decision had to be made about what sort to use and they picked the one that most people want. Isn't that what we should be expecting in a product that caters to its users? I see in other comments you've suggested that there should be a separate option for choosing between lexicographical sort and natural sort. But in the past, when lexicographical sort was the only option, why weren't you complaining about it being user-hostile to only have one option then? Why is it only when the default is something you're personally not used to that it warrants complaint? And where do we stop, do we have separate controls for every single sortable string field to determine whether it should be sorted lexicographically or naturally? Or just the name field? Don't you think that is going to lead to interface bloat?
- wrs 1y agoTo answer the question in the article, I’m pretty sure Windows Explorer (and probably File Manager before that) has sorted filenames this way for at least 30 years.
- rozab 1y agoI can confirm that this does not happen in Windows 98, but does happen in Windows XP.
- deleted 1y ago[deleted]
- epistasis 1y agoThis is reminding me of the whole "Worse is better" essay and debate: https://news.ycombinator.com/item?id=27916370 https://news.ycombinator.com/item?id=27916370 The author wants the "worse" sort, one based on ASCII/Unicode codepoints, without any intelligence for numbers that 99% of GUI users want. For their purposes, they've assumed something about the implementation, to the point that a convenience feature is actually a misfeature for them. But the author here is probably a developer, or close to one, so they do not represent the needs of most people using computers. Understanding the target audience for your product results in very different design decisions. Better is better might be great for products, but worse is better is probably better for systems that need to grow and evolve.
- BeFlatXIII 1y ago> The author wants the "worse" sort, one based on ASCII/Unicode codepoints, without any intelligence for numbers that 99% of GUI users want. I want the author's opinion on how caplital and lowercase letters should be sorted. Do they follow strict ASCII/Unicode codepoints, or do they normalize into actual alphabetical order and sort upper/lower within each letter?
- jowea 1y agoAsciibetical sorting
- _9ptr 1y agoAnd where do you sort the letter ä? (After a is correct in German, but I think Swedish does it differently.)
- dvdkon 1y agoThis feels like the right moment to mention "ch", which is considered a letter in orthodox Czech, sorted between "h" and "i". The problem is, you can't reliably distinguish between "ch"-the-letter and "ch" as just "c" and "h" combined, which are present in loan words but also some original Czech compound words. So if you're doing it "properly", sorting strings in Czech involves understanding the etymology of every word.
- dcomp 1y agoI think the algorithm is probably incorrect. A number starting with 0 should be treated lexically not numerically. Otherwise you have a situation where img_1_01.jpg and img_01_1.jpg does not have a complete ordering.
- gregates 1y agoIt wouldn't be the first time widely-used software sorted numbers by a function that does not produce a total ordering. For example, Excel: https://gregat.es/excel-numeric-order-transitivity/ https://gregat.es/excel-numeric-order-transitivity/
- re 1y ago> Otherwise you have a situation where img_1_01.jpg and img_01_1.jpg does not have a complete ordering. (Good) "natural sort" implementations generally have ways of handling ties like this. It's similar to the problem of case-insensitive sort over case sensitive sets.
- crazygringo 1y agoThat's not the issue. The issue here is that one camera appends milliseconds to the seconds without a separator, and the other uses a separator. So of course the ones that include milliseconds look like bigger numbers and get sorted last. Leading zeros aren't the issue here.
- ano-ther 1y ago"The Tyranny of the Marginal User" strikes again: https://nothinghuman.substack.com/p/the-tyranny-of-the-marginal-user https://nothinghuman.substack.com/p/the-tyranny-of-the-margi...
- donalhunt 1y agoI've encountered a tangential problem to this with package versioning on Linux distros. Thankfully it was not too hard to write an algorithm to compare versions (thanks AI!).
- skwee357 1y agoI got used to naming files/folders with leading zeros when I want them to be sorted alphabetically (for example payslips/invoices, etc). But I'm a tech guy, I know what does "alphabetically" mean in the tech world. And it probably is not what common folks mean when they think "alphabetically" outside the tech world. Edit: in fact, if I recall correctly, the proper term for this kind of sort (the one OP wants) is alphanumeric sort.
- vslira 1y agoshameless plug (though I don't get a cent out of this, of course): https://blog.vslira.net/2025/03/a-neat-approach-for-sortable-versioned.html https://blog.vslira.net/2025/03/a-neat-approach-for-sortable...
- lblume 1y agoI also got used to it, but especially when writing short scripts that generate numbered files it gets annoying to have to pad with zeroes every time, and also precommit to a specific amount of digits you want to allow (finding a compromise between adding a ridiculous amount, like 20, and using only 4 despite knowing the script might one day surpass 10⁵ files). The natural numbers are ordered. Let me use its ordering instead of having to rely on an ad-hoc lexicographic fixed-length tuple representation of decimal digits, without any padding. My position is that numbers in filenames should always be considered atomically unless explicitly instructed otherwise. If there were no issues of backwards compatibility, I would thus advocate for changing ls. Eza (maintained fork of Exa, Rust-based ls alternative) actually does sort this way by default, much to my delight.
- tonytamps 1y agoWhen I say Afferbeck Lauder [1] I mean "alphabetical order" [1] https://en.m.wikipedia.org/wiki/Afferbeck_Lauder https://en.m.wikipedia.org/wiki/Afferbeck_Lauder
- adolph 1y agoThose are beautiful, thank you for posting them
- perching_aix 1y agoTLDR: the author found out about natural ordering [0], i.e. treating a sequence of digits as a number while sorting. Usually preferable, except when not. Just like distinguishing between upper- and lowercase letters, and other misery. [0] https://en.wikipedia.org/wiki/Natural_sort_order https://en.wikipedia.org/wiki/Natural_sort_order
- thaumasiotes 1y ago> Usually preferable, except when not. Just like distinguishing between upper- and lowercase letters When would it be preferable to distinguish between capital and lowercase letters? Wiktionary does it religiously, and it always makes the entries worse. Want to know what something means in German? Well, that's on a separate page. Do you want to look something up while using your phone? Don't be stupid; use a desktop that won't autocapitalize the first letter you type in.
- perching_aix 1y agoWhen you're dealing with unix-y Git repositories for example. If you mean more from a user perspective, it really depends. For registry keys for example, since they're interacted with programmatically for the most part, I was expecting them to be case-sensitive. They're case-insensitive though, so that was a bit of a whiplash.
- cesarb 1y ago> But nope, this is not it, because the good old ls sorts my files correctly Did the author try "ls -v"? It would probably give the exact same order these file managers used.
- echohack5 1y agoThis was a fun thing to realize in my early days of programming in Delphi. I guess the author will soon realize why old systems name things ticket "00001" and so on.
- vslira 1y agoHa! I had the exact same realization on MacOS. Extremely annoying behavior.
- cubefox 1y agoBy the way, there seems to be a "standard" way to sort strings: > Unicode Technical Report #10 also specifies the Default Unicode Collation Element Table (DUCET). This data file specifies a default collation ordering. https://en.wikipedia.org/wiki/Unicode_collation_algorithm https://en.wikipedia.org/wiki/Unicode_collation_algorithm I assume this mainly aims at giving a reasonable compromise between the different dictionary and phone book sorting rules of various languages (and even locales), which should give reasonable results for most languages. I assume this also puts "Alice2" before "Alice10".
- t-3 1y ago> I assume this also puts "Alice2" before "Alice10". It doesn't (per https://www.unicode.org/reports/tr10/#Non-Goals https://www.unicode.org/reports/tr10/#Non-Goals): > 1.9.2 Non-Goals > > The Default Unicode Collation Element Table (DUCET) explicitly does not provide for the following features: > [ ... ] > Numeric formatting: numbers composed of a string of digits or other numerics will not necessarily sort in numerical order.
- cubefox 1y agoOh, that surprises me.
- zarzavat 1y agoEven if you are a file naming Einstein and you always zero pad your integers to exactly right length, we have this thing called the internet, where you can download other people's files. Your OCD is not my OCD.
- deleted 1y ago[deleted]
- mcswell 1y agoBe glad you don't have to deal with non-ASCII characters: acute/grave/tilde/umlaut/diaresis/etc. accented characters, dotted vs. dotless 'i' (Turkish), barred i (an 'i' with a sort of dash through the middle, used in some languages for a sort of schwa-like vowel), thorn, not to mention non-Roman characters. And different languages sort the same characters differently, so you can't just pay attention to their Unicode values. (@cubefox has a post here pointing to the Unicode Consortium's doc about sorting)
- deleted 1y ago[deleted]
- t-3 1y agoRenaming things to make them queue correctly (I usually couldn't care less about visual sorting, I use a terminal) is by far my #1 task by LoC and frequency of occurence, and by far the most annoying. Metadata can be very helpful to obviate this issue, but it usually just leads to another problem where you now need metadata editors and readers in addition to the "user-visible" name metadata. It's frustrating.
- kens 1y agoSorting so "foo9" is before "foo10" is called natural sort. I found out about natural sort a week ago and I am thrilled that my programs now print their output in a sensible order. Give natural sort a try and see if it improves your life too :-) I found the magic two lines of Python to do a natural sort here, by the way: https://stackoverflow.com/questions/11150239/natural-sorting/11150413#11150413 https://stackoverflow.com/questions/11150239/natural-sorting...
- jcynix 1y agoNatural sort is an Option in sort(1): for i in $(seq 2 10) ; do touch img_$i-hn.txt done ls img_* | sort -V img_2-hn.txt img_3-hn.txt img_4-hn.txt img_5-hn.txt img_6-hn.txt img_7-hn.txt img_8-hn.txt img_9-hn.txt img_10-hn.txt And we have "sort -h" to sort the output of e.g. "du -sh *" properly. Edit: formatting and add sort -h
- Marha01 1y agoWhy is the author so perplexed by it? https://en.wikipedia.org/wiki/Natural_sort_order https://en.wikipedia.org/wiki/Natural_sort_order It's simply natural sorting. I don't see what is so controversial about it.
- lblume 1y agoI find the term "natural" to be inadequate here, there is nothing natural about sorting strings in this particular fashion compared to another. It should be given a more descriptive name, like "Number-aware alphabetic" or something like that as to actually give a hint about what it does.
- DonHopkins 1y agoSorting by name (collation) is waaay tricker than simply figuring out how to parse the numbers. The International Components for Unicode library implements the Unicode Collation Algorithm, which depends on the language code and region of the locale, and looks up the quirks for each locale in the Common Locale Data Repository. It's a much better idea to just use the standard ICU library or platform specific libraries (which are often build on ICU like JavaScript's Intl.Collator), instead of trying to hot dog it by rolling your own. International Components for Unicode https://en.wikipedia.org/wiki/International_Components_for_Unicode https://en.wikipedia.org/wiki/International_Components_for_U... >ICU provides the following services: Unicode text handling, full character properties, and character set conversions; Unicode regular expressions; full Unicode sets; character, word, and line boundaries; language-sensitive collation and searching; normalization, upper and lowercase conversion, and script transliterations; comprehensive locale data and resource bundle architecture via the Common Locale Data Repository (CLDR); multiple calendars and time zones; and rule-based formatting and parsing of dates, times, numbers, currencies, and messages. Unicode Collation Algorithm https://en.wikipedia.org/wiki/Unicode_collation_algorithm https://en.wikipedia.org/wiki/Unicode_collation_algorithm >The Unicode collation algorithm (UCA) is an algorithm defined in Unicode Technical Report #10, which is a customizable method to produce binary keys from strings representing text in any writing system and language that can be represented with Unicode. These keys can then be efficiently compared byte by byte in order to collate or sort them according to the rules of the language, with options for ignoring case, accents, etc.[1] >Unicode Technical Report #10 also specifies the Default Unicode Collation Element Table (DUCET). This data file specifies a default collation ordering. The DUCET is customizable for different languages,[1][2] and some such customizations can be found in the Unicode Common Locale Data Repository (CLDR).[3] Common Locale Data Repository https://en.wikipedia.org/wiki/Common_Locale_Data_Repository https://en.wikipedia.org/wiki/Common_Locale_Data_Repository >The Common Locale Data Repository (CLDR) is a project of the Unicode Consortium to provide locale data in XML format for use in computer applications. CLDR contains locale-specific information that an operating system will typically provide to applications. CLDR is written in the Locale Data Markup Language (LDML). >Among the types of data that CLDR includes are the following: Translations for language names Translations for territory and country names Translations for currency names, including singular/plural modifications Translations for weekday, month, era, period of day, in full and abbreviated forms Translations for time zones and example cities (or similar) for time zones Translations for calendar fields Patterns for formatting/parsing dates or times of day Exemplar sets of characters used for writing the language Patterns for formatting/parsing numbers Rules for language-adapted collation Rules for spelling out numbers as words Rules for formatting numbers in traditional numeral systems (such as Roman and Armenian numerals) Rules for transliteration between scripts, much of it based on BGN/PCGN romanization Tricky collation examples: sv-SE (Swedish): å, ä, ö are separate letters at the end of the alphabet, not variants of a or o. de-DE (German): ä, ö, ü may sort as ae, oe, ue in some contexts, or as distinct letters. ß sometimes sorts as ss. tr-TR (Turkish): dotted i (i) and dotless ı are different letters; I sorts with ı, not with i. es-ES (Spanish): traditionally ch and ll were treated as single letters with their own place in the alphabet. cs-CZ (Czech): ch still counts as a unique letter, sorted after h. da-DK / no-NO (Danish/Norwegian): ø comes after z. is-IS (Icelandic): þ (“thorn”) is part of the alphabet, after z. fr-FR (French): accents usually ignored in sorting, so é = e, but not always depending on collation settings. el-GR (Modern Greek): tonos accents, final sigma ς vs. σ, etc. nl-NL (Dutch): the digraph “ij” is often treated as a single letter, and capitalized as “IJ”. In dictionaries and phone books it often sorts as a single letter under “I”, but sometimes is listed after “X” depending on tradition. Then you get into non-Latin languages like, Chinese, Japanese, and Korean collation, which gets hairy with radicals, kana order, and stroke count. Also different locales have different ways of representing numbers, like switching between "," and "." as separators and decimal points. ICU supports integer only "natural" numeric collation, so anything more complicated like versions, floating point, negative numbers, hex, thousands separators, fractions, roman numerals, etc, you'd have to build on top of ICU. ICU doesn't support incomprehensible dead languages like Latin or Ancient Greek (it does however support French ;). It does support Roman numeral formatting, but not collation, which would be pretty tricky and ambiguous. https://www.youtube.com/watch?v=sKWvTlLMB-Y https://www.youtube.com/watch?v=sKWvTlLMB-Y A nuanced but common example that ICU/UCA/CLDR helps with is a menu to select the current locale: you have to translate each language's name into the current locale, and also sort them in the current locale. On top of different collations they can also have totally different spellings, like "United States of America" is "Verenigde Staten van Amerika" in Dutch. This makes it challenging for users to find their own language when the locale is set wrong! You just can't win. Not to mention emojis! Which comes first: The chicken or the egg? The taco or the poop? Also, the Mac Finder switches ":" and "/" for historical reasons (HFS used to use ":" as a directory separator instead of "/"), so you can create a file name like "9/11 Attack" in the Finder, which actually gets the underlying Unix filename "9:11 Attack". Don't believe me? Rename a file in the Finder to include a slash, which you know is impossible to represent as a Unix file name. Then go "ls" the directory in the shell. The Mac Finder weirdly collates "/" after "9" because under the hood it’s really storing it as ":", which sorts before "0". But it also has other punctuation collating inconsistencies, sorting "," and ";" and others after "0" too. Definitely not ASCII order -- I'm not sure what rules it uses, but it's different than "ls". However, while it's generally true you can't have "/" in Unix file names, NFS used to trustingly let clients rename Unix files to include a "/" in their name, which the Gator Box AppleTalk/Ethernet gateway let you do with the Mac Finder (pre OS/X), which would silently corrupt your "dump" backups on the Unix NFS server, so you would not learn about it until you tried to retrieve your files and "restore" crashed. https://news.ycombinator.com/item?id=31821646 https://news.ycombinator.com/item?id=31821646 >Another reason that NFS sucks: Anyone remember the Gator Box? It enabled you to trick NFS into putting slashes into the names of files and directories, which seemed to work at the time, but came back to totally fuck you later when you tried to restore a dump of your file system. >The NFS protocol itself didn't disallow slashes in file names, so the NFS server would accept them without question from any client, silently corrupting the file system without any warning. Thanks, NFS!
- nilslindemann 1y agoIn Total Commander, there is a function in the options to sort strict by numerical char code. It will sort those files correctly. Unfortunately, it will also sort "10.txt" before "2.txt". --- In all file managers, I miss an API point where one can give a userdefined sorting function for the file and folder list.
- lblume 1y agoWhat do you mean by "Unfortunately"? This appears to be the only correct conclusion from the algorithm you selected, you can't eat the cake and have it too. Regarding your second point, that's not really what a graphical file manager is for, I think. At this point (likely even earlier) you would be better off just writing a simple script in the scripting language of your choice. (If going for something fancy, you could also implement a FUSE based on symlinks for the original files, where the filename is prepended by a sort key. This would work for every major file manager and you could manipulate the files in mostly the same way as before.)
- nilslindemann 1y agoParagraph 1: I speak of a sorting method which splits the filename at the boundaries between numbers and non-numbers, and sorts by the parts of the resulting tuple, the numbers naturally (10 comes after 2) and the rest by numerical char code. Paragraph 2: I am not sure what you mean here with writing a script. The graphical file manager shall sort its file list using the sorting function I hand over to it. "That's not really what a graphical file manager is for". Says who? Every software which has a plugin system does that, why should a file manager not?
- magicalhippo 1y agoPlex team, are you reading this? For some inexplicable reason, Plex just throws its hand up on non-ASCII characters and puts them first. In Norway we have three extra letters, æøå, and they're at the end of the alphabet after z. But in Plex, I have Øystein Sunde[1] placed before any other in my music library. Now in the 1990s I would forgive US software for such a thing, but it's 2025... [1]: https://en.wikipedia.org/wiki/%C3%98ystein_Sunde https://en.wikipedia.org/wiki/%C3%98ystein_Sunde
- ryukoposting 1y agoPor qué no los dos? Call lexicographic order "sort by name" as it's called now, and call dumb character-by-character sort "plain" or something like that. I'm not a designer, maybe there are more intuitive names, but come on. This isn't an intractable problem.
- yujzgzc 1y agoI thought this was going to be a deep dive into what "alphabetical" means and how that's itself not a universal term between locales, what with so many different collation preferences.
- lblume 1y agoThat would likely have been a more useful article for the average developer. It is extremely hard to be aware of all the ways strings of different locales can defy our intuitions.
- dusted 1y agoThis got me riled up to the point where I blew a gasket and just can't.. I agree with the article. I liked computers better when everyone hated them, and for the reasons they hated them..
- lblume 1y agoI agree that the base functionality of just sorting character by character can be occasionally useful. However I would really be interested in seeing why you believe this to be the correct choice for user-facing graphical file managers, as its evident problems with typical usage seem more salient compared to the edge cases as illustrated in the article.
- dusted 1y agoBecause I value predictability more than convenience, I'm too stupid to figure out the set of rules deemed "the right way of sorting", I can't even.. the possible ways it can be implemented is nearly infinite, it makes me anxious just to think about it. Sorting by numerical value is simple, I can understand that, I know the ascii table, I can predict what happens, I know where to look for stuff, meaning I can quickly find out where stuff is, because I don't have to look a bit here, then there, and then think "oh, maybe if it starts with a letter, then has a space and then a number they will order them by word - number" but if they start with a word but has no number they won't be sorted along with those, but if they st... nope.
- Taek 1y agoThis is one of the big ways that LLMs are going to change the game for UX. Your operating system is going to have some sort of 'butler', which knows all of your preferences, and the butler will go through the APIs and man files and informational dialogs of every app you use and auto-configure them. Then if you want something to change, just ask the butler. If the app is open source and doesn't support the requested feature, the butler might even be able to code it up.
- paholg 1y agoI think the real issue here is that two Android phones take photos with incompatible naming schemes. I am sure that at some point someone thought the milliseconds should or should not be separated from the seconds and made that change without thinking through the consequences.
- bapak 1y agoMaybe it's just me but I don't miss this at all: Image-1.jpg Image-11.jpg Image-2.jpg The only time natural sort bit me was with nonsensical names like <md5>.jpg
- RHSeeger 1y agoI think it depends on the person. That order is exactly what I expect and want.
- sneak 1y agoMore importantly, it is how computers work, and how computers have worked for many decades. Anyone with experience expects them to work this way. Trying to be clever to cater to the inexperienced only harms both groups.
- crazygringo 1y agoComputers have been sorting with natural sort for decades. By now, it is "how computers work". Were you under the impression this was something new?
- f33d5173 1y agoI was very surprised by it when I noticed it a year or so ago. What's interesting is that when it works, eg you have a directory with numbers from 1-10, you don't really notice it. It isn't until it bites you in the ass, eg your downloads folder with a bunch long numeric strings, some in hex, where you want to find one and suddely it's not where you expect. I used a gui software some years ago that distinguished between version sort and alphabetic sort. It would be handy to have a toggle.
- crazygringo 1y agoYou prefer looking at photos in that weirdly particular shuffled order that isn't the order they were taken in?
- AlienRobot 1y agoI have the same problem on Nemo. More specifically, I had made a small app that displayed files of a directory in alphabetical order, and then when I look at it in Nemo it isn't the same order because I didn't implement their smart algorithm.
- lblume 1y agoI fail to see how this is a "problem"? You implemented a sorting mechanism that was useful to your application, while Nemo implemented another which as this thread demonstrates seems to be much more useful and intuitive for the average user. This is also of course not specific to Nemo, as no 'modern' file manager on Linux sorts filenames like it's 1980 and all you are able to feasibly do is step through the bytes.
- AlienRobot 1y agoNemo isn't some random app a teenager made. It's the default file manager of a desktop OS. I expect it to cover more use cases. I expect the same for other file managers on Linux. Although I must say I'm generally let down by Linux software.
- notatoad 1y ago>Well, apparently all these operating systems have decided that no, users are too dumb and they cannot possibly understand what alphabetical order means. i really really hate this framing, and i see it far too often. no, the operating system developers did not make a value judgement about their users. they observed their users to find out what behaviour was expected, and they designed the behaviour of the system to match the behaviour that the majority of users expect. and then you made an incorrect assumption about how the system works, and decided that your incorrect assumption means everybody else is dumb and you're the only smart person in this situation?
- JoBrad 1y agoI rename all of my photos upon import using the created date, formatted as `YYYY-MM-DD kk:mm:ss`. But it would frankly be great if most file browsers just let me sort photos based on metadata. But then I just end up in a dedicated photo browser, instead.
- andriamanitra 1y agoThe so-called "natural" sort makes sense for version numbers and enumeration (without zero-padding) but I'm more often dealing with file names with a datetime (like in the article), a hexadecimal hash, or just randomized string of characters that includes numbers. In those cases "natural" sort makes it harder to find the file you're looking for. Even when files are enumerated it's pretty rare to have more than 9 parts and no zero-padding, whereas there are almost always multiple consecutive digits in the use cases for which "natural" sort is not a good fit for. It just feels like a bad default, at least for a programmer's workload.
- re 1y ago> a hexadecimal hash I agree with you on this point. > a datetime AFAICT, natural sort shouldn't ever make datetimes harder to find, unless they are formatted inconsistently, as in the author's case. Suppose one camera wrote dates as 20250928 and another as 2025-09-28. ASCIIbetical sort would do nothing to help here. Natural sort can even improve things over ASCII sort, for instance if someone is stuck with a format like "28/9/2025" or "September 2 2025"
- brimstedt 1y agoIsnt the author confusing "alphabetical sorting" with "ASCII sorting"? Afaik there is no universal way to handle numbers in alphabetical lists. Sometimes numbers some before letters, sometimes after, etc. A digit is not a part of the alphabet, right?
- Wowfunhappy 1y ago> Isnt the author confusing "alphabetical sorting" with "ASCII sorting"? But it's actually not ASCII sorting either! ASCII sorting would mean 'Z' comes before 'a' and I assume even the author doesn't want that! No matter what, there are going to be hidden tricks!
- userbinator 1y agoBut it's actually not ASCII sorting either! ASCII sorting would mean 'Z' comes before 'a' and I assume even the author doesn't want that! I don't know about the author, but that's exactly what many others who know about ASCII expect, including me. Digits, then uppercase, then lowercase.
- phendrenad2 1y agoI think if we (in our industry in general) had REAL agile, and not pseudo-waterfall "the designers design it, the engineers implement it, QA QA's it, and then we lay everyone off because they're no longer needed" (but, loophole alert! We did daily standups and used Jira, so it was "agile" the whole time!), then we'd have a snowball's chance in hell of actually having a reasonable solution to this. Off the top of my head, this seems like something that should be a setting in control panel. But, because everyone assumes (contra to agile) that the designers "got it right the first time", this kind of improvement can't happen.
- jeroenhd 1y agoIf you don't like the default natural sorting order, you can just change it in Dolphin. Settings > Configure Dolphin > View > Content Display > select anything other than "natural". You can even pick if you want case sensitivity or not. The OS doesn't think you're too stupid to understand sorting, it relies on you being smart enough to figure out where the setting is located. In this case, four levels deep is probably too much to ask from users if they will write an entire blog post like this before finding the toggle.
- bloak 1y agoThe correct sorting algorithm is described here: https://manpages.debian.org/stretch/dpkg-dev/deb-version.5.en.html https://manpages.debian.org/stretch/dpkg-dev/deb-version.5.e... I was joking. Really I would sort file names lexicographically. But the way Debian sorts version numbers is interesting and seems like a good way of handling that particular situation.
- userbinator 1y agoThe most irritating circumstance for this is looking for files named with a hash: 3ea4f... ... 97dce... ... 126b9... This is one of the settings I immediately turn off on Windows via the registry key mentioned in the other comments here. I miss the time when computers did what you told them to, instead of trying to read your mind. These days, it's more like "trying to change your mind". I absolutely hate the "the user is wrong" authoritarian mentality that unfortunately has infected a ton of software, even open-source.
- krick 1y agoExactly. This is even more annoying when it isn't exactly a hash, but some gibberish you cannot really make sense of, which does have a numeric section in them: like a user ID, or unix time, or who knows what else it could be, but you are trying to visually find a file abcd89764237 somewhere after abcd683426834, and it isn't evident why you cannot, unit you notice that the latter has more digits in its "ID" for some reason.
- antonyh 1y agoIt looks like GTK & KDE both suffer from this - I get this behaviour in Thunar and in Dolphin. This is the kind of thing that makes me lose sleep. It's the same on MacOS too, at least in the latest version.
- TacticalCoder 1y ago[dead]
- amelius 1y agoAnother problem which annoys me to no end is that most file managers and file selection boxes put directories before files. This makes it hard to find the file that was most recently changed, for example. Which is an action that is extremely common. (In fact, why does my file manager not have a most-recently-used shortcut?)
- ivanjermakov 1y agoHave you tried sorting by created date instead? What is the point of relying on filename if that's not what you want.
- mzmzmzm 1y agoThere's a Group Policy setting in Windows: Computer Configuration\Administrative Templates\Windows Components\File Explorer\Turn off numerical sorting in File Explorer Group Policy has so many essential settings I hurry to change with every isntall. I wish Windows would expose more of them to the user in ordinary settings.
- pif 1y agoBro forgot that ls has an option to obtain the same sorting as the one he doesn't like. ls has it because it solves a real user need.
- kiitos 1y agols sorts filenames strictly lexicographically, comparing character by character, so e.g. "055436307" is compared as the characters "0", "5", "5", etc. so it sorts before "121134" because "0" is less than "1". if all compared characters match and one string ends, the shorter one comes first. Symbols like _ are just more characters, and their position relative to digits depends on the locale’s collation table. Google Drive uses ICU collation with the numeric option enabled, which treats each consecutive sequence of digits inside the filename as an integer. so "055436307" is parsed as the number 55,436,307, while "121134" is parsed as 121,134. and since 121,134 < 55,436,307 then "121134..." comes before "055436307..." even though lexicographic order would suggest the opposite. and i think when two digit runs have the same numeric value, the shorter run comes first; if runs are equal and the string continues, then normal character comparison resumes, including any underscores or suffixes
- IshKebab 1y agoIt feels like this algorithm could be improved though. If a number has leading zeros you probably don't want to sort it numerically. That said the author's situation where it's numerical and different lengths seems likely rare enough that it probably isn't worth complicating things.
- debazel 1y agoThe leading zero isn't an issue because it will sort correctly under both systems. The issue OP is having is that he's adding random numbers after the hhmmss section. If instead he added a delimiter before the random number the files would sort correctly under both systems as well, e.g. hhmmss_num.
- IshKebab 1y agoYes that's what I said: > where it's numerical and different lengths
- kiitos 1y agofor me the takeaway here isn't that sorting should use some smarter/better mechanism for inferring semantic intent from filenames, but rather that sorting should not try to infer semantic intent from filenames in the first place
- themafia 1y agoI'm often yelling at the software on my computer: "STOP TRYING TO HELP ME!" It's like having a toddler help you make a meal. It wants to be involved and recognized so badly. Meanwhile I'm starving and just want to get the food done as quickly as is possible and I'm constantly tripping over this little ball of misguided efforts. Please. Stop trying to be smarter than me. You often can't, and when you get it wrong, you make it measurably worse. If you insist on doing this please give me the "Expert Mode" setting back so I can flatly disable ALL OF IT with one click.
- ggm 1y agoIf only we could represent sort order by some structural form of decision logic which also embeds encoding a regular .. well.. expression matching a pattern..
- benoau 1y agoEarlier this year I submitted a bug to VSCode about sorting Playwright tests in the alphabetical-and-numerical order that VSCode favours, after Playwright told me it was a VSCode issue. Some people rushed to fix this as I'd done some diving into the issue and presented the relevant information and code, so now VSCode's Playwright test list uses the same sorting mechanism as the rest of VSCode. Sadly, the underlying Playwright does not receive that order from VSCode so it still actually runs sequentially-numbered tests in strict alphabetical order. :(
- cyberax 1y agoHeh. One of the bugs that once caused me to bang my head against the wall was caused by the Estonian language. Its alphabet has Z following S and Š. So the "foolproof" regexp to match the letters '[a-za-Z]' was misfiring for some entries.
- avodonosov 1y agoDumb tools are more robust.
- makeitdouble 1y agoMany commentors are positing the "clever" sort is what 99% of the user's want, but I really doubt it has been properly checked beyond the original PO's hunch and at most some user panel with pre-sampled data. Most of these decisions are early default behaviors that stay there as long as users aren't clamoring for change, and TBH I can't imagine most users to have a self emerging strong opinion on how alphabetical sort should be working.
- Anamon 1y agoThis case is really the opposite, though. Sorting strictly by character value was the default for decades because it was the quickest and easiest to implement. Only with computers hitting the really big mainstream, and millions of complaints from users who don't know the ASCII table by heart, or have ever heard of ASCII, did tools start to implement the kind of ordering logic that any user who isn't a developer would expect.
- SigmundA 1y agoOur users loved when we added "natural" sort, which was pain in the ass in the db, but ultimately no big deal. They absolutely do not care or understand the difference between alphabetical and numerical and natural, what they care about is 10 should not come before 9 in "Item 10" vs "Item 9". Whatever pedantic argument you have that natural is not alphabetic will lose you sales, your users do not care and want numbers to make sense in sorting.
- seriocomic 1y agoMore fascinating for me is this discussion thread, where there's legitimate debate around the need/expectation for alphabetical sorting to match/include lexical sorting. I'm personally in the "want lexical as part of alphabetical" - as 'photo19' should come after 'photo2' in my expectations, but the number of cases cited where this doesn't/shouldn't work is enough to justify a degree of contextual or situation awareness that most systems and interfaces simply aren't designed to cater for (file-systems vs photo-storage applications).
- kccqzy 1y agoIn case anyone else strongly disagrees with the author and wants to implement the Microsoft/Google/KDE/etc behavior, Google conveniently has this open sourced: https://github.com/google/closure-library/blob/b312823ec5f84239ff1db7526f4a75cba0420a33/closure/goog/string/string.js#L385 https://github.com/google/closure-library/blob/b312823ec5f84... How did I find it? Well I wanted to implement it a while ago and I found it in Closure a library I was already using.
- NoPicklez 1y agoIf they truly want alphabetical order wouldn't it be that the 9 is "nine" and 1 is "one" and therefore nine would be before one? Otherwise they mean lexicographical where they only look at the left most value and sort that. You can't ask for something to be alphabetical and expect it to sort numerically.
- fsckboy 1y agoit's weird to me that all the people declaring that they know what the average user wants to see, don't also suggest that the computer should rename files it encounters as necessary to give the user what the user wants. if we don't have to collate as dictated by ascii, why should we expect users to live within the bounds of file names with dotted extensions? you think users care whether something is a jpg or a png? do users want to see .MOV and .mov next to each other (not because sort, because one camera programmer did it that way for an ancient DOS filesystem, and another didn't.) (unix, btw, never required that users live with dotted extensions, that was a digital knockoff/cpm/microsoft thing that you didn't understand so all your new tools enforce it even though you never had to put your code in a .c file, that was just for your convenience as the user whose needs must be respected) so, we have to have "computery filenames" but we should violate "computery sorting"? how incredibly close-minded of you, you have no idea or basis to know what users want to see. oh, and the solemnity with which you make these proclamations, ok, don't get me started on that.
- whatgoodisaroad 1y agoi would argue that when you say "alphabetical order" you mean "lexicographic order"
- darkhorn 1y ago> I don’t know when this became the norm, to be honest I have not used a normal graphical file manager in a long time. If I remember correctly Windows 98 was sorting alphabetically. Then Windows XP strted to take numbers into consideration.
- layer8 1y agoFWIW, for Windows Explorer the numerical sort order can be disabled by setting the DWORD value HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows\Currentversion\Policies\Explorer\NoStrCmpLogical in the registry to 1.
- bhandziuk 1y agoI honestly thought Explorer was broken and have been looking into 3rd party file browsers for Windows because this has been driving me so nuts. Thank you!
- layer8 1y agoThere’s also the account-specific HKEY_CURRENT_USER\Software\Microsoft\Windows\CurrentVersion\Policies\Explorer\ NoStrCmpLogical, which I meant to mention, but mixed them up by mistake.
- dzuc 1y agoI don't know why I was surprised to learn this but there is a standard for alphabetical order. The NISO Guidelines for Alphabetical Arrangement of Letters and Sorting of Numerals and Other Symbols: https://www.niso.org/sites/default/files/2017-08/tr03.pdf https://www.niso.org/sites/default/files/2017-08/tr03.pdf
- Nevermark 1y agoConvenient-to-select settings should always include: Sort: In Alphabetical Order In Alphanumeric Order In Alphabetic-Word Order In Right-Aligned Alphabetic Order Randomly Sometimes Never By Hash Very Fast In the Background In the Foreground In the Underground In the Cloud Yes With Bubbles No Strong Opinion Of On YYYY-MM-DD HH-MM-SS: [SELECT] Repeat: [SELECT] With Random Site Free Download Sort Extension: [SELECT] Let Facebook Emergency Backup Sort [SELECT] Who Sort?
- fastball 1y agoNumbers aren't part of "the alphabet", so sorting digits within a string by the numeric value makes just as much sense (and is what most users want most of the time) as treating digits as isolated characters (what OP wants). As an aside, this is also the reason why ISO 8601 is the best date format – it sorts the same way whether you do it alphabetically or lexicographically.
- dmichulke 1y agoI have the same issue with "15 minutes before" instead of "2025-09-29 01:13:30". (Which is wrong once the site doesn't update) Needless to say, those are all "features" dumbing us down in the long run. A philosophical side question: I want to opt out of this but I can't. So is this is case where my peers are limiting my intellectual development? I.e. preventing me from a) doing the time calculations in my head, b) writing my software such that is uses leading zeros?
- ssivark 1y agoUnfortunately, it's not so simple, especially once you go beyond the ASCII. Dylan Beattie has this brilliant talk [1] where he points out how even the "systems" in human language involve a pile of quirks rather than any simple clean rules, and many of those rules are conflicting and the appropriate order of precedence depends on the context. Eg: the correct sorting order for the same sets of strings might even depend on the geography in which the question is asked! If you haven't had to deal with it previously, you'd be flabbergasted at how many foot-guns there are in such a simple question as alphabetical sorting, even without involving numeric components in strings. [1] There's no such thing as plain text https://www.youtube.com/watch?v=ajfb5LSbQVM https://www.youtube.com/watch?v=ajfb5LSbQVM
- d1sxeyes 1y agoI am surprised how many people are comfortable calling sorting numbers alphabetical sorting (including TFA). In true alphabetical sorting, sorting numbers is undefined behaviour. Both of these sorting methods are valid extensions of alphabetical sorting, and which you prefer is just that: a preference. So actually when he says ‘alphabetical order’, he does not, in fact, mean ‘alphabetical order’.
- franky47 1y agoCursed alphabetical sorting of numbers: 8 5 4 9 1 7 6 3 2 0 Can you guess what it is?
- chithanh 1y agoThey are sorted by their Unicode character names obviously U+0038 DIGIT EIGHT ... U+0030 DIGIT ZERO
- adornKey 1y agoHistorically I'd like to add that FileNames are just sequences of bytes that come with a few restrictions. They don't even have an encoding you can use to sort something. Windows FileNames look like UTF-16, but they can be truncated. You can't convert them to UTF-8 and back without loss. (For that you need WTF-8) Once you use random FileNames you'll start to notice...
- merelysounds 1y ago> I miss the time when computers did what you told them to, instead of trying to read your mind. This could be an illusion, or at least something difficult to evaluate; the operator is less likely to notice the situations when the computer successfully “reads their mind”. Also, I guess new users (i.e. those unfamiliar with previous behavior) won’t care as much about wrong assumptions; they will only learn that one doesn’t need a leading zero.
- thomasahle 1y ago> I miss the time when computers did what you told them to, instead of trying to read your mind. You haven't seen anything yet. Get ready for "Sort by AI" which will try to interpret the content of your images to sort them based on what you'll want to look at next. Incidentally, in this case AI would have sorted them the way you want: These look like photos straight from a phone, with filenames in the form: IMG_YYYYMMDD_HHMMSS... So the natural way to sort them is *chronologically*, by the timestamp embedded in the filename. If we do that, the order becomes: 1. `IMG_20250820_055436307.jpg` — Aug 20, 05:54:36 2. `IMG_20250820_092016029_HDR.jpg` — Aug 20, 09:20:16 3. `IMG_20250820_092440966_HDR.jpg` — Aug 20, 09:24:40 4. `IMG_20250820_092832138_HDR.jpg` — Aug 20, 09:28:32 5. `IMG_20250820_095716_607.jpg` — Aug 20, 09:57:16 6. `IMG_20250820_103857_991.jpg` — Aug 20, 10:38:57 7. `IMG_20250820_103903_811.jpg` — Aug 20, 10:39:03 That order reflects the actual sequence the photos were taken.
- ____tom____ 1y agoThere are quite a few more rules for sorting that can be applied - it's not just numbers, and numbers don't always work the way you describe. There is "Dictionary Order", "Phone book order", and a few other standards. (Dictionary order is not lexicographic order, even if the two are now commonly conflated). A simple rule that most still know is a book titled "The Book", should be sorted under "Book, The". They have variations on how special characters sort, how abbreviations are handled, and even have differences in numbers. For example, in phone book order, "21st Century" sorts under “Twenty-first”, not "21". And, of course, non-English languages add all sorts of other rules. This tends to get ignored these days, as lexical sorts are so much easier to implement, that people forget there are other, preferred options.
- navigate8310 1y agoThis is also the case with Excel. If numbers are stored in general formatted columns and you sort by A to Z, you'll get 1, 10, 11, 12, ..., 2, 20, 21 and so on
- Sammi 1y ago> I have also found a setting to fix Dolphin’s behavior, but it was very much buried into its many configuration options. KDE wins again. It's my favorite desktop environment, because it has defaults that are friendly to noobs, but it also get out of your way and lets you change things if you want. The trend is for other desktop environments to be either/or. Either they are super simple and noob friendly, or they are super technical and have a steep learning curve and you get to configure everything - but only via text config. Maybe Cosmic looks like it's going the same route as KDE, where it's trying to bridge the gap.
- antonyh 1y agoThanks, now I cannot unsee this. Thunar has this broken sort order too, and I've no idea how to make it sort file names with hash values 'properly' - by which I mean the same as `ls` which broadly speaking on my system is 0 to 9 then a-z case insensitive. Instead I have an order of starting character that goes 1,4,5,7,9,2,3,7,8,9,4,6,1,2,.. etc etc which is utterly useless as a sort. I've always thought the sort was weird but couldn't quite figure out why (I usually sort by date descending). Another non-productive thing to figure out and fix.
- bashkiddie 1y ago> not every single piece of software fucks up something as basic as string sorting it is neither basic nor simple. Have you ever heard of UTF-8 and locales? Here is an exercise for the curious reader: Pick any UTF-8 string "a", and another one "b", so that in increasing lexicographical order "a" sorts after "a+b" ("a" concatenated by "b"). ("a" > "a+b")
- edude03 1y agoI feel like it's not intelligence or lack there of, it's that implementing sort with a[i] < b[i] is the simplest way to do it. Putting 9 before 10 would require some kind of windowing since otherwise you'd be comparing 9 and 1, and of course 1 is smaller.
- bhandziuk 1y agoThis must be why, when I have a folder in Win11 full of files with GUIDs as names, they are never in the order I expect. Windows seems to sort them randomly but there must be some sub-sequence of numbers that it's deciding are the important ones and sorting off those. For me I'd much rather just sort left to right alphabetical.
- LLyaudet 1y agoNothing new about lexicographic order versus natural sort order. However, for people like me who love "creeping featurism", the UI and UX could be improved. First, both lexicographic order and natural sort order are not absolute, they are relative to an underlying character order, itself relying on character grouping algorithm (it is not the same thing to say a byte is a character and to group bytes to have an utf8 character or an utf32LE character). So far, the UI is just an "arrow" upward or downward next to Name column or things like that: easy as pie, one or two clicks next to a column name and you have the chosen order. But if you copy what is available in ERP with web apps like an ERP done with Django or another framework, you can: - sort on many columns, each column sort being one element of a bigger lexicographic order on the chosen columns, - propose many distinct sorts for the same column: most UIs stop at "increasing/decreasing" order, but with a modal, or one or more selects, you can propose to chose your wanted sort with more than two possible values. For example, if dealing with strings, most UIs only propose increasing/decreasing natural sort order grouping with UTF8 characters (GNU/Linux) or UTF16 characters (Windows) and default system collation. But instead you could treat any string as a sequence of bytes and select your character grouping algorithm (including a cascade : try to group as UTF8, if fails, try as UTF32LE, etc., if all fails use each byte as a character (not an ASCII one, UTF8 would have worked, an ISO-8859 for example, see collation after) ... you can start laughing ;) XD I'm still serious but I also know it is quite funny :) ), then select your collation, then select your sort between lexicographic and natural. This is usually not needed, but would give full control to the user. And I love "creeping featurism" and giving full control to the user :). Typically, if one day an AI can code a variant of Gnome and Nautilus in a matter of hours, then with this kind of knowledge, if you know how to ask, you could have such a complex and juicy UI that is functional in a matter of hours :). Maybe one day we will all have custom-made OSs instead of ready-made OSs, and people with knowledge and a taste for complexity will have options mere mortals never thought of :). To return on the current ground, distinct sorts is common in ERPs with "type" or "status" columns: imagine you have an ERP with a webpage for today deliveries, and in some office someone needs the list of today deliveries with those that are already delivered at the top, in another office someone else needs those that are still currently in delivery at the top, and again someone else needs those that have an anomaly at the top, 3 distinct orders at least. I made up this example, and it is simple enough to think that a filter could replace an order, since I only talked about the top. But I can tell you such things do exist in real world applications.
- angg 1y ago> I miss the time when computers did what you told them to, instead of trying to read your mind. amen