8 ms·
Dijkstra: Why numbering should start at zero
- lucumo 17y agoFor my MSc thesis I'm programming some stuff for the Cell B.E. processor. It provides AltiVec vector instructions which allow quick parallel computation. It's great for getting nice speed ups, but it requires a lot of thought and bit twiddling from time to time. It has two C "intrinsics" called vec_mule() and vec_mulo(). These take in two vectors of integer datatypes (char, short) for multiplication and produces one vector with items that are twice as big. char becomes short, short becomes long. To make room for these larger datatypes it only multiplies the even (vec_mule) or the odd (vec_mulo) items. It's a nice construction, until you start to wonder which elements are even. Is "element 0" even, or is the "first element" odd? Should I really have to wonder about that? I have never questioned the utter coolness of starting at index 0, until I encountered this. I do not know why we start at 0 if the rest of the world starts at 1, but it seems silly and without reason.
- jacquesm 17y agoWell, the rest of the world also starts at 0: 1 sheep -> the number of sheep present when there is one 0 sheep -> the number of sheep present when there are none Because you normally only enumerate things that are there it seems like you are starting at '1', but really you start relative to the situation where there are 'none'. You can't have 'negative' presence of anything so it took a while longer to realize that there might be negative numbers too. As for element 0 being odd or even, it's even. So, it isn't silly and it has a reason. 0 is one of the biggest inventions in the history of mankind, programmers have given it it's natural place because we could not make our computers work if we didn't do that. Imagine binary logic based on '1' and '2'....
- lucumo 17y ago> Well, the rest of the world also starts at 0 Yes, they do, when they have no elements. But if an array has an element 0, it actually has 1 element, not 0. Whereas if I have a sheep 1, I have 1 sheep. There is no sheep 0. I think it would be good to note that I have nothing against 0 as a number. It's useful for many purposes including to indicate that a set actually has 0 elements, but I'm not so sure it's such a good idea to actually call the first element element 0.
- jacquesm 17y agoYou are confusing 'rank' with 'count'.
- lucumo 17y agoNot that I know of. Normal people and programmers do not count the same way. | last | | rank | count ------+------+------- sheep | 2 | 2 sheep | 1 | 1 sheep | - | 0 ------+------+------- array | 1 | 2 array | 0 | 1 array | - | 0
- jacquesm 17y agoProgrammers are pretty normal people :) When counting sheep in a field I'll count 1 sheep, 2 sheep etc. When counting variables I'll say one variable, two variables and so on. No difference there. When placing items in an array I say this value goes in to slot '0', this will go into slot '1'. The slots are labelled with their index. And I will say there are now two values stored in the array, one at 'position' 0 and one at 'position' 1. I think the whole problem disappears when you think of 'rank' as labels and count as the actual number of elements. The first sheep is just a label you stick on a sheep to identify it, that makes it different from the other sheep.
- thunk 17y agoEr, that's all worng. The question isn't number of, but index of. The index of the first element in a collection is 0. The index of the first sheep in a flock is (conventionally) 1. The total elements up to index n (inclusive) is n+1 -- the opposite parity -- which can be initially confusing.
- jacquesm 17y agoAn index is a label, a convenient way of finding something when you remember where you left it. You can start your indexing at any offset, you don't even need to use numbers (you could for instance look at an associative array where you store stuff based on a hash of some value). index = position, label The count is the number of occupied slots and is totally independent of the positions or labels. Otherwise how would you count the number of elements in a sparse matrix ?
- wcarss 17y agoIn this particular case though, where you have operations depending on evenness/oddness of some attribute related to position in a set (either position by index, or count) and the index starts with 0, there is absolutely confusion to be had over which implementation would be used. Your logic is sound, but it's counter-intuitive to think of Obj[2] as a member of 'mulo', an odd number, "because 2 is really the 3rd element". A person implementing the system could be logical by your ideal (making 2 odd) or logical by the ideal of common sense and readability (making 2 even). Either could validly hold in a person's mind. Given the ambiguity, I think it just shows that you shouldn't base operations on oddness/evenness, or should rigidly and obviously define them if you must use them.
- Radix 17y agoI'm happy to see this. I should have submitted it when I found it. The version I found appears to be handwritten by Dijkstra and I like it better that way. http://www.cs.utexas.edu/~EWD/ewd08xx/EWD831.PDF http://www.cs.utexas.edu/~EWD/ewd08xx/EWD831.PDF
- thunk 17y agoAbsolutely. His penmanship is as thoughtful as his bearing. I just hate pdf links.
- JMostert 17y agoDid you read Dijkstra's article? He's making the case for 0, though he's doing it a little abstractly. Simply put, if you start numbering at 1, you are setting yourself up for more boundary problems and off-by-one errors than if you start at 0 (and by extension, inclusive lower bounds and exclusive upper bounds). That's not to say that some algorithms are not in fact easier expressed by numbering things from 1, just that they're not the majority. Oh, and obviously 0 is even. Why? Because 0 mod 2 = 0, 0 is evenly divided by 2, and that's what "even" means. If you need more intuition, though: 1 is indisputably odd, and even and odd numbers alternate, so 0 is even. The rest is philosophy -- you can probably find definitions for "odd" and "even" where 0 is a problem. That's fine, but those don't help. Ignore them. Mathematically there's no problem whatsoever. The trickiness only comes if you insist in thinking in terms of "the first element", rather than "element number 0". Some people use "zeroth", but this seems to invite more confusion because it induces two meanings for all the other ordinal forms -- if there's a zeroth element, does that mean the "first" element is in fact the second element? Best to avoid ordinals altogether -- you usually don't need any more than "first" and "last" anyway.
- dkersten 17y ago"obviously 0 is even" Sounds like a bit of an odd statement to me.
- lucumo 17y ago> Did you read Dijkstra's article? Yes, I did. Does it hurt to offer a different perspective? I feel it doesn't. Approaching a subject with an open mind or from a different angle tends to be good for discussion. Offering a contradicting opinion or observation helps everyone to understand an issue better. > Oh, and obviously 0 is even. I might counter this with a similar question you asked me: "did you read my post?" I didn't make a case that 0 was odd, I asked how I should look at it: as element 0, or as the first element. If you look at it as element 0, the first element of the vector is an even element. If, on the other hand, you look at it as the first element, you will think it is an odd element. This is the case I made, not that 0 is even.
- JMostert 17y ago
- TweedHeads 17y agoI don't care if computers think in 0s and 1s or if arrays should start at 0 because it is the way computers think. Computers were made to serve us and as such they should translate all their inner thoughts to more human consumable data. Five is 5, not 101. So no, numbering should start at ONE, even if most programmers have already been hardwired to start counting from zero.
- thunk 17y agoBut he was making the point that there are valid human reasons for zero-based indexing - e.g. keeping the lower bound within the natural numbers; having the upper bound be the number of elements in the preceding sequence, etc. The "way computers think" doesn't enter the discussion.
- TweedHeads 17y ago"keeping the lower bound within the natural numbers; having the upper bound be the number of elements in the preceding sequence, etc." Read your post and understand the same can be said of using one-based indexing.
- thunk 17y agoThat's false. You'd be forced to sacrifice the exclusive upper bound, which would then force you to sacrifice the difference between the upper and lower bounds being the number of elts in the collection. Unless the lower bound was exclusive, in which case it could be an unnatural number. No, the only way there's no lump in the carpet is his way.
- scott_s 17y agoWe have this convention for the benefit of humans. Indexing starting at 0 has the convenient property that when we modulus an arbitrary number by the size of a range, the result is a valid index into that range. This comes up when implementing things that map a larger set onto a smaller set, such as buffers or hash tables. We could still achieve the same effect if we started indexing at 1, but it would require more code, and it would be less clear.
- edw519 17y agoI write in primarily in 3 languages, 2 start indexes at 0 and the third starts at 1. Not withstanding Dijkstra's logical arguments, I have adopted starting at 1 for 2 reasons: 1. I don't have to think about "shifting" data in order to look at it in the third language. 2. I don't have to "think" at all. It's intuitive that the first element is "1". I just insert a null in the first position of any array in a 0-starting language.
- cjbos 17y agoI'm a actionscript developer at my day job and I would totally take up this approach if I didn't have to deal with screen/stage coordinates all the time. You see, the backend of our project is in Coldfusion which has 1 based indexes for Array's and it has led to head-butting screen confusion when requesting the current index from the server. Then the visual side is always 1 based when showing data to the end-user, I'm always thinking I'm the problem in the middle. There have been numerous occasions where QA have called me to ask why the first slide in a slideshow is called "Slide 0"!
- loup-vaillant 17y agoIf you ever have to pack multidimentional data in a flat array, I wish you luck, then. In this case, starting at 1 messes up lookup and insertion. You have to insert "-1" or "+1" pretty often. Too hard to get right in my opinion, I prefer the easier way: to start at 0.
- pbhjpbhj 17y agoShouldn't the "index start value" just be a compilation directive?
- embeddedradical 17y agoi couldn't bring myself to stick null pointers here and there for the convenience of syntax :P i see a lot about code being written for humans, but i'm not onboard with it. elegance in how it executes is #1, and elegance of maintenance is #2.
- compay 17y agoPeople sometimes complain about Lua's tables, which use 1 as their first index when treated like arrays. I suppose it could make some algorithms more complicated but in my day to day programming it's never made any difference to me so far.
- JBiserkov 17y agoI wish I was shown this in my "introduction to programming course". All they told us was "because that's how things are. [MEMORIZE IT!]" Perhaps I should read more of Dijkstra's writings. Edit: I love when directory indexing is NOT forbidden. http://www.cs.utexas.edu/~EWD/ewd08xx/ http://www.cs.utexas.edu/~EWD/ewd08xx/
- Locke1689 17y agoIt's fairly straightforward if you're a C programmer and think of arrays as contiguous blocks of memory. int * a = (int * )malloc(5); a is a pointer to the beginning of an integer (default 4 byte) array. a[i] == a + i == ((uint8_t * )a)+i*sizeof(int). Therefore, a[0] == a + 0 == the first element in a contiguous block of memory.
- discojesus 17y agowhich would also illustrate what K&R said about array indices just being syntactic sugar for pointer arithmetic. My mind has been blown.
- gjm11 17y agoIf that blows your mind, you might want to avoid looking at the following: "foo"[2] == ("foo"+2) == (2+"foo") == 2["foo"]. And yes, you really can do that: in general a[b] and b[a] are exactly the same thing in C.
- scott_s 17y agoMore C array pointer fun: *("abcdefg" + 3) == 'd'
- bhousel 17y agosure as long as you don't care about stuff like unicode..
- hc 17y agonot quite. "foo"[2] is a char, ("foo"+2) is a string.
- gjm11 17y agooh, bother, I forgot what HN does to asterisks. In what I wrote above, look for the bit in italics and mentally insert asterisks on each side of it. It was right when I typed it in, I promise...
- BillGoates 17y agoI am starting to think Dijkstra is about the worst thing ever happening to software development. He arguments that option A is the least ugly one, but forgets that computers have maximum numbers. According to Dijkstra, you should test if a variable is a valid integer like: 0 <= var < MaxInt + 1. The best possible outcome of such a test would be a compiler error. His argument that sequences defined as min <= i < (max + 1) because (max + 1) - min = total number of items is just silly. Maybe true for maths, not so for programmers. When reading code, you want to know if the code is valid, not how many items there are in a list. And reading i <= Max instead of i < (max + 1) is simpler. Secondly, his article is about 0 or 1 based arrays, not random selections. And in case of 1 based arrays, the max = total number of elements. So option C is best. About whether 0 or 1 based arrays are better. They both have their uses. 0 based for spatial coordinates, 1 based for lists and character positions.
- thunk 17y ago1) The memo has nothing to do with testing for MAXINT. Just test it inclusively. 2) You frequently want to know the length of a subseq, and you frequently want successive subseq ranges to dovetail. 3) What in the world does rand() have to do with this?
- BillGoates 17y ago1) He is using real world math arguments why computer arrays should be 0 or 1 based. His example seems valid, but once you use Maxint it fails, proving his argument wrong. 2) Yes, but that has nothing to do with anything I wrote. Also calculating the length of a subset out of a 0 or 1 based array is the same, making your point completely irrelevant. 3) Where in the world do you read anything about rand()?
- thunk 17y ago1) MAXINT is already a special case, so testing inclusively doesn't introduce extra accidental complexity. 2) If you test the upper bound inclusively, then the upper bound of the previous subseq is not the lower bound of the next subseq. 3) I misinterpreted GP. Scratch that :)
- nishta 17y ago"Consider now the subsequences starting at the smallest natural number: inclusion of the upper bound would then force the latter to be unnatural by the time the sequence has shrunk to the empty one." Maybe it is too late, but I don't understand this sentence. A sequence starting at the smallest natural number would be: 0, 1, 2. The variant Dijkstra is arguing about is: 0 <= i <= 2. How can the latter be unnatural? And what does he mean with "shrunk"?
- jongraehl 17y ago[0,1) = (0) [0,0) = () [0,0] = (0) [0,x] = () Probably you'd say x=-1 but if you're using unsigned indices, then you can't distinguish the biggest possible sequence from the empty one (this is true of all schemes, actually, but at least the other don't require negative numbers).