15 ms·
Null References: The Billion Dollar Mistake
- seventh-chord 7y ago"Making everything a reference: The Billion Dollar Mistake" is the talk I want to see
- x3ro 7y agoCan you elaborate? I can't remember the last time I thought "oh darn it why is this a reference", but I can think of a billion problems I've had with nulls in jvm languages
- emsy 7y agoCache incoherency, which will cost us more and more performance as CPUs will improve slower in the future.
- gameswithgo 7y agothere are a few completely different ways to interpret this, can you explain?
- seventh-chord 7y agoin languages like c, rust or go, where you can put arbitrary data on the stack, it seems to me as if such issues are less common because you dont have to worry about initializing pointers and allocating memory unless you actually want to put something on the heap. Thus if you make everything a reference in your language its no wonder you run into issues like null-pointers more often
- DaiPlusPlus 7y agoWith stack allocation you then encounter problems with object lifetime. Rust solves this problem by binding references to scope, and Go solves this by invisibility changing an allocation to the heap (and uses ref-counting? I think?). I wish C had a feature that would let you allocate something on the stack and then return to the parent stack frame without popping the stack-pointer - that would be handy for self-contained object-constructors.
- cm2187 7y agoPlus also doesn’t the stack need to be small to fit into the CPU cache?
- seventh-chord 7y agoThere is absolutely no requirement from the hardware that the stack be any particular size
- DaiPlusPlus 7y agoOn the Windows desktop, the default stack size is 1MB. In IIS-hosted applications the default stack size is reduced to 250KB due to the popularity of the now-outdated programming trope of "one thread per request (per-connection)". On x86 Linux the default stack size is 2MB - which seems generous.
- cm2187 7y agoI am not thinking requirement by rather performance wise.
- kragen 7y agoIf you start allocating multi-kilobyte objects in your stack frames, they are not going to fit into L1.
- zozbot234 7y agoGo uses fully-general GC, not reference counting. Obligate reference counting is used in other languages such as Swift, probably with worse throughput than obligate tracing GC.
- giulianob 7y agoRegardless of whether it's on the stack or heap the point still stands. If all your objects are randomly allocated then an array is just references to those objects and will start out null. If you're using value types then your array of objects will never be null (empty instead) and you will benefit from CPU caching the data.
- dorfsmay 7y agoEverything in Python is a reference, and there's no null pointer issues.
- auxym 7y agoI've certainly had some "None" errors in Python. I think the difference comes from dynamic vs static typing. In Python, you sort of get into the habit of "defensive" programming: checking inputs to your function, catching Nones, etc. In java, you tend to rely more on the type system. If it typechecks/compiles, there's a good chance it's OK. That is, until you get a null value that's not handled. That's the root issue I think: If null is an acceptable value per the type, then the same type system should force you to handle it. As do the type systems in ML languages for option types, for example.
- JakobProgsch 7y agoThe first line is why I'm not 100% convinced of the severity of this mistake compared to the alternatives. The problem fundamentally is the use of magic values/numbers to represent the concept of "no value". You don't need explicit language support to have that concept and the bugs it causes. I guess having that as an intrinsic concept in the language makes it more likely that people use it badly. On the other hand debuggers etc. also intrinsically understand this and segfaults due to null pointers are usually very easy to localize once you see them. On the other hand if a "bad programmer" introduced their own magic non-value in a supposedly safe language, debugging that becomes way more confusing.
- andrepd 7y agoNo, that's not the "fundamental problem". The fundamental problem is a type system that lies. A "pointer to string" is not actually a pointer to a string, it's a pointer to a string or to nothing. If your api returns a pointer of the latter type, it should signal this by making the return type "maybe-pointer to string" (although it has the same memory representation as "pointer to string"). Then, if the user tries to dereference a maybe-pointer (that is, to use a maybe-pointer as a pointer), the type system can statically catch this and make it a simple type check failure compilation error. The user must first check if it's null through a function that casts a maybe-pointer to a pointer. Nothing about this precludes the usage of sentinel values.
- kragen 7y agoYou may be interested in http://canonical.org/~kragen/memory-models http://canonical.org/~kragen/memory-models then. I don't think it's necessarily a mistake but it's definitely taken for granted far too much.
- x3ro 7y agoThis comes up again and again in one form or the other, yet new languages still seem to be making the same mistake. Of all languages I've touched, Rust seems to be the only one that mostly circumvents this problem. Are there other good examples?
- gameswithgo 7y agorust, f#, ocaml, latest version of c# has an option to sort of get rid of nulls, zig
- hawkice 7y agoHaskell, notoriously. I believe it pioneered the ergonomics of the alternatives used elsewhere.
- cmrdporcupine 7y agoAFAIK Standard ML predates Haskell and it has an option type.
- dunefox 7y agoML is even older than SML and has algebraic data types.
- jrockway 7y agoI assume two reasons, efficiency and because an efficient implementation of mutable state would have the same problem. Right now, a single sentinel value makes a pointer null or not null (0x0 is null, everything else is not null). This is exactly how you'd implement a stricter type, like "Maybe". Encoded as a 64-bit integer, "Nothing" would be represented as 0x00000000 and "Just foo" would be represented as 0xfoo. No object may be stored at the sentinel value, 0x00000000. Exactly the same as what we have now, and provides no assurances that 0xfoo is actually a valid object. Meanwhile, Haskell which "doesn't have null" crashes for exactly the same reason your non-Haskell program crashes with a null pointer exception: f :: Num a => Maybe a -> Maybe a f (Just x) = Just (x + 41) This blows up at runtime when you call f Nothing, because f Nothing is defined as "bottom", which crashes the program when evaluated. It's exactly the same as langages with null pointers: func f(x *int) *int { result := *x + 41 return &result } And the solution is the same, your linter or whatever has to tell you "hey maybe you should implement the Nothing case" or "hey maybe you should check the null pointer". Where I'm going with this is that you need to develop entirely new datatypes and have an even stricter type system than Haskell. Maybe Rust is doing this, but it's hard. We all know null is a problem, but calling null something else doesn't make the problems go away.
- rzwitserloot 7y agoThis old chestnut again. There is an inherent problem in designing processes and writing code to capture them: The notion of not-a-value. There are great many ways to solve them. The most common ones are 'null' and 'Optional[T]'. Neither just makes the problem magically go away. If a process is designed (or a programmer writes it) thinking that 'ah, well, here, not-a-value cannot happen', but it can, then.. you have a bug. Some language features might make it possible to help reduce how often it occurs, but eliminate it? I don't think so. Imagine, for example, in an Optional based language, that you just map the optional to a lambda to execute on the optional, and the behaviour of the optional is to then simply silently do nothing if it's optional.none. That'd be a much harder to find bug than a nullpointer error. (errors with stack traces pointing at the problem are obviously vastly superior to mysterious do-nothing behaviour with no logs or traces of any sort!). Some other creative solutions: * [Pony](https://www.ponylang.io/ https://www.ponylang.io/) tries to be very careful about registering when an object is 'valid' and when it isn't, and when you write code, you have to say which state the objects you interact with can be in. This lets you avoid a lot of the issues... but pony is quite experimental. * In java you can annotate any usage of a type with nullity info, and then compiler linter tools will simply tell you that you have failed to take into account a potential null value. You are then free to ignore these warnings if you're just writing test code, or know better. Avoids clogging up the works with optional, but as the java ecosystem shows, you can't just snap your fingers and make 30 years of massive community effort instantaneously instantly be festooned with 'might-not-hold-a-value' style information. At least the annotation style gives the hope of being backwards compatible (to be clear, optional, for java? Really bad idea). * in ObjC, if you send a message to a null pointer, it silently does nothing, in contrast to virtually all other languages with null types where attempting to message a null ref causes an error or even a core dump. * Just write better APIs. Have objects that represent blank state (empty strings, empty collections, perhaps dummy streams which provide no bytes / elements, etc). For example, in java: Java's map (a dictionary implementation) has the `.get(key)` method which returns the value associated with that key, and returns `null` if there is no such value. About 6 years ago another method was added in a backwards compatible fashion (so, all java map implementations got this automatically): `getOrDefault(key, defaultValue)`. This one returns the provided default value if key isn't in the map. You'd think optionals provide a general mechanism for this, but, in scala, you have both: There's `someMap get(key)` which returns an optional, so to get the 'give me a default value' behaviour, that'd be `someMap.get(key).getOrElse(defaultValue)`, but maps in scala also have the java shortcut: `someMap.getOrElse(key, defaultValue)`. Sufficient thought in your APIs mostly obviates the issues. null is not a milion dollar mistake. It is a solution to an intrinsic problem with advantages and disadvantages over other solutions.
- agumonkey 7y agoShould every domain have a Nil element instead ?
- fhars 7y agoNo, obviously not. Every domain having a Nil element is exactly the problem null references have introduced (at least for the call by reference parts of the affected languages).
- agumonkey 7y agoNull is a single nil for all, I meant having a null per domain would force people to think of what it means to have nothing in that field and handle it. Maybe I'm too naive.
- augusto2112 7y agoSometimes you don't want to allow a value to be null at all, but with null references you can't represent that at the language level.
- agumonkey 7y agoBut for numbers, a zero is not considered null, because it was handled in the operators rules.
- RickJWagner 7y agoNo comment on the Null References, but I will say I love the time-index provided for the video. I wish every video had these!
- olliej 7y agoNull termination is still easily much worse. At least the general case of null dereferences today (less so earlier) is a page fault.
- Matthias247 7y agoOut of all possible gotchas in programming languages I still find null pointers the easiest one to discover and fix. You directly see when and where it happens, and the fix is usally straightforward. Compared to that invalid pointers (stale references) are a lot more painful, since programs might continue to work for a while. Managed languages do at least prevent those. Multithreading issues are imho the biggest pain points, since they are introduced so easily and often go unnoticed for a long time. The amount of languages that prevent those is unfortunately not that big (Rust plus pure singlethreaded languages like JS plus pure functional languages).
- smt88 7y ago> You directly see when and where it happens, and the fix is usally straightforward. This is not true in most dynamic languages, especially ones where I/O is not typed. You have to be extremely dilligent about verifying input. JavaScript comes to mind.
- Supermancho 7y ago> You have to be extremely dilligent about verifying input. That's true of all languages. Null references are a problem of low effort development. Calling it a billion dollar mistake is sensationalist hand-wringing. It accidentally highlighted how carelessly most programs are written, implying that without it developers wouldn't be checking inputs as strictly, because they wouldn't need to. Yes it's another type, but lots of languages have a nil/null and there hasn't been a demonstrative reason to pull it.
- smt88 7y agoWell to be clear, most modern languages with reasonable type systems will force you to explicitly verify the type of your input. C#, for example, forces you to cast your JSON before using it. If the cast fails because you got the class definition wrong, you get an error (like a constructor error).
- zeendo 7y ago> there hasn't been a demonstrative reason to pull it" is not the reason it's not "been pulled Languages are hard to change and backwards compatibility is paramount. Hell, some languages support null just for interoperability (i.e. Scala) when they would have otherwise not allowed it when they were created. Null isn't expressive and is historical baggage. At this point "billion dollar" is probably an understatement. I wonder how many people that have spent significant time writing in languages that allow null and those that don't prefer having null? I, for one, wouldn't willingly go back to a language that allows null.
- decafbad 7y agoC.A.R Hoare couldn't foresee consequences 55 years ago. That's a small mistake. We should blame language designers who didn't bother to handle the problem after it's been obvious.
- littlecranky67 7y agoLot of mainstream languages nowadays support non-nullable types, i.e. TypeScript and C# (taken from F#).
- deleted 7y ago[deleted]
- microcolonel 7y agoNull references are not a mistake, they make perfect sense. Letting nullable types be dereferenced directly is the mistake. Null references are at the core of a great number of sensible datastructures, and they're a natural fit for conventional computers.
- int_19h 7y agoThere are two separate concepts here that often gets conflated. There's null reference in a sense of a special pointer value (usually all bits set to 0) that means "this doesn't point to anything". That's a useful low-level tool that allows for compact representation of many important data structure. And then there's null reference in a sense of type systems. To be more specific, "null reference" here is really a shortening of "every reference in the type system is implicitly nullable". And that is the billion dollar mistake. An explicitly nullable reference type that requires explicit check on dereference, or option types, that use null pointers under the hood, are obviously not the problem.
- dooglius 7y ago> "null reference" here is really a shortening of "every reference in the type system is implicitly nullable" I don't think this is quite accurate, there are definitely cases where non-null pointers are required (e.g. dereferencing). It's more correct to say that the type system does not explicitly indicate whether a pointer might be null or not.
- microcolonel 7y agoJust consider all pointers null until proven otherwise, shouldn't be that hard to do something like this in static analysis. Even if a reference is non-null, you still have to wonder if it's valid.
- throwaway2048 7y agoTo put it a bit more compactly, why is Boolean logic "True, False, Null"
- teh_klev 7y agoInfoQ has some gems, but their video content presentation is terrible (tiny box, or full screen): https://www.youtube.com/watch?v=YYkOWzrO3xg https://www.youtube.com/watch?v=YYkOWzrO3xg
- LorenPechtel 7y agoCount me amongst those who do not think they're a mistake. You need to indicate no-data-here in some fashion. If you try to use that no-data in some fashion having your program blow up from a null reference is a feature to me--in the vast majority of cases it's better go boom than silently continue doing something wrong. In the few where that's not the case you can trap the exception and go on. The real solution is what has been done with C# in recent years--have the compiler track whether a field can contain a null or not and squawk if you try to dereference something you haven't checked. That causes it to blow up in the best place--compile time rather than runtime.
- kroltan 7y agoYes, which is where types like `Optional` come in. If you make a language where null doesn't exist by default, but still provide a standard way of indicating non-presence, you get the advantage of compile-time correctness checking. Also, the compiler can still optimize the `(hasValue, value)` tuple into a possibly-0 pointer when the type of the value is a pointer. (which by the way, is exactly what Rust does, among others)
- iCarrot 7y agoThey are called Nullable types in C# and must be declared with `?` after the type. But, Nullable<T>.HasValue check is not forced and Nullable<T>.Value will throw a different exception instead if it is null (InvalidOperationException).
- to11mtm 7y agoWell, Depends on which type you are referring to (Which is part of what pains me with nullable ref types, as nice as it is to have) If it's a value type (T), ? will make it Nullable<T> and provide the behavior described. Reference types however can always be null, and do not have a .HasValue as exampled above. However newer versions of C# let you declare nullable references on a compiler level, but rather than HasValue/Value you still have to do the null check and instead can bypass via the new deref operator (!)
- mlangenberg 7y agoI was expecting someone to mention the Crystal programming language. In Crystal, types are non-nilable and null references will be caught at compile time. https://crystal-lang.org/2013/07/13/null-pointer-exception.html https://crystal-lang.org/2013/07/13/null-pointer-exception.h... I certainly recognize that many bugs in Ruby programs announce themselves as `NoMethodError: undefined method '' for nil:NilClass`. So to be able to catch that before releasing code is a very welcoming addition in my opinion.