7 ms·
Low-Lock Singletons In D
- brazzy 13y agoDoesn't D have the ability to do stuff on class initialization? Or does it not initialize classes lazily? In Java, this is how a threadsafe Singleton looks like: private static MySingleton instance = new MySingleton(); It will be initialized when the class is first used. The whole double-checked locking and TLS thing is IMO a huge pointless smartassery contest in VM lawyering hinging on the eminently false premise that having singletons in a heavily multithreaded environment initialized sooner than they're used is a problem that anyone needs solved. Whereas in reality, 90% of all applications don't need lazy loading, and for 90% of those that do, the above is lazy enough.
- spand 13y agoI think you missed a static keyword in that line of code.
- brazzy 13y agoYou're right; fixed.
- andralex 13y agoThat requires an acquire on every access, so it's generally slower.
- brazzy 13y agoNo, it doesn't. The JVM guarantees class initialization to be threadsafe (the lock is implicit), so you don't need mutual exclusion to access the reference - it's initialized before any thread could attempt to access it.
- WalterBright 13y agoThis is an elegant solution to the double checked locking bug http://en.wikipedia.org/wiki/Double-checked_locking http://en.wikipedia.org/wiki/Double-checked_locking, incurring insignificant overhead.
- rbehrends 13y agoUnfortunately, the space overhead is not quite that insignificant. It requires at least one bit of storage per thread; this does not matter much for a singleton class, but if you want to implement more general write-once data structures (such as arrays or hash tables), it's not negligible. Somewhat less problematic is that fast TLS is not natively available on all architectures (in particular, Mac OS X). Luckily, this can be worked around relatively easily.
- WalterBright 13y agoD does its own TLS implementation on OS X. But now that newer versions of OS X have native TLS, we'll be scrapping ours and using the native one.
- jamesaguilar 13y agoOr you could use atomic swaps on the initialized bit. That's both cheaper and simpler. psuedocode: member atomic32 init member mutex lock member obj ptr membarrier() if (atomic_load(&init)) { return ptr } else { locked_init() atomic_save(&init, 1) membarrier() } Not an expert, use this at your own risk.
- vy8vWJlco 13y agoIf you have atomic swaps (hardware "lock-free"), the global shared instance pointer should be enough - since it either holds a valid pointer or it doesn't.
- limmeau 13y agoWhat do you swap into the pointer? A new instance?
- vy8vWJlco 13y agoIf the pointer is 0 through a normal test, you would allocate the resource and do a compare-and-swap to record the allocation if the pointer is still 0. As noted beside this comment, you might wind up allocating several times in parallel (which may be undesirable too), but you would never clobber an existing allocation or change the instance that is dispensed.
- jamesaguilar 13y agoNo, you don't want two threads trying to allocate the object at once. You still need a locked init.
- vy8vWJlco 13y agoYes, there would be an allocation race, so if you cannot accept allocating unnecessarily in a discardable/recoverable way (but also one that preserves the semantics of the singleton: one and only one instance returned over its life), you would still need a flag. Though you might instead use an "incomplete"/place-holder reference value to simulate the flag and prevent duplication of effort/wasting resources while the allocation happens - all in the space of the pointer.
- vy8vWJlco 13y agoI don't think thread-local variables are necessary. Consider a global flag that gets instantiated and synchronized across threads/cores to a state of 0 such that if it ever gets set to 1, it was after publishing and synchronizing the shared instance via a global pointer. This pointer, and the global flag, will not revert for the rest of the singleton's life (noted in the article) - so there is nothing to synchronize past that - and it is only for the guarding flag's 0 (unallocated) case that there is a question if things are synchronized. That check and potential allocation are what need a standard lock (provided by D's syncronized block), which the example already provides.
- rbehrends 13y agoTo get such a global flag with the desired behavior, you need explicit memory barriers. Otherwise, you cannot guarantee that it won't become visible to another processor before the object's data itself becomes visible. The problem here is that the memory accesses for the path that bypasses the synchronization block are not guaranteed to occur in the correct order without barriers.
- vy8vWJlco 13y agoYep - it should be enough to force a coherent view for all as part of the initialization.
- rbehrends 13y agoYes, it's well-known that you can fix the problem with memory barriers. But the article's point was that such memory barriers can be avoided entirely for the common execution path if you have fast thread-local storage (because you can make sure that every thread is forced through the full synchronization block once before accessing anything). Given that memory barriers can be expensive on some processors, that's a not insignificant win.
- vy8vWJlco 13y agoDoes D's synchronization not contain memory barriers (I don't know)? (Edited)