5 ms·
I think I've just come to accept that sychronization is the pain point in any language. It's callbacks, promises, and the single event loop in nodejs. It's chan
by hacknat 11y ago
I think I've just come to accept that sychronization is the pain point in any language. It's callbacks, promises, and the single event loop in nodejs. It's channels in golang.
No one can come up with a single abstraction for synchronization without it failing in some regard. I code in go quite a bit and I just try to avoid synchronization like the plague. Are there gripes I have with the language? Sure, CS theory states that a thread safe hash table can perform just about as well as a none-thread safe, so why don't we have one in go? However...
Coming up with a valid case where a language's synchronization primitive fails and then flaming it as an anti-pattern (for the clicks and the attention, I presume) is trolling and stupid.
- yetihehe 11y ago> No one can come up with a single abstraction for synchronization without it failing in some regard. Erlang did. Or at least it's as close as possible.
- hacknat 11y agoI'm not saying Erlang isn't great, but if you need to pass a large datastructure around between Erlang processes then copy message passing starts to be a lot and you need to share memory. You can do it in Erlang, but I'd hardly call it great, and you're avoiding the sync primitive that Erlang offers.
- catnaroek 11y agoHow about Rust's “share by transferring ownership”? (0) In the general case, whatever object you give to a third party, you don't own anymore. And the type checker enforces this. (1) Unless the object's type supports shallow copying, in which case, you get to keep a usable copy after the move. (2) If the object's type doesn't support shallow copying, but supports deep cloning, you can also keep a copy [well, clone], but only if you explicitly request it. This ensures that communication is always safe, and never more expensive than it needs to be. --- Sorry, I can't post a proper reply because I'm “submitting too fast”, so I'll reply here... The solution consists of multiple steps: (0) Wrap the resource in a RWLock [read-write lock: http://doc.rust-lang.org/std/sync/struct.RwLock.html http://doc.rust-lang.org/std/sync/struct.RwLock.html], which can be either locked by multiple readers or by a single writer. (1) The RWLock itself can't be cloned, so wrap it in an Arc [atomically reference-counted pointer: http://doc.rust-lang.org/std/sync/struct.Arc.html http://doc.rust-lang.org/std/sync/struct.Arc.html], which can be cloned. (2) Clone and send to as many parties as you wish. --- I still can't post a proper reply, so... Rust's ownership and borrowing system is precisely what makes RWLock and Arc work correctly.
- hacknat 11y agoWhat if you want multiple readers at once, and a writer thrown in once in a while? Edit: Okay, my point was that the sync primitives of most languages alone can't save you and you're using RWLock in your example, so clearly ownership by itself doesn't solve everything, right? That's the point I'm trying to make. Edit2: Hmm, I'll have to check that out. I don't know that I would call Rust's ownership model super easy to reason about, but it is nice that the compiler prevents you from doing so much stupid $#^&.
- pcwalton 11y ago> Okay, my point was that the sync primitives of most languages alone can't save you and you're using RWLock in your example, so clearly ownership by itself doesn't solve everything, right? The thing is that Rust ensures that you take the locks properly. It's an compile-time error to forget to take the lock or to forget to release the lock†. You can't access the guarded data without doing that. † For lock release, it's technically possible to hold onto a lock forever by intentionally creating cycles and leaking, but you really have to go out of your way to do so and it never happens in practice.
- azth 11y ago> Hmm, I'll have to check that out. I don't know that I would call Rust's ownership model super easy to reason about, but it is nice that the compiler prevents you from doing so much stupid $#^&. It's much better get compile time errors than deal with very hard to reproduce data races.
- kazinator 11y agoOnly, as usual, in situations when all else is equal. By the way, on a related note, data races themselves are easier to reproduce than the visible negative consequences of those races on the execution of that program. That's the basis of tools like the "Helgrind" tool in Valgrind. That is to say, we can determine that some data is being accessed without a consistently held lock even when that access is working fine by dumb luck. We don't need an accident to prove that racing was going on, in other words. :)
- felixgallo 11y agoErlang lifts sufficiently large binaries into refs, which isn't perfect but pragmatically helps a lot with that problem.
- jerf 11y agoI've been bitten by the fact that Erlang lacks a channel-like primitive. You've got half-a-dozen "pool" abstractions on github because it's actually sorta hard to run a pool on pure asynchronous messages when there is absolutely no way to send a message out to "somebody", the way Go channels can have multiple listeners. I know that would only work on a local node but there's already a couple of functions that have already penetrated that abstraction anyhow. You also have to deal with mailboxes filling up, still have problems with single processes becoming bottlenecks, and the whole system is pervasively dynamically typed which is fine until it isn't. It is pretty good, but it's not the best possible. (Neither is Go. I still like Erlang's default of async messages better in a lot of ways. I wish there was a way to get synchronous messages to multiple possible listeners somehow in Erlang, but I still think async is the better default.)
- yetihehe 11y ago> You've got half-a-dozen "pool" abstractions on github because it's actually sorta hard to run a pool on pure asynchronous messages when there is absolutely no way to send a message out to "somebody" You can store receivers in ets table and implement any type of selection algorithm you want or have some process which selects workers. There is no default method, because one default method is not good for everyone and people will complain that it's not good for them. Implementing pools is easy in erlang, I've done tailored implementations for several projects. > You also have to deal with mailboxes filling up Yeah, unless you implement back-pressure mechanism like waiting for confirmation of receiving. In ALL systems you have to deal with filling queues. > I wish there was a way to get synchronous messages to multiple possible listeners somehow in Erlang You can implement receiver which waits for messages and exits when all are received or after timeout, it's trivial in erlang but I haven't needed it yet. Here is a simple example: receive_multi(Acc,0) -> Acc; receive_multi(Acc,Num) -> receive {special,Data} -> receive_multi([Data|Acc],Num-1) after 5000 -> Acc end.
- jerf 11y ago"You can store receivers in ets table and implement any type of selection algorithm you want or have some process which selects workers." Your process that selects workers has no mechanism for telling which are already busy. It is easy to implement a pool in Erlang where you may accidentally select a busy worker when there's a free one available. Unfortunately, due to the nature of the network and the way computations work at scale, that's actually worse than it sounds; if one of the pool members gets tied up, legitimately or otherwise, in a long request, it will keep getting requests that it ignores until done, unnecessarily upping the latency of those other requests, possibly past the tolerance of the rest of the system. "You can implement receiver which waits for messages and exits when all are received or after timeout, it's trivial in erlang but I haven't needed it yet." That's the opposite of the direction I was talking about. You can't turn that around trivially. You can fling N messages out to N listeners, you can fling a message out to what always boils down to a random selection of N listeners (any attempt to be more clever requires coordination which requires creating a one-process bottleneck), but there is no way to say "Here's a message, let the first one of these N processes that gets to it take it". You wouldn't have so many pool implementations if they weren't trying to get around this problem. It would actually be relatively easy to solve in the runtime but you can't bodge it in at the Erlang level; you simply lack the necessary primitives.
- zzzcpan 11y ago> I think I've just come to accept that sychronization is the pain point in any language. No, it's not. Everything is easier with event loops, because everything is always synchronized. And since it is, there is no need for concurrent hash tables, locks, channels, you name it. There is also no more shutdown and cancellation problems, you get them for free and easier than anything. The only thing left is a __consistent__ API with callbacks. But as long as you go with higher order functions you are not going to have any problems.
- hacknat 11y agoWhat if you need to do a compute intensive task on a large data structure? You know you might need to take advantage of more than one core and sharing memory between the threads will be difficult. Assuming you're talking about nodeJS, nodeJS serializes and deserializes objects in and out of C++ land in order to do compute intensive tasks. Hardly a catch all! Are event loops good at some things? Of course! Are the good at everything. Are you high?
- zzzcpan 11y agoWell, no, I'm not talking about nodejs. Just in general, about event loops in programming languages. > What if you need to do a compute intensive task on a large data structure? That's a very specialized thing, not something general, that everyone needs. But either way there is no problem abstracting it away with higher order functions in event loops. However, everyone will most definitely need networking and doing networking by sharing memory between threads is very very hard. Event loops are much easier for that.
- jerf 11y agoEither your event handlers are going to be called in a nondeterministic order, or they won't. If they are going to be called in a nondeterministic order, you still have access control issues and can get yourself into all sorts of concurrency-style problems. If they aren't going to be called in a nondeterministic order, perhaps because you just have a single cascade of events (open socket, write this, get that, close socket), then in a language like Go you just write the "synchronous"-looking code, and you don't have to write the code as if it's evented. You have only marginally more sharing problems than the event loop. Raw usage of event loops are a false path. They solve very few problems and introduce far more.
- nostrademons 11y agoBecause concurrency is hard. You can't reason about concurrent programs the way you can about sequential ones, and no abstraction is going to completely fix that. After having worked with it a fair bit, however, I'm beginning to really like Promises + async/await (as in ES7, Python 3.4, and C#). It manages to keep most of the concurrency explicit while still letting you use language mechanisms like semicolons, local variables, and try/catch for sequencing. If you make sure your promises are pure, you can also avoid the race conditions & composability problems of shared state + mutexes. (Although that requirement is easier said than done...it'll be interesting to see what Rust's single-writer multiple-reader ownership system brings to the mix.)