6 ms·
While I think it's great that more and more developers are being exposed to building network applications using async I/O (which I guess is "evented" now) via N
by jamwt 16y ago
While I think it's great that more and more developers are being exposed to building network applications using async I/O (which I guess is "evented" now) via Node.js, I think it's worthwhile to point out that the state of the art has moved well beyond these kind of callback frameworks. The reason is simple: it sucks programming in callback patterns on serious, large projects. You end up with lots of routines that are 6 or 7 callback chained together--and don't forget to attach error callbacks as well at each point.
In the Python world, for example, eventlet, gevent, and diesel (disclosure: my project) all use coroutines to achieve very high performance asynchronous I/O that's still written in a "synchronous" style. And before that, CCP games has been using stackless python to do coroutine-based networking for the servers behind the very popular MMORPG eve online (http://www.tentonhammer.com/node/10044 http://www.tentonhammer.com/node/10044).
Even slicker, in languages like Erlang and Haskell (via forkIO + the select/epoll I/O manager), support for asynchronous networking using "blocking-style" code is a fundamental language/compiler/vm feature. When you write any code in these languages, you're transparently taking advantage of the same scaling characteristics node.js and its ilk provide.
So, I'd humbly suggest to anyone getting serious about doing this kind of programming that they investigate some of these systems (which are usually built by weary callback-style async I/O veterans) and leapfrog the rest of their peers. Consider Node.js a gateway drug.
Linkage:
http://eventlet.net/ http://eventlet.net/
http://github.org/jamwt/diesel http://github.org/jamwt/diesel
http://en.wikipedia.org/wiki/Asynchronous_I/O#Light-weight_processes_or_threads http://en.wikipedia.org/wiki/Asynchronous_I/O#Light-weight_p...
http://www.galois.com/~dons/slides/a-scalable-io-manager-for-ghc.pdf http://www.galois.com/~dons/slides/a-scalable-io-manager-for...
- icey 16y agoI think developers get attracted to frameworks because of fun examples like the one linked in the posting. Until there are multitudes of interesting examples using other tech, I think it's going to be an uphill battle to attract significant numbers of developers to them. Since many of the Node.js demos like this one are open-source it's much easier for developers to download working code and make modifications to see how everything works. The technically superior solutions often get overlooked for the most accessible ones.
- jorangreef 16y agoRe: "The technically superior solutions often get overlooked for the most accessible ones." Are there any server-side Javascript environments being overlooked? And would it be fair to say that these are technically superior to V8 + Node?
- alec 16y agoHe's not talking about technically superior server-side Javascript environments, he's talking about langauges and frameworks that provide more than just a callback model.
- dylanz 16y agoI definitely agree when it comes to the code scalability of something like Node.js (although, I haven't built gargantuan javascript applications before, but am going to assume that it's not necessarily a pleasant experience). Diesel looks great. When looking into alternatives in Erlang/Haskell/Scala, or XMPP/BOSH, either it's not easy to find quick how-to examples, or I'm looking in the wrong places. Any fingers in the right direction would be welcome.
- tibbe 16y agoThe Glasgow Haskell Compiler (GHC) uses async I/O to implement all I/O actions (e.g. printing to stdout, opening a socket) and it's completely invisible to the programmer. Bryan O'Sullivan and I have rewritten the implementation used in GHC and GHC 6.14 will offer much better scalability (using epoll/kqueue) than previous version. Just fork of a thread per connection (using forkIO) and just call normal I/O actions. The implementation will multiplex those forkIO threads onto a single OS thread that calls epoll/kqueue.
- tdmackey 16y agoCoroutines bring their own headaches and baggage. To say that callbacks are behind "state of the art" is a little misplaced,; its just a different way of doing things. With coroutines you have to worry about IO all over the place and have to make your functions coroutine safe and Ryan Dahl, the node.js creator, will argue with you all day long about coroutines vs callbacks. I do agree, though, that developers should not ignore other ways of doing things, just don't discredit the callback way of doing things as it can be useful.
- jerf 16y ago"With coroutines you have to worry about IO all over the place and have to make your functions coroutine safe and Ryan Dahl," No, no, a thousand times no. Node.js partisans really need to start actually using Erlang or Haskell for a little while before spouting this canned line off. You do not have to jump through enormous hoops to deal with IO in Haskell or Erlang, it just works. Callbacks are behind the state of the art. Coroutines or generators or any other cooperative-multitasking primitives in languages that didn't support them from day one and have enormous sets of libraries not "cooperative-aware" are behind the state of the art, too. This is all just "cooperative multithreading" again and I am yet to see anyone explain why this time is going to turn out any different than last time we tried cooperative multithreading.
- codahale 16y agoWould that I had more upvotes for this. The actual cutting edge of event-driven server architecture[1] looks something like this: server = do { sys_call_1; fork client; server; } client = do { sys_call_2; } (And even then it has marginal benefits in terms of throughput and latency compared to threaded implementations.) [1] Li and Zdancewic. Combining events and threads for scalable network services. Proc. 2007 PLDI (2007).
- aredridel 16y agoIt's really a question of how much state you're storing, and how you're dealing with that. Many languages and many runtimes allocate stacks in megabytes, and store every local variable in it. With a callback-based system, stacks are short, and you explicitly carry that state. You KNOW what state you're storing, and you can see it easily. There's some benefit to that. And the pattern for aggregating functions is different: Receive a starting event, emit a done event -- you aggregate processes into sets of events, not into function calls. So yeah, it's not going to follow some of the same patterns that non-event-driven code follows, and event-driven code is going to look rather different than callback-passing code. I think Ryan's right about coroutines, too: Coming back to vastly different state after a simple function call is basically going back to the days of using global variables.
- creationix 16y agoWhether you like the model or not, the fact is that JavaScript with it's callback/event based model is what browsers understand. And the internet isn't going anywhere anytime soon. Node is an attempt at using this successful model on the server too where is can solve massive scalability issues using the exact paradigm front-end devs already know. Also while there are many technical ways to make code look blocking, but really be running other events under the hood, it's this exact implicit running of "other stuff" that makes writing threaded code so hard. You have to assume that things can change between every function call because you don't know if somewhere down the chain it's doing pseudo-blocking. JavaScript's model is simple, you provide callbacks and you know exactly at what boundaries things can happen asynchronously. In summary, node is one way of doing it. We think it's a better model and it's proven itself in the browser. If you think another model is better, than by all means go for it! Let the leapfrogging begin ;)
- _delirium 16y agoI agree this is currently the best way of using similar models on both the client and server side. It'll be interesting to see if anything changes, on either side or both, when webworkers (browser threads) get more widespread support. At least it's likely that more direct comparisons of the programming models will be possible in the future.
- jamwt 16y agoI think javascript is a fine little language. But: I don't make my server-side language decisions based on what happens to be in the browser. To think that a language required for use in such a constrained environment is just _coincidentally_ also the best language to use server-side where people have been doing event-based programming for decades seems... unlikely. It would be as ridiculous as asserting that we should all use postscript in the browser because that's the standard we've been using in printers for the last 20 years. I suppose if you make the argument that browsers dictate that developers _must_ learn javascript, and therefore, with Node.js they can also program server-side without ever needing to learn a second language, that is a sound argument. I don't, however, know many good programmer that spend a career knowing exactly one language.
- folletto 16y agoProgrammer-mentality is somewhat different from other mentalities, but it's still bound to simplicity laws. The technologies you are describing were already there, but it's not just a problem of technology, feature set or execution power. It's also a matter of simplicity. Node.js is working because it gives so much power with a very simple approach. And by approach I mean also how much is simple to understand it and start producing something useful. That's exactly the reason why we use more abstract languages, and that's the reason why Node.js is getting a lot of attention recently. If you don't know any language, do you think it's simpler to start with Erlang or Haskell, or with JavaScript? I don't have any proof, but my bet is surely on JavaScript. :) And even the most obvious things are important. Because maybe a uber-developer can ignore the small details, but most of the programmers aren't uber, they just want to develop easily and happily. Every little detail matters. Just see how many steps you need to install Erlang on your machine, and how many steps you need to do the same with node. It seems stupid, I know. But when you sum every detail... it matters. :) At the same time, this gives power to the uber-programmers out there to bring on more cutting-edge solutions when they need them, and feed the "simple" level with their discoveries and experience, making the language and frameworks evolve. If you're right, one day we are going to have simple coroutines - in the complex, environmental and social sense expressed above. Maybe even in Node.js, because in the end JavaScript 1.7 afaik supports yield and V8 could implement it in the future. ;)
- gcampbell 16y agoHow many steps you need to install Erlang on your machine: "sudo apt-get install erlang"?
- silentbicycle 16y agoFirst off, is everyone running your Linux distro? ("sudo pkg_add erlang" works for me, though.) Besides, are you seriously implying that the hard part of getting started with Erlang is installing it? A lot of developers aren't willing to sink time into learning a language that actually has new ideas, you know? Hell, it's not even OO. :)
- Andi 16y agoUse Step, man: http://github.com/creationix/step http://github.com/creationix/step
- creationix 16y agoI have written extensively on this topic, here are a few articles. Here is the progression leading up to the development of Step. http://howtonode.org/control-flow http://howtonode.org/control-flow http://howtonode.org/control-flow-part-ii http://howtonode.org/control-flow-part-ii http://howtonode.org/control-flow-part-iii http://howtonode.org/control-flow-part-iii http://howtonode.org/do-it-fast http://howtonode.org/do-it-fast http://howtonode.org/step-of-conductor http://howtonode.org/step-of-conductor
- IsaacSchlueter 16y agoI've heard this before. "Oh, callbacks are fine for little things like this hello world demo, but for Big Programs, you need (threads/coroutines/etc.) because they're Serious Business." That might be true. Maybe for Big Programs, something else is better. But I think the problem is that setting out to build a Big Program for Serious Business is Doing It Wrong. The best frameworks (or, at least, my favorite frameworks) are collections of small tools with consistent interfaces and interchangeable parts. If you can manage callbacks for a 100-200 LOC program, then you can build a Big Program out of such modules with a little bit of forethought. Instead of setting out to write a Big Program, why not try to figure out how to express your Big Problem in terms of a bunch of Little Problems, and then come up with a way to have Little Programs assemble. Then, assemble the Little Programs to solve the Little Problems. There is no solveable Big Problem which cannot be broken down into some finite set of Little Problems. Callbacks make it natural to solve problems in this manner. I actively develop a several thousand line NodeJS project. It's quite nice, actually.
- jamwt 16y agoI don't know how to completely address your post b/c it responds to assertions I didn't make. I said nothing about "Serious Business"--and bundling threads and coroutines together is like bundling horseshoes and bicycles. I made no defense of threads. Callback style, within an I/O handler, is (almost always) a concession to the event loop. When I'm writing a program, I write: routine(): result = do_one_thing() do_next_thing(result) do_whatever() With callback based I/O, you need the routine to end so the "reactor" up the call stack can get control again to do I/O on your behalf: routine_1(): return make_some_io_request(callback=routine_2) routine_2(): return make_another_io_request(callback=routine_3) routine_3(): return send_response_to_original_socket() Now, if you were desperate, you could just call over to the reactor like this: routine(): resp = do_io_for_me() next_resp = do_more_io_for_me(resp) return transform(next_resp) .. but, there are (at least) two problems with this: 1. You're making the stack deeper every time you "call over" to the reactor... you will stack overflow eventually 2. In most languages, the reactor doesn't have a way to "call back into" your stack frame and resume it. Coroutines are a way to "call over" to the reactor, have your stack state frozen (in heap space, without going deeper in the call stack), and then "unfreeze it" later on, when the reactor has completed the IO. The greenlet package provides this for Python; Haskell and Erlang have native support for this. And, this is ultimately how you want to program. A(); B(); C(). Sure, there are exceptions for UI driven callbacks, and callback patterns aren't all bad, but typically in network IO cases, the callbacks are explicitly a concession to the I/O loop's need to be scheduled, made necessary by a missing language or vm feature. And: cleaner code benefits 200 LOC apps as much as it does 200k LOC apps. A good design is a good design at any scale.
- mikealrogers 16y agoI've said it before and I'll say it again: if callbacks are so hard how come all these liberal arts majors write a shit load of working jQuery code? I think the real problem is that callbacks are too simple for people that want to obsess about engineering new primitives and programming styles.
- jorangreef 16y agoI agree with you that a synchronous programming interface is easier than an asynchronous programming interface. However, none of the systems you suggest make it possible to write code that can be run by both the server and the browser. If you were writing web applications, which represent the typical use-case for Node, then this would be a prerequisite.
- allertonm 16y agoIt can hardly be a "prerequisite" given that most web applications built to date have managed to do without it. It's a convenience, at least intuitively, but in practice one has to wonder what parts of your code run in both browser and server and do async I/O.
- jorangreef 16y agoFirstly, most web applications to date do not: 1. Work offline. 2. Provide sub-20ms interface switching response rates. 3. Protect the interface against extended network or server failure. 4. Buffer against poor network performance. So I would agree with you that they have "managed to do without it" (code-sharing) so far. But they have done so only by avoiding responsibility for these goals, by considering them a convenience, or by not considering them at all. Secondly, I can tell you from experience that a non-Javascript server-side framework can do the above but only with difficulty. Existing popular frameworks do not provide "convention-over-configuration" assistance in terms of straddling the network or exposing Model definitions to Javascript etc. They render interfaces on the server-side, rather than letting controllers and views interact on the client for example. Indeed it would not even be fair to expect these frameworks to meet the above requirements, simply because they were not designed to do so. Eventually one must arrive at a realization that web applications (as distinct from "web sites") will mostly be written in the same language AND framework on both client and server. Fear not, this will probably be more fun.
- allertonm 16y agoI simply do not see why your goals 1-4 require a) the same language on the server as the client and b) the use of a non-blocking I/O model at the server. It appears to be a leap of faith on your part. While marshalling of model objects to JSON is obviously going to be easier to achieve if your model objects are Javascript objects, marshalling to JSON is trivially achieveable in almost all programming languages. To suggest that developers should - to achieve this microscopic payoff - choose to build servers using a language with a crippled concurrency model (i.e none) is laughable.