8 ms·
Hmm, let's read this one. > "March 28, 2019" > "Fibers: the Most Elegant Windows API" > "The Windows API — a.k.a. Win32 — is notorious for being clunky, ugly
by uasm 8y ago
Hmm, let's read this one.
> "March 28, 2019"
> "Fibers: the Most Elegant Windows API"
> "The Windows API — a.k.a. Win32 — is notorious for being clunky, ugly, and lacking good taste. Microsoft has done a pretty commendable job with backwards compatibility, but the trade-off is that the API is filled to the brim with historical cruft. Every hasty, poor design over the decades is carried forward forever, and, in many cases, even built upon, which essentially doubles down on past mistakes. POSIX certainly has its own ugly corners, but those are the exceptions. In the Windows API, elegance is the exception."
I'm glad OP just discovered in mid-2019 an API that has been there since Windows XP (ca. 2001). Why then, does OP start by crapping all over the Windows codebase?
You think the Fibers API is elegant? Have a look at IOCP. IO in Windows in general actually, since the days of Windows 95. Event management. Threading and parallelism. The kernel subsystem. The APIs in Windows have remained stable, consistent, forward and backwards compatible - from pretty much day one.
- deleted 8y ago[deleted]
- teh_klev 8y ago> I'm glad OP just discovered in mid-2019 an API that has been there since Windows XP Can I bring to your attention the second para: "That’s why, when I recently revisited the Fibers API, I was pleasantly surprised." I don't think it's the author's first rodeo. The article reads like a retrospective look at the fibers API and how well it's aged.
- userbinator 8y agoIt's been there since NT 3.51, according to https://www.geoffchappell.com/studies/windows/win32/kernel32/api/index.htm https://www.geoffchappell.com/studies/windows/win32/kernel32... and my copy of Win32.hlp (from back when MS documentation was actually proofread before release...)
- PaulHoule 8y agoIOCP is particularly strong when you compare it to select, poll, epoll, kpoll, and all of the other attempts to make asyncio work in Linux.
- machinecoffee 8y agoIOCP in Windows is a good API (although a bit non-obvious), but to be fair, they did have a few tries and missteps before they got there.
- deleted 8y ago[deleted]
- zzzcpan 8y agoIt's not. IOCP is the worst of all of them.
- kentonv 8y agoEh. I used to think that, until I actually wrote event systems based on all of the above. epoll turns out to be my favorite. Buffer allocation with IOCP feels incredibly ugly. You must allocate a buffer before you initiate a read(), and then you must leave that buffer alone until the read() completes. You would think that this has the advantage that the kernel doesn't need to allocate its own buffer, and you get some sort of zero-copy magic where bytes land directly in userspace. But that's not really the case. The socket still needs a buffer on the kernel side in case bytes arrive while no read() is pending. So now you have to allocate a buffer for each socket and the kernel also has to allocate a buffer for each socket, which seems like a waste. This rabbit hole gets deeper. What happens if the socket receives two packets in rapid succession? In a naive implementation, the first packet completes the read(), and signals the completion on the IOCP. The app now has to process that event and start a new read() when it's ready. But in the meantime, the second packet arrives. Whoops, no read() is pending, so now the packet has to go into a kernel-side buffer only to be copied later. But that's a naive implementation. Apparently, in reality, the kernel implements some sort of nagle-like algorithm where it tries to wait a bit for additional packets before it actually signals completion of a read(). But this introduces delay, and many projects (e.g. Chrome) have discovered this delay is rather harmful to certain kinds of performance. I read somewhere that Chrome and others have given up on IOCP and use WSAPoll instead -- but I can't remember where I read this because all the lore about Windows event handling is hidden in random forum threads and Github gists rather than proper documentation. It seems to me that the right way to do what Windows was trying to do here would be for userspace to allocate a ring buffer for the kernel to use, and then the app and the kernel would coordinate the start and end pointers of the ring buffer. Then the kernel needs no buffer of its own; it can always deliver to userspace. If the buffer fills up, the kernel can do exactly what it would do in the case of a regular kernel-side buffer filling up -- apply backpressure and force the peer to retransmit later.
- jstimpfle 8y agoI had a hard time with IOCP and eventuall ditched it because I didn't find a way to avoid creating a separate read buffer for each and every handle. Suppose I want to watch 10K connections, or files, or such. 4K should be a reasonable read buffer size. Do I really need to spend 40M of memory just to read from the handles simultaneously?
- kbenson 8y agoI'm confused (but I'm also not familiar with the API). It sounds like you're saying you want to be able to read from 10k connections simultaneously while allocating less than a byte of buffer space to each, so I assume I'm misunderstanding some aspect of what you're trying to say (or some aspect of the API that puts this into the correct context).
- jstimpfle 8y agoNo I'm saying I want to read from N handles simultaneously using M buffers, where M <= N. Of course, the OS shouldn't write to one of my read buffers that currently contains other buffered data. But that doesn't mean that it has to be N = M, since that is a waste. I guess M = 10 * number of OS threads should usually be plenty to achieve good performance. But I'm happy to be corrected by people more experienced with systems performance.
- kbenson 8y agoI believe I understand. By not allocating in the api call itself you can pre-allocate your own buffers and keep track of which are in-use and not in-use yourself and use some subset of the total amount that would be allocated in each api call be reusing buffers after a particular connection is done with it. That makes sense.
- zzzcpan 8y agoOn unixy systems kernel manages buffers by itself, using as little memory as possible, not allocating anything when no data is received, and userspace doesn't need to allocate any memory either until it is notified and decides to non-blockingly read the data from the kernel. Meaning that waiting for data from connections costs nothing, while on windows it requires preallocating buffer per connection. Basically the same problem that synchronous APIs have. The myth that IOCP is elegant has to die.
- spapas82 8y agoI think this is a must read for new developers: https://www.joelonsoftware.com/2004/06/13/how-microsoft-lost-the-api-war/ https://www.joelonsoftware.com/2004/06/13/how-microsoft-lost... (yes it's from 2004)