7 ms·
Rob Pike discovers sendfile(2)
- unwind 16y agoFor the un-initiated, sendfile() is a system call that sends data between two file descriptors. The intent is to make the kernel do the read/write cycle instead of the application (user-level) code, thereby cutting down the number of times the data needs to be mapped between kernel and userspace memory spaces. The manual page: <http://linux.die.net/man/2/sendfile> http://linux.die.net/man/2/sendfile>. A related Linux Journal article: <http://www.linuxjournal.com/article/6345> http://www.linuxjournal.com/article/6345>.
- jrockway 16y agoYeah, it's all about context switching, which is just unnecessary work. Doesn't matter if you are serving 10 dynamic HTTP requests a day on a 16-core super machine. Does matter if you are serving the same file 100,000 times a second from your phone :) (I do wish it worked for any fd to any other fd, because I have to write that code myself rather frequently. Example: copying data from a pty to the real terminal.)
- tonfa 16y agoWould splice help? http://en.wikipedia.org/wiki/Splice_(system_call) http://en.wikipedia.org/wiki/Splice_(system_call)
- FooBarWidget 16y agoUnfortunately not, splice requires that one of the fds is a file. I want to forward data from a socket to another.
- wmf 16y agoThat's not what the man page says. AFAIK HAProxy splices from one socket to another.
- fragmede 16y agoIn 2.6.31, support was added to > Allow splice(2) to work when both the input and the output is a pipe. http://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.6.git;a=commitdiff;h=7c77f0b3f9208c339a4b40737bb2cb0f0319bb8d http://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.6...
- alexgartrell 16y agoThis goes half way toward solving the problem. What most people want is a pure socket->socket. As evidenced by haproxy, you must still use a pipe intermediary, requiring multiple data copies (on the plus side, it's still fast). ret = splice(fd, NULL, b->pipe->prod, NULL, max, SPLICE_F_MOVE|SPLICE_F_NONBLOCK); edited to add (because I couldn't reply): It almost certainly is faster than read/write, because it's straight memcpy's (which are pretty fast on modern hardware), instead of memcpy's, context switching, special read()/write() logic, and other stuff.
- FooBarWidget 16y agoIs that still faster than userspace forwarding with read() and write()?
- caf 16y agoThe pipe intermediary doesn't necessarily add a copy. You SPLICE_F_MOVE the data from the first fd to the pipe, then SPLICE_F_MOVE it from the pipe to the second fd. If either or both of those can be done zero-copy, they will be. The pipe intermediary is just a way of holding onto the reference-counted pages.
- ithkuil 16y agoThe sendfile manpage says: " Presently (Linux 2.6.9): in_fd, must correspond to a file which supports mmap(2)-like operations (i.e., it cannot be a socket); and out_fd must refer to a socket. " so apparently you cannot use sendfile for your scenario.
- berntb 16y agoI hadn't seen splice, it was cool. But what really blew my mind was the discussion of a system call on Wikipedia. :-) To put it into the wikipedia is obvious I guess, but I just fell in love with the net again. Thank you. Edit: I really should do some low level stuff again.
- gruseom 16y agoI'm glad you posted that. It links to the following superb email thread about slice(), tee(), and other zery-copy magic: http://kerneltrap.org/node/6505 http://kerneltrap.org/node/6505
- leif 16y agoIt's not about context switching, it's about having to copy data from kernel space to user space and back. If there's no processing done between the calls to read() and write(), then there's no point copying it to user space.
- sophacles 16y agoread() then write() is expensive even using zero-copy techniques because there are at minimum 4 context switches for every chunk of data. Further, every time the process stops running (another is scheduled say) the copy stops. With sendfile as a system call, this is not a problem as the kernel is working at this transfer every time it is running. (i.e. every context switch and every interrupt -- actually with dma, it the send could be happening even when the kernel is not running...)
- aliguori 16y agoNo, this is not correct. The idea of sendfile() is that you can send a physical address (probably something in the buffer cache) directly to a network adapter in a zero copy manner. That's current impossible with a send()/write() to a file descriptor in Linux because you can't construct an SKB from a userspace address without forcefully pinning memory. Pinning memory from userspace is a privileged operation. OTOH, since buffer cache is not part of the memory of a process, you can obtain a physical address of it. There's no magic in the kernel that avoids context switches. If a kernel thread has to switch to another kernel thread, that's still a context switch. And syscalls are ridiculously cheap on modern hardware. sysenter is like < 100 cycles. That said, there's no theoretical advantage to doing this via sendfile verses doing a send() from an mmap()'d file. If Linux was smarter, it could realize instantly that the address given to send is file backed and could construct an SKB from the physical memory without pinning. This is what the reference to a "5 minute hack" was vs. Rob Pike's claim that the interface exists to work around a problem in Linux. He's completely correct.
- sophacles 16y agoI'm not disagreeing with this. However, the specified algorithm which I was posting about was read then write. It is what the parent said, not mmap then send. The two sets of calls have very different semantics as you said. For the case of data = read(fd, blocksize, dataptr); while (data > 0) { write(sock, data, ...); data = read(fd, blocksize, dataptr); } My statement holds true. Even if the context switching is minimal overhead, there are more context switches caused by my code, and further blocking and other processes add more delays.
- jacquesm 16y agoRemove extra > from url to make it work. Think 'netcat'.
- jacobolus 16y agoIt’s really annoying that such bugs in news.arc’s renderer never get fixed. Surrounding a URL in <> is the official RFC-recommended way to do things. From RFC 2396 (1998): > In practice, URI are delimited in a variety of ways, but usually within double-quotes "http://test.com/ http://test.com/, angle brackets <http://test.com/> http://test.com/>, or just using whitespace [...] Using <> angle brackets around each URI is especially recommended as a delimiting style for URI that contain whitespace. http://www.ietf.org/rfc/rfc2396.txt http://www.ietf.org/rfc/rfc2396.txt
- jbarham 16y ago> sendfile() is a system call that sends data between two file descriptors Not quite. As Pike complains, the first file descriptor must be mmap-able and the second must be a socket. I think his objection is that such narrowly applicable system calls do not belong in what should be a general purpose API.
- jbjohns 16y agoOn Linux. I don't believe all Unix OSes have that requirement.
- loewenskind 16y ago>It can be written in a few lines of efficient user code. I'm not sure he's understanding what this is. There is no copying needed here at all. The kernel could make the hard drive write to a place in memory, have the NIC read from that place and just manage interrupts between the two. The kernel wouldn't have to touch the data at all. This may not be how Linux does it (given the requirement for a memmap'able file descriptor) but that would be possible at least. I don't think you could do anything near this in user code.
- acqq 16y agoBy having mmap you avoid copying from the kernel RAM space to the user RAM space, mmap makes both sharing the same chunk, therefore the restriction for the thing being mmapable. If you don't allow mmap the system will certainly have to copy between kernel and user memory (in both directions with separate read and write calls).
- loewenskind 16y agoNo, sendfile takes a file descriptor and forwards the contents to a socket. It wouldn't have to mmap the file because the user code never touches the file data. This is probably why other OSes that support sendfile don't have this restriction. >If you don't allow mmap the system will certainly have to copy between kernel and user memory Why? All the kernel has to do is set the DMA location for the drive to be a kernel buffer and follow any inodes when the file is split up on the disk. The NIC driver can just be told to load the data from those locations directly so no kernel copying at all, much less into user space.
- acqq 16y ago> Why? I wrote about mmap in the context of separate write and read calls (i.e. no senddata call exists), not in the context of senddata.
- kmavm 16y agoRob Pike deserves the benefit of your doubt when discussing UNIX implementations. http://en.wikipedia.org/wiki/Rob_Pike http://en.wikipedia.org/wiki/Rob_Pike The read(2)/write(2) solution from userspace would involve no copying as well, assuming a competent kernel implementation. The only "overheads" to speak of relative to what's possible with sendfile("2") would be those associated with writing one contiguous page table entry for every 4KB of data to set up the mappings; since sendfile(2) requires that you use mmap'able input, it probably incurs the same overheads. Fine, let's say doing the right thing with read(2)/write(2) is hard, even really hard, and sendfile(2) is faster today. Making expedient shortcuts to performance in the system call API has historically not turned out well in UNIX. People write software which depends on this interface, and that software may well outlive any existing hardware and its quirks.
- ithkuil 16y agoThe unnecessary data copying problem, as Robert Pike suggests, can be also solved by a more generic Zero-Copy approach, instead of adding a specific single purpose system call. http://www.linuxjournal.com/article/6345 http://www.linuxjournal.com/article/6345 http://kerneltrap.org/node/294 http://kerneltrap.org/node/294 http://www.cs.duke.edu/ari/trapeze/freenix/node6.html http://www.cs.duke.edu/ari/trapeze/freenix/node6.html It has to be noted however that often the term Zero-Copy is used to describe a technique which avoids memory copying by employing virtual memory remapping. VM tricks are also expensive because, depending on the architecture, it might require flushing the TLBs and impact subsequent memory accesses. The advantage of this way of zero copy approach thus depends on several factors such as the amount of data being transferred to the kernel's buffers. I don't have any recent data regarding real word performances, any references are welcome. However it's far from being self-evident that VM tricks can rival the performance of a dedicated 'sendfile' like system call.
- tptacek 16y agoFor the uninitiated, the TLB is the thing that keeps your MMU hardware from having to trawl through the page directory in memory every time it accesses a virtual address; it's a cache, and you generally want to avoid flushing it.
- mfukar 16y agoRob Pike should read about D-Bus in the kernel[1], next. Maybe he'll have some comments. No, I mispoke, I think he'll definitely have some comments. [1] http://git.collabora.co.uk/?p=user/alban/linux-2.6.35.y/.git;a=summary http://git.collabora.co.uk/?p=user/alban/linux-2.6.35.y/.git...
- wmf 16y agoThere is a legitimate debate about the performance vs. elegance tradeoff that sendfile represents. Putting D-Bus in the kernel is just a bad idea.
- mfukar 16y agoI disagree. I think it's nothing but a good idea compared to its current state, for the following reasons: - IPC in userspace is insecure, and currently easily subverted by rootkits. - When a system is under heavy load, the dbus daemon becomes a choke point for scheduling. "Spammy" processes can prevent messages from reaching higher priority processes. - It is a lifesaver for embedded devices with poor multitasking capabilities. Think about ucLinux. - It's not for everyone. It's like PF_RING for packet capturing. - It makes sense not only for performance, but also reliability. - Even though DBus can be improved in several other aspects, I think it's nice that someone sat down and wrote this instead of bikeshedding on some website. 2c
- crux_ 16y agoI was just doing research on using sendfile's modern zero-copy replacement(s), splice & tee: http://kerneltrap.org/node/6505 http://kerneltrap.org/node/6505 They were linked elsewhere here, but worth repeating in a top-level comment. :) Possibly a handy tool for the rare times you need to squeeze blood out of a stone.
- luckydude 16y agoHi, I'm the guy who came up with the splice idea. It's based on what I learned doing this: www.connectathon.org/talks96/bds.pdf which was for the EIS (Earth Imaging System) project, a government effort to image the earth about 15 years ago. That project eventually had 200Mhz MIPS SMP boxes moving data through NFS at close to 1Gbyte/sec 24x7. So far as I know, nobody else has ever come close to that even with 10x faster CPUs. Most of the people in this thread pretty clearly don't understand the issues involved, Rob included (sorry, Rob, go talk to Greg). Moving lots and lots of data very quickly precludes looking at each byte by the CPU. The only thing that should look at each byte is a DMA engine. Sendfile(2) is a hack, that's true. It is a subset of what I imagined splice(2) could be (actually splice(3), the syscalls are pull(2) and push(2)). But it's a necessary hack. Jens' splice() implementation was a start but wasn't really what I imagined for splice(), to really go there you need to rethink how the OS thinks about moving data. Unless the buyin is pervasive splice() is sort of a wart.
- deleted 16y ago[deleted]
- fragmede 16y ago> So far as I know, nobody else has ever come close to that even with 10x faster CPUs. With 10-gigabit becoming commonplace in high-end server rooms, dual 10-gigabit cards hitting the market, and 100-gigabit on its way, I've got to ask, what you mean by no one has come close? (Certainly, 1 GByte/sec is achievable; splice is a large part of that.)
- luckydude 16y agoThe big print giveth, the fine print taketh away :) Just because you have a 10 gigabit pipe (which is 1.2Gbyte/sec) that doesn't mean you can fill it. And filling it with a benchmark is different than filling it with NFS traffic. So far as I know, nobody has come close to what SGI could do years ago, i.e, stream 100's of MB/sec of data off the disk and out the network and keep doing it. I've tried with Linux to build a disk array that would do gigabit, not 10-gigabit, and couldn't come anywhere close. I think about 60MB/sec was where things started to fail. I'd love to hear about a pair of boxes that could fill multiple gigabit pipes with the following benchmark: $ tar cf - big_data | rsh server "tar xf -" You can set any block size you want in tar, I don't care, the files can be reasonably large, etc. If you manage to fill one pipe, then try it with 10-gigabit and let's see what that does. I'd love to be wrong but so far as I can tell, we're nowhere near what SGI could do.