5 ms·
Their open conect FreeBSD boxes are nuts. The amount of data they spew is crazy. 1/3 of all internet traffic. One service. Mind blowing.
by X86BSD 9y ago
Their open conect FreeBSD boxes are nuts. The amount of data they spew is crazy. 1/3 of all internet traffic. One service. Mind blowing.
- toomuchtodo 9y agoIt gives me hope that someone could still go out and challenge YouTube using a stable of these at various internet exchanges and a central object store/web front end.
- X86BSD 9y agoI think YouTube simply has first to market mindshare. I personally use and prefer Vimeo. It's a better experience and quality. I don't think anything technical or special keeps YouTube number one other than first to market mindshare.
- reificator 9y agoPart of that mindshare is that I only go to Vimeo when I can't find something on YouTube, and then when I find it on Vimeo it's incomplete and in approximately 240p. It's not fair to Vimeo, because they actually had a copy and YouTube didn't, but because of multiple experiences like that I was surprised to see a positive remark about their quality.
- analogic 9y agojust hope there's at least enough competition to keep them halfway honest
- blocked_again 9y ago> I don't think anything technical or special keeps YouTube number one other than first to market mindshare It's the youtube content creators like PewDiPie, Casey Niestat, Kurzgesagt etc that makes Youtube special just like that community that makes Hacker news special.
- StillBored 9y agoI suspect (having worked on an application handing > 100Gbit of I/O per node) that the OS choice doesn't really matter that much. That is because in my case, the data path portion basically talked directly to a couple PCIe boards. It bypassed the entirety of the kernel outside of some setup API's to claim memory/interrupts/etc. That meant the transfer limits generally came down to lack of PCIe or memory bandwidth (depending on which generation of machine/configuration we were using). The CPU's in the machines spent 99.99% of their time running code we wrote. Despite the talents of most OS developers, generic OS/driver code is not optimized for absolute performance in one case, rather it tends to be tuned to perform well over a wide range of situations. The general goal is to be a fair arbitrator of system resources to multiple competing processes. Further, most general purpose OS's are under the assumption that I/O is slow or low bandwidth. Take the entirety of the linux filesystem/block layer/scsi layer, which is written under the assumption that the system is attached to a high latency low bandwidth spinning disk, so burning a few cycles coalescing requests, or handling the page cache isn't a big deal. That code doesn't scale when you plug it into a NVMe disk with 2GB/sec of bandwidth, much less a storage network with 100GB/sec of IO bandwidth. Anyway, if you throw all these assumptions away and ignore modern "best practices" development models of assembling piles of unrelated libraries to solve a task, you end up with really lean (probably fits in the L1i cache) software that can perform two or three orders of magnitude faster than similar code written using modern methods.
- luckydude 9y agoThat sounds like a ftp-like measurement of throughput, and yeah, what you said will work for that just fine. Netflix connections are typically about 1mbit/sec each (older apps open up ~4 connections per video for reasons that are no longer valid but the apps aren't all updated). So to fill a 100Gbit pipe they have 100,000 connections running at the same time. Which makes filling that pipe super super impressive.
- StillBored 9y agoIn our case we were doing a fair amount of data manipulation, so it wasn't strictly a case of pushing the data through, although we had higher bandwidth per stream. But, there are a bunch of different ways to solve the problems. I guess how impressive it is depends on they have gone about solving their particular cases. There is a fair number of network accelerators that offload individual stream level management to little cores running on the network adapter itself. Cavium, EzChip and now even companies like mellanox are playing in this space https://www.enterprisetech.com/2017/10/04/mellanox-ethernetarm-nics-lighten-cpu-burden/ https://www.enterprisetech.com/2017/10/04/mellanox-etherneta.... So, i'm not sure the impressive parts are necessarily in the stream counts but what they must be doing to "align" (for lack of a better term) them. AKA the trade offs between keeping a few seconds of a video stream in RAM vs sourcing it from disk/wherever so that multiple users streams are aligned to avoid having to hit a secondary storage medium. In netflix's case I suspect that requiring fairly large buffers on the endpoint allow them to get away with a much lower QoS metric on any given stream. Put another way, at least the few times I've watched netflix's bandwidth usage, it seems to be bursty. It blasts a few 10's of MB/s of data and then sits idle for a few seconds while the stream plays and then you get another chunk.