7 ms·
An Uber-like CDN
- akamaka 4y agoWill the server-side code for this be available as open source, so that it can be used by someone who wants to run their own CDN but doesn’t want to be part of the network?
- mranton 4y agoright now no, but we have an idea to opensource everything.
- focusedone 4y agoIgnorant question - How well does this scale? Are peers all necessarily running in major cloud environments? Do you imagine seeing peers in the form of smaller players with beefy home setups?
- mranton 4y ago>How well does this scale? Are peers all necessarily running in major cloud environments? The network scales very well. The network is divided into multiple GeoZones. We can add remove these zones easily. GeoZones are not connected to each other, so there is no single point of failure. Actually the major cloud envs are the most expensive possible variants for Peers. Bare metal servers are much cheaper in our case. >Do you imagine seeing peers in the form of smaller players with beefy home setups? Home internet has a couple of disadvantages: - traffic is not "unlimited". Xfinity doesn't allow you more than 1 Tb in and out traffic per month - home internet is asymmetrical: you could have 1 Gbps in and only 5 Mbps out.
- focusedone 4y agoSo a small, regional ISP with a bunch of unused outbound bandwidth and spare rack space could be a good peer in this scheme, right?
- mranton 4y agocorrect
- amalcon 4y agoSomething doesn't add up here. They're actively telling third parties to lease cloud servers with large amounts of free bandwidth, and allocate them to this purpose. I see four possibilities: - They are paying the third party less than what it costs to rent the server, in which case the third party is losing money and shouldn't be doing this. - They are paying the third party at least what it costs to rent the server, in which case why aren't they just renting the server themselves? They could probably negotiate a bulk discount. - They are trying to hide this activity from the cloud server provider by involving third parties. While this may not be a ToS violation as such, I'd guess that if they're hiding it, it is something the cloud providers will want to crack down on if it becomes significant. - This entire scheme is actually a plan to get access to countries with weird regulatory requirements, e.g. foreign companies are not usually allowed to own servers in China. This one's pretty legit actually, but it's a really complex version of the scheme. I think at least one of those must be true, and none of them would be a good sign for the future. Edit: I guess "paying the third party less than their costs" and/or "not making a profit themselves" might be what they mean by "Uber-like", so it would just be an elaborate joke?
- boh 4y agoI guess it really is Uber-like.
- asciii 4y ago> Something doesn't add up here. Agreed, I'm mildly annoyed that farba.js repo has not been updated since 2021 (no further info) + grammar issues [1] [1] https://github.com/FarbaCDN/farba-public/blob/master/src/farba.js#L30 https://github.com/FarbaCDN/farba-public/blob/master/src/far...
- mranton 4y ago>Something doesn't add up here. They're actively telling third parties to lease cloud servers with large amounts of free bandwidth: Bare-metal servers are much better is our case. Please note, bandwidth is not free. You can follow the links and see the actual price for it. If LeaseWeb tells you: 1 Gbps dedicated/unlimited per $143/month, it means you can use it 24x7. >- They are paying the third party less than what it costs to rent the server we don't know prices in all datacenters all over the world. That's why the _Peer_ sets the price for his service.
- amadeuspagel 4y agoWhat problem does this solve? Cloudflare is already free.
- yamtaddle 4y agoNot if you hit large enough scale, or serve too much non-"Web" content.
- _joel 4y agoNot for serious volumes, there's also a load more feaatures only paid and enterprise accounts get.
- mranton 4y agoI'm collecting use cases when Cloudflare stopped service for a "non-web content" rule violation: https://news.ycombinator.com/item?id=34639212 https://news.ycombinator.com/item?id=34639212
- jsnell 4y agoThere's a bunch of really obvious objections to this idea around performance, privacy, economics, UX, and reliability. Impressively, the post did not manage to address any of them up front. 1. Performance. In terms of latency, this is adding an extra round-trip to each page request (to get the metadata). It's also likely going to load different resources from different servers (hostname, IP) preventing connection reuse. I guess they can't even start fetching data until the full page has loaded? And it feels like this scheme would make it much harder to reliably cache resources, since the browser will cache by URL while they'd struggle to make sure the URLs are stable across calls. 2. Privacy. Their threat model only addresses malicious peers changing content, not malicious peers trying to track users. 3. Economics. If running peer servers really was profitable, the company would do it themselves rather than outsource it. Given the "where can you get a server" section, it's not even that they're expecting this to just be running on spare capacity. The problem with the setup is that incentives of the peers are badly misaligned with the network, the clients, and the end-user. A peer wants to maximize traffic; the clients and end-users want to minimize it. So a peer would be incentivized to set cache-control headers to prevent caching, to increase traffic. Likewise a peer would be incentivized to only keep copies of the most accessed resources, to be able to serve as high a proportion of the traffic as possible with a given disk budget, while all the other stakeholders would like each file to be available just N+1 times in each region. 4. UX. URLs will be really crappy compared to a proper CDN. Right-click copy a link to an image and send it to somebody? It'll initially be to some super-dodgy URL, and stop working in a day or two as the CDN node starts caching different content. And AMP showed that people do actually care about URLs. It's pretty hard to take this seriously.
- deleted 4y ago[deleted]
- mranton 4y ago>1. Performance. In terms of latency, this is adding an extra round-trip to each page request (to get the metadata). It's also likely going to load different resources from different servers (hostname, IP) preventing connection reuse. I guess they can't even start fetching data until the full page has loaded? fair point. We use same technique as lazy-load (you have to wait the full page has loaded). >3. Economics. If running peer servers really was profitable Uber doesn't drive taxis himself, airBnB doesn't own property. This is how Uber-economy works. We believe it works in the CDN industry as well. We just don't want to be another CDN provider.
- eastdakota 4y agoTravis built RedSwoosh (https://en.m.wikipedia.org/wiki/Red_Swoosh https://en.m.wikipedia.org/wiki/Red_Swoosh) and then went on to build Uber. Now someone saw Uber and claims it inspired them to rebuild RedSwoosh. ¯\_(ツ)_/¯
- dedalus 4y agoQuite an honor to see Matthew Prince (CEO of Cloudflare) chime here. Akamai acquired RedSwoosh precisely on the same economics promised and if this were economical/viable they would already have deployed it, no? I did hack with Travis on RedSwoosh code base optimizing his TCP Zipper algorithm which was pretty cool BTW
- mranton 4y ago>The Red Swoosh peercasting tool is a _browser_extension_ that caches data, reflecting and sharing files delivered through the "Swoosh network" I don't know how it works in details, but it looks like you need to install a browser plugin to enable the p2p network here. Our solution works transparently for the end user.
- dedalus 4y agoThe extension simply rewrites the URLs and there's a way to "swoosh" your links which works transparent for the end user: https://web.archive.org/web/20050111015750/http://redswoosh.com/product.php https://web.archive.org/web/20050111015750/http://redswoosh....
- uncertainrhymes 4y agoThere are several red flags here, not the least of which is that a CDN isn't just about bandwidth, it is about disk. If a cache node doesn't have the object, it has to go to origin to get it. That has a cost at origin, which most CDN customers want to minimize. Cache management (eviction, invalidation, etc) is done where? By the look of that diagram, you are actually doubling your origin traffic just to push to the cache after serving the original client request. Perhaps subsequent requests could get served from their 'peer' but the hit rate will be abysmal. Maybe there is more to it than they describe here, but I don't see how this can work. Are they going to rewrite the object in a way that doesn't require the host certificate?
- mranton 4y ago>a CDN isn't just about bandwidth, it is about _disk_ The disk performance - is the weakest part of the Peer. It is a part you can't guarantee any QoS on a cheap server. The idea is to keep popular cache on Peers in memory. Everything else we will host from a limited number of regular servers (belong to Farba), with high-performance disk subsystem. >Cache management (eviction, invalidation, etc) is done where? yes. You can invalidate either a single URL or a folder with a wildcard >you are actually doubling your origin traffic just to push to the cache after serving the original client request this is something we can proxy and do only one request per file. Current implementation was done for simplicity and reliability. >but I don't see how this can work. Are they going to rewrite the object in a way that doesn't require the host certificate? you can easily check this up. Please open the network console on our demo-page and see how it works: - we have two different certificates (one for the main site, for for the Peer) - everything works flawlessly
- uncertainrhymes 4y agoYour example servers have 16-32 GB of memory, which is basically nothing. If you ever ramp to real world traffic the peer will be constantly evicting objects (if you use a normal LRU). Even if you backstop that with a mid-tier belonging to Farba to protect the origin, what kind of cache hit rate are you expecting at the peer?
- yamrzou 4y agoRelated: P2P CDN: https://hn.algolia.com/?query=p2p%20cdn https://hn.algolia.com/?query=p2p%20cdn, https://hn.algolia.com/?query=peer%20to%20peer%20cdn https://hn.algolia.com/?query=peer%20to%20peer%20cdn
- mranton 4y agokind of. They are WebRTC-based CDN.
- todd3834 4y agoHow does this compare to IPFS? I was playing around with IPFS and it seems like a lower risk higher node version of this. Performance wasn’t half bad either. I guess it depends on which node you end up downloading from and whether there was a nearby cache hit. P
- mranton 4y agoIPFS is very slow. It's like a cold storage. Not good for web-cache at all. Though Cloudlare provides gateways for IPFS, and again, we have a CDN layer here. It would be nice to see IPFS performance numbers to compare.
- danielrhodes 4y agoThe value of this largely depends on what you're optimizing for. If your constraint is you need a lot of cheap bandwidth but don't care that much about latency, there are many places to find it (find a cheap host, set up Varnish, you're good to go). However, a lot of people use CDNs because they are fast. They want the latency to be ultra low and they want it to be reliable. This is where the author's solution falls short: CDNs can get premium quality bandwidth because they have more scale than their customers. Sure you can find servers from various hosts in lots of different data centers, but you might not get the latency or quality of bandwidth you were looking for. This can be critical for some applications like streaming or serving images. If your constraint is speed, you almost have to go with a CDN (even if you are huge). If it's volume, you are in a better position. If your volume is small, you might as well go with something off the shelf like Cloudfront or Cloudflare. If you are ultra huge like YouTube or Netflix, then you have to peer directly with ISPs.
- datadeft 4y ago>> At Farba, we are building a zero-cost CDN infrastructure where all edge servers belong to anonymous Peers. Zero-cost CDN for 2 USD/TB! Sounds like a great deal to me. On a more serious note, is price the problem with other CDNs?
- emj 4y agoCoral CDN 2004-2015 was ment to be p2p, why that can be a problem is mentioned in comments here and in this paper from Princeton. https://www.cs.princeton.edu/~mfreed/docs/coral-nsdi10.pdf https://www.cs.princeton.edu/~mfreed/docs/coral-nsdi10.pdf
- nektro 4y agoIf folks have the spare bandwidth they should run Tor nodes
- jedberg 4y agoThis sounds a lot like CoralCDN. You should look them up, they had some good papers. We actually used CoralCDN on reddit for a while, where we'd send users to the CoralCDN link of the webpage they were trying to get to instead of the actual page. It sadly didn't work very well. Here is one of their papers: https://www.usenix.org/legacy/event/nsdi10/tech/full_papers/freedman.pdf https://www.usenix.org/legacy/event/nsdi10/tech/full_papers/...
- nso 4y agoOne of the problems I have with the VP of the product discussed is that while I trust CloudFlare, Akami or any of the big players, to run a tight network -- with high and STABLE performance -- I can't say the same about a random vendor (uber driver?) in this network. Was the issues you were seeing in this general area?
- deleted 4y ago[deleted]
- frankreyes 4y agoReverse Tor?
- jefftk 4y agoThis says it's a system for delivering content "efficiently", but it doesn't look like this manages that. Normally loading a resource is a round trip to your server, which might be a bit inefficient because your server could be on the other side of the world. But what they replace it with is: 1. Round trip to Farba to get metadata. 2. Round trip to your "driver" to get the resource. 3. Computing the hash of the resource to ensure the "driver" didn't cheat. It could still be worth it as a way of saving money on bandwidth, but I'd be quite surprised if you got the speed benefits of a CDN.
- mranton 4y agoeverything works in our sandbox. We need a couple of customers to collect real data. I'll definitely going to post again with these numbers.
- bastawhiz 4y agoWhat stops a peer from, say, injecting malicious JavaScript into pages? Or going rogue and serving porn for every request for a JPEG? What if they only do it for every 1000th request? You could send a hash of the file to the client from the coordination server and check it. If the hash isn't valid, you can re-request the file from another peer. But I don't think you can trust the client to report abuse: a malicious client could report innocent peers (e.g., a vindictive Chrome extension). Or, the site owner could manipulate the coordination JavaScript to report the downloaded file as invalid/corrupt/tampered with, and then just not re-request the file (and there's nothing you can do to stop them). So in that case, the site owner is managing to not pay for the service (assuming you don't bill for downloads that were reported as invalid). But on top of all of that, this is _slow_. It's slow on my (quite fast) internet. If you're using this, it's because you're trying to get cheap bandwidth, not because you value performance. The whole point of a proper CDN is to improve performance by spreading out load and moving the edge server physically closer to the user. This might be putting the edge server physically closer to the user, but the mechanism to do that seems to take almost as long as the request/download of the asset itself.
- tantalor 4y ago> Since the requested file comes from an untrusted source, the next step is to verify the integrity of the file. The script calculates the hash sum and compares it to the digest from the JSON reply. If the sums match, the script displays the downloaded file to the User.
- bastawhiz 4y agoAnd if they don't? Or if the client says they don't?
- mranton 4y ago>Or, the site owner could manipulate the coordination JavaScript to report the downloaded file as invalid/corrupt/tampered with, and then just not re-request the file (and there's nothing you can do to stop them). So in that case, the site owner is managing to not pay for the service (assuming you don't bill for downloads that were reported as invalid). This fraud is possible. We collect logs from three sources: - the Balancer (belong to Farba) - the Peer - the User We can analyze every single case of suspicious behavior.
- PaulHoule 4y agoIt’s latency that matters and you can’t get back ms you lost because you are using some server that somebody located in a second rate location.
- gabereiser 4y agoThe math doesn’t work. If I sign up for a server and host a node and charge $1 more than I’m paying, it would just make more sense to do that yourself and save yourself the dollar.
- solidsnack9000 4y agoTo everyone saying, this business does not make sense because they could just buy the servers and serve it themselves and make all the money, please consider that the same argument applies to any intermediated business whatever: to Uber; to wholesalers; to franchise businesses like McDonald's. The reason businesses like this work in general is because they break a big, complicated business into smaller businesses where there is specialization. There may not be be meaningful potential for specialization here; but that is the argument to have. For example, it may be that Farba can be a leaner business if they don't have to think about how to get the best pricing in dozens of different hosting markets.
- taco_philips 4y ago[dead]
- MarkSweep 4y agoIt seems a little odd to use GitHub as the CDN for your JavaScript library. Like what if this got really popular, what’s to stop GutHub for shutting it down for abusing their hosting? If I were a customer for a CDN, I would feel more comfortable if they owned their own domain and could control how (and if!) their JavaScript is available to customers. Anyways, other than that it’s a fun concept. I like the P2P nature of it.
- iudqnolq 4y agoThat's insane. This JS has to be loaded before they can start any of the actual requests. It should be on the most performant host possible.
- louison11 4y agoThis isn’t good for SEO, which is highly dependent on speed. The CDN data is loaded via JavaScript, meaning crawlers might not load it. I don’t know what this is trying to solve, CloudFlare pretty much already fixes the egress charges problem.
- sgammon 4y agoas a web user, no thanks, I do not trust your peers or really you even. I trust cloudflare because they have earned that trust over decades. as a web developer, no thanks, I do not want to trust my app’s speed to your profit incentive or margin.