4 ms·
Thanks for satisfying my curiosity! Also, congrats on your success! > Yes, my user-facing servers are proxying the files to the users. I've never operated a s
by themulticaster 5y ago
Thanks for satisfying my curiosity! Also, congrats on your success!
> Yes, my user-facing servers are proxying the files to the users.
I've never operated a service as large as yours, so take my question with a grain of salt:
I'm wondering whether it would make sense to split off the actual file front-end servers from the user-facing servers (going for a redirect approach instead of proxying), since the requirements for serving the UI (low latency, low bandwidth) are so different from the file serving requirements (high bandwidth, but latency is not an issue). In theory, the traffic load from the files could negatively impact the UI latency leading to perceived sluggishness of the website. But perhaps that's not an issue in practice?
Since you mentioned elsewhere that you wanted to move to content delivery: What kind of content delivery do you have in mind? At the moment I can only think of either classic CDNs (but that's a few order of magnitudes larger) or ads (but that's an entirely different area).
- Fornax96 5y agoProxying the files has a number of benefits. - The first is that I can have all my API endpoints under one domain. This simplifies downloading as you don't need to make a separate request to figure out where the file is stored. - The storage servers that Hetzner sells only have 1 Gbps bandwidth. That runs out very quickly when a file goes viral. The 10 Gbps caching servers do a lot of heavy lifting here, this makes sure the disks in the storage nodes last longer. - I can also decide to switch to a different storage system on my storage nodes when I want. I have been considering to deploy reed-solomon encoding for a while. That would make it impossible to link directly to a single storage server as a single file would also be distributed. - Sending out this much data uses a lot of RAM for TCP send buffers. Installing this RAM on a single content delivery node is cheaper than installing it on every storage server. To prevent the bandwidth load from affecting the UI speed I have a rate limiter on the download API which slows down when the uplink reaches 95% capacity. This way there is always some bandwidth left for the HTML and database communications. With regards to content delivery: I want to use pixeldrain to serve static files. Nothing like the fancy site-wrapping tech that cloudflare uses. The idea is that users can have a file tree on pixeldrain somewhat like dropbox. They can copy the direct download link to that file and use it to embed videos, audio and pictures in their own websites. Because this is a lot simpler than other CDN services I can offer it at a very competitive price.
- ddorian43 5y ago> That would make it impossible to link directly to a single storage server as a single file would also be distributed. Check out https://github.com/chrislusf/seaweedfs/ https://github.com/chrislusf/seaweedfs/ implementation of reed solomon. Small files can still be served from 1 server. It's also efficient for small files, which a image store requires.
- Fornax96 5y agoI have been looking at seaweedfs for a while. I suffer from a pretty bad case of NIH Syndrome, so my gut feeling says I should implement it myself. But seaweed is such a good fit that I should probably just try it. I especially want to know how well it deals with adding / removing servers, data loss and network instability. The beautiful thing with writing my own solution is that it takes at most 10 minutes to diagnose and fix a problem when it occurs.
- ddorian43 5y agoYou can contribute to it and learn it along the way. Pretty flexible.
- chrislusf 5y agoI have the same syndrome. But I feel seaweedfs should be easy enough to be picked up. If you have any questions, just let me know.
- Fornax96 5y agoHey, I was trying out Seaweed yesterday. I ran into an issue with HTTPS, it seems to be not very well supported at the moment. I managed to get it running mostly secure, but replication didn't work because the volumes were calling eachother with HTTP instead of HTTPS. The HTTP request is hardcoded here: https://github.com/chrislusf/seaweedfs/blob/43fd11278ef811856551904b531bcc91821d0c9f/weed/topology/store_replicate.go#L56 https://github.com/chrislusf/seaweedfs/blob/43fd11278ef81185... and probably in a lot of other places. I even tried setting the address of the volume to https in the startup script, but then it makes a request to http://https://volume1.example.com http://https://volume1.example.com and it still fails. I also noticed that the master API is still available over HTTP even when HTTPS is enabled. I can make an issue for these things if you want. These issues are currently the only showstopper for me. I need to have every endpoint on TLS with peer verification enabled. If you can get it fixed I will gladly continue testing seaweedfs and support you on Patreon :-)