12 ms·
How Discord Resizes 150M Images Every Day with Go and C++
- throwthisawayt 9y agoDid it seem to anyone else that sticking to Python would have been way easier? It didn’t seem like any of the performance gains were through Golang.
- deepnotderp 9y agoYeah, either a good Python JIT or Cython would have been fine honestly. I never understood the obsession with "python is slow" when you can recover almost all of the performance with a good JIT or Cython (in many/most cases).
- victor106 9y agoWhat are some good Python JIT's that are worth trying out?
- dr_zoidberg 9y agoPyPy.
- victor106 9y agoThanks.
- Drdrdrq 9y agoYes. Or simply profiling the app and optimizing sore spots would have helped too. It seems to me there was no real reason to move from Python to Go, apart from preference.
- stmw 9y agoFrom reading this, seems HTTP handling speed was important to them? which Go is probably better for. Also, interfacing Python to C/C++ is pretty unpleasant.
- jononor 9y agoIn Python they already had an extremely fast library with bindings available.
- smaili 9y agoI believe this little piece answers your question: > We likely could have addressed this behavior in Image Proxy, but we had been experimenting with using more Go, and it seemed like a good place to try Go out. At the heart of if, they were looking for opportunities to use more Go in their stack and they deemed this situation as a fit.
- prophesi 9y agoAnd they ended up open-sourcing the library they built, so it's a win on all sides.
- fleitz 9y agoThe age old solution in search of a problem.
- Karrot_Kream 9y agoI think that's a bit reductionist, no? There are many reasons they may have been searching for moving to Go. Off the top of my head I can think of: 1. Static typing increasing confidence and velocity 2. Better developer-facing tooling increasing velocity 3. More employees knowledgeable about Go than Python 4. More enthusiasm (and therefore faster velocity) around Go development. The blog post was about the engineering challenges they faced and how they solved them and I think it was a great write-up in that regard. The post wasn't about why they switched this service from Python to Go.
- fleitz 9y agoIt might be, then again I see a lot of wheel reinvention in tech / NIH syndrome. I'm the kind of hacker who if a service runs out of memory every 2 hours, writes a crontab to restart it every hour after X random minutes so they don't all restart at the same time. It gets a lot of eye rolls from the other engineers searching for perfection, but it tends to produce services quickly that are highly reliable. And look now the engineers who like chaos monkey don't even have to set that up. It's built in. It looks like most of the savings were in switching from pillow to opencv, something that thumbor already does. https://github.com/thumbor/opencv-engine https://github.com/thumbor/opencv-engine
- deleted 9y ago[deleted]
- harikb 9y agoI understand this is a personal preference, but having spent a good amount time with both Python and Go, FWIW I would also choose Go if I were solving the same problem.
- detaro 9y agoI don't think the article gives us the data to know this. Where did the latency spikes in the original implementation come from? Would fixing them have required a complete rewrite of the Python parts anyways?
- gourou 9y agoLink to the resulting open-source project: https://github.com/discordapp/lilliput https://github.com/discordapp/lilliput
- caltrops 9y agoI’d be very worried about a security issue with the unsafe C++ code. You really have to run this kind of complex parsing in a disposable containerized environment to do it safely. Or do everything carefully and in a memory safe language.
- bri3d 9y agoI'm not sure why this is being downvoted - image processing is one of the most dangerous parts of a common consumer-facing web software stack. By and large this is because image container formats are poorly documented, overly broad, and rely on a lot of tricky binary parsing that's easy to mess up in an unsafe programming language. It's also one of the most obvious ingress points for untrusted binary data uploaded by an end-user, which is always going to be dangerous. See the persistent, years-long trend where mobile devices and game consoles get exploited via some combination of libtiff and libpng.
- pmelendez 9y agoTrue (and I didn't downvote by the way), but a "memory safe" language might not be as helpful as people might think. Most of memory managed languages still rely on native libraries to perform image processing, if at the end you are using libpng and there is an exploit on it, it doesn't matter if you are using python or C++, both code base would have the same exploit if it is not explicitly mitigated in the logic.
- deleted 9y ago[deleted]
- GuB-42 9y agoThe downvote is probably because the comment implied that the issue is that the image processing is done in "unsafe" C++ and that another language should have been used. However, there isn't much choice. Performance is very important in image processing, so much that many libraries contain hand-written assembly. In the article, it says that 90% of processing power is dedicated to it. Using a safer language in a safe way could completely kill performance and significantly increase the costs.
- JepZ 9y agoAnybody knows how well libvips https://github.com/DAddYE/vips https://github.com/DAddYE/vips compares to liliput performance wise?
- b1naryth1ef 9y agovips (Go binding) is included in the benchmarks mentioned in the post, but at the time of running them (~10 months ago) vips pulled 51482954 ns/op on a 1024x1024 test image, where as pillow-simd managed 3324135.3035 ns/op.
- CapacitorSet 9y agoFor ease of reading, that's respectively 51 ms and 3 ms.
- deleted 9y ago[deleted]
- JepZ 9y agoThanks :-) Looks like I didn't scroll properly when I looked at that file. My bad :-/
- Xeoncross 9y agoWhy don't more companies resize images client-side first using <canvas> and then save the server some work by only asking it to verify the result by - resizing to the same size - removing metadata This results in much faster transfer (10x less bandwidth used often for mobile uploads) and reduces server load by "farming" out the work to the clients. https://developer.mozilla.org/en-US/docs/Web/API/CanvasRenderingContext2D/drawImage https://developer.mozilla.org/en-US/docs/Web/API/CanvasRende... # Edit: On Keeping Full Resolution Images Some people mention having original highest-resolution images are important. I don't think that is true for most applications. Most apps don't need hi-resolution history as much as current, live engagement so older photos being smaller isn't a big deal. As technology moves on you simply start allowing higher-res uploads. Youtube, facebook, and others have done this fine as the older stuff is replaced with the new/current/now() content. In fact, even our highest resolution images are still low-quality for the future. Pick a good max size for your site (4k?) and resize everything down to that. In a year, bump it up to 6k, then 10k, etc... Keeping costs low has it's benefits, especially for us startups. Now if you have massive collateral, then knock yourself out.
- CamTin 9y agoBecause then you don't get to build out fun infrastructure like this and write it up in your company blog.
- Xeoncross 9y agoSure you do! Don't you know that adding "Free" to a title increases the ROI by 40%? "How Discord Resizes 150M Images Every Day for Free"
- fleitz 9y agoAgreed. I wish the posts contained a "it cost X developer hours to recreate thumbor or $$$ total, and we saved Y dollars per month" meaning in approximately 15 years we'll have broken even on this investment. Oh yeah and we don't even do intelligent resizing like thumbor does.
- 9y ago
- sgk284 9y agoThere is already an (unofficial Google) image proxy written in Go that is quite fast, does caching (local or backed by S3/GCS), and does other nice things like smart cropping: https://github.com/willnorris/imageproxy https://github.com/willnorris/imageproxy Seemed like a lot of unnecessary work for them to reimplement a service from scratch without gaining any major perf benefits over their existing one and without leaning on an existing well-known and well-built foundation.
- bpicolo 9y agohttps://github.com/thoas/picfit https://github.com/thoas/picfit is another golang lib for this, and it's pretty mature at this point. The one thing these don't support though is smarter cropping that takes into account image contents, which takes enough cpu power to require preprocessing
- brian-armstrong 9y agoAuthor of the blog post here - it looks like what you linked does its image resizing in pure Go. In our testing we found these libraries are significantly slower than the C++ resize libraries. I would guess we would need at least 10x as many instances if we used that resizer, though probably a lot more
- deleted 9y ago[deleted]
- devwastaken 9y agoHow is the security? Any sort of image processing is a potential exploitation point. I see it says it uses the 'mature' libjpeg-turbo and libpng libraries,along with giflib for .gifs, but even with full trust of those, the C code, patches, and changes ontop could be more exploitation points. You can look through Imagemagick alone to see all the fun things possible when seemingly basic processing turns into exploits. https://www.cvedetails.com/vulnerability-list/vendor_id-1749/Imagemagick.html https://www.cvedetails.com/vulnerability-list/vendor_id-1749...
- Buttons840 9y agoWow really? Is there room for another image processing library? Is ImageMagic poorly written or is image manipulation inherently risky?
- abiox 9y ago> Is there room for another (...) seems to me that there is no limit to available room. well, i suppose we're capped by the collective capacity of local storage and storage service providers.
- bri3d 9y agoImageMagick is notoriously questionable. It was originally written, I believe, as a local command-line tool for users to work with their own images, so security and untrusted input were not primary concerns. Additionally, image manipulation is inherently challenging - not even due to the actual manipulation of image pixel data, but due to the proliferation of complex image container formats which require binary data manipulation and byte copying in performance-critical code. This is a minefield for secure programming practices because it puts at direct odds performance and sanity checking, as well as encouraging pointer and memory arithmetic and unsafe access.
- tedunangst 9y agoImageMagick is a particularly poor choice because it will try parsing a thousand formats your users will never upload. That's a lot of code to leave exposed to the internet.
- kylehotchkiss 9y agoI wish Cloudfront supported resize parameters so we wouldn't have to keep buildings these or paying a lot for Imgix.
- fleitz 9y agoHow much would you pay for an image resizing service? I'd been thinking for a while of putting a fleet of autoscaled thumbor boxes behind cloudfront and making a billing API for it.
- deleted 9y ago[deleted]
- kylehotchkiss 9y agoImgix's $10 minimum is so much for a personal site with maybe 500 uniques a month. If you're going for a service like that, think of people like me who host on s3/cloudfront for $.20/month. But let people scale up to millions of pageviews a month. Don't need anything fancy. Just w=? h=? would be great, developers can handle the DPI stuff with sourceset tags.
- manigandham 9y agoCloudinary is free. https://cloudinary.com/pricing https://cloudinary.com/pricing
- abeach222 9y agoYou can use lambda edge functions for this. They recently announced support for query string parameters. https://aws.amazon.com/about-aws/whats-new/2017/10/lambda-at-edge-now-provides-access-to-query-string-parameters-country-and-device-type-headers/ https://aws.amazon.com/about-aws/whats-new/2017/10/lambda-at... I have built an image resizing service around this with go and libvips. With go libvips, s3gof3r, you can load s3 images directly into a buffer, pass to libvips, and serve without writing to disk. Basically, you can use edge functions with your origin as the above go service.
- linkmotif 9y ago> Today, Media Proxy operates with a median per-image resize of 25ms and a median total response latency of 85ms. It resizes more than 150 million images every day. Media Proxy runs on an autoscaled GCE group of n1-standard-16 host type, peaking at 12 instances on a typical day. Awesome! <3
- deleted 9y ago[deleted]
- manigandham 9y agoNice, but why? https://cloudinary.com https://cloudinary.com, https://www.imgix.com https://www.imgix.com, or https://www.filestack.com https://www.filestack.com already exist and are well worth it for 99% of apps. Even at scale, it really doesn't cost that much to have someone else do it. You can use a thin proxy through your existing CDN if you want to save on their bandwidth fees. Also http://thumbor.org http://thumbor.org and https://imageresizing.net https://imageresizing.net if you want a library to host yourself which are already very fast and well tested. Put them in a docker container on a kubernetes cluster and it's all done in an hour.
- zitterbewegung 9y agoMaybe it’s because that they don’t want a dependency on a external service that could go down ?
- StreamBright 9y agoSo you have an internal dependency that could go down?
- Volt 9y agoIt's better when you can do something about it.
- StreamBright 9y agoIt depends.
- manigandham 9y agoIt's images... seems like a very low risk situation, especially when they are served from a CDN.
- deleted 9y ago[deleted]
- 0xbear 9y agoThat’s 1700 images per second. Doable on one (beefy) box. 3 to account for the diurnal cycle. Am I supposed to be impressed?
- mbrumlow 9y agoI don't get why you are being down voted. This is almost exactly what I thought. It's just not that much data given the state is computer hardware. Where I work we have single nodes processing near that much data a hour -- these are beefy systems though.
- 0xbear 9y agoPeople have just drunk so much “cheap commodity hardware” kool aid by now, they don’t realize there are cheaper and easier ways of doing things now, assuming you have devs who can code and tune for performance. Same with “big data”. Most people have sub-1T datasets. You simply don’t need Spark or anything custom for that.
- brian-armstrong 9y agoCan you link to which resize library you're using? We'd love to see a 90% further reduction in instances
- mbrumlow 9y agoSorry to be confusing, I am not resizing images. Just working with data sets as large as what I image 150M images would be. The software I am working on takes point and time backups of computers and uploads them to "the cloud", I mean servers in a data center. There they can be virtualized with a click of a button in mass or one at a time, and near instantly. This involves transfering, encrypting, compression and creating checksum of terabytes of data a hour (per node). While not exactly resizing images, I would image the computational power was on par with the service described. The entire system has about 4 PB or 8 PB in it right now, as backups are pruned (based on what people will pay for storage). My software has a ton of space to grow and become better, but I think a better story would have been how discord handles 150M images a hour. If anything bandwidth acquiring the source image would be what I would consider the largest problem, not the CPU time to resize. In fact as long as your resize code slightly faster than the download then streaming it in and out would put your bottleneck entirely on bandwidth. I will also note I am not a fan of libraries :p but that is not what this is about. EDIT: Also kudos to you, somebody criticized your post and you had the best response one could have. Inquiring minds are awesome.
- Const-me 9y agoI wonder why people implement such things on CPU? PCI express is ~100 gbit/sec, much faster than any network interface. Internally, a GPU can resize these images by an order of magnitude faster than that, see the fillrate columns in the GPU spec.
- malikNF 9y agoMost probably its because of the time it takes to push the image on the the GPU and then back to the CPU.
- acdha 9y agoThis isn't just resampling an image: decoding a variety of image (and even video) formats, decompressing the selected frame, performing the actual resize, and then compressing the result. If the resample doesn't save more than the setup overhead, it'd be an immediate loss. Even if it does, there's an engineering cost since you now need to make sure that all of your servers have GPUs available, your chosen implementation code supports all of them with acceptable quality and error handling, etc. Since the GPU hardware has become commonplace, there's definitely a lot more attention on using it in the server space and I think it'll become common in the next few years but that has a migration cost for early adopters since you're hitting less mature projects for critical functions. Internet-facing image processing has a bunch of tedious but important work handling format variations and errors (it'll be reported as a bug in your software if the image opens in a browser and/or photoshop), making sure that you handle gamma/colorspace consistently, etc. If you're trying to get production-ready server out the door, it's really tempting not to deal with any of that once you hit the point where it's fast enough that engineering time costs more than the server savings.
- Const-me 9y ago> This isn't just resampling an image GPUs can do that, too: http://fastcompression.com/products/jpeg/cuda-jpeg.htm http://fastcompression.com/products/jpeg/cuda-jpeg.htm > you now need to make sure that all of your servers have GPUs available OP is running on google’s cloud: “n1-standard-16 host type, peaking at 12 instances on a typical day.” That instance costs $0.76/hour. Adding NVIDIA Tesla K80 is $0.7 extra. > it's really tempting not to deal with any of that Yeah, that’s understandable. But the original article dealt with a lot of strange technologies to get the performance they want. And ended up doing much slower, performance wise, than what’s possible with a GPU.
- ymse 9y agoThis post reminded me of a very old article from Yahoo/Tumblr explaining how they were (ab)using Ceph to generate thumbnails on the fly as pictures were uploaded using the Ceph OSD plugin interface. Unfortunately the post seems to have disappeared from the internet (it was probably around 6 years ago), so here are some other teasers: https://yahooeng.tumblr.com/post/116391291701/yahoo-cloud-object-store-object-storage-at https://yahooeng.tumblr.com/post/116391291701/yahoo-cloud-ob... https://ceph.com/geen-categorie/dynamic-object-interfaces-with-lua/ https://ceph.com/geen-categorie/dynamic-object-interfaces-wi... Disclaimer: not affiliated with Ceph apart from being a happy sysadmin.
- noahdesu 9y agoHere is a link to a talk I gave last month describing how to use Lua to generate thumbnails remotely in the Ceph/RADOS OSD servers. Talk is from Lua workshop 2017. Relevant content begins at 15m40s. https://youtu.be/bGQc-PpJAyk?t=15m40s https://youtu.be/bGQc-PpJAyk?t=15m40s
- tuananh 9y agois there any open source project img proxy that can do this? eg: instead of this http://localhost:8080/https://octodex.github.com/images/codercat.jpg http://localhost:8080/https://octodex.github.com/images/code... we can create alias like octo and url will become this http://localhost:8080/octo/images/codercat.jpg http://localhost:8080/octo/images/codercat.jpg