8 ms·
Machine learning saves us $1.7M a year on document previews
- wsuen 6y agoHi folks, author here. I am very excited to share this post about how we use machine learning at Dropbox to optimize our document preview generation process. Enjoy! I'll be online for the next hour or so; happy to answer any questions about ML at Dropbox.
- setib 6y agoI work in an innovation ML-oriented lab, and we have a hard time identifying use cases with real added values. So I wondered: who had the initiative to use ML in riviera, the riviera team or the ML team? How do you collabore between the two teams/worlds (production team and data science team)?
- wsuen 6y agoHi setib, great question. The original idea to use heuristics for preview cost reduction came out of a Hack Week project. This led to an initial brainstorm meeting between the ML team and the Previews team about what this might look like as a full-fledged ML product. From the beginning the ML team's focus was on providing measurable impact to our Previews stakeholders. One thing that helped us collaborate effectively was being transparent about the process and unknowns of ML (which are different from the constraints of non-ML software engineering). We openly shared our process and results, including experimental outcomes that did not work as well as planned and that we did not roll into production. We also worked closely with Previews to define rollout, monitoring, and maintenance processes that would reduce ops load on their side and provide clear escalation paths should something unexpected happen. Consistent and clear communication helps build trust. On their side, the Previews team has been an amazing ML partner, and it was a joy to work with them.
- andy99 6y agoI'm curious to know the answer to this question as well. I have done a fair bit of work working with organizations to identify ML use cases. When we looked at it from a business process perspective, honestly it didn't go very well. Trying to find company process specific interventions, especially in the format of building a funnel to priortize which to move forward, rarely surfaces unique or game changing ideas. We usually ended up generating a list of things where either ML played a minimal role, something more simple would have been better, or you'd need AGI. What I've seen work better is a product approach, where ML is incorporated as a feature (rarely but possible the centerpiece) of a full solution for an industry that provides a new way of doing something and the value that comes with it. The caveat is that this is hard and takes up front R&D and product market fit research that any product would. It doesn't happen in a series of workshops with representatives from the business. This Dropbox story is an obvious counterexample, and really looks like the mythical "low hanging fruit" that we always want to identify in ideation workshops. But I'd be careful trying to generalize a process for identifying ML use cases from it.
- Jugurtha 6y ago@andy99, @setib: we're a boutique that helps large organization in different sectors and industries with machine learning. Energy, banking, telcos, retail, transportation, etc. These organizations have different maturity levels and their functions expect different deliverables. The organizations range on the maturity level from "We want to use AI, can you help us?" to "We have an internal machine learning and data science team that's overbooked, can you help?" to "We have an internal team, but you worked on [domain] for one of your project and we'd like your expertise". For the expectations, you can deal with an extremely technical team that tells you: I want something that spits JSON. I'll send your service this payload and I expect this payload. So that's a tiny part. Sometimes, you have to build everything: data acquisition, develop and train models, make a web application for their domain experts with all the bells and whistles, admin, roles, data management, etc. I wrote about some of the problems we hit here[0]. The point is that, finding these problems is an effort that requires a certain skill/process and goodwill from the clients. We worked on a variety of problems. - [0]: https://news.ycombinator.com/item?id=25871632 https://news.ycombinator.com/item?id=25871632
- 6y ago
- wsuen 6y agoI've gotta run, I'll take a look later if other questions come in!
- junippor 6y agoHaven't read the doc. But is it just me or does that seem tiny, considering how large dropbox is?
- heipei 6y agoIt does seem tiny, and my first thought was "how many dollars did they burn to save those $1.7M", but that was one of the first things they evaluated, and both the research phase and operational burden of running the service seem to be relatively small so that the investment definitely paid off. It's great that they're talking real numbers, loved the post in general!
- RcouF1uZ4gsC 6y ago> We used the “percentage-rejected” metric minus the false negatives to ballpark the $1.7 million total annual savings. I think this may too sanguine for the false negatives in that it ignore latency sensitivity. Generally, batch processing (like preview generation during pre-warming) is cheaper than latency sensitive processing (like preview generation when the user is waiting for it). If you don’t take that into account, you can be misled by your cost metrics.
- wsuen 6y agoHi there. Fortunately for the ML team, the Previews team kept a detailed cost breakdown for different preview types - including on the fly generation cost, async generation cost, and cost to serve a preview. Our ballpark accounted for some of the varying costs, though the difference is not significant.
- mehrdada 6y agoThe technical details are interesting but the emphasis on "1.7M savings" screams misdirection of resources, considering the salaries of SWEs/ML engineers and more importantly opportunity cost of them to deploy to an optimization task.
- wsuen 6y agoHi mehrdada. In the article, we discuss how to evaluate tradeoffs of ML projects. One of these tradeoffs is cost of development and deployment vs. cost of not developing a solution. In our particular case, the tradeoff made sense.
- jeffbee 6y agoIt does sound like it would barely have broken even when considering the opportunity cost of the highly-compensated developers who had to write it, which they ignore in the article. It goes against the "rules of thumb" i learned at Google, which suggest that an engineer would break even if they saved [redacted but huge number] CPUs per year, and should only choose problems that promise to save 10x that or better.
- bagels 6y agoWas totally going to point this out as well. It's entirely possible that their ML team cost nearly this much to begin with.
- laluser 6y agoWe're looking at this from a very narrow lens. A few of these wins a year in a small team will end up paying for itself fairly quickly.
- piyh 6y agoLet's say X engineers making 200k a year work on this. This is a 5x return on your money in 5 years if it took 8 people working the full year to complete it. Sounds like a solid business case to me.
- dheera 6y ago
- joosters 6y agoIs the cost saving really measuring the right thing? Instead of comparing against the cost of pre-generating and caching every file preview, shouldn’t they be comparing against he cost of adding enough infrastructure (or just optimising their preview code) to make on-the-fly preview generation acceptably fast?
- wsuen 6y agoHi joosters, thanks for the question. It is always a good practice to ask if we are solving the right problem. The decision to prewarm is ultimately a product decision to give users a better experience across the many surfaces where they encounter file previews. There are limitations to how fast previews can be generated on the fly, even under optimal conditions (unlimited top-of-the-line hardware, max of 1 request at a time). For instance, optimal on-the-fly preview generation for more expensive files (say, a video that is an hour long or a gigapixel image) can add 10s of seconds to tti -- not good from a user perspective! Given this constraint, we wanted to optimize where we spend on the preview generation infrastructure without negatively impacting the user experience. We chose to do this with a combination of heuristics and the ML solution described in the article.
- catmanjan 6y agoCouldn't you just load the first frame of the video, and aren't most image formats optimized for fast thumbnailing? I take your point that there are times when it will be slower, but are you saying that you assumed ML would be faster than on the fly, or did you actually check?
- ddorian43 6y ago> Couldn't you just load the first frame of the video No, often it's black. You have to do scene-detection if you want meaningful previews/thumbnails, can't always go random about it. And the original files are stored in HDD and not SSD/RAM. Source: I don't work there.
- mushufasa 6y agoFYI there's also the off-the-shelf https://filepreviews.io https://filepreviews.io for this
- appleflaxen 6y agocool product, but the blog hasn't seen a post since 2017. is it active?
- mushufasa 6y agoit's a side project of active OSS hero jpadilla https://github.com/jpadilla https://github.com/jpadilla. the service is active but he doesn't do marketing for it. he did talk about it on a podcast a couple months ago https://jpadilla.com/2020/01/10/podcast-djangochat-ep-45/ https://jpadilla.com/2020/01/10/podcast-djangochat-ep-45/
- electricshampo1 6y agoA similar problem previously written about for Google Drive (which files to suggest to the user and possible prefetch) Quick Access: Building a Smart Experience for Google Drive (2017) https://research.google/pubs/pub46184/ https://research.google/pubs/pub46184/ and a recent follow up Improving Recommendation Quality in Google Drive: https://storage.googleapis.com/pub-tools-public-publication-data/pdf/8b3a830b943f7b3e16d51963f9e7aadd51b84960.pdf https://storage.googleapis.com/pub-tools-public-publication-...
- sjg007 6y agoHmm.. how does it compare to a heuristic to say only generate previews for recently uploaded/changed files for active users who may share a lot?
- mdoms 6y agoDropbox are currently laying off 315 employees but I imagine any process improvement that can save 5-10 full time employee equivalents is more than welcome.
- lspears 6y agoThe more interesting problem is learning doc -> image directly. With Dropbox's scale of data, seems feasible for some data types.
- caturopath 6y agoI'm sure that couldn't leak anything important.
- spullara 6y agoThis doesn't seem like enough savings to justify paying a team to maintain it.
- deleted 6y ago[deleted]
- sk5t 6y agoSetting aside the Dropbox-proprietary services, I'd be most interested to see what features/categories you picked from the file metadata, how those were prepped or rotated for training, what training/prediction algorithms you tried and how you picked a winner. Also curious to see the note that observing the "reason" was important here--isn't this a case where a total black box would be sufficient?
- milleramp 6y agoOk not related, I wonder how much paper we could save by showing (by default) a print preview. 1.7m trees? Profit for humanity...
- ollien 6y agoWhat kind of "signals" are being fed into this model? Are we talking like, scroll position relative to a file, and things like that?
- danielscrubs 6y agoI have some questions. How much will the false negatives cost in customer satisfaction? How many FTEs will you need to maintain and improve this? How many more customers did you get from the PR?
- spondyl 6y agoUnfortunately it won't save me from unsubscribing because of the ever increasing upsell tactics and dialogue prompts by the Dropbox application. I just want a simple file storage :(
- twoslide 6y agoI just did the same. I moved to Google, even though I’m no fan. Cost is about 70% less.
- numpad0 6y agoThey won’t do that because simple file storage is always “abused” and will have to be shut down. There always has to be stressers to limit use cases to casual ones.
- skinkestek 6y agoAbused how? By using their paid quota? I know "unlimited" options are abused, but Dropbox isn't unlimited?
- numpad0 6y agoWhatever ways incompatible with business models, not necessarily with hostile/malicious/abusive intent. Maybe "abused" can be replaced with "used for procedurally generated data" or something.
- skinkestek 6y agoIf I rent 2 TB and use 2TB it doesn't matter what I use it for, does it? I would agree if we talked about Backblaze with "unlimited" backup for next-to-nothing, but this is Dropbox which is rather pricey and also limited to a very specific number of (Tera)Bytes. If they cannot deliver on their promises, shame on them.
- numpad0 6y ago