3 ms·
The machine learning team is responsible for all search and discovery features that users engage with when searching our inventory, some of the highest traffic,
by mlthoughts2018 6y ago
The machine learning team is responsible for all search and discovery features that users engage with when searching our inventory, some of the highest traffic, uptime demanding services I’ve seen anywhere in my 15 year career. We also run all of the image and language processing that occurs in real time in the image editor post upload, along with a HA queue system that aggregates requests to route them to shared GPUs on the backend. In our case, many machine learning systems are on the critical path for customers, and we have a significant platform team within machine learning specifically to ensure we can support this, meet uptime demands, design auto-scaling solutions in our on-prem data centers, execute failover and disaster recovery (and drills to simulate these) and more. And so does every other team.
- throwanem 6y agoOkay, that makes sense. Why do you assume, or at least give the very strong impression of assuming, that no one else does anything like this? Or even just that Etsy doesn't?
- mlthoughts2018 6y agoBecause they just wrote a big article about how they created a culture, developed special vocab (“slush”), and intentionally plan on a multi-year basis all around a code freeze that is not just a day, not just a weekend, but lasts for weeks and affects all systems (which either means no systems are in a state of proving their resiliency, or leadership doesn’t trust engineers, or both), and they are proud of this. That is a whole mess of giant neon red flags that something is wrong here regarding engineering for resilience and trusting your deployment system to catch errors. Additionally, for a freeze lasting weeks, unfreezing and resuming merges from some backlog of frozen changes would likely introduce far greater risk to revenue or to customers than managing a trustworthy deployment workflow during high traffic events. I would even question the fundamental premise that this is safer or beneficial to customers at all. It seems much, much more likely to be a “cover your ass” method to prevent any possible blowback from outages due to unremediated tech debt that the weak leadership is not capable of protecting as a worthwhile priority for quarterly goals.