4 ms·
So not a technical problem after all.
by asynchronous 3y ago
So not a technical problem after all.
- ggm 3y agoIf technical goes to scaling and effects related to scaling, maybe they just decided to let the people with >5m inodes incur the technical penalty, which is almost always speed? it was always a cost issue, in some sense. If they designed the entire FS model behind drive to have inode limits per entity, then there's a small technology issue bumping that number maybe. Quota systems can be quite painful, in days gone past the Robert Elz Quota model on UFS demanded reboots for some changes, which at scale in a google sized entity is a rolling surface of state change.
- vlovich123 3y agoTheir explanation of the 5M limit is because of inode limitations when they materialize the content via the companion client?
- ggm 3y ago5m is an odd number. 24 bit numberfield would get you 16m uniques. I can't see how they needed to steal bits to get to 23 or 22. So there must be some other factor at play here apart from 'we used 3 bytes in the gRPC model' Perhaps the Windows exposed API assumed somewhere around 5m objects, and they care about the compatibility greatly.
- perbu 3y agoI don't think the number is based on any technical limitation. It could be a number based on usage. Say 99.9% of the users use less than 5M files. So my pruning 0.1% of the user base they are getting rid of half the number of files, saving X millions of dollars in cost. Nobody wants to cater to the difficult users that are doing stuff like this.
- hnfong 3y agoThat's interesting, because if your argument is valid, they wouldn't have to roll back the change.
- ggm 3y agoif filescount() > 5m then.. had to be removed. There is no quick release in a google scale monrepo and QA process.
- Tommstein 3y agoAfter that pruning, you now have a new 0.1% of users using more files than the other 99.9%. Might as well prune that 0.1% of file hogs too. Rinse and repeat until you have no users left.
- Cthulhu_ 3y agoMaybe they just keep a safe margin so they can add their own metadata inodes or something. Or it's arbitrary.
- vlovich123 3y agoSure. But all that would mean is that you can only have 5M files materialized locally. I don’t see why drive itself would be limited at 5M. Very strange for this to have any real technical limit.
- Freedom2 3y ago[dead]
- deleted 3y ago[deleted]
- martius 3y agoDisclaimer: SRE at Google, not on Drive. I didn't look into this and don't know what really happened, but I can guess. It's more likely a tradeoff. Either you set this limit and protect your service from an identified scalability limit with the current architecture, or you plan for a (possibly long and expensive) redesign to get rid of this bottleneck. My guess is that a group of SRE and devs identified the risk, listed their options, evaluated the impact on users (eg: which fraction of users have more than 5M files) and assumed that the change would be mostly unnoticed and that the error message would be enough to push users with more than 5M files in their drive account to do some cleanup. This group of people underestimated the impact and maybe chose to not involve product managers. It's also possible that the decision was rushed because of recent layoffs or any other random event which pushed engineers to act quickly. With the bad press, leadership got involved, decision to rollback was taken. In my team, we would have a retrospective doc to discuss the issue (not exactly a postmortem, as this process has specific requirements which would not be applicable to this case). I think this is a easy mistake to do even with very good intentions, and I can see myself doing it.
- soundsgoodtome 3y agoIs PM not consulted for changes like this? I feel like this is something that a competent PM would ask eng to pause on while they determined actual user impact and lines up messaging.
- martius 3y agoShort answer, yes, a PM should have been consulted and give their approval for the change. In the scenario I described above (again, it's a guess, I don't know what actually happened), it's possible that the PM was bypassed because the engineers for a reason they thought was good. For example: * they didn't even think about involving a PM because that's something you have never done, * the people who wrote/reviewed the change assumed the conversation already happened, * they were pressured to move fast to mitigate an imminent or existing problem (performance, scalability, ...), * the PM who should have made this call "left" the company and didn't (get a chance to) hand-off their responsibilities. I guess what I mean with these comments is that sometimes there are misses like this. Maybe it's a sign of a systematic failure and that internal processes should be improved, but I don't think it means that the company is fundamentally unable to handle these changes correctly or can't make the right technical or product decisions.