6 ms·
Fly.io Infra log: week-by-week record of what the team does
- shafyy 2y agoI like the general idea, but it feels a bit weird to describe what each specific person did and naming them. Only thinking about working on this team gives me anxiety. Also, the weekly cadence is a bit high?
- redrove 2y agoExactly, this only made me feel sorry for the people working there. Weekly updates by naming and shaming every team member would cause me so much anxiety even if they were just internal. I’ve been on SRE/infra teams my entire career.
- lbotos 2y agoI'm not on SRE teams so I'm curious how your team works: - Is it remote? - What is the update cadence? - Is this information shared with Manager and "manager communicates team status"?
- tptacek 2y agoWe talked with the team before this went up --- originally, it had people's full names on it, which nobody liked. I think a lot of people here have had bad experiences in other company cultures that don't apply here; we're a profoundly bottom-up team (we don't have product managers, ICs [across the board] own chunks of the roadmap, EMs don't assign work) and signing names to things that get done probably makes more sense here than it does elsewhere. That's not to say we don't have internal dysfunctions, or that we don't do things that annoy engineers. Just that that our dysfunctions are different. Weekly feels right to me. We'll see how it goes.
- kungfufrog 2y agoJust out of curiosity, and since this is vaguely related to the content, how has Fly.io adjusted to SOC 2 when it comes to infrastructure changes and deployment? I read a while ago about how you guys passed SOC 2 which was an enlightening read, but I'm curious what systems you have put in place to keep velocity going and how you manage change more generally across infrastructure? My current experience is that once we started transition to SOC 2 compliance and brought in new processes everything crawled to a halt, we now have multiple layers of reviews and scheduling, and the bureaucratic process is followed at the cost of expediency and doing sensible things in a sensible way (i.e., everything must follow the process, no matter how trivial or significant.)
- tptacek 2y agoWe have a "SOC2 review" regime, for a security-sensitive subset of our code. The very most important thing to remember about SOC2 is that your auditors are attesting that you meet a standard that you yourself set. If you set a standard for yourself that every single change anywhere is going to be reviewed in triplicate, that's what auditors are going to check you on. So the key is to be extremely deliberate about the standard you're setting in your initial Type I. Every auditor is going to want to see some kind of change review, but the purpose of that change review is something you determine, not them. For us, the SOC2-auditable change review process is about ensuring that people merging PRs aren't 3 raccoons in a trenchcoat; it's not a vulnerability management process. The trap people fall into is that SOC2 is often the point at which they start paying attention to security process as a whole, and they led SOC2 lead security process for them. No! Death! Be thinking about security from day 1, and have clarity about what subset of your security process you want SOC2 to measure and track.
- kungfufrog 2y agoThank you so much for the considered response. Very insightful and it's given me some food for thought.
- lopkeny12ko 2y agoIf you are a member of the team, and are actually getting things done every week, what is there to be anxious about? It's a great accountability mechanism.
- shafyy 2y agoIt would just stress me out. But I guess everbody is different.
- mrmincent 2y agoI like the first half of describing the issues faced during the week - it’s something I’ve been thinking of doing for my team, and the format is easy to digest. I really don’t understand the point of the second half in publicly disclosing what each individual has done though. I hate the idea that team members would feel like they might have to rush stuff out so they look as productive as someone else, or that they’d feel not as good as others.
- usrme 2y agoI like the spirit of the second half to demonstrate what the team is up to, but I too wondered about potentional one-upsmanship problems. I'll definitely start following this to learn more though.
- iamkonstantin 2y agoIt reads more like a retrospective for each team member? It wasn’t clear if it’s opt-in to have your status update posted publicly but I can imagine (and hope) it’s okay not to partake.
- aranw 2y agoFor you the second half probably doesn't make much sense but if these documents are also shared internally it would be really useful for others at the company to see what the team and individuals have been up too
- hackernewds 2y agothat seems micromanage-y and fraught with lack of psychological safety
- aranw 2y agoBut what if they are just used as internal updates? Surely you share updates with what you've done in some way?
- tptacek 2y ago
- fidotron 2y agoWas it public before that fly.io had pieces hosted on OVH?
- sudhirj 2y agoIs there any issue with OVH in particular?
- fidotron 2y agoI'd invert that - by far the lowest latency ping I've seen for an external service was fly.io (single digit ms), and I never worked out quite where it was. Edit: to be clear, I mean ping from my house, not another data center.
- dns_snek 2y agoAre you very close to that particular datacenter by any chance?
- fidotron 2y agoAm in Montreal so assume they have one here practically connected to my ISP. But similar local hosting has been twice as slow, hence the oddity.
- xena 2y agoYep, there's a PoP in Montreal. Deploy to region YUL to use it.
- fidotron 2y agoYeah, testing that is how I saw the impressive figures. Like many others though I cannot live with the fly unbounded pricing model or difficulties with data persistence, otherwise I would absolutely have been there a year ago. EDIT: In fairness to fly it appears they have updated both their pricing plans and data persistence situation since I was making these decisions.
- kulor 2y agoFly does a good job of illustrating requisite battles with entropy when operating complex systems
- deleted 2y ago[deleted]
- tossandthrow 2y agowe have had serious problems with fly. namely that flycast addresses could not connect or resolve. this issue was present in multiple in all regions we tried. and has been going on for months. in the end we left fly postgres and manage everything ourselves. fly has a lot of issues with their platform. from my perspective this reads like what I would usually do as a part of a development team to the strategic teams when things does not work. the only issues is that I am a customer at fly, not a strategic relation - we simply move on the other platforms.
- CSSer 2y agoWhere’d you end up moving to, out of curiosity? I’ve kind of been in the same boat.
- tossandthrow 2y agoWe just manage our own postgres container on Fly and use the <app-name>.internal. We only use fly for a dev / testing environment. Production and staging run on a more mature platform.
- nik736 2y ago..."off of special-purpose hosts on OVH"... Would be interesting to know, for what exactly Fly is using OVH.
- tptacek 2y agoThe metrics cluster, and that's it.
- leonheld 2y agoAdding to the sentiment for other comments: please reconsider the bottom half portion where you name the actions of each individual. There's a lot of hidden work in software development and the only person who should care about individual actions is the PO/Manager/whoever, not the end-customer or stakeholder.
- deleted 2y ago[deleted]
- siamese_puff 2y ago[flagged]
- deleted 2y ago[deleted]
- tonetheman 2y agoYeah that is weird hustle porn or at least the bottom part of that page is. I would not want my name on that page if I worked there. Tech companies are weird. The top portion of the page is great.
- deleted 2y ago[deleted]
- CoolCold 2y agoInteresting, thanks! * something about infra after all, something to learn from, dm-clone - TIL * largest team and just 6 persons...whoa! * on "A Registry machine unexpectedly reached a storage limit, disrupting deploys that pushed to that registry for about 20 minutes." - quick reaction assuming it required some manual steps to do after issue has been detected * naming is simple "infra engineers", not SRE, not DevOps, not DevOps Developer, not even plain Sysadmins Added blog to my RSS reader, should be useful and fun
- tptacek 2y agoThe team is way bigger than 6 people.