12 ms·
In many ways you can think of large, long-living tech companies not unlike old cities like, say, London or Paris. The buildings and roads you see are built on
by ctur 4y ago
In many ways you can think of large, long-living tech companies not unlike old cities like, say, London or Paris. The buildings and roads you see are built on top of older buildings and ruins. The streets are weirdly shaped and intersect at odd angles because they were made hundreds of years before and adapted over time as needs evolved. There are catacombs underneath sidewalks and no one genuinely understands it all nor does a single reference exist that explains it.
Literally everything predates everyone who lives there. Generations and generations of original designers, architects, and laborers have arrived, plied their trade, and moved away. There are people who are experts in certain parts, and who can build a new skyscraper at any given spot, but it is just layering and organic growth.
The emergent complexity of centuries of being lived in and adapted belies easy understanding.
Large tech companies are similar. You just can't understand how "it" all works. If you were to build it from scratch, perhaps you could, because it would be simpler and clearer, but nothing was made with the current state in mind. It evolved and adapted over time.
So reading this, I am not surprised. I think you'd get the same answer about many other aspects of data, code, system history, etc at any other 10+ yr old tech giant.
- TomSwirly 4y agoWhat you appear to be saying is that large technology companies are designed from the ground up around flouting the laws about data retention. Why is this OK?
- borbulon 4y agoyeah "favela" is usually the metaphor I use, but it's the same general gist, if a little messier.
- CaptainTaboo 4y agoIt's not just the tech giants. I work for a company that is fairly dominant in its particular industry (If you're in the US, you've probably done business with one of its customers). Much of our growth has come through the acquisition of other companies, each of which gathers its own user and customer data. From my perspective in IT operations, I don't have any visibility into how we capture user and customer data, but I know we're big enough to track a user across the various products and websites, but still siloed enough that internally we still refer to the various business units as if they were separate organizations, each holding its own subset of the data. To sum up, this description of Facebook's internals doesn't surprise me at all.
- kixiQu 4y agoI work for a large long-living tech company. Where legal compliance and security are concerned, "it is just layering and organic growth" is not something you get to say. If we were speculating about such a thing on a coffee break, maybe we'd get the kind of answer given here, but if a single reference doesn't exist for something that is necessary for compliance, people get paged until it does. (There may be other things taken as seriously I'm not thinking of, but compliance and security are the two I've seen drag people out of their beds by their ankles)
- encoderer 4y agoYes and it’s a good thing that most industries do not carry this governance burden.
- mschuster91 4y agoIf more industries did carry this governance burden, we'd see a lot less cases of companies getting hacked left and right. For way too many companies, "IT" is just another cost center that doesn't bring in any profits and so is barely kept at life support levels. And when something inevitably happens, it's rarely the C-level that gets to go to the unemployment office, it's more one poor soul who has complained for years that he needs more staff and budget to at least get the most egregious bullshit fixed.
- phpisthebest 4y agoI doubt that, many of the companies that have been hacked were governed by these regulations, and they spent money governing to the regulations not governing to security which is not the same thing If I was burdened by regulation and red tape I can assure you my org would be less secure not more so
- seti0Cha 4y agoHah, I might have agreed before I worked on a credit card processing system. There were plenty of rules alright, but it's possible to abide by all the rules without actually making a secure system. Some of the rules made it harder to create a secure system because they did not reflect the state of the art, or they made adding additional security too burdensome to implement. In the end, you can't force competence through rules.
- markstos 4y agoYou can hire one or more data privacy people whose whole job it is to track this. They get an inventory of all the systems and talk to the managers of every system, find the answers and document them. Meta is choosing not to know or pretending not to know. They have the resources to know where your data goes.
- UncleMeat 4y agoI think you are underestimating the complexity of this task by about three orders of magnitude. This is the way to do it, but you can't do it for a company of this scale with a handful of people.
- RHSeeger 4y agoBut pretending they can only afford "a handful of people" to work on the problem is a little silly, too. They don't have the answers because they choose not to devote the resources to get the answers; not because they are unable to devote the resources.
- UncleMeat 4y agoI agree with that. Companies like Facebook should have a large org responsible for this problem. My point is that the problem is way more complex than some might think.
- citiguy 4y agoAgreed - when I worked in the financial industry there was something called data lifecycle management. Since keeping data around indefinitely meant you could be held liable at any time for anything that happened in the distant past, companies are motivated to track data and delete anything that's no longer necessary for business as soon as possible. No such law exists for Facebook, so they don't care. Someone needs to pass a law with a big fine attached and then things will change.
- ImPostingOnHN 4y ago
- mistrial9 4y agothis useful excuse-insight will be provided to all compliance personnel at the next staff meeting, in writing.
- BiteCode_dev 4y agoMaybe, but one can find out. You can follow the chain of calls of any software, and they have access to everything, as well as the whole history. It's like when in an interview Buffets says it's very hard to know where all the money is in the financial system. Sure, it's a complex system, but it's more of a matter of incentive.
- withinboredom 4y agoYou can follow the chains to a degree, but that assumes you have access. At work, you can follow things until it gets to a spool file, but you won’t have access to the system/code that reads the file.
- IanCal 4y agoUnless it was built with this in mind, you absolutely don't have access to this. Full data lineage is a complex task. A very simple example would just be that a database is destructively updated by some script at some time. What's the log here? What ran? Can you actually replicate it? Do you have full backup of the data before? It's so easy to lose this.
- ACow_Adonis 4y agoAs someone who's arguably professionally employed to do things like tell people where money is in the financial system, no, you can't. In the same way you can't enumerate the atoms or molecules in a human body. sure, you can come up with good estimates and summaries that are useful, but not the actual full and truthful state of affairs. In addition, you are a self-referential participant in the system: so acting in an attempt to observe such likely immediately makes your information at least partially incorrect. Additionally, if we can't get full/transparency or introspection into machine learning models or neural networks (and we can't because i don't believe we even possess the theoretical knowledge of how to do so), it should also be wilfully apparent that we can't actually know the true lineage or dependencies of systems that implement such in production. Not to mention the difficulties as the inputs/outputs of such pass in and out of the notional entity or system being observed.
- dvfjsdhgfv 4y agoAs far as code goes, it's true for many companies. As for data, it was similar in many large European companies in the pre-GDPR era. Today, one crucial question is always asked: when you work with data, is this personal data? If it is, you need to deal it with a special way. Personal data is both an asset and a liability. Most companies found a way do do it. It was a long and often very painful process, involving everything from data entry to backup processes and procedures, but somehow we managed to do it.
- deleted 4y ago[deleted]
- adrianh 4y agoThis is a seductive notion, but it's indeed possible for a tech company to understand how a single piece of data goes through its systems. It might take a while, and it might involve weeks of code spelunking and dozens of conversations with engineers — but it is indeed possible. A codebase is not encumbered with the physical constaints of an old city. It's not like trying to figure out the physical state of a certain area of London 30 meters deep — which might be impossible without digging and therefore disrupting other structures in the area. A codebase is simply instructions, which are readable and understandable (with the exception of black-box machine-learning models, but those at least have defined inputs and outputs).
- leros 4y agoThe thing with tech though is that it's easy to add new connections, connecting anything to anything. You might have two independent datasets that can't be stored together due to privacy or security regulations. The teams who work on that know and understand that. But some other team working on something completely unrelated might join those datasets for some non-nefarious purpose, which then creates a new dataset with those data combined. Now other teams who also don't know about the regulations might build on top of that combined dataset. They might even generate new data fields from the data that's not supposed to be joined. This stuff can propogate through layers of systems, queries, service calls, analytical engines, offline spreadsheets, etc. It's very difficult to keep track while keeping your engineering teams independent and nimble.
- rufus_foreman 4y ago>> It might take a while, and it might involve weeks of code spelunking and dozens of conversations with engineers — but it is indeed possible If it takes a while it will be out of date by the time it is complete. A single request might directly impact a dozen separate systems, which in turn call other services. You won't be able to find anyone who knows much past the layer they directly interact with. All of these services are being worked on. Some older ones are in the process of being deprecated in favor of newer ones. Someone is already working on the design of the replacements for the newer ones. Changes are being pushed weekly if not daily. If you send a request and then send it again a minute later, it might go through a completely different set of services based on a feature toggle. You can't step into the same river twice.
- Jenk 4y agoI agree with the point you are making, but I disagree with the conclusion. A company like Meta that has been building and advancing the field of ML, AI, and even falling foul of the ethics/morality of large scale public manipulation and political campaigning has already demonstrated they have the wherewithal to find where their data exists. Period.
- citiguy 4y agoThe issue is more that they don't care. They don't really need to track it so they don't.
- liampulles 4y agoTo extend this wonderful metaphor, perhaps the least terrible solution is for the "city workers" to constantly log any issues as they come across them in their normal day-to-day work. It is the responsibility of the city then to proactively go and fix those issues as they come up.
- jollyllama 4y agoTo build on your metaphor though, there are cabbies with "the knowledge" in London, and probably similar folks in Paris. The way I have seen this work in tech is, you have to keep the people who know where the bodies are buried around for a long time, and others will use them as a resource to inform their efforts. The alternative, of letting them walk over the years, is that large sections of your stack will have "here be dragons" written over them and become no-go zones where your only option is to route not only engineering but business efforts around them.
- slenk 4y agoNot tech giants that care about compliance...it's not undoable. They are just poor stewards.
- hkgjjgjfjfjfjf 4y ago
- hkgjjgjfjfjfjf 4y ago
- rockinghigh 4y ago> I think you'd get the same answer about many other aspects of data, code, system history, etc at any other 10+ yr old tech giant. Processes to access personal data vary wildly among tech companies. At Apple, when a machine learning engineer wants to store and use data on the server side, they need to go through layers of approval from lawyers. Even when approved, it often comes with serious constraints about what can be done with the data. Meta is a lot more lax.
- weare138 4y agoThat's just a truism at any tech company. The issue is Facebook/Meta was misleading about what data is being gathered and the "Download Your Information" didn't actually show you what data FB had gathered about you. By their own admission they don't even know yet users and regulatory agencies were knowingly misled into believing that information was accurate.
- Rackedup 4y ago> The buildings and roads you see are built on top of older buildings and ruins It's easy to delete old data that is never accessed, so I don't get your point...