4 ms·
Curious what use-cases you're envisioning here. Having worked with the system described in the past, it's quite flexible: data can be easily pulled out into mo
by Shog9 6y ago
Curious what use-cases you're envisioning here.
Having worked with the system described in the past, it's quite flexible: data can be easily pulled out into more specialized systems for aggregation or trend analysis, while the raw logs remain quickly accessible over long enough periods of time to allow for digging into everything from support cases involving a single person to monitoring distributed attacks. Access can be controlled and monitored with reasonable granularity, and training new folks to use it is as easy as teaching some basic Select queries.
Not gonna suggest it's the most efficient system (if nothing else, SQLServer wastes a TON of disk space), but it minimizes complexity while deftly avoiding the choice between throwing away information when it's still needed and keeping everything in flat files which are too unwieldy to be used.
- oneplane 6y agoIt's not that the system itself isn't functional at all, but think about cost, flexibility (i.e. the migration as posted) and support. If you have to move that much data at once because your RDBMS doesn't work in any other way you're going to be in trouble, even if you get the 50k costing SSDs. All of the things SQL server does can be done with not-SQL-server things, but you suddenly gain the capability to take shards offline and do multi-version migration, multi-host migration, multi-storage-backend migration. You can query in SQL, but also GraphQL. You can use Lucene search, and you can use ES-specific queries. You can have multiple write hosts, you can have tiered storage on application-level, OS-level and SAN-level and they can actually work together. Again, it's not that SQL server doesn't do anything, it's just that it's probably not the best plan for time series access logs.
- Shog9 6y agoI'm sure there are better options. OTOH... I was at SO for about 9 years, and the system described existed for all of it - predating GraphQL, predating even ElasticSearch. There's stuff built on top of it that was never envisioned when it was created, and stuff that wouldn't have been possible without it. At this point, moving to something else would likely be a far more costly investment just in terms of requirements analysis than the migration described in the blog post. Would it pay off? Maybe! But for a critical bit of infrastructure, that's something you put an awful lot of careful thought and research into before you venture to do anything... And meanwhile, that infrastructure has to be kept running.
- oneplane 6y agoOne important facet that isn't much technology related at all is the knowledge pool you can work with. Every time a generic component is implemented non-standard you essentially end up with unusable knowledge for both existing people leaving for another job and new people coming in with the common knowledge. While most companies that are old enough to predate the iPhone usually have a lot of in-house technology that predates conventions and generic implementations, there is plenty of movement towards commonalty instead of keeping that special snowflake stuff in; at least at the places where I have/had influence. When you freeze an in-house tool in time you essentially rob your own people of progress. The solution isn't rip-and-replace, but a process of continuous improvement is good for everyone.