13 ms·
A log/event processing pipeline you can't have (2019)
- traceroute66 6y agoIt was all making for such an interesting read until the last paragraph. I don't know about anyone else, but I have this inherent hatred of company marketing material disguised as blog posts. If you are going to write a decent blog post, then write a decent blog post. If people are curious about the author they can look them up (and their affiliation). Don't turn it into a sales pitch.
- MaulingMonkey 6y agoWhile I share the annoyance for stealth sales pitches: 1. it appears to have been added 2 months after the initial posting rather than a cynical cash grab. 2. it appears the resulting company has pivoted to VPN stuff and is no longer in the log/event processing business anyways?
- tmd83 6y agoThat's what I was wondering. Tailscale doesn't do logs did it stared with that :(. I would loved to see there logging solution. I like the writeup though I need to do a second pass to fully review everything. The biggest thing that just suddenly hit me was 5 TB/day is really 60MB/s.
- ithkuil 6y agoThe Company's mission statement is "Simplifying the long tail of software development". Perhaps networking is only of the possibly many building blocks they plan to tackle to simplify software development for the normal companies.
- ghj 6y agoHe wrote the post on 2019-02-16 and didn't put in the edit until 2019-04-26 so the causality probably happened the other way around. It seemed like this is a career summary from his time at Google Fiber (which was killed by google). It sucks when your baby gets killed for reasons outside of your control. You want those years of your life to mean something, for the work to live on in some form (hence the "Please, please, steal these ideas"). He probably realized after publishing that the post wasn't enough to give him closure so he started a company to keep working on it instead. (I don't know why I am psychoanalyzing a random blog. I am probably projecting really hard.)
- chestervonwinch 6y agoFiber was killed? Google won’t stop hassling us about getting it here in Austin, TX!
- throwaway88955 6y agoYou got it :) simply this company has been receiving advertising on HN that is worth hundreds of dollars since last December. Check it out yourself https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=tailscale&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... Fake organized upvoting rings, artificial placement on the homepage if the post isn't getting organic upvotes, re-upping, etc... Literally 100% of posts about this company got to the homepage in a website where 99.99999% of posts don't get even 2 upvotes no matter how great they are. It's corruption that could result in jail but hey, in tech you can get away with this kind of corruption
- apenwarr 6y ago[Author here] As others have pointed out, it wasn't originally a sales pitch and we later pivoted to something else. I agree it came across in a kind of annoying way. I've added a new update at the end to describe what happened, for anyone who is curious.
- cpach 6y agoThank you. Good luck with Tailscale! Very cool service.
- winrid 6y agoFun read, nice to see some "down to Earth" engineering. :)
- ezekiel68 6y agoLoved this. Great, actionable advice that's still applicable over a year later -- and a true geek's sense of humor. The hidden contrarian in all of us cheers along with his trials and triumphs. I didn't mind the soft-sell final paragraph at all, since he gave away the keys to the kingdom in the rest of the article anyway.
- RasmusWL 6y agoI was excited to look at the company, but turns out they pivoted away from doing log processing :(
- deleted 6y ago[deleted]
- user5994461 6y agoGood read, except the part where the author says there are no existing solutions for processing logs. There are quite a few robust scalable ones. syslog-ng, logstash or fluentd on the host to collect and aggregate logs. (logstash/fluentd can parse text messages with regex and handle a hundred different things like s3/kafka/http but they are much more resource intensive). kibana or graylog to centralize logs and search, the storage is elasticsearch. A simple syslog-ng on the devices could probably do the job. Little known fact about syslog, it can reliably forward messages over TCP, logs are numbered, have retries and syslog-ng can do DNS load balancing.
- camdencheek 6y agoThe flexibility of Fluentd and the community around it are great, but ultimately the resource intensity proved to be too high and too unpredictable for our use case. We ended up building our own general purpose log parser in Go to support our hosted log monitoring platform. Though it's still somewhat young, performance is great (beats Fluent Bit in most of our benchmarks), and it's nearly as flexible as Fluentd in terms of configuration. If you want to check it out, we recently open-sourced it and are always looking for more feedback: https://github.com/observIQ/stanza https://github.com/observIQ/stanza
- d4rti 6y agoelasticsearch would not cope with that volume of data on the hardware described.
- user5994461 6y agoWell, the article doesn't mention the hardware. Only one machine for 5 TB a day which doesn't add up as far as I am concerned. If they were storing everything in S3, it's possible to do something similar with ElasticSearch for the same order of costs (maybe three times?), the money going to EBS storage instead. ElasticSearch allows to query and visualize logs, which plain S3 storage doesn't, it's worth a bit more IMO.
- Spivak 6y ago
- gbrown_ 6y ago> So, the pages are still around when the system reboots. ... > The kernel notices that a previous dmesg buffer is already in that spot in RAM (because of a valid signature or checksum or whatever) and decides to append to that buffer instead of starting fresh. This sounds like it should be very unreliable. Perhaps it works in practice but I couldn't see myself relying on such a mechanism.
- baruch 6y agoI've worked on a system that relied on it to do fast upgrade reboots and it worked. I'll have no problem relying on this again.
- detaro 6y agoWhy? In a software-triggered warm reset you wouldn't loose data to lack of refresh, so it doesn't seem that different than recovering from an interrupted on-disk log or whatever.
- fluential 6y agoGreat article based on real life experience. I have been building logging and protective monitoring pipelines for a while now. From my experience if it comes to log shipping from hosts rsyslog + relp + disk assited in-memory asynchronous queues are preferred, most of the time you just only have network i/o as logs would not touch disk. The idea is to ship logs off the device ASAP as well as destination acts as a sink server capable to handle most of the spikes withouth stressing local source. All done via rsyslog which also wraps actual logs into json format locally. The glue could be syslog-tag. At the other end you could have ELK stack and logstash using json_lines codec input (pretty fast) structuring data further to your likings. Just looking into metrics now the avg time for logs showing in ELK is 7-200ms (the latency comes mostly from specific reads happening against the ES cluster). As ELK is always the slowest component, dropping logs compressed in-memory directly onto disk is also an option. One thing to note is that RELP can produce extra duplicates which are easily handled by inserting into Elasticsearch using specific document ID (some performance penalty) which could be some unique hash computed on (log content, timestamp, host) etc. With this in place you can also easily "replay" stream of logs to fill potential gaps. This type of setup scales really good as well. Edit: typos
- wwright 6y agoWe have a similar setup, but I've been generating UUIDs for each event at logging time to avoid any computational overhead for deduping.
- cbsmith 6y agoUnless they are type1 uuid's, there's probably just as much computational overhead going on in the uuid creation. ;-)
- wwright 6y agoWell, in our case, the event generators are almost never CPU bound but the logstash nodes are. Generating a UUID (which is mostly just waiting for the kernel to hand you some bytes, which allows other threads time to do useful work) before sending has no opportunity cost and frees up the more-valuable logstash resources. (These are micro-optimizations anyway.)