5 ms·
Fluentd: a high performance unified logging layer
- cwyers 11y ago> We will go through the installation process, basic setup, listen to events through the HTTP interface, and look at a simple use case of storing HTTP events into a MongoDB database. ...and I'm out.
- kiyoto 11y agoOne of the maintainers of Fluentd here. Care to elaborate?
- deleted 11y ago[deleted]
- wyaeld 11y agostep1: build useful tool for serious people step2: mention mongo step3: enjoy having the room to yourself :-P
- kiyoto 11y agoI'm not here to really defend MongoDB, but the amazing thing about MongoDB is that, despite all the vitriol against it, it continues to be used at companies and projects that are far greater than what many naysayers ever touch: Stripe and Wish.com immediately come to my mind. Also, I am giving MongoDB the benefit of the doubt per Curt Monash's law of databases: Rule 1. It takes at least 7 years to build a database Rule 2. You are not an exception to Rule 1.
- dijit 11y agoFacebook and Google(?) still use mysql. I would still use mariadb or postgresql before mysql. just because the big boys use it, doesn't mean it's good.
- cwyers 11y ago> the amazing thing about MongoDB is that, despite all the vitriol against it, it continues to be used at companies and projects that are far greater than what many naysayers ever touch: Stripe and Wish.com immediately come to my mind. I honestly don't know if that mitigates against a lot of the MongoDB hate you see around here (including my post above). If I was at a place like Stripe and that money to pay Aphyr to battle-test all my database systems and an ops staff the handle all the issues that he found, then I could maybe deploy MongoDB successfully. I'd also have the sort of really, really hard problems that require me to spend that kind of resources on the problem. MongoDB is a much, much worse choice for those of us who DON'T have those needs and resources, though.
- ploxiln 11y agoI worked at bitly for a couple of years, we used mongodb for one of our secondary datastores. We had the expertise to keep it working, but we really hated it. It's this funny situation where some engineer starts off using it because it's very convenient to run, and maybe the data model is convenient. You scale it up a bit, debug it, tune it, stuff you'd have to do with any system. At the point when you're sure mongo really sucks, it's too late. The business and even technical management side always prioritizes some other project over replacing it, even though you spend a surprising amount of time each month maintaining the mongodb cluster (performance degredation that compaction doesn't completely fix, and other stuff...). I get the sense that with 3.0.3 or so, mongodb is a real database now. But it's been years of pain and false advertising. I'd still always vote against it. (Even though at my current place some people started using it for a service... :(
- flohofwoe 11y agofluentd is one of the more enjoyable logging systems I worked with. None of the above (HTTP, MongoDB, etc...) is needed at all, you can just as well setup a logfile watcher as input, which sends the log messages over to a logfile writer on another machine, all within 5..10 minutes on 2 VMs. Or you can directly write messages (JSON or MessagePack) to a TCP socket (MessagePack-formatted messages to TCP socket is basically the lowest level you get). Messages have a tag for filtering and routing. It's basically just input plugins connected to output plugins, and you build a message distribution network from that. And it is extremely easy to get into.
- Sanddancer 11y agoThis feels like they are trying very hard to avoid calling their syslog a syslog, probably with good reason. A lot of distributions provide a very minimalist build of whichever syslog daemon they've decided to use, and as such, people get the idea that syslogs can't write to databases, or parse json, or listen on pipes or any other number of things modern syslogs can do.
- kiyoto 11y ago>This feels like they are trying very hard to avoid calling their syslog a syslog I _wish_ rsyslog and their friends were really that easy to extend, and I say this as a maintainer of Fluentd. Fluentd came about precisely because syslog family of data collectors fall short in certain ways: 1. tag-based data routing: as you get more and more data sources, it becomes very important to keep track of what goes where. In my view, this is one of the key reasons Fluentd is used at many companies, and why Kafka has become popular on the message queue side (topic-based stream modeling) 2. Extensibility: afaik, it's not all that intuitive for most programmers and sysadmins to extend and add new inputs, outputs, filters, etc. for rsyslog and/or syslog-ng. Admittedly, this is a subjective point, but looking at both Logstash and Fluentd's vast lists of plugins [1][2], I feel justified to make this claim. 3. Configurable transport logic: I've never met anyone who is happy with syslog's buffering and/or failovers. Because we've heard so much about this particular problem, when we were building Fluentd (...4 years ago), we took extra care to make buffering and failover easy to configure and extensible. Happy to answer more questions =D [1] https://www.fluentd.org/plugins/all https://www.fluentd.org/plugins/all [2] https://github.com/logstash-plugins https://github.com/logstash-plugins
- dozzie 11y agoWhere did you get the idea that message transporter is a syslog derivative? syslog is meant to collect logs from daemons, which typically run locally. Remote logging (a.k.a. 514/udp) does not make syslog a general purpose transport. Fluentd, on the other hand, does very little to handle logs. It receives, processes and forwards structured messages. Those can be logs, of course, once you apply liblognorm or grok, but it's far from be the only thing. I wouldn't use, for example, syslog to transport monitoring data or inventory, but Fluentd matches the task very well.
- vpeters25 11y agoFor years I've been considering coding a tamper proof logger, something where each entry has a hash that depends on the entry's log and the hash of the previous entry. This could help detect potential system compromise. I haven't really taken the time to look for a logger with such a feature, it would be nice to know if fluentd has something like this.
- kiyoto 11y agoThat's an interesting idea: right now, we do not have this. However, it should not be too difficult to implement it either as a filter or output plugin.
- ploxiln 11y agosystemd's "journald" does this, it's called "Sealing", and it's enabled by default. I turn it off on my stand-alone systems, it's a waste of processing on them IMHO (particularly if you don't backup the sealing key)
- marcusmartins 11y agoHeka by Mozilla is another alternative - http://hekad.readthedocs.org/en/latest/ http://hekad.readthedocs.org/en/latest/. I have been running in production to ship docker logs to our Elasticsearch cluster.
- kiyoto 11y agoInteresting. Do you run hekad on the host machine?
- jedisct1 11y agoHeka's really nice, and the lua sandbox makes it easy to write new codecs. We (OVH) recently tried Heka to push syslog data into Kafka, but it eventually kept crashing under load. Other alternatives were way to slow for our needs. So we ended up writing a simple tool, not as flexible as Heka but about 10 times faster https://github.com/jedisct1/flowgger https://github.com/jedisct1/flowgger
- core0 11y agoWe'd love to use a ruby-based solution like this, but the docs say it will lose data whenever the receiving end crashes. Any plans to fix that? The way it was described in the docs gave me the impression there is no acknowledgement of network writes - if that's true won't even clean shutdowns lose data sometimes?
- kiyoto 11y ago>We'd love to use a ruby-based solution like this, but the docs say it will lose data whenever the receiving end crashes. Any plans to fix that? Where does it say this? I don't think this was ever the case for Fluentd. >The way it was described in the docs gave me the impression there is no acknowledgement of network writes - if that's true won't even clean shutdowns lose data sometimes? This is not true. All writes are acknowledged over TCP, at least between Fluentd and Fluentd.
- core0 11y ago> Where does it say this? I don't think this was ever the case for Fluentd. It's in http://docs.fluentd.org/articles/high-availability#forwarder-failure http://docs.fluentd.org/articles/high-availability#forwarder..., which says: However, possible message loss scenarios do exist: The process [log forwarder’s fluentd from the paragraph above] dies immediately after receiving the events, but before writing them into the buffer. Is this document out of date?
- kiyoto 11y ago>Is this document out of date? No, in the described case, the message can get lost. This is a really unlikely scenario though. The only real-world case that I know of first-hand is using file buffer and somehow being unable to write to disk, possibly because the disk is full. Something like that can be prevented by a fairly routine set of server monitoring alerts.
- frsyuki 11y agoACK of network transfer is available ("require_ack_response" option). This option ends up choice of at-most-once semantics vs. at-least-once semantics. You need to choose and you can choose. Fuentd provides "buffer_type file" to buffer records on disk. Shutting down won't loose data. If you need to choose memory buffer for performance reasons, fluentd enables "flush_at_shutdown" option by default. You would also want to use <secondary> feature. This lets you to write a buffer chunk to another storage if the primary destination is not available "retry_limit" times. Those concerns would be solved by the document: http://docs.fluentd.org/articles/out_forward#buffered-output-parameters http://docs.fluentd.org/articles/out_forward#buffered-output...
- riquito 11y agoI'm just starting to adopt fluentd but I'm scared by the fact that the "official" drivers have different interfaces and unclear leadership. e.g. https://github.com/fluent/fluent-logger-php https://github.com/fluent/fluent-logger-php https://github.com/fluent/fluent-logger-python https://github.com/fluent/fluent-logger-python and a bit of drama https://github.com/fluent/fluent-logger-php/issues/36 https://github.com/fluent/fluent-logger-php/issues/36
- edsiper2 11y agonote: I am one of Fluentd maintainers. There is nothing to be scared: - Fluentd is an enterprise solution already adopted by thousands of users. - Fluentd is sponsored and made by Treasure Data[0] where we collect around 800k events per second. - You have to make a difference between what is official and what is third party. Fluentd have more than 300 plugins and is likely that you will find some differences on how each extension is used, but at the end everything is compatible. We make sure to maintain a clean list[1] of functional and maintained extensions. - We lead the project and we invite you to reach us anytime through our mailing list or other communication channel[2] reach us anytime :) [0] http://www.teasuredata.com http://www.teasuredata.com [1] http://www.fluentd.org/plugins http://www.fluentd.org/plugins [2] http://www.fluentd.org/community http://www.fluentd.org/community
- elcct 11y agoI started building something of similar concept a while ago. But had to pause the development for some time. It is written in Go, so much easier to install etc. It has plugins for Hadoop, Mongo, RabbitMQ, File and stdout :) Can take data from tail, HTTP, RabbitMQ, heartbeat and other chains. https://github.com/mikeszltd/chainsd https://github.com/mikeszltd/chainsd
- jedisct1 11y ago"much easier to install etc." Fluentd is extremely easy to install. They provide packages with everything you need; you don't even have to install Ruby beforehand.
- elcct 11y agoSure, that is not the best point, but I like to have as little stuff as possible installed and polluting servers with Ruby doesn't feel right to me. Of course this is just a personal taste.
- IMTDb 11y agoIs there any benefit/disadvantage between fluentd and logstash ? I am not using either one, but I'll need to centralize my logs soon. My understanding tells me that these are two very similar projects, but I might be wrong.
- annnnd 11y ago> Its built-in reliability through memory and file-based buffering to prevent inter-node data loss have... I wonder if we can expect "Call me maybe - fluentd" from Aphyr soon. ;)
- jedisct1 11y agoI've been using Fluentd for years, and it's a super useful tool. It initially had some memory leaks that prevented us to use it in production, but it's now very stable. Writing new input/output plugins is extremely simple. And yes, it's written in Ruby, but give it a spin before judging; it's fast enough for most needs.
- kolev 11y agoIsn't Fluentd in Ruby though? It's 2015 and we need something like this in Go [0] [1] or Rust [2]. [0] Heka: https://hekad.readthedocs.org/ https://hekad.readthedocs.org/ [1] Chainsd: https://github.com/mikeszltd/chainsd https://github.com/mikeszltd/chainsd [2] Flogger: https://github.com/jedisct1/flowgger https://github.com/jedisct1/flowgger
- jedisct1 11y agoFluentd is pretty fast, with the critical parts being written in C. It's fast enough for most applications while providing a lot of flexibility.
- kolev 11y agoWell, written in Go or Rust, it could have both the readability of Ruby, the performance of C/C++. Having critical parts in C and the rest in Ruby is worse.
- allengeorge 11y agoWhy does the language it's written in matter? If it has the features you want, the reliability you want, and performs well given the load you're applying - that should be enough, no? It's not like we're embedding it into an app we're writing.
- 102030485868 11y agoWell, in terms of maintenance it's a little bit more work. Sure, it has the features and reliability. But does that really justify having to maintain a completely new environment? Maybe it does, maybe it doesn't; it depends. No you may not be embedding it in your app, but it's now a part of your stack. You'll need to keep an eye on updates, etc for a completely different environment. Plus there are other factors, like approved languages. Certain companies only allow using languages X and Y. Don't even think about language Z. I had to re-write an 80loc Python script to a much larger Perl one because I just didn't grok Perl all too well at the time. It didn't matter that Python was installed on the system. All that mattered was that Perl was approved and Python was not.