4 ms·
elasticsearch would not cope with that volume of data on the hardware described.
by d4rti 6y ago
elasticsearch would not cope with that volume of data on the hardware described.
- user5994461 6y agoWell, the article doesn't mention the hardware. Only one machine for 5 TB a day which doesn't add up as far as I am concerned. If they were storing everything in S3, it's possible to do something similar with ElasticSearch for the same order of costs (maybe three times?), the money going to EBS storage instead. ElasticSearch allows to query and visualize logs, which plain S3 storage doesn't, it's worth a bit more IMO.
- Spivak 6y agoI think the author's math tracks. Our pipeline looks something like hosts -> rsyslog collector -> kafka -> custom crunching -> elastic which is a bit inefficient and we try to go hosts -> kafka when we can but a lot of stuff only supports rsyslog so it's there and simple enough. The rsyslog collector and our crunchers are teeeny tiny compared to the rest of the pipeline and can chew through up to a week of backlog in a few hours. The bottleneck is the network for us and if we upped the pipe to 10G we could probably get away with a single host.
- kakoni 6y agoWell there is clickhouse. And projects like this https://github.com/flant/loghouse https://github.com/flant/loghouse
- cbsmith 6y agoI missed where a volume of data was described that was outside elastic search's capacity. Seems unlikely that it "would not cope" regardless of configuration.