8 ms·
Show HN: My Single-File Python Script I Used to Replace Splunk in My Startup
My immediate reaction to today's news that Splunk was being acquired was to comment in the HN discussion for that story:
"I hated Splunk so much that I spent a couple days a few months ago writing a single 1200 line python script that does absolutely everything I need in terms of automatic log collection, ingestion, and analysis from a fleet of cloud instances. It pulls in all the log lines, enriches them with useful metadata like the IP address of the instance, the machine name, the log source, the datetime, etc. and stores it all in SQlite, which it then exposes to a very convenient web interface using Datasette.
I put it in a cronjob and it's infinitely better (at least for my purposes) than Splunk, which is just a total nightmare to use, and can be customized super easily and quickly. My coworkers all prefer it to Splunk as well. And oh yeah, it's totally free instead of costing my company thousands of dollars a year! If I owned CSCO stock I would sell it-- this deal shows incredibly bad judgment."
I had been meaning to clean it up a bit and open-source it but never got around to it. However, someone asked today in response to my comment if I had released it, so I figured now would be a good time to go through it and clean it up, move the constants to an .env file, and create a README.
This code is obviously tailored to my own requirements for my project, but if you know Python, it's extremely straightforward to customize it for your own logs (plus, some of the logs are generic, like systemd logs, and the output of netstat/ss/lsof, which it combines to get a table of open connections by process over time for each machine-- extremely useful for finding code that is leaking connections!). And I also included the actual sample log files from my project that correspond to the parsing functions in the code, so you can easily reason by analogy to adapt it to your own log files.
As many people pointed out in responses to my comment, this is obviously not a real replacement for Splunk for enterprise users who are ingesting terabytes a day from thousands of machines and hundreds of sources. If it were, hopefully someone would be paying me $28 billion for it instead of me giving it away for free! But if you don't have a huge number of machines and really hate using Splunk while wasting thousands of dollars, this might be for you.
- rgrieselhuber 3y agoThanks for sharing.
- appleaday1 3y agoThanks for the share, I still find it hilarious how Python is by default installed on most distros, I was working on some compression tools and by default the os didn't come with the ability zip/unzip toolsets, but the python standard library zipfile did. https://docs.python.org/3/library/zipfile.html https://docs.python.org/3/library/zipfile.html
- codetrotter 3y ago> by default the os didn't come with the ability zip/unzip Some versions of tar are able to extract zip files. Try tar xf somefile.zip It might or might not work with the version in your OS
- BerislavLopac 3y agoYou don't even have to write a custom script around the library: python -m zipfile -e monty.zip target-dir/ https://docs.python.org/3/library/zipfile.html#command-line-interface https://docs.python.org/3/library/zipfile.html#command-line-...
- surfingdino 3y ago"Hell hath no fury like a Python dev annoyed" :-) Thank you for sharing!
- ozfive 3y agoHell hath no fury like any dev annoyed by something they could build themselves...
- hliyan 3y agoA long long time ago, I used a series of tail -f's and unix pipes to aggregate logs, and grep, less and awk to analyse them. There were about 20 different services written in C++, each producing over 1GB of logs each day. Managed to debug some fairly complex algorithmic trading bugs. Twenty years later, I still can't fathom why we're spending so much money on Splunk, DataDog an the like.
- dsXLII 3y agoVolume. 1GB of data per day is rounding error. If you have tens of thousands of servers, each generating hundreds of gigabytes of data per day, tail -f and grep don't scale especially well.
- vincnetas 3y ago100GB of logs per day? what kind of applications are that chatty?
- getrealyall 3y agoAnd I bet a hang glider can't fly from New York to Paris, either! The nerve! Recall that the poster said this was for a small startup. If you're Google, by all means, use Google logging tools. If you aren't, then solve the problem you have, not the problem your résumé needs.
- Okkef 3y ago"log files of several several gigabytes ... Process them in minutes" Thanks but no thanks.
- eigenvalue 3y agoI misspoke there-- meant to say: "The application has been tested with log files several gigabytes in size from dozens of machines and can process all of it in minutes." That's the time it takes to connect to 20+ machines, download multiple gigs of log files from all of them, and parse/ingest all the data into a sqlite. If you have a big machine with a lot of cores and a lot of RAM, it's incredibly performant for what it does.
- rollcat 3y ago"This simple tool solves X at my org" is probably the most underrated type of project. There's not enough room to overcomplicate something that isn't a core part of the business, it must be practical to maintain, simple&stupid enough so that onboarding is not a hurdle, etc. I encourage everyone to share your "splunk in 1kloc of Python" projects! Some of my own: - https://github.com/rollcat/judo https://github.com/rollcat/judo is Ansible without Python or YAML - https://github.com/rollcat/zfs-autosnap https://github.com/rollcat/zfs-autosnap manages rolling ZFS snapshots
- eigenvalue 3y agoThanks, based on the dismissive replies to my original comment in the Splunk acquisition discussion, I thought this would get a lot of hostile takes saying that it was dumb, that I reinvented the wheel because I didn't want to spend 2 weeks trying to figure out opentelemetry nonsense and tools X, Y, and Z, that it was trivial, that it wouldn't scale, etc. But people are actually being surprisingly nice and friendly! I guess people just really hate Splunk!
- egwor 3y agoI suggest you sell it to Oracle, get some popcorn and watch the Cisco vs Oracle log war begin!
- RexM 3y agomired in antitrust lawsuits
- _aavaa_ 3y ago> reinvented the wheel I hate this meme. It's as if cars, trains, and airplanes all use the same wheels. Or that wheels under my stove, my tiny filing dresser, and my shopping cart are all the same. Oh yeah, re-inventing the wheel, what a stupid idea and something we obviously don't frequently do and for good reasons. This meme is almost as bad as the horrible misquoted "premature optimisation is the root of all evil".
- 3y ago
- nurettin 3y ago> If I owned CSCO stock I would sell it-- this deal shows incredibly bad judgment." That may be so, but beware that acquisitions usually increase stock price rather than decrease it.
- candiddevmike 3y agoIs that overtime or immediately after acquisition? https://www.google.com/finance/quote/CSCO:NASDAQ?&window=5D https://www.google.com/finance/quote/CSCO:NASDAQ?&window=5D
- nurettin 3y agoThe results should show up pretty quickly. Maybe that was pricing in the acquisition, or the acquisition is nothing compared to the whole company and that's quarterly earnings. Not sure.
- KaiserPro 3y agoI used to work at a splunk shop. It was used for alerting, graphing & prediction. It was critical to how the company functioned. There was lots of stuff that relied on splunk, and we had splunk specialists who knew the magic splunkQL to get the graph/data they wanted. However, we managed to remove most of the need for splunk by using graphite/grafana. It took about 2-3 years but it meant that non techs could create dashbaords and alerts. As someone once told me, splunk is the most expensive way you can ignore your data.
- oblvious-earth 3y agoQuickly skimming some points that would irratate me if I had to maintain this script: * Importing Paramiko but regularly call `ssh` via subprocess * Unused functions like `execute_network_commands_func` * Sharing state via a global instead of creating a class Overall it's fit for purpose, but makes a lot of assumptions about the host and client machines. As you said in the thread you're running a very small number of servers (less than 30). I've written similiar things over the years and they are great for what you need. When I heavily used Splunk (back in 2013) I was in an application production support team that managed over 100 productions servers for over a dozen applications, there were dozens of other teams in similar situations across the company. The Splunk instance was managed by a central team, minimal assumptions about the client environment, had well defined permissions, understood common and essoteric logging formats, and could reinterpret the log structure at query time. A script like this is not competiting in that kind of situation.
- msto 3y ago* barely any comments and not a single docstring in the entire kiloline file
- eigenvalue 3y agoI find comments annoying to read and write and distracting. I’d rather fit more code on the screen at once and instead focus on making the variable names and function names really descriptive and clear so you immediately grasp what it’s doing from context alone. Nowadays, if you really need comments to tell you what code is doing, you can just throw it into ChatGPT and get it that way.
- fastasucan 3y agoI really like the tool you made, and appreciate helping your company save money as well! I don't think it matter that this isn't a perfect fit for everyone else (as you said, this was something you made to solve your problem) - but boy do I disagree with the "variable and function names really descriptive and clear so you immediate grasp what its doing from context alone". What is a descriptive function or variable name is extremely dependent on how familiar you are with the context the program functions in. Using `execute_network_commands_func`from above - this descriptive name say nothing about what network commands that are executed. With docstrings it would be so easy to detail input and output of this function.
- catlover76 3y agoThat's cool! I disagree though that the deal shows any bad judgment on Cisco's part; the gravamen of whether the acquisition was good is not whether many software developers can quickly develop replacements for their own use-cases, or how ergonomic the software is, but whether Splunk is a profitable business with a bunch of paying subscriptions/contracts that aren't going to go away any time soon.
- CliffStoll 3y agoOh, how I wish I had your scripts (and insights!) when I was analyzing Unix logs in 1986, looking for the footprints of an intruder...
- mercer 3y agoI was about to ask you to get around the campfire and tell the story again, but I see other commenters got ahead of me :). I'll be getting another Klein bottle soon for a gift, if you still do those :). Hope you're doing well!
- CliffStoll 3y agoMy smiles to you Mercer: it's fun to look back over my shoulder to a slowly vanishing time, when the Arpanet backbone ran at 4800 baud and a 1 megabyte Unix workstation was hot stuff...
- neilk 3y agoYou should write about that sometime! /s
- h0p3 3y agoI'm kinda glad you didn't; it might have made the book I read as a kid (and again as an adult, and again with my offspring) less interesting somehow.
- CliffStoll 3y agoUh, yes, h0p3 ... I didn't exactly start on that adventure thinking I'd write a book. Chasing after those hackers was orthogonal to my work in astronomy and the Keck telescope.
- h0p3 3y agoYes, sir. I appreciate that. I think you made that very clear in the book as well. I'll agree that having OP's tooling back then would likely have been quite useful to you and others. I'm often terrible with words. What I meant to say was: thank you for writing the book. Your story has been an important part of my family's lives for three generations (your work is also mentioned 3 times in my ℍ𝕪𝕡𝕖𝕣𝔱𝔢𝔵𝔱, and, prominently in my record of reaching out to others out of the blue [I've a habit of knocking on doors with low success rates]). Never thought I'd have the chance to say that to you. =D. `/salute`. Thank you, sir.
- xorcist 3y agoFetching logs regularly sounds hard? Wouldn't you need to keep track of the position of all files, with heuristics around file rotations? And if something catastrophic happens, the most interesting data would be in that last block which couldn't be polled? Normally you'd avoid all that complexity by shipping logs the other way, sending from each machine. That way you can keep state locally should you need to. All unix-like systems do this out of the box, and almost all software supports the syslog protocol to directly stream logs. But you can also use something like filebeat and a bunch of other modern alternatives. The analyzer can then run locally on the log server and a whole lot of complexity just disappears.
- eigenvalue 3y agoI considered doing it the way you described, but then you need to deploy software on every single one of your machines and make sure it's running, that it's not accidentally using up 99% of your CPU (I've had bad experiences with the monitoring agents for Splunk and Netdata misbehaving and slowing down the machines and causing problems), etc. Whereas with the "pull" approach I used in my tool, you don't need to deploy ANY software to the machines you are monitoring-- you just connect with SSH and grab the files you need and do all the work on your control node.
- xorcist 3y agoOne way or the other, your hosts are running your application, and you are already deploying software on every single host. But I hear you with some of the agents. That's why I mentioned syslog. It's already there, it's supported by most logging packages, and it's dead simple. No additional software required. All text. What it doesn't do is structured logging, but analyzing on the log host is often enough. Agents aren't all that bad however, and you're likely already running some agent like icinga or zabbix for regular monitoring.
- hjvalente 3y agohaving a Netdata agent taking your machine's CPU to 99% shouldn't happen, not sure when was the list time you tried it but a lot of recent improvements have been done on the Netdata Agent also, with Netdata you can achieve the same architecture design using a Netdata Parent that could be your "control node" and to where you stream the metrics of the nodes you want to keep running with as less load as possible - you can even offload the health engine and the machine learning take a look at https://learn.netdata.cloud/docs/streaming/ https://learn.netdata.cloud/docs/streaming/ and https://learn.netdata.cloud/docs/configuring/how-to-optimize-the-netdata-agent-s-performance https://learn.netdata.cloud/docs/configuring/how-to-optimize...
- simonw 3y agoI love this! Log analysis isn't one of the core use-cases for Datasette, but I've done my own experiments with it that have worked pretty well - anything up to 10GB or so of data is likely to work just fine if you pipe it into SQLite, and you could go a lot larger than that with a bit of tuning. I added some features to my sqlite-utils CLI tool a while back to help with log ingestion as well: https://simonwillison.net/2022/Jan/11/sqlite-utils/ https://simonwillison.net/2022/Jan/11/sqlite-utils/
- dingdong33 3y ago[flagged]
- runjake 3y agoNeat! Definitely a better solution for single source logs. Splunk is ridiculous and Cisco acquiring it isn't going to make that better. For others with a bit more complex needs, take a look at the free (or paid) versions of Graylog Open[1]. It's really improved over the years. I had messed with Graylog in it's early days but was turned off by it. A few years back, I encountered someone doing some neat stuff with it. It looked much improved. I stood up a "pilot project" to test, and it's now been running for years and several different people use it for their areas of responsibility. It does log collection/transforming and graphing and dashboarding and we use the everloving crap out of it at work. I wish I could publicly post some of the stuff we're doing with it. It takes input from just about any source. 1. https://graylog.org/products/source-available/ https://graylog.org/products/source-available/
- languagehacker 3y agoAh, the hubris of the single developer who believes they can replace a battle-tested product from a company with innumerable decades of combined human effort. Glad it works for you!
- cooper_ganglia 3y agoI mean, it seems like it's working for them. Not every single startup needs the same solutions as a larger company, especially when the solution is as expensive as Splunk!
- occamsrazorwit 3y agoThis is kinda ironic. The founders of Splunk originally created the product, because they realized that every sysadmin used their own, single-file script to analyze logs. They cleaned the scripts up and productionized what emerged. The reason Splunk grew to a billion-dollar company in the first place was that those sysadmins preferred to switch over to something that was more enterprise-grade. Life is cyclical.
- prabhatsharma 3y agoSo it requires redis. Would have loved if it was just a simple script or binary.