10 ms·
Show HN: Log collector that runs on a $4 VPS
Hey guys, I'm building erlog to try and solve problems with logging. While trying to add logs to my application, I couldn't find any lightweight log platform which was easy to set up without adding tons of dependencies to my code, or configuring 10,000 files.
ErLog is just a simple go web server which batch inserts json logs into an sqlite3 server. Through tuning sqlite3 and batching inserts, I find I can get around 8k log insertions/sec which is fast enough for small projects.
This is just an MVP, and I plan to add more features once I talk to users. If anyone has any problems with logging, feel free to leave a comment and I'd love to help you out.
- unxdfa 4y agoI see your idea but you could drop the JSON and use rsyslogd + logrotate + grep? You can grep 10 gig files on a $5 VPS easily and quickly! I can't speak for a $4 one ;)
- tiagod 4y agoIf you use grep you'll be doing the same expensive operation every time, following files naively will fail after rotation, etc... And if you use something like Loki it's easier to integrate with other tools to react to the logs
- unxdfa 4y agoIt’s potentially a premature optimisation to not do that expensive operation every time. Loki and brethren have a significant infrastructure cost and cognitive load to consider. I speak from experience and know where the ROI appears and it’s far from the use case specified here.
- ilyt 4y ago> following files naively will fail after rotation, ...so what you're saying they have to write "tail -F" instead of "tail". > If you use grep you'll be doing the same expensive operation every time if you have ingest that low it barely matters. Modern grep replacements are pretty fast
- mekster 4y agoWhy do people like to stick to inefficient ancient method like grep for log viewing? Try tools like Metabase and see how it makes your log reading far better.
- Hamuko 4y agoI feel like if you're going to use "$4 VPS" as a quantifier, you could at least specify which $4 VPS is being used.
- teruakohatu 4y agoDO's 512mb basic VPS starts at $4, so I am guessing it is that.
- benatkin 4y agoI don't think it is. That one is shared vCPU and I've been hearing about a single vCPU one.
- sgt 4y agoLook at this: https://www.hetzner.com/cloud https://www.hetzner.com/cloud More like $5 but still, 1 vCPU, 2GB RAM, 20GB NVMe storage. Closer to $4 USD if you let go of IPv4 in favor of IPv6 only. Edit: Looks like that's also a shared vCPU.
- harisamin 4y agoAh cool! Somewhat related I built a json log query tool recently using rust and SQLite. Didn’t build the server part of it https://github.com/hamin/jlq https://github.com/hamin/jlq
- Nevin1901 4y agoThat's really cool. I might rewrite some of your code in go and use it in erlog for searching (and give you credit of course). How did you come up with the idea for jlq? It seems like it solved a pretty cool use case.
- keroro 4y agoIf anyones looking for similar services Im using vector.dev to move logs around & it works great & has a ton of sources/destinations pre-configured.
- Dachande663 4y agoI’ve found the hard part is not so much the collection of logs (especially at this scale), but the eventual querying. If you’ve got an unknown set of fields been logged, queries very quickly devolve into lots of slow table scans or needing materialised views that start hampering your ingest rate. I settled on a happy/ok midpoint recently whereby I dump logs in a redis queue using filebeat as it’s very simple. Then have a really simple queue consumer that dumps the logs into clickhouse using a schema Uber detailed (split keys and values), so queries can be pretty quick even over arbitrary fields. 30,00 logs an hour and I can normally search for anything in under a second.
- metadat 4y agoWhat are the hardware requirements/ / resources dedicated to pull this off?
- Dachande663 4y agoRuns off a 2GB digital ocean box, which I think is $10 now? It’s probably incredibly boring to describe, but I think that’s why it just tends to work. The whole thing took an afternoon to write (in PHP of all things too).
- simonw 4y agoThis genuinely sounds the opposite of boring to me. I'd love to read a full, detailed description of this, hacky PHP scripts included!
- Dachande663 4y agoI've added some more info https://news.ycombinator.com/item?id=34771486 https://news.ycombinator.com/item?id=34771486
- mr-karan 4y agoI've a similar pipeline to yours (for the storage part). I use vector.dev for collecting and aggregating logs, enriching with metadata (cluster, env, AWS tags) and then finally storing them in a Clickhouse table. Do you use any particular UI/Frontend tool for querying these logs?
- withinboredom 4y agoNeat! Have you considered using query params instead of bodies, then just piping the access logs to a spool (no program actually on the server, just return an empty file). Then your program can just read from the spool and dump them into sqlite. That should tremendously improve throughput, at the expense of some latency.
- Nevin1901 4y agoThat's a really good idea, thanks for suggesting it. I'll try implementing it. I'm hoping the main bottleneck is with inserting the logs into SQLite, so using a spool might help
- folmar 4y agoSorry, but I don't see the selling point yet. Rsyslog has omlibdbi module that send your data to sqlite. It can consume pretty much any standard protocol on input, is already available and battle proven.
- aninteger 4y agoI'm doing something similar with a $5 VPS, but with fastcgi/c++/sqlite3. I then have a cronjob that then aggregates error logs, generates an summary and posts to a Slack channel. Personally I wish I didn't have to write it, but it works.
- Nevin1901 4y agoOne of my eventual goals with erlog is actually doing observability (eg: it'll send you reports if logs/metrics deviate from the norm), so it's really interesting to see you had this problem.
- sgt 4y agoImagine what we could do with modern hardware if programs were as efficient as your typical C++/SQLite combo!
- marcrosoft 4y agoWoah cool. I did the same thing. I Made a poor man’s small scale splunk replacement with SQLite json and go. I used the built in json and full text search extensions.
- Thaxll 4y agoYou could have just used Filebeat? It's also in Go and it's pretty easy to use. https://www.elastic.co/guide/en/beats/filebeat/current/filebeat-input-httpjson.html https://www.elastic.co/guide/en/beats/filebeat/current/fileb...
- mekster 4y agoI think Vector really shines with its VRL language to parse and enrich data. It's well thought out with buffering for network errors and throwing errors on parsing instead of silently discarding. https://vector.dev/docs/reference/vrl/ https://vector.dev/docs/reference/vrl/
- cnkk 4y agoI am been using vector.dev for a long time now. It is also easy to setup. And it looks similar to your idea.
- rsdbdr203 4y agoThis is exactly why I build log-store. Can easily handle 60k logs/sec, but I think more importantly is the query interface. Commands to help you extract value from your logs, including custom commands written in Python. Free through '23 is my motto... Just a solo founder looking for feedback.
- recck 4y agoI came across this a few months ago and have been following pretty closely. Having been using this locally in a Docker container has been painless. The UI is definitely iterating quickly, but the time-to-first-log was impressive! Happy to keep using it.
- binwiederhier 4y agolog-store [1] is pretty neat. Thanks for making it. It's super powerful and easy to use. There's a learning curve with the query language, but it's super cool once you figure it out. [1] https://log-store.com/ https://log-store.com/
- spsesk117 4y agoDisclaimer: I am friends with the founder of log-store. I have been beta testing it for a while for small scale (~50 million non-nested json objects) log aggregation it's working beautifully for this case. It's a no nonsense solution that is seemless to integrate and operate. On the ops side, it's painless to setup, maintain, and push logs to. On the user side, its extremely fast and straight forward. End users are not fumbling their way through a monster UI like Kibana, access to information they need is straight forward and uncluttered. I can't speak to it's suitability in a 1TB logs/day situation, but for a small scale straight forward log agg. tool I can't recommend it enough.
- andymac4182 4y agoI have been using https://datalust.co/ https://datalust.co/ to handle this. It scales really well down and up to how much you want to spend. It comes with existing integrations with a lot of libraries and formats and a CLI to push data from file based logs to their service. They have just added a new parser and query engine written in Rust to get the best performance out of your instance. https://news.ycombinator.com/item?id=34758674 https://news.ycombinator.com/item?id=34758674
- peterpost2 4y agoSecond this, seq is incredibly handy and easy to query. Performance could be better though
- maybesimpler 4y agoYou could also not write your own server. Just configure OpenResty, write some simple LUA to push to the redis queue. Then consume the queue via your language of choice to write to your store(clickhouse).
- remram 4y agoMay be more widely applicable for personal servers: lnav, an advanced log file viewer for the terminal: https://lnav.org/ https://lnav.org/ It uses SQLite internally but can parse log files in many formats on the fly. C++, BSD license, discussed 1 month ago: https://news.ycombinator.com/item?id=34243520 https://news.ycombinator.com/item?id=34243520
- Weryj 4y agoI run a self hosted version of Sentry.io on a NUC at home and a relay on a VPS, the. Use Tailscale to connect the two. If you have an old computer at home, using a VPS as the gateway is always a good option. Edit: you can then use the VPS as a exit node for internet.
- arjvik 4y agoI'm working on a project where I'm handling simultaneous connections to a bunch of peers. What's the best way to log messages to trace the flow of requests through my system when multiple code paths are running asynchronously (NodeJS, so I can't simply get a thread ID)?
- Groxx 4y agoWith tracing libraries, e.g. https://opentracing.io/ https://opentracing.io/
- viraptor 4y agoThat and some backend... most are SaaS though. The only self hosted one I know of is Grafana Traces/Tempo. I mean, you can just log the trace/span/parent IDs for each request, but that's a bit painful to deal with later.
- ilyt 4y agojaeger + opentracing/opentelemetry libs are easy enough. When testing Jaeger can just work as in-memory database, or put it to some other storage like elasticsearch
- vbezhenar 4y agoLogs must be stored in S3, it's no-brainer. Disk storage is too expensive. Logging system should be designed for S3 from ground up IMO.
- addandsubtract 4y agoHow is S3 the cheapest option? Backblaze is $0.005 per GB and Hetzner sells storage boxes for less than €0.0038 per GB.
- deleted 4y ago[deleted]
- aejnsn 4y agoReplace what “S3” with “S3-compatible object storage”.
- zo1 4y agoAWS is almost never the answer unless you work at making it cheap and work for you. It's a black hole of insanely convoluted billing and filled with snake oil salesmen (dev ops priests) that'll complicate it so much that you don't even know what you're paying for or why you even need it. And S3 is just a gateway drug into this whole mess, stay far away from it kids.
- vbezhenar 4y agoI'm talking about S3 API. Every cloud I'm aware of, provides S3-compatible object storage. And this storage is much cheaper than VM-attached storage.
- Jamie9912 4y agoI think he was being sarcastic
- vbezhenar 4y agoI'm talking about storage volumes which are attached to the VM versus object storage which provides S3-compatible API. You can use S3 API to access Backblaze. I'm not experienced with Hetzner Storage Box but I don't think that you can just attach it to your storage VM as fast storage. You can mount it as samba store but I think that's a recipe for disaster.
- int0x2e 4y agoI strongly urge people to try something like Application Insights. It's not dirt cheap, but not that expensive, and lets you collect anything you'd want and query your telemetry/logs retroactively extremely flexibly. It's just great.
- ilyt 4y ago...uh, just rsyslog and files ? I think it can even write to SQLite