7 ms·
Show HN: A Better Log Service
Hello everyone, there are many log services available and this is my attempt at a better one.
Most online logging tools feature convoluted UIs, arbitrary mandatory fields, questionable AI/insights, complex pricing, etc. I hope my application fixes most of these issues. It also has some nice features, such as automatic Geo IP checks and public dashboards.
Although I've created lots of software, this is my first open source application (MIT license), the tutorial for selfhosting is hopefully sufficient! Most of my development career has been with C#, NodeJS and PHP. For this project I've used PHP (8.3) which is an absolute joy to work with. The architecture is very scalable, but I've only tested up to a few billion logs. The current version is used in production for a few months now. Hope you enjoy/fork it as you see fit!
- that_guy_iain 2y agoThis looks very interesting! My suggestion for the self-hosting is to create docker images and use docker-compose. The self-hosting currently is a bit of effort to setup. I also wonder if PHP is a good language for this. For the UI, yea that's fine and makes sense. But for the log processor that's going to need to handle a high throughput which PHP just isn't good at. For the same resources, you can have Go doing thousands of requests per second vs PHP doing hundreds of requests per second.
- herbst 2y agoOn the other side some people (me) are happy to have an actual self hosting setup and not being forced to use a docker setup with unknown overhead.
- majkinetor 2y agoNo benefit using go over C#, IMO, and I am also baffled by the switch
- that_guy_iain 2y agoI just used go as an example, any compiled language would be better.
- withinboredom 2y agoI highly suspect it wouldn't be better in brainfuck...
- majkinetor 2y agoIt uses Clickhouse, though, which should be xtremelly fast for this.
- that_guy_iain 2y agoYes. But PHP still needs to process it before it goes to Clickhouse. PHP is the bottleneck.
- wiseowise 2y ago[flagged]
- axelthegerman 2y agoIf that "bottleneck" is thousands of requests per second then it doesn't really matter for smaller deployments does it? (Which seems to be the target audience and not FAANG) I'm not a big fan when folks call out languages as bottlenecks when they have no proof on the actual overhead and how much faster it would be in another language.
- that_guy_iain 2y agoTo tweak a PHP deployment to handle hundreds of requests per second which is very very realistic for a basic logging for a mid-sized application you're looking at having a very beefy server setup. Most PHP deployments barely reach a hundred per server. And this is an open source project is should be designed to handle basic production workloads which it could but it'll cost you a bunch more than if you used the correct languages. > I'm not a big fan when folks call out languages as bottlenecks when they have no proof on the actual overhead and how much faster it would be in another language. Honestly, I thought it was so obvious that an interpreted language is not good for high throughput endpoints that it didn't need to be proven. I also thought it was obvious that a logging system is going to handle lots and lots of data. It could be easily proven by doing a bunch of work but obviously there is no point in me proving it.
- 2y ago
- withinboredom 2y agoPHP is arguably the best solution here. If a log ingestion process breaks everything, no other logs are harmed (a default shared-nothing architecture). Using something like Go, C#, etc, it might be "faster" but less resilient -- or more complex to handle the resiliency. > But for the log processor that's going to need to handle a high throughput which PHP just isn't good at. I'm sorry, but wut? PHP is probably one of the fastest languages out there if you can ignore frameworks. It's backed by some of the most tuned C code out there and should be just about as fast as C for most tasks. The only reason it is not is due to the function call overhead -- which is by-far the slowest aspect of PHP. > you can have Go doing thousands of requests per second vs PHP doing hundreds of requests per second. This is mostly due to nginx and friends ... There is frankenphp (a frontend for php running in caddy which is written in go) which can easily handle 80k+ requests per second.
- that_guy_iain 2y agoI'm going to have to also reply with, sorry but what?! PHP is one of the fastest-interpreted languages. But compiled are going to be faster than interpreted pretty much everytime. It loses benchmarks against every language. That's not to mention it's slowed down by the fact it have to rebuild everything per request. As a PHP developer for 15+ years, I can tell you what PHP is good at and what PHP is not good at. High throughput API endpoints such as log ingestion are not a good fit for PHP. Your argument that if it breaks it's fine. Yea, who wants a log system that will only log some of your logs? No one. It's not mission critical but it's pretty important to keep working if you want to keep your system working. And in fact, some places it is a legal requirement.
- withinboredom 2y ago> It loses benchmarks against every language. Every language loses benchmarks against every other language. That's not surprising. Since you didn't provide a specific benchmark, it's hard to say why it lost. > High throughput API endpoints such as log ingestion are not a good fit for PHP. I disagree; but ultimately, it depends on how you're doing it. You can beat or exceed compiled languages in some cases. PHP allows some low-level stuff directly implemented in C and also the high-level stuff you're used to in interpreted languages.
- williebeek 2y agoThanks for the tip, I will check if inserting rows with Go is any faster. For reference, inserting a log takes three steps, first the log data is stored in a Redis Stream (memory), a number of logs are taken from the stream and saved to disk and finally inserted in batches in ClickHouse. I've created it so you can take the ClickHouse server offline without losing any data (it will be inserted later). For reference, moving about 4k logs from memory to disk takes less than 0.1 second. This is a real log from one of the webservers: Start new cron loop: 2024-12-18 08:11:16.397...stored 3818 rows in /var/www/txtlog/txtlog/tmp/txtlog.rows.2024-12-18_081116397_ES2gnY3fVc (0.0652 seconds). Storing this data in ClickHouse takes a bit more than 0.1 second: Start new cron loop: 2024-12-18 08:11:17.124...parsing file /var/www/txtlog/txtlog/tmp/txtlog.rows.2024-12-18_081116397_ES2gnY3fVc * Inserting 3818 row(s) on database server 1...0.137 seconds (approx. 3021.15 KB). * Removed /var/www/txtlog/txtlog/tmp/txtlog.rows.2024-12-18_081116397_ES2gnY3fVc As for Docker, I'm too much of a Docker noob but I appreciate the suggestion.
- ryanianian 2y agoPHP trivially scales up to multiple nodes behind an LB. You're really only limited by your backend storage connection count and throughput. Go and friends may make for more efficient resource utilization, but it will be marginal in the grand scheme of things unless there are plans to do massively different things. As it is this code is very simple. I haven't used PHP in 15 years and I was able to trace through this from front-end to back-end in less than 3 minutes. To me it look like a really great level of complexity for the problem it solves. Keep it up, OP.
- that_guy_iain 2y agoYou can but that costs more money... > Keep it up, OP. Live in the real world. No one wants to have a fleet of servers for their logging infra when there are options to run it on a single server.
- hipadev23 2y ago> PHP doing hundreds of requests per second. You may want to update your understanding of PHP and Go's speed . Both of your estimates are off by a couple orders of magnitude on commodity hardware. There are also numerous ways to make PHP extremely fast today (e.g. swoole, ngx_php, or frankenphp) instead of the 1999 best practice of apache with mod_php. Go is absolutely an excellent choice, but your opinion on PHP is quite dated. Here are benchmarks for numerous Go (green) and PHP (blue) web frameworks: https://www.techempower.com/benchmarks/#hw=ph&test=fortune§ion=data-r22&l=zijnjz-cn3 https://www.techempower.com/benchmarks/#hw=ph&test=fortune&s...
- p_ing 2y agoAs soon as you add C#, ASP.NET Core shoots to the top of the Fortune stack.
- deleted 2y ago[deleted]
- that_guy_iain 2y agoWhat you're talking about is generally not considered production-ready. While you can use these tools you will almost certainly run into problems. I know this because as an active PHP developer for over a decade I'm very much paying attention to that field of PHP. What we see here is a classic case of benchmarks saying one thing when the reality of production code says something else. Also, I used go as a generic example of compiled languages. But what we see is production-grade Go languages outperforming non-production-ready experimental PHP tooling. And if we go to look at all of them https://www.techempower.com/benchmarks/#hw=ph&test=fortune§ion=data-r22 https://www.techempower.com/benchmarks/#hw=ph&test=fortune&s... We'll see that even the experimental PHP solution is 43 and being beat out by compiled languages.
- hipadev23 2y agoNobody is suggesting PHP beats compiled. We’re arguing with you about your utter lack of expertise in the language, knowledge of the ecosystem and “production-ready” status of the many options, and your overall coding ability when it comes to PHP.
- majkinetor 2y agoWe need glorified (rip)grep instead of ELKS and friends, which have huge learning curve. I welcome this effort.
- remram 2y agoI'm pretty satisfied with Loki. It just ingests the logs and offers a powerful query language to extract data at query time (e.g. parse JSON, run regexes, plot). It can store data in a local folder or S3-compatible storage. I also gave up on configuring ELK in the past...
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- valyala 2y agoIf you like grep, you'll like https://docs.victoriametrics.com/victorialogs/querying/#command-line https://docs.victoriametrics.com/victorialogs/querying/#comm...
- bdcravens 2y agoVery nice. A lot of the complexity you described is why I've settled on using CloudWatch logs for anything I have on AWS. I don't need a fancy UI, just a powerful querying language for investigation and debugging. With that said, it would be nice to see at least some mechanism for building aggregates queries (for example, 4* results in the last 24 hours by user) but if it's ClickHouse underneath, I assume that's easy using standard ClickHouse tools.
- stingraycharles 2y agoI hate how Cloudwatch itself is so fragmented, and they have three different query languages for logs. It’s all cognitive overhead I don’t want to learn.
- bdcravens 2y agoI will say that the language isn't the most intuitive, and a project like this one with some simply querying with the (presumed) ability to drop down to SQL for power use is probably the ideal solution. (Doable with CloudWatch logs and Athena, but that's another can of complex worms)
- infecto 2y agoI would be happy to pay a premium for a better cloudwatch. For me it is always not intuitive which I am sure is driven by limited use.
- deleted 2y ago[deleted]
- drchaim 2y agoDon't tell me why, but I've developed an instinct that recognizes solutions that use Clickhouse under the hood :)
- HatchedLake721 2y agoTell us more!
- williebeek 2y agoIt uses both, MySQL for the metadata and ClickHouse for the logs. The selfhost page explains a bit more about the architecture. edit: the connection to ClickHouse uses the MySQL driver, this is actually a very nice CH feature, you can connect to CH using the regular mysql or postgresql client tools. The PHP MySQL PDO driver works seamlessly. One catch, using advanced features like CH query timeouts requires a CTE function, check the model/txtlogrowdb.php file if you're interested.
- stingraycharles 2y agoSeems to be using MySQL instead? https://github.com/WillieBeek/txtlog/blob/master/txtlog/database/db.php https://github.com/WillieBeek/txtlog/blob/master/txtlog/data...
- deleted 2y ago[deleted]
- dobin 2y agoPretty unrelated, but i like how it displays large amount of potentially diverse JSON events. Would need some better filtering and sorting, hiding of keys etc. Products which do this well are Elastic and Splunk, but are too heavy for my taste.
- szundi 2y agoI always played with the idea that the logs could be viewed as packets of some protocol and use wireshark to filter them and view related logs as a “stream” like view that wireshark provides
- rednafi 2y agoThis is nice. At work, we use Datadog for logging, and I have previously used CloudWatch, Splunk, and Honeycomb. Among these, only Honeycomb makes implementing canonical log lines [1] easier. I want arbitrarily wide, structured logs [2] without paying exorbitant costs for cardinality. Our Datadog costs are outrageous, and it seems like no one cares at this point. Pydantic Logfire is also doing some good work in Python-specific environments. I use both Python and Go, but Logfire wasn’t as ergonomic in Go. [1]: https://stripe.com/blog/canonical-log-lines https://stripe.com/blog/canonical-log-lines [2]: https://www.honeycomb.io/blog/structured-events-basis-observability https://www.honeycomb.io/blog/structured-events-basis-observ...
- deleted 2y ago[deleted]
- valyala 2y agoTry VictoriaLogs [1]. It supports wide events with hundreds of fields (aka canonical logs). [1] https://docs.victoriametrics.com/victorialogs/ https://docs.victoriametrics.com/victorialogs/
- reacharavindh 2y agoMy current log solution that is based on Clickhouse I’m tinkering with in free time in Victorialogs. https://docs.victoriametrics.com/victorialogs/ https://docs.victoriametrics.com/victorialogs/
- deleted 2y ago[deleted]
- mooreds 2y agoI've heard good things about Axiom[0], especially for high scale needs. 0: https://axiom.co/ https://axiom.co/
- deleted 2y ago[deleted]
- mdaniel 2y agoIf you like them, please submit the link on its own, and not to take away from someone's MIT "Show HN" to plug a non open source project
- mooreds 2y agoFair enough. Thanks for the feedback.
- theogravity 2y agoI added support for it just now because of this comment. Had it on my list of things to integrate with since I saw the comment. Thanks! https://loglayer.dev/transports/axiom.html https://loglayer.dev/transports/axiom.html
- mdaniel 2y agoWhat in the world does this mean? https://txtlog.net/doc#:~:text=use%20your%20local%20time%20when%20inserting%20logs https://txtlog.net/doc#:~:text=use%20your%20local%20time%20w... That's made twice as bad by the "we throw away Z because you were just kidding by including it". That leads me to believe that any RFC 3339 that isn't automatically Z (e.g. 1996-12-19T16:39:57-08:00 <https://datatracker.ietf.org/doc/html/rfc3339#section-5.8 https://datatracker.ietf.org/doc/html/rfc3339#section-5.8>) is ... well, I don't know what it's going to do but it likely won't be good It also appears that your documentation is currently a very verbose version of an OpenAPI spec, so you may save your readers some trouble by actually publishing one, with the added advantage that they come with a "Try it" button in the OpenAPI renders That would allow you to save the natural language parts for describing things that are not API-centric (such as the "but WWWWHHHHYYY mysql AND clickhouse" that you alluded to elsewhere but wasn't mentioned at all in /doc nor /selfhost)
- tyingq 2y agoThe date treatment isn't great, but the repo seems to indicate it's existed as a public thing for 22 days. So perhaps just an early compromise to get it working.
- mdaniel 2y agoFor all the folks championing how awesome PHP is in this thread, one would surely hope it has rfc3339 aware date parsing, no? But I guess that <https://www.php.net/manual-lookup.php?pattern=rfc%203339&scope=quickref https://www.php.net/manual-lookup.php?pattern=rfc%203339&sco...> and <https://www.php.net/manual-lookup.php?pattern=iso8601&scope=quickref https://www.php.net/manual-lookup.php?pattern=iso8601&scope=...> both being :shruggle: doesn't do it any favors. However, it seems it is just a search stupidity because https://www.php.net/manual/en/datetimeimmutable.createfromformat.php#:~:text=datetimeinterface%3A%3Aiso8601 https://www.php.net/manual/en/datetimeimmutable.createfromfo... I do love this, since it 100% squares with my mental model of PHP's approach to life: you're holding it wrong https://www.php.net/manual/en/function.date-parse-from-format.php#129810 https://www.php.net/manual/en/function.date-parse-from-forma...
- adriand 2y agoI’m curious about the open source nature of this and how you / people in general manage a project where you are hosting it and need to maintain its security, but are also presumably merging pull requests as people contribute to the project. I would be quite paranoid about this, ie concerned that someone might slip a line of code in with the intent of breaching the service that I would not catch during code review. I know this is true of any open source project but it feels especially fraught when you are also hosting it and letting people sign up and pay for it. I’m wondering if you or others have experience with this and what approaches and practices mitigate this risk.
- gabeio 2y agoJust because a project is “open source” doesn’t actually mean you must accept or even merge PRs from others. After reading others pointing this out my opinion of managing open source projects have significantly changed. Of course, you can entertain PRs and see if the idea behind them is sound but not accept the raw code from others and implement the features they way you envision instead. Keep in mind it’s always possible to have a vulnerability without anyone else’s assistance. This is especially true if you use dependencies, as you don’t keep track of every line of code they add.
- withinboredom 2y ago> This is especially true if you use dependencies, as you don’t keep track of every line of code they add. You absolutely should vendor your dependencies and review them before accepting the new version. Even though they are dependencies, you are ultimately responsible for using them. "They are just dependencies" doesn't absolve you of responsibility.
- dlln 2y agoGreat points about dependencies and reviewing PRs. In addition to manual reviews, layering security tools within your CI/CD pipeline is key. Tools like static code analyzers, dependency scanners, and security linters help catch vulnerabilities early. Open source can also be a valuable way to uncover security gaps, but having a secure channel for reporting vulnerabilities is crucial to address them quickly. Leveraging techniques like Content Security Policies (CSPs) adds extra layers of protection, promoting proactive security throughout development and deployment.
- hk1337 2y agoIt's a minor thing but I would remove the jQuery dependency. You're not doing much with that plain javascript couldn't do just as well if not better. Plain JS has come a long way since jQuery first came out.
- nesarkvechnep 2y agoSome people praised Go as a better language for the use case than PHP. I’d say Elixir is even better. It can handle massive concurrency easy, can be made distributed easy, has a built-in, in-memory, key-value store (ETS), and is probably the best high-level language for anything that’s facing the network.
- lukevp 2y agoI've really been interested in learning more about Elixir and how it accomplishes these things, because I constantly hear the same opinions from others. Do you have some good resources you'd recommend for getting started with Elixir for a principal engineer that wants to understand these at-scale issues and how Elixir solves them better than other languages?
- nesarkvechnep 2y agoYes, two books. To get a feel for the language - “Elixir in Action” by Sasa Juric. To discover how Elixir and the platform it’s built on excel in scalability and fault-tolerance - “Designing for Scalability with Erlang/OTP” by Francesco Cesarini.
- thomquaid 2y agothe 'easy to use' / 'view' was very nice. if you could add the actual session logs in it would be amazing.
- TripleChecker 2y agoIt looks like that's a PHP codebase. I'm curious why one should use this solution instead of more performant Go/Rust log backends? Also, one of the login links takes you to a 404 page: https://triplechecker.com/s/jDTmQa/txtlog.net https://triplechecker.com/s/jDTmQa/txtlog.net
- giraffe_lady 2y agoThey said > Most of my development career has been with C#, NodeJS and PHP and then > The architecture is very scalable, but I've only tested up to a few billion logs.
- deleted 2y ago[deleted]
- theogravity 2y ago[flagged]
- deleted 2y ago[deleted]
- piterrro 2y ago> there are many log services available and this is my attempt at a better one. Out of curiosity, can you describe how your service is better than others? >I hope my application fixes most of these issues Do you care to elaborate on the "how"?
- juanisimo 2y ago[flagged]