11 ms·
Show HN: Open-source APM with support for tracing, metrics, and logs
Uptrace is an all-in-one tool that supports distributed tracing, metrics, and logs. It uses OpenTelelemetry observability framework to collect data and ClickHouse database to store it.
You can ingest data using OpenTelemetry Protocol (OTLP), Vector Logs, and Zipkin API. You can also use OpenTelemetry Collector to collect Prometheus metrics or receive data from Jaeger, X Ray, Apache, PostgreSQL, MySQL and many more.
The latest Uptrace release introduces support for OpenTelemetry Metrics which includes:
- User interface to build table-based and grid-based dashboards.
- Pre-built dashboard templates for Golang, Redis, PostgreSQL, MySQL, and host metrics.
- Metrics monitoring aka alerting rules inspired by Prometheus.
- Notifications via email/Slack/PagerDuty using AlertManager integration.
There are 2 quick ways to try Uptrace:
- Using the Docker container - https://github.com/uptrace/uptrace/tree/master/example/docker https://github.com/uptrace/uptrace/tree/master/example/docke...
- Using the public demo - https://app.uptrace.dev/play https://app.uptrace.dev/play
I will be happy to answer your questions in the comments.
- wdb 4y agoHow does it compare to Opstrace? (www.opstrace.com)
- PeterZaitsev 4y agoFalse Advertising! BSL Licensed is not Open Source. To Be fair Utrace restrictions are relatively light but it is still Source Available project not Open Source
- ram_rar 4y agocan you elaborate more on why clickhouse for backend? And what challenges if any are you facing with clickhouse?
- vmihailenco 4y agoYou can ingest data using OpenTelemetry Protocol (OTLP), Vector Logs, and Zipkin API. You can also use OpenTelemetry Collector to collect Prometheus metrics or receive data from Jaeger, X Ray, Apache, PostgreSQL, MySQL and many more. The latest Uptrace release introduces support for OpenTelemetry Metrics which includes: - User interface to build table-based and grid-based dashboards. - Pre-built dashboard templates for Golang, Redis, PostgreSQL, MySQL, and host metrics. - Metrics monitoring aka alerting rules inspired by Prometheus. - Notifications via email/Slack/PagerDuty using AlertManager integration. There are 2 quick ways to try Uptrace: - Using the Docker container - https://github.com/uptrace/uptrace/tree/master/example/docker https://github.com/uptrace/uptrace/tree/master/example/docke... - Using the public demo - https://app.uptrace.dev/play https://app.uptrace.dev/play I will be happy to answer your questions in the comments.
- 0JzW 4y agothis looks amazing! i would definitely like to use this for log monitoring. however, i have a question. is it possible to get logs for individual docker containers?
- vmihailenco 4y agoIt is possible using Vector Logs which Uptrace supports out-of-the-box, for example: - https://vector.dev/docs/reference/configuration/sources/docker_logs/ https://vector.dev/docs/reference/configuration/sources/dock... - https://uptrace.dev/get/ingest/vector.html https://uptrace.dev/get/ingest/vector.html If you are having troubles making it work, feel free to open an issue on Github and I will provide a complete example.
- nik736 4y agoNice! Exactly what I've been looking for, will give it a try for sure. Sentry eats a lot of resources so I was looking for an alternative.
- vmihailenco 4y agoThanks! Don't hesitate to send any feedback you have so we have a chance to improve :)
- derN3rd 4y agoNice to see so many new projects in the area of APM in the last few months. We recently tried Signoz and Grafana Tempo and while I can't say something about uptrace yet (will definitely try it out) I want to list some pros and cons about them. Grafana Tempo Pros: - Easy and smooth integration into our existing Grafana instance, no additional frontend needed - No new storage engine needed (No additional Clickhouse, Postgres, etc) as it saves its data to S3 - Supports OTLP Cons: - Search is limited by param size and unique params (as its baked to be indexed) - Ingestion is not in real time, but configurable (time to finish span) Signoz: Pros: - Supports OTLP - Integrates Logs and Metrics within the same service (for Grafana you need Loki then) - Supports real time querying Cons: - Uses new storage engines (or extends the software stack) with adding ClickHouse - Adds an additional frontend (might not be relevant for everyone) - Doesn't provide SSO yet, so you need to manage users differently Interesting to see, that UpTrace also chose ClickHouse (btw I love ClickHouse!) Some questions: - Can I easily disable certain features? (e.g. alerting) - Is there support for SSO for self-hosted installation? - Are there any recommendations for scaling (e.g. benchmarks) on how many spans/s are supported on what hardware? Thanks in advance!
- pranay01 4y agothanks for the mention. I am one of the maintainers at SigNoz [1]. Thanks for laying out the points in Pro section. We also recently launched logs witg v0.11.0 so you may want to give it a try again - we now have have metrics, logs and traces in a single app. Would love to understand a few points in more details you have mentioned in Cons for SigNoz > - Uses new storage engines (or extends the software stack) with adding ClickHouse Can you explain a bit more on the concern here? > - Doesn't provide SSO yet, so you need to manage users differently This is in our roadmap and we will be shipping it soon. [1] https://github.com/SigNoz/signoz https://github.com/SigNoz/signoz
- vmihailenco 4y agoThanks for the feedback! >- Are there any recommendations for scaling (e.g. benchmarks) on how many spans/s are supported on what hardware? With Uptrace, I was able to achieve 10k spans / second on a single core by running the binary with GOMAXPROCS=1. That is 1-3 terrabytes of compressed data each month which is more than most users need. Practically, you are limited by the $$$ you are willing to spend on ClickHouse servers, not by Uptrace ingestion speed. So my recommendation is to scale Uptrace vertically by throwing more cores at it. That will allow you to go very very far. >- Is there support for SSO for self-hosted installation? So far the only way to add news users is via the YAML config. We are considering to add a REST API or a CLI tool for the same purpose, but it is not clear how that would work with the YAML. Regarding the SSO, it would be nice if you can provide an app that already does that so we can better estimate the complexity. But so far we don't have such plans. >- Can I easily disable certain features? (e.g. alerting) Yes, most YAML sections just can be removed / commented out to disable the feature.
- edf13 4y agoLooks nice... I'm a bit out of touch in this space but my last solution for similar would be Datadog. How does this compare?
- vmihailenco 4y agoDataDog has a high learning curve and can be rather expensive if you need to monitor a lot of hosts and microservices. Uptrace tries hard to stay simple while providing almost the same set of features. It can also be self-hosted without paying anything which can save a lot of money. Uptrace aims to be an open source alternative to DataDog, but realistically we are not there yet.
- jonasdevops 4y agoBefore you try, please make sure you are comfortable with their license - https://github.com/uptrace/uptrace/blob/master/LICENSE https://github.com/uptrace/uptrace/blob/master/LICENSE (Business Source License 1.1), which as License says "The Business Source License (this document, or the “License”) is not an Open Source license"
- vmihailenco 4y agoIt is the same license used by MariaDB, Sentry, CockroachDB, Couchbase and many others. Technically it is not open source and instead is called source available, but you can enjoy pretty much the same benefits. Out of curiosity, what makes you uncomfortable about the license?
- fmajid 4y agoEr, MariaDB is GPL2, and will forever be since it is derived from MySQL. I'm guessing they aligned the license terms on those of ClickHouse, which is the underlying data store for Uptrace. From my understanding, if you use Uptrace and ClickHouse to manage your internal telemetry and don't offer it to clients, you should be fine. Still, non-standard license terms give pause, as there is always the possibility they will be restricted further in a bait-and-switch operation like that done by MongoDB or ElasticSearch.
- vmihailenco 4y ago>the possibility they will be restricted further It is true for all licenses, for example, it is possible to keep old code available under old permissive license but release all new code under a more restricted license. None is safe! :)
- jillesvangurp 4y agoThis not correct. There's a difference between OSS projects with shared copyright (every individual contributor holds the copyright to their contributions) and oss projects where the copyright is required to be transferred to some company. In the second case, this company then holds the copyright to the entire source tree and can re-license it at will. Mongo and Elastic did this and they are good examples of why you should not transfer copyright because they then have the right to re-license your contributions as they please. That's the reason they insist on copyright transfers. They reserve the right to change the license on future versions of the software. You can still use the old versions under the license that applied at the time. So, Opensearch is a fork of the last Apache 2.0 licensed version of Elasticsearch. Versions after that are licensed under a non OSS license whereas Opensearch is a proper open source project where copyright belongs to individual contributors, which ironically is still mostly Elastic plus whatever individual opensearch contributors added. Most projects don't insist on copyright transfers however and given enough external contributors it becomes increasingly hard for them to get permission to change he license. Regardless of the license, there is absolutely zero chance of something like mysql, linux, or other long existing OSS projects ever being re-licensed because it would require tracing down tens of thousands of copyright holders (or their surviving relatives) to get permission for that most of whom would probably not be willing to do that. This is so impractical that it will never happen. And even if it happened, anyone could continue using and contributing under the old license. All you'd have is a fork that is cut off from those contributions (because licenses like Gpl v2 don't allow mixing with proprietary code). So given a permission you will never get, you'd have a fork that is effectively yours of an original source tree that still belongs to all the original contributors and is licensed under the original license. OSS done properly builds communities that exist for as long as people continue to be willing to use and contribute to the software. Some OSS projects are now decades old.
- tmd83 4y agoI wonder if anyone can answer some question on distributed tracing for me. The difference between old days of APM vs. tracing as I understand is two things. 1. Originally APM was single process and it was language aware, usually do sampling stacktrace to find where times are being taken and some very well know place to instrument for exact timing say response time or query time. Tracers are more working by instrumenting methods of framework/servers/runtime at well known point and getting the timing. In man ways it's a lot more coarse as it might know of a hot loop that I have in my code. But it can trace very well with exact timing at framework boundary like web, cache, db etc. 2. The APM were primarily single process and couldn't really show a different service/process which doesn't work in a micro-service/distributed world. The way I understand it is that Tracers would allow me to narrow down to the service/component very easily. Whether I can find out why that component is slow might not be as easy (not sure what granularity tracing happens inside a component). I wonder if this understanding of mine is correct. The second thing I am really unsure about is sampling and overhead. What's the usual overhead of a single tracing (I know it's variable) but generally are they more expensive at a single request level. Also do they usually sample and is there a good/recommend way to sample this. I forgot exactly who but (probably NewRelic) was saying they collect all traces (like every request?) and discard if they are not anomalous (to save on storage). But does that mean taking a trace is very cheap? And is that end of the request sampling decision something that's common or that's a totally unique capability some have.
- vmihailenco 4y agoMy understanding is that APM became or always was a marketing term which is used rather freely. For that reason I try to avoid it, but search engines love it and I don't know a better alternative. >Whether I can find out why that component is slow might not be as easy (not sure what granularity tracing happens inside a component). It is true that you can't always guess what operation going to be slow and instrument it, but it is almost always a network or a database call. There is still no way to tell *why* it is slow, but the more data you have the more hints you get. >What's the usual overhead of a single tracing Depending on what is your base comparison point the answer can be very different. Usually, you trace or instrument network/filesystem/database calls and in those cases the overhead is negligible (few percents at most). >But does that mean taking a trace is very cheap? What you've described is tail-based sampling and it only helps with reducing storage requirements. It does not reduce the sampling overhead. Check https://uptrace.dev/opentelemetry/sampling.html https://uptrace.dev/opentelemetry/sampling.html But is taking a trace cheap? Definitely. Billions of traces? Not so. >request sampling decision something that's common or that's a totally unique capability some have. It is a common practice to reduce cost of sampling when you have billions of traces, but it is an advanced feature because it requires backends to buffer incoming spans in memory so you can decide if the trace is anomalous or not. Besides, you can't store only anomalous traces because you will lose a lot of useful details and you can't really detect anomaly without knowing what is the norm. Hopefully that helps.
- KronisLV 4y agoThis seems like a pretty cool project! Currently using Apache Skywalking myself, because it's reasonably simple to get up and running, as well as integrate with some of the more popular stacks: https://skywalking.apache.org/ https://skywalking.apache.org/ I do wonder how ClickHouse (which Uptrace uses) would compare with something like ElasticSearch (which is used by Skywalking and some others) and how badly/well an attempt to use something like MariaDB/MySQL/PostgreSQL for a similar workload would actually go. I mean, something like Matomo Analytics already uses a traditional RDBMS for storing its data, albeit it might be an order of magnitude or two off from the typical APM solution.
- vmihailenco 4y agoWhen compared with ElasticSearch, ClickHouse can handle the same amount data using 10x less resources and that is not an exaggeration. It is even worse with MariaDB/MySQL/PostgreSQL. I guess ElasticSearch is still relevant when it comes to searching text, but ClickHouse is much faster when it comes to filtering and analyzing the data. Give ClickHouse a try and you won't be disappointed. https://benchmark.clickhouse.com/ https://benchmark.clickhouse.com/
- bovermyer 4y agoI see lots of new tracing options these days, and that seems to have taken over the "APM" term. I still have yet to see new profiling options. When I think of APM, I think of CPU profiling and automatic instrumentation of black box systems, not request tracing. I should be able to see which function calls are slow/problematic, without having to add code to the application.
- nijave 4y agoI think there is some interesting work being done with eBPF in the profiling space
- vmihailenco 4y agoI use memory profiling with Go and it is indeed very useful. I think that whatever Go already provides is enough, but I guess Uptrace could try to automate some things and/or provide some fancy UI. But I find CPU profiling a lot less useful, because production profiles tend to be too broad and it is hard to make sense from them, for example, Uptrace profile mostly consists of memory allocations and network calls. So I would not say that CPU profiling is superior/better/can replace tracing.
- bovermyer 4y agoTracing is a piece of the puzzle and necessary. Profiling occupies a different part of the monitoring ecosystem.
- tylergetsay 4y agoI think the log interface should be optimized for keyboard navigation and larger screens. On my 4k monitor it only takes up 1/2 the width and only shows 10 lines at a time, id expect closer to ~100
- vmihailenco 4y agoThanks for the feedback. Any projects that you could recommend that do it right?
- xfer 4y agoAnyways to export dashboard for public viewing, maybe even static image? It looks like all drawing is done client side at present.
- vmihailenco 4y agoEcharts which we use supports exporting charts as images so it probably can be added relatively easy. Embedding is another possible option.
- xyzzy_plugh 4y agoI've been out of the loop for a while but... > OpenTelemetry Protocol (OTLP) > OTLP > OLTP I'm going back to bed.