29 ms·
Influxdb made the switch from Go to Rust
- tsak 3y agoSeems like they have their priorities right. /s https://news.ycombinator.com/item?id=36657829 https://news.ycombinator.com/item?id=36657829
- vb-8448 3y agoCan someone explain what's the InfluxData's market? Or how they make/plan to make money? If we speak about metrics, Prometheus just win.
- GauntletWizard 3y agoI am about the biggest Prometheus stan that you can find, but I will not mock or denigrate influx db. They are a runner-up, but there is room in this market for runner-ups, and their feature set does match some that Prometheus is not good at.
- vb-8448 3y agoDon't get me wrong, I don't denigrate Influx DB. They have some interesting features that prometheus doesn't have, I just don't think that this features can lead to a mass adoption(and money). Let's put this way: is there any killer feature that can ditch most of the Prometheus installations in their favor?
- ilyt 3y agoAnd it *has* to be feature, not just "it's faster", because they are (much) faster time-series databases that support PromQL and are near-drop-in to your infrastructure.
- ilyt 3y agoI think they missed one thing: Nobody wants to get locked in with the runner-up. If they on top of that provided some compatibility layer for Prometheus/PromQL, like few other competitors did, then the prospective enterprise client have warm fuzzies that if they don't like it they don't need to rewrite entirety of their stack to work with something else. People could also use the existing ecosystem and "just plug it in", even replacing Prometheus instances they might have. In the end people want to ingest the metrics and display it in Grafana. They don't need another visualisation solution that has less support and documentation. They don't want to learn new weird query language that is simulatenously more verbose and less readable than PromQL or even influxQL from v1.
- manicennui 3y agoI think you'd find that Prometheus is way behind a number of solutions in the corporate world, including Graphite, DogStatsD, etc. Popularity on HN does not translate to real world popularity.
- vb-8448 3y agoMost of the companies in the corporate world have more than a monitoring system, at least 3 in my experience: sysadmins have at least one, network guys another one and one for legacy platforms (mainframes, as400 and so on). Plus, in the last years with the rise of Kubernetes, most corporates have at least one cluster with Prometheus monitoring it. Always, base on my experience, Prometheus is very popular, but the adoption is not so wide due to its `oss nature`: CTOs want someone to blame when things go wrong.
- hagen1778 3y agoThat's only a matter of time when they hire younger engineers who are familiar with modern monitoring systems and eager to apply their knowledge in practice.
- manicennui 3y agoWhat is it that you think prometheus offers over other solutions? It is more likely that the younger engineer is going to learn that companies don't care about what is popular on HN.
- hagen1778 3y ago> What is it that you think prometheus offers over other solutions? I like Prometheus and think this is a great piece of software. But even if we won't go into actual details, Prometheus is baked in into Kubernetes monitoring [0]. That's the first monitoring system young engineers will meet with when learning k8s. Although, k8s and Prometheus are both CNCF projects which means both of them will be promoted in synergy with each other. > It is more likely that the younger engineer is going to learn that companies don't care about what is popular on HN. This is not what I think younger engineers do :) [0] https://kubernetes.io/docs/tasks/debug/debug-cluster/resource-usage-monitoring/ https://kubernetes.io/docs/tasks/debug/debug-cluster/resourc...
- jeroenhd 3y agoI see Influx pop up a lot in communities like Home Assistant for doing time series data collection. I'm using Postgres for that myself (not a great experience but I knew that when I was too lazy to find alternatives during setup) but people seem very pleased with its performance. I imagine similar dashboard services that don't necessarily work well in Prometheus are a good market for these types of databases. Prometheus is nice, but I don't think it's suitable for all Influx use cases (and vice versa).
- iworshipfaangs2 3y agoDo you know if there is any particular advantage of Influx over Prometheus for IOT stuff? I also have noticed that Influx is way more popular in that space, but I don’t know whether the reasons are technical or just social (more tutorials, more shared experience, etc).
- 46Bit 3y agoInflux is a full time series database. It's less opinionated than Prometheus.
- ilyt 3y agoVictoriaMetrics all-in-one binary is honestly everything that I need in my home IoT things. Prometheus-compatible interface for query, a bunch of ingest protocols, smaller memory usage than InfluxDB (v1, haven't tested v2 coz new language have less grafana support). Options to scale too
- lakomen 3y agoInflux is a wholesome solution where Prometheus is just the time series database
- vb-8448 3y agoMaybe i miss something, but the version 3 is just the database, they abandoned the TICK stack some time ago.
- justin_oaks 3y agoAnd I thank them for it. I was using the TICK stack a few years ago and the switch to TIG (Telegraf, InfluDB, and Grafana) has been a breath of fresh air. Telegraf and InfluxDB are solid, but Chronograf was behind Grafana in usability and features. Kapacitor was pretty rough. The language it used was hard to write, the docs were barebones, somewhat confusing, and sometimes inaccurate. When I switched off of Kapacitor, the CPU usage on the server dropped significantly too. So I'm guessing Kapacitor wasn't too CPU friendly either.
- Too 3y agoCan you expand what you are missing? Prometheus is not just the db itself, it’s the ecosystem around it. You’ve got service-discovery, alertmanager and basically every application in existence having a /metrics endpoint and some pre-made Grafana dashboard.
- capableweb 3y agoAs is tradition, hosted service: https://www.influxdata.com/influxdb-pricing/ https://www.influxdata.com/influxdb-pricing/ Basically either you manage it yourself, or you pay them to do Serverless/Dedicated/Clustered hosted setup for you.
- vb-8448 3y agoYeah, I know, but there are a lot of competition in this field. Every cloud provider, Grafana and other players have their own managed Prometheus compatible solution.
- influx 3y agoThey seem all over the place. https://www.theregister.com/2023/07/11/influxdata_apologizes_for_ending_services/ https://www.theregister.com/2023/07/11/influxdata_apologizes...
- gabereiser 3y agoThere in lies the problem. With shifting priorities and shifting strategies I don’t see them actually knowing what their market is.
- cmrdporcupine 3y agoThis I guess is the mystery to me. Projects like this are great. But how are people funding them? (AKA how can I get my own paws on some investment $$ to work on cool database^Wother-systems-level tech? I'll even promise to try to make $$)
- matsemann 3y agoPrometheus only handle aggregated data, though. While with influx you can store the events themselves with labels etc. While Prometheus is often good enough for standard metrics, it is just things it can't handle.
- ilyt 3y agoAFAIK only difference is that in Influx a given row can have more than one value so say "interface traffic" would be interface_if_octets host=router,instance=eth0,rx=123584,tx=213956 while in prometheus it would be interface_if_octets host=router,instance=eth0,type=rx 123584 interface_if_octets host=router,instance=eth0,type=tx 213956 which in theory yes it is more compact but it gave me more annoyances than advantages during querying >While Prometheus is often good enough for standard metrics, it is just things it can't handle. My experience is that just anything made to ingest and analyze logs ends up mediocre for metrics and vice versa. I don't think I've seen single product that did both well or efficiently. So I'd rather have good metrics and just use ELK/Graylog/whatever else for logs.
- matsemann 3y agoNot really the biggest difference, I feel. In Prometheus, you only have the data you scrape. So lets say you scrape every minute. Then all you have is a counter per pod increasing from some value to something else. With influx you can save every event. So you know exactly when it happened, the unique labels for that event etc. It's a completely different paradigm. So in Prometheus you have myCounter,pod=1,time=20:23,value=1000 myCounter,pod=2,time=20:23,value=500 myCounter,pod=1,time=20:24,value=1100 myCounter,pod=2,time=20:24,value=700 so all you know is that some event happened 100 times on pod1 and 200 times on pod2 the last minute. But with influx you could have a row for every single event. Of course that explodes the query time in comparison, but allows you to do much more with the data if needed.
- hagen1778 3y ago> Prometheus only handle aggregated data, though. That's not true. You're referring to pull-based approach for metrics collection. It has its tradeoffs (like fixed interval scraping), but has a lot of benefits too (like higher reliability). Check the following link [0] from VictoriaMetrics docs, which supports both push and pull approaches. Prometheus also gained push support this year, though. However, the main difference between Prometheus-like systems (Thanos, Mimir, VictoriaMetrics) and more traditional DBs for time series like InfluxDB or TimescaleDB is that first are designed to reflect system's state, and last are designed to reflect system's events. That's the main difference in paradigm, data model, and query languages. There is a reason why PromQL is so easy in 99% of cases, and so complex and annoying when users want to express what they get used to in traditional databases. I'm saying this because I went through creating a Grafana datasource for ClickHouse [1] and I felt how complicated it is to express a most straightforward PromQL query in SQL, and vice versa. If you'd like to learn more about differences between common queries for plotting time series in PromQL and SQL see my talk here [2]. [0] https://docs.victoriametrics.com/keyConcepts.html#write-data https://docs.victoriametrics.com/keyConcepts.html#write-data [1] https://grafana.com/grafana/plugins/vertamedia-clickhouse-datasource/ https://grafana.com/grafana/plugins/vertamedia-clickhouse-da... [2] https://youtu.be/_zORxrgLtec?t=835 https://youtu.be/_zORxrgLtec?t=835
- pauldix 3y agoSo it sounds like you're asking about our use cases. We have customers across almost every vertical. But what's most common are application and server monitoring, sensor data (industrial, rockets, satellites, etc.), and network monitoring. Metrics is certainly one use case that people pay us for. With v3, we expect that real-time analytics and some more data warehousing types of use cases will become interesting. We always envisioned InfluxDB as a store for observational data of all kinds, not just metrics. On how we make money, we sell our products. We have at this time: - InfluxDB v1 Enterprise (a self-managed, clustered implementation of InfluxDB) - InfluxDB v1 Cloud (Enterprise, but as as single-tenant managed service. We still run this for hundreds of customers) - InfluxDB v2 Cloud (multi-tenant, usage based, we're running this for thousands of customers) - InfluxDB v3 Serverless (multi-tenant, usage based v3) - InfluxDB v3 Cloud Dedicated (single-tenant, resource based pricing) - InfluxDB v3 Clustered (self-managed v3, clustered database) We'll have single server versions in the future, but we're a bit off from that. Right now our focus is on continuing support for our v1 and v2 customers, and further developing our v3 products for new customers and customers that want to migrate over.
- atombender 3y agoDoes InfluxDB support event-based data models? For example, imagine something simple an SQL database's query log. Each query is an event, with data such as latency, rows, block I/O and so on, and metadata such as the database, the full SQL query, and so on. This is the kind of use case where traditional "measurement"-based time series databases like Prometheus aren't a good fit, because you have huge column cardinality for the labels (one value per query). Meanwhile, more general-purpose columns databases like ClickHouse and BigQuery have no issues with this type of data.
- pauldix 3y agoInfluxDB v3 is built to handle this kind of data. It's a columnar database, using object storage and Parquet files for persistence.
- techn00 3y agotl;dr - No garbage collector - Fearless concurrency (thanks Rust compiler) - Performance - Error handling - Crates - they thought they were gonna use C++ and wanted interop (ended up not using C++?) - ecosystem: Apache Arrow DataFusion - "I thought that if we're going to rewrite most of the database anyway, we might as well do it in the best language choice in 2020" But the real reason might be: "Rust good, Go bad" /s
- tsak 3y agoRust gets you on the frontpage!
- valenterry 3y agoGo does too. Those two feel like the most hyped languages currently on HN.
- CharlieDigital 3y agoNo love for C#
- leosanchez 3y agoI love C# :)
- e-master 3y agoC# is lovely, and with eventual first class AOT support I believe it will become more widespread
- demi56 3y agoNo F# ?
- adamors 3y agoGo isn’t really hyped these days, Zig maybe. But Go is old reliable now (which is good).
- mardifoufs 3y agoI love influx but damn do they like moving (too?) fast and quickly changing stuff. In a way, it's pretty cool since it means that they don't get stuck with bad decisions for backwards compatibility reasons, but it's a bit of a roller coaster for users. Not sure what's the best solution though. Having a "stable" but fundamentally limited product (I guess influxdb v1) or breaking stuff in hopes of ending up with a way better technical foundation.
- DotaFan 3y agoSounds like tech I wouldn't wanna depend on.
- mnahkies 3y agoI've always had a soft spot for influxdb after using it for a self hosted datadog/newrelic etc solution many (6+) years ago with great success. Still use it in conjunction with telegraf and grafana for personal project monitoring, but I've not brought myself to upgrade from the 1.x series. Hopefully it's improved, but last time I tried upgrading I found the UX in grafana to be subpar on the newer versions, as I recall you lost the autocomplete/UI to build your queries. Obviously grafana is it's own project but feels like they (influx) should invest more resource in areas like this to encourage people to upgrade - if you're going to do major upgrades make sure they have feature parity
- justin_oaks 3y agoLike you, I've stuck with Influx v1, Telegraf and Grafana. My policy is to upgrade only when there are significant reasons to. When I evaluated InfluxDB 2, there were no major reasons for me to switch. Of course, the data ingested in my case is relatively small. YMMV. I looked at TimescaleDB but at the time there was no easy way to get data from Telegraf to TimescaleDB. Telegraf finally merged code that allows writes to Postgres databases, but it took like 3 years to do that. Ultimately, I still stuck with InfluxDB v1 because sending data to it via the InfluxDB line protocol is so simple. I have a couple of bash scripts that use awk to transform command output to Influx line protocol and send it to InfluxDB. It's just so simple. I love it. I love learning about new things, but the InfluxDB v1 keeps working fine so I may not switch from it until something forces me to do it.
- zzzeek 3y ago> So this isn't the approach I'd recommend for this kind of project, but we started fresh from scratch. wow, the from scratch rewrite. I can't even imagine that for a major piece of software
- pizzafeelsright 3y agoThis is my dream job.
- brabel 3y agoThe author was looking for an excuse to use Rust for something since 2018: https://www.influxdata.com/blog/rust-can-be-difficult-to-learn-and-frustrating-but-its-also-the-most-exciting-thing-in-software-development-in-a-long-time/ https://www.influxdata.com/blog/rust-can-be-difficult-to-lea... The rewrite started in 2020... they rationalize now, but it's pretty clear they just really wanted Rust and found the reasons they list later as a post-decision justification. Nothing wrong with that, if you don't mind risking the future of your business on a risky rewrite... though if I had a job working on the Go code base and my employer suddenly announced we should drop everything and start a rewrite, so go learn Rust, which has a much smaller job pool in my area as far as I can see, I guess I would've been really pissed off and would leave as fast as possible.
- smabie 3y agoSmaller job pool? How does learning Rust reduce your chances of getting a new job? Learning new stuff increases your chances, not decreasea them.
- pauldix 3y agoTo be fair, we didn't drop everything and do a rewrite. Over the last 3.5 years (the length of time for this project), our total engineering team has ranged from 50-90 people. For the first year it was me and two other people. Then for the 2 years following that it was 9 people total. It wasn't until late last year that we made the decision to go all in on the rewrite and made that the focus of everyone in engineering. And we did that because we had 4 years of experience trying to get v2 and Flux to be successful, with modest results. Most of the time we were developing this version, we were spending massively more engineering effort on developing v2 or maintaining v1 for our customers.
- Dowwie 3y agoWould be great to see an in- depth blog post by Andrew and team about Rust, the bad and the good. They didn't just build a system but one that was optimized for performance. What were the major challenges during the rewrite? Have you optimized CI build times?
- Patrickmi 3y agoPeople are not talking that InfluxData is NO more just a time series database, this isn’t just a language change but feature additions and with the level of massive dependence on C++ libraries it’s pretty foolish to continue using Go
- lakomen 3y agohttps://github.com/influxdata/influxdb https://github.com/influxdata/influxdb The non-reddit link target
- tekla 3y agoIs it possible to do HA with Influxdb OSS yet?
- johnbellone 3y agoFocus on making money and less on replatforming for a second time.
- pphysch 3y agoReplatforming to chase the hype cycle and sell to non-technical CTOs can be an effective money-making strategy.
- adamnemecek 3y agoRust is not hype.
- louwrentius 3y agoAs a side note, the Flux language that they introduced in v2 never seems to have taken off as there are a few (2) public Grafana dashboards made with it, whereas the older influx language has around 1345 currently. Unfortunately I was stupid enough to build my dashboard on flux, which I’m really sorry to say I dislike quite a bit, while still wanting to be respectful for the people who build the stuff. All that said, I think that influx is a great tool although I’m mostly using it for personal projects and haven’t run anything at scale.
- ilyt 3y agoI tried it and it was more complex to do anything compared to PromQL. Like, it looks super powerful for complex queries I will never need to make...
- louwrentius 3y agoYes, and in another HN post the Influx founder stated (me paraphrasing) it’s and it’s all about SQL compatibility.
- tptacek 3y agoThe backstory here is they were doing a rewrite anyways, for reasons that had not much to do with languages; they expected to write some C++ for the new version. Rust was the right call for them.
- bachmeier 3y agoThis was discussed on HN at the time (2020): https://news.ycombinator.com/item?id=25049253 https://news.ycombinator.com/item?id=25049253 At some point HN is going to have to decide if it's the Rust subreddit or a news site.
- mardifoufs 3y agoI think they just completed the transition which is why it came up again.
- threatofrain 3y agoRust is one of the few languages that have a chance to climb out of the hobbyist/academic/ultra-niche range, so it's interesting for me to hear about developments towards the direction of reaching mainstream status. I'd say the same thing about Zig but with less strength. After that nobody wants to hear about Java.
- n8henrie 3y agoI started using HN because I'm a rust fanboy and there was a lot of rust content (same with lobsters). I'm glad to say there is a lot of other HN content that interests me, I might never have known. Funny enough, in contrast to when I joined, the pendulum seems to have swung, and comments disparaging rust seem to be en vogue.
- PaulWaldman 3y agoThe issue is that InfluxDB is an infrastructure product. Changing the core impacts the way users interact with the product. If Figma decided to change their backend, it could be transparent to users. Opinions could be different if first they implemented a complete compatibility layer, Flux included, prior to making the migration.
- Frankmartin321 3y ago[dead]
- inv2004 3y ago[flagged]
- d3w4s9 3y agoThis is really surprising for a commercial company. Not about the choice of language itself, but the fact that they prioritize such rewrites over features, similar to the concerns from other commenters. They mention "performance" and "garbage collector" and "error handling" which are almost technical details. A company can make good money for quite a while as long as the performance is "good enough", and usually focuses on adding features instead of worrying about any rewriting/re-architect until there is a bottleneck or issues start to significantly slow down development. How "successful" such a rewrite still needs to be seen, but this is risky for most companies/products.
- voiper1 3y agoThe first comment is from the Cofounder and he lists the _features_ that motivated a rewrite: >Then there's the question of why we did a rewrite at all. We wanted to get at some important requirements: >Unlimited cardinality >Analytics queries against time series at the performance of a columnar DB >Use object store as the durability layer for historical data (i.e. separate compute from storage) >SQL and broader ecosystem compatibility >All of that stuff taken together meant that we'd be rewriting most of the core of the database. ...
- d3w4s9 3y agoSure, the write definitely could have helped releasing these new features, but the language transition itself takes non-trivial effort and could have been used for adding additional features instead. I haven't seen any evidence that such a language transition is desperately needed. Also I said deprioritize adding features, not stop adding features, mind you. The core question here is what kind of benefit we are seeing by changing the language. If you are a fan of Rust and want to rewrite all your just because of that and you are the CTO, whatever, what can I say. But the title here is a bit clickbait-y and I really don't see much reflection on the languages themselves and how to balance the business needs for such a transition.
- pauldix 3y agoIn this case, the features we kept getting asked for by our customers necessitated a change in underlying database architecture. I talk about that quite a bit in the reddit thread. I totally agree that a rewrite is risky. It's not something I'd choose to do again, but at the time we didn't really see any way around rewriting the bulk of the database (even if we kept it implemented in Go). Using Rust and the Arrow ecosystem of projects (Parquet, DataFusion, Flight) meant that there were a ton of things we didn't have to do from scratch. One of our staff engineers, Andrew Lamb, has called it a toolkit for building databases. Thanks in part to his contributions, I think he's right.
- hu3 3y agoIs there a comparison of the new InfluxDB against https://clickhouse.com https://clickhouse.com? I ask because ClickHouse is quite hot at the moment from my experience in consulting and that seems to be reflected in Google Trends [1]. And there are some startups relying on ClickHouse for their log/monitoring products like https://signoz.io https://signoz.io and https://hyperdx.io https://hyperdx.io. [1] https://trends.google.com/trends/explore?date=all&q=ClickHouse,influxdb&hl=en-GB https://trends.google.com/trends/explore?date=all&q=ClickHou...
- matt3210 3y agoHow can they know what is defined behavior and what is undefined? How can it be proven? Rust has no spec and is therefore all undefined behavior. </cope>
- cyber1 3y agoThis is intriguing. Interesting, how does this new Influx engine compete in terms of performance with VictoriaMetrics (which is written in Go and really fast)? They moved their entire stack from Go to Rust, rewrote the system from the ground, and spent a lot of time on it, I guess this is a big cost. Is it worth it?
- hagen1778 3y agoIf I'm reading this [0] right, there will be no a standalone OS influxdb 3.0 version. So there's no point in comparing. I also wonder if it would be allowed to publish benchmarks of ENT version by 3rd-parties. [0] https://www.influxdata.com/blog/the-plan-for-influxdb-3-0-open-source/ https://www.influxdata.com/blog/the-plan-for-influxdb-3-0-op...
- brokensoul98 3y ago[dead]
- exabrial 3y agoI love InfluxDB 1.x and the TICK stack. They abandoned a beautiful piece of software to chase shiny things with 2.x... sad to see them do it again. Someone needs to pickup the original ideas of 1.x since they can't seem to stay focused, as their marketshare is ripe for grabbing.
- pauldix 3y agoWith 3.0 we've worked hard to pull in the v1 API. Both the v1 write and query endpoints are supported in v3. There may some gaps here and there, but our goal is to make it so that all the things that worked with v1 will work with v3. The only two exceptions would be the subscription API that Kapacitor used and Continuous Queries. We warned people off from using both of those features in v1 as they didn't work well under load. With v3, I prefer to think of it as us doubling down on core database performance and functionality. With v2 we tried to create this whole development platform. V3 brings our focus back to the core database, which I think will yield better results for everyone.
- jacobgorm 3y agoMy problem with Rust is that compilation is too slow, as is downloading the gazillion crates needed to go anywhere. When I heard it was going to replace C I was expecting similar build times, but in reality Rust builds much slower than C++. With C/C++ (and CMake + Ninja) it seemed we were finally getting to a point where incremental builds would complete before hitting the 400ms attention span "Doherty Threshold", and now it seems we are going back to the days of having to spend our time sword-fighting (xkcd) during slow builds.
- jokethrowaway 3y agoIt's way better compared to the early days and we have incremental compilation in dev. I don't consider it a problem anymore. People got the memo that slapping Serialize on everything has a cost an unless you're doing huge native dependencies which require long compilation time, it's pretty snappy.
- asah 3y agoHow big are y'all that you need InfluxDB? For one database that receives 100M 600 byte JSON records/day, a single node AWS PostgreSQL RDS instance is handling it effortlessly, and the DBA work is very part-time. We keep year+ of summaries and 48 hours of detail, unloading the rest to S3 as parquet files, queryable by Athena if we need. AFAICT, we're spending <$4K/mon all in, including backups. p.s. a buddy at a top-3 TV streaming service is also doing this for logging all viewing activity, but with Aurora.
- asah 3y agoupdate: it's actually more like $1500/mon and that includes an RDS proxy, failover instance and online backups with one-click restore.
- lbruder 3y agowe're storing 120M records a day in influx 1.8, offloading cold data to S3, all on a single m5.xlarge instance that runs other backend services on the side. Less than 400 bucks per month overall aws cost, with lots of other VMs in there. We could use RDS too, but why change it to something more expensive if it already works...
- asah 3y agothanks! killer example. how are you dealing with failover? how are you dealing with interactive reporting? how large are your records? is the s3 data in cold storage? how much is stored in S3 right now? is it warm-enough to query via (e.g.) athena?
- Zuiii 3y agoWe're in the process of doing the opposite. Wish we had chosen the right tool for the job from day one.