4 ms·
I think the most often mentioned problems mentioned are pollution of Hetzner addresses by shady people (might be addressed with "exits" from AWS / Cloudflare) a
by Keyframe 1y ago
I think the most often mentioned problems mentioned are pollution of Hetzner addresses by shady people (might be addressed with "exits" from AWS / Cloudflare) and you are running on hardware which does tend to fail / needs upgrades. Were there some concerns on those from you?
Also, Loki! How do you handle memory hunger on loki reader for those pesky long range queries, and are there alternatives?
- sksjvsla 1y agoPollution: We front everything user-facing through Cloudflare, so external users (and bots) don’t interact directly with our Hetzner/OVH IPs. We lock down our IPs at the firewall (ufw + Cloudflare IP allowlisting) so only trusted sources can even connect at L4. Failures/upgrades: We provision with Terraform, so spinning up replacements or adding capacity is fast and deterministic. We monitor hardware metrics via Prometheus and node exporter to get early warnings. So far (9 months in) no hardware failure, but it’s a risk we offset through this automation + design. Apps are mostly data-less and we have (frequently tested) disaster recovery for the database. Loki: We’re handling the memory hunger by • Distinguishing retention limits and index retention • Tuning query concurrency and max memory usage via Loki'’'s config + systemd resource limits. • Use Promtail-style labels + structured logging so queries can filter early rather than regex the whole log content. • Where we need true deep history search, we offload to object store access tools or simple grep of backups — we treat Loki as operational logs + nearline, not as an archive search engine.
- Keyframe 1y agoThanks for thorough answer! Seems like you've platformized(!) yourself to an extent, have you considered going full on with k8s on top of metal (their machines) to offset some of the concerns about hardware?
- sksjvsla 1y agoThanks for the compliment. We used AWS EKS in the old days and we never liked the extreme complexity of it. With two Spring Boot apps, a database and Redis running across Ubuntu servers, we found simpler tools to distribute and scale workloads. Since compute is dirt cheap, we over-provision and sleep well. We have live alerts and quarterly reviews (just looking at a dashboard!) to assess if we balance things well. K8s on EKS was not pleasant, I wanna make sure I never learn how much worse it can get across European VPS providers.
- NewJazz 1y agoHmm, what was so unpleasant about EKS if you don't mind my asking?
- mdaniel 1y agoI'm guessing the answer is going to center around the word "complexity" cited in their original comment. That is: I would guess it's YAGNI more than EKS itself There's an ongoing thread (one of many) exploring the different perspectives on that debate: https://news.ycombinator.com/item?id=44317825 https://news.ycombinator.com/item?id=44317825
- mystifyingpoi 1y agoHe said "in the old days" so probably before addons, managed nodegroups or auto mode. This must have been hell.
- NewJazz 1y agoAh. Yeah I'm setting things up with managed node groups and it doesn't seem so bad so far... Waiting for the other shoe to drop after so much doom saying though haha. Luckily we removed the need for anything stateful, so I can ignore the EBS-CSI shortcomings. Also trying to keep it simple/minimal when it comes to ingress and networking.
- mystifyingpoi 1y agoCool. I've never had any issue with EBS CSI driver itself, the biggest issue were idiosyncracies of EBS itself, like the mount count limit or availability zone requirements. These need ugly workarounds, like limiting your volumes to a single AZ, so, no HA. On the other side, their VPC CNI plugin and their ingress controller are pretty much set and forget.
- NewJazz 1y agoYeah basically it is that EBS limitation (AZ-specific) combined with autoscaling causing quirky failure modes unless you are careful about how you set things up. https://github.com/kubernetes-sigs/aws-ebs-csi-driver/issues/923 https://github.com/kubernetes-sigs/aws-ebs-csi-driver/issues...
- sksjvsla 1y agoA good alternatives for Loki is Victoria. Popular, way more performant and reputable but we went with Loki because of the relative size and diversity of maintainers between the two projects. Your points are super valid and we worked around it as mentioned above.
- chuckadams 1y agoQuickwit is also worth a look, along with its log collector companion Vector. I think at least Vector was a YC company before they got shlorped up by Datadog, but they're still both actively maintained open source.
- TZubiri 1y agohttps://en.wikipedia.org/wiki/Sybil_attack https://en.wikipedia.org/wiki/Sybil_attack One of the advantages of more expensive providers seems to be that they have good reputation due to a de facto PoW mechanism.
- sksjvsla 1y agoDepends on the use case, right? I don’t accept traffic from random Hetzner IPs — only Cloudflare’s IPs are allowed. The only potential indirect risks is if your Hetzner VPS IP range gets blacklisted (because some Hetzner clients abuse it for Sybil attacks or spam). Or if Hetzner infrastructure was heavily abused, their upstream or internal networking could (in theory) experience congestion or IP reputation problems — but this is very unlikely to affect your individual VPS performance. This depends on what you are doing on Hetzner and how you restrict access but for an ISO-27001 certified enterprise app, I believe this is extremely unlikely.
- liampulles 1y ago(Not OP): On the loki question: yeah our project had a similar issue. I did a lot of playing around with the loki configuration, and what you'll discover by reading their blogs on Loki performance is that the indexing settings they recommend are not the ones that are used by default in helm (and probably other deployment configurations). Once I did some reconfiguration, added read specific instances, and implemented their other recommendations - we did see much better performance. Just remember: their interest is that you buy their cloud service, not in giving an out-of-the-box great experience on their open source stuff.