27 ms·
Show HN: SadServers – Test your Linux troubleshooting skills
Hello, I'm building SadServers.com, a SaaS where users can test their Linux troubleshooting skills on real Linux servers in a "Capture the Flag" fashion.
I hope this is useful, to learn more about the project please see https://github.com/fduran/sadservers https://github.com/fduran/sadservers
- Timja 4y agoThe idea is really cool, but all I see is "Waiting for server..." and nothing happens.
- kiyundai 4y agoThat's the trick you failed the first challenge : "Did you try to turn it off and on again?"
- bravetraveler 4y agoCommenting to give this a try later, I've routinely been the person to get these kinds of gremlins escalated I've long wanted for some sort of mock, "things are broken - I want to see how you think" approach for sysad
- shagie 4y agoIn the "tricks of hacker news" - 188 points by fduran 3 hours ago | unvote | flag | hide | past | favorite | 68 comments If you click 'favorite' it will save it to your favorites list. This is a publicly visible list - yours is https://news.ycombinator.com/favorites?id=bravetraveler https://news.ycombinator.com/favorites?id=bravetraveler and mine is https://news.ycombinator.com/favorites?id=shagie https://news.ycombinator.com/favorites?id=shagie which makes it easy to get a bookmark type style functionality within HN. As I tend to favorite less often than I comment, it makes it easier to find those things I want to find again.
- bravetraveler 4y agoMuch appreciated! I'm woeful about using not using features like this, it's a character fault at this point. The HN interface too tends to just have my eyes filter out those links... but that's no defense. Especially good to know that it's publicly viewable! Not that I'm particularly worried of being outed by anything I favorite here, it's just good to be mindful of the data we make and where it goes.
- vermon 4y agoSeems like it's out of capacity: An error occurred (VcpuLimitExceeded) when calling the RunInstances operation: You have requested more vCPU capacity than your current vCPU limit of 64 allows for the instance bucket that the specified instance type belongs to. Please visit http://aws.amazon.com/contact-us/ec2-request to request an adjustment to this limit. Maybe something like https://leaningtech.com/webvm-server-less-x86-virtual-machines-in-the-browser/ https://leaningtech.com/webvm-server-less-x86-virtual-machin... would be cheaper and more reliable for this kind of thing?
- fduran 4y agoYes, HN effect lol-sob. Mitigation: reducing servers life time temporarily so more people can try.
- warent 4y agoUsually I roll my eyes when someone posts their own website to HN and it crashes under load. But given the nature and complexity of yours I think there's room for understanding and patience :)
- fduran 4y agoThanks, I did some stress-testing and infra is scalable enough but I forgot about the AWS quotas, my bad. Quota increase requested and servers are killed off so hopefully "soon" the issue will go away.
- Nextgrid 4y agoScaling this service without breaking the bank could become its own "sad server" scenario. I'd start by moving the test VMs to bare-metal servers running libvirt. You can get a 128GB RAM server for ~110 EUR and that should be able to run around 120 concurrent VMs assuming 1GB of RAM to each (CPU isn't a major issue in this case).
- deeblering4 4y ago> It's also my not-so-secret hope that a sophisticated enough version of SadServers could be used by tech companies (or for companies that carry on job interviews on their behalf) to automate or facilitate the Linux troubleshooting interview section. Yup, that's what I was afraid of.
- fduran 4y agoThat doesn't mean that I'd charge individual users :-) Heck, I'm not even asking for an email (and I had to do extra session management coding for that).
- aliqot 4y agoI knew this is where it was headed :/
- lbotos 4y agoWhy are you afraid of this? My org has run a hands-on technical exam with a stack of linux admin basics (I won't enumerate them here because people do their research) but they are based on real problems we've had and the feedback is overwhelmingly "this was one of the best technical interviews I've ever had." We ask the engineer who is proctoring the interview to think about the following question: Would you want to pair with that engineer again? If that answer is no, then we probably won't go further because pairing with engineers to troubleshoot is what we do every day. Some great resumes have died with not knowing how to see what's running on port 80.
- Pr0ject217 4y agoCool!
- computershit 4y agoI love this idea, I'll definitely try it out when provisioning for scenario machines is up again. Nice work.
- apawloski 4y agoBased on your architecture diagram it looks like you're spinning up an instance per-user? As you're probably finding now, you will hit AWS limits quickly. You might instead want to have a smaller pool of (larger) servers that you run co-resident VMs on with https://firecracker-microvm.github.io/ https://firecracker-microvm.github.io/. That will avoid account limits and also keep your AWS costs more predictable.
- yamtaddle 4y agoJust run them in Linux VMs with WASM, on the users' browsers. Make them all pay for it with higher utility bills and greater wear & tear on their hardware. trollface.jpg
- freeone3000 4y agoThis is actually a good idea for this -- the user wants the education, they can pay for it with their own hardware. Keep your costs low!
- cogman10 4y agoProbably a better experience for everyone. You just have to distribute the image (rather than running vms) and the user gets instantaneous responses.
- b20000 4y agodid you read up on the problems with leetcode?
- fduran 4y agoHi, not sure what the question means, I came up with the scenarios not copying from leetcode if that's what you mean.
- pxc 4y agoI think they mean 'are you aware of the limitations of Leetcode-like tests and the downsides of their (over)use in hiring processes?' (FWIW I think this is a very cool and fun educational project regardless of what usefulness it might or might not have in IT hiring decisions, and I'm looking forward to playing with it)
- yubiox 4y agoCan't get to the first problem because of HN hug but anyway there are fake ways to "solve" it like renaming the logfile (what they test for solved is provided).
- BossingAround 4y agoThis is a self-test, not a certification. The goal is not to defeat the verification goal, but to learn something. So yeah, it's perfectly acceptable that the tests are not bullet-proof.
- Timja 4y agoDepends on how the broken program writes to the log. If it does while true; do echo hello >> bad.log; done Then renaming bad.log will not solve the challenge.
- teddyh 4y agoReplace it with a symlink to /dev/null! Or /dev/full if we feel like it. (Yes, these are bad solutions, since the instructions explicitly said to stop the process which is writing.)
- fduran 4y agoThere are ways to cheat but not so simple; there's a script that checks for the solution and a hash of the script is checked for modifications.
- bm-rf 4y agoI'm assuming you're spinning up an EC2 instance for each lab. What do you think about using pre-built docker images for each challenge instead? that way they can spin up in just a couple of seconds. Might also be cheaper?
- bravetraveler 4y agoNot a bad idea but something to consider; this limits the options for kernel level things quite considerably
- clvx 4y agoprobably lxd would be better.
- fduran 4y agoI wanted to do full VMs rather than Docker images but yes I could do Docker images or dedicated big instances with VMs on top like somebody else is suggesting.
- x258wang_hn 4y ago
- dugmartin 4y agoI'd suggest integrating https://bellard.org/jslinux/ https://bellard.org/jslinux/ and running the VM in the browser if you can - then you can scale without running out of resources.
- m00dy 4y agoor linux kernel port on webassembly.
- fduran 4y agoThanks, I've been looking at WASM, for ex https://github.com/snaplet/postgres-wasm/tree/main/packages/buildroot https://github.com/snaplet/postgres-wasm/tree/main/packages/... , it would certainly simplify everything to "download a fat file".
- jodrellblank 4y agoHave you seen https://copy.sh/v86/ https://copy.sh/v86/ ? It doesn't run as fast as jslinux but is BSD Licensed, on Github, and supports resuming the VM from a snapshot. https://github.com/copy/v86 https://github.com/copy/v86
- fduran 4y agoDidn't know about this, thanks!
- andrewmcwatters 4y agoMy only feedback is that this is unrealistic because today developers wouldn’t try to debug something, they’d just destroy the instance, push a commit and hope it fixed something infra related then recreate it. Why would you need to understand how something works? Just use containers. /s
- sshd 4y agoThis is so sad but so true!
- edmcnulty101 4y agoIf its dumb and it works it's not dumb.
- vsareto 4y agoDevelopers just need to understand everything because we need developers to do everything and meet all deadlines. We wouldn't dare consider a support role that could troubleshoot it because then there would be no point to having developers that can do everything! /s
- cube00 4y agoSupport doesn't deliver features, we need new features! /s
- grepLeigh 4y agoIf most developers can't debug a VM, then anyone who can will be able to charge a premium. If you have a proficiency in ops, remember that the next time you negotiate a compensation package. [Edited my compensation numbers to avoid down votes - yikes]
- andrewmcwatters 4y agoI feel like you definitely have to target particular companies and more specifically specific titles and skills to offer to do so. My guess is trying to sell high end services as a "principal software engineer" isn't going to be enough to justify that cash comp to a lot of people hiring.
- mewse-hn 4y agoCompleted the first challenge and it was a lot of fun - spoiler I've never had to use the 'lsof' command before.
- nobody9999 4y ago>Completed the first challenge and it was a lot of fun - spoiler I've never had to use the 'lsof' command before. I've been waiting a while for the "sad server" to come up for me and read the scenario (saint john) whilst waiting. lsof was the first thing that came to mind after reading the scenario. I guess that once I actually get a "sad server" I'll make it "happy" quickly :)
- N3Xxus_6 4y agoWell this sucks I wanted to try it lol. It's timing out for me or throws an error.
- hotpotamus 4y agoAre you familiar with Trueability? https://www.trueability.com/ https://www.trueability.com/ It seems like this is a similar SaaS.
- fduran 4y agoDidn't know about this one. There's quite a few labs/sandbox SaaS but what I've seen so far is that they are more for training with a "follow the recipe" model (do this do that to configure something, rather than "this (real) server is broken, fix it (with possibly different solutions)" which imho is more real-life and useful.
- hotpotamus 4y agoI believe the company was founded by some coworkers of mine way back when at Rackspace who often interviewed Linux admins with a lab VM and I assume they just automated the setup and spun it off as their own business. At least that's what happened as far as I can tell; I didn't know the parties involved.
- DeathArrow 4y ago>Practice for your next SRE/DevOps interview. Are SREs and DevOps tasked with administration of operating systems?
- BossingAround 4y agoHow do you automate something you can't do manually?
- dsr_ 4y agoIf you have ops somewhere in your responsibilities, then yes.
- asmr 4y agoBoth SRE and DevOps are essentially evolved sysadmin roles. The DevOps philosophy is cross-functional and many sysadmins have adopted a DevOps approach. The latest edition of the classic sysadmin book "The Practice of System and Network Administration" is now centered around DevOps.
- KaiserPro 4y ago> Are SREs and DevOps tasked with administration of operating systems? yes, eventually. you can dress it up in all the fancy terms that you like. but devops and SREs are sysadmins with better PR. its critical that SREs understand _how_ to debug a system, so that they can work out how to put in fixes, and or design better systems.
- jabroni_salad 4y agodepends on what layer the issue is happening at. I know everyone thinks the OS has been abstracted away but my ticket queue says otherwise. "yaml engineering" is just a control surface, I still need to pop the hood often.
- jen_h 4y agoYeah. Random data point: One of my most favorite SRE interviews ever (serious fun!) involved hands-on troubleshooting that eventually required gdb.
- grepLeigh 4y agoVery cool! This reminds me of the ops challenge @ Slack. I'm not sure if they still do this, but the SRE/platform infra interview used to involve a VM running a malfunctioning LAMP stack. You'd get SSH access to the VM, then submit a diagnostic report of what was broken (and how you fixed it). Reminded me of how Red Hat used to run their certification test (RHCE). I probably still have the live CDs for my RHCE laying around somewhere.
- stevekemp 4y agoI've had interviews like that in the past, and really enjoyed them. Much better than "Draw an architecture diagram for how you'd handle a serverless IoT application" - where you lose points, silenly, because you didn't pick something the interviewer expected you to do. Usually a simple combination of immutable files, SELinux policies, and types in configuration files were enough for most of the challenges. Though now and again you'd find they'd given you a server with packages removed, or not yet installed.
- fduran 4y agoOh that reminds me, I loved the original Stripe CTF, it's been 10 years already! https://twitter.com/fduran/status/240321390698442753 https://twitter.com/fduran/status/240321390698442753
- PanosJee 4y agoHack The Box -> Fix The Box
- BossingAround 4y agoI'd love to get the actual VM content offline, packaged as Vagrantfiles or Containerfiles. Love the idea though! Go to Pluralsight and pitch it to them :)
- fduran 4y agoA few people have suggested offering content offline as a Docker image etc, good idea, thanks.
- yapril 4y agoCan I download the images so I could run it on my own machine ? I'd really appreciate, I've got an interview very soon :)
- lagrange77 4y agoReally cool idea. After choosing a problem, the endpoint you poll at https://sadservers.com/celery-progress/xxxx https://sadservers.com/celery-progress/xxxx repeatedly returns {pending: true, current: 0, total: 100, percent: 0} for me.
- fduran 4y agoyes good catch (I should forbid internet access to this end point), poor queue is waiting on VM up but there's no quota left until other VMs are garbaged-collected.
- jer0me 4y agoNew challenge: Fix SadServers’ sad servers
- vetkat 4y agoAnd while we’re at it, we might as well write a wrapper around low-upvote Server Fault questions in the hope that they attract more attention when the problem is gamified.
- 10g1k 4y ago"Have you turned it off and on again?"
- imwillofficial 4y agoThis is badass, just what I need!
- arwt 4y agoInteresting idea! Looking forward to trying this once some VMs are available. :-)
- diffcheck 4y agoThe tasks loading infinitely, is it a zero challenge?
- sylvainkalache 4y agoVery cool project! I was the founder of a school training software engineers, we had an infrastructure track that got a lot of our students to land SRE positions. When asking employers for feedback about our grads, one feedback kept coming: they lack experience when it comes to troubleshooting. So I went on a quest to simulate that infra debugging while in an academic context. I came up with the idea of giving students broken servers. I used Docker container and would setup a simple workload and mess it up with classic issues. Needless to say students generally did not like it :) debugging isn’t fun. But it did help a lot.
- fzyzcjy 4y agoThis looks interesting! But it keeps loading forever saying "Your server is being created" (hit VM limit again?)
- ASalazarMX 4y agoI only want to say that I love the name SadServers. Strongly memorable.