14 ms·
Solaris to Linux Migration 2017
- unethical_ban 9y agoThis is a great set of information comparing features and tools between the two ecosystems. I like it, and wish more were available for Linux -> BSD and even lower level command tool comparisons like apt-get vs. yum/dnf. In fact, this works a general purpose intro into several important OS concepts from an ops and kernel hacker perspective. My only surprise is that this is written as a specific response to Oracle Solaris' demise. From that specific perspective, how many target viewers are there? 10? Illumos isn't losing contributors, and there are still several active Illumos distros. Nevertheless, interesting.
- brendangregg 9y agoYes, I hope to write one for Solaris -> BSD. This post was written for the illumos community as well.
- rhizome 9y agoYou might be interested in the long-lived Unix Rosetta Stone: http://bhami.com/rosetta.html http://bhami.com/rosetta.html
- nailer 9y agoAlso the 2017 version - covers current-gen OSs including BSD, SmartOS, newer Linux distros and Windows powershell: https://certsimple.com/rosetta-stone https://certsimple.com/rosetta-stone
- holydude 9y agoIt is actually sad to see that you had to write this. Funny how the most used / popular technology and a mismanagement from a single company can crush other competing tech. It is frightening how much of what was invested in Solaris is now lost because of it.
- espadrine 9y ago> It's a bit early for me to say which is better nowadays on Linux, ZFS or btrfs, but my company is certainly learning the answer by running the same production workload on both. I suspect we'll share findings in a later blog post. I am eager to read this piece! Even though I am afraid to see it confirm that btrfs still struggles to catch up… The 2016 bcachefs benchmarks[0] are a mixed bag. [0]: https://evilpiepirate.org/~kent/benchmark-full-results-2016-04-19/terse https://evilpiepirate.org/~kent/benchmark-full-results-2016-...
- OllyMason 9y agoAlso eager to read this piece
- jsiepkes 9y agoIs BTRFS really an option now that RedHat has decided to pull the plug on their BTRFS development? That basically leaves Oracle and Suse I think? As far as I can tell the future of BTRFS doesn't look good. Facebook using it doesn't mean anything since they are probably using it for distributed applications. Meaning the entire box (including BTRFS) can just die and the cluster won't be impacted. I really can't imagine they are using BTRFS on every node in their cluster.
- geofft 9y ago> Facebook using it doesn't mean anything since they are probably using it for distributed applications. Meaning the entire box (including BTRFS) can just die and the cluster won't be impacted. I don't think that follows for lots of reasons: - If enough of your boxes die that you lose quorum (whether from filesystem instability or from unrelated causes like hardware glitches), your cluster is impacted. So, at the least, if you expect your boxes to die at an abnormally high rate, you have to have an abnormally high number of them to maintain service. - Filesystem instability is (I think) much less random than hardware glitches. If a workload causes your filesystem to crash on one machine, recovering and retrying it on the next machine will probably also make it crash. So you may not even be able to save your service by throwing more nodes at the problem. A bad filesystem will probably actually break your service. - Crashes cause a performance impact, because you have to replay the request and you have fewer machines in the cluster until your crashed node reboots. It would take an extraordinarily fast filesystem to be a net performance win if it's even somewhat crashy. - Most importantly, distributed systems generally only help you if you get clean crashes, as in power failure, network disconnects, etc. If you have silent data corruption, or some amount of data corruption leading up to a crash later, or a filesystem that can't fsck properly, your average distributed system is going to deal very poorly. See Ganesan et al., "Redundancy Does Not Imply Fault Tolerance: Analysis of Distributed Storage Reactions to Single Errors and Corruptions", https://www.usenix.org/system/files/conference/fast17/fast17-ganesan.pdf https://www.usenix.org/system/files/conference/fast17/fast17... So it's very doubtful that Facebook has decided that it's okay that btrfs is crashy because they're running it in distributed systems only.
- jsiepkes 9y agoNice article! Though I do think the article could have more clearly noted that Linux containers are not meant as security boundaries. It doesn't explicitly say it but it is a very important distinction. Unlike FreeBSD jails and Solaris Zones. You can't run multiple docker tennant's safely on the same hardware. Docker is basically the equivalent of a sign which says: "don't walk on the grass" as opposed to an actual wall which FreeBSD jails and Solaris zones have. Now if you have a very homogene environment (say you are deploying hundreds of instances of the exact same app) then this is probably fine. Docker is primarily a deployment tool. If your an organization which runs all kinds of applications (with varying levels of security quality) that's an entirely different story.
- okket 9y agoMaybe relevant: "Setting the Record Straight: containers vs. Zones vs. Jails vs. VMs" https://blog.jessfraz.com/post/containers-zones-jails-vms/ https://blog.jessfraz.com/post/containers-zones-jails-vms/ Some discussion about this article here on HN: https://news.ycombinator.com/item?id=13982620 https://news.ycombinator.com/item?id=13982620 (160 days ago, 235 comments)
- jsiepkes 9y agoTrue, and that very long post basically says with many words: "Yes, Linux namespace (docker) isn't as secure as FreeBSD jails or Solaris Zones but security is not the problem docker solves. Docker solves a deployment problem, not a security problem."
- gvb 9y agoComparing containers (a concept) with VM/jails/zones is a non-sequitur. To quote: A “container” is just a term people use to describe a combination of Linux namespaces and cgroups. Linux namespaces and cgroups ARE first class objects. NOT containers. [...] VMs, Jails, and Zones are if you bought the legos already put together AND glued. So it’s basically the Death Star and you don’t have to do any work you get it pre-assembled out of the box. You can’t even take it apart. Containers come with just the pieces so while the box says to build the Death Star, you are not tied to that. You can build two boats connected by a flipping ocean and no one is going to stop you. Docker is a bunch of boxes floating on a flipping ocean[1]. They could have made a deathstar, but they chose not to. [1] https://www.google.com/search?q=docker+logo https://www.google.com/search?q=docker+logo
- dwheeler 9y agoIf you can't access it directly, here's a cached version: https://web.archive.org/web/20170905181357/http://www.brendangregg.com/blog/2017-09-05/solaris-to-linux-2017.html https://web.archive.org/web/20170905181357/http://www.brenda...
- Ologn 9y ago> Crash Dump Analysis...In an environment like ours (patched LTS kernels running in VMs), panics are rare. As the order of magnitude of systems administered increases, rare changes to occasional changes to frequent. Especially when it is not running in a VM. Also, from time to time you just get a really bad version of a distro kernel, or some off piece of hardware that is ubiquitous in your setup, and these crashes become more frequent and serious. (Recent example of a distro kernel bug - https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1674838 https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1674838 . I foolishly upgraded to Ubuntu 17.04 on its release in stead of letting it get banged around for a few weeks. For the next five weeks it crashed my desktop about once a day, until a fix was rolled out in Ubuntu proposed) Most companies I've worked at want to have some official support channels, so usually we'd be running RHEL, and if I was seeing the same crash more than once I'd probably send the crash to Red Hat, and if the crash pointed to the system, then the server maker (HP, Dell...) or hardware driver maker (QLogic, Avago/Broadcom). Solaris crash dumps worked really well though - they worked smoothly for years before kdump was merged into the Linux kernel. It is one of those cases where you benefited from the hardware and software both being made by the same company.
- tytso 9y agoCrash dumps don't matter as much if your distributed architecture has to account for hardware failures. (Or VM failures, or network hiccups, etc.) Kernel developers still have to use crash dumps to root-cause an individual crash, but crash dumps are most useful for extremely hard-to-reproduce crashes that are rare (but if you are using the "Pet" model as opposed to the "Cattle" model, even a single failure of a critical DB instances can't be tolerated). For crashes that are easy to trigger, crash dumps are useful, but they are much less critical to figure out what's going on. If your distributed architecture can tolerate rare crashes, then you might not even consider worth the support contract cost to root cause and fix every last kernel crash. Yes, it's ugly. But if you are administrating a very large number of systems, this can be a very useful way of looking at the world.
- rodgerd 9y ago> As the order of magnitude of systems administered increases, rare changes to occasional changes to frequent. Especially when it is not running in a VM. I think the perf engineer for Netflix is quite aware of this.
- severino 9y agoSolaris support ends in November, 2034. Yeah, 17 years from now. No need to hurry ;-)
- pjmlp 9y agoGiven that they fired everyone, who do you think is going to give that support and fix bugs?
- severino 9y agoMy comment was a joke, obviously. However, it's not that difficult to provide that kind of support and bugfixing. It's not like developing something new.
- mrpippy 9y agoThey didn't quite fire everyone: Alan Coopersmith (https://twitter.com/alanc/status/904366563976896512 https://twitter.com/alanc/status/904366563976896512) is still present. I'm sure lots of support and bug fixes will be neglected and probably moved offshore though. Also, it certainly seems like Oracle Solaris 11.3 (released in fall 2015) will be the last publicly available version. Between-release updates (SRUs) have always been for paying customers only, but now it seems like there will never be another release.
- tannhaeuser 9y agoHm, when was the next Unix timestamp range overflow again?
- cmurf 9y agoSmartOS might be easier for the Solaris familiar, looking to deploy Linux containers rather than go fully Linux.
- yellowapple 9y agoDoes SmartOS actually support Linux containers (aside from the obvious approach of running containers in a Linux VM)? Last I checked, SmartOS just used the word "container" to refer to Solaris Zones.
- qaq 9y agoyou checked a long time ago :)
- notnarb 9y agoSmartOS has code specifically in place for using 'lx' branded zones as wrappers for docker containers (which is distinct from its support of kvm) https://github.com/joyent/smartos-live/blob/master/src/dockerinit/README.md https://github.com/joyent/smartos-live/blob/master/src/docke... https://www.cruwe.de/2016/01/27/using-docker-on-smartos-hypervisors.html https://www.cruwe.de/2016/01/27/using-docker-on-smartos-hype... It's been a while since I've looked in to it, but if memory serves they are using docker filesystem snapshots without modification and running them on a thin translation layer of Linux system calls to Solaris system calls. Hard to find anything backing this up, so I could be way off the mark as to how it's implemented. EDIT: forgot that's what 'lx' zones are: zones which allow the execution of Linux binaries
- agentile 9y agoOmniOS anyone? https://omnios.omniti.com/ https://omnios.omniti.com/
- ssvss 9y agoArticle from 4 months back, > OmniTI will be suspending active development of OmniOS https://news.ycombinator.com/item?id=14179006 https://news.ycombinator.com/item?id=14179006
- realworldview 9y agohttps://www.theregister.co.uk/2017/07/13/open_source_community_resurrects_omnios/ https://www.theregister.co.uk/2017/07/13/open_source_communi...
- ptrott2017 9y agoSince OmniTI pulled back from the project - there has now been a reboot by the community. Regular releases started in early July with regular updates following with support for USB, ISO and PXE versions. The most recent release was August-28-2017. See http://www.omniosce.org/ http://www.omniosce.org/ for more info.
- ptrott2017 9y agoSee http://www.omniosce.org/ http://www.omniosce.org/ for most recent community release
- Annatar 9y agoImage packaging system, no preinstall, postinstall, preremove or postremove scripting, and a mishmash of Python and C, with that horrible automated installer thrown in for good measure, are you kidding? No way. SmartOS or bust.
- SrslyJosh 9y ago> If you absolutely can't stand systemd or SMF, there is BSD, which doesn't use them. You should probably talk to someone who knows systemd very well first, because they can explain in detail why you should like it. I can't imagine trying to sell anything with the phrase "why you should like it". SMF certainly doesn't need that kind of condescending pitch--it just fucking works and doesn't get in your way.
- scott_karana 9y agoIt helps that SMF doesn't include reimplementations of DNS resolvers and NTP clients, both of which have caused security flaws in systemd. :/
- bonzini 9y agoYou certainly can only use systemd's pid1 together with ntpd (systemd only does SNTP, not full-blown NTP) and no caching resolver. In fact it's the default in most Linux distributions. The only mandatory pieces if you use systemd as pid1 are udevd and journald. Misinformation like this is exactly why Brendan said you should ask an actual user of systemd.
- scott_karana 9y agoI actually like systemd as an init system, and use it that way on all my current machines. (I'm a fan of journald so far too.) However, the mere fact that (s)NTP and DNS are re-implemented in the same codebase is still unsettling to me. Nor is it misinformation to mention their existence, or the bugs/limitations caused by the duplication of effort. http://people.canonical.com/~ubuntu-security/cve/2017/CVE-2017-9445.html http://people.canonical.com/~ubuntu-security/cve/2017/CVE-20... https://github.com/coreos/bugs/issues/391 https://github.com/coreos/bugs/issues/391 (Looks like the sNTP client hasn't caused security flaws, though. I was wrong there.)
- bonzini 9y agoThey are independent, you can use one without the other. They are just residing in the same repository and build system.
- Veratyr 9y ago> Linux has also been developing its own ZFS-like filesystem, btrfs. Since it's been developed in the open (unlike early ZFS), people tried earlier ("IS EXPERIMENTAL") versions that had serious issues, which gave it something of a bad reputation. It's much better nowadays, and has been integrated in the Linux kernel tree (fs/btrfs), where it is maintained and improved along with the kernel code. Since ZFS is an add-on developed out-of-tree, it will always be harder to get the same level of attention. https://btrfs.wiki.kernel.org/index.php/Status https://btrfs.wiki.kernel.org/index.php/Status So long as there exists code in BTRFS marked "Unstable" (RAID56), I refuse to treat BTRFS as production ready. If it's not ready, fix it or remove it. I consistently run into issues even when using BTRFS in the "mostly OK" RAID1 mode. I don't buy the implication that "it will always be harder to get the same level of attention" will lead to BTRFS being better maintained either. ZFS has most of the same features plus a few extra and unlike BTRFS, they're actually stable and don't break. I'm no ZFS fanboy (my hopes are pinned solidly on bcachefs) but BTRFS just doesn't seem ready for any real use from my experience with it so far and it confuses me. Are BTRFS proponents living in a different reality to me where it doesn't constantly break? EDIT: I realize on writing this that it I might sound more critical of the actual article than I really am. I think his points are mostly fair but I feel this particular line paints BTRFS to have a brighter, more production-ready future than I believe is likely given my experiences with it. BTRFS proponents also rarely point out the issues I have with it so I worry they're not aware of them.
- the8472 9y ago> ZFS has most of the same features plus a few extra The ZFS feature set is not a strict superset of what btrfs offers. The ability to online-restripe between almost any layout combination is quite useful for example. So is on-demand deduplication, which is also far less resource-intensive than ZFS dedup.
- Veratyr 9y ago> The ability to online-restripe between almost any layout combination is quite useful for example This is true but since the only stable replication options on BTRFS are RAID1 and single, this online restripe is of very limited usefulness. I'd _really_ like BTRFS to fix its issues so I could use this reshaping (it's the main feature I'm missing from current filesystems) but it's been years and replication is still unstable.
- cbhl 9y ago> "Xen is a type 1 hypervisor that runs on bare metal, and KVM is type 2 that runs as processes in a host OS." Is it not the other way around, that KVM runs on bare metal (and needs processor support) while Xen runs as processes (and needs special kernel binaries)?
- detaro 9y agoNo, it's not. It's true that KVM needs processor support: it kind of adds a special process type that the kernel runs in an virtualized environment, through the hardware virtualization features. The linux kernel of the host schedules the execution of the VMs. Xen has a small hypervisor running on bare metal. It can run both unmodified guests using hardware support or modified guests where hardware access is replaced with direct calls into the hypervisor (paravirtualization). The small hypervisor schedules the execution of the VMs. For access to devices it cooperates with a special virtual machine (dom0), which has full access to the hardware, runs the drivers and multiplexes access for the other VMs - the hypervisor is really primarily scheduling and passing data between domains, very micro-kernel like. Dom0 needs kernel features to fulfill that role.
- cbhl 9y agoToday I learned something. Thanks!
- liuw 9y agoNo. See the Classification section. https://en.m.wikipedia.org/wiki/Hypervisor https://en.m.wikipedia.org/wiki/Hypervisor
- brendangregg 9y agoI updated this because, as someone pointed out, type 2 is no longer a good description of KVM. It uses kernel modules that access devices directly, so it's not strictly type 2 where it all runs as a process. Maybe I should just stop using the "types".
- cthalupa 9y ago
- luckydude 9y agoI'm ex-sun, kernel group. I'm retired, early at 55, but after reading Brendan's write up, I'd work for that guy. Holy smokes, he is all over it. Reminds me of me when I was in the groove. Brendan if you read this, I'm old school, not good at all the ruby on rails etc, but systems, yeah, pretty good. Not looking for money, looking for working with smart people, I can be paid in stock. If it works out, great, if it doesn't still great because I like smart people. Sorry for making it about me, you all should read his post, it's someone who is completely in the groove, has breadth and depth, knows about systems. These people are rare, but hugely valuable when you need to scale stuff up.
- jamesmishra 9y agoIf you are indeed interested in coming out of retirement, Brendan Gregg does performance engineering at Netflix. It might be worth reaching out. http://www.brendangregg.com/blog/2017-05-16/working-at-netflix-2017.html http://www.brendangregg.com/blog/2017-05-16/working-at-netfl...
- luckydude 9y agoI'm kinda old and burned out but thanks for the link. I'll reach out. It could be a boatload of fun.
- rurban 9y agoIt's even much more fun than described in this article, because Netflix mostly works with FreeBSD not Linux. So no NIH syndrom, proper dtrace (no eBFS hacks), proper ZFS, proper tooling, easy kernel maintenance.
- gribbly 9y agoAs I've seen it described here before (by Brendan Gregg ?) Netflix uses FreeBSD for the CDN servers (streaming the video), and Linux for everything else, browsing Netflix, encoding, etc.
- sandGorgon 9y agoI think two things need to be mentioned : 1. ZFS is officially supported by Canonical on Ubuntu as part of their support plans. 2. Docker over raw containers or zones.