3 ms·
> No Sans Ephemeral disks? Or persistent local? > Our own stack Very cool ^_^ How do you deal with geographic zones? Are they silod? > Bandwidth Ok, makes
by asharp 15y ago
> No Sans
Ephemeral disks? Or persistent local?
> Our own stack
Very cool ^_^ How do you deal with geographic zones? Are they silod?
> Bandwidth
Ok, makes perfect sense. Thanks.
> Securing KVM
Do you use cfgroups/selinux to deal with compromise of a kvm domain? I've seen quite a few vulnerabilities coming out on the debian/etc. security mailing lists.
- comice 15y ago> Ephemeral disks? Or persistent local? Persistent local disks (hardware raid6 15k rpm). More storage options on the roadmap too. > Very cool ^_^ How do you deal with geographic zones? Are they silod? Our zones are different datacenters in different buildings, with completely different power supplies, UPSes and backup generators. > Do you use cfgroups/selinux to deal with compromise of a kvm domain? cgroups currently, selinux in development. (full disclosure: I'm a Brightbox bod too!)
- asharp 15y ago> persistent disks How do you then deal with people who create an instance and then don't run it. Unless i'm mistaken, you'd be forced to either unbalance for storage or VM usage. > cgroups currently How do you protect the kernel from something like CVE-2011-2212? Quite cool otherwise :)
- rednaught 15y agoNot sure what distro he is using but Debian and RHEL have patches for this. http://security-tracker.debian.org/tracker/source-package/qemu-kvm http://security-tracker.debian.org/tracker/source-package/qe... http://security-tracker.debian.org/tracker/CVE-2011-2212 http://security-tracker.debian.org/tracker/CVE-2011-2212 https://rhn.redhat.com/errata/RHSA-2011-0919.html https://rhn.redhat.com/errata/RHSA-2011-0919.html
- asharp 15y agoI know of this. The interesting question it brings up is how do you keep a cloud like this patched and up to date without dropping SLA?
- rednaught 15y agoPatching is normally considered part of scheduled or emergency maintenance and therefore doesn't count against the SLA for uptime. This is fairly standard in the hosting/ISP world. So much of this can be automated now that it is not a problem. As a provider myself, I allow customers to pick their patch day/time. They can even manually push patches themselves and be present to test when the service comes back up. Proactive maintenance(datacenter, networking, hardware, OS, and appptack/utils) should be considered a way of life these days if you're a provider. If customers don't understand or agree with that, then there are plenty of providers who don't keep up-to-date offerings that they can migrate to.
- comice 15y agoYou start paying for an instance as soon as you create it, whether you're running it or not. To keep a disk image around for an instance without paying for the instance, you can use snapshots, which use a different storage system. We protect against those types of problems at the moment with keeping patched up to date. Xen doesn't solve this problem either, it has had it's own share of vulnerabilities with these kinds of repercussions. Even selinux only mitigates some of the risks - not all. A combination of mandatory access control and a good update, audit and monitoring strategy is the best approach imo.
- asharp 15y agoInteresting choice, and I'd imagine that it sidesteps all of the balance issues. On the other hand, what now separates you from a VPS provider with an API? (like say Linode) Interesting. I agree with your approach. When I was looking at KVM, I noticed its rather insane surface area (Everything is in the kernel as a kernel module) and that almost all of the vulnerabilities found seem to stem from that arch decision. As an example have a look at this: nelhage.com/talks/kvm-defcon-2011.pdf I was just wondering if you knew of any way to secure that down, or any way to patch quickly enough that you don't break SLA by having to forever reboot people. Just generally given that we've had CVE-2011-2212 CVE-2011-2527 CVE-2011-1751 CVE-2011-0011 CVE-2011-1750 for kvm itself in the last few months, so say you have 5 critical bugs in kvm a year. If you need to restart everybody to patch and it takes a minute or so to restart per vm and you have say 40 vms/box (random numbers for the sake of argument), that's 3 hours of downtime/year not counting kernel upgrades/etc. That means the best you can do is 99.9% uptime not even considering transit failure/dc power failure/etc. So my question is how do you deal with that?
- comice 15y agoCVE-2011-2527 and 0011 are local attacks only, guest can't trigger it. CVE-2011-2212, 1751 and 1750 are serious, but mitigated in various ways making actual exploitation difficult (and conspicuous). selinux will mitigate it even further. Remember that all of these fixes are upgrades to userspace - which means there are many more upgrade options than if they were kernel or hypervisor (think the equivalent of a live migrate to the same host, as an example). Btw, Brightbox happened to be the original reporters of CVE-2011-0011 :)