9 ms·
Persisting state between AWS EC2 spot instances
- deleted 9y ago[deleted]
- stonewhite 9y agoWe normally utilize spots with Spotinst + Elasticbeanstalk. Our billing looked great ever since. This solution looks good, yet only applies to single instance scenarios. I presume this kind of thinking might move forward with EFS + chroot for an actual scalable solution that cannot be ran on Elasticbeanstalk.
- olegkikin 9y agoIf you don't care about reliability, why not just get a cheap and powerful VPS? Paying $90/month for that machine is madness. I pay $6/month for 6GB RAM, 4 cores, 50GB disk.
- yjftsjthsd-h 9y agoPerhaps integration with other AWS services?
- mayank 9y agoAWS Lightsail is AWS’s option there.
- blibble 9y agowhere are you getting that for $6?
- amq 9y agoNot quite $6, but close: https://www.scaleway.com/pricing/ https://www.scaleway.com/pricing/
- nik736 9y agoIf you don't need any connection scaleway is a good option, since they tend to absolutely not care about their network quality and reliability at all.
- olegkikin 9y agoHostus.us The deal was found on LowEndBox, not sure if it's still available, but there are many other ones.
- deivid 9y agoWhere? I'm using Digital Ocean and it'd be way more expensive for that kind of configuration.
- sacheendra 9y agoDigital Ocean is still a premium provider. I would look at providers like OVH and even cheaper (Treudler, TransIP, RamNode, etc.) For example, an SSD with 2 vCPUs, 8GB RAM and 40GB SSD is 13.49$ per month from OVH.
- kuschku 9y agoHere’s a list of providers by cost: https://git.io/vps https://git.io/vps (PS: Don’t use DigitalOcean, they tend to steal your credit if they feel like it. Lost 100 bucks "promotional credit" that way with only a few days notice)
- dewey 9y agoThey expired some credits that I haven't used but after asking they just restored them and I could use them. Asking helps.
- kuschku 9y agoAsking is not a solution. This is a question of trust. I have to trust that DO will keep my data safe, that, if the US government would be after my data, DO would prevent them from accessing it. I have to trust that DO won’t access my data. How am I supposed to trust my, and my user’s personally identifying data, to a company that just like that revokes credit, without warning, and says "well, if you ask nicely, you can get it back"?
- Johnny555 9y agoBilling practices and data privacy seem like completely different subjects and I'd be surprised if there's much correlation between the two.
- dagw 9y agoIf you don't care about reliability, why not just get a cheap and powerful VPS? Personally, because my needs aren't constant. I might need two cores for two months followed by 100 cores for a week.
- yjftsjthsd-h 9y agoTLDR: Attach EBS volume and use that to store Docker containers. I suppose it's a decent solution if you don't want to deal with prefixes.
- amq 9y agoWouldn't it be simpler to have the smallest possible instance run an NFS server? This would also have an additional bonus of scalability. Edit: or use AWS EFS
- otterley 9y agoEFS is far more expensive than EBS. Price it out; you'll see.
- samstave 9y agoAnd while we are talking about costs, make sure you check for unused WBA volumes frequently, as you still pay for them if they aren't attached/used - and sometimes a dev will create a provisioned iops drive and forget to delete it and you pay a lot for those volumes..
- Johnny555 9y agoIt is 3X more expensive ($0.30/gb vs $0.10/gb for us-east), but it's replicated across AZ's (so is more durable than EBS which is only replicated within an AZ), and you only pay for what you use, you don't need to overprovision the EBS volume to account for peak dataset size. And since it's shared, you don't need to replicate data across multiple nodes... so if 10 compute nodes needs access to the data set, they can all just read it from the same EFS filesystem, no need to download it 10 times to each compute node. So EFS can still be very cost effective compared to EBS.
- otterley 9y agoAre you counting the impact on the ENI's available bandwidth and additional instance costs needed for more network throughput? As I understand it, EFS requests are issued through the front end interface, while EBS requests go through the storage backplane interface. Also, NFS has different behavior with respect to buffer caching that needs to be taken into account. It often does not cache as effectively as block storage does.
- manigandham 9y ago
- Pirate-of-SV 9y agoTo make this even more streamlined you'd tag the volumes and discover the volumes with `aws ec2 describe-volumes` and filter unattached volumes with the magic tag.
- sevagh 9y agoThere's a handful of tag-based automatic EBS volume attachers out there: * https://github.com/sevagh/goat https://github.com/sevagh/goat (my own) * https://github.com/UKHomeOffice/smilodon https://github.com/UKHomeOffice/smilodon
- tuananh 9y agowe use k8s at work. i just have to create PVC and when spot instance terminated along with the container; new container will be created and mount the PVC again automatically.
- otterley 9y agoEven if you don't use spot instances, the technique of using separate EBS volumes to hold state is useful (and well-known). Ordinary on-demand instances can also be terminated prematurely due to hardware failure or other issues, so storing state on a non-root volume should be considered a best current practice for any instance type.
- manigandham 9y agoPersistent storage remains a complicated problem. Attaching volumes on the fly with docker volume abstraction works well enough for most cloud workloads, whether on-demand or spot, but it's still easy to run into problems. This is leading to rapid progress in clustered/distributed filesystems and it's even built into the Linux kernel now with OrangeFS [1]. There are also commercial companies like Avere [2] who make filers that run on object storage with sophisticated caching to provide a fast networked but durable filesystem. Kubernetes is also changing the game with container-native storage. This seems to be the most promising model for the future as K8S can take care of orchestrating all the complexities of replicas and stateful containers while storage is just another container-based service using whatever volumes are available to the nodes underneath. Portworx [3] is the great commercial option today with Rook and OpenEBS [4] catching up quickly. 1. http://www.orangefs.org http://www.orangefs.org 2. http://www.averesystems.com/products/products-overview http://www.averesystems.com/products/products-overview 3. https://portworx.com https://portworx.com 4. https://github.com/openebs/openebs https://github.com/openebs/openebs
- manigandham 9y agoAlso want to highlight that AWS will now allow spot instances to just be stopped instead of terminated, so only compute power is removed but data is persisted automatically as long as you use EBS root/attached volumes. https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec2-spot-can-now-stop-and-start-your-spot-instances/ https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec...
- objectivefs 9y agoUsing a clustered/distributed filesystem definitively simplifies persisting the state between EC2 spot instances. It also makes it easier to scale out the work load when you need more instances accessing the same data. To add to your list: there is also ObjectiveFS[1] that integrates well with AWS (uses S3 for storage, works with IAM roles, etc) and EC2 spot instances. [1]. https://objectivefs.com https://objectivefs.com
- manigandham 9y ago
- fulafel 9y agoThere's a mechanism exactly for this purpouse in Linux: pivot_root. It's used in the standard boot process to switch from the initrd (initial ramdisk) environment to the real system root. ec2-spotter classic uses this, but you can also make a pivoting AMI of your favourite Linux distribution. One thing to watch out for is how to keep the OS automatic kernel updates working. AMIs are rarely updated and you're going to have a "damn vulnerable linux" if you don't get the updates just after booting a new image.
- js4all 9y agoWhen you are using Kubernetes, you won't have to deal with this yourself. The Cluster will move pods from nodes that are stopped because the spot price is exceeded. Ideally place nodes at different bids. So there will be a performance hit but no outage. With the new AWS start/stop feature [1] nodes will come up again when the spot price sinks. 1) https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec2-spot-can-now-stop-and-start-your-spot-instances/ https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec...
- ramanan 9y agoWell, one easy way when using Ubuntu-like distributions is to simply place your `/home` folder on a separate (persistent) EBS volume [1]. With a few on-boot scripts to attach-volumes / start-containers, it should be fairly easy to get going as well. [1] https://engineering.semantics3.com/the-instance-is-dead-long-live-the-instance-8b159f25f70a https://engineering.semantics3.com/the-instance-is-dead-long...
- TrickyRick 9y agoThis was exactly what I was thinking, why complicate things by replacing the root volume when one can simply mount the disk to any other directory and point the application there?
- raverbashing 9y agoIs it just me or to me spot instances should deal with work and not storage, and hence your (stateful) units of work should be in a Queue/DB? (in a non-spot instance) Attaching and detaching volumes is a good idea but I wouldn't use that to keep state
- solatic 9y agoOP is offering some very dangerous advice. Twenty years ago, software was hosted on fragile single-node servers with fragile, physical hard disks. Programmers would read and write files directly from and to the disk, and learn the hard way that this left their systems susceptible to corruption in case things crashed in the middle of a write. So behold! People began to use relational databases which offered ACID guarantees and were designed from the ground up to solve that problem. Now we have a resource (spot instances) whose unreliability is a featured design constraint and OP's advice is to just mount the block storage over the network and everything will be fine? Here's hoping OP is taking frequent snapshots of their volumes because it sure sounds like data corruption is practically a statistical guarantee if you take OP's advice without considering exactly how state is being saved on that EBS volume.
- jen20 9y agoThis pattern is a lot safer if you use ZFS. Spot instances don't just disappear though, you get notification and have a chance to perform shutdown actions, except in the case of hardware failure - which is the same with non-spot instances.
- solatic 9y ago- EBS, being block storage, doesn't recognize the filesystem format on top of it, and therefore doesn't recognize if you formatted the block storage as ZFS and therefore will not use ZFS snapshots when using Amazon's native EBS snapshotting. If you wish to use ZFS snapshots, you have to build that on top of what Amazon gives you, along with all the other aspects of ZFS storage, i.e. building a ZFS storage pool from separate EBS volumes. I mean, it would be nice if Amazon had a hosted ZFS solution, but so far, doesn't seem like it. - Yes, you get a notification, but it's a proprietary notification scheme that your application must be designed to poll for. Why can't Amazon use standard signals like SIGPWR to indicate imminent shutdown? - Just because it isn't smart for non-spot instances doesn't suddenly make it smart for spot instances ;)
- subway 9y ago
- sciurus 9y agoThe author goes to great lengths to come up with a way for the software that was running on a terminated spot instance to be relaunched using the same root filesystem on a new spot instance, but they never explain why they need to do exactly this. Maybe they already ran everything in Docker containers on CoreOS, so their solution isn't a big shift, but I strongly suspect they could find a simpler way to save and restore state if they got over this obsession with preserving the root filesystem their software sees.
- bdcravens 9y agoSpot instances can now "stop" instead of "terminate" when you get priced out, persisting the attached EBS volumes: https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec2-spot-can-now-stop-and-start-your-spot-instances/ https://aws.amazon.com/about-aws/whats-new/2017/09/amazon-ec...
- alex_duf 9y agoIt sounds wrong to try to keep the state across two ec2 instances. If you find yourself in that situation, try pushing your state outside the ec2 instance a bit harder. (dynamodb, s3 etc...) You will get a lot of benefit out of it, but may lose in performance, which is fine in 99% of the cases.
- likelynew 9y agoI don't know why all the comments are saying this is bad idea. For me, one of thing for I use EC2 is deep learning. I just use spot GPU instance, attach overlayroot volume and launch jupyter notebook in it. Other things like google dataflow is not useful to me due to the price and the process of installing packages. I can also think of many other use cases for using some persistence volume for some manual task.
- archgoon 9y agoSo I was pleasantly surprised to discover that for the last several years, spot instances have provided a mechanism that give you 2 minutes notice prior to shutdown: http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-interruptions.html http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-inte... Learn something new everyday. :) https://aws.amazon.com/blogs/aws/new-ec2-spot-instance-termination-notices/ https://aws.amazon.com/blogs/aws/new-ec2-spot-instance-termi...
- bdcravens 9y agoSee my top-level comment - you can now set "shutdown" behavior to stop instead of terminate (though 2-minute notice still useful)
- jdchernofsky 9y agoOr you could just use Spotinst: https://spotinst.com/ https://spotinst.com/