7 ms·
Umm sorry, no. I'm a storage engineer with one of the leading enterprise vendors. SSDs are already "short-stroked." They have much more space available intern
by usernew 3y ago
Umm sorry, no. I'm a storage engineer with one of the leading enterprise vendors. SSDs are already "short-stroked." They have much more space available internally than you can see through the controller externally. The more "enterprisey" the drive, the more of that hidden space there is.
A wise sysadmin buys a drive with more of that hidden internal space, instead of getting a larger drive with less hidden space. The logic inside the drive is much better at zeroing/out and background trim on that internal hidden space, than host-addressable blocks that are currently unused. That's because the drive has no idea if you're about to use them, while the hidden space is guaranteed to be always free.
In fact - a fun fact. Flash has been used as a Write buffer on storage arrays for a couple of decades.
- wmf 3y agoA wise sysadmin buys a drive with more of that hidden internal space, instead of getting a larger drive with less hidden space. The pricing of enterprise SSDs is... not exactly fair for the performance you get.
- usernew 3y agothat's like complaining a bentley is too expensive and not good at moving your piano. when you get fined $1mil/min by the SEC for downtime, and data loss or corruption can cost you in the hundreds of millions, or someone dying at a hospital, enterprise gear is cheap. reliability is the key and for what you are paying. now yes, there may be a specific consumer drive more reliable than a specific enterprise drive. so which do you buy? well the enterprise array vendor tested the crap out of everything in every combination and workload and environment, and picked one for you, and it comes with full support and SLAs that you can use to meet gov regulations. performance for an enterprise drive is not something you usually consider, at all. in fact, did you know that when I quote a storage array, I can't even specify the drive type of vendor, and who knows what will get shipped? Ionly specify drive size. your performance comes from all your workloads clumped together, spread over a thousand of these drives connected with infiniband, with dedupe and compression on the backend, and hundreds of terabytes of RAM. the perf of an individual drive is not relevant. but yes, sticking it into an AMD server you bought on newegg when they had a sale is not its purpose, and is a very bad deal. now when we talk about wise sysadmins, we're talking about guys who know their stuff, and do "important big stuff." not a guy at a small business ordering from newegg. and that guy - he shouldn't be coming up with any storage policies because he lacks the needed large-scale experience. I sold a 1PB usable-effective (after 3x dedupe/compression) space storage array last year. It was $2mil, after a 65% discount. If one time in its 5year lifecycle the "enterprisey" stuff on that array prevents about 10 seconds of downtime, it paid for itself.
- wongarsu 3y agoSure, and judging by System Z sales there's still a demand for earthquake-resistant machines with hot-swappable CPUs. But for most of us, infrequent downtime is acceptable, and most single-machine downtimes are either automatically mitigated or have very limited impact. In that scenario, getting a prosumer SSD and over-provisioning it can be a sensible choice.
- geraldhh 3y ago* Hetzner joined the chat
- renonce 3y agoThere are many other ways to get reliability at the software level, usually via redundancy. Even enterprise level SSDs doesn't guarantee you free of data losses and downtimes, you have to back them up regularly and preferably set up high availability for the database.
- usernew 3y agocool. so I have a few thousand VMs of various OSs and hypervisors, a few hundred home-built applications with an average of 10 components each, some mainframe, some ibm-i, a bunch of linux, solaris, aix, and HP-UX server, and I'm using mysql, sql server, postgres, oracle, and a couple of DB2. there may be an informix DB or two somewhere - I don't know, I manage storage. about ten different volume managers, about 30 different filesystems - and some raw disks of course. 10PB total space, it needs to be metro cluster replicated ten miles away for active/active, and async in a different state can be about a minute behind max. So, write me some software that's going to make all that work and make the storage highly available. make sure it works for everything. you're exposed to solutions that make that work, daily. when you swipe your credit card buying condoms - you have no idea how much happens so your transaction doesn't get lost, corrupted, or errored out.
- the_third_wave 3y ago
- bradfa 3y agoIt depends on what kind of performance you're looking for and how you value that performance. Under sustained long duration (ie: hours to days) of continuous full load performing small block writes, even the highest rated consumer SSD drives will begin to show significantly reduced throughput in comparison to pretty much any enterprise SSD drives. Additionally, it's totally normal to see >=5 drive writes per day and 5 year warranties on enterprise drives. Consumer drives usually are rated significantly below 1 drive write per day and rarely for more than 3 years of warranty. If you're performing a lot of writes, you're going to need to replace worn out SSDs and that's not free in a business (remote hands, downtime, etc). So buying less durable storage has costs which don't show up on the original purchase invoice but do need to be factored in within a business setting. The warranty period is the best indication of how confident the drive vendor is about drive durability.
- Retric 3y agoMore total space for the same price changes that equation. Spending 2+X per GB means rather than 5x the same X it’s 5 times 1/2 or less total space. And that’s before you consider write amplification issues with having dramatically less total SSD space. Enterprise SSD’s have a few benefits, but they shouldn’t be your default choice for all servers.
- vetinari 3y ago> Consumer drives ... and rarely for more than 3 years of warranty. Looking at my list of consumer SSDs and their warranties: Samsung EVO 850: 2y Samsung EVO 960: 3y Samsung EVO 860: 5y Samsung EVO 970: 5y Samsung PRO 970: 5y For comparison with spinning rust HDDs: Seagate Ironwolf: 3y WD Red (both EFAX & EFRX): 3y
- bradfa 3y agoBut take a wider sample of consumer SSDs, a large majority are <=3 years. And look at enterprise SSDs. Seagate Ironwolf and WD Red are not enterprise spinny hard disk drives, instead look at Seagate X18 or WD HC560.
- marginalia_nu 3y agoMost of this hides the problem. It still exists though. I have a piece of logic in my indexing code that essentially transposes an ~100 Gb multi-value dictionary on disk. If you do this the naive way with random writes, the write amplification makes it a complete non-starter. All the caching layers and buffers fill up with completely disjointed 8 byte writes and it takes ages to write. What I've ended up doing is to in an intermediate stage write the data to be written into a series of files, containing pairs of offsets and data (up to like 100Mb each); and then going over the files one by one and essentially evaluating them as assembly instructions. Both passes have relatively good data locality, and despite essentially writing 2.5X as much shit to disk, it takes hours rather than weeks to do this operation.
- sumtechguy 3y agoShort stroke in a SSD era is more about trying to keep your data within a page of a SSD. Depending on the flash drive you are using the controller may or may not do some of it for you. One of the easiest perf gains I can sometimes get from a program that is write intensive is to put a small amount of write buffer into the mix. Depending on your OS that can help a lot too. Basically keep my code out of the kernel and off the bus and keep the write block to something that closely resembles what the drive considers a block. If you are doing a bunch of small 8 byte writes randomly in your file you probably are having a bad time as you may quickly exceed the amount of buffer the drive has for that sort of thing. It will start backing you off which bleeds up into the kernel space then very quickly into your program. Keeping them together can help, as you found out. I see this sort of issue in a lot of programs. As it is a dead easy problem to make in your program. You need to write something out you just splat it out somewhere. With a dozens 1/4/8/16 byte writes instead of one big write. Basically not thinking about how that data is getting into your files. Most of the time that is just fine and not that big of a deal. But as your data set grows or you want better perf you have to worry about it. I usually use something like filemon and can see what is going on. You can see the pattern where there will be a large stack of I/O with hundreds of very small read and writes. While SSDs are an order of magnitude faster than the older drives. They still have their command structure and kernel context switching you have to deal with. You in some cases want to minimize that as it can become a large portion of writing and reading data. As with most optimizations (in this case the drive is faster) we just moved where the bottleneck is (to the kernel and bus typically).
- bcrl 3y agoI am well aware that flash is overprovisioned. One consumer SSDs the overprovisioning is quite small (maybe 5-10%), so short stroking the SSD will have a much greater effect than on enterprise SSDs. It also has the nice side effect of giving the drive more free space to work with which allows the garbage collection room to be more gradual and efficient, which can improve performance for more write intensive workloads. Actually, many storage arrays use battery backed RAM, not flash as a write buffer. Flash does not have the endurance needed to serve as the write buffer for a large storage array. Some products that use battery backed RAM for this purpose will dump the contents of that RAM onto flash. I worked on a messaging appliance that used that approach, albeit with supercaps to provide the hardware time to dump 4GB of DRAM onto a compact flash card. Supercaps were easier to monitor and maintain. There are also plenty of hardware RAID cards that use batteries for their write buffer as well. Edit: there are also persistent memories like MRAM that don't need power to retain their contents which are used in this space as well.
- usernew 3y agoyes, big arrays (symmetrix, shark) have battery backed everything - enough to take all the cache and dump it to disk when power is lost. terabytes of RAM. they also often use use a flash cache behind the RAM - that was the primary use for Optane. The reason for this, is sustained writes instead of peaks. If you let more and more sit in cache, eventually you have write folding. That results in less write IO on the backend, and you're able to have a higher sustained peak. smaller cheaper arrays (unity, pure, isilon) and HCI (nutanix, vxrail) that don't have terabytes of RAM - they have gigabytes, and pretty much always use flash as a write cache. In fact, I cannot name one that doesn't. No one in enterprise storage cares about the indurance of flash. All that means is that twice a year, a vendor engineer comes out to replace a few flash drives under your support contract. And flash has been used as a write cache behind a smaller RAM cache, for two decades, by all major storage vendors.