7 ms·
It doesn't work like that. 2TB drive will always be faster than an 8TB drive. The amount of data has no effect when compared to the physical attributes of the d
by lykron 10y ago
It doesn't work like that. 2TB drive will always be faster than an 8TB drive. The amount of data has no effect when compared to the physical attributes of the drive. More platters will increase the response time.
Ceph seems to offer Tiering, which would move frequently accessed data into a faster tier while the infrequent data to a slower tier.
- matt4077 10y agoThough 2TB of data on an 8TB drive does mean only 1/4 as many requests hit it, right?
- lykron 10y agoI'm not sure I understand your question. A 2TB/2TB disk will have the same number of requests as a 2TB/8TB disk, as they both have the same amount of data. If you are talking physically, 2TB/8TB would theoretically be faster than a 2TB/2TB disk if the performance attributes were the same. But a 2TB HDD will have a faster average seek time than an 8TB HDD due to physical design. Any performance gains of only partially filling a drive would probably be offset by the slower overall drive.
- Dylan16807 10y ago> a 2TB HDD will have a faster average seek time than an 8TB HDD due to physical design. Any performance gains of only partially filling a drive would probably be offset by the slower overall drive. I'm skeptical here. So your minimum seek time goes up on the 8TB because it's harder to align. But your maximum seek time should drop tremendously because the drive arm only has to move along the outer fifth of the drive. And your throughput should be great because of the doubled linear density combined with using only the outer edge.
- lykron 10y agoI'm not saying that you wouldn't see any performance gain. 2TB on the outer track will be faster than 8TBs on the same 8TB disk, but I'm saying any gains will be lost due the dense nature of the drives. A quick google search shows that there are marginal gains on the outer track vs the inner, but that is only on sequential workloads. For something like GitLab, the workloads would be anything but. http://superuser.com/questions/643013/are-partitions-to-the-inner-outer-edge-significantly-faster http://superuser.com/questions/643013/are-partitions-to-the-... https://www.pythian.com/blog/hard-drive-inner-or-outer/ https://www.pythian.com/blog/hard-drive-inner-or-outer/
- Dylan16807 10y agoIgnore the part about where the partition is, then. 1. If I look at 2TB vs. 8TB HGST drives, their seek times are 8ms and 9ms respectively. But if you're only using a quarter of the 8TB drive, the drive arm needs to move less than a quarter as much. Won't that save at least 1ms? 2. The 8TB drive has a lot more data per inch, and it's spinning at the same speed. Once a read or write begins, it's going to finish a lot faster. 3. Here's a benchmark putting 8TB ahead in every single way http://hdd.userbenchmark.com/Compare/HGST-Ultrastar-He8-Helium-8TB-vs-Hitachi-UltraStar-7K4000-2TB/m29773vsm10357 http://hdd.userbenchmark.com/Compare/HGST-Ultrastar-He8-Heli...
- toomuchtodo 10y ago> Ceph seems to offer Tiering, which would move frequently accessed data into a faster tier while the infrequent data to a slower tier. By "Tiering", is this moving data between different drive types? Or by moving the data to different parts of the platter to optimize for rotational physics?
- lykron 10y agoBy Tiering, I'm talking about moving blocks of data from slower to faster or visa verse. If I have 10TB of 10k IOPS storage and 100TB of 1k IOPS storage in a tiered setup, data that is frequently accessed would reside in the 10k IOPS tier while less frequently accessed data would be in the 1k tier. In this case, the blocks of popular repositories would be stored in SSD, while the blocks of your side project that you haven't touched in 4 years would be on the HDD. You still have access to it, it might just take a bit longer to clone. Ceph can probably explain it better than I can. http://docs.ceph.com/docs/jewel/rados/operations/cache-tiering/ http://docs.ceph.com/docs/jewel/rados/operations/cache-tieri...
- sytse 10y agoThat is pretty awesome. Should the SSD's for the fast storage be on the OSD nodes that also have the HDD's or should it be separate OSD nodes?
- lykron 10y agoThat would be something I would test. I don't know Ceph, so I would be taking a shot in the dark. I would guess it would not make much of a difference as everything is block level. I, personally, would do 1x PCIe SSD for cache, 1x 2/4TB SSD, and 2x 4TB HDD for each storage node. Edit. If Ceph is smart enough, it would be aware of the tiers present on the node, and would tier blocks on that node. So a Block on node A will stay on node A.
- illumin8 10y agoThe fact that you're asking this question on Hacker News leaves little doubt in my mind that you and your team are not prepared for this (running bare metal). I read the entire article, and while you talked about having a backup system (in a single server, no less!) that can restore your dataset in a reasonable amount of time, you have no capability for disaster recovery. What happens when a plumbing leak in the datacenter causes water to dump into your server racks? How long did it take you to acquire servers, build the racks and get them ready for you to host customer data? Can your business withstand that amount of downtime (days, weeks, months?) and still operate? These questions are the ones you need to be asking. In other words, double your budget because you'll need a DR site in another datacenter.