5 ms·
For comparison here's the nvme smart-log for my 2.5 year old X1 Carbon running Linux - used for VSCode/compiles and browsing/ssh etc. carbonx ~ sudo nvme
by blinkingled 6y ago
For comparison here's the nvme smart-log for my 2.5 year old X1 Carbon running Linux - used for VSCode/compiles and browsing/ssh etc.
carbonx ~ sudo nvme smart-log /dev/nvme0n1
Smart Log for NVME device:nvme0n1 namespace-id:ffffffff
critical_warning : 0
temperature : 29 C
available_spare : 100%
available_spare_threshold : 10%
percentage_used : 0%
endurance group critical warning summary: 0
data_units_read : 12,862,645
data_units_written : 16,193,903
host_read_commands : 124,284,317
host_write_commands : 218,295,105
controller_busy_time : 567
power_cycles : 3,078
power_on_hours : 558
unsafe_shutdowns : 209
media_errors : 0
So around 8.2 TB written in 2.5 years of daily use. And I have had syncthing running on it past few months to sync some git repos as an experiment. (And I run a rolling release distro so lots of updates.) In comparison the M1 data posted here looks out of whack.
I generally use bpftrace to find if anything keeps writing to my SSD - for the most part I don't find any misbehaving programs on modern Linux distros. Assuming dtrace still works on M1 Macs you might be able to find what is writing to the disk.
- eulers_secret 6y agoAnother Linux user data point: XPS13(9310)/512GB nvme/16GB RAM/2GB swapfile/kubuntu 20.10/kernel 5.11 (in use since ~Dec 5, 2020): SMART/Health Information (NVMe Log 0x02) Critical Warning: 0x00 Temperature: 33 Celsius Available Spare: 100% Available Spare Threshold: 50% Percentage Used: 0% Data Units Read: 631,319 [323 GB] Data Units Written: 1,157,883 [592 GB] Host Read Commands: 4,650,820 Host Write Commands: 7,329,774 Controller Busy Time: 16 Power Cycles: 37 Power On Hours: 45 Unsafe Shutdowns: 15 Media and Data Integrity Errors: 0 Error Information Log Entries: 1 Warning Comp. Temperature Time: 0 Critical Comp. Temperature Time: 0 I do Linux kernel dev and have done quite a few kernel compiles at this point (using Ubuntu's full config) as well as general browsing and some Docker workloads and Spotify... Judging by 'Power On Hours' combined with 'Data Units Written', looks like there's a bug to me. As a side, I don't know why 'Available Spare Threshold' is 50%, and the 1 'Error Information Log Entries' appears to be successful results of some test.
- ryan29 6y agoI have the same question here. How do you only have 45 power on hours in 80 days? Does anyone know why power on hours seem to be under reported so often?
- frant-hartm 6y agoJust a guess, but maybe it doesn't count time in low power modes?
- curryst 6y agoAs it relates to the Mac's, I'm starting to wonder if their SMART reporting is just faulty. Power on Hours seems to be definitively wrong, I'm starting to wonder if all of the SMART data is bogus and there is no issue.
- the8472 6y agoOn linux there are a bunch of "laptop mode" configurations that can minimize the time the disk will be woken up to actually write out data. The price is that you'll lose the last N minutes of work on a hard crash. And it only works when you have enough RAM to avoid swapping and keep dirty pages in memory. And your workloads don't explicitly call fsync. But if setup correctly your drive will spend most of its time sleeping.
- Filligree 6y ago> And your workloads don't explicitly call fsync. If you run ZFS, you can set sync=disabled for your filesystems. This will disable fsync. Unlike most (all?) other filesystems, that's actually safe. ZFS doesn't reorder writes between transaction groups, so after a crash you'll get a consistent state from however many minutes ago. (However, txgs have a time limit of 5 seconds by default. You also need to increase that.)
- ryan29 6y agoHow do you only have 558 PoH with ~900 days of daily use? Do you only use it for 30m per day? Here's a 512GB WD Black from my home lab server. It runs ~10 super light usage VMs with Docker stacks that include 2 GitLab instances, 2 GitLab runners, 3 Nextcloud instances, 3 Redmine instances, 1 Gitea instance, 2 Drone runners, 1 Minio instance, 1 Nexus instance, 1 Emby instance (w/transcoding), and various reverse proxies, etc.. The write endurance is supposed to be 300TBW, so it should really be over 20% used but says 0%. SMART/Health Information (NVMe Log 0x02) Critical Warning: 0x00 Temperature: 38 Celsius Available Spare: 100% Available Spare Threshold: 10% Percentage Used: 0% Data Units Read: 95,066,245 [48.6 TB] Data Units Written: 135,910,315 [69.5 TB] Host Read Commands: 972,466,473 Host Write Commands: 2,504,003,547 Power On Hours: 17,383 Error Information (NVMe Log 0x01, max 256 entries) No Errors Logged Compare it to a 500MB Crucial MX500 under the same load which is supposed to have 180TBW: Model Family: Crucial/Micron BX/MX1/2/3/500, M5/600, 1100 SSDs Device Model: CT500MX500SSD1 Serial Number: Sector Sizes: 512 bytes logical, 4096 bytes physical Rotation Rate: Solid State Device Form Factor: 2.5 inches ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE 9 Power_On_Hours 0x0032 100 100 000 Old_age Always - 11225 173 Ave_Block-Erase_Count 0x0032 032 032 000 Old_age Always - 1032 194 Temperature_Celsius 0x0022 067 044 000 Old_age Always - 33 (Min/Max 0/56) 202 Percent_Lifetime_Remain 0x0030 032 032 001 Old_age Offline - 68 246 Total_Host_Sector_Write 0x0032 100 100 000 Old_age Always - 59166051308 That's 68% used after ~30TBW (59,166,051,308 sectors*512 = 30,293,018,269,696 bytes). What I've learned from trying to diagnose those MX500s is that TBW doesn't really matter all that much. It's the P/E cycles that really count. For example, the MX500s are rated for 1500 erase cycles (#173 above) IIRC. I've also become skeptical of many SMART implementations. I know the PoH on the Crucials is incorrect because when they were new they were reporting 45 days of PoH on a system with 76 days of uptime. So if the manufacturers can't get something as simple as PoH right, how can anything be trusted?
- read_if_gay_ 6y ago
- ltultraweight 6y agoAnother Linux datapoint, X1 Carbon 6th Generation, daily use. SMART/Health Information (NVMe Log 0x02) Critical Warning: 0x00 Temperature: 35 Celsius Available Spare: 100% Available Spare Threshold: 10% Percentage Used: 25% Data Units Read: 22,242,132 [11.3 TB] Data Units Written: 74,693,212 [38.2 TB] Host Read Commands: 540,582,518 Host Write Commands: 1,857,635,922 Controller Busy Time: 5,131 Power Cycles: 884 Power On Hours: 3,497 Unsafe Shutdowns: 261 Media and Data Integrity Errors: 0 Error Information Log Entries: 882 Warning Comp. Temperature Time: 0 Critical Comp. Temperature Time: 0 Temperature Sensor 1: 35 Celsius Temperature Sensor 2: 37 Celsius