AskOverflow.Dev

AskOverflow.Dev Logo AskOverflow.Dev Logo

AskOverflow.Dev Navigation

  • 主页
  • 系统&网络
  • Ubuntu
  • Unix
  • DBA
  • Computer
  • Coding
  • LangChain

Mobile menu

Close
  • 主页
  • 系统&网络
    • 最新
    • 热门
    • 标签
  • Ubuntu
    • 最新
    • 热门
    • 标签
  • Unix
    • 最新
    • 标签
  • DBA
    • 最新
    • 标签
  • Computer
    • 最新
    • 标签
  • Coding
    • 最新
    • 标签
主页 / server / 问题

问题[hdd](server)

Martin Hope
fi11222
Asked: 2022-02-23 23:43:10 +0800 CST

SATA 错误出现在 Journalctl 中,而 SMART 诊断正常 - 主板问题?

  • 0

在注意到异常长的磁盘操作延迟后,我查找了 journalctl,这就是我发现的:

Feb 22 14:02:11.711182 Onan01 kernel: ata10: hard resetting link
Feb 22 14:02:12.186958 Onan01 kernel: ata10: SATA link up 1.5 Gbps (SStatus 113 SControl 310)
Feb 22 14:02:12.187044 Onan01 kernel: ata10.00: configured for UDMA/33
Feb 22 14:02:12.187068 Onan01 kernel: ata10: EH complete
Feb 22 14:02:22.782960 Onan01 kernel: ata10: SATA link up 1.5 Gbps (SStatus 113 SControl 310)
Feb 22 14:02:22.783033 Onan01 kernel: ata10.00: configured for UDMA/33
Feb 22 14:03:27.472083 Onan01 kernel: ata10.00: exception Emask 0x0 SAct 0x0 SErr 0xd0000 action 0x6 frozen
Feb 22 14:03:27.472241 Onan01 kernel: ata10: SError: { PHYRdyChg CommWake 10B8B }
Feb 22 14:03:27.472271 Onan01 kernel: ata10.00: failed command: WRITE DMA EXT
Feb 22 14:03:27.472300 Onan01 kernel: ata10.00: cmd 35/00:18:00:35:44/00:00:74:00:00/e0 tag 14 dma 12288 out
                                               res 40/00:01:00:4f:c2/00:00:00:00:00/00 Emask 0x4 (timeout)
Feb 22 14:03:27.472323 Onan01 kernel: ata10.00: status: { DRDY }
Feb 22 14:03:27.472345 Onan01 kernel: ata10: hard resetting link
Feb 22 14:03:27.950979 Onan01 kernel: ata10: SATA link up 1.5 Gbps (SStatus 113 SControl 310)
Feb 22 14:03:27.951084 Onan01 kernel: ata10.00: configured for UDMA/33
Feb 22 14:03:27.951113 Onan01 kernel: ata10: EH complete
Feb 22 14:04:03.852081 Onan01 kernel: ata10.00: exception Emask 0x10 SAct 0x0 SErr 0x40d0000 action 0xe frozen
Feb 22 14:04:03.852242 Onan01 kernel: ata10.00: irq_stat 0x00000040, connection status changed
Feb 22 14:04:03.852274 Onan01 kernel: ata10: SError: { PHYRdyChg CommWake 10B8B DevExch }
Feb 22 14:04:03.852301 Onan01 kernel: ata10.00: failed command: WRITE DMA EXT
Feb 22 14:04:03.852325 Onan01 kernel: ata10.00: cmd 35/00:38:58:35:44/00:00:74:00:00/e0 tag 17 dma 28672 out
                                               res 50/00:00:38:23:00/00:00:ac:00:00/e0 Emask 0x10 (ATA bus error)
Feb 22 14:04:03.852357 Onan01 kernel: ata10.00: status: { DRDY }

第一种错误(超时)似乎比第二种错误(ATA 总线错误)更频繁。每个都有不少。SATA 通道ata10连接到 WD Caviar Green HDD。

此磁盘上的 SMART 诊断显然是干净的:

sudo smartctl --all /dev/sdf1
smartctl 7.1 2019-12-30 r5022 [x86_64-linux-5.4.0-100-generic] (local build)
Copyright (C) 2002-19, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Device Model:     WDC WD20EZAZ-00GGJB0
Serial Number:    WD-WXT1A29LE265
LU WWN Device Id: 5 0014ee 211b07a4f
Firmware Version: 80.00A80
User Capacity:    2,000,398,934,016 bytes [2.00 TB]
Sector Sizes:     512 bytes logical, 4096 bytes physical
Rotation Rate:    5400 rpm
Form Factor:      3.5 inches
Device is:        Not in smartctl database [for details use: -P showall]
ATA Version is:   ACS-3 T13/2161-D revision 5
SATA Version is:  SATA 3.1, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is:    Wed Feb 23 11:37:14 2022 IST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x00) Offline data collection activity
                    was never started.
                    Auto Offline Data Collection: Disabled.
Self-test execution status:      (   0) The previous self-test routine completed
                    without error or no self-test has ever 
                    been run.
Total time to complete Offline 
data collection:        (32520) seconds.
Offline data collection
capabilities:            (0x7b) SMART execute Offline immediate.
                    Auto Offline data collection on/off support.
                    Suspend Offline collection upon new
                    command.
                    Offline surface scan supported.
                    Self-test supported.
                    Conveyance Self-test supported.
                    Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                    power-saving mode.
                    Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                    General Purpose Logging supported.
Short self-test routine 
recommended polling time:    (   2) minutes.
Extended self-test routine
recommended polling time:    ( 103) minutes.
Conveyance self-test routine
recommended polling time:    (   2) minutes.
SCT capabilities:          (0x3035) SCT Status supported.
                    SCT Feature Control supported.
                    SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  1 Raw_Read_Error_Rate     0x002f   200   200   051    Pre-fail  Always       -       0
  3 Spin_Up_Time            0x0027   184   170   021    Pre-fail  Always       -       1783
  4 Start_Stop_Count        0x0032   099   099   000    Old_age   Always       -       1573
  5 Reallocated_Sector_Ct   0x0033   200   200   140    Pre-fail  Always       -       0
  7 Seek_Error_Rate         0x002e   200   200   000    Old_age   Always       -       0
  9 Power_On_Hours          0x0032   083   083   000    Old_age   Always       -       13100
 10 Spin_Retry_Count        0x0032   100   100   000    Old_age   Always       -       0
 11 Calibration_Retry_Count 0x0032   100   100   000    Old_age   Always       -       0
 12 Power_Cycle_Count       0x0032   099   099   000    Old_age   Always       -       1524
192 Power-Off_Retract_Count 0x0032   199   199   000    Old_age   Always       -       761
193 Load_Cycle_Count        0x0032   147   147   000    Old_age   Always       -       160779
194 Temperature_Celsius     0x0022   115   104   000    Old_age   Always       -       28
196 Reallocated_Event_Count 0x0032   200   200   000    Old_age   Always       -       0
197 Current_Pending_Sector  0x0032   200   200   000    Old_age   Always       -       0
198 Offline_Uncorrectable   0x0030   100   253   000    Old_age   Offline      -       0
199 UDMA_CRC_Error_Count    0x0032   200   200   000    Old_age   Always       -       0
200 Multi_Zone_Error_Rate   0x0008   100   253   000    Old_age   Offline      -       0

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Short offline       Completed without error       00%     13100         -
# 2  Short offline       Completed without error       00%     13099         -

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

一件奇怪的事情是,长时间的 SMART 测试似乎无法正常工作。它们从进度 90% 直接完成(没有 80%、70% 等),之后,它们不会出现在“SMART 自测日志”部分中。

我连续两天经历了文件操作延迟。重新启动后,问题似乎消失了,然后又回来了。具体来说,问题表现为复制或移动文件的长时间延迟以及 LibreOffice 挂起文件保存。知道导致此类错误的原因是什么吗?

操作系统:Ubuntu 20.04

处理器:锐龙3

MB:技嘉 X570 UD

smart sata hdd
  • 1 个回答
  • 86 Views
Martin Hope
Alexander Kolodziej
Asked: 2021-09-30 10:24:39 +0800 CST

小型 esx 主机的磁盘:ssd raid5 还是 hdd raid10?[复制]

  • 0
这个问题在这里已经有了答案:
你能帮我做容量规划吗? (3 个回答)
11 个月前关闭。

我即将更换运行大约 20 个虚拟机(大部分处于空闲状态)的 ESXi 主机。今天,外部 iSCSI 盒 raid6(总共 7Tb)只使用了大约 800Gb。

新服务器将仅具有内部磁盘。2Tb 就足够了。它很可能是戴尔 R640(最多 8 个光盘)。ESXi 将位于一个简单的 raid1(2 个小磁盘)上。

但是虚拟机存储呢?我应该去哪个?(HDD 稍微便宜一些,但差别不大)。

A. 3 个 SSD (960Gb) Raid5 = 2Tb

B. 6 个 HDD (600Gb 10K) Raid10 = 几乎 2Tb

服务器当然会有一个专用的raid-controller (perc)。

raid ssd vmware-esxi hdd dell-perc
  • 1 个回答
  • 241 Views
Martin Hope
HCSF
Asked: 2021-03-02 23:50:17 +0800 CST

即使指定了`fio`也不会超时?

  • 1
# fio --name=random-write --directory=/mnt/test/ --ioengine=posixaio --rw=randwrite -bs=4k --numjobs=1 --size=4g -iodepth=1 -runtime=600 --time_based --end_fsync=1
random-write: (g=0): rw=randwrite, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=posixaio, iodepth=1
fio-3.7
Starting 1 process
random-write: Laying out IO file (1 file / 4096MiB)
Jobs: 1 (f=1): [w(1)][100.0%][r=0KiB/s,w=0KiB/s][r=0,w=0 IOPS][eta 00m:00s]    

知道为什么它在我设置的 60 分钟而不是 600 秒后返回吗?

我检查了dmesg,没有错误:

[Mon Mar  1 20:53:36 2021] XFS (sda2): Mounting V5 Filesystem
[Mon Mar  1 20:53:37 2021] XFS (sda2): Starting recovery (logdev: internal)
[Mon Mar  1 20:53:45 2021] XFS (sda2): Ending recovery (logdev: internal)

我同时在同一个盒子上的另一个驱动器(而不是 SSD)上运行了相同的命令,它按时完成并返回。

提前致谢!

filesystems xfs centos7 hdd fio
  • 1 个回答
  • 213 Views

Sidebar

Stats

  • 问题 205573
  • 回答 270741
  • 最佳答案 135370
  • 用户 68524
  • 热门
  • 回答
  • Marko Smith

    新安装后 postgres 的默认超级用户用户名/密码是什么?

    • 5 个回答
  • Marko Smith

    SFTP 使用什么端口?

    • 6 个回答
  • Marko Smith

    命令行列出 Windows Active Directory 组中的用户?

    • 9 个回答
  • Marko Smith

    什么是 Pem 文件,它与其他 OpenSSL 生成的密钥文件格式有何不同?

    • 3 个回答
  • Marko Smith

    如何确定bash变量是否为空?

    • 15 个回答
  • Martin Hope
    Tom Feiner 如何按大小对 du -h 输出进行排序 2009-02-26 05:42:42 +0800 CST
  • Martin Hope
    Noah Goodrich 什么是 Pem 文件,它与其他 OpenSSL 生成的密钥文件格式有何不同? 2009-05-19 18:24:42 +0800 CST
  • Martin Hope
    Brent 如何确定bash变量是否为空? 2009-05-13 09:54:48 +0800 CST
  • Martin Hope
    cletus 您如何找到在 Windows 中打开文件的进程? 2009-05-01 16:47:16 +0800 CST

热门标签

linux nginx windows networking ubuntu domain-name-system amazon-web-services active-directory apache-2.4 ssh

Explore

  • 主页
  • 问题
    • 最新
    • 热门
  • 标签
  • 帮助

Footer

AskOverflow.Dev

关于我们

  • 关于我们
  • 联系我们

Legal Stuff

  • Privacy Policy

Language

  • Pt
  • Server
  • Unix

© 2023 AskOverflow.DEV All Rights Reserve