************************ crashinfo ************************* /exports/testreports/42551/testresults/replay-ost-single-zfs-centos7_x86_64-centos7_x86_64/oleg423-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.bensW/vmlinux [TAINTED] DUMPFILE: /exports/testreports/42551/testresults/replay-ost-single-zfs-centos7_x86_64-centos7_x86_64/oleg423-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Wed May 8 18:27:46 EDT 2024 UPTIME: 00:26:55 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 393 NODENAME: oleg423-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 29 last 5s 60 last 60s 77 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 389 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 20609 z_null_iss 0 ms (no user stack) 20695 mmp 0 ms (no user stack) 484 kworker/2:2 0 ms (no user stack) 3213 monitor_thread 27 ms (no user stack) 3208 socknal_reaper 31 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 719554 2.7 GB 75% of TOTAL MEM USED 235513 920 MB 24% of TOTAL MEM SHARED 8418 32.9 MB 0% of TOTAL MEM BUFFERS 5096 19.9 MB 0% of TOTAL MEM CACHED 66923 261.4 MB 7% of TOTAL MEM SLAB 22698 88.7 MB 2% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 59853 233.8 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- CLOSE 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 1612.3 eth0 n/a 3.5 RSS_TOTAL=52572 pages, %mem= 0.8 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137668380 ffff880137679000 sysfs sysfs /sys ffff880137668540 ffff880139944000 proc proc /proc ffff880137668700 ffff880137678000 devtmpfs devtmpfs /dev ffff8801376688c0 ffff8800b523e800 securityfs securityfs /sys/kernel/security ffff880137668a80 ffff880137679800 tmpfs tmpfs /dev/shm ffff880137668c40 ffff88012b278800 devpts devpts /dev/pts ffff880137668e00 ffff88013767a000 tmpfs tmpfs /run ffff880137668fc0 ffff88013767a800 tmpfs tmpfs /sys/fs/cgroup ffff880137669180 ffff88013767b000 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669340 ffff88013767b800 pstore pstore /sys/fs/pstore ffff880137669500 ffff88013767d800 cgroup cgroup /sys/fs/cgroup/devices ffff8801376696c0 ffff88013767d000 cgroup cgroup /sys/fs/cgroup/pids ffff880137669880 ffff88013767c800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff880137669a40 ffff88013767c000 cgroup cgroup /sys/fs/cgroup/freezer ffff880137669c00 ffff88013767e000 cgroup cgroup /sys/fs/cgroup/perf_event ffff880137669dc0 ffff88013767e800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012a2f6000 ffff88013767f000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a2f61c0 ffff88013767f800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a2f6380 ffff88012a358000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a2f6540 ffff88012a358800 cgroup cgroup /sys/fs/cgroup/cpuset ffff880137012700 ffff8800b528d800 configfs configfs /sys/kernel/config ffff880137012c40 ffff8800b528f800 ext4 /dev/nbd0 / ffff88012b271c00 ffff8800b528c000 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff880137012e00 ffff88012a35d800 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880138ccac40 ffff88012b27a000 mqueue mqueue /dev/mqueue ffff88012a2f6700 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff88012a2f68c0 ffff8800b402e000 hugetlbfs hugetlbfs /dev/hugepages ffff8800b4a6a380 ffff88012a32c000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff880137012fc0 ffff880136978000 ramfs none /mnt ffff880138ccae00 ffff8800b4a10000 squashfs /dev/vda /home/green/git/lustre-release ffff880137013180 ffff880129cd5800 tmpfs none /var/lib/stateless/writable ffff88012a2f6c40 ffff880129cd5800 tmpfs none /var/cache/man ffff88012a2f6e00 ffff880129cd5800 tmpfs none /var/log ffff880138ccafc0 ffff880129cd5800 tmpfs none /var/lib/dbus ffff880137013340 ffff880129cd5800 tmpfs none /tmp ffff880137013500 ffff880129cd5800 tmpfs none /var/lib/dhclient ffff8800b4a6a8c0 ffff880129cd5800 tmpfs none /var/tmp ffff88012a2f6fc0 ffff880129cd5800 tmpfs none /var/lib/NetworkManager ffff8800b4a6aa80 ffff880129cd5800 tmpfs none /var/lib/systemd/random-seed ffff880138ccb180 ffff880129cd5800 tmpfs none /var/spool ffff8801370136c0 ffff880129cd5800 tmpfs none /var/lib/nfs ffff880138ccb340 ffff880129cd5800 tmpfs none /var/lib/gssproxy ffff880138ccb500 ffff880129cd5800 tmpfs none /var/lib/logrotate ffff880137013880 ffff880129cd5800 tmpfs none /etc ffff880138ccb6c0 ffff880129cd5800 tmpfs none /var/lib/rsyslog ffff8800b4a6ac40 ffff880129cd5800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff8800b4a6ae00 ffff8800b1c14800 nfs4 192.168.200.253:/exports/state/oleg423-server.virtnet /var/lib/stateless/state ffff880138ccb880 ffff8800b1c14800 nfs4 192.168.200.253:/exports/state/oleg423-server.virtnet /boot ffff880138ccba40 ffff8800b1c14800 nfs4 192.168.200.253:/exports/state/oleg423-server.virtnet /etc/etc/kdump.conf ffff880138ccbc00 ffff8800b528c000 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff88012a2f7dc0 ffff8800b4a10000 squashfs /dev/vda /usr/sbin/mount.lustre ffff88012b271340 ffff880130f72800 lustre lustre-mdt1/mdt1 /mnt/lustre-mds1 ffff8800b2709880 ffff880091a6d800 lustre lustre-ost2/ost2 /mnt/lustre-ost2 ffff8800b2708c40 ffff88012c097000 lustre lustre-ost1/ost1 /mnt/lustre-ost1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 124.142263] Lustre: 3216:0:(client.c:2343:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1715205759/real 1715205759] req@ffff88008e12dc00 x1798523516371264/t0(0) o13->lustre-OST0000-osc-MDT0000@0@lo:7/4 lens 224/368 e 0 to 1 dl 1715205775 ref 1 fl Rpc:XQr/200/ffffffff rc 0/-1 job:'osp-pre-0-0.0' uid:0 gid:0 [ 125.136029] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 125.334688] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 125.334696] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.204.123@tcp (at 0@lo) [ 125.412015] Lustre: DEBUG MARKER: oleg423-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 127.211416] Lustre: lustre-MDT0000: haven't heard from client ee4bd752-b8f0-43ed-9118-7b6e5a164da5 (at 192.168.204.23@tcp) in 31 seconds. I think it's dead, and I am evicting it. exp ffff880090005800, cur 1715205778 expire 1715205748 last 1715205747 [ 127.814387] Lustre: DEBUG MARKER: oleg423-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 128.228564] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 132.780222] Lustre: DEBUG MARKER: == replay-ost-single test 1: touch ======================= 18:03:03 (1715205783) [ 133.303028] Lustre: Failing over lustre-OST0000 [ 133.314651] Lustre: server umount lustre-OST0000 complete [ 133.491640] LustreError: 137-5: lustre-OST0000: not available for connect from 192.168.204.23@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 133.499936] LustreError: Skipped 1 previous similar message [ 134.029367] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 138.494462] LustreError: 137-5: lustre-OST0000: not available for connect from 192.168.204.23@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 138.500984] LustreError: Skipped 1 previous similar message [ 145.090194] Lustre: lustre-OST0000: Imperative Recovery not enabled, recovery window 60-180 [ 146.163630] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 146.411666] Lustre: DEBUG MARKER: oleg423-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 146.783528] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.204.123@tcp (at 0@lo) [ 146.783535] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 148.733977] Lustre: DEBUG MARKER: oleg423-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 149.122707] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 153.657266] Lustre: DEBUG MARKER: == replay-ost-single test 2: |x| 10 open(O_CREAT)s ======= 18:03:24 (1715205804) [ 154.235311] Lustre: Failing over lustre-OST0000 [ 154.257951] Lustre: server umount lustre-OST0000 complete [ 155.085936] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 155.094157] LustreError: 137-5: lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 155.099672] LustreError: Skipped 3 previous similar messages [ 165.991868] Lustre: lustre-OST0000: Imperative Recovery not enabled, recovery window 60-180 [ 165.998141] mount.lustre (20873) used greatest stack depth: 10168 bytes left [ 167.230192] Lustre: DEBUG MARKER: oleg423-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 167.523820] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 236.637596] Lustre: lustre-OST0000: recovery is timed out, evict stale exports [ 236.639811] Lustre: 20908:0:(genops.c:1528:class_disconnect_stale_exports()) lustre-OST0000: disconnect stale client 0cd8d0fc-d34b-41b6-917f-f4d700291109@192.168.204.23@tcp [ 236.644893] Lustre: lustre-OST0000: disconnecting 1 stale clients [ 236.667259] Lustre: 20908:0:(ldlm_lib.c:2874:target_recovery_thread()) too long recovery - read logs [ 236.667342] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.204.123@tcp (at 0@lo) [ 236.674702] LustreError: dumping log to /tmp/lustre-log.1715205887.20908 [ 236.685149] Lustre: lustre-OST0000: Recovery over after 1:10, of 2 clients 1 recovered and 1 was evicted. ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.19s (real) 6.67s (CPU), Child processes: 5.50s