************************ crashinfo ************************* /exports/testreports/49254/testresults/replay-ost-single-zfs-centos7_x86_64-centos7_x86_64/oleg224-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ KERNEL: /tmp/crash-anaysis.XEvVU/vmlinux [TAINTED] DUMPFILE: /exports/testreports/49254/testresults/replay-ost-single-zfs-centos7_x86_64-centos7_x86_64/oleg224-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Fri Feb 14 07:26:13 EST 2025 UPTIME: 00:47:01 LOAD AVERAGE: 0.00, 0.01, 0.07 TASKS: 338 NODENAME: oleg224-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 35 last 5s 73 last 60s 86 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 334 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 20 rcuos/1 0 ms (no user stack) 9 rcu_sched 0 ms (no user stack) 27 rcuos/2 0 ms (no user stack) 8417 ll_ost_create00 1 ms (no user stack) 1 systemd 2 ms /usr/lib/systemd/systemd --switched-root --system --deserialize 22 +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 733391 2.8 GB 76% of TOTAL MEM USED 221676 865.9 MB 23% of TOTAL MEM SHARED 10034 39.2 MB 1% of TOTAL MEM BUFFERS 5165 20.2 MB 0% of TOTAL MEM CACHED 74917 292.6 MB 7% of TOTAL MEM SLAB 17768 69.4 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 64907 253.5 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 10 LISTEN 3 NAGLE disabled (TCP_NODELAY): 8 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 18 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 2818.7 eth0 n/a 2.4 RSS_TOTAL=66228 pages, %mem= 1.1 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137668540 ffff88012b376800 sysfs sysfs /sys ffff880137668700 ffff880139944000 proc proc /proc ffff8801376688c0 ffff880137678000 devtmpfs devtmpfs /dev ffff880137668a80 ffff88012b376000 securityfs securityfs /sys/kernel/security ffff880137668c40 ffff88012b377000 tmpfs tmpfs /dev/shm ffff880137668e00 ffff88013771f800 devpts devpts /dev/pts ffff880137668fc0 ffff88012b377800 tmpfs tmpfs /run ffff880137669180 ffff88012a2f8000 tmpfs tmpfs /sys/fs/cgroup ffff880137669340 ffff88012a2f8800 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669500 ffff88012a2f9000 pstore pstore /sys/fs/pstore ffff8801376696c0 ffff88012a2fb000 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff880137669880 ffff88012a2fa800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff880137669a40 ffff88012a2fa000 cgroup cgroup /sys/fs/cgroup/memory ffff880137669c00 ffff88012a2f9800 cgroup cgroup /sys/fs/cgroup/devices ffff880137669dc0 ffff88012a2fb800 cgroup cgroup /sys/fs/cgroup/blkio ffff88012aa0c000 ffff88012a2fc000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012aa0c1c0 ffff88012a2fc800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012aa0c380 ffff88012a2fd000 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012aa0c540 ffff88012a2fd800 cgroup cgroup /sys/fs/cgroup/pids ffff88012aa0c700 ffff88012a2fe000 cgroup cgroup /sys/fs/cgroup/cpuset ffff880138ccba40 ffff8800b4c30000 configfs configfs /sys/kernel/config ffff8800b4cd0000 ffff880137325000 ext4 /dev/nbd0 / ffff880138ccbc00 ffff8800b5237000 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88013700f340 ffff8800b4c34800 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff88013700f500 ffff88012b291000 mqueue mqueue /dev/mqueue ffff88012aa0cc40 ffff880129fb4800 hugetlbfs hugetlbfs /dev/hugepages ffff880138ccbdc0 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff88013700f6c0 ffff880137323800 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff88013700fa40 ffff880129f62800 ramfs none /mnt ffff88013700fc00 ffff880137324800 tmpfs none /var/lib/stateless/writable ffff88012aa0ce00 ffff8800b4c31000 squashfs /dev/vda /home/green/git/lustre-release ffff88013700e1c0 ffff880137324800 tmpfs none /var/cache/man ffff8800b2132000 ffff880137324800 tmpfs none /var/log ffff88012aa0cfc0 ffff880137324800 tmpfs none /var/lib/dbus ffff8800b4cd0700 ffff880137324800 tmpfs none /tmp ffff8800b4cd08c0 ffff880137324800 tmpfs none /var/lib/dhclient ffff88012aa0d180 ffff880137324800 tmpfs none /var/tmp ffff880138ccb6c0 ffff880137324800 tmpfs none /var/lib/NetworkManager ffff88012aa0d340 ffff880137324800 tmpfs none /var/lib/systemd/random-seed ffff88012aa0d500 ffff880137324800 tmpfs none /var/spool ffff8800b4cd0a80 ffff880137324800 tmpfs none /var/lib/nfs ffff8800b21321c0 ffff880137324800 tmpfs none /var/lib/gssproxy ffff8800b2132380 ffff880137324800 tmpfs none /var/lib/logrotate ffff8800b2132540 ffff880137324800 tmpfs none /etc ffff8800b4efc000 ffff880137324800 tmpfs none /var/lib/rsyslog ffff8800b2132700 ffff880137324800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff8800b4cd0c40 ffff880129f63000 nfs4 192.168.200.253:/exports/state/oleg224-server.virtnet /var/lib/stateless/state ffff8800b4cd1180 ffff880129f63000 nfs4 192.168.200.253:/exports/state/oleg224-server.virtnet /boot ffff8800b2132a80 ffff880129f63000 nfs4 192.168.200.253:/exports/state/oleg224-server.virtnet /etc/etc/kdump.conf ffff8800b2132c40 ffff8800b5237000 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff88012aa0d880 ffff8800b4c31000 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800a7a55500 ffff8800a8b4c800 lustre lustre-mdt1/mdt1 /mnt/lustre-mds1 ffff8800b1af8380 ffff880094424800 lustre lustre-ost2/ost2 /mnt/lustre-ost2 ffff8800a7a55880 ffff880094587800 lustre lustre-ost1/ost1 /mnt/lustre-ost1 ffff8800b1af9340 ffff880094535800 tmpfs tmpfs /run/user/0 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 489.754542] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 489.757512] LustreError: 2104:0:(ldlm_lib.c:3250:target_send_reply_msg()) @@@ dropping reply req@ffff88011b4e0e00 x1824032746522752/t4294967398(0) o36->ed515c55-7bac-403e-8683-0749441b9c5a@192.168.202.24@tcp:61/0 lens 504/456 e 0 to 0 dl 1739533651 ref 1 fl Interpret:/200/0 rc 0/0 job:'rm.0' uid:0 gid:0 [ 505.789053] Lustre: lustre-MDT0000: Client ed515c55-7bac-403e-8683-0749441b9c5a (at 192.168.202.24@tcp) reconnecting [ 505.812399] Lustre: 7102:0:(mdt_recovery.c:128:mdt_req_from_lrd()) @@@ restoring transno req@ffff88008d349180 x1824032746522752/t4294967398(0) o36->ed515c55-7bac-403e-8683-0749441b9c5a@192.168.202.24@tcp:77/0 lens 504/2888 e 0 to 0 dl 1739533667 ref 1 fl Interpret:/202/0 rc 0/0 job:'rm.0' uid:0 gid:0 [ 506.754579] Lustre: Failing over lustre-OST0000 [ 506.785734] Lustre: server umount lustre-OST0000 complete [ 506.978806] LustreError: lustre-OST0000-osc-MDT0000: operation ost_statfs to node 0@lo failed: rc = -107 [ 506.981800] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 506.989179] LustreError: 8764:0:(ldlm_lib.c:1094:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 506.998358] LustreError: 8764:0:(ldlm_lib.c:1094:target_handle_connect()) Skipped 8 previous similar messages [ 519.638093] Lustre: 4648:0:(mgc_request_server.c:553:mgc_llog_local_copy()) MGC192.168.202.124@tcp: no remote llog for lustre-sptlrpc, check MGS config [ 519.719430] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 520.821577] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 521.160197] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 521.160254] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.202.124@tcp (at 0@lo) [ 521.650750] Lustre: DEBUG MARKER: oleg224-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 524.966159] Lustre: DEBUG MARKER: oleg224-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 525.563958] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 527.745224] Lustre: DEBUG MARKER: oleg224-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 2008.738691] Lustre: DEBUG MARKER: replay-ost-single test_6: @@@@@@ FAIL: OST1 recovery not completed [ 2012.861023] Lustre: DEBUG MARKER: == replay-ost-single test 7: Fail OST before obd_destroy ========================================================== 07:12:43 (1739535163) [ 2024.127108] Lustre: DEBUG MARKER: before: 7536640 after_dd: 7532544 took 2 seconds [ 2025.171179] LustreError: 9361:0:(osd_handler.c:698:osd_ro()) lustre-OST0000: *** setting device osd-zfs read-only *** [ 2025.446830] Lustre: DEBUG MARKER: ost1 REPLAY BARRIER on lustre-OST0000 [ 2026.074034] Lustre: Failing over lustre-OST0000 [ 2026.096602] Lustre: server umount lustre-OST0000 complete [ 2028.238948] LustreError: 8415:0:(ldlm_lib.c:1094:target_handle_connect()) lustre-OST0000: not available for connect from 192.168.202.24@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 2028.245691] LustreError: 8415:0:(ldlm_lib.c:1094:target_handle_connect()) Skipped 4 previous similar messages [ 2028.627453] LustreError: lustre-OST0000-osc-MDT0000: operation ost_statfs to node 0@lo failed: rc = -107 [ 2028.632759] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2038.734290] Lustre: 10196:0:(mgc_request_server.c:553:mgc_llog_local_copy()) MGC192.168.202.124@tcp: no remote llog for lustre-sptlrpc, check MGS config [ 2038.813426] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 2038.821348] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 2039.937615] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 2040.226435] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 2040.226474] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.202.124@tcp (at 0@lo) [ 2040.748532] Lustre: DEBUG MARKER: oleg224-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 2043.881646] Lustre: DEBUG MARKER: oleg224-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 2044.487157] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 2046.743464] Lustre: DEBUG MARKER: oleg224-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.69s (real) 7.01s (CPU), Child processes: 5.62s