************************ crashinfo ************************* /exports/testreports/49254/testresults/sanity-flr-zfs-centos7_x86_64-centos7_x86_64/oleg431-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.ZcCNV/vmlinux [TAINTED] DUMPFILE: /exports/testreports/49254/testresults/sanity-flr-zfs-centos7_x86_64-centos7_x86_64/oleg431-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Fri Feb 14 07:46:00 EST 2025 UPTIME: 01:06:47 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 340 NODENAME: oleg431-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2400 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 28 last 5s 69 last 60s 83 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 336 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 924 in:imjournal 0 ms /usr/sbin/rsyslogd -n 30395 z_null_int 0 ms (no user stack) 11 rcuos/0 0 ms (no user stack) 30394 z_null_iss 0 ms (no user stack) 30449 mmp 0 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 731032 2.8 GB 76% of TOTAL MEM USED 224035 875.1 MB 23% of TOTAL MEM SHARED 9547 37.3 MB 0% of TOTAL MEM BUFFERS 5165 20.2 MB 0% of TOTAL MEM CACHED 75018 293 MB 7% of TOTAL MEM SLAB 17184 67.1 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 64977 253.8 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 10 LISTEN 3 NAGLE disabled (TCP_NODELAY): 8 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 18 LISTEN 8 Raw sockets info -------------------- CLOSE 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 4004.0 eth0 n/a 4.8 RSS_TOTAL=68596 pages, %mem= 1.0 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff8801376688c0 ffff8800b5213800 sysfs sysfs /sys ffff880137668a80 ffff880139944000 proc proc /proc ffff880137668c40 ffff880137678000 devtmpfs devtmpfs /dev ffff880137668e00 ffff8800b5213000 securityfs securityfs /sys/kernel/security ffff880137668fc0 ffff8800b5214000 tmpfs tmpfs /dev/shm ffff880137669180 ffff880137325800 devpts devpts /dev/pts ffff880137669340 ffff8800b5214800 tmpfs tmpfs /run ffff880137669500 ffff8800b5215000 tmpfs tmpfs /sys/fs/cgroup ffff8801376696c0 ffff8800b5215800 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669880 ffff8800b5216000 pstore pstore /sys/fs/pstore ffff880137669a40 ffff88012a348000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff880137669c00 ffff88012a348800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff880137669dc0 ffff88012a349000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a360000 ffff88012a349800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a3601c0 ffff88012a34a000 cgroup cgroup /sys/fs/cgroup/pids ffff88012a360380 ffff88012a34a800 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a360540 ffff88012a34b000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a360700 ffff88012a34b800 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a3608c0 ffff88012a34c000 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a360a80 ffff88012a34c800 cgroup cgroup /sys/fs/cgroup/devices ffff88012b2b68c0 ffff8800b51fd000 configfs configfs /sys/kernel/config ffff88012b2b6c40 ffff8800b51fb000 ext4 /dev/nbd0 / ffff8800b4c1c000 ffff8800b53a5000 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012b2b6e00 ffff88012b523800 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff88012b2fa700 ffff880136428800 mqueue mqueue /dev/mqueue ffff880138ccae00 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff8800b4c1c380 ffff880129f8d800 hugetlbfs hugetlbfs /dev/hugepages ffff88012b2b6fc0 ffff88012b522000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff88012b2b7340 ffff88012b522800 ramfs none /mnt ffff880138ccafc0 ffff880129f89800 squashfs /dev/vda /home/green/git/lustre-release ffff8800b4c1c540 ffff8800b408d800 tmpfs none /var/lib/stateless/writable ffff8800b4c1c8c0 ffff8800b408d800 tmpfs none /var/cache/man ffff88012b2b7500 ffff8800b408d800 tmpfs none /var/log ffff88012b2b76c0 ffff8800b408d800 tmpfs none /var/lib/dbus ffff88012b2fa8c0 ffff8800b408d800 tmpfs none /tmp ffff880138ccb180 ffff8800b408d800 tmpfs none /var/lib/dhclient ffff88012b2faa80 ffff8800b408d800 tmpfs none /var/tmp ffff88012b2b7880 ffff8800b408d800 tmpfs none /var/lib/NetworkManager ffff88012b2fac40 ffff8800b408d800 tmpfs none /var/lib/systemd/random-seed ffff88012b2fae00 ffff8800b408d800 tmpfs none /var/spool ffff88012b2fafc0 ffff8800b408d800 tmpfs none /var/lib/nfs ffff880138ccb340 ffff8800b408d800 tmpfs none /var/lib/gssproxy ffff88012b2fb180 ffff8800b408d800 tmpfs none /var/lib/logrotate ffff880138ccb500 ffff8800b408d800 tmpfs none /etc ffff88012b2fb340 ffff8800b408d800 tmpfs none /var/lib/rsyslog ffff88012b2b7a40 ffff8800b408d800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff88012b2b7c00 ffff8800b4182000 nfs4 192.168.200.253:/exports/state/oleg431-server.virtnet /var/lib/stateless/state ffff88012b2b6700 ffff8800b4182000 nfs4 192.168.200.253:/exports/state/oleg431-server.virtnet /boot ffff88012b2b7dc0 ffff8800b4182000 nfs4 192.168.200.253:/exports/state/oleg431-server.virtnet /etc/etc/kdump.conf ffff8800b4c1ca80 ffff8800b53a5000 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b4c1d6c0 ffff880129f89800 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800b1969340 ffff880129f8e000 lustre lustre-mdt1/mdt1 /mnt/lustre-mds1 ffff8800b4c1c1c0 ffff8800933ad800 lustre lustre-ost2/ost2 /mnt/lustre-ost2 ffff8800b4c1d180 ffff880092c0b000 lustre lustre-ost1/ost1 /mnt/lustre-ost1 ffff8800b18b7500 ffff8800929d7000 tmpfs tmpfs /run/user/0 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 261.113683] hrtimer: interrupt took 5399270 ns [ 261.624834] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 262.086258] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 262.112291] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.204.131@tcp (at 0@lo) [ 264.133123] Lustre: DEBUG MARKER: oleg431-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 266.573030] Lustre: DEBUG MARKER: oleg431-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 1747.467087] Lustre: Failing over lustre-OST0001 [ 1747.489550] Lustre: server umount lustre-OST0001 complete [ 1747.801997] Lustre: lustre-OST0001-osc-MDT0000: Connection to lustre-OST0001 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1747.809361] LustreError: 8466:0:(ldlm_lib.c:1094:target_handle_connect()) lustre-OST0001: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 1747.816329] LustreError: 8466:0:(ldlm_lib.c:1094:target_handle_connect()) Skipped 4 previous similar messages [ 1749.105874] Lustre: DEBUG MARKER: oleg431-client.virtnet: executing wait_import_state (DISCONN|IDLE) osc.lustre-OST0001-osc-ffff88012a454000.ost_server_uuid 50 [ 1749.511159] Lustre: DEBUG MARKER: osc.lustre-OST0001-osc-ffff88012a454000.ost_server_uuid in IDLE state after 0 sec [ 1749.541343] LustreError: 8466:0:(ldlm_lib.c:1094:target_handle_connect()) lustre-OST0001: not available for connect from 192.168.204.31@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 1751.994782] Lustre: 25710:0:(mgc_request_server.c:553:mgc_llog_local_copy()) MGC192.168.204.131@tcp: no remote llog for lustre-sptlrpc, check MGS config [ 1752.046747] Lustre: lustre-OST0001: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 1752.052232] Lustre: lustre-OST0001: in recovery but waiting for the first client to connect [ 1753.204727] Lustre: lustre-OST0001: Will be in recovery for at least 1:00, or until 1 client reconnects [ 1753.207667] Lustre: lustre-OST0001: Denying connection for new client 91e608ab-79e9-4885-a316-1635841c82a7 (at 192.168.204.31@tcp), waiting for 1 known clients (0 recovered, 0 in progress, and 0 evicted) to recover in 0:59 [ 1753.250038] Lustre: lustre-OST0001: Recovery over after 0:01, of 1 clients 1 recovered and 0 were evicted. [ 1753.250129] Lustre: lustre-OST0001-osc-MDT0000: Connection restored to 192.168.204.131@tcp (at 0@lo) [ 1753.444612] Lustre: DEBUG MARKER: oleg431-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1754.629972] Lustre: DEBUG MARKER: oleg431-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3233.080659] Lustre: DEBUG MARKER: sanity-flr test_31: @@@@@@ FAIL: test_31 failed with 1 [ 3236.387596] Lustre: DEBUG MARKER: == sanity-flr test 32: data should be mirrored to newly created mirror ========================================================== 07:33:08 (1739536388) [ 3237.333561] Lustre: Failing over lustre-OST0000 [ 3237.360109] Lustre: server umount lustre-OST0000 complete [ 3238.791840] Lustre: DEBUG MARKER: oleg431-client.virtnet: executing wait_import_state (DISCONN|IDLE) osc.lustre-OST0000-osc-ffff88012a454000.ost_server_uuid 50 [ 3239.417842] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3239.421957] LustreError: 16971:0:(ldlm_lib.c:1094:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3241.949942] LustreError: 10332:0:(ldlm_lib.c:1094:target_handle_connect()) lustre-OST0000: not available for connect from 192.168.204.31@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3243.176865] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-ffff88012a454000.ost_server_uuid in DISCONN state after 4 sec [ 3244.339356] Lustre: 30629:0:(mgc_request_server.c:553:mgc_llog_local_copy()) MGC192.168.204.131@tcp: no remote llog for lustre-sptlrpc, check MGS config [ 3244.374787] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 3244.378894] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 3245.584764] Lustre: DEBUG MARKER: oleg431-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3246.240226] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 3246.336804] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 3246.336821] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.204.131@tcp (at 0@lo) [ 3246.635217] Lustre: DEBUG MARKER: oleg431-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.93s (real) 7.18s (CPU), Child processes: 5.71s