************************ crashinfo ************************* /exports/testreports/42551/testresults/replay-ost-single-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg108-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.hEDhT/vmlinux [TAINTED] DUMPFILE: /exports/testreports/42551/testresults/replay-ost-single-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg108-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Wed May 8 18:27:46 EDT 2024 UPTIME: 00:26:55 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 245 NODENAME: oleg108-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 27 last 5s 68 last 60s 76 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 241 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 3686 ptlrpcd_00_03 0 ms (no user stack) 11225 ll_ost_create00 0 ms (no user stack) 3685 ptlrpcd_00_02 0 ms (no user stack) 34 rcuos/3 0 ms (no user stack) 11217 osp-pre-0-0 0 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 743668 2.8 GB 77% of TOTAL MEM USED 211399 825.8 MB 22% of TOTAL MEM SHARED 10526 41.1 MB 1% of TOTAL MEM BUFFERS 7846 30.6 MB 0% of TOTAL MEM CACHED 67991 265.6 MB 7% of TOTAL MEM SLAB 15436 60.3 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 60186 235.1 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- CLOSE 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 1612.5 eth0 n/a 1.6 RSS_TOTAL=54072 pages, %mem= 0.9 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137668380 ffff880137679000 sysfs sysfs /sys ffff880137668540 ffff880139944000 proc proc /proc ffff880137668700 ffff880137678000 devtmpfs devtmpfs /dev ffff8801376688c0 ffff88012b396000 securityfs securityfs /sys/kernel/security ffff880137668a80 ffff880137679800 tmpfs tmpfs /dev/shm ffff880137668c40 ffff880137171000 devpts devpts /dev/pts ffff880137668e00 ffff88013767a000 tmpfs tmpfs /run ffff880137668fc0 ffff88013767a800 tmpfs tmpfs /sys/fs/cgroup ffff880137669180 ffff88013767b000 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669340 ffff88013767b800 pstore pstore /sys/fs/pstore ffff880137669500 ffff88013767d800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff8801376696c0 ffff88013767d000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff880137669880 ffff88013767c800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff880137669a40 ffff88013767c000 cgroup cgroup /sys/fs/cgroup/perf_event ffff880137669c00 ffff88013767e000 cgroup cgroup /sys/fs/cgroup/cpuset ffff880137669dc0 ffff88013767e800 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a362000 ffff88013767f000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a3621c0 ffff88013767f800 cgroup cgroup /sys/fs/cgroup/pids ffff88012a362380 ffff88012a380000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a362540 ffff88012a380800 cgroup cgroup /sys/fs/cgroup/devices ffff8801377f3340 ffff880129c2c000 configfs configfs /sys/kernel/config ffff88012a3628c0 ffff88012a387000 ext4 /dev/nbd0 / ffff8801377f36c0 ffff880129c2c800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012a362fc0 ffff8800b5219800 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880138ccba40 ffff88012b397800 hugetlbfs hugetlbfs /dev/hugepages ffff8801377f3880 ffff880137172800 mqueue mqueue /dev/mqueue ffff8801377f3a40 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff880129eac1c0 ffff8800b4176000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff8801377f3c00 ffff880129c5b800 ramfs none /mnt ffff880129eac540 ffff8800b2462800 squashfs /dev/vda /home/green/git/lustre-release ffff88012a363180 ffff880129c2b000 tmpfs none /var/lib/stateless/writable ffff88012a363340 ffff880129c2b000 tmpfs none /var/cache/man ffff880138ccbdc0 ffff880129c2b000 tmpfs none /var/log ffff88012a363500 ffff880129c2b000 tmpfs none /var/lib/dbus ffff88012a3636c0 ffff880129c2b000 tmpfs none /tmp ffff88012a363880 ffff880129c2b000 tmpfs none /var/lib/dhclient ffff88012a363a40 ffff880129c2b000 tmpfs none /var/tmp ffff88012a363c00 ffff880129c2b000 tmpfs none /var/lib/NetworkManager ffff88012a363dc0 ffff880129c2b000 tmpfs none /var/lib/systemd/random-seed ffff880129fc0000 ffff880129c2b000 tmpfs none /var/spool ffff88012a362e00 ffff880129c2b000 tmpfs none /var/lib/nfs ffff88012a362c40 ffff880129c2b000 tmpfs none /var/lib/gssproxy ffff880129eac700 ffff880129c2b000 tmpfs none /var/lib/logrotate ffff880129fc01c0 ffff880129c2b000 tmpfs none /etc ffff88012a362a80 ffff880129c2b000 tmpfs none /var/lib/rsyslog ffff8800b24f8000 ffff880129c2b000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff8801377f3dc0 ffff8800b4176800 nfs4 192.168.200.253:/exports/state/oleg108-server.virtnet /var/lib/stateless/state ffff880129eac8c0 ffff8800b4176800 nfs4 192.168.200.253:/exports/state/oleg108-server.virtnet /boot ffff8800b24f81c0 ffff8800b4176800 nfs4 192.168.200.253:/exports/state/oleg108-server.virtnet /etc/etc/kdump.conf ffff880129eaca80 ffff880129c2c800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b1dc3a40 ffff8800b2462800 squashfs /dev/vda /usr/sbin/mount.lustre ffff880129fc0c40 ffff880129c5c800 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 ffff880129fc0fc0 ffff88011e08f800 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff880129fc16c0 ffff88012d069000 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 ffff8800b435fa40 ffff88009aa3e000 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 136.752743] Lustre: lustre-MDT0001: haven't heard from client afc3f8d6-cd1c-44a4-b03f-34b8dc2fb94c (at 192.168.201.8@tcp) in 32 seconds. I think it's dead, and I am evicting it. exp ffff88012cbe0800, cur 1715205787 expire 1715205757 last 1715205755 [ 140.505100] Lustre: DEBUG MARKER: == replay-ost-single test 1: touch ======================= 18:03:11 (1715205791) [ 140.974806] Lustre: Failing over lustre-OST0000 [ 140.987058] Lustre: server umount lustre-OST0000 complete [ 141.764828] LustreError: 137-5: lustre-OST0000: not available for connect from 192.168.201.8@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 142.271567] LustreError: 11-0: lustre-OST0000-osc-MDT0000: operation ost_destroy to node 0@lo failed: rc = -107 [ 142.273961] LustreError: Skipped 1 previous similar message [ 142.275374] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 151.776335] LustreError: 137-5: lustre-OST0000: not available for connect from 192.168.201.8@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 151.779852] LustreError: Skipped 5 previous similar messages [ 152.859326] LDISKFS-fs (dm-2): file extents enabled, maximum tree depth=5 [ 152.863010] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: user_xattr,acl,no_mbcache,nodelalloc [ 152.902624] Lustre: lustre-OST0000: Imperative Recovery not enabled, recovery window 60-180 [ 154.115390] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 154.136769] Lustre: DEBUG MARKER: oleg108-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 154.232291] Lustre: lustre-OST0000-osc-MDT0001: Connection restored to 192.168.201.108@tcp (at 0@lo) [ 154.232297] Lustre: lustre-OST0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 154.238236] Lustre: Skipped 1 previous similar message [ 156.346268] Lustre: DEBUG MARKER: oleg108-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 156.703178] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 160.841986] Lustre: DEBUG MARKER: == replay-ost-single test 2: |x| 10 open(O_CREAT)s ======= 18:03:31 (1715205811) [ 161.384011] Lustre: Failing over lustre-OST0000 [ 161.398637] Lustre: server umount lustre-OST0000 complete [ 162.911135] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 162.915229] Lustre: Skipped 2 previous similar messages [ 167.919062] LustreError: 137-5: lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 167.924189] LustreError: Skipped 6 previous similar messages [ 173.400580] LDISKFS-fs (dm-2): file extents enabled, maximum tree depth=5 [ 173.404482] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: user_xattr,acl,no_mbcache,nodelalloc [ 173.444769] Lustre: lustre-OST0000: Imperative Recovery not enabled, recovery window 60-180 [ 174.504642] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 174.706896] Lustre: DEBUG MARKER: oleg108-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 243.559646] Lustre: lustre-OST0000: recovery is timed out, evict stale exports [ 243.562639] Lustre: 20700:0:(genops.c:1528:class_disconnect_stale_exports()) lustre-OST0000: disconnect stale client 898362d5-f8d8-46df-a2a1-c4c53eb8380a@192.168.201.8@tcp [ 243.566638] Lustre: lustre-OST0000: disconnecting 1 stale clients [ 243.570660] Lustre: 20700:0:(ldlm_lib.c:2874:target_recovery_thread()) too long recovery - read logs [ 243.570690] Lustre: lustre-OST0000-osc-MDT0001: Connection restored to 192.168.201.108@tcp (at 0@lo) [ 243.570691] Lustre: Skipped 1 previous similar message [ 243.577798] LustreError: dumping log to /tmp/lustre-log.1715205893.20700 [ 243.591104] Lustre: lustre-OST0000: Recovery over after 1:10, of 3 clients 2 recovered and 1 was evicted. ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.62s (real) 6.62s (CPU), Child processes: 5.96s