************************ crashinfo ************************* /exports/testreports/51008/testresults/recovery-small-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg126-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.DtykS/vmlinux [TAINTED] DUMPFILE: /exports/testreports/51008/testresults/recovery-small-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg126-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Wed Apr 23 22:58:08 EDT 2025 UPTIME: 01:57:00 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 270 NODENAME: oleg126-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 25 last 5s 60 last 60s 76 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 266 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 13 watchdog/0 0 ms (no user stack) 28 watchdog/3 0 ms (no user stack) 14 watchdog/1 0 ms (no user stack) 21 watchdog/2 0 ms (no user stack) 4516 l2arc_feed 48 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 725603 2.8 GB 75% of TOTAL MEM USED 229464 896.3 MB 24% of TOTAL MEM SHARED 18240 71.2 MB 1% of TOTAL MEM BUFFERS 15151 59.2 MB 1% of TOTAL MEM CACHED 72543 283.4 MB 7% of TOTAL MEM SLAB 17674 69 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 60155 235 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 7017.3 eth0 n/a 3.0 RSS_TOTAL=58060 pages, %mem= 0.8 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137668380 ffff880137679000 sysfs sysfs /sys ffff880137668540 ffff880139944000 proc proc /proc ffff880137668700 ffff880137678000 devtmpfs devtmpfs /dev ffff8801376688c0 ffff8800b5238000 securityfs securityfs /sys/kernel/security ffff880137668a80 ffff880137679800 tmpfs tmpfs /dev/shm ffff880137668c40 ffff88013710a800 devpts devpts /dev/pts ffff880137668e00 ffff88013767a000 tmpfs tmpfs /run ffff880137668fc0 ffff88013767a800 tmpfs tmpfs /sys/fs/cgroup ffff880137669180 ffff88013767b000 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669340 ffff88013767b800 pstore pstore /sys/fs/pstore ffff880138ccba40 ffff8800b51de000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff880138ccbc00 ffff8800b51dd800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff880138ccbdc0 ffff8800b51dd000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a2a2000 ffff8800b51dc800 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a2a21c0 ffff8800b51de800 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a2a2380 ffff8800b51df000 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a2a2540 ffff8800b51df800 cgroup cgroup /sys/fs/cgroup/pids ffff88012a2a2700 ffff88012a378000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a2a28c0 ffff88012a378800 cgroup cgroup /sys/fs/cgroup/devices ffff88012a2a2a80 ffff88012a379000 cgroup cgroup /sys/fs/cgroup/blkio ffff8800b4cbc1c0 ffff8800b5241800 configfs configfs /sys/kernel/config ffff8800b4cbc380 ffff8800b5245000 ext4 /dev/nbd0 / ffff880137669500 ffff8800b5240800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012a308540 ffff880129f8d000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff88012a2a3dc0 ffff88013710c000 mqueue mqueue /dev/mqueue ffff88012a308700 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff8800b4cbc000 ffff88012a37e000 hugetlbfs hugetlbfs /dev/hugepages ffff88012a3088c0 ffff880129f8f800 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff880137669dc0 ffff8800b417d800 ramfs none /mnt ffff880138ccb880 ffff8800b523e000 tmpfs none /var/lib/stateless/writable ffff88012a308a80 ffff880129de4800 squashfs /dev/vda /home/green/git/lustre-release ffff8800b4cbc540 ffff8800b523e000 tmpfs none /var/cache/man ffff88012a308c40 ffff8800b523e000 tmpfs none /var/log ffff8800b4cbc700 ffff8800b523e000 tmpfs none /var/lib/dbus ffff88012a308e00 ffff8800b523e000 tmpfs none /tmp ffff8800b4cbc8c0 ffff8800b523e000 tmpfs none /var/lib/dhclient ffff880137669a40 ffff8800b523e000 tmpfs none /var/tmp ffff880137669880 ffff8800b523e000 tmpfs none /var/lib/NetworkManager ffff88012a308fc0 ffff8800b523e000 tmpfs none /var/lib/systemd/random-seed ffff88012a309180 ffff8800b523e000 tmpfs none /var/spool ffff8800b4cbca80 ffff8800b523e000 tmpfs none /var/lib/nfs ffff88012a309340 ffff8800b523e000 tmpfs none /var/lib/gssproxy ffff8800b4cbcc40 ffff8800b523e000 tmpfs none /var/lib/logrotate ffff88012a309500 ffff8800b523e000 tmpfs none /etc ffff8800b4cbce00 ffff8800b523e000 tmpfs none /var/lib/rsyslog ffff8800b4cbcfc0 ffff8800b523e000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff8801376696c0 ffff8800b20fc800 nfs4 192.168.200.253:/exports/state/oleg126-server.virtnet /var/lib/stateless/state ffff8800b4cbd340 ffff8800b20fc800 nfs4 192.168.200.253:/exports/state/oleg126-server.virtnet /boot ffff8800b4cbd180 ffff8800b20fc800 nfs4 192.168.200.253:/exports/state/oleg126-server.virtnet /etc/etc/kdump.conf ffff8800b4cbd500 ffff8800b5240800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b23556c0 ffff880129de4800 squashfs /dev/vda /usr/sbin/mount.lustre ffff88012b495dc0 ffff8800a7b06800 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff88012b495880 ffff8800a00f6800 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 ffff8800b4121dc0 ffff880087468000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 ffff8800b41201c0 ffff8800ac24d800 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 3366.730271] Lustre: 11245:0:(genops.c:1791:obd_export_evict_by_uuid()) lustre-OST0000: evicting 3acca340-df38-459f-a205-b4283e5c387f at adminstrative request [ 3366.735780] LustreError: 7201:0:(ldlm_lockd.c:2951:ldlm_bl_thread_exports()) cfs_fail_timeout id 31e sleeping for 4000ms [ 3368.516117] LustreError: 3610:0:(ldlm_lockd.c:1425:ldlm_handle_enqueue()) cfs_fail_timeout id 31e awake [ 3368.519137] LustreError: 3610:0:(ldlm_lockd.c:1447:ldlm_handle_enqueue()) ### lock on destroyed export ffff88012e8d3000 ns: filter-lustre-OST0000_UUID lock: ffff88012e9fae00/0x8ade9935cbaaa9dd lrc: 3/0,0 mode: --/PW res: [0x280000401:0x2e63:0x0].0x0 rrc: 4 type: EXT [0->4095] (req 0->4095) gid 0 flags: 0x70000000020020 nid: 192.168.201.26@tcp remote: 0x1a72cd63db758f9 expref: 3 pid: 3610 timeout: 0 lvb_type: 0 [ 3369.122125] LustreError: 3611:0:(ldlm_lockd.c:1425:ldlm_handle_enqueue()) cfs_fail_timeout interrupted [ 3375.682281] Lustre: lustre-MDT0001: haven't heard from client 735deee5-5992-4d53-97cf-9802a154aad3 (at 192.168.201.26@tcp) in 102 seconds. I think it's dead, and I am evicting it. exp ffff8800918e6800, cur 1745459843 deadline 1745459841 last 1745459741 [ 3378.184128] Lustre: lustre-OST0000: Client 55b93203-161a-4eba-829d-26d77ac3acbf (at 192.168.201.26@tcp) reconnecting [ 3378.189760] Lustre: Skipped 6 previous similar messages [ 3380.282216] Lustre: DEBUG MARKER: == recovery-small test 66: lock enqueue re-send vs client eviction ========================================================== 21:57:27 (1745459847) [ 3380.620131] Lustre: *** cfs_fail_loc=157, val=2147483648*** [ 3380.621756] LustreError: 18744:0:(ldlm_lib.c:3260:target_send_reply_msg()) @@@ dropping reply req@ffff88008ed0d880 x1830243962449280/t0(0) o101->3acca340-df38-459f-a205-b4283e5c387f@192.168.201.26@tcp:318/0 lens 576/688 e 0 to 0 dl 1745459903 ref 1 fl Interpret:/200/0 rc 0/0 job:'stat.0' uid:0 gid:0 [ 3382.673394] LustreError: 25393:0:(mdt_handler.c:2326:mdt_getattr_name_lock()) cfs_fail_timeout id 136 sleeping for 40000ms [ 3384.908699] Lustre: 11775:0:(genops.c:1791:obd_export_evict_by_uuid()) lustre-MDT0000: evicting 3acca340-df38-459f-a205-b4283e5c387f at adminstrative request [ 3385.279143] LustreError: 25393:0:(mdt_handler.c:2326:mdt_getattr_name_lock()) cfs_fail_timeout interrupted [ 3385.281949] LustreError: 25393:0:(mdt_handler.c:2326:mdt_getattr_name_lock()) Skipped 1 previous similar message [ 3386.634271] Lustre: DEBUG MARKER: == recovery-small test 67: connect vs import invalidate race ========================================================== 21:57:34 (1745459854) [ 3388.905893] Lustre: 12165:0:(genops.c:1791:obd_export_evict_by_uuid()) lustre-MDT0000: evicting 3acca340-df38-459f-a205-b4283e5c387f at adminstrative request [ 3402.068327] Lustre: DEBUG MARKER: == recovery-small test 100: IR: Make sure normal recovery still works w/o IR ========================================================== 21:57:49 (1745459869) [ 3403.415311] Lustre: Failing over lustre-OST0000 [ 3403.447627] Lustre: server umount lustre-OST0000 complete [ 3404.417893] LustreError: lustre-OST0000-osc-MDT0001: operation ost_statfs to node 0@lo failed: rc = -107 [ 3404.420328] LustreError: Skipped 2 previous similar messages [ 3415.463947] LDISKFS-fs (dm-2): file extents enabled, maximum tree depth=5 [ 3415.468029] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: user_xattr,acl,no_mbcache,nodelalloc [ 3415.491268] Lustre: 13296:0:(mgc_request_server.c:550:mgc_llog_local_copy()) MGC192.168.201.126@tcp: no remote llog for lustre-sptlrpc, check MGS config [ 3416.796956] Lustre: DEBUG MARKER: oleg126-server.virtnet: executing set_default_debug -1 all [ 3421.281798] Lustre: DEBUG MARKER: oleg126-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 3421.634150] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 3424.691620] Lustre: DEBUG MARKER: == recovery-small test 101a: IR: Make sure IR works w/o normal recovery ========================================================== 21:58:12 (1745459892) [ 3425.609748] Lustre: Failing over lustre-OST0000 [ 3425.636820] Lustre: server umount lustre-OST0000 complete [ 3437.557456] LDISKFS-fs (dm-2): file extents enabled, maximum tree depth=5 [ 3437.560729] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: user_xattr,acl,no_mbcache,nodelalloc [ 3437.579780] Lustre: 15194:0:(mgc_request_server.c:550:mgc_llog_local_copy()) MGC192.168.201.126@tcp: no remote llog for lustre-sptlrpc, check MGS config [ 3437.625602] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 3439.066047] Lustre: DEBUG MARKER: oleg126-server.virtnet: executing set_default_debug -1 all [ 3548.594122] Lustre: lustre-OST0000: recovery is timed out, evict stale exports [ 3548.596744] Lustre: 15207:0:(genops.c:1618:class_disconnect_stale_exports()) lustre-OST0000: disconnect stale client 3acca340-df38-459f-a205-b4283e5c387f@ [ 3548.601214] Lustre: 15207:0:(genops.c:1618:class_disconnect_stale_exports()) Skipped 1 previous similar message [ 3548.605941] Lustre: lustre-OST0000: disconnecting 1 stale clients ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 11.88s (real) 6.39s (CPU), Child processes: 5.45s