************************ crashinfo ************************* /exports/testreports/47154/testresults/recovery-small-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg124-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.1vlsN/vmlinux [TAINTED] DUMPFILE: /exports/testreports/47154/testresults/recovery-small-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg124-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Wed Nov 13 13:15:43 EST 2024 UPTIME: 01:08:31 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 247 NODENAME: oleg124-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 14 last 5s 59 last 60s 75 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 243 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 9 rcu_sched 0 ms (no user stack) 824 gmain 0 ms /usr/sbin/NetworkManager --no-daemon 4513 dbuf_evict 0 ms (no user stack) 34 rcuos/3 0 ms (no user stack) 4511 arc_reap 262 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 739203 2.8 GB 77% of TOTAL MEM USED 215864 843.2 MB 22% of TOTAL MEM SHARED 10941 42.7 MB 1% of TOTAL MEM BUFFERS 8500 33.2 MB 0% of TOTAL MEM CACHED 71187 278.1 MB 7% of TOTAL MEM SLAB 15319 59.8 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 58979 230.4 MB 7% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 4108.4 eth0 n/a 4.7 RSS_TOTAL=51140 pages, %mem= 0.8 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff88012eb4c000 ffff880137300800 sysfs sysfs /sys ffff88012eb4c1c0 ffff880139944000 proc proc /proc ffff88012eb4c380 ffff880137678000 devtmpfs devtmpfs /dev ffff88012eb4c540 ffff8800b5201800 securityfs securityfs /sys/kernel/security ffff88012eb4c700 ffff880137301000 tmpfs tmpfs /dev/shm ffff88012eb4c8c0 ffff88013771f000 devpts devpts /dev/pts ffff88012eb4ca80 ffff880137301800 tmpfs tmpfs /run ffff88012eb4cc40 ffff880137302000 tmpfs tmpfs /sys/fs/cgroup ffff88012eb4ce00 ffff880137302800 cgroup cgroup /sys/fs/cgroup/systemd ffff88012eb4cfc0 ffff880137303000 pstore pstore /sys/fs/pstore ffff8801377f6540 ffff88012b3ba800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff8801377f6700 ffff88012b3bb000 cgroup cgroup /sys/fs/cgroup/memory ffff8801377f68c0 ffff88012b3bb800 cgroup cgroup /sys/fs/cgroup/pids ffff8801377f6a80 ffff88012b3bc000 cgroup cgroup /sys/fs/cgroup/devices ffff8801377f6c40 ffff88012b3bc800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff8801377f6e00 ffff88012b3bd000 cgroup cgroup /sys/fs/cgroup/cpuset ffff8801377f6fc0 ffff88012b3bd800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff8801377f7180 ffff88012b3be000 cgroup cgroup /sys/fs/cgroup/perf_event ffff8801377f7340 ffff88012b3be800 cgroup cgroup /sys/fs/cgroup/freezer ffff8801377f7500 ffff88012b3bf000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012eb4d180 ffff880137303800 configfs configfs /sys/kernel/config ffff88012eb4d340 ffff880137307800 ext4 /dev/nbd0 / ffff8801376688c0 ffff8800b5203800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012eb4da40 ffff8800b410f800 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880137668a80 ffff8800b51c1800 hugetlbfs hugetlbfs /dev/hugepages ffff880137668c40 ffff88012b290800 mqueue mqueue /dev/mqueue ffff880137668e00 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff88012eb4dc00 ffff8800b410e000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff880137669180 ffff8800b51c4000 ramfs none /mnt ffff88012eb4ddc0 ffff8800b4147800 squashfs /dev/vda /home/green/git/lustre-release ffff8801377f7a40 ffff8800b4b06000 tmpfs none /var/lib/stateless/writable ffff880137669340 ffff8800b4b06000 tmpfs none /var/cache/man ffff88012eb4d880 ffff8800b4b06000 tmpfs none /var/log ffff880137669500 ffff8800b4b06000 tmpfs none /var/lib/dbus ffff8801376696c0 ffff8800b4b06000 tmpfs none /tmp ffff880129f2c700 ffff8800b4b06000 tmpfs none /var/lib/dhclient ffff88012eb4d6c0 ffff8800b4b06000 tmpfs none /var/tmp ffff880137669880 ffff8800b4b06000 tmpfs none /var/lib/NetworkManager ffff88012eb4d500 ffff8800b4b06000 tmpfs none /var/lib/systemd/random-seed ffff880129f2c8c0 ffff8800b4b06000 tmpfs none /var/spool ffff8800b2398000 ffff8800b4b06000 tmpfs none /var/lib/nfs ffff880137669a40 ffff8800b4b06000 tmpfs none /var/lib/gssproxy ffff8800b23981c0 ffff8800b4b06000 tmpfs none /var/lib/logrotate ffff880129f2ca80 ffff8800b4b06000 tmpfs none /etc ffff880137669c00 ffff8800b4b06000 tmpfs none /var/lib/rsyslog ffff880137669dc0 ffff8800b4b06000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff8800b2398380 ffff8800b4956000 nfs4 192.168.200.253:/exports/state/oleg124-server.virtnet /var/lib/stateless/state ffff8800b23988c0 ffff8800b4956000 nfs4 192.168.200.253:/exports/state/oleg124-server.virtnet /boot ffff880137668700 ffff8800b4956000 nfs4 192.168.200.253:/exports/state/oleg124-server.virtnet /etc/etc/kdump.conf ffff880138ccb340 ffff8800b5203800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b19b1dc0 ffff8800b4147800 squashfs /dev/vda /usr/sbin/mount.lustre ffff880129f2c1c0 ffff880084a3d800 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 ffff880129f2d500 ffff88008c7f6000 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff880129f2dc00 ffff8800b4b05800 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff880129f2d340 ffff880094b47000 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 413.321430] Lustre: Skipped 1 previous similar message [ 413.323381] Pid: 9552, comm: ll_ost00_002 3.10.0-7.9-debug #1 SMP Sat Mar 26 23:28:42 EDT 2022 [ 413.326194] Call Trace: [ 413.327442] [<0>] ptlrpc_set_wait+0x7cf/0x850 [ptlrpc] [ 413.329251] [<0>] ldlm_run_ast_work+0xe3/0x400 [ptlrpc] [ 413.331847] [<0>] ldlm_handle_conflict_lock+0x70/0x300 [ptlrpc] [ 413.333310] [<0>] ldlm_lock_enqueue+0x27d/0x930 [ptlrpc] [ 413.334725] [<0>] ldlm_cli_enqueue_local+0x36e/0x890 [ptlrpc] [ 413.336567] [<0>] ofd_destroy_by_fid+0x19c/0x610 [ofd] [ 413.338050] [<0>] ofd_destroy_hdl+0x20c/0xae0 [ofd] [ 413.339515] [<0>] tgt_request_handle+0x74e/0x1a60 [ptlrpc] [ 413.341186] [<0>] ptlrpc_server_handle_request+0x281/0xce0 [ptlrpc] [ 413.342832] [<0>] ptlrpc_main+0xc7e/0x1690 [ptlrpc] [ 413.344340] [<0>] kthread+0xe4/0xf0 [ 413.345334] [<0>] ret_from_fork_nospec_begin+0x7/0x21 [ 413.346686] [<0>] 0xfffffffffffffffe [ 417.924120] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1731518034/real 1731518034] req@ffff880091b86a00 x1815627865181568/t0(0) o104->lustre-MDT0000@192.168.201.24@tcp:15/16 lens 328/224 e 0 to 1 dl 1731518050 ref 1 fl Rpc:XQr/2/ffffffff rc 0/-1 job:'' uid:4294967295 gid:4294967295 [ 417.931591] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 433.933053] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1731518050/real 1731518050] req@ffff880091b86a00 x1815627865181568/t0(0) o104->lustre-MDT0000@192.168.201.24@tcp:15/16 lens 328/224 e 0 to 1 dl 1731518066 ref 1 fl Rpc:XQr/2/ffffffff rc 0/-1 job:'' uid:4294967295 gid:4294967295 [ 433.941818] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 3 previous similar messages [ 449.945093] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1731518066/real 1731518066] req@ffff880091b86a00 x1815627865181568/t0(0) o104->lustre-MDT0000@192.168.201.24@tcp:15/16 lens 328/224 e 0 to 1 dl 1731518082 ref 1 fl Rpc:XQr/2/ffffffff rc 0/-1 job:'' uid:4294967295 gid:4294967295 [ 449.954920] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 3 previous similar messages [ 481.957043] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1731518098/real 1731518098] req@ffff880091b86a00 x1815627865181568/t0(0) o104->lustre-MDT0000@192.168.201.24@tcp:15/16 lens 328/224 e 0 to 1 dl 1731518114 ref 1 fl Rpc:XQr/2/ffffffff rc 0/-1 job:'' uid:4294967295 gid:4294967295 [ 481.963353] Lustre: 7208:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 7 previous similar messages [ 481.965543] LustreError: 7208:0:(ldlm_lockd.c:759:ldlm_handle_ast_error()) ### client (nid 192.168.201.24@tcp) failed to reply to blocking AST (req@ffff880091b86a00 x1815627865181568 status 0 rc -110), evict it ns: mdt-lustre-MDT0000_UUID lock: ffff8800a0a86f40/0x7e41158294fbf98a lrc: 4/0,0 mode: PR/PR res: [0x200000007:0x1:0x0].0x0 bits 0x13/0x0 rrc: 3 type: IBT gid 0 flags: 0x60200400000020 nid: 192.168.201.24@tcp remote: 0x2f94399faf9e3d5 expref: 9 pid: 7208 timeout: 565 lvb_type: 0 [ 481.975427] LustreError: lustre-MDT0000: A client on nid 192.168.201.24@tcp was evicted due to a lock blocking callback time out: rc -110 [ 481.978489] LustreError: 7198:0:(ldlm_lockd.c:241:expired_lock_main()) ### lock callback timer expired after 16s: evicting client at 192.168.201.24@tcp ns: mdt-lustre-MDT0000_UUID lock: ffff8800a0a86f40/0x7e41158294fbf98a lrc: 3/0,0 mode: PR/PR res: [0x200000007:0x1:0x0].0x0 bits 0x13/0x0 rrc: 3 type: IBT gid 0 flags: 0x60200400000020 nid: 192.168.201.24@tcp remote: 0x2f94399faf9e3d5 expref: 10 pid: 7208 timeout: 0 lvb_type: 0 [ 481.994338] Lustre: mdt00_002: service thread pid 7208 completed after 112.091s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 484.014683] Lustre: DEBUG MARKER: == recovery-small test 10b: re-send BL AST =============== 12:15:17 (1731518117) [ 485.243073] LustreError: 11034:0:(ldlm_lockd.c:759:ldlm_handle_ast_error()) ### client (nid 192.168.201.24@tcp) failed to reply to blocking AST (req@ffff88009152ce00 x1815627865184512 status 0 rc -110), evict it ns: filter-lustre-OST0001_UUID lock: ffff880091b83840/0x7e41158294fbf856 lrc: 4/0,0 mode: PW/PW res: [0x2c0000401:0x4:0x0].0x0 rrc: 3 type: EXT [0->18446744073709551615] (req 0->4194303) gid 0 flags: 0x60000400030020 nid: 192.168.201.24@tcp remote: 0x2f94399faf9e373 expref: 7 pid: 9550 timeout: 568 lvb_type: 0 [ 485.243080] LustreError: lustre-OST0001: A client on nid 192.168.201.24@tcp was evicted due to a lock blocking callback time out: rc -110 [ 485.243149] LustreError: 7198:0:(ldlm_lockd.c:241:expired_lock_main()) ### lock callback timer expired after 16s: evicting client at 192.168.201.24@tcp ns: filter-lustre-OST0000_UUID lock: ffff880084a0db00/0x7e41158294fbf8aa lrc: 3/0,0 mode: PW/PW res: [0x280000401:0x4:0x0].0x0 rrc: 3 type: EXT [0->18446744073709551615] (req 0->4095) gid 0 flags: 0x60000400030020 nid: 192.168.201.24@tcp remote: 0x2f94399faf9e396 expref: 7 pid: 9550 timeout: 0 lvb_type: 0 [ 485.244650] Lustre: ll_ost00_002: service thread pid 9552 completed after 112.015s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 485.247179] Lustre: ll_ost00_001: service thread pid 9551 completed after 112.018s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 485.272276] LustreError: 11034:0:(ldlm_lockd.c:759:ldlm_handle_ast_error()) Skipped 2 previous similar messages [ 485.278125] Lustre: ll_ost00_004: service thread pid 11034 completed after 112.049s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 501.665780] Lustre: DEBUG MARKER: == recovery-small test 10c: re-send BL AST vs reconnect race (LU-5569) ========================================================== 12:15:34 (1731518134) [ 502.751519] Lustre: lustre-MDT0001: Client 45cf2ef9-f1b7-4660-85c4-f4d6f1981a61 (at 192.168.201.24@tcp) reconnecting [ 502.754804] Lustre: Skipped 2 previous similar messages [ 504.165870] Lustre: DEBUG MARKER: == recovery-small test 10d: test failed blocking ast ===== 12:15:37 (1731518137) ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 11.89s (real) 6.32s (CPU), Child processes: 5.51s