************************ crashinfo ************************* /exports/testreports/51830/testresults/racer-ldiskfs-DNE-centos7_x86_64-centos7_x86_64-retry1/oleg233-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.2WLbv/vmlinux [TAINTED] DUMPFILE: /exports/testreports/51830/testresults/racer-ldiskfs-DNE-centos7_x86_64-centos7_x86_64-retry1/oleg233-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Wed May 21 02:50:22 EDT 2025 UPTIME: 00:16:56 LOAD AVERAGE: 0.00, 0.37, 0.74 TASKS: 320 NODENAME: oleg233-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 59 last 5s 72 last 60s 92 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 316 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 905 in:imjournal 0 ms /usr/sbin/rsyslogd -n 7194 ldlm_cn00_000 0 ms (no user stack) 16122 mdt_rdpg00_002 0 ms (no user stack) 46 kworker/1:1 0 ms (no user stack) 9553 ll_ost_create00 3 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 653972 2.5 GB 68% of TOTAL MEM USED 301095 1.1 GB 31% of TOTAL MEM SHARED 43877 171.4 MB 4% of TOTAL MEM BUFFERS 41509 162.1 MB 4% of TOTAL MEM CACHED 106613 416.5 MB 11% of TOTAL MEM SLAB 26151 102.2 MB 2% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 58972 230.4 MB 7% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- CLOSE 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 1014.3 eth0 n/a 1.1 RSS_TOTAL=52524 pages, %mem= 0.9 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff88012a380000 ffff88012a388000 sysfs sysfs /sys ffff88012a3801c0 ffff880139944000 proc proc /proc ffff88012a380380 ffff880137678000 devtmpfs devtmpfs /dev ffff88012a380540 ffff8800b5216800 securityfs securityfs /sys/kernel/security ffff88012a380700 ffff88012a388800 tmpfs tmpfs /dev/shm ffff88012a3808c0 ffff880137306000 devpts devpts /dev/pts ffff88012a380a80 ffff88012a389000 tmpfs tmpfs /run ffff88012a380c40 ffff88012a389800 tmpfs tmpfs /sys/fs/cgroup ffff88012a380e00 ffff88012a38a000 cgroup cgroup /sys/fs/cgroup/systemd ffff88012a380fc0 ffff88012a38a800 pstore pstore /sys/fs/pstore ffff88012a381180 ffff88012a38c800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012a381340 ffff88012a38c000 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a381500 ffff88012a38b800 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a3816c0 ffff88012a38b000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a381880 ffff88012a38d000 cgroup cgroup /sys/fs/cgroup/devices ffff88012a381a40 ffff88012a38d800 cgroup cgroup /sys/fs/cgroup/pids ffff88012a381c00 ffff88012a38e000 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a381dc0 ffff88012a38e800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012a3bc000 ffff88012a38f000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a3bc1c0 ffff88012a38f800 cgroup cgroup /sys/fs/cgroup/memory ffff88012a3bc380 ffff8800b5253000 configfs configfs /sys/kernel/config ffff88012b2b4a80 ffff8800b7c07000 ext4 /dev/nbd0 / ffff88012a3bc8c0 ffff8800b51e3800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012b2b5180 ffff8800b51e6000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880138ccbdc0 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff8800b7b26000 ffff88012b230800 hugetlbfs hugetlbfs /dev/hugepages ffff880137668a80 ffff880137307800 mqueue mqueue /dev/mqueue ffff88012b2b5340 ffff8800b4151800 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff88012b2b56c0 ffff8800b4153000 ramfs none /mnt ffff88012a3bca80 ffff8800b7c06800 tmpfs none /var/lib/stateless/writable ffff88012a3bcc40 ffff8800b7c07800 squashfs /dev/vda /home/green/git/lustre-release ffff88012a3bce00 ffff8800b7c06800 tmpfs none /var/cache/man ffff88012a3bcfc0 ffff8800b7c06800 tmpfs none /var/log ffff880137668c40 ffff8800b7c06800 tmpfs none /var/lib/dbus ffff8800b7b26380 ffff8800b7c06800 tmpfs none /tmp ffff880137668e00 ffff8800b7c06800 tmpfs none /var/lib/dhclient ffff88012a3bd180 ffff8800b7c06800 tmpfs none /var/tmp ffff880137668fc0 ffff8800b7c06800 tmpfs none /var/lib/NetworkManager ffff880137669180 ffff8800b7c06800 tmpfs none /var/lib/systemd/random-seed ffff88012a3bd340 ffff8800b7c06800 tmpfs none /var/spool ffff8800b7b26540 ffff8800b7c06800 tmpfs none /var/lib/nfs ffff88012a3bd500 ffff8800b7c06800 tmpfs none /var/lib/gssproxy ffff88012a3bd6c0 ffff8800b7c06800 tmpfs none /var/lib/logrotate ffff88012a3bd880 ffff8800b7c06800 tmpfs none /etc ffff88012b2b5880 ffff8800b7c06800 tmpfs none /var/lib/rsyslog ffff88012a3bda40 ffff8800b7c06800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff88012a3bdc00 ffff880129c5b800 nfs4 192.168.200.253:/exports/state/oleg233-server.virtnet /var/lib/stateless/state ffff88012a3bddc0 ffff880129c5b800 nfs4 192.168.200.253:/exports/state/oleg233-server.virtnet /boot ffff88012b8fe000 ffff880129c5b800 nfs4 192.168.200.253:/exports/state/oleg233-server.virtnet /etc/etc/kdump.conf ffff88012b8fe1c0 ffff8800b51e3800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff88012b8ff880 ffff8800b7c07800 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800b7b26a80 ffff8800b4a25000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 ffff8800b7b27180 ffff880082139000 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff8800b7b26c40 ffff88009899c800 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff8800b7b26fc0 ffff8800997c0800 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 371.411392] [<0>] mdt_object_lock_internal+0x1b3/0x470 [mdt] [ 371.413598] [<0>] mdt_object_lock+0x88/0x1c0 [mdt] [ 371.415040] [<0>] mdt_getattr_name_lock+0xc4a/0x2cc0 [mdt] [ 371.416108] [<0>] mdt_intent_getattr+0x2cc/0x4e0 [mdt] [ 371.417621] [<0>] mdt_intent_opc.constprop.74+0x211/0xc60 [mdt] [ 371.419255] [<0>] mdt_intent_policy+0x10f/0x460 [mdt] [ 371.420595] [<0>] ldlm_lock_enqueue+0x397/0x980 [ptlrpc] [ 371.421829] [<0>] ldlm_handle_enqueue+0x547/0x1900 [ptlrpc] [ 371.423154] [<0>] tgt_enqueue+0x68/0x240 [ptlrpc] [ 371.424576] [<0>] tgt_request_handle+0x74e/0x1a60 [ptlrpc] [ 371.425917] [<0>] ptlrpc_server_handle_request+0x257/0xcd0 [ptlrpc] [ 371.427515] [<0>] ptlrpc_main+0xc61/0x1640 [ptlrpc] [ 371.428796] [<0>] kthread+0xe4/0xf0 [ 371.430063] [<0>] ret_from_fork_nospec_begin+0x7/0x21 [ 371.431843] [<0>] 0xfffffffffffffffe [ 371.661292] Lustre: mdt00_024: service thread pid 16383 was inactive for 40.016 seconds. Watchdog stack traces are limited to 3 per 300 seconds, skipping this one. [ 371.661294] Lustre: mdt00_006: service thread pid 16123 was inactive for 40.016 seconds. Watchdog stack traces are limited to 3 per 300 seconds, skipping this one. [ 371.661301] Lustre: Skipped 8 previous similar messages [ 371.672736] Lustre: Skipped 2 previous similar messages [ 431.053306] LustreError: 7199:0:(ldlm_lockd.c:257:expired_lock_main()) ### lock callback timer expired after 100s: evicting client at 192.168.202.33@tcp ns: mdt-lustre-MDT0001_UUID lock: ffff8800972e1e00/0x478f77978a2911a9 lrc: 3/0,0 mode: CR/CR res: [0x240000403:0x42f8:0x0].0x0 bits 0xa/0x0 rrc: 5 type: IBT gid 0 flags: 0x60200400000020 nid: 192.168.202.33@tcp remote: 0x3a1f858721849053 expref: 2195 pid: 16127 timeout: 430 lvb_type: 0 [ 431.090710] Lustre: mdt00_023: service thread pid 16225 completed after 100.095s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.092032] Lustre: mdt00_014: service thread pid 16131 completed after 100.078s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.092186] LustreError: 16153:0:(ldlm_lockd.c:2550:ldlm_cancel_handler()) ldlm_cancel from 192.168.202.33@tcp arrived at 1747809636 with bad export cookie 5156471591094174383 [ 431.096539] LustreError: 16384:0:(ldlm_lockd.c:1447:ldlm_handle_enqueue()) ### lock on destroyed export ffff88009a059000 ns: mdt-lustre-MDT0001_UUID lock: ffff880133aa6000/0x478f77978a2984c7 lrc: 3/0,0 mode: PR/PR res: [0x240000402:0x1:0x0].0x0 bits 0x12/0x0 rrc: 18 type: IBT gid 0 flags: 0x50200000000000 nid: 192.168.202.33@tcp remote: 0x3a1f85872184b35a expref: 16 pid: 16384 timeout: 0 lvb_type: 0 [ 431.096599] Lustre: mdt00_006: service thread pid 16123 completed after 99.451s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096603] Lustre: mdt00_025: service thread pid 16384 completed after 99.432s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096700] Lustre: mdt00_018: service thread pid 16135 completed after 99.451s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096715] Lustre: mdt00_024: service thread pid 16383 completed after 99.451s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096736] Lustre: mdt_io00_009: service thread pid 16243 completed after 100.030s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096779] Lustre: mdt00_002: service thread pid 7209 completed after 99.717s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096827] Lustre: mdt00_000: service thread pid 7207 completed after 99.477s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096890] Lustre: mdt00_004: service thread pid 11574 completed after 99.477s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.096926] Lustre: mdt00_020: service thread pid 16138 completed after 99.703s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.097016] Lustre: mdt00_021: service thread pid 16139 completed after 99.477s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.097410] Lustre: mdt00_003: service thread pid 9572 completed after 99.718s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.098277] Lustre: mdt00_019: service thread pid 16137 completed after 99.479s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.100321] Lustre: mdt00_016: service thread pid 16133 completed after 99.720s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.100393] Lustre: mdt00_015: service thread pid 16132 completed after 99.725s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.100447] Lustre: mdt00_012: service thread pid 16129 completed after 99.714s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 431.101392] Lustre: mdt00_010: service thread pid 16127 completed after 99.708s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 11.59s (real) 6.38s (CPU), Child processes: 5.17s