************************ crashinfo ************************* /exports/testreports/51288/testresults/racer-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg435-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.ORfmO/vmlinux [TAINTED] DUMPFILE: /exports/testreports/51288/testresults/racer-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg435-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Thu May 1 16:30:34 EDT 2025 UPTIME: 00:33:39 LOAD AVERAGE: 0.00, 0.03, 0.38 TASKS: 316 NODENAME: oleg435-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2400 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 25 last 5s 62 last 60s 84 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 312 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 50 kworker/1:1 0 ms (no user stack) 49 kworker/0:1 0 ms (no user stack) 17643 kworker/3:1 0 ms (no user stack) 46 kworker/2:1 0 ms (no user stack) 3765 ptlrpcd_00_01 3 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 677582 2.6 GB 70% of TOTAL MEM USED 277485 1.1 GB 29% of TOTAL MEM SHARED 31288 122.2 MB 3% of TOTAL MEM BUFFERS 28969 113.2 MB 3% of TOTAL MEM CACHED 98035 382.9 MB 10% of TOTAL MEM SLAB 22947 89.6 MB 2% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 59008 230.5 MB 7% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 2017.2 eth0 n/a 0.6 RSS_TOTAL=51424 pages, %mem= 0.8 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137668380 ffff880137678800 sysfs sysfs /sys ffff880137668540 ffff880139944000 proc proc /proc ffff880137668700 ffff880137678000 devtmpfs devtmpfs /dev ffff8801376688c0 ffff8800b5213000 securityfs securityfs /sys/kernel/security ffff880137668a80 ffff880137679000 tmpfs tmpfs /dev/shm ffff880137668c40 ffff88012b298800 devpts devpts /dev/pts ffff880137668e00 ffff880137679800 tmpfs tmpfs /run ffff880137668fc0 ffff88013767a000 tmpfs tmpfs /sys/fs/cgroup ffff880137669180 ffff88013767a800 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669340 ffff88013767b000 pstore pstore /sys/fs/pstore ffff880137669500 ffff88013767d000 cgroup cgroup /sys/fs/cgroup/hugetlb ffff8801376696c0 ffff88013767c800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff880137669880 ffff88013767c000 cgroup cgroup /sys/fs/cgroup/pids ffff880137669a40 ffff88013767b800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff880137669c00 ffff88013767d800 cgroup cgroup /sys/fs/cgroup/devices ffff880137669dc0 ffff88013767e000 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a330000 ffff88013767e800 cgroup cgroup /sys/fs/cgroup/memory ffff88012a3301c0 ffff88013767f000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a330380 ffff88013767f800 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a330540 ffff88012a370000 cgroup cgroup /sys/fs/cgroup/perf_event ffff880137012c40 ffff88012b3ed800 configfs configfs /sys/kernel/config ffff880137012e00 ffff88012b3ee000 ext4 /dev/nbd0 / ffff88012b2ce8c0 ffff8800b5216800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012a3308c0 ffff8800b45d0000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880138ccbc00 ffff88012a377000 hugetlbfs hugetlbfs /dev/hugepages ffff88012a330a80 ffff880137310800 mqueue mqueue /dev/mqueue ffff88012a330c40 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff88012b2cea80 ffff8800b453a000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff88012b2cec40 ffff8800b3e6e000 ramfs none /mnt ffff88012b2cee00 ffff8800b3e6a800 tmpfs none /var/lib/stateless/writable ffff88012a330e00 ffff8800b478b000 squashfs /dev/vda /home/green/git/lustre-release ffff88012b2cefc0 ffff8800b3e6a800 tmpfs none /var/cache/man ffff88012a330fc0 ffff8800b3e6a800 tmpfs none /var/log ffff88012a331180 ffff8800b3e6a800 tmpfs none /var/lib/dbus ffff88012b2cf180 ffff8800b3e6a800 tmpfs none /tmp ffff88012a331340 ffff8800b3e6a800 tmpfs none /var/lib/dhclient ffff88012a331500 ffff8800b3e6a800 tmpfs none /var/tmp ffff88012a3316c0 ffff8800b3e6a800 tmpfs none /var/lib/NetworkManager ffff88012a331880 ffff8800b3e6a800 tmpfs none /var/lib/systemd/random-seed ffff88012b2cf340 ffff8800b3e6a800 tmpfs none /var/spool ffff88012b2cf500 ffff8800b3e6a800 tmpfs none /var/lib/nfs ffff88012a331a40 ffff8800b3e6a800 tmpfs none /var/lib/gssproxy ffff88012a331c00 ffff8800b3e6a800 tmpfs none /var/lib/logrotate ffff88012a331dc0 ffff8800b3e6a800 tmpfs none /etc ffff8800b1c5e000 ffff8800b3e6a800 tmpfs none /var/lib/rsyslog ffff8800b1c5e1c0 ffff8800b3e6a800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff88012b2cf6c0 ffff8800b453d800 nfs4 192.168.200.253:/exports/state/oleg435-server.virtnet /var/lib/stateless/state ffff88012b2cfa40 ffff8800b453d800 nfs4 192.168.200.253:/exports/state/oleg435-server.virtnet /boot ffff8800b1c5e380 ffff8800b453d800 nfs4 192.168.200.253:/exports/state/oleg435-server.virtnet /etc/etc/kdump.conf ffff8800b1c5e540 ffff8800b5216800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff88012a322c40 ffff8800b478b000 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800b1aa3880 ffff88012a376000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 ffff8800b1aa2380 ffff8800950fd000 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff8800b1aa3dc0 ffff8800997a9000 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff8800b3cda540 ffff8800a6391800 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 497.161433] [<0>] 0xfffffffffffffffe [ 498.055737] Lustre: mdt_io00_010: service thread pid 16474 was inactive for 40.083 seconds. Watchdog stack traces are limited to 3 per 300 seconds, skipping this one. [ 498.062388] Lustre: Skipped 19 previous similar messages [ 501.511738] Lustre: mdt_io00_013: service thread pid 16480 was inactive for 40.004 seconds. Watchdog stack traces are limited to 3 per 300 seconds, skipping this one. [ 501.516538] Lustre: Skipped 1 previous similar message [ 535.560068] Lustre: mdt_io00_007: service thread pid 16204 was inactive for 72.273 seconds. Watchdog stack traces are limited to 3 per 300 seconds, skipping this one. [ 557.576246] LustreError: 7218:0:(ldlm_lockd.c:257:expired_lock_main()) ### lock callback timer expired after 101s: evicting client at 192.168.204.35@tcp ns: mdt-lustre-MDT0001_UUID lock: ffff88009d2ee200/0x1e73857d435475ea lrc: 3/0,0 mode: PR/PR res: [0x240000403:0x2254:0x0].0x0 bits 0x13/0x0 rrc: 5 type: IBT gid 0 flags: 0x60200400000020 nid: 192.168.204.35@tcp remote: 0xa7d7db0ff071cc0c expref: 1387 pid: 16164 timeout: 556 lvb_type: 0 [ 557.636833] LustreError: 16160:0:(ldlm_lockd.c:1447:ldlm_handle_enqueue()) ### lock on destroyed export ffff880084ad8000 ns: mdt-lustre-MDT0001_UUID lock: ffff880072330c00/0x1e73857d4354a7a7 lrc: 3/0,0 mode: PR/PR res: [0x240000402:0x240f:0x0].0x0 bits 0x1b/0x0 rrc: 4 type: IBT gid 0 flags: 0x50200400000020 nid: 192.168.204.35@tcp remote: 0xa7d7db0ff071da98 expref: 1276 pid: 16160 timeout: 0 lvb_type: 0 [ 557.643551] Lustre: mdt00_016: service thread pid 16167 completed after 100.262s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.648728] Lustre: mdt_io00_000: service thread pid 7240 completed after 100.504s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.648897] Lustre: mdt00_008: service thread pid 16158 completed after 100.174s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.649041] Lustre: mdt00_017: service thread pid 16168 completed after 100.189s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.649162] Lustre: mdt00_019: service thread pid 16228 completed after 100.190s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.649279] Lustre: mdt00_021: service thread pid 16237 completed after 100.190s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.649400] Lustre: mdt00_011: service thread pid 16161 completed after 100.316s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.649674] Lustre: mdt00_003: service thread pid 9592 completed after 100.190s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.649797] Lustre: mdt00_020: service thread pid 16236 completed after 100.190s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.651638] Lustre: mdt_io00_008: service thread pid 16205 completed after 100.497s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.659543] Lustre: mdt_io00_002: service thread pid 7242 completed after 100.502s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.660090] Lustre: mdt00_012: service thread pid 16163 completed after 100.316s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.660351] Lustre: mdt00_007: service thread pid 16157 completed after 100.383s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.660456] Lustre: mdt00_018: service thread pid 16169 completed after 100.395s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.666561] LustreError: 7241:0:(client.c:1282:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff880071dd0000 x1830949458860288/t0(0) o104->lustre-MDT0001@192.168.204.35@tcp:15/16 lens 328/224 e 0 to 0 dl 0 ref 1 fl Rpc:QU/0/ffffffff rc 0/-1 job:'' uid:4294967295 gid:4294967295 [ 557.678708] LustreError: 7213:0:(ldlm_lockd.c:2549:ldlm_cancel_handler()) ldlm_cancel from 192.168.204.35@tcp arrived at 1746129972 with bad export cookie 2194244216500901200 [ 557.792887] LustreError: 16160:0:(ldlm_lockd.c:1447:ldlm_handle_enqueue()) Skipped 7 previous similar messages [ 557.799156] Lustre: mdt00_009: service thread pid 16159 completed after 100.496s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.799842] Lustre: mdt00_010: service thread pid 16160 completed after 100.426s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.816505] Lustre: mdt00_000: service thread pid 7227 completed after 100.684s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.838237] Lustre: mdt_io00_001: service thread pid 7241 completed after 100.653s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.838960] Lustre: mdt_io00_009: service thread pid 16326 completed after 100.639s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.849596] Lustre: mdt_io00_004: service thread pid 16187 completed after 100.503s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.855226] Lustre: mdt_io00_006: service thread pid 16199 completed after 100.489s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.880142] Lustre: mdt_io00_005: service thread pid 16198 completed after 99.937s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.887598] Lustre: mdt_io00_010: service thread pid 16474 completed after 99.915s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.888973] LustreError: 16475:0:(mdt_reint.c:2523:mdt_reint_migrate()) lustre-MDT0000: migrate [0x200000402:0x2:0x0]/9 failed: rc = -2 [ 557.888976] LustreError: 16475:0:(mdt_reint.c:2523:mdt_reint_migrate()) Skipped 47 previous similar messages [ 557.889037] Lustre: mdt_io00_011: service thread pid 16475 completed after 99.887s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.897528] Lustre: mdt_io00_003: service thread pid 16174 completed after 99.711s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.911496] Lustre: mdt_io00_013: service thread pid 16480 completed after 96.404s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 557.922872] Lustre: mdt_io00_007: service thread pid 16204 completed after 94.636s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 13.32s (real) 7.10s (CPU), Child processes: 6.13s