************************ crashinfo ************************* /exports/testreports/47322/testresults/conf-sanity3-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg104-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.T36A1/vmlinux [TAINTED] DUMPFILE: /exports/testreports/47322/testresults/conf-sanity3-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg104-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Sun Nov 24 03:40:05 EST 2024 UPTIME: 02:30:19 LOAD AVERAGE: 0.74, 0.70, 0.66 TASKS: 162 NODENAME: oleg104-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 35 last 5s 50 last 60s 65 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 158 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 19473 kworker/0:1 0 ms (no user stack) 11 rcuos/0 0 ms (no user stack) 20 rcuos/1 0 ms (no user stack) 34 rcuos/3 0 ms (no user stack) 9 rcu_sched 0 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 812209 3.1 GB 85% of TOTAL MEM USED 142858 558 MB 14% of TOTAL MEM SHARED 9323 36.4 MB 0% of TOTAL MEM BUFFERS 6655 26 MB 0% of TOTAL MEM CACHED 75454 294.7 MB 7% of TOTAL MEM SLAB 16017 62.6 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 62691 244.9 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 10 LISTEN 3 NAGLE disabled (TCP_NODELAY): 8 user_data set (NFS etc.): 8 Unusual Situations: Doing Retransmission: 1 (run xportshow --retrans for details) UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 18 LISTEN 8 Raw sockets info -------------------- CLOSE 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 9016.4 eth0 n/a 0.2 RSS_TOTAL=56632 pages, %mem= 0.9 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff88012a352000 ffff880137320800 sysfs sysfs /sys ffff88012a3521c0 ffff880139944000 proc proc /proc ffff88012a352380 ffff880137678000 devtmpfs devtmpfs /dev ffff88012a352540 ffff8800b5219800 securityfs securityfs /sys/kernel/security ffff88012a352700 ffff880137321000 tmpfs tmpfs /dev/shm ffff88012a3528c0 ffff880137189800 devpts devpts /dev/pts ffff88012a352a80 ffff880137321800 tmpfs tmpfs /run ffff88012a352c40 ffff880137322000 tmpfs tmpfs /sys/fs/cgroup ffff88012a352e00 ffff880137322800 cgroup cgroup /sys/fs/cgroup/systemd ffff88012a352fc0 ffff880137323000 pstore pstore /sys/fs/pstore ffff88012a353180 ffff880137325000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012a353340 ffff880137324800 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a353500 ffff880137324000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a3536c0 ffff880137323800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a353880 ffff880137325800 cgroup cgroup /sys/fs/cgroup/pids ffff88012a353a40 ffff880137326000 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012a353c00 ffff880137326800 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a353dc0 ffff880137327000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a35e000 ffff880137327800 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a35e1c0 ffff88012a388000 cgroup cgroup /sys/fs/cgroup/devices ffff880137012fc0 ffff880129c10800 configfs configfs /sys/kernel/config ffff880137013180 ffff880129c12800 ext4 /dev/nbd0 / ffff880137668c40 ffff8800b521a800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012a35e380 ffff88012a38f000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880137013880 ffff88013718b000 mqueue mqueue /dev/mqueue ffff88012a35e540 ffff88012a389000 hugetlbfs hugetlbfs /dev/hugepages ffff880137668e00 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff880137668fc0 ffff880129cbb000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff880138ccbc00 ffff8800b51d7800 ramfs none /mnt ffff880137669180 ffff880129c14800 squashfs /dev/vda /home/green/git/lustre-release ffff880137013a40 ffff880129d1d000 tmpfs none /var/lib/stateless/writable ffff880137669340 ffff880129d1d000 tmpfs none /var/cache/man ffff880137669500 ffff880129d1d000 tmpfs none /var/log ffff8801376696c0 ffff880129d1d000 tmpfs none /var/lib/dbus ffff88012a35e700 ffff880129d1d000 tmpfs none /tmp ffff880137669880 ffff880129d1d000 tmpfs none /var/lib/dhclient ffff880137669a40 ffff880129d1d000 tmpfs none /var/tmp ffff880137669c00 ffff880129d1d000 tmpfs none /var/lib/NetworkManager ffff88012a35e8c0 ffff880129d1d000 tmpfs none /var/lib/systemd/random-seed ffff88012a35ea80 ffff880129d1d000 tmpfs none /var/spool ffff880137669dc0 ffff880129d1d000 tmpfs none /var/lib/nfs ffff880137668700 ffff880129d1d000 tmpfs none /var/lib/gssproxy ffff8800b21aa000 ffff880129d1d000 tmpfs none /var/lib/logrotate ffff8800b21aa1c0 ffff880129d1d000 tmpfs none /etc ffff8800b21aa380 ffff880129d1d000 tmpfs none /var/lib/rsyslog ffff8800b21aa540 ffff880129d1d000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff880137013c00 ffff8800b41ec800 nfs4 192.168.200.253:/exports/state/oleg104-server.virtnet /var/lib/stateless/state ffff8800b21aa700 ffff8800b41ec800 nfs4 192.168.200.253:/exports/state/oleg104-server.virtnet /boot ffff880137013dc0 ffff8800b41ec800 nfs4 192.168.200.253:/exports/state/oleg104-server.virtnet /etc/etc/kdump.conf ffff8800b21aa8c0 ffff8800b521a800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff88012e2dd180 ffff880129c14800 squashfs /dev/vda /usr/sbin/mount.lustre ffff88012e2dda40 ffff88009b759800 tmpfs tmpfs /run/user/0 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 8962.767709] Lustre: DEBUG MARKER: oleg104-server.virtnet: executing set_default_debug -1 all [ 8962.853271] LustreError: 16779:0:(mgc_request.c:1451:mgc_apply_recover_logs()) mgc: cannot find UUID by nid '192.168.201.104@tcp': rc = -2 [ 8962.855650] LustreError: 16779:0:(mgc_request.c:1451:mgc_apply_recover_logs()) Skipped 1 previous similar message [ 8962.857477] Lustre: 16779:0:(mgc_request.c:1628:mgc_process_recover_log()) MGC192.168.201.104@tcp: error processing lustre-mdtir log recovery: rc = -2 [ 8962.860144] Lustre: 16779:0:(mgc_request.c:1628:mgc_process_recover_log()) Skipped 1 previous similar message [ 8962.862079] Lustre: 16779:0:(mgc_request.c:1900:mgc_process_log()) MGC192.168.201.104@tcp: IR log lustre-mdtir failed, not fatal: rc = -2 [ 8962.864592] Lustre: 16779:0:(mgc_request.c:1900:mgc_process_log()) Skipped 1 previous similar message [ 8965.141281] Lustre: lustre-OST0003: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 8965.143558] Lustre: lustre-OST0003: Denying connection for new client 8efe2386-23b4-411e-b55b-0d3b23d8df5d (at 192.168.201.4@tcp), waiting for 2 known clients (0 recovered, 0 in progress, and 0 evicted) to recover in 0:59 [ 8966.334286] Lustre: lustre-OST0003-osc-MDT0001: Connection restored to 0@lo (at 0@lo) [ 8966.334295] Lustre: lustre-OST0003: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 8974.791283] Lustre: Failing over lustre-OST0003 [ 8974.807688] Lustre: server umount lustre-OST0003 complete [ 8975.163779] LustreError: lustre-OST0003: not available for connect from 192.168.201.4@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 8975.168988] LustreError: Skipped 2 previous similar messages [ 8975.822954] Lustre: DEBUG MARKER: == conf-sanity test 153a: bypass invalid NIDs quickly ==== 03:39:22 (1732437562) [ 8976.348410] Lustre: lustre-OST0003-osc-MDT0001: Connection to lustre-OST0003 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 8976.351752] Lustre: Skipped 2 previous similar messages [ 8986.365329] LustreError: lustre-OST0003: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 8986.374515] LustreError: Skipped 5 previous similar messages [ 8991.051656] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 8991.057001] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 8991.062023] Lustre: Skipped 2 previous similar messages [ 8992.752148] Lustre: server umount lustre-MDT0000 complete [ 8994.052127] LustreError: 18108:0:(ldlm_lockd.c:2575:ldlm_cancel_handler()) ldlm_cancel from 0@lo arrived at 1732437580 with bad export cookie 5278162864651093894 [ 8994.056000] LustreError: MGC192.168.201.104@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 8994.064859] LustreError: 18108:0:(ldlm_lockd.c:2575:ldlm_cancel_handler()) Skipped 4 previous similar messages [ 8996.380574] Lustre: lustre-MDT0001-lwp-OST0000: Connection to lustre-MDT0001 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 8996.387591] Lustre: Skipped 5 previous similar messages [ 9000.213573] Lustre: server umount lustre-MDT0001 complete [ 9001.662733] Lustre: server umount lustre-OST0000 complete [ 9002.991666] Lustre: server umount lustre-OST0001 complete [ 9004.732955] Lustre: DEBUG MARKER: oleg104-server.virtnet: executing set_hostid [ 9007.010007] Lustre: DEBUG MARKER: oleg104-server.virtnet: executing load_modules_local [ 9009.962828] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: errors=remount-ro [ 9012.328875] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: errors=remount-ro [ 9014.368820] LDISKFS-fs (dm-2): file extents enabled, maximum tree depth=5 [ 9014.374445] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro [ 9016.648747] LDISKFS-fs (dm-3): file extents enabled, maximum tree depth=5 [ 9016.657622] LDISKFS-fs (dm-3): mounted filesystem with ordered data mode. Opts: errors=remount-ro ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.31s (real) 6.29s (CPU), Child processes: 5.94s