************************ crashinfo ************************* /exports/testreports/42551/testresults/replay-dual-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg422-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.mxc2G/vmlinux [TAINTED] DUMPFILE: /exports/testreports/42551/testresults/replay-dual-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg422-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Wed May 8 18:48:16 EDT 2024 UPTIME: 00:47:28 LOAD AVERAGE: 0.12, 0.44, 0.58 TASKS: 271 NODENAME: oleg422-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 27 last 5s 63 last 60s 73 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 267 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 25681 kworker/2:2 0 ms (no user stack) 4442 l2arc_feed 0 ms (no user stack) 34 rcuos/3 0 ms (no user stack) 272 kworker/3:2 0 ms (no user stack) 1 systemd 0 ms /usr/lib/systemd/systemd --switched-root --system --deserialize 22 +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 716082 2.7 GB 74% of TOTAL MEM USED 238985 933.5 MB 25% of TOTAL MEM SHARED 16346 63.9 MB 1% of TOTAL MEM BUFFERS 12864 50.2 MB 1% of TOTAL MEM CACHED 81733 319.3 MB 8% of TOTAL MEM SLAB 18975 74.1 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 73508 287.1 MB 9% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 2845.6 eth0 n/a 2.1 RSS_TOTAL=55416 pages, %mem= 0.9 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff8801376688c0 ffff88012a2e0000 sysfs sysfs /sys ffff880137668a80 ffff880139944000 proc proc /proc ffff880137668c40 ffff880137678000 devtmpfs devtmpfs /dev ffff880137668e00 ffff88013730f800 securityfs securityfs /sys/kernel/security ffff880137668fc0 ffff88012a2e0800 tmpfs tmpfs /dev/shm ffff880137669180 ffff88013771a800 devpts devpts /dev/pts ffff880137669340 ffff88012a2e1000 tmpfs tmpfs /run ffff880137669500 ffff88012a2e1800 tmpfs tmpfs /sys/fs/cgroup ffff8801376696c0 ffff88012a2e2000 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669880 ffff88012a2e2800 pstore pstore /sys/fs/pstore ffff880137669a40 ffff88012a2e4800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff880137669c00 ffff88012a2e4000 cgroup cgroup /sys/fs/cgroup/perf_event ffff880137669dc0 ffff88012a2e3800 cgroup cgroup /sys/fs/cgroup/devices ffff88012a30c000 ffff88012a2e3000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a30c1c0 ffff88012a2e5000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a30c380 ffff88012a2e5800 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a30c540 ffff88012a2e6000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a30c700 ffff88012a2e6800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a30c8c0 ffff88012a2e7000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012a30ca80 ffff88012a2e7800 cgroup cgroup /sys/fs/cgroup/pids ffff88012b2a2380 ffff8800b5246800 configfs configfs /sys/kernel/config ffff880129e42540 ffff8800b51ed800 ext4 /dev/nbd0 / ffff88012a33a700 ffff88012a356000 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012a33ac40 ffff8800b4106000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff88012b2a2540 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff88012a33ae00 ffff88012b2b8000 mqueue mqueue /dev/mqueue ffff88012b2a2700 ffff880129ea9800 hugetlbfs hugetlbfs /dev/hugepages ffff88012a30cc40 ffff8800b4b75800 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff88012a33afc0 ffff8800b4100000 ramfs none /mnt ffff88012a30ce00 ffff8800b526d000 squashfs /dev/vda /home/green/git/lustre-release ffff88012a30cfc0 ffff8800b5269000 tmpfs none /var/lib/stateless/writable ffff88012b2a28c0 ffff8800b5269000 tmpfs none /var/cache/man ffff88012a30d180 ffff8800b5269000 tmpfs none /var/log ffff88012a30d340 ffff8800b5269000 tmpfs none /var/lib/dbus ffff880129e42a80 ffff8800b5269000 tmpfs none /tmp ffff88012b2a2a80 ffff8800b5269000 tmpfs none /var/lib/dhclient ffff88012a33b180 ffff8800b5269000 tmpfs none /var/tmp ffff88012a33b340 ffff8800b5269000 tmpfs none /var/lib/NetworkManager ffff88012a33b500 ffff8800b5269000 tmpfs none /var/lib/systemd/random-seed ffff88012a30d500 ffff8800b5269000 tmpfs none /var/spool ffff88012a30d6c0 ffff8800b5269000 tmpfs none /var/lib/nfs ffff88012a30d880 ffff8800b5269000 tmpfs none /var/lib/gssproxy ffff88012a30da40 ffff8800b5269000 tmpfs none /var/lib/logrotate ffff880129e42c40 ffff8800b5269000 tmpfs none /etc ffff880129e42e00 ffff8800b5269000 tmpfs none /var/lib/rsyslog ffff880129e42fc0 ffff8800b5269000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff88012a33b6c0 ffff880129c99800 nfs4 192.168.200.253:/exports/state/oleg422-server.virtnet /var/lib/stateless/state ffff88012b2a2c40 ffff880129c99800 nfs4 192.168.200.253:/exports/state/oleg422-server.virtnet /boot ffff88012b2a2e00 ffff880129c99800 nfs4 192.168.200.253:/exports/state/oleg422-server.virtnet /etc/etc/kdump.conf ffff88012b2a2fc0 ffff88012a356000 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800ad845880 ffff8800b526d000 squashfs /dev/vda /usr/sbin/mount.lustre ffff880129e43500 ffff880131066000 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 ffff880129e43180 ffff8800921b1000 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff8800b422f500 ffff8800920b6800 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff8800b422ec40 ffff88009ad4f000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 2532.062270] LDISKFS-fs (dm-0): recovery complete [ 2532.064034] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 2532.096270] LustreError: 166-1: MGC192.168.204.122@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 2532.099245] LustreError: Skipped 14 previous similar messages [ 2533.176828] Lustre: DEBUG MARKER: oleg422-server.virtnet: executing set_default_debug -1 all [ 2533.574639] LustreError: 7120:0:(ldlm_lockd.c:261:expired_lock_main()) ### lock callback timer expired after 101s: evicting client at 192.168.204.22@tcp ns: filter-lustre-OST0000_UUID lock: ffff88009afc7cc0/0x479edc5800ec3e91 lrc: 3/0,0 mode: PW/PW res: [0x280000401:0x863:0x0].0x0 rrc: 3 type: EXT [0->18446744073709551615] (req 0->18446744073709551615) gid 0 flags: 0x60000400000020 nid: 192.168.204.22@tcp remote: 0x3fdc44f0e7c2bb50 expref: 454 pid: 18911 timeout: 2532 lvb_type: 0 [ 2539.905531] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2611 to 0x2c0000401:2657) [ 2539.905556] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2643 to 0x280000401:2689) [ 2540.734829] Lustre: DEBUG MARKER: oleg422-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 2541.273603] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 2547.479923] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 2550.053686] Lustre: DEBUG MARKER: test_26 fail mds2 4 times [ 2550.724317] Lustre: Failing over lustre-MDT0001 [ 2550.813312] Lustre: server umount lustre-MDT0001 complete [ 2564.945544] LDISKFS-fs (dm-1): recovery complete [ 2564.948255] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,acl,no_mbcache,nodelalloc [ 2564.967332] Lustre: lustre-OST0000: haven't heard from client a6acc7d7-a7c2-4e0d-9613-98b785fa9b34 (at 192.168.204.22@tcp) in 31 seconds. I think it's dead, and I am evicting it. exp ffff88007c012000, cur 1715208212 expire 1715208182 last 1715208181 [ 2564.986702] Lustre: Skipped 2 previous similar messages [ 2566.222446] Lustre: DEBUG MARKER: oleg422-server.virtnet: executing set_default_debug -1 all [ 2570.601334] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:353 to 0x280000400:385) [ 2570.601346] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:354 to 0x2c0000400:385) [ 2571.467629] Lustre: DEBUG MARKER: oleg422-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid [ 2572.007256] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 2573.355092] Lustre: DEBUG MARKER: replay-dual test_26: @@@@@@ FAIL: dbench failed with rc 1 [ 2578.047892] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 2580.605972] Lustre: DEBUG MARKER: test_26 fail mds1 5 times [ 2581.303054] Lustre: Failing over lustre-MDT0000 [ 2581.409808] Lustre: lustre-MDT0000: Not available for connect from 192.168.204.22@tcp (stopping) [ 2581.413456] Lustre: Skipped 14 previous similar messages [ 2585.095646] LustreError: 11-0: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 2585.100431] LustreError: Skipped 2 previous similar messages [ 2587.312401] Lustre: server umount lustre-MDT0000 complete [ 2601.535587] LDISKFS-fs (dm-0): recovery complete [ 2601.537923] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 2602.787197] Lustre: DEBUG MARKER: oleg422-server.virtnet: executing set_default_debug -1 all [ 2607.587822] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2890 to 0x280000401:2913) [ 2607.588844] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2858 to 0x2c0000401:2881) [ 2608.461930] Lustre: DEBUG MARKER: oleg422-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 2609.057284] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 2610.918577] Lustre: DEBUG MARKER: replay-dual test_26: @@@@@@ FAIL: dbench 6353 missing ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.24s (real) 6.56s (CPU), Child processes: 5.67s