************************ crashinfo ************************* /exports/testreports/42551/testresults/sanity-scrub-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg351-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.IScUw/vmlinux [TAINTED] DUMPFILE: /exports/testreports/42551/testresults/sanity-scrub-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg351-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Wed May 8 19:07:51 EDT 2024 UPTIME: 01:07:03 LOAD AVERAGE: 0.01, 0.02, 0.05 TASKS: 259 NODENAME: oleg351-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 14 last 5s 65 last 60s 78 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 255 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 9 rcu_sched 0 ms (no user stack) 20 rcuos/1 0 ms (no user stack) 1 systemd 0 ms /usr/lib/systemd/systemd --switched-root --system --deserialize 21 21786 ldlm_bl_01 54 ms (no user stack) 49 kworker/0:1 57 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 739449 2.8 GB 77% of TOTAL MEM USED 215618 842.3 MB 22% of TOTAL MEM SHARED 10940 42.7 MB 1% of TOTAL MEM BUFFERS 9122 35.6 MB 0% of TOTAL MEM CACHED 67848 265 MB 7% of TOTAL MEM SLAB 16728 65.3 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 59986 234.3 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 8 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 4020.4 eth0 n/a 1.3 RSS_TOTAL=49228 pages, %mem= 0.8 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137668380 ffff880137678800 sysfs sysfs /sys ffff880137668540 ffff880139944000 proc proc /proc ffff880137668700 ffff880137678000 devtmpfs devtmpfs /dev ffff8801376688c0 ffff8800b51ee000 securityfs securityfs /sys/kernel/security ffff880137668a80 ffff880137679000 tmpfs tmpfs /dev/shm ffff880137668c40 ffff88012b270000 devpts devpts /dev/pts ffff880137668e00 ffff880137679800 tmpfs tmpfs /run ffff880137668fc0 ffff88013767a000 tmpfs tmpfs /sys/fs/cgroup ffff880137669180 ffff88013767a800 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669340 ffff88013767b000 pstore pstore /sys/fs/pstore ffff880137669500 ffff88013767d000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff8801376696c0 ffff88013767c800 cgroup cgroup /sys/fs/cgroup/memory ffff880137669880 ffff88013767c000 cgroup cgroup /sys/fs/cgroup/pids ffff880137669a40 ffff88013767b800 cgroup cgroup /sys/fs/cgroup/perf_event ffff880137669c00 ffff88013767d800 cgroup cgroup /sys/fs/cgroup/cpuset ffff880137669dc0 ffff88013767e000 cgroup cgroup /sys/fs/cgroup/devices ffff88012a39e000 ffff88013767e800 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a39e1c0 ffff88013767f000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a39e380 ffff88013767f800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012a39e540 ffff88012a3b0000 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a39e8c0 ffff88012a3b7000 configfs configfs /sys/kernel/config ffff88012a39ea80 ffff88012a3b5000 ext4 /dev/nbd0 / ffff8800b4a9a000 ffff8800b48be000 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff8801377f7500 ffff88012a3b2000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff88012a39f180 ffff8800b49d5800 hugetlbfs hugetlbfs /dev/hugepages ffff8800b4a9a1c0 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff8800b4a9a380 ffff88012b271800 mqueue mqueue /dev/mqueue ffff8801377f76c0 ffff880129c33000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff880138ccbc00 ffff8800b48bd000 ramfs none /mnt ffff880138ccbdc0 ffff8800b48bc000 squashfs /dev/vda /home/green/git/lustre-release ffff88012a39f340 ffff8800b49d7000 tmpfs none /var/lib/stateless/writable ffff8800b4a9a540 ffff8800b49d7000 tmpfs none /var/cache/man ffff8801377f7a40 ffff8800b49d7000 tmpfs none /var/log ffff8800b1ea81c0 ffff8800b49d7000 tmpfs none /var/lib/dbus ffff88012a39f500 ffff8800b49d7000 tmpfs none /tmp ffff8800b1ea8380 ffff8800b49d7000 tmpfs none /var/lib/dhclient ffff88012a39f6c0 ffff8800b49d7000 tmpfs none /var/tmp ffff8800b1ea8540 ffff8800b49d7000 tmpfs none /var/lib/NetworkManager ffff88012a39f880 ffff8800b49d7000 tmpfs none /var/lib/systemd/random-seed ffff88012a39fa40 ffff8800b49d7000 tmpfs none /var/spool ffff8801377f7c00 ffff8800b49d7000 tmpfs none /var/lib/nfs ffff8801377f7dc0 ffff8800b49d7000 tmpfs none /var/lib/gssproxy ffff88012a39fc00 ffff8800b49d7000 tmpfs none /var/lib/logrotate ffff88012a39fdc0 ffff8800b49d7000 tmpfs none /etc ffff8800b4a9a700 ffff8800b49d7000 tmpfs none /var/lib/rsyslog ffff8800b4a9a8c0 ffff8800b49d7000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff8801377f6e00 ffff880129ea1000 nfs4 192.168.200.253:/exports/state/oleg351-server.virtnet /var/lib/stateless/state ffff88012a39efc0 ffff880129ea1000 nfs4 192.168.200.253:/exports/state/oleg351-server.virtnet /boot ffff8800b4a9ae00 ffff880129ea1000 nfs4 192.168.200.253:/exports/state/oleg351-server.virtnet /etc/etc/kdump.conf ffff88012a39ee00 ffff8800b48be000 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b1e19180 ffff8800b48bc000 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800b1ea8000 ffff8800b1c5b000 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff8800b1ea9dc0 ffff88009abee000 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 ffff880129c25500 ffff88009a486000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 ffff880129c25880 ffff880099941000 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 691.317690] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: (null) [ 693.565393] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: (null) [ 696.086161] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: errors=remount-ro [ 696.330096] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: (null) [ 699.241149] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,user_xattr,no_mbcache,nodelalloc [ 699.245594] Lustre: lustre-MDT0000: reset Object Index mappings [ 706.193978] Lustre: 18382:0:(client.c:2343:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1715206338/real 1715206338] req@ffff88009dd48e00 x1798523632259456/t0(0) o400->lustre-MDT0000-lwp-OST0000@0@lo:12/10 lens 224/224 e 0 to 1 dl 1715206354 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 706.194017] Lustre: lustre-MDT0001-lwp-OST0000: Connection to lustre-MDT0001 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 706.194019] Lustre: Skipped 8 previous similar messages [ 706.205819] Lustre: 18382:0:(client.c:2343:ptlrpc_expire_one_request()) Skipped 53 previous similar messages [ 712.208174] LustreError: 18380:0:(client.c:1291:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff880082b5d180 x1798523632260864/t0(0) o250->MGC192.168.203.151@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 712.213672] LustreError: 18380:0:(client.c:1291:ptlrpc_import_delay_req()) Skipped 3 previous similar messages [ 712.302410] LustreError: 137-5: lustre-MDT0001: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 712.306770] LustreError: Skipped 4 previous similar messages [ 712.323840] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 713.007559] Lustre: DEBUG MARKER: oleg351-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 715.142656] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,acl,no_mbcache,nodelalloc [ 715.207942] LustreError: 11-0: lustre-MDT0000-osp-MDT0001: operation mds_connect to node 0@lo failed: rc = -114 [ 715.210372] LustreError: Skipped 1 previous similar message [ 715.233398] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:617 to 0x280000400:641) [ 715.233400] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:618 to 0x2c0000400:641) [ 715.937458] Lustre: DEBUG MARKER: oleg351-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 716.226391] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 1 client reconnects [ 716.227250] Lustre: lustre-MDT0001-lwp-OST0001: Connection restored to (at 0@lo) [ 716.227251] Lustre: Skipped 15 previous similar messages [ 716.233036] Lustre: lustre-MDT0000: Recovery over after 0:01, of 1 clients 1 recovered and 0 were evicted. [ 716.248899] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:649 to 0x2c0000401:673) [ 716.252035] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:650 to 0x280000401:673) [ 717.857003] Lustre: lustre-MDT0000: trigger partial OI scrub for RPC inconsistency, checking FID [0x200005222:0x1:0x0]/64002: rc = 0 [ 720.882669] Lustre: lustre-MDT0001: trigger OI scrub by RPC for [0x240004a51:0x1:0x0]/32002 with flags 0x52: rc = 0 [ 781.316193] Lustre: DEBUG MARKER: == sanity-scrub test 4d: FID in LMA mismatch with object FID won't block create ========================================================== 18:13:48 (1715206428) [ 783.575300] Lustre: *** cfs_fail_loc=19b, val=0*** [ 783.576649] Lustre: Skipped 1 previous similar message [ 788.545964] Lustre: DEBUG MARKER: == sanity-scrub test 4e: FID reuse can be fixed ========== 18:13:56 (1715206436) [ 789.546302] Lustre: *** cfs_fail_loc=1a0, val=0*** [ 789.547816] Lustre: Skipped 455 previous similar messages [ 793.579266] Lustre: DEBUG MARKER: == sanity-scrub test 5: OI scrub state machine =========== 18:14:01 (1715206441) [ 825.394423] Lustre: lustre-OST0000: haven't heard from client b14e56f3-a207-4e86-a6bf-684b4bda64a7 (at 192.168.203.51@tcp) in 33 seconds. I think it's dead, and I am evicting it. exp ffff88012f67d800, cur 1715206473 expire 1715206443 last 1715206440 [ 825.404124] LustreError: 24060:0:(client.c:1281:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff88009f011180 x1798523632332736/t0(0) o104->lustre-OST0000@192.168.203.51@tcp:15/16 lens 328/224 e 0 to 0 dl 0 ref 1 fl Rpc:QU/0/ffffffff rc 0/-1 job:'' uid:4294967295 gid:4294967295 [ 826.402750] Lustre: lustre-MDT0001: haven't heard from client b14e56f3-a207-4e86-a6bf-684b4bda64a7 (at 192.168.203.51@tcp) in 34 seconds. I think it's dead, and I am evicting it. exp ffff880099944000, cur 1715206474 expire 1715206444 last 1715206440 ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 11.65s (real) 6.19s (CPU), Child processes: 5.43s