************************ crashinfo ************************* /exports/testreports/49254/testresults/sanity-quota-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg116-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.DHKBK/vmlinux [TAINTED] DUMPFILE: /exports/testreports/49254/testresults/sanity-quota-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg116-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Fri Feb 14 08:03:14 EST 2025 UPTIME: 01:20:16 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 268 NODENAME: oleg116-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 31 last 5s 73 last 60s 88 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 264 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 3758 lnet_discovery 0 ms (no user stack) 28 watchdog/3 0 ms (no user stack) 13 watchdog/0 0 ms (no user stack) 14 watchdog/1 0 ms (no user stack) 21 watchdog/2 17 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 729955 2.8 GB 76% of TOTAL MEM USED 225112 879.3 MB 23% of TOTAL MEM SHARED 13638 53.3 MB 1% of TOTAL MEM BUFFERS 10587 41.4 MB 1% of TOTAL MEM CACHED 73618 287.6 MB 7% of TOTAL MEM SLAB 16204 63.3 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 62811 245.4 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 10 LISTEN 3 NAGLE disabled (TCP_NODELAY): 8 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 18 LISTEN 8 Raw sockets info -------------------- CLOSE 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 4813.7 eth0 n/a 2.2 RSS_TOTAL=69000 pages, %mem= 1.1 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137668380 ffff880137679000 sysfs sysfs /sys ffff880137668540 ffff880139944000 proc proc /proc ffff880137668700 ffff880137678000 devtmpfs devtmpfs /dev ffff8801376688c0 ffff880137315000 securityfs securityfs /sys/kernel/security ffff880137668a80 ffff880137679800 tmpfs tmpfs /dev/shm ffff880137668c40 ffff880137102800 devpts devpts /dev/pts ffff880137668e00 ffff88013767a000 tmpfs tmpfs /run ffff880137668fc0 ffff88013767a800 tmpfs tmpfs /sys/fs/cgroup ffff880137669180 ffff88013767b000 cgroup cgroup /sys/fs/cgroup/systemd ffff880137669340 ffff88013767b800 pstore pstore /sys/fs/pstore ffff880138ccba40 ffff88012a330000 cgroup cgroup /sys/fs/cgroup/blkio ffff880138ccbc00 ffff88012a330800 cgroup cgroup /sys/fs/cgroup/cpuset ffff880138ccbdc0 ffff88012a331000 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a2b2000 ffff88012a331800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012a2b21c0 ffff88012a332000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a2b2380 ffff88012a332800 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a2b2540 ffff88012a333000 cgroup cgroup /sys/fs/cgroup/pids ffff88012a2b2700 ffff88012a333800 cgroup cgroup /sys/fs/cgroup/devices ffff88012a2b28c0 ffff88012a334000 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012a2b2a80 ffff88012a334800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff8800b4c921c0 ffff8800b51fc800 configfs configfs /sys/kernel/config ffff8800b4c92380 ffff8800b51fd800 ext4 /dev/nbd0 / ffff880137669500 ffff8800b51fd000 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff8800b4c92540 ffff880129c89000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff8800b4c92700 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff88012a2b3dc0 ffff8800b4171800 hugetlbfs hugetlbfs /dev/hugepages ffff8800b4c928c0 ffff880137104000 mqueue mqueue /dev/mqueue ffff88012a2f2a80 ffff8800b427c000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff8800b4c92a80 ffff8800b41c3800 ramfs none /mnt ffff8801376696c0 ffff880137315800 tmpfs none /var/lib/stateless/writable ffff880137669880 ffff880137317000 squashfs /dev/vda /home/green/git/lustre-release ffff8800b4c92c40 ffff880137315800 tmpfs none /var/cache/man ffff880137669a40 ffff880137315800 tmpfs none /var/log ffff8800b4c92e00 ffff880137315800 tmpfs none /var/lib/dbus ffff8800b4c92fc0 ffff880137315800 tmpfs none /tmp ffff8800b4c93180 ffff880137315800 tmpfs none /var/lib/dhclient ffff8800b4c93340 ffff880137315800 tmpfs none /var/tmp ffff8800b4c93500 ffff880137315800 tmpfs none /var/lib/NetworkManager ffff8800b4c936c0 ffff880137315800 tmpfs none /var/lib/systemd/random-seed ffff8800b4c93880 ffff880137315800 tmpfs none /var/spool ffff8800b4c93a40 ffff880137315800 tmpfs none /var/lib/nfs ffff880137669c00 ffff880137315800 tmpfs none /var/lib/gssproxy ffff8800b4c93c00 ffff880137315800 tmpfs none /var/lib/logrotate ffff8800b4c93dc0 ffff880137315800 tmpfs none /etc ffff8800b4c92000 ffff880137315800 tmpfs none /var/lib/rsyslog ffff8800b1880000 ffff880137315800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff8800b18801c0 ffff8800b4174000 nfs4 192.168.200.253:/exports/state/oleg116-server.virtnet /var/lib/stateless/state ffff8800b18808c0 ffff8800b4174000 nfs4 192.168.200.253:/exports/state/oleg116-server.virtnet /boot ffff88012a2f2e00 ffff8800b4174000 nfs4 192.168.200.253:/exports/state/oleg116-server.virtnet /etc/etc/kdump.conf ffff8800b1880a80 ffff8800b51fd000 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8801373a8fc0 ffff880137317000 squashfs /dev/vda /usr/sbin/mount.lustre ffff88012c46e000 ffff88013767c000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 ffff88012c46e700 ffff880129c8f000 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff88012c46f880 ffff88009d01a800 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 ffff8801373a9a40 ffff8800b1b9c000 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff8800b4302700 ffff88009d018800 tmpfs tmpfs /run/user/0 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 1155.620248] Lustre: Skipped 10 previous similar messages [ 1158.561369] Lustre: 19627:0:(service.c:1438:ptlrpc_at_send_early_reply()) @@@ Could not add any time (5/5), not sending early reply req@ffff88009c669500 x1824032976637568/t0(0) o4->dce837a3-00e1-45e7-aebe-d67697a189b4@192.168.201.16@tcp:196/0 lens 488/448 e 1 to 0 dl 1739534541 ref 2 fl Interpret:/200/0 rc 0/0 job:'dd.60000' uid:60000 gid:60000 [ 1159.548665] Lustre: 22862:0:(client.c:2346:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1739534521/real 1739534521] req@ffff88013175c380 x1824032982088320/t0(0) o601->lustre-MDT0000-lwp-OST0001@0@lo:23/10 lens 336/336 e 0 to 1 dl 1739534537 ref 2 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'ll_ost_io00_003.0' uid:0 gid:0 [ 1159.566709] Lustre: lustre-MDT0000-lwp-OST0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1159.576700] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0001_UUID (at 0@lo) reconnecting [ 1159.586957] Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 192.168.201.116@tcp (at 0@lo) [ 1159.646571] LustreError: 16439:0:(service.c:2122:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1824032982097280 [ 1164.577948] Lustre: *** cfs_fail_loc=513, val=601*** [ 1164.582984] Lustre: Skipped 49 previous similar messages [ 1175.598424] Lustre: 3762:0:(client.c:2346:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1739534537/real 1739534537] req@ffff88009cf32a00 x1824032982097280/t0(0) o601->lustre-MDT0000-lwp-OST0001@0@lo:23/10 lens 336/336 e 0 to 1 dl 1739534553 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'ll_ost_io00_003.0' uid:0 gid:0 [ 1175.617702] Lustre: lustre-MDT0000-lwp-OST0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1175.625353] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0001_UUID (at 0@lo) reconnecting [ 1175.637303] Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 192.168.201.116@tcp (at 0@lo) [ 1178.241069] LustreError: 16439:0:(service.c:2122:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1824032982107392 [ 1180.641947] Lustre: *** cfs_fail_loc=513, val=601*** [ 1180.649818] Lustre: Skipped 87 previous similar messages [ 1194.240404] Lustre: 22862:0:(client.c:2346:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1739534555/real 1739534555] req@ffff88009a9e9180 x1824032982107392/t0(0) o601->lustre-MDT0000-lwp-OST0001@0@lo:23/10 lens 336/336 e 0 to 1 dl 1739534571 ref 2 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'ll_ost_io00_003.0' uid:0 gid:0 [ 1194.251424] Lustre: lustre-MDT0000-lwp-OST0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1194.259558] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0001_UUID (at 0@lo) reconnecting [ 1194.264090] Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 192.168.201.116@tcp (at 0@lo) [ 1212.578403] Lustre: DEBUG MARKER: == sanity-quota test 7a: Quota reintegration (global index) ========================================================== 07:03:09 (1739534589) [ 1220.241123] Lustre: Failing over lustre-OST0000 [ 1220.278650] Lustre: server umount lustre-OST0000 complete [ 1221.361667] LustreError: lustre-OST0000-osc-MDT0001: operation ost_statfs to node 0@lo failed: rc = -107 [ 1221.361708] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1221.370091] LustreError: Skipped 1 previous similar message [ 1222.571148] LustreError: 2098:0:(qsd_reint.c:617:qqi_reint_delayed()) lustre-MDT0000: Delaying reintegration for qtype:0 until pending updates are flushed. [ 1222.573882] LustreError: 2098:0:(qsd_reint.c:617:qqi_reint_delayed()) Skipped 11 previous similar messages [ 1224.839929] LDISKFS-fs (dm-2): file extents enabled, maximum tree depth=5 [ 1224.843176] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: user_xattr,acl,no_mbcache,nodelalloc [ 1224.865121] Lustre: 2604:0:(mgc_request_server.c:553:mgc_llog_local_copy()) MGC192.168.201.116@tcp: no remote llog for lustre-sptlrpc, check MGS config [ 1224.912589] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 1224.919980] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 1226.001043] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 1226.176381] Lustre: DEBUG MARKER: oleg116-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1226.905101] Lustre: lustre-OST0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 1226.905161] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.201.116@tcp (at 0@lo) [ 1228.208110] Lustre: DEBUG MARKER: oleg116-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 2710.270718] Lustre: DEBUG MARKER: oleg116-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 4191.272138] Lustre: DEBUG MARKER: oleg116-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.22s (real) 6.55s (CPU), Child processes: 5.65s