************************ crashinfo ************************* /exports/testreports/49254/testresults/sanity-hsm-zfs-centos7_x86_64-centos7_x86_64/oleg405-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.EUaVI/vmlinux [TAINTED] DUMPFILE: /exports/testreports/49254/testresults/sanity-hsm-zfs-centos7_x86_64-centos7_x86_64/oleg405-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Fri Feb 14 08:03:01 EST 2025 UPTIME: 01:23:48 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 356 NODENAME: oleg405-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2400 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 54 last 5s 68 last 60s 77 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 352 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 7148 ldlm_bl_02 0 ms (no user stack) 14136 kworker/1:0 0 ms (no user stack) 20 rcuos/1 0 ms (no user stack) 9 rcu_sched 0 ms (no user stack) 21 watchdog/2 5 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 635211 2.4 GB 66% of TOTAL MEM USED 319856 1.2 GB 33% of TOTAL MEM SHARED 10373 40.5 MB 1% of TOTAL MEM BUFFERS 5165 20.2 MB 0% of TOTAL MEM CACHED 71723 280.2 MB 7% of TOTAL MEM SLAB 22318 87.2 MB 2% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 60000 234.4 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 5025.2 eth0 n/a 2.2 RSS_TOTAL=56128 pages, %mem= 0.9 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880137012fc0 ffff8800b525a000 sysfs sysfs /sys ffff880137013180 ffff880139944000 proc proc /proc ffff880137013340 ffff880137678000 devtmpfs devtmpfs /dev ffff880137013500 ffff8800b51fc800 securityfs securityfs /sys/kernel/security ffff8801370136c0 ffff8800b525a800 tmpfs tmpfs /dev/shm ffff880137013880 ffff880137143000 devpts devpts /dev/pts ffff880137013a40 ffff8800b525b000 tmpfs tmpfs /run ffff880137013c00 ffff8800b525b800 tmpfs tmpfs /sys/fs/cgroup ffff880137013dc0 ffff8800b525c000 cgroup cgroup /sys/fs/cgroup/systemd ffff88012aa5e000 ffff8800b525c800 pstore pstore /sys/fs/pstore ffff88012aa5e1c0 ffff8800b525e800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012aa5e380 ffff8800b525e000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012aa5e540 ffff8800b525d800 cgroup cgroup /sys/fs/cgroup/memory ffff88012aa5e700 ffff8800b525d000 cgroup cgroup /sys/fs/cgroup/devices ffff88012aa5e8c0 ffff8800b525f000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012aa5ea80 ffff8800b525f800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012aa5ec40 ffff88012a378000 cgroup cgroup /sys/fs/cgroup/pids ffff88012aa5ee00 ffff88012a378800 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012aa5efc0 ffff88012a379000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012aa5f180 ffff88012a379800 cgroup cgroup /sys/fs/cgroup/perf_event ffff880137668380 ffff88013767f800 configfs configfs /sys/kernel/config ffff880137669880 ffff88013767c800 ext4 /dev/nbd0 / ffff8800b53ec1c0 ffff88012b365800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff880138ccac40 ffff880129dc1000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880137669c00 ffff8800b49cf000 hugetlbfs hugetlbfs /dev/hugepages ffff880137669dc0 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff880137012e00 ffff880137144800 mqueue mqueue /dev/mqueue ffff880138ccae00 ffff880129dc2000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff8800b53461c0 ffff8800b49cb000 ramfs none /mnt ffff8800b53ec000 ffff8800b4068000 tmpfs none /var/lib/stateless/writable ffff88012aa5fa40 ffff8800b48e0800 squashfs /dev/vda /home/green/git/lustre-release ffff8800b53ec380 ffff8800b4068000 tmpfs none /var/cache/man ffff8800b53ec540 ffff8800b4068000 tmpfs none /var/log ffff880138ccafc0 ffff8800b4068000 tmpfs none /var/lib/dbus ffff88012aa5fc00 ffff8800b4068000 tmpfs none /tmp ffff8800b5346540 ffff8800b4068000 tmpfs none /var/lib/dhclient ffff880138ccb180 ffff8800b4068000 tmpfs none /var/tmp ffff8800b53ec700 ffff8800b4068000 tmpfs none /var/lib/NetworkManager ffff8800b5346700 ffff8800b4068000 tmpfs none /var/lib/systemd/random-seed ffff880138ccb340 ffff8800b4068000 tmpfs none /var/spool ffff88012aa5fdc0 ffff8800b4068000 tmpfs none /var/lib/nfs ffff8800b53468c0 ffff8800b4068000 tmpfs none /var/lib/gssproxy ffff8800b53ec8c0 ffff8800b4068000 tmpfs none /var/lib/logrotate ffff8800b5346a80 ffff8800b4068000 tmpfs none /etc ffff88012aa5f880 ffff8800b4068000 tmpfs none /var/lib/rsyslog ffff8800b53eca80 ffff8800b4068000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff88012aa5f6c0 ffff8800b409a000 nfs4 192.168.200.253:/exports/state/oleg405-server.virtnet /var/lib/stateless/state ffff8800b42b41c0 ffff8800b409a000 nfs4 192.168.200.253:/exports/state/oleg405-server.virtnet /boot ffff8800b42b4380 ffff8800b409a000 nfs4 192.168.200.253:/exports/state/oleg405-server.virtnet /etc/etc/kdump.conf ffff8800b53ece00 ffff88012b365800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b25828c0 ffff8800b48e0800 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800b42b5500 ffff88012dc4c800 lustre lustre-ost1/ost1 /mnt/lustre-ost1 ffff8800b42b4000 ffff880133c08800 lustre lustre-ost2/ost2 /mnt/lustre-ost2 ffff8800b42b4fc0 ffff88012f9b3800 lustre lustre-mdt1/mdt1 /mnt/lustre-mds1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 1909.799435] Lustre: Skipped 1 previous similar message [ 1910.794623] Lustre: 3313:0:(client.c:2346:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1739535047/real 1739535047] req@ffff8801345ab100 x1824032770188672/t0(0) o400->lustre-MDT0000-lwp-OST0001@0@lo:12/10 lens 224/224 e 0 to 1 dl 1739535063 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 1910.804480] Lustre: 3313:0:(client.c:2346:ptlrpc_expire_one_request()) Skipped 3 previous similar messages [ 1911.566148] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 1911.588922] Lustre: lustre-MDT0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 1911.606821] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:707 to 0x240000400:737) [ 1911.606828] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:707 to 0x280000400:737) [ 1912.537763] Lustre: DEBUG MARKER: oleg405-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1912.992209] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1920.601897] Lustre: DEBUG MARKER: == sanity-hsm test 408: Verify fiemap on release file ==== 07:11:14 (1739535074) [ 1921.175704] Lustre: DEBUG MARKER: SKIP: sanity-hsm test_408 ORI-366/LU-1941: FIEMAP unimplemented on ZFS [ 1921.750108] Lustre: DEBUG MARKER: == sanity-hsm test 409a: Coordinator should not stop when in use ========================================================== 07:11:15 (1739535075) [ 1923.742200] LustreError: 13483:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 sleeping for 5000ms [ 1928.745572] LustreError: 13483:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 awake [ 1934.559464] Lustre: DEBUG MARKER: == sanity-hsm test 409b: getattr released file with CDT stopped after remount ========================================================== 07:11:28 (1739535088) [ 1935.445588] Lustre: Modifying parameter lustre.mdt.lustre-MDT0000.hsm_control in log params [ 1935.449337] Lustre: Skipped 2 previous similar messages [ 1956.256294] Lustre: Failing over lustre-MDT0000 [ 1956.378560] Lustre: server umount lustre-MDT0000 complete [ 1968.499303] LustreError: MGC192.168.204.105@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 1968.608095] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1968.615672] Lustre: Skipped 1 previous similar message [ 1968.617831] LustreError: 3315:0:(import.c:709:ptlrpc_connect_import_locked()) already connecting [ 1968.651867] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 1968.676056] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 1969.581729] Lustre: DEBUG MARKER: oleg405-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1971.405299] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 1971.673538] Lustre: lustre-MDT0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 1971.691938] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:740 to 0x280000400:769) [ 1971.692010] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:741 to 0x240000400:769) [ 1972.677027] Lustre: DEBUG MARKER: oleg405-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1973.169799] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1973.635875] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to (at 0@lo) [ 1973.638615] Lustre: Skipped 1 previous similar message [ 1975.618583] Lustre: 3315:0:(client.c:2346:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1739535112/real 1739535112] req@ffff880084d02a00 x1824032770215936/t0(0) o400->lustre-MDT0000-lwp-OST0001@0@lo:12/10 lens 224/224 e 0 to 1 dl 1739535128 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 1975.635009] Lustre: 3315:0:(client.c:2346:ptlrpc_expire_one_request()) Skipped 6 previous similar messages [ 1997.095177] Lustre: DEBUG MARKER: == sanity-hsm test 410: lfs data_version -s allows release of force-archived file ========================================================== 07:12:30 (1739535150) [ 1999.673944] Lustre: DEBUG MARKER: == sanity-hsm test 411: hsm_ops rbac role ================ 07:12:33 (1739535153) [ 2054.584610] Lustre: DEBUG MARKER: == sanity-hsm test 500: various LLAPI HSM tests ========== 07:13:28 (1739535208) [ 2070.080927] Lustre: HSM agent fca61502-7ab5-4909-b47d-bbd1b0a1a25d already registered ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 12.27s (real) 6.83s (CPU), Child processes: 5.39s