************************ crashinfo ************************* /exports/testreports/49655/testresults/sanity-hsm-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg428-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.Rj66w/vmlinux [TAINTED] DUMPFILE: /exports/testreports/49655/testresults/sanity-hsm-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg428-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Thu Feb 27 12:43:03 EST 2025 UPTIME: 01:23:39 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 271 NODENAME: oleg428-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2400 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 18 last 5s 62 last 60s 77 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 267 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 7195 ldlm_bl_02 0 ms (no user stack) 16 ksoftirqd/1 0 ms (no user stack) 31237 kworker/1:0 0 ms (no user stack) 3757 lnet_discovery 0 ms (no user stack) 21242 hsm_cdtr 47 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 646991 2.5 GB 67% of TOTAL MEM USED 308076 1.2 GB 32% of TOTAL MEM SHARED 11731 45.8 MB 1% of TOTAL MEM BUFFERS 8920 34.8 MB 0% of TOTAL MEM CACHED 157390 614.8 MB 16% of TOTAL MEM SLAB 17957 70.1 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 61028 238.4 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 5016.3 eth0 n/a 1.9 RSS_TOTAL=52308 pages, %mem= 0.8 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff88012a3a0000 ffff88012a3a8000 sysfs sysfs /sys ffff88012a3a01c0 ffff880139944000 proc proc /proc ffff88012a3a0380 ffff880137678000 devtmpfs devtmpfs /dev ffff88012a3a0540 ffff8800b51bf800 securityfs securityfs /sys/kernel/security ffff88012a3a0700 ffff88012a3a8800 tmpfs tmpfs /dev/shm ffff88012a3a08c0 ffff88013771f000 devpts devpts /dev/pts ffff88012a3a0a80 ffff88012a3a9000 tmpfs tmpfs /run ffff88012a3a0c40 ffff88012a3a9800 tmpfs tmpfs /sys/fs/cgroup ffff88012a3a0e00 ffff88012a3aa000 cgroup cgroup /sys/fs/cgroup/systemd ffff88012a3a0fc0 ffff88012a3aa800 pstore pstore /sys/fs/pstore ffff88012a3a1180 ffff88012a3ac800 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a3a1340 ffff88012a3ac000 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012a3a1500 ffff88012a3ab800 cgroup cgroup /sys/fs/cgroup/pids ffff88012a3a16c0 ffff88012a3ab000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a3a1880 ffff88012a3ad000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012a3a1a40 ffff88012a3ad800 cgroup cgroup /sys/fs/cgroup/devices ffff88012a3a1c00 ffff88012a3ae000 cgroup cgroup /sys/fs/cgroup/memory ffff88012a3a1dc0 ffff88012a3ae800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a3a2000 ffff88012a3af000 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a3a21c0 ffff88012a3af800 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a3a2540 ffff8800b5225800 configfs configfs /sys/kernel/config ffff880137668a80 ffff8800b4064800 ext4 /dev/nbd0 / ffff88012a3a36c0 ffff8800b5226800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff8800b526c540 ffff8800b40cc800 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff88012a3a3880 ffff8800b4895800 hugetlbfs hugetlbfs /dev/hugepages ffff880137668c40 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff880137668e00 ffff88012b2a0800 mqueue mqueue /dev/mqueue ffff8800b526c700 ffff8800b40ce000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff88012a3a3c00 ffff8800b5222000 ramfs none /mnt ffff880138ccba40 ffff8800b497f000 squashfs /dev/vda /home/green/git/lustre-release ffff88012a3a2380 ffff8800b4896000 tmpfs none /var/lib/stateless/writable ffff880137668fc0 ffff8800b4896000 tmpfs none /var/cache/man ffff880137669180 ffff8800b4896000 tmpfs none /var/log ffff880137669340 ffff8800b4896000 tmpfs none /var/lib/dbus ffff880137669500 ffff8800b4896000 tmpfs none /tmp ffff8801376696c0 ffff8800b4896000 tmpfs none /var/lib/dhclient ffff880130f0a000 ffff8800b4896000 tmpfs none /var/tmp ffff880130f0a1c0 ffff8800b4896000 tmpfs none /var/lib/NetworkManager ffff880130f0a380 ffff8800b4896000 tmpfs none /var/lib/systemd/random-seed ffff880130f0a540 ffff8800b4896000 tmpfs none /var/spool ffff880130f0a700 ffff8800b4896000 tmpfs none /var/lib/nfs ffff880137669880 ffff8800b4896000 tmpfs none /var/lib/gssproxy ffff8800b526c8c0 ffff8800b4896000 tmpfs none /var/lib/logrotate ffff880130f0a8c0 ffff8800b4896000 tmpfs none /etc ffff880130f0aa80 ffff8800b4896000 tmpfs none /var/lib/rsyslog ffff880130f0ac40 ffff8800b4896000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff880130f0ae00 ffff88012a3e2800 nfs4 192.168.200.253:/exports/state/oleg428-server.virtnet /var/lib/stateless/state ffff880130f0b500 ffff88012a3e2800 nfs4 192.168.200.253:/exports/state/oleg428-server.virtnet /boot ffff8800b526ca80 ffff88012a3e2800 nfs4 192.168.200.253:/exports/state/oleg428-server.virtnet /etc/etc/kdump.conf ffff880137669a40 ffff8800b5226800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b1945dc0 ffff8800b497f000 squashfs /dev/vda /usr/sbin/mount.lustre ffff88012d013a40 ffff8800a86a3000 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff88012d0128c0 ffff88009f036000 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 ffff8800b526d340 ffff8800b4063800 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff8800b526da40 ffff880093b03000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 1742.196151] LustreError: MGC192.168.204.128@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 1742.299842] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 1742.322364] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 1743.382601] Lustre: DEBUG MARKER: oleg428-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1745.911542] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 1747.309243] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 192.168.204.128@tcp (at 0@lo) [ 1747.313031] Lustre: Skipped 2 previous similar messages [ 1747.324331] Lustre: lustre-MDT0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 1747.346612] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:487 to 0x2c0000401:513) [ 1747.346621] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:486 to 0x280000401:513) [ 1748.452561] Lustre: DEBUG MARKER: oleg428-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1749.050890] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1757.617897] Lustre: DEBUG MARKER: == sanity-hsm test 408: Verify fiemap on release file ==== 11:48:41 (1740674921) [ 1761.922621] Lustre: DEBUG MARKER: == sanity-hsm test 409a: Coordinator should not stop when in use ========================================================== 11:48:45 (1740674925) [ 1762.808437] LustreError: 7206:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 sleeping for 5000ms [ 1767.816737] LustreError: 7206:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 awake [ 1775.997414] Lustre: DEBUG MARKER: == sanity-hsm test 409b: getattr released file with CDT stopped after remount ========================================================== 11:48:59 (1740674939) [ 1778.261008] Lustre: Modifying parameter lustre.mdt.lustre-MDT0000.hsm_control in log params [ 1778.264164] Lustre: Skipped 1 previous similar message [ 1799.792667] Lustre: Failing over lustre-MDT0000 [ 1799.880621] Lustre: server umount lustre-MDT0000 complete [ 1802.396154] Lustre: lustre-MDT0000-lwp-OST0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1802.401303] Lustre: Skipped 2 previous similar messages [ 1812.384450] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 1812.431872] LustreError: MGC192.168.204.128@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 1812.501289] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 1812.512296] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 1813.349377] Lustre: DEBUG MARKER: oleg428-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1816.023898] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 1817.517463] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to (at 0@lo) [ 1817.521085] Lustre: Skipped 3 previous similar messages [ 1817.529355] Lustre: lustre-MDT0000: Recovery over after 0:02, of 3 clients 3 recovered and 0 were evicted. [ 1817.547024] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:517 to 0x280000401:545) [ 1817.547560] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:516 to 0x2c0000401:545) [ 1818.615837] Lustre: DEBUG MARKER: oleg428-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1819.097306] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1843.331485] Lustre: DEBUG MARKER: == sanity-hsm test 410: lfs data_version -s allows release of force-archived file ========================================================== 11:50:07 (1740675007) [ 1845.344489] Lustre: DEBUG MARKER: == sanity-hsm test 411: hsm_ops rbac role ================ 11:50:09 (1740675009) [ 1897.610191] Lustre: DEBUG MARKER: == sanity-hsm test 500: various LLAPI HSM tests ========== 11:51:01 (1740675061) [ 1919.056822] Lustre: HSM agent 56e79868-4d1c-4acb-b305-bc69d3211e05 already registered ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 11.16s (real) 6.03s (CPU), Child processes: 5.10s