************************ crashinfo ************************* /exports/testreports/48990/testresults/sanity-hsm-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg412-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.y9kKc/vmlinux [TAINTED] DUMPFILE: /exports/testreports/48990/testresults/sanity-hsm-ldiskfs-DNE-centos7_x86_64-centos7_x86_64/oleg412-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Sun Feb 2 03:23:51 EST 2025 UPTIME: 01:23:37 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 277 NODENAME: oleg412-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2400 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 15 last 5s 65 last 60s 84 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 273 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 4514 l2arc_feed 0 ms (no user stack) 20 rcuos/1 0 ms (no user stack) 9 rcu_sched 0 ms (no user stack) 1 systemd 5 ms /usr/lib/systemd/systemd --switched-root --system --deserialize 22 3760 monitor_thread 120 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 645918 2.5 GB 67% of TOTAL MEM USED 309149 1.2 GB 32% of TOTAL MEM SHARED 11659 45.5 MB 1% of TOTAL MEM BUFFERS 8905 34.8 MB 0% of TOTAL MEM CACHED 157392 614.8 MB 16% of TOTAL MEM SLAB 18067 70.6 MB 1% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 61031 238.4 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 5015.1 eth0 n/a 1.9 RSS_TOTAL=54692 pages, %mem= 0.9 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff88012a310000 ffff880137300800 sysfs sysfs /sys ffff88012a3101c0 ffff880139944000 proc proc /proc ffff88012a310380 ffff880137678000 devtmpfs devtmpfs /dev ffff88012a310540 ffff880137678800 securityfs securityfs /sys/kernel/security ffff88012a310700 ffff880137301000 tmpfs tmpfs /dev/shm ffff88012a3108c0 ffff88013771f000 devpts devpts /dev/pts ffff88012a310a80 ffff880137301800 tmpfs tmpfs /run ffff88012a310c40 ffff880137302000 tmpfs tmpfs /sys/fs/cgroup ffff88012a310e00 ffff880137302800 cgroup cgroup /sys/fs/cgroup/systemd ffff88012a310fc0 ffff880137303000 pstore pstore /sys/fs/pstore ffff88012a311180 ffff880137305000 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a311340 ffff880137304800 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a311500 ffff880137304000 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012a3116c0 ffff880137303800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff88012a311880 ffff880137305800 cgroup cgroup /sys/fs/cgroup/devices ffff88012a311a40 ffff880137306000 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a311c00 ffff880137306800 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a311dc0 ffff880137307000 cgroup cgroup /sys/fs/cgroup/pids ffff88012a312000 ffff880137307800 cgroup cgroup /sys/fs/cgroup/memory ffff88012a3121c0 ffff88012a358000 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a312540 ffff88012a35e000 configfs configfs /sys/kernel/config ffff8801377f6700 ffff8800b5200800 ext4 /dev/nbd0 / ffff88012a3128c0 ffff88012a35c000 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff8801377f68c0 ffff88013767c000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff8801377f6a80 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff880137668540 ffff8800b51cb000 hugetlbfs hugetlbfs /dev/hugepages ffff88012a312e00 ffff88012b2b0800 mqueue mqueue /dev/mqueue ffff880137668700 ffff8800b40bc000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff8801377f6e00 ffff8800b4141800 ramfs none /mnt ffff88012a312fc0 ffff880129f09800 squashfs /dev/vda /home/green/git/lustre-release ffff8801376688c0 ffff88012a35c800 tmpfs none /var/lib/stateless/writable ffff880137668a80 ffff88012a35c800 tmpfs none /var/cache/man ffff880129dce8c0 ffff88012a35c800 tmpfs none /var/log ffff880137668c40 ffff88012a35c800 tmpfs none /var/lib/dbus ffff880129dcea80 ffff88012a35c800 tmpfs none /tmp ffff880137668e00 ffff88012a35c800 tmpfs none /var/lib/dhclient ffff880137668fc0 ffff88012a35c800 tmpfs none /var/tmp ffff880137669180 ffff88012a35c800 tmpfs none /var/lib/NetworkManager ffff880129dcec40 ffff88012a35c800 tmpfs none /var/lib/systemd/random-seed ffff88012a313180 ffff88012a35c800 tmpfs none /var/spool ffff880137669340 ffff88012a35c800 tmpfs none /var/lib/nfs ffff880137669500 ffff88012a35c800 tmpfs none /var/lib/gssproxy ffff88012a313340 ffff88012a35c800 tmpfs none /var/lib/logrotate ffff8801376696c0 ffff88012a35c800 tmpfs none /etc ffff88012a313500 ffff88012a35c800 tmpfs none /var/lib/rsyslog ffff880137669880 ffff88012a35c800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff880137669a40 ffff880129f13800 nfs4 192.168.200.253:/exports/state/oleg412-server.virtnet /var/lib/stateless/state ffff88012a313a40 ffff880129f13800 nfs4 192.168.200.253:/exports/state/oleg412-server.virtnet /boot ffff8801377f6fc0 ffff880129f13800 nfs4 192.168.200.253:/exports/state/oleg412-server.virtnet /etc/etc/kdump.conf ffff880129dcee00 ffff88012a35c000 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b528fc00 ffff880129f09800 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800b1880e00 ffff88013767f000 lustre /dev/mapper/ost1_flakey /mnt/lustre-ost1 ffff880129d86e00 ffff88012d4cd800 lustre /dev/mapper/ost2_flakey /mnt/lustre-ost2 ffff8800b1881500 ffff880122889000 lustre /dev/mapper/mds2_flakey /mnt/lustre-mds2 ffff8801376d16c0 ffff88006ef28000 lustre /dev/mapper/mds1_flakey /mnt/lustre-mds1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 1785.558397] LustreError: MGC192.168.204.112@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 1785.625669] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 1785.638419] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 1786.700510] Lustre: DEBUG MARKER: oleg412-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1788.635621] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 1790.635142] Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to (at 0@lo) [ 1790.639229] Lustre: Skipped 2 previous similar messages [ 1790.648155] Lustre: lustre-MDT0000: Recovery over after 0:02, of 3 clients 3 recovered and 0 were evicted. [ 1790.672995] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:487 to 0x280000401:513) [ 1790.672999] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:486 to 0x2c0000401:513) [ 1791.835511] Lustre: DEBUG MARKER: oleg412-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1792.375947] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1800.502099] Lustre: DEBUG MARKER: == sanity-hsm test 408: Verify fiemap on release file ==== 02:30:13 (1738481413) [ 1805.088100] Lustre: DEBUG MARKER: == sanity-hsm test 409a: Coordinator should not stop when in use ========================================================== 02:30:18 (1738481418) [ 1806.010044] LustreError: 14030:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 sleeping for 5000ms [ 1811.017524] LustreError: 14030:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 awake [ 1819.796438] Lustre: DEBUG MARKER: == sanity-hsm test 409b: getattr released file with CDT stopped after remount ========================================================== 02:30:33 (1738481433) [ 1822.200476] Lustre: Modifying parameter lustre.mdt.lustre-MDT0000.hsm_control in log params [ 1822.204767] Lustre: Skipped 1 previous similar message [ 1843.847192] Lustre: Failing over lustre-MDT0000 [ 1843.926100] Lustre: server umount lustre-MDT0000 complete [ 1845.721417] Lustre: lustre-MDT0000-lwp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1845.731219] Lustre: Skipped 6 previous similar messages [ 1856.279393] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 1856.306248] LustreError: MGC192.168.204.112@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 1856.366400] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 1856.375670] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 1857.138143] Lustre: DEBUG MARKER: oleg412-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1858.746863] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 1861.370944] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 192.168.204.112@tcp (at 0@lo) [ 1861.375097] Lustre: Skipped 3 previous similar messages [ 1861.384898] Lustre: lustre-MDT0000: Recovery over after 0:02, of 3 clients 3 recovered and 0 were evicted. [ 1861.417096] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:517 to 0x2c0000401:545) [ 1861.417135] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:516 to 0x280000401:545) [ 1862.298185] Lustre: DEBUG MARKER: oleg412-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1862.723878] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1886.888636] Lustre: DEBUG MARKER: == sanity-hsm test 410: lfs data_version -s allows release of force-archived file ========================================================== 02:31:40 (1738481500) [ 1888.589199] Lustre: DEBUG MARKER: == sanity-hsm test 411: hsm_ops rbac role ================ 02:31:41 (1738481501) [ 1941.988869] Lustre: DEBUG MARKER: == sanity-hsm test 500: various LLAPI HSM tests ========== 02:32:35 (1738481555) [ 1956.922833] Lustre: HSM agent 44719430-2282-48fa-b7fc-3f94eb97fe5d already registered ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 11.05s (real) 5.98s (CPU), Child processes: 5.04s