************************ crashinfo ************************* /exports/testreports/49655/testresults/sanity-hsm-zfs-centos7_x86_64-centos7_x86_64/oleg420-server-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.N6iCe/vmlinux [TAINTED] DUMPFILE: /exports/testreports/49655/testresults/sanity-hsm-zfs-centos7_x86_64-centos7_x86_64/oleg420-server-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Thu Feb 27 12:43:03 EST 2025 UPTIME: 01:23:37 LOAD AVERAGE: 0.00, 0.01, 0.05 TASKS: 360 NODENAME: oleg420-server.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2400 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 46 last 5s 62 last 60s 77 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 356 TASK_RUNNING 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 11 rcuos/0 0 ms (no user stack) 9 rcu_sched 0 ms (no user stack) 13846 mmp 0 ms (no user stack) 9449 mmp 0 ms (no user stack) 3289 monitor_thread 42 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955067 3.6 GB ---- FREE 633944 2.4 GB 66% of TOTAL MEM USED 321123 1.2 GB 33% of TOTAL MEM SHARED 10362 40.5 MB 1% of TOTAL MEM BUFFERS 5162 20.2 MB 0% of TOTAL MEM CACHED 72745 284.2 MB 7% of TOTAL MEM SLAB 22337 87.3 MB 2% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739676 2.8 GB ---- COMMITTED 61072 238.6 MB 8% of TOTAL LIMIT +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 9 LISTEN 3 NAGLE disabled (TCP_NODELAY): 7 user_data set (NFS etc.): 8 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 17 LISTEN 8 Raw sockets info -------------------- CLOSE 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 5015.0 eth0 n/a 1.7 RSS_TOTAL=55564 pages, %mem= 0.9 +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff88012a340000 ffff88012a348000 sysfs sysfs /sys ffff88012a3401c0 ffff880139944000 proc proc /proc ffff88012a340380 ffff880137678000 devtmpfs devtmpfs /dev ffff88012a340540 ffff8800b54cf800 securityfs securityfs /sys/kernel/security ffff88012a340700 ffff88012a348800 tmpfs tmpfs /dev/shm ffff88012a3408c0 ffff8801372f6000 devpts devpts /dev/pts ffff88012a340a80 ffff88012a349000 tmpfs tmpfs /run ffff88012a340c40 ffff88012a349800 tmpfs tmpfs /sys/fs/cgroup ffff88012a340e00 ffff88012a34a000 cgroup cgroup /sys/fs/cgroup/systemd ffff88012a340fc0 ffff88012a34a800 pstore pstore /sys/fs/pstore ffff88012a341180 ffff88012a34c800 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012a341340 ffff88012a34c000 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012a341500 ffff88012a34b800 cgroup cgroup /sys/fs/cgroup/memory ffff88012a3416c0 ffff88012a34b000 cgroup cgroup /sys/fs/cgroup/perf_event ffff88012a341880 ffff88012a34d000 cgroup cgroup /sys/fs/cgroup/devices ffff88012a341a40 ffff88012a34d800 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012a341c00 ffff88012a34e000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012a341dc0 ffff88012a34e800 cgroup cgroup /sys/fs/cgroup/blkio ffff88012a2c4000 ffff88012a34f000 cgroup cgroup /sys/fs/cgroup/pids ffff88012a2c41c0 ffff88012a34f800 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff880138ccb500 ffff8800b4da0800 configfs configfs /sys/kernel/config ffff88012b262a80 ffff8800b7c1c800 ext4 /dev/nbd0 / ffff880129f28700 ffff88012a312800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012b262c40 ffff880129d6d000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff88012a2c4c40 ffff88012a311800 hugetlbfs hugetlbfs /dev/hugepages ffff88012b262e00 ffff8801372f7800 mqueue mqueue /dev/mqueue ffff88012b262fc0 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff880137668380 ffff880129fa5000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff880129f288c0 ffff8800b53f9000 ramfs none /mnt ffff88012a2c4e00 ffff880129d6e800 tmpfs none /var/lib/stateless/writable ffff88012b263340 ffff880129d12000 squashfs /dev/vda /home/green/git/lustre-release ffff880137668540 ffff880129d6e800 tmpfs none /var/cache/man ffff880129f28c40 ffff880129d6e800 tmpfs none /var/log ffff88012a2c4fc0 ffff880129d6e800 tmpfs none /var/lib/dbus ffff880137668700 ffff880129d6e800 tmpfs none /tmp ffff88012a2c5180 ffff880129d6e800 tmpfs none /var/lib/dhclient ffff88012b263500 ffff880129d6e800 tmpfs none /var/tmp ffff88012a2c5340 ffff880129d6e800 tmpfs none /var/lib/NetworkManager ffff8801376688c0 ffff880129d6e800 tmpfs none /var/lib/systemd/random-seed ffff88012b2636c0 ffff880129d6e800 tmpfs none /var/spool ffff880129f28e00 ffff880129d6e800 tmpfs none /var/lib/nfs ffff880137668a80 ffff880129d6e800 tmpfs none /var/lib/gssproxy ffff880129f28fc0 ffff880129d6e800 tmpfs none /var/lib/logrotate ffff880129f29180 ffff880129d6e800 tmpfs none /etc ffff88012a2c5500 ffff880129d6e800 tmpfs none /var/lib/rsyslog ffff880137668c40 ffff880129d6e800 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff88012a2c56c0 ffff8800b408c800 nfs4 192.168.200.253:/exports/state/oleg420-server.virtnet /var/lib/stateless/state ffff88012a2c5a40 ffff8800b408c800 nfs4 192.168.200.253:/exports/state/oleg420-server.virtnet /boot ffff880137668e00 ffff8800b408c800 nfs4 192.168.200.253:/exports/state/oleg420-server.virtnet /etc/etc/kdump.conf ffff88012a2c5c00 ffff88012a312800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff8800b1f59dc0 ffff880129d12000 squashfs /dev/vda /usr/sbin/mount.lustre ffff8800b1fcb180 ffff880130103800 lustre lustre-ost1/ost1 /mnt/lustre-ost1 ffff8800a4d988c0 ffff880099f35800 lustre lustre-ost2/ost2 /mnt/lustre-ost2 ffff8800a4d99500 ffff880099f36000 lustre lustre-mdt1/mdt1 /mnt/lustre-mds1 +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 1561.476965] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 1562.069676] Lustre: DEBUG MARKER: oleg420-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1566.265797] Lustre: 3294:0:(client.c:2348:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1740674715/real 1740674715] req@ffff880137362680 x1825228135261952/t0(0) o400->lustre-MDT0000-lwp-OST0000@0@lo:12/10 lens 224/224 e 0 to 1 dl 1740674731 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 1566.273795] Lustre: 3294:0:(client.c:2348:ptlrpc_expire_one_request()) Skipped 11 previous similar messages [ 1566.292790] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to 192.168.204.120@tcp (at 0@lo) [ 1566.294913] Lustre: Skipped 1 previous similar message [ 1566.512994] Lustre: lustre-MDT0000: Recovery over after 0:05, of 2 clients 2 recovered and 0 were evicted. [ 1566.528757] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:707 to 0x280000400:737) [ 1566.528762] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:707 to 0x240000400:737) [ 1567.332050] Lustre: DEBUG MARKER: oleg420-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1567.736820] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1577.295554] Lustre: DEBUG MARKER: == sanity-hsm test 408: Verify fiemap on release file ==== 11:45:41 (1740674741) [ 1577.778914] Lustre: DEBUG MARKER: SKIP: sanity-hsm test_408 ORI-366/LU-1941: FIEMAP unimplemented on ZFS [ 1578.243394] Lustre: DEBUG MARKER: == sanity-hsm test 409a: Coordinator should not stop when in use ========================================================== 11:45:42 (1740674742) [ 1578.894056] LustreError: 10724:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 sleeping for 5000ms [ 1583.897763] LustreError: 10724:0:(mdt_hsm_cdt_client.c:366:mdt_hsm_register_hal()) cfs_fail_timeout id 164 awake [ 1591.521942] Lustre: DEBUG MARKER: == sanity-hsm test 409b: getattr released file with CDT stopped after remount ========================================================== 11:45:55 (1740674755) [ 1593.955686] Lustre: Modifying parameter lustre.mdt.lustre-MDT0000.hsm_control in log params [ 1614.825682] Lustre: Failing over lustre-MDT0000 [ 1614.947705] Lustre: server umount lustre-MDT0000 complete [ 1626.863443] LustreError: MGC192.168.204.120@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 1626.958329] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 1626.968070] Lustre: Skipped 1 previous similar message [ 1626.997969] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 1627.024252] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 1627.865079] Lustre: DEBUG MARKER: oleg420-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 1629.678316] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 1631.646024] Lustre: lustre-MDT0000: Recovery over after 0:02, of 2 clients 2 recovered and 0 were evicted. [ 1631.663594] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:740 to 0x280000400:769) [ 1631.663857] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:741 to 0x240000400:769) [ 1631.957743] Lustre: 3293:0:(client.c:2348:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1740674781/real 1740674781] req@ffff880131cc4000 x1825228135289600/t0(0) o400->lustre-MDT0000-lwp-OST0001@0@lo:12/10 lens 224/224 e 0 to 1 dl 1740674797 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 1631.968089] Lustre: 3293:0:(client.c:2348:ptlrpc_expire_one_request()) Skipped 5 previous similar messages [ 1631.988826] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to 192.168.204.120@tcp (at 0@lo) [ 1631.991333] Lustre: Skipped 1 previous similar message [ 1632.728826] Lustre: DEBUG MARKER: oleg420-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 1633.321647] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 1657.602847] Lustre: DEBUG MARKER: == sanity-hsm test 410: lfs data_version -s allows release of force-archived file ========================================================== 11:47:02 (1740674822) [ 1659.755720] Lustre: DEBUG MARKER: == sanity-hsm test 411: hsm_ops rbac role ================ 11:47:04 (1740674824) [ 1712.531111] Lustre: DEBUG MARKER: == sanity-hsm test 500: various LLAPI HSM tests ========== 11:47:56 (1740674876) [ 1728.159393] Lustre: HSM agent 3491f002-86bb-4b85-961c-2ebafb7a84dc already registered ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details ------------------------------------------------------------------------------ ** Execution took 11.80s (real) 6.75s (CPU), Child processes: 5.02s