************************ crashinfo ************************* /exports/testreports/47845/testresults/sanity-benchmark-special10-ldiskfs-centos7_x86_64-centos7_x86_64-retry1/oleg140-client-timeout-core (3.10.0-7.9-debug) +==========================+ | *** Crashinfo v1.3.7 *** | +==========================+ +++WARNING+++ PARTIAL DUMP with size(vmcore) < 25% size(RAM) KERNEL: /tmp/crash-anaysis.BwUFY/vmlinux [TAINTED] DUMPFILE: /exports/testreports/47845/testresults/sanity-benchmark-special10-ldiskfs-centos7_x86_64-centos7_x86_64-retry1/oleg140-client-timeout-core [PARTIAL DUMP] CPUS: 4 DATE: Thu Dec 12 11:07:12 EST 2024 UPTIME: 01:04:58 LOAD AVERAGE: 1.00, 1.01, 1.05 TASKS: 154 NODENAME: oleg140-client.virtnet RELEASE: 3.10.0-7.9-debug VERSION: #1 SMP Sat Mar 26 23:28:42 EDT 2022 MACHINE: x86_64 (2399 Mhz) MEMORY: 4 GB PANIC: "" +--------------------------+ >------------------------| Per-cpu Stacks ('bt -a') |------------------------< +--------------------------+ -- CPU#0 -- PID=0 CPU=0 CMD=swapper/0 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 rest_init+0x8e #7 start_kernel+0x456 #8 x86_64_start_reservations+0x2a #9 x86_64_start_kernel+0x152 #10 start_cpu+0x5 -- CPU#1 -- PID=0 CPU=1 CMD=swapper/1 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#2 -- PID=0 CPU=2 CMD=swapper/2 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 -- CPU#3 -- PID=0 CPU=3 CMD=swapper/3 #-1 native_safe_halt+0xb, 449 bytes of data #0 default_idle+0x1e #1 default_enter_idle+0x45 #2 cpuidle_enter_state+0x40 #3 cpuidle_idle_call+0xd8 #4 arch_cpu_idle+0xe #5 cpu_startup_entry+0x14a #6 start_secondary+0x1eb #7 start_cpu+0x5 +--------------------------------+ >---------------------| How This Dump Has Been Created |---------------------< +--------------------------------+ Cannot identify the specific condition that triggered vmcore +---------------+ >------------------------------| Tasks Summary |------------------------------< +---------------+ Number of Threads That Ran Recently ----------------------------------- last second 16 last 5s 24 last 60s 32 ----- Total Numbers of Threads per State ------ TASK_INTERRUPTIBLE 149 TASK_RUNNING 1 TASK_UNINTERRUPTIBLE|TASK_DEAD 1 +++WARNING+++ There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details +-----------------------+ >--------------------------| 5 Most Recent Threads |--------------------------< +-----------------------+ PID CMD Age ARGS ----- -------------- ------ ---------------------------- 922 in:imjournal 0 ms /usr/sbin/rsyslogd -n 1971 ptlrpcd_rcv 0 ms (no user stack) 1963 socknal_reaper 0 ms (no user stack) 1138 sshd 39 ms sshd: root@pts/0 24 kworker/2:0 70 ms (no user stack) +------------------------+ >-------------------------| Memory Usage (kmem -i) |-------------------------< +------------------------+ PAGES TOTAL PERCENTAGE TOTAL MEM 955079 3.6 GB ---- FREE 362468 1.4 GB 37% of TOTAL MEM USED 592611 2.3 GB 62% of TOTAL MEM SHARED 473682 1.8 GB 49% of TOTAL MEM BUFFERS 5150 20.1 MB 0% of TOTAL MEM CACHED 516487 2 GB 54% of TOTAL MEM SLAB 40208 157.1 MB 4% of TOTAL MEM TOTAL HUGE 0 0 ---- HUGE FREE 0 0 0% of TOTAL HUGE TOTAL SWAP 262143 1024 MB ---- SWAP USED 0 0 0% of TOTAL SWAP SWAP FREE 262143 1024 MB 100% of TOTAL SWAP COMMIT LIMIT 739682 2.8 GB ---- COMMITTED 61650 240.8 MB 8% of TOTAL LIMIT +++three oldest UNINTERRUPTIBLE threads ... ran 3481s ago PID=10504 CPU=1 CMD=bonnie++ #0 __schedule+0x2e2 #1 schedule+0x29 #2 schedule_timeout+0x209 #3 io_schedule_timeout+0xad #4 io_schedule+0x18 #5 bit_wait_io+0x11 #6 __wait_on_bit_lock+0x5f #7 __lock_page_killable+0x74 #8 do_generic_file_read.constprop.52+0x3c9 #9 generic_file_aio_read+0x1c9 #10 do_file_read_iter+0x8ca #11 ll_file_read_iter+0x2d #12 ll_file_aio_read+0x1b1 #13 ll_file_read+0x111 #14 vfs_read+0xb9 #15 sys_read+0x7f #16 system_call_fastpath+0x1f, 477 bytes of data +-------------------------------+ >----------------------| Scheduler Runqueues (per CPU) |----------------------< +-------------------------------+ ---+ CPU=0 ---- | CURRENT TASK , CMD=swapper/0 ---+ CPU=1 ---- | CURRENT TASK , CMD=swapper/1 ---+ CPU=2 ---- | CURRENT TASK , CMD=swapper/2 ---+ CPU=3 ---- | CURRENT TASK , CMD=swapper/3 +------------------------+ >-------------------------| Network Status Summary |-------------------------< +------------------------+ TCP Connection Info ------------------- ESTABLISHED 3 LISTEN 3 NAGLE disabled (TCP_NODELAY): 3 user_data set (NFS etc.): 1 UDP Connection Info ------------------- 2 UDP sockets, 0 in ESTABLISHED Unix Connection Info ------------------------ ESTABLISHED 26 CLOSE 18 LISTEN 8 Raw sockets info -------------------- ESTABLISHED 1 Interfaces Info --------------- How long ago (in seconds) interfaces transmitted/received? Name RX TX ---- ---------- --------- lo n/a 3892.6 eth0 n/a 0.0 RSS_TOTAL=74328 pages, %mem= 1.2 +++WARNING+++ Run 'hanginfo' to get more details +------------+ >-------------------------------| Mounted FS |-------------------------------< +------------+ MOUNT SUPERBLK TYPE DEVNAME DIRNAME ffff880138cca000 ffff880139940800 rootfs rootfs / ffff880138ccaa80 ffff88012b397800 sysfs sysfs /sys ffff880138ccac40 ffff880139944000 proc proc /proc ffff880138ccae00 ffff880137678000 devtmpfs devtmpfs /dev ffff880138ccafc0 ffff8800b6d40800 securityfs securityfs /sys/kernel/security ffff880138ccb180 ffff88012aa78000 tmpfs tmpfs /dev/shm ffff880138ccb340 ffff8801370db000 devpts devpts /dev/pts ffff880138ccb500 ffff88012aa78800 tmpfs tmpfs /run ffff880138ccb6c0 ffff88012aa79000 tmpfs tmpfs /sys/fs/cgroup ffff880138ccb880 ffff88012aa79800 cgroup cgroup /sys/fs/cgroup/systemd ffff880138ccba40 ffff88012aa7a000 pstore pstore /sys/fs/pstore ffff880138ccbc00 ffff88012aa7c000 cgroup cgroup /sys/fs/cgroup/net_cls,net_prio ffff880138ccbdc0 ffff88012aa7b800 cgroup cgroup /sys/fs/cgroup/cpu,cpuacct ffff88012aaaa000 ffff88012aa7b000 cgroup cgroup /sys/fs/cgroup/hugetlb ffff88012aaaa1c0 ffff88012aa7a800 cgroup cgroup /sys/fs/cgroup/devices ffff88012aaaa380 ffff88012aa7c800 cgroup cgroup /sys/fs/cgroup/pids ffff88012aaaa540 ffff88012aa7d000 cgroup cgroup /sys/fs/cgroup/freezer ffff88012aaaa700 ffff88012aa7d800 cgroup cgroup /sys/fs/cgroup/cpuset ffff88012aaaa8c0 ffff88012aa7e000 cgroup cgroup /sys/fs/cgroup/blkio ffff88012aaaaa80 ffff88012aa7e800 cgroup cgroup /sys/fs/cgroup/memory ffff88012aaaac40 ffff88012aa7f000 cgroup cgroup /sys/fs/cgroup/perf_event ffff8801370988c0 ffff8800b6100000 configfs configfs /sys/kernel/config ffff880137098c40 ffff8800b6104000 ext4 /dev/nbd0 / ffff880137669500 ffff88012b310800 rpc_pipefs rpc_pipefs /var/lib/nfs/rpc_pipefs ffff88012aaaafc0 ffff8800b7b5b000 autofs systemd-1 /proc/sys/fs/binfmt_misc ffff880137098e00 ffff8801370a3800 mqueue mqueue /dev/mqueue ffff8801376696c0 ffff88012aabc800 hugetlbfs hugetlbfs /dev/hugepages ffff88012b26ae00 ffff880139947800 debugfs debugfs /sys/kernel/debug ffff880137098fc0 ffff88012b317000 binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc/ ffff880137669a40 ffff88012aaba000 ramfs none /mnt ffff880137099340 ffff8800b6325000 tmpfs none /var/lib/stateless/writable ffff880137099500 ffff8800b6321800 squashfs /dev/vda /home/green/git/lustre-release ffff880137669dc0 ffff8800b6325000 tmpfs none /var/cache/man ffff8801370996c0 ffff8800b6325000 tmpfs none /var/log ffff88012b26afc0 ffff8800b6325000 tmpfs none /var/lib/dbus ffff8800b6e62000 ffff8800b6325000 tmpfs none /tmp ffff88012b26b180 ffff8800b6325000 tmpfs none /var/lib/dhclient ffff88012aaab180 ffff8800b6325000 tmpfs none /var/tmp ffff8800b6e621c0 ffff8800b6325000 tmpfs none /var/lib/NetworkManager ffff880137099880 ffff8800b6325000 tmpfs none /var/lib/systemd/random-seed ffff8800b6e62380 ffff8800b6325000 tmpfs none /var/spool ffff88012aaab340 ffff8800b6325000 tmpfs none /var/lib/nfs ffff88012aaab500 ffff8800b6325000 tmpfs none /var/lib/gssproxy ffff88012b26b340 ffff8800b6325000 tmpfs none /var/lib/logrotate ffff88012aaab6c0 ffff8800b6325000 tmpfs none /etc ffff880137099a40 ffff8800b6325000 tmpfs none /var/lib/rsyslog ffff88012b26b500 ffff8800b6325000 tmpfs none /var/lib/dhclient/var/lib/dhclient ffff88012b26b6c0 ffff8800b6324000 nfs4 192.168.200.253:/exports/state/oleg140-client.virtnet /var/lib/stateless/state ffff880137099dc0 ffff8800b6324000 nfs4 192.168.200.253:/exports/state/oleg140-client.virtnet /boot ffff88012b26b880 ffff8800b6324000 nfs4 192.168.200.253:/exports/state/oleg140-client.virtnet /etc/etc/kdump.conf ffff8800b6e62540 ffff88012b310800 rpc_pipefs sunrpc /var/lib/nfs/var/lib/nfs/rpc_pipefs ffff88012a5441c0 ffff88013730a000 nfs4 192.168.200.253://exports/testreports/47845/testresults/sanity-benchmark-special10-ldiskfs-centos7_x86_64-centos7_x86_64-retry1 /tmp/tmp/testlogs ffff8800b3fbe700 ffff8800b6d40000 tmpfs tmpfs /run/user/0 ffff8801377f5340 ffff8800b6321800 squashfs /dev/vda /usr/sbin/mount.lustre ffff88012e1d9dc0 ffff8800b41c4000 lustre 192.168.201.140@tcp:/lustre /mnt/lustre +-------------------------------+ >----------------------| Last 40 lines of dmesg buffer |----------------------< +-------------------------------+ [ 99.982788] Lustre: DEBUG MARKER: Using TIMEOUT=20 [ 110.318567] Lustre: DEBUG MARKER: oleg140-client.virtnet: executing check_logdir /tmp/testlogs/ [ 111.117754] Lustre: DEBUG MARKER: oleg140-client.virtnet: executing yml_node [ 112.225315] Lustre: DEBUG MARKER: Client: 2.16.0.RC5 [ 112.881933] Lustre: DEBUG MARKER: MDS: 2.16.0.RC5 [ 113.507714] Lustre: DEBUG MARKER: OSS: 2.16.0.RC5 [ 113.911052] Lustre: DEBUG MARKER: -----============= acceptance-small: sanity-benchmark ============----- Thu Dec 12 10:04:06 EST 2024 [ 119.116587] Lustre: DEBUG MARKER: excepting tests: [ 119.633474] Lustre: DEBUG MARKER: skipping tests SLOW=no: iozone [ 120.138333] Lustre: DEBUG MARKER: === sanity-benchmark: start setup 10:04:12 (1734015852) === [ 121.206362] Lustre: DEBUG MARKER: oleg140-client.virtnet: executing check_config_client /mnt/lustre [ 123.015359] Lustre: lustre-OST0000-osc-ffff8800b41c4000: disconnect after 24s idle [ 128.263587] Lustre: DEBUG MARKER: Using TIMEOUT=20 [ 133.670037] Lustre: DEBUG MARKER: === sanity-benchmark: finish setup 10:04:26 (1734015866) === [ 134.443927] Lustre: DEBUG MARKER: == sanity-benchmark test dbench: dbench ================== 10:04:26 (1734015866) [ 138.560838] random: crng init done [ 259.689441] hrtimer: interrupt took 13908030 ns [ 284.563531] dbench (9459) used greatest stack depth: 10368 bytes left [ 290.397682] LNet: 9707:0:(debug.c:314:libcfs_debug_str2mask()) using a numerical debug mask is deprecated [ 295.502790] Lustre: DEBUG MARKER: == sanity-benchmark test bonnie: bonnie++ ================ 10:07:07 (1734016027) [ 296.769376] Lustre: DEBUG MARKER: min OST has 3569592kB available, using 5354388kB file size [ 433.479282] Lustre: 1974:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1734016150/real 1734016150] req@ffff8800a017a300 x1818247359199616/t0(0) o4->lustre-OST0000-osc-ffff8800b41c4000@192.168.201.140@tcp:6/4 lens 488/448 e 0 to 1 dl 1734016166 ref 2 fl Rpc:XQr/200/ffffffff rc 0/-1 job:'bonnie++.500' uid:500 gid:500 [ 433.493229] Lustre: lustre-OST0000-osc-ffff8800b41c4000: Connection to lustre-OST0000 (at 192.168.201.140@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 434.500265] Lustre: 1972:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1734016151/real 1734016151] req@ffff8800a007fb80 x1818247359202048/t0(0) o400->lustre-MDT0000-mdc-ffff8800b41c4000@192.168.201.140@tcp:12/10 lens 224/224 e 0 to 1 dl 1734016167 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 434.500274] Lustre: 1973:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1734016151/real 1734016151] req@ffff8800a007ed80 x1818247359201920/t0(0) o400->MGC192.168.201.140@tcp@192.168.201.140@tcp:26/25 lens 224/224 e 0 to 1 dl 1734016167 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 434.500279] Lustre: 1973:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 6 previous similar messages [ 434.500306] LustreError: MGC192.168.201.140@tcp: Connection to MGS (at 192.168.201.140@tcp) was lost; in progress operations using this service will fail [ 434.552473] Lustre: 1972:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 8 previous similar messages [ 434.556489] Lustre: lustre-MDT0000-mdc-ffff8800b41c4000: Connection to lustre-MDT0000 (at 192.168.201.140@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 434.563169] Lustre: Skipped 1 previous similar message [ 439.500252] Lustre: 1973:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1734016156/real 1734016156] req@ffff8800a007e680 x1818247359202688/t0(0) o400->lustre-OST0000-osc-ffff8800b41c4000@192.168.201.140@tcp:28/4 lens 224/224 e 0 to 1 dl 1734016172 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 439.514946] Lustre: 1973:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 3 previous similar messages [ 444.519273] Lustre: 1973:0:(client.c:2363:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1734016161/real 1734016161] req@ffff8800a007e300 x1818247359203200/t0(0) o400->lustre-OST0000-osc-ffff8800b41c4000@192.168.201.140@tcp:28/4 lens 224/224 e 0 to 1 dl 1734016177 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 [ 444.533831] Lustre: 1973:0:(client.c:2363:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 474.101367] LustreError: 1967:0:(events.c:211:client_bulk_callback()) event type 1, status -5, desc ffff880062700000 [ 474.101419] LNetError: 1963:0:(socklnd.c:1585:ksocknal_destroy_conn()) Completing partial receive from 12345-192.168.201.140@tcp[2], ip 192.168.201.140:988, with error, wanted: 658472, left: 658472, last alive is 57 secs ago [ 474.116050] LustreError: 1967:0:(events.c:211:client_bulk_callback()) Skipped 1 previous similar message [ 1618.299244] LustreError: 4741:0:(osc_cache.c:908:osc_extent_wait()) extent ffff8800a3bb5000@{[71680 -> 72703/72703], [3|0|+|rpc|wiY|ffff880095c4d860], [4218880|1024|+|-|ffff8800a03d3180|1024|ffff8801373b0000]} lustre-OST0000-osc-ffff8800b41c4000: wait ext to 0 timedout, recovery in progress? [ 1618.299264] LustreError: 9472:0:(osc_cache.c:908:osc_extent_wait()) ### extent: ffff8800a3bb5180 ns: lustre-OST0001-osc-ffff8800b41c4000 lock: ffff8800a03d0240/0xa843d1e1638d6af1 lrc: 9/0,0 mode: PW/PW res: [0x280000400:0xac1:0x0].0x0 rrc: 2 type: EXT [0->18446744073709551615] (req 0->8191) gid 0 flags: 0x800029400000000 nid: local remote: 0x43f07197e5998f4c expref: -99 pid: 10504 timeout: 0 lvb_type: 1 [ 1618.327088] LustreError: 4741:0:(osc_cache.c:908:osc_extent_wait()) Skipped 1 previous similar message ****************************************************************************** ************************ A Summary Of Problems Found ************************* ****************************************************************************** -------------------- A list of all +++WARNING+++ messages -------------------- PARTIAL DUMP with size(vmcore) < 25% size(RAM) There are 3 threads running in their own namespaces Use 'taskinfo --ns' to get more details Run 'hanginfo' to get more details ------------------------------------------------------------------------------ ** Execution took 11.09s (real) 5.77s (CPU), Child processes: 5.30s