[ 2769.142119] Lustre: *** cfs_fail_loc=513, val=601*** [ 2769.547085] LustreError: 6521:0:(service.c:2346:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1876424818216320 [ 2769.889095] Lustre: *** cfs_fail_loc=513, val=601*** [ 2769.892050] Lustre: Skipped 30 previous similar messages [ 2770.911724] Lustre: *** cfs_fail_loc=513, val=601*** [ 2770.915973] Lustre: Skipped 6 previous similar messages [ 2774.496205] Lustre: *** cfs_fail_loc=513, val=601*** [ 2774.502684] Lustre: Skipped 6 previous similar messages [ 2779.617217] Lustre: *** cfs_fail_loc=513, val=601*** [ 2779.619025] Lustre: Skipped 9 previous similar messages [ 2785.247275] Lustre: 42294:0:(service.c:1612:ptlrpc_at_send_early_reply()) @@@ Could not add any time (4/4), not sending early reply req@ffff8ab245b52a00 x1876424803565824/t0(0) o4->e4fb8066-e18f-4cab-963a-68940383e459@192.168.203.44@tcp:590/0 lens 488/448 e 1 to 0 dl 1789500835 ref 2 fl Interpret:/600/0 rc 0/0 job:'dd.60000' uid:60000 gid:60000 projid:0 [ 2785.759668] Lustre: 19496:0:(client.c:2504:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1789500815/real 1789500815] req@ffff8ab241f68000 x1876424818216320/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1789500831 ref 2 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'ll_ost_io00_003.0' uid:0 gid:0 projid:4294967295 [ 2785.780426] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2785.805426] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2785.818555] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to 0@lo (at 0@lo) [ 2785.879609] LustreError: 16575:0:(service.c:2346:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1876424818225280 [ 2789.861752] Lustre: *** cfs_fail_loc=513, val=601*** [ 2789.871137] Lustre: Skipped 46 previous similar messages [ 2794.465233] LustreError: 16572:0:(service.c:2346:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1876424818228480 [ 2794.486248] LustreError: 16572:0:(service.c:2346:ptlrpc_server_handle_req_in()) Skipped 2 previous similar messages [ 2801.103148] Lustre: 3644:0:(client.c:2504:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1789500831/real 1789500831] req@ffff8ab245be0700 x1876424818225280/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1789500847 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'ll_ost_io00_005.0' uid:0 gid:0 projid:4294967295 [ 2801.132309] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2801.149573] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2801.164054] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to 0@lo (at 0@lo) [ 2804.240210] LustreError: 6520:0:(service.c:2346:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1876424818233600 [ 2806.239653] Lustre: *** cfs_fail_loc=513, val=601*** [ 2806.249553] Lustre: Skipped 79 previous similar messages [ 2810.335142] Lustre: 3644:0:(client.c:2504:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1789500840/real 1789500840] req@ffff8ab244ceca80 x1876424818228352/t0(0) o601->lustre-MDT0000-lwp-MDT0001@0@lo:23/10 lens 336/336 e 0 to 1 dl 1789500856 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'lquota_wb_lustr.0' uid:0 gid:0 projid:4294967295 [ 2810.336463] Lustre: lustre-MDT0000-lwp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2810.374840] Lustre: 3644:0:(client.c:2504:ptlrpc_expire_one_request()) Skipped 1 previous similar message [ 2810.420583] Lustre: lustre-MDT0000: Received new LWP connection from 0@lo, keep former export from same NID [ 2810.437643] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 2810.847931] LustreError: 16575:0:(service.c:2346:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1876424818237952 [ 2810.861549] LustreError: 16575:0:(service.c:2346:ptlrpc_server_handle_req_in()) Skipped 3 previous similar messages [ 2820.575424] Lustre: 19496:0:(client.c:2504:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1789500850/real 1789500850] req@ffff8ab245be0700 x1876424818233600/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1789500866 ref 2 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'ll_ost_io00_003.0' uid:0 gid:0 projid:4294967295 [ 2820.621128] Lustre: 19496:0:(client.c:2504:ptlrpc_expire_one_request()) Skipped 1 previous similar message [ 2820.631340] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2820.659476] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2820.678608] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to 0@lo (at 0@lo) [ 2826.703200] Lustre: 3645:0:(client.c:2504:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1789500856/real 1789500856] req@ffff8ab24aed2d80 x1876424818238208/t0(0) o601->lustre-MDT0000-lwp-MDT0001@0@lo:23/10 lens 336/336 e 0 to 1 dl 1789500872 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'lquota_wb_lustr.0' uid:0 gid:0 projid:4294967295 [ 2826.721068] Lustre: lustre-MDT0000-lwp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2826.725648] Lustre: 3645:0:(client.c:2504:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 2826.755641] Lustre: Skipped 1 previous similar message [ 2826.769795] Lustre: lustre-MDT0000: Received new LWP connection from 0@lo, keep former export from same NID [ 2826.770651] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0001_UUID (at 0@lo) reconnecting [ 2826.783253] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 2851.380809] Lustre: DEBUG MARKER: == sanity-quota test 7a: Quota reintegration (global index) ========================================================== 15:34:55 (1789500895) [ 2875.845959] Lustre: Failing over lustre-OST0000 [ 2875.960646] Lustre: server umount lustre-OST0000 complete [ 2877.927969] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2883.053692] LustreError: 42582:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 2883.082208] LustreError: 42582:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 5 previous similar messages [ 2885.899367] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 2886.155158] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 2886.186727] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 2888.034173] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 2888.220457] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 0@lo (at 0@lo) [ 2888.220601] Lustre: lustre-OST0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 2888.226676] Lustre: Skipped 1 previous similar message [ 2891.044675] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 2897.732559] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 2903.069438] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 2908.893992] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 2913.555974] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 2919.031848] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 2924.907621] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 2946.165842] Lustre: Failing over lustre-OST0000 [ 2946.249419] Lustre: server umount lustre-OST0000 complete [ 2947.553411] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2947.556479] LustreError: 42582:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 2947.564958] Lustre: Skipped 2 previous similar messages [ 2947.577253] LustreError: 42582:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 1 previous similar message [ 2952.675725] LustreError: 8965:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 2952.682832] LustreError: 8965:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 2 previous similar messages [ 2953.678859] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 2954.109187] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 2954.137945] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 2955.304175] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 2960.223420] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 2968.908768] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3025.505919] Lustre: lustre-OST0000: recovery is timed out, evict stale exports [ 3025.518273] Lustre: 88557:0:(genops.c:1600:class_disconnect_stale_exports()) lustre-OST0000: disconnect stale client e4fb8066-e18f-4cab-963a-68940383e459@ [ 3025.526774] Lustre: lustre-OST0000: disconnecting 1 stale clients [ 3025.539141] Lustre: lustre-OST0000: Recovery over after 1:10, of 3 clients 2 recovered and 1 was evicted. [ 3025.539434] Lustre: lustre-OST0000-osc-MDT0001: Connection restored to 0@lo (at 0@lo) [ 3025.558959] Lustre: Skipped 1 previous similar message [ 3034.574671] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3039.245822] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3043.998887] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3048.489289] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3053.690265] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3082.621558] Lustre: DEBUG MARKER: == sanity-quota test 7b: Quota reintegration (slave index) ========================================================== 15:38:47 (1789501127) [ 3110.660128] Lustre: *** cfs_fail_loc=a02, val=0*** [ 3116.870274] LustreError: 3644:0:(qsd_handler.c:298:qsd_req_completion()) $$$ DQACQ failed with -5, flags:0x8 qsd:lustre-OST0000 qtype:usr lqe: ffff8ab242a24300 id:60000 enforced:1 granted: 1024 pending:0 waiting:0 req:1 usage: 2048 qunit:0 qtune:0 edquot:0 default:no revoke:0 [ 3116.893776] Lustre: Failing over lustre-OST0000 [ 3116.956834] Lustre: server umount lustre-OST0000 complete [ 3118.050404] LustreError: lustre-OST0000-osc-MDT0001: operation ost_statfs to node 0@lo failed: rc = -107 [ 3118.051443] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3118.054416] LustreError: Skipped 1 previous similar message [ 3118.092184] Lustre: Skipped 1 previous similar message [ 3118.105231] LustreError: 26540:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3118.129741] LustreError: 26540:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 1 previous similar message [ 3123.865404] LustreError: 8436:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 192.168.203.44@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3123.886577] LustreError: 8436:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 3 previous similar messages [ 3125.163445] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 3125.550941] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 3125.575962] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 3126.700339] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 3127.016375] Lustre: lustre-OST0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 3127.016968] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 0@lo (at 0@lo) [ 3127.049738] Lustre: Skipped 1 previous similar message [ 3132.302462] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3139.646162] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3145.233287] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3150.939618] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3156.194066] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3160.983362] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3166.359225] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3203.385329] Lustre: DEBUG MARKER: == sanity-quota test 7c: Quota reintegration (restart mds during reintegration) ========================================================== 15:40:47 (1789501247) [ 3223.956067] LustreError: 98280:0:(qsd_reint.c:488:qsd_reint_main()) cfs_fail_timeout id a03 sleeping for 10000ms [ 3223.964113] LustreError: 98280:0:(qsd_reint.c:488:qsd_reint_main()) Skipped 1 previous similar message [ 3225.757072] Lustre: Failing over lustre-MDT0000 [ 3226.148839] Lustre: server umount lustre-MDT0000 complete [ 3226.296849] LustreError: 6517:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-MDT0000: not available for connect from 192.168.203.44@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3226.324100] LustreError: 6517:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 2 previous similar messages [ 3227.107260] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 3227.119960] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3229.118768] LustreError: 98281:0:(qsd_reint.c:488:qsd_reint_main()) cfs_fail_timeout interrupted [ 3229.130608] LustreError: 98281:0:(qsd_reint.c:488:qsd_reint_main()) Skipped 1 previous similar message [ 3235.324287] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3235.419427] LustreError: MGC192.168.203.144@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 3235.706291] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 3235.757221] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 3236.510844] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 3240.704733] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3240.948567] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 3240.957949] Lustre: Skipped 1 previous similar message [ 3240.969507] LustreError: 3643:0:(ldlm_resource.c:1207:ldlm_resource_complain()) lustre-MDT0000-lwp-OST0001: namespace resource [0x200000006:0x2020000:0x0].0x0 (ffff8ab245a95f00) refcount nonzero (1) after lock cleanup; forcing cleanup. [ 3241.055520] Lustre: 98285:0:(qsd_reint.c:249:qsd_reint_index()) lustre-OST0001: index version for fid [0x200000005:0x100c:0x0] is 0, but index isn't empty (1) [ 3241.126047] Lustre: lustre-MDT0000: Recovery over after 0:05, of 2 clients 2 recovered and 0 were evicted. [ 3241.192592] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:130 to 0x2c0000401:161) [ 3241.194841] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:134 to 0x280000401:161) [ 3249.421378] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3344.287648] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3349.334724] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3355.468732] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3363.565516] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3371.216690] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3400.410818] Lustre: DEBUG MARKER: == sanity-quota test 7d: Quota reintegration (Transfer index in multiple bulks) ========================================================== 15:44:04 (1789501444) [ 3422.264908] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3428.745518] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3466.015846] Lustre: DEBUG MARKER: == sanity-quota test 7e: Quota reintegration (inode limits) ========================================================== 15:45:10 (1789501510) [ 3495.812474] Lustre: Failing over lustre-MDT0001 [ 3496.128121] Lustre: server umount lustre-MDT0001 complete [ 3496.928625] LustreError: lustre-MDT0001-osp-MDT0000: operation mds_statfs to node 0@lo failed: rc = -107 [ 3496.934082] Lustre: lustre-MDT0001-lwp-OST0000: Connection to lustre-MDT0001 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3496.943821] LustreError: 6522:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-MDT0001: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3496.943844] LustreError: 6522:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 10 previous similar messages [ 3496.994332] Lustre: Skipped 5 previous similar messages [ 3509.224389] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3509.726328] Lustre: lustre-MDT0001: Imperative Recovery not enabled, recovery window 60-180 [ 3509.774661] Lustre: lustre-MDT0001: in recovery but waiting for the first client to connect [ 3511.821435] Lustre: lustre-MDT0001: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 3514.866524] Lustre: lustre-MDT0001-lwp-OST0001: Connection restored to 0@lo (at 0@lo) [ 3514.884522] Lustre: Skipped 3 previous similar messages [ 3514.889887] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3514.904730] Lustre: lustre-MDT0001: Recovery over after 0:03, of 2 clients 2 recovered and 0 were evicted. [ 3514.964729] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:3 to 0x2c0000400:33) [ 3523.314037] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 3529.063293] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0001.recovery_status 1475 [ 3535.120752] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 3541.106816] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0001.recovery_status 1475 [ 3547.534794] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 3553.318425] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0001.recovery_status 1475 [ 3628.635968] Lustre: Failing over lustre-MDT0001 [ 3628.953430] Lustre: server umount lustre-MDT0001 complete [ 3629.739877] LustreError: 6547:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-MDT0001: not available for connect from 192.168.203.44@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3629.757345] LustreError: 6547:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 11 previous similar messages [ 3632.628291] LustreError: lustre-MDT0001-osp-MDT0000: operation mds_statfs to node 0@lo failed: rc = -107 [ 3637.406424] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3637.744067] Lustre: lustre-MDT0001: Not available for connect from 0@lo (not set up) [ 3637.753230] Lustre: Skipped 2 previous similar messages [ 3638.017741] Lustre: lustre-MDT0001: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 3638.081566] Lustre: lustre-MDT0001: in recovery but waiting for the first client to connect [ 3639.715131] Lustre: lustre-MDT0001: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 3642.604146] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3643.385245] Lustre: lustre-MDT0001: Recovery over after 0:04, of 2 clients 2 recovered and 0 were evicted. [ 3643.454379] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:3 to 0x2c0000400:65) [ 3650.986534] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 3655.863438] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0001.recovery_status 1475 [ 3661.084612] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 3666.200644] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0001.recovery_status 1475 [ 3672.096346] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 3678.497910] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0001.recovery_status 1475 [ 3756.579047] Lustre: DEBUG MARKER: == sanity-quota test 7f: Quota reintegration automatically ========================================================== 15:50:00 (1789501800) [ 3768.317634] Lustre: *** cfs_fail_loc=a11, val=0*** [ 3774.574251] Lustre: *** cfs_fail_loc=a11, val=0*** [ 3861.631687] Lustre: 115659:0:(qsd_reint.c:249:qsd_reint_index()) lustre-MDT0001: index version for fid [0x200000005:0x1004:0x0] is 0, but index isn't empty (1) [ 3864.364744] Lustre: DEBUG MARKER: == sanity-quota test 8: Run dbench with quota enabled ==== 15:51:48 (1789501908) [ 4050.899642] Lustre: DEBUG MARKER: == sanity-quota test 9: Block limit larger than 4GB (b10707) ========================================================== 15:54:55 (1789502095) [ 4052.233181] Lustre: DEBUG MARKER: OST0_SIZE: 3601408 required: 4900000 [ 4057.773567] Lustre: DEBUG MARKER: == sanity-quota test 10: Test quota for root user ======== 15:55:02 (1789502102) [ 4093.226641] Lustre: DEBUG MARKER: == sanity-quota test 11: Chown/chgrp ignores quota ======= 15:55:37 (1789502137) [ 4131.919797] Lustre: DEBUG MARKER: == sanity-quota test 12a: Block quota rebalancing ======== 15:56:16 (1789502176) [ 4193.648487] Lustre: DEBUG MARKER: == sanity-quota test 12b: Inode quota rebalancing ======== 15:57:18 (1789502238) [ 4322.327763] Lustre: DEBUG MARKER: == sanity-quota test 13: Cancel per-ID lock in the LRU list ========================================================== 15:59:26 (1789502366) [ 4379.668740] Lustre: DEBUG MARKER: == sanity-quota test 14: check panic in qmt_site_recalc_cb ========================================================== 16:00:23 (1789502423) [ 4400.439644] Lustre: Failing over lustre-OST0000 [ 4400.707649] Lustre: server umount lustre-OST0000 complete [ 4401.130257] LustreError: lustre-OST0000-osc-MDT0001: operation ost_statfs to node 0@lo failed: rc = -107 [ 4401.146128] Lustre: lustre-OST0000-osc-MDT0001: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 4401.162398] Lustre: Skipped 3 previous similar messages [ 4401.175757] LustreError: 42296:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 4401.193428] LustreError: 42296:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 5 previous similar messages [ 4411.367650] LustreError: 42587:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 4411.386170] LustreError: 42587:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 5 previous similar messages [ 4413.929848] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 4414.359256] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 4414.390930] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 4415.693646] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 4416.317786] Lustre: lustre-OST0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 4416.320779] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 0@lo (at 0@lo) [ 4416.349170] Lustre: Skipped 5 previous similar messages [ 4421.806840] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4457.078378] Lustre: DEBUG MARKER: == sanity-quota test 15: Set over 4T block quota ========= 16:01:41 (1789502501) [ 4483.052649] Lustre: DEBUG MARKER: == sanity-quota test 16a: lfs quota should skip the inactive MDT/OST ========================================================== 16:02:07 (1789502527) [ 4504.786968] Lustre: lustre-MDT0001: Client e4fb8066-e18f-4cab-963a-68940383e459 (at 192.168.203.44@tcp) reconnecting [ 4526.187853] Lustre: DEBUG MARKER: == sanity-quota test 16b: lfs quota should skip the nonexistent MDT/OST ========================================================== 16:02:50 (1789502570) [ 4527.810997] Lustre: DEBUG MARKER: SKIP: sanity-quota test_16b needs >= 3 MDTs [ 4529.559115] Lustre: DEBUG MARKER: == sanity-quota test 16c: lfs quota should preserve usage with an unavailable OST ========================================================== 16:02:53 (1789502573) [ 4543.799688] Lustre: setting import lustre-OST0000_UUID INACTIVE by administrator request [ 4544.647364] Lustre: setting import lustre-OST0000_UUID INACTIVE by administrator request [ 4546.906718] Lustre: Failing over lustre-OST0000 [ 4546.985126] Lustre: server umount lustre-OST0000 complete [ 4548.283417] LustreError: 42584:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-OST0000: not available for connect from 192.168.203.44@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 4548.326242] LustreError: 42584:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 1 previous similar message [ 4561.606905] Lustre: DEBUG MARKER: oleg344-client.virtnet: executing wait_import_state (DISCONN|IDLE) osc.lustre-OST0000-osc-ffff9587c62db800.ost_server_uuid 50 [ 4563.810393] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-ffff9587c62db800.ost_server_uuid in DISCONN state after 0 sec [ 4572.572979] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 4572.839447] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 4572.862186] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 4574.459784] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 4574.469740] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 0 in progress, and 0 evicted) to recover in 1:00 [ 4575.404891] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 0 in progress, and 0 evicted) to recover in 0:59 [ 4578.175181] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4580.507983] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 0 in progress, and 0 evicted) to recover in 0:53 [ 4581.162807] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 4581.179188] Lustre: Skipped 1 previous similar message [ 4581.187870] LustreError: lustre-OST0000-osc-MDT0000: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 4581.204964] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 0@lo (at 0@lo) [ 4581.208500] Lustre: Skipped 1 previous similar message [ 4581.223508] LustreError: 8437:0:(tgt_handler.c:534:tgt_filter_recovery_request()) @@@ not permitted during recovery req@ffff8ab376778000 x1876424820773504/t0(0) o13->lustre-MDT0000-mdtlov_UUID@0@lo:128/0 lens 224/0 e 0 to 0 dl 1789502638 ref 1 fl Interpret:/200/ffffffff rc 0/-1 job:'osp-pre-0-0.0' uid:0 gid:0 projid:4294967295 [ 4581.232352] LustreError: lustre-OST0000-osc-MDT0000: operation ost_get_info to node 0@lo failed: rc = -11 [ 4581.269744] LustreError: 8437:0:(tgt_handler.c:534:tgt_filter_recovery_request()) Skipped 1 previous similar message [ 4581.299258] LustreError: Skipped 1 previous similar message [ 4581.980480] LustreError: lustre-OST0000-osc-MDT0001: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 4581.991569] LustreError: 42582:0:(tgt_handler.c:534:tgt_filter_recovery_request()) @@@ not permitted during recovery req@ffff8ab248fce680 x1876424820774016/t0(0) o7->lustre-MDT0001-mdtlov_UUID@0@lo:128/0 lens 264/0 e 0 to 0 dl 1789502638 ref 1 fl Interpret:/200/ffffffff rc 0/-1 job:'osp-pre-0-1.0' uid:0 gid:0 projid:4294967295 [ 4582.008903] LustreError: 42582:0:(tgt_handler.c:534:tgt_filter_recovery_request()) Skipped 1 previous similar message [ 4585.436269] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 4585.626128] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 2 in progress, and 0 evicted) to recover in 0:58 [ 4590.745079] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 2 in progress, and 0 evicted) to recover in 0:53 [ 4600.993440] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 2 in progress, and 0 evicted) to recover in 0:43 [ 4601.023523] Lustre: Skipped 1 previous similar message [ 4621.474299] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 2 in progress, and 0 evicted) to recover in 0:23 [ 4621.493456] Lustre: Skipped 3 previous similar messages [ 4644.500318] Lustre: lustre-OST0000: recovery is timed out, evict stale exports [ 4644.521602] Lustre: 132332:0:(genops.c:1600:class_disconnect_stale_exports()) lustre-OST0000: disconnect stale client e4fb8066-e18f-4cab-963a-68940383e459@ [ 4644.558236] Lustre: lustre-OST0000: disconnecting 1 stale clients [ 4657.307057] Lustre: lustre-OST0000: Denying connection for new client 0d5d2d2f-8bf1-4154-9e47-62442b189ee6 (at 192.168.203.44@tcp), waiting for 3 known clients (0 recovered, 2 in progress, and 1 evicted) to recover in 0:17 [ 4657.332048] Lustre: Skipped 6 previous similar messages [ 4674.500137] Lustre: lustre-OST0000: recovery is timed out, evict stale exports [ 4674.512190] Lustre: 132332:0:(genops.c:1600:class_disconnect_stale_exports()) lustre-OST0000: disconnect stale client lustre-MDT0000-mdtlov_UUID@0@lo [ 4674.533747] Lustre: lustre-OST0000: disconnecting 2 stale clients [ 4674.555038] Lustre: lustre-OST0000: Recovery over after 1:40, of 3 clients 0 recovered and 3 were evicted. [ 4678.634235] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 4678.667425] Lustre: Skipped 1 previous similar message [ 4678.670816] LustreError: lustre-OST0000-osc-MDT0000: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 4678.698686] LustreError: lustre-OST0000-osc-MDT0001: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 4678.702166] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 0@lo (at 0@lo) [ 4678.729840] Lustre: Skipped 1 previous similar message [ 4680.720692] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing wait_import_state FULL os[cp].lustre-OST0000-osc-MDT0000.ost_server_uuid 50 [ 4734.850736] Lustre: DEBUG MARKER: rpc test_16c: @@@@@@ FAIL: can't put import for os[cp].lustre-OST0000-osc-MDT0000.ost_server_uuid into FULL state after 50 sec, have [ 4736.087301] Lustre: DEBUG MARKER: oleg344-server.virtnet: executing check_logdir /tmp/test_logs/1789502726 [ 4740.705777] Lustre: DEBUG MARKER: sanity-quota test_16c: @@@@@@ FAIL: mds1: import is not in FULL state after 50