[ 2887.649396] Lustre: *** cfs_fail_loc=513, val=0*** [ 2888.340316] Lustre: *** cfs_fail_loc=513, val=0*** [ 2888.342282] Lustre: Skipped 6 previous similar messages [ 2889.032464] LustreError: 5537:0:(service.c:2127:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1846868950205952 [ 2890.210849] Lustre: *** cfs_fail_loc=513, val=0*** [ 2890.212279] Lustre: Skipped 15 previous similar messages [ 2892.768718] Lustre: *** cfs_fail_loc=513, val=0*** [ 2895.328723] Lustre: 35294:0:(client.c:2295:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1761314288/real 1761314288] req@0000000041369206 x1846868950205952/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1761314295 ref 2 fl Rpc:XNQr/0/ffffffff rc 0/-1 job:'ll_ost_io00_006.0' [ 2895.366697] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2895.383468] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2895.409767] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to (at 0@lo) [ 2895.525525] LustreError: 5537:0:(service.c:2127:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1846868950206912 [ 2897.888854] Lustre: *** cfs_fail_loc=513, val=0*** [ 2897.890533] Lustre: Skipped 15 previous similar messages [ 2902.944217] Lustre: 3016:0:(client.c:2295:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1761314295/real 1761314295] req@000000008ed306ff x1846868950206912/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1761314302 ref 1 fl Rpc:XNQr/0/ffffffff rc 0/-1 job:'ll_ost_io00_006.0' [ 2902.981997] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2903.011755] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2903.020630] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to (at 0@lo) [ 2905.037069] LustreError: 5537:0:(service.c:2127:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1846868950208000 [ 2908.131120] Lustre: *** cfs_fail_loc=513, val=0*** [ 2908.134151] Lustre: Skipped 27 previous similar messages [ 2912.226059] Lustre: 35294:0:(client.c:2295:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1761314304/real 1761314304] req@00000000237ec0f3 x1846868950208000/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1761314311 ref 2 fl Rpc:XNQr/0/ffffffff rc 0/-1 job:'ll_ost_io00_006.0' [ 2912.251881] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2912.280277] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2912.285267] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to (at 0@lo) [ 2915.334915] LustreError: 5537:0:(service.c:2127:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1846868950209088 [ 2922.464510] Lustre: 35294:0:(client.c:2295:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1761314315/real 1761314315] req@00000000013494be x1846868950209088/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1761314322 ref 2 fl Rpc:XNQr/0/ffffffff rc 0/-1 job:'ll_ost_io00_006.0' [ 2922.524882] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2922.562275] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2922.574971] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to (at 0@lo) [ 2926.048938] Lustre: *** cfs_fail_loc=513, val=0*** [ 2926.050939] Lustre: Skipped 39 previous similar messages [ 2926.603912] LustreError: 5537:0:(service.c:2127:ptlrpc_server_handle_req_in()) drop incoming rpc opc 601, x1846868950210368 [ 2933.730039] Lustre: 35294:0:(client.c:2295:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1761314326/real 1761314326] req@00000000de432192 x1846868950210368/t0(0) o601->lustre-MDT0000-lwp-OST0000@0@lo:23/10 lens 336/336 e 0 to 1 dl 1761314333 ref 2 fl Rpc:XNQr/0/ffffffff rc 0/-1 job:'ll_ost_io00_006.0' [ 2933.786989] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 2933.814520] Lustre: lustre-MDT0000: Client lustre-MDT0000-lwp-OST0000_UUID (at 0@lo) reconnecting [ 2933.832381] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to (at 0@lo) [ 2993.293684] Lustre: DEBUG MARKER: == sanity-quota test 7a: Quota reintegration (global index) ========================================================== 09:59:52 (1761314392) [ 3022.028639] Lustre: Failing over lustre-OST0000 [ 3024.215230] Lustre: server umount lustre-OST0000 complete [ 3024.457970] LustreError: 11-0: lustre-OST0000-osc-MDT0000: operation ost_setattr to node 0@lo failed: rc = -107 [ 3024.461141] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3024.475482] LustreError: 137-5: lustre-OST0000_UUID: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3025.892985] LustreError: 137-5: lustre-OST0000_UUID: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3030.448309] LustreError: 68093:0:(qsd_reint.c:635:qqi_reint_delayed()) lustre-MDT0000: Delaying reintegration for qtype:0 until pending updates are flushed. [ 3030.458381] LustreError: 68093:0:(qsd_reint.c:635:qqi_reint_delayed()) Skipped 8 previous similar messages [ 3031.012530] LustreError: 137-5: lustre-OST0000_UUID: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3031.029981] LustreError: Skipped 1 previous similar message [ 3036.136258] LustreError: 137-5: lustre-OST0000_UUID: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3036.144857] LustreError: Skipped 1 previous similar message [ 3036.392240] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 3036.402384] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 3037.493942] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 3038.025728] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.202.130@tcp (at 0@lo) [ 3038.025739] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 3038.045623] Lustre: lustre-OST0000: deleting orphan objects from 0x0:137 to 0x0:161 [ 3039.754333] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all 8 [ 3047.118617] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3051.826984] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3056.596948] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3061.788400] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3068.257723] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3074.617053] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3114.643757] Lustre: Failing over lustre-OST0000 [ 3114.739928] Lustre: server umount lustre-OST0000 complete [ 3114.984770] LustreError: 11-0: lustre-OST0000-osc-MDT0000: operation ost_statfs to node 0@lo failed: rc = -107 [ 3114.996985] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3115.024768] LustreError: 137-5: lustre-OST0000_UUID: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3122.124623] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 3122.140187] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 3124.091840] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 1 client reconnects [ 3124.164196] Lustre: lustre-OST0000: Recovery over after 0:01, of 1 clients 1 recovered and 0 were evicted. [ 3124.166438] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.202.130@tcp (at 0@lo) [ 3124.223807] Lustre: lustre-OST0000: deleting orphan objects from 0x0:137 to 0x0:193 [ 3126.282402] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all 8 [ 3134.155554] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3140.356070] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3145.800845] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3150.169739] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3156.641860] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3161.310390] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3207.054475] Lustre: DEBUG MARKER: == sanity-quota test 7b: Quota reintegration (slave index) ========================================================== 10:03:26 (1761314606) [ 3236.284406] LustreError: 76723:0:(qsd_reint.c:635:qqi_reint_delayed()) lustre-MDT0000: Delaying reintegration for qtype:0 until pending updates are flushed. [ 3236.302613] LustreError: 76723:0:(qsd_reint.c:635:qqi_reint_delayed()) Skipped 8 previous similar messages [ 3238.782905] Lustre: *** cfs_fail_loc=a02, val=0*** [ 3244.278379] Lustre: Failing over lustre-OST0000 [ 3244.346815] Lustre: server umount lustre-OST0000 complete [ 3245.031610] Lustre: lustre-OST0000-osc-MDT0000: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3245.038197] LustreError: 137-5: lustre-OST0000_UUID: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 3245.044043] LustreError: Skipped 1 previous similar message [ 3251.102617] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 3251.118904] Lustre: lustre-OST0000: in recovery but waiting for the first client to connect [ 3251.329824] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 3252.658470] Lustre: lustre-OST0000: Recovery over after 0:01, of 2 clients 2 recovered and 0 were evicted. [ 3252.671914] Lustre: lustre-OST0000-osc-MDT0000: Connection restored to 192.168.202.130@tcp (at 0@lo) [ 3252.694559] Lustre: lustre-OST0000: deleting orphan objects from 0x0:195 to 0x0:225 [ 3254.672601] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all 8 [ 3261.614824] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3266.061526] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3270.883947] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3275.923254] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3283.112675] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3288.474641] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3353.868350] Lustre: DEBUG MARKER: == sanity-quota test 7c: Quota reintegration (restart mds during reintegration) ========================================================== 10:05:53 (1761314753) [ 3375.106661] LustreError: 82055:0:(qsd_reint.c:635:qqi_reint_delayed()) lustre-MDT0000: Delaying reintegration for qtype:0 until pending updates are flushed. [ 3375.110853] LustreError: 82055:0:(qsd_reint.c:635:qqi_reint_delayed()) Skipped 11 previous similar messages [ 3377.561237] LustreError: 82221:0:(fail.c:138:__cfs_fail_timeout_set()) cfs_fail_timeout id a03 sleeping for 10000ms [ 3377.571918] LustreError: 82221:0:(fail.c:138:__cfs_fail_timeout_set()) Skipped 5 previous similar messages [ 3379.110386] Lustre: Failing over lustre-MDT0000 [ 3379.172234] Lustre: lustre-MDT0000-lwp-OST0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3379.198764] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 3379.203330] Lustre: Skipped 1 previous similar message [ 3379.531284] Lustre: server umount lustre-MDT0000 complete [ 3382.552335] LustreError: 82224:0:(fail.c:144:__cfs_fail_timeout_set()) cfs_fail_timeout interrupted [ 3387.703708] LustreError: 166-1: MGC192.168.202.130@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 3387.723360] Lustre: Evicted from MGS (at 192.168.202.130@tcp) after server handle changed from 0x82e77f4105dbc6e1 to 0x82e77f4105dc5255 [ 3387.737221] Lustre: MGC192.168.202.130@tcp: Connection restored to 192.168.202.130@tcp (at 0@lo) [ 3388.136207] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 3388.246426] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 3392.343695] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all 8 [ 3395.776595] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 1 client reconnects [ 3395.847746] Lustre: lustre-MDT0000: Recovery over after 0:01, of 1 clients 1 recovered and 0 were evicted. [ 3395.894648] Lustre: lustre-OST0001: deleting orphan objects from 0x0:115 to 0x0:161 [ 3396.261561] LustreError: 3015:0:(ldlm_resource.c:1127:ldlm_resource_complain()) lustre-MDT0000-lwp-OST0001: namespace resource [0x200000006:0x2020000:0x0].0x0 (00000000b82b49cb) refcount nonzero (1) after lock cleanup; forcing cleanup. [ 3399.367322] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3499.810120] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3504.429277] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3509.696882] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3516.640743] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3520.881701] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3563.535218] Lustre: DEBUG MARKER: == sanity-quota test 7d: Quota reintegration (Transfer index in multiple bulks) ========================================================== 10:09:22 (1761314962) [ 3580.599230] LustreError: 89500:0:(qsd_reint.c:635:qqi_reint_delayed()) lustre-OST0001: Delaying reintegration for qtype:0 until pending updates are flushed. [ 3580.604723] LustreError: 89500:0:(qsd_reint.c:635:qqi_reint_delayed()) Skipped 11 previous similar messages [ 3587.840792] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0000.recovery_status 1475 [ 3592.334487] Lustre: DEBUG MARKER: oleg230-server.virtnet: executing _wait_recovery_complete *.lustre-OST0001.recovery_status 1475 [ 3667.184353] Lustre: DEBUG MARKER: == sanity-quota test 7e: Quota reintegration (inode limits) ========================================================== 10:11:06 (1761315066) [ 3668.516811] Lustre: DEBUG MARKER: SKIP: sanity-quota test_7e needs >= 2 MDTs [ 3670.190929] Lustre: DEBUG MARKER: == sanity-quota test 8: Run dbench with quota enabled ==== 10:11:09 (1761315069) [ 3688.240552] LustreError: 91896:0:(qsd_reint.c:635:qqi_reint_delayed()) lustre-OST0001: Delaying reintegration for qtype:1 until pending updates are flushed. [ 3688.247322] LustreError: 91896:0:(qsd_reint.c:635:qqi_reint_delayed()) Skipped 3 previous similar messages [ 3876.165064] Lustre: DEBUG MARKER: SKIP: sanity-quota test_9 skipping SLOW test 9 [ 3878.313577] Lustre: DEBUG MARKER: == sanity-quota test 10: Test quota for root user ======== 10:14:37 (1761315277) [ 3928.001872] Lustre: DEBUG MARKER: == sanity-quota test 11: Chown/chgrp ignores quota ======= 10:15:26 (1761315326) [ 3987.546925] Lustre: DEBUG MARKER: SKIP: sanity-quota test_12a skipping SLOW test 12a [ 3988.898851] Lustre: DEBUG MARKER: == sanity-quota test 12b: Inode quota rebalancing ======== 10:16:28 (1761315388) [ 3990.141408] Lustre: DEBUG MARKER: SKIP: sanity-quota test_12b needs >= 2 MDTs [ 3991.868300] Lustre: DEBUG MARKER: == sanity-quota test 13: Cancel per-ID lock in the LRU list ========================================================== 10:16:31 (1761315391) [ 4033.864263] Lustre: DEBUG MARKER: sanity-quota test_13: @@@@@@ FAIL: lock count(0) isn't 1