[ 3736.260905] Lustre: Failing over lustre-MDT0000 [ 3738.592798] Lustre: server umount lustre-MDT0000 complete [ 3752.383760] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 3752.388120] LDISKFS-fs (dm-0): recovery complete [ 3752.396914] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3752.768878] Lustre: lustre-MDT0000: Aborting client recovery [ 3752.772858] LustreError: 107286:0:(ldlm_lib.c:3004:target_stop_recovery_thread()) lustre-MDT0000: Aborting recovery [ 3752.780800] Lustre: 107318:0:(ldlm_lib.c:2404:target_recovery_overseer()) recovery is aborted, evict exports in recovery [ 3752.784197] Lustre: 107318:0:(ldlm_lib.c:2404:target_recovery_overseer()) Skipped 2 previous similar messages [ 3752.788300] Lustre: 107318:0:(genops.c:1622:class_disconnect_stale_exports()) lustre-MDT0000: disconnect stale client lustre-MDT0001-mdtlov_UUID@ [ 3752.795604] Lustre: 107318:0:(genops.c:1622:class_disconnect_stale_exports()) Skipped 1 previous similar message [ 3752.801420] Lustre: lustre-MDT0000: disconnecting 2 stale clients [ 3752.806621] Lustre: lustre-MDT0000-osd: cancel update llog [0x200017b00:0x1:0x0] [ 3752.817212] Lustre: lustre-MDT0001-osp-MDT0000: cancel update llog [0x2400007ea:0x1:0x0] [ 3752.863826] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:1543 to 0x2c0000401:1633) [ 3752.872688] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:1571 to 0x280000401:1633) [ 3757.510167] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3758.069265] LustreError: lustre-MDT0000-osp-MDT0001: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 3780.786334] Lustre: DEBUG MARKER: == replay-single test 38: test recovery from unlink llog (test llog_gen_rec) ========================================================== 17:28:30 (1786829310) [ 3814.090897] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 3816.009120] Lustre: Failing over lustre-MDT0000 [ 3816.404360] Lustre: server umount lustre-MDT0000 complete [ 3838.187295] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 3838.188690] LDISKFS-fs (dm-0): recovery complete [ 3838.195300] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3850.863138] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2034 to 0x280000401:2049) [ 3850.870231] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2034 to 0x2c0000401:2049) [ 3851.337431] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3861.527323] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 3863.418347] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 3886.539321] Lustre: DEBUG MARKER: == replay-single test 39: test recovery from unlink llog (test llog_gen_rec) ========================================================== 17:30:16 (1786829416) [ 3912.750651] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 3923.848719] Lustre: Failing over lustre-MDT0000 [ 3924.377779] Lustre: server umount lustre-MDT0000 complete [ 3945.940046] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 3945.945389] LDISKFS-fs (dm-0): recovery complete [ 3945.950860] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3953.120798] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c875e68700 x1873622492023808/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 3953.144743] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 46 previous similar messages [ 3957.393643] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3962.245272] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2450 to 0x280000401:2465) [ 3962.246392] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2450 to 0x2c0000401:2465) [ 3968.229203] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 3969.765866] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 3991.075464] Lustre: DEBUG MARKER: == replay-single test 41: read from a valid osc while other oscs are invalid ========================================================== 17:32:01 (1786829521) [ 3993.225820] Lustre: setting import lustre-OST0001_UUID INACTIVE by administrator request [ 3994.283989] Lustre: lustre-OST0001-osc-MDT0000: Connection to lustre-OST0001 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3994.296487] Lustre: Skipped 37 previous similar messages [ 3994.303344] Lustre: lustre-OST0001: Client lustre-MDT0000-mdtlov_UUID (at 0@lo) reconnecting [ 3994.312146] LustreError: lustre-OST0001-osc-MDT0000: This client was evicted by lustre-OST0001; in progress operations using this service will fail. [ 4001.378505] Lustre: DEBUG MARKER: == replay-single test 42: recovery after ost failure ===== 17:32:11 (1786829531) [ 4027.477965] Lustre: DEBUG MARKER: ost1 REPLAY BARRIER on lustre-OST0000 [ 4042.593734] Lustre: Failing over lustre-OST0000 [ 4042.771665] Lustre: server umount lustre-OST0000 complete [ 4044.259151] LustreError: 8435:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-OST0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 4044.286309] LustreError: 8435:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 182 previous similar messages [ 4066.817867] LDISKFS-fs (dm-2): 3 truncates cleaned up [ 4066.820207] LDISKFS-fs (dm-2): recovery complete [ 4066.832825] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 4068.469125] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 4068.495624] Lustre: Skipped 4 previous similar messages [ 4069.231659] Lustre: lustre-OST0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 4069.246378] Lustre: Skipped 4 previous similar messages [ 4074.347367] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4135.098385] Lustre: DEBUG MARKER: == replay-single test 43: mds osc import failure during recovery; don't LBUG ========================================================== 17:34:25 (1786829665) [ 4143.156519] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 4145.755474] Lustre: Failing over lustre-MDT0000 [ 4145.995958] Lustre: server umount lustre-MDT0000 complete [ 4165.605112] Lustre: 3647:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786829680/real 1786829680] req@ffff99c875de5500 x1873622492335488/t0(0) o400->MGC192.168.201.120@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786829696 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4165.638032] Lustre: 3647:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 3 previous similar messages [ 4165.655228] LustreError: MGC192.168.201.120@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 4165.663568] LustreError: Skipped 9 previous similar messages [ 4170.718230] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 4170.722034] LDISKFS-fs (dm-0): recovery complete [ 4170.734449] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4176.237334] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 4176.241744] Lustre: Skipped 7 previous similar messages [ 4176.287956] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 4176.290907] Lustre: Skipped 18 previous similar messages [ 4181.237985] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4181.570970] Lustre: *** cfs_fail_loc=204, val=2147483648*** [ 4181.574176] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2867 to 0x280000401:2913) [ 4191.252578] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4192.995539] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4196.831267] LustreError: 116738:0:(osp_precreate.c:974:osp_precreate_cleanup_orphans()) lustre-OST0001-osc-MDT0000: cannot cleanup orphans: rc = -11 [ 4196.833259] Lustre: lustre-OST0001: Client lustre-MDT0000-mdtlov_UUID (at 0@lo) reconnecting [ 4197.855648] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2866 to 0x2c0000401:2881) [ 4212.388722] Lustre: DEBUG MARKER: == replay-single test 44a: race in target handle connect ========================================================== 17:35:42 (1786829742) [ 4216.806620] LustreError: 6520:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4221.922459] LustreError: 6520:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4221.932141] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4221.983309] LustreError: 43909:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 waking [ 4223.820451] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4229.088369] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4229.099558] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4230.827795] LustreError: 6522:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4236.255288] LustreError: 6522:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4236.259492] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4237.560264] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4242.913487] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4244.635192] LustreError: 11771:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4250.079351] LustreError: 11771:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4250.082282] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4250.085522] Lustre: Skipped 1 previous similar message [ 4258.288048] LustreError: 11771:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4258.295188] LustreError: 11771:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 1 previous similar message [ 4263.391151] LustreError: 11771:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4263.393683] LustreError: 11771:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 1 previous similar message [ 4270.047560] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4270.057924] Lustre: Skipped 2 previous similar messages [ 4279.850872] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4279.859174] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 2 previous similar messages [ 4284.897419] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4284.908850] LustreError: 6521:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 2 previous similar messages [ 4294.115721] Lustre: DEBUG MARKER: == replay-single test 44b: race in target handle connect ========================================================== 17:37:04 (1786829824) [ 4295.747344] LustreError: 11771:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4306.469634] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4310.499920] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4315.700177] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4318.220992] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4325.934228] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4325.940787] Lustre: Skipped 1 previous similar message [ 4335.799164] LustreError: 11771:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4335.807048] Lustre: 11771:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff99c85e43ed80 x1873622468399104/t0(0) o38->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:0/0 lens 520/416 e 0 to 0 dl 1786829847 ref 1 fl Complete:H/200/0 rc 0/0 job:'lctl.0' uid:0 gid:0 projid:4294967295 [ 4336.167203] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4336.179995] Lustre: Skipped 3 previous similar messages [ 4336.184366] LustreError: 67185:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4361.764106] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4361.771719] Lustre: Skipped 1 previous similar message [ 4376.191130] LustreError: 67185:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4376.199643] Lustre: 67185:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff99c875de1c00 x1873622468403200/t0(0) o38->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:0/0 lens 520/416 e 0 to 0 dl 1786829887 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4377.130482] LustreError: 11771:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4401.706529] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4401.722284] Lustre: Skipped 3 previous similar messages [ 4417.159900] LustreError: 11771:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4417.162870] Lustre: 11771:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff99c875d4f800 x1873622468406400/t0(0) o38->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:0/0 lens 520/416 e 0 to 0 dl 1786829928 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4422.191916] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4422.207945] Lustre: Skipped 1 previous similar message [ 4422.215418] LustreError: 11771:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4447.791159] Lustre: lustre-MDT0000: Export ffff99c87f399000 already connecting from 192.168.201.20@tcp [ 4447.797462] Lustre: Skipped 4 previous similar messages [ 4462.319449] LustreError: 11771:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4462.330936] Lustre: 11771:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff99c84fbff100 x1873622468409984/t0(0) o38->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:0/0 lens 520/416 e 0 to 0 dl 1786829973 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4463.143038] LustreError: 17699:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4503.199216] LustreError: 17699:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4503.205498] Lustre: 17699:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff99c875c01880 x1873622468413184/t0(0) o38->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:0/0 lens 520/416 e 0 to 0 dl 1786830014 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4508.215954] LustreError: 6522:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4523.488352] LustreError: 6522:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout interrupted [ 4527.726421] Lustre: DEBUG MARKER: == replay-single test 44c: race in target handle connect ========================================================== 17:40:58 (1786830058) [ 4534.559120] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 4538.130664] Lustre: Failing over lustre-MDT0000 [ 4538.451161] Lustre: server umount lustre-MDT0000 complete [ 4549.954283] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 4549.956648] LDISKFS-fs (dm-0): recovery complete [ 4549.972845] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4550.312230] Lustre: *** cfs_fail_loc=712, val=0*** [ 4550.317498] LustreError: 43904:0:(service.c:1394:ptlrpc_check_req()) @@@ Invalid replay without recovery req@ffff99c85e747800 x1873622492531456/t0(0) o400->lustre-MDT0000-mdtlov_UUID@0@lo:0/0 lens 224/0 e 0 to 0 dl 0 ref 1 fl New:/2c0/ffffffff rc 0/-1 job:'ptlrpcd_rcv.0' uid:0 gid:0 projid:4294967295 [ 4550.492992] Lustre: lustre-MDT0000: Aborting client recovery [ 4550.495804] LustreError: 121589:0:(ldlm_lib.c:3004:target_stop_recovery_thread()) lustre-MDT0000: Aborting recovery [ 4550.513071] Lustre: 121621:0:(ldlm_lib.c:2404:target_recovery_overseer()) recovery is aborted, evict exports in recovery [ 4550.524891] Lustre: 121621:0:(ldlm_lib.c:2404:target_recovery_overseer()) Skipped 2 previous similar messages [ 4550.531312] Lustre: 121621:0:(genops.c:1622:class_disconnect_stale_exports()) lustre-MDT0000: disconnect stale client lustre-MDT0001-mdtlov_UUID@ [ 4550.536301] Lustre: 121621:0:(genops.c:1622:class_disconnect_stale_exports()) Skipped 1 previous similar message [ 4550.542752] Lustre: lustre-MDT0000: disconnecting 2 stale clients [ 4550.548900] Lustre: lustre-MDT0000-osd: cancel update llog [0x2000182d0:0x1:0x0] [ 4550.562874] Lustre: lustre-MDT0001-osp-MDT0000: cancel update llog [0x2400007eb:0x1:0x0] [ 4550.602654] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2866 to 0x2c0000401:2913) [ 4550.606230] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2867 to 0x280000401:2945) [ 4554.866238] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4555.765262] LustreError: lustre-MDT0000-osp-MDT0001: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 4555.778395] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 4555.789305] Lustre: Skipped 23 previous similar messages [ 4571.072774] Lustre: Failing over lustre-MDT0000 [ 4571.104579] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 4571.109827] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 4571.112676] LustreError: Skipped 4 previous similar messages [ 4571.129708] Lustre: Skipped 3 previous similar messages [ 4573.371929] Lustre: server umount lustre-MDT0000 complete [ 4593.556105] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4606.488489] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4607.525466] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2867 to 0x280000401:2977) [ 4607.525738] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2866 to 0x2c0000401:2945) [ 4617.602962] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4619.316942] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4628.400061] Lustre: DEBUG MARKER: == replay-single test 45: Handle failed close ============ 17:42:38 (1786830158) [ 4628.544782] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 4628.561833] Lustre: Skipped 2 previous similar messages [ 4636.906511] Lustre: DEBUG MARKER: == replay-single test 46: Don't leak file handle after open resend (3325) ========================================================== 17:42:47 (1786830167) [ 4637.897691] Lustre: *** cfs_fail_loc=122, val=2147483648*** [ 4637.904474] LustreError: 6530:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c85e6d8e00 x1873622468477440/t0(0) o700->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:370/0 lens 264/248 e 0 to 0 dl 1786830180 ref 1 fl Interpret:/200/0 rc 0/0 job:'touch.0' uid:0 gid:0 projid:4294967295 [ 4658.206329] Lustre: Failing over lustre-MDT0000 [ 4658.559501] Lustre: server umount lustre-MDT0000 complete [ 4658.660362] Lustre: lustre-MDT0000-lwp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 4658.669491] LustreError: 6526:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 4658.670074] Lustre: Skipped 17 previous similar messages [ 4658.692354] LustreError: 6526:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 101 previous similar messages [ 4677.328318] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4679.651776] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 4679.657337] Lustre: Skipped 2 previous similar messages [ 4683.276202] Lustre: lustre-MDT0000: Recovery over after 0:04, of 2 clients 2 recovered and 0 were evicted. [ 4683.289821] Lustre: Skipped 2 previous similar messages [ 4683.372522] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2947 to 0x2c0000401:2977) [ 4683.375631] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:2979 to 0x280000401:3009) [ 4683.701925] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4694.961501] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4696.879390] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4710.473262] Lustre: DEBUG MARKER: == replay-single test 47: MDS->OSC failure during precreate cleanup (2824) ========================================================== 17:44:00 (1786830240) [ 4713.588585] Lustre: Failing over lustre-OST0000 [ 4713.707380] Lustre: server umount lustre-OST0000 complete [ 4732.864536] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 4738.921633] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4749.967655] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 4752.124765] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 4825.854307] Lustre: DEBUG MARKER: == replay-single test 48: MDS->OSC failure during precreate cleanup (2824) ========================================================== 17:45:55 (1786830355) [ 4834.990613] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 4838.101242] Lustre: Failing over lustre-MDT0000 [ 4838.560949] Lustre: server umount lustre-MDT0000 complete [ 4857.314781] Lustre: 3644:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786830372/real 1786830372] req@ffff99c879910000 x1873622492690304/t0(0) o400->MGC192.168.201.120@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786830388 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4857.359918] Lustre: 3644:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 4857.381471] LustreError: MGC192.168.201.120@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 4857.393316] LustreError: Skipped 3 previous similar messages [ 4861.225857] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 4861.228460] LDISKFS-fs (dm-0): recovery complete [ 4861.235289] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4867.560082] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xf03e36cc1e26fc03 [ 4867.820810] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 4867.833950] Lustre: Skipped 4 previous similar messages [ 4867.867658] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 4867.884026] Lustre: Skipped 6 previous similar messages [ 4871.585254] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4873.863827] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3030 to 0x280000401:3073) [ 4873.895717] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2998 to 0x2c0000401:3041) [ 4945.559090] Lustre: DEBUG MARKER: == replay-single test 50: Double OSC recovery, don't LASSERT (3812) ========================================================== 17:47:56 (1786830476) [ 4947.379392] Lustre: lustre-OST0000: Client lustre-MDT0000-mdtlov_UUID (at 0@lo) reconnecting [ 4947.382266] Lustre: Skipped 2 previous similar messages [ 4959.684370] Lustre: DEBUG MARKER: == replay-single test 52: time out lock replay (3764) ==== 17:48:10 (1786830490) [ 4962.633896] Lustre: Failing over lustre-MDT0000 [ 4962.853303] Lustre: server umount lustre-MDT0000 complete [ 4980.540961] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4990.947361] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c74228c700 x1873622492766976/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4991.014146] Lustre: lustre-MDT0000: Not available for connect from 192.168.201.20@tcp (not set up) [ 4995.546943] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4996.594099] Lustre: *** cfs_fail_loc=157, val=2147483648*** [ 4996.604340] LustreError: 129351:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c875e54380 x1873622468612608/t0(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:13/0 lens 328/344 e 0 to 0 dl 1786830578 ref 1 fl Complete:/240/0 rc 0/0 job:'ldlm_lock_repla.0' uid:0 gid:0 projid:4294967295 [ 5052.966949] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnected, waiting for 2 clients in recovery for 1:06 [ 5053.002690] Lustre: 129351:0:(ldlm_lib.c:2085:extend_recovery_timer()) lustre-MDT0000: extended recovery timer reached hard limit: 180, extend: 1 [ 5053.081978] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3052 to 0x2c0000401:3073) [ 5053.082280] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3085 to 0x280000401:3105) [ 5059.494704] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5061.489211] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5071.683666] Lustre: DEBUG MARKER: == replay-single test 53a: |X| close request while two MDC requests in flight ========================================================== 17:50:02 (1786830602) [ 5074.133865] Lustre: *** cfs_fail_loc=115, val=2147483648*** [ 5083.509211] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5085.409974] Lustre: Failing over lustre-MDT0000 [ 5085.612861] Lustre: server umount lustre-MDT0000 complete [ 5109.047972] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5109.053086] LDISKFS-fs (dm-0): recovery complete [ 5109.061812] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5115.368418] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c87f3c3100 x1873622492826112/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5121.080042] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3075 to 0x2c0000401:3105) [ 5121.080403] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3085 to 0x280000401:3137) [ 5121.183824] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5131.505877] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5133.155124] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5143.692561] Lustre: DEBUG MARKER: == replay-single test 53b: |X| open request while two MDC requests in flight ========================================================== 17:51:14 (1786830674) [ 5145.180279] Lustre: *** cfs_fail_loc=107, val=2147483648*** [ 5154.807612] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5156.984565] Lustre: Failing over lustre-MDT0000 [ 5157.366588] Lustre: server umount lustre-MDT0000 complete [ 5181.798811] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5181.805266] LDISKFS-fs (dm-0): recovery complete [ 5181.824034] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5193.504398] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5194.221060] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 5194.229329] Lustre: Skipped 27 previous similar messages [ 5194.311688] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3139 to 0x280000401:3169) [ 5194.314357] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3075 to 0x2c0000401:3137) [ 5203.810210] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5205.444960] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5214.488627] Lustre: DEBUG MARKER: == replay-single test 53c: |X| open request and close request while two MDC requests in flight ========================================================== 17:52:24 (1786830744) [ 5215.794867] Lustre: *** cfs_fail_loc=107, val=2147483648*** [ 5217.941691] Lustre: *** cfs_fail_loc=115, val=2147483648*** [ 5224.793909] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5226.764594] Lustre: Failing over lustre-MDT0000 [ 5226.954598] Lustre: server umount lustre-MDT0000 complete [ 5230.054352] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 5230.061076] LustreError: Skipped 2 previous similar messages [ 5250.160217] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5250.162291] LDISKFS-fs (dm-0): recovery complete [ 5250.167450] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5262.408185] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3139 to 0x280000401:3201) [ 5262.409543] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3139 to 0x2c0000401:3169) [ 5263.320453] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5276.561278] Lustre: DEBUG MARKER: == replay-single test 53d: close reply while two MDC requests in flight ========================================================== 17:53:26 (1786830806) [ 5278.949908] Lustre: *** cfs_fail_loc=13b, val=315*** [ 5278.956113] Lustre: *** cfs_fail_loc=13b, val=2147483648*** [ 5278.957865] LustreError: 6523:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c877d35f80 x1873622468678016/t257698037777(0) o35->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:256/0 lens 392/456 e 0 to 0 dl 1786830821 ref 1 fl Interpret:/600/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5282.312070] Lustre: Failing over lustre-MDT0000 [ 5282.784558] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 5282.800952] Lustre: Skipped 28 previous similar messages [ 5282.803379] LustreError: 6526:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 5282.831537] LustreError: 6526:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 200 previous similar messages [ 5282.971946] Lustre: server umount lustre-MDT0000 complete [ 5302.269769] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5305.392912] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 5305.409349] Lustre: Skipped 6 previous similar messages [ 5307.495568] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5307.891322] Lustre: lustre-MDT0000: Recovery over after 0:02, of 2 clients 2 recovered and 0 were evicted. [ 5307.905528] Lustre: Skipped 6 previous similar messages [ 5307.951677] Lustre: 43561:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c881af9500 x1873622468678016/t257698037777(0) o35->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:285/0 lens 392/456 e 0 to 0 dl 1786830850 ref 1 fl Interpret:/602/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5307.973034] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3139 to 0x2c0000401:3201) [ 5307.973238] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3203 to 0x280000401:3233) [ 5317.920287] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5319.993781] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5329.800033] Lustre: DEBUG MARKER: == replay-single test 53e: |X| open reply while two MDC requests in flight ========================================================== 17:54:19 (1786830859) [ 5331.009801] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 5331.016629] LustreError: 67185:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c8765ddf80 x1873622468694272/t261993005072(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:352/0 lens 504/448 e 0 to 0 dl 1786830917 ref 1 fl Interpret:/200/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5340.551660] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5342.432399] Lustre: Failing over lustre-MDT0000 [ 5342.691689] Lustre: server umount lustre-MDT0000 complete [ 5365.606633] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5365.608830] LDISKFS-fs (dm-0): recovery complete [ 5365.615926] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5375.631596] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5376.066759] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3203 to 0x280000401:3265) [ 5376.075332] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3203 to 0x2c0000401:3233) [ 5376.099128] Lustre: 6520:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c875c01500 x1873622468694272/t261993005072(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:397/0 lens 504/2880 e 0 to 0 dl 1786830962 ref 1 fl Interpret:/202/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5385.545695] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5387.252376] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5395.685610] Lustre: DEBUG MARKER: == replay-single test 53f: |X| open reply and close reply while two MDC requests in flight ========================================================== 17:55:26 (1786830926) [ 5396.939168] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 5396.943185] LustreError: 6521:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c875f89c00 x1873622468710656/t266287972368(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:418/0 lens 504/448 e 0 to 0 dl 1786830983 ref 1 fl Interpret:/200/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5398.907436] Lustre: *** cfs_fail_loc=13b, val=315*** [ 5406.006234] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5407.824831] Lustre: Failing over lustre-MDT0000 [ 5408.356855] Lustre: server umount lustre-MDT0000 complete [ 5431.716497] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5431.722130] LDISKFS-fs (dm-0): recovery complete [ 5431.748297] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5444.138888] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3203 to 0x2c0000401:3265) [ 5444.144487] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3267 to 0x280000401:3297) [ 5444.167184] Lustre: 6522:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c875ffca80 x1873622468710656/t266287972368(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:465/0 lens 504/2880 e 0 to 0 dl 1786831030 ref 1 fl Interpret:/202/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5444.191960] Lustre: 6522:0:(mdt_recovery.c:102:mdt_req_from_lrd()) Skipped 1 previous similar message [ 5444.343300] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5457.293773] Lustre: DEBUG MARKER: == replay-single test 53g: |X| drop open reply and close request while close and open are both in flight ========================================================== 17:56:27 (1786830987) [ 5458.609684] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 5458.611417] Lustre: Skipped 1 previous similar message [ 5458.621342] LustreError: 6521:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c84f975500 x1873622468726144/t270582939664(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:479/0 lens 504/448 e 0 to 0 dl 1786831044 ref 1 fl Interpret:/200/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5458.652506] LustreError: 6521:0:(ldlm_lib.c:3346:target_send_reply_msg()) Skipped 1 previous similar message [ 5460.664886] Lustre: *** cfs_fail_loc=115, val=2147483648*** [ 5469.645140] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5471.773672] Lustre: Failing over lustre-MDT0000 [ 5472.194585] Lustre: server umount lustre-MDT0000 complete [ 5491.168715] Lustre: 3644:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786831006/real 1786831006] req@ffff99c8708e2a00 x1873622493029888/t0(0) o400->MGC192.168.201.120@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786831022 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5491.242031] Lustre: 3644:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 6 previous similar messages [ 5491.257700] LustreError: MGC192.168.201.120@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 5491.290446] LustreError: Skipped 7 previous similar messages [ 5496.414940] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5496.423021] LDISKFS-fs (dm-0): recovery complete [ 5496.430502] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5501.113798] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 5501.119601] Lustre: Skipped 7 previous similar messages [ 5501.143484] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 5501.165576] Lustre: Skipped 7 previous similar messages [ 5505.859528] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5506.590122] Lustre: 6522:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c744738700 x1873622468726144/t270582939664(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:527/0 lens 504/2880 e 0 to 0 dl 1786831092 ref 1 fl Interpret:/202/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5506.607296] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3267 to 0x2c0000401:3297) [ 5506.612558] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3267 to 0x280000401:3329) [ 5517.270582] Lustre: DEBUG MARKER: == replay-single test 53h: open request and close reply while two MDC requests in flight ========================================================== 17:57:27 (1786831047) [ 5518.218365] Lustre: *** cfs_fail_loc=107, val=2147483648*** [ 5520.231206] Lustre: *** cfs_fail_loc=13b, val=315*** [ 5520.235705] Lustre: *** cfs_fail_loc=13b, val=2147483648*** [ 5520.244548] LustreError: 6523:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c8558dd050 x1873622468741248/t274877906960(0) o35->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:497/0 lens 392/456 e 0 to 0 dl 1786831062 ref 1 fl Interpret:/600/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5529.712531] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5531.584737] Lustre: Failing over lustre-MDT0000 [ 5531.863885] Lustre: server umount lustre-MDT0000 complete [ 5557.332836] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5557.338268] LDISKFS-fs (dm-0): recovery complete [ 5557.354067] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5563.449398] Lustre: 43561:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c879903b80 x1873622468741248/t274877906960(0) o35->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:540/0 lens 392/456 e 0 to 0 dl 1786831105 ref 1 fl Interpret:/602/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5563.478634] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3331 to 0x280000401:3361) [ 5563.488174] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3267 to 0x2c0000401:3329) [ 5564.419780] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5578.182355] Lustre: DEBUG MARKER: == replay-single test 55: let MDS_CHECK_RESENT return the original return code instead of 0 ========================================================== 17:58:28 (1786831108) [ 5579.348905] Lustre: *** cfs_fail_loc=12b, val=2147483991*** [ 5579.357365] LustreError: 6522:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c743d2c380 x1873622468754048/t279172874255(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:600/0 lens 664/608 e 0 to 0 dl 1786831165 ref 1 fl Interpret:/600/0 rc 301/0 job:'touch.0' uid:0 gid:0 projid:0 [ 5639.718254] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnecting [ 5639.742160] Lustre: Skipped 1 previous similar message [ 5639.764436] Lustre: 6521:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c84439c000 x1873622468754048/t279172874255(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:661/0 lens 664/3488 e 0 to 0 dl 1786831226 ref 1 fl Interpret:/602/0 rc 0/0 job:'touch.0' uid:0 gid:0 projid:0 [ 5647.292339] Lustre: DEBUG MARKER: == replay-single test 56: don't replay a symlink open request (3440) ========================================================== 17:59:37 (1786831177) [ 5656.260040] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5658.438214] Lustre: Failing over lustre-MDT0000 [ 5658.927808] Lustre: server umount lustre-MDT0000 complete [ 5686.398847] LDISKFS-fs (dm-0): 4 truncates cleaned up [ 5686.400956] LDISKFS-fs (dm-0): recovery complete [ 5686.417969] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5701.600548] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c875d4f100 x1873622493140224/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5706.845446] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5707.332142] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3331 to 0x2c0000401:3361) [ 5707.334065] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3331 to 0x280000401:3393) [ 5717.547805] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5719.017292] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5737.396265] Lustre: DEBUG MARKER: == replay-single test 57: test recovery from llog for setattr op ========================================================== 18:01:08 (1786831268) [ 5744.472363] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5746.225305] Lustre: Failing over lustre-MDT0000 [ 5746.515194] Lustre: server umount lustre-MDT0000 complete [ 5768.223816] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5768.225344] LDISKFS-fs (dm-0): recovery complete [ 5768.229871] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5778.739200] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5779.532946] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:3331 to 0x280000401:3425) [ 5779.534287] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3363 to 0x2c0000401:3393) [ 5786.911963] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5788.142445] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5793.775713] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 5805.822304] Lustre: DEBUG MARKER: == replay-single test 58a: test recovery from llog for setattr op (test llog_gen_rec) ========================================================== 18:02:15 (1786831335) [ 5851.240908] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5853.390485] Lustre: Failing over lustre-MDT0000 [ 5854.132995] Lustre: server umount lustre-MDT0000 complete [ 5856.238377] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 5856.253591] LustreError: Skipped 7 previous similar messages [ 5875.072181] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5875.074016] LDISKFS-fs (dm-0): recovery complete [ 5875.080313] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5882.340977] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xf03e36cc1e28a8a0 [ 5882.353394] Lustre: MGC192.168.201.120@tcp: Connection restored to 0@lo (at 0@lo) [ 5882.360449] Lustre: Skipped 35 previous similar messages [ 5882.409660] Lustre: lustre-MDT0000: Not available for connect from 192.168.201.20@tcp (not set up) [ 5886.560564] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5888.279531] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:4676 to 0x280000401:4705) [ 5888.280355] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4644 to 0x2c0000401:4673) [ 5894.436252] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5895.912651] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5954.368044] Lustre: DEBUG MARKER: == replay-single test 58b: test replay of setxattr op ==== 18:04:44 (1786831484) [ 5965.558450] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5967.736472] Lustre: Failing over lustre-MDT0000 [ 5968.123769] Lustre: server umount lustre-MDT0000 complete [ 5969.911600] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 5969.923225] Lustre: Skipped 29 previous similar messages [ 5969.929534] LustreError: 6521:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 5969.929544] LustreError: 6521:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 265 previous similar messages [ 5990.497899] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5990.503541] LDISKFS-fs (dm-0): recovery complete [ 5990.509527] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5996.000077] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c850a1b800 x1873622493942400/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5996.028091] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 2 previous similar messages [ 5997.611544] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 5997.620249] Lustre: Skipped 7 previous similar messages [ 6001.032097] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6001.733799] Lustre: lustre-MDT0000: Recovery over after 0:04, of 3 clients 3 recovered and 0 were evicted. [ 6001.739168] Lustre: Skipped 7 previous similar messages [ 6001.781404] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:4707 to 0x280000401:4737) [ 6001.782257] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4644 to 0x2c0000401:4705) [ 6012.509952] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 6015.200742] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6030.038203] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount FULL mgc.*.mgs_server_uuid 1475 0 [ 6032.105267] Lustre: DEBUG MARKER: mgc.*.mgs_server_uuid in FULL state after 0 sec [ 6040.664364] Lustre: DEBUG MARKER: == replay-single test 58c: resend/reconstruct setxattr op ========================================================== 18:06:10 (1786831570) [ 6048.816749] Lustre: *** cfs_fail_loc=123, val=2147483648*** [ 6112.076028] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 6112.081260] Lustre: Skipped 1 previous similar message [ 6112.088309] LustreError: 11771:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c851846450 x1873622471564928/t296352743435(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:378/0 lens 66040/440 e 0 to 0 dl 1786831698 ref 1 fl Interpret:/600/0 rc 0/0 job:'setfattr.0' uid:0 gid:0 projid:0 [ 6171.196773] Lustre: 17699:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c881dee680 x1873622471564928/t296352743435(0) o36->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:437/0 lens 66040/440 e 0 to 0 dl 1786831757 ref 1 fl Interpret:/602/0 rc 0/0 job:'setfattr.0' uid:0 gid:0 projid:0 [ 6183.859671] Lustre: DEBUG MARKER: SKIP: replay-single test_59 skipping ALWAYS excluded test 59 [ 6185.726883] Lustre: DEBUG MARKER: == replay-single test 60: test llog post recovery init vs llog unlink ========================================================== 18:08:35 (1786831715) [ 6201.566455] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 6204.532413] Lustre: Failing over lustre-MDT0000 [ 6205.033616] Lustre: server umount lustre-MDT0000 complete [ 6221.663101] Lustre: 3647:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786831737/real 1786831737] req@ffff99c875e54380 x1873622494051712/t0(0) o400->MGC192.168.201.120@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786831753 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 6221.681680] Lustre: 3647:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 5 previous similar messages [ 6221.687447] LustreError: MGC192.168.201.120@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 6221.692874] LustreError: Skipped 5 previous similar messages [ 6227.496604] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 6227.498760] LDISKFS-fs (dm-0): recovery complete [ 6227.533186] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6232.039382] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xf03e36cc1e2b8ab0 [ 6232.261227] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 6232.273413] Lustre: Skipped 5 previous similar messages [ 6232.300396] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 6232.312726] Lustre: Skipped 5 previous similar messages [ 6236.717500] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6239.134675] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4806 to 0x2c0000401:4833) [ 6239.138466] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:4839 to 0x280000401:4865) [ 6249.154807] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 6251.344610] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6262.807626] Lustre: DEBUG MARKER: == replay-single test 61a: test race llog recovery vs llog cleanup ========================================================== 18:09:52 (1786831792) [ 6286.443611] Lustre: DEBUG MARKER: ost1 REPLAY BARRIER on lustre-OST0000 [ 6298.950584] Lustre: Failing over lustre-OST0000 [ 6299.073338] Lustre: server umount lustre-OST0000 complete [ 6323.215994] LDISKFS-fs (dm-2): 3 truncates cleaned up [ 6323.223779] LDISKFS-fs (dm-2): recovery complete [ 6323.246516] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 6331.672846] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6346.983630] Lustre: Failing over lustre-OST0000 [ 6347.155473] Lustre: server umount lustre-OST0000 complete [ 6368.017711] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 6377.026646] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6390.690727] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 6392.865820] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 6436.371460] Lustre: DEBUG MARKER: == replay-single test 61b: test race mds llog sync vs llog cleanup ========================================================== 18:12:46 (1786831966) [ 6439.793778] Lustre: Failing over lustre-MDT0000 [ 6439.861873] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 6440.073883] Lustre: server umount lustre-MDT0000 complete [ 6459.844525] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6465.449270] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6465.604450] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:4839 to 0x280000401:4897) [ 6465.609183] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4806 to 0x2c0000401:4865) [ 6480.335606] Lustre: Failing over lustre-MDT0000 [ 6480.599067] Lustre: server umount lustre-MDT0000 complete [ 6500.190469] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6505.081834] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6505.981235] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 6505.990139] Lustre: Skipped 21 previous similar messages [ 6506.099214] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:4839 to 0x280000401:4929) [ 6506.099443] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4806 to 0x2c0000401:4897) [ 6515.682322] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 6517.336392] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6527.206943] Lustre: DEBUG MARKER: == replay-single test 61c: test race mds llog sync vs llog cleanup ========================================================== 18:14:17 (1786832057) [ 6543.269468] Lustre: Failing over lustre-OST0000 [ 6543.368893] Lustre: server umount lustre-OST0000 complete [ 6543.850090] LustreError: lustre-OST0000-osc-MDT0001: operation ost_statfs to node 0@lo failed: rc = -107 [ 6543.863300] LustreError: Skipped 2 previous similar messages [ 6564.120550] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 6572.210138] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6583.354440] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 6585.366030] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 6596.208668] Lustre: DEBUG MARKER: == replay-single test 61d: error in llog_setup should cleanup the llog context correctly ========================================================== 18:15:26 (1786832126) [ 6598.367571] Lustre: Failing over lustre-MDT0000 [ 6598.859333] Lustre: server umount lustre-MDT0000 complete [ 6600.159974] Lustre: lustre-MDT0000-lwp-OST0000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 6600.167220] LustreError: 6526:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 6600.179217] Lustre: Skipped 21 previous similar messages [ 6600.196932] LustreError: 6526:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 134 previous similar messages [ 6608.045026] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6608.124960] Lustre: *** cfs_fail_loc=605, val=0*** [ 6608.128692] LustreError: 164940:0:(llog_obd.c:192:llog_setup()) MGS: ctxt 0 lop_setup=ffffffffc0a3c4b0 failed: rc = -95 [ 6608.134681] LustreError: 164940:0:(obd_config.c:845:class_setup()) setup MGS failed (-95) [ 6608.137492] LustreError: 164940:0:(obd_mount.c:259:lustre_start_simple()) MGS setup error -95 [ 6608.141151] LustreError: 164940:0:(tgt_mount.c:116:server_deregister_mount()) MGS not registered [ 6608.144113] LustreError: Failed to start MGS 'MGS' (-95). Is the 'mgs' module loaded? [ 6608.147509] LustreError: 164940:0:(tgt_mount.c:2129:server_put_super()) no obd lustre-MDT0000 [ 6608.159133] Lustre: server umount lustre-MDT0000 complete [ 6608.161222] LustreError: 164940:0:(super25.c:179:lustre_fill_super()) llite: Unable to mount : rc = -95 [ 6615.578377] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6618.664138] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 6618.673584] Lustre: Skipped 6 previous similar messages [ 6619.868929] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6621.218293] Lustre: lustre-MDT0000: Recovery over after 0:03, of 2 clients 2 recovered and 0 were evicted. [ 6621.223138] Lustre: Skipped 6 previous similar messages [ 6621.257614] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4899 to 0x2c0000401:4929) [ 6621.260098] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:4931 to 0x280000401:4961) [ 6628.665368] Lustre: DEBUG MARKER: == replay-single test 62: don't mis-drop resent replay === 18:15:59 (1786832159) [ 6636.133928] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 6639.618224] Lustre: Failing over lustre-MDT0000 [ 6639.931575] Lustre: server umount lustre-MDT0000 complete [ 6663.960503] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 6663.962852] LDISKFS-fs (dm-0): recovery complete [ 6663.973985] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6667.232140] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c85e701180 x1873622494434048/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 6667.295700] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 5 previous similar messages [ 6670.905074] Lustre: *** cfs_fail_loc=707, val=0*** [ 6674.300969] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6730.313175] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnected, waiting for 2 clients in recovery for 0:55 [ 6730.400941] Lustre: 167112:0:(ldlm_lib.c:2085:extend_recovery_timer()) lustre-MDT0000: extended recovery timer reached hard limit: 180, extend: 1 [ 6730.414787] Lustre: 167112:0:(ldlm_lib.c:2085:extend_recovery_timer()) Skipped 1 previous similar message [ 6730.886896] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4942 to 0x2c0000401:4961) [ 6730.888696] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:4975 to 0x280000401:4993) [ 6738.787874] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 6741.592503] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6757.403241] Lustre: DEBUG MARKER: == replay-single test 65a: AT: verify early replies ====== 18:18:06 (1786832286) [ 6791.855408] LustreError: 6520:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c870a85c00 x1873622472471168/t0(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:259/0 lens 664/0 e 0 to 0 dl 1786832334 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6791.870121] LustreError: 6520:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 11000ms [ 6802.927201] LustreError: 6520:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6802.997497] LustreError: 6524:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c877ea9880 x1873622472472448/t0(0) o35->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:270/0 lens 392/0 e 0 to 0 dl 1786832345 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6807.828377] LustreError: 6522:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c877eaa680 x1873622472478848/t0(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:314/0 lens 576/0 e 0 to 0 dl 1786832389 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'unlinkmany.0' uid:0 gid:0 projid:0 [ 6807.866880] LustreError: 6522:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 27 previous similar messages [ 6813.165713] LustreError: 43910:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c881c03100 x1873622494510208/t0(0) o6->lustre-MDT0000-mdtlov_UUID@0@lo:280/0 lens 544/0 e 0 to 0 dl 1786832355 ref 1 fl Interpret:/200/ffffffff rc 0/-1 job:'osp-syn-1-0.0' uid:0 gid:0 projid:4294967295 [ 6813.213937] LustreError: 43910:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 24 previous similar messages [ 6818.297324] LustreError: 43914:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c877eaaa00 x1873622494512896/t0(0) o6->lustre-MDT0000-mdtlov_UUID@0@lo:285/0 lens 544/0 e 0 to 0 dl 1786832360 ref 1 fl Interpret:/200/ffffffff rc 0/-1 job:'osp-syn-0-0.0' uid:0 gid:0 projid:4294967295 [ 6818.307830] LustreError: 43914:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 3 previous similar messages [ 6825.942128] Lustre: DEBUG MARKER: == replay-single test 65b: AT: verify early replies on packed reply / bulk ========================================================== 18:19:15 (1786832355) [ 6859.943278] LustreError: 8441:0:(tgt_handler.c:2830:tgt_brw_write()) cfs_fail_timeout id 224 sleeping for 11000ms [ 6870.951130] LustreError: 8441:0:(tgt_handler.c:2830:tgt_brw_write()) cfs_fail_timeout id 224 awake [ 6882.231508] Lustre: DEBUG MARKER: == replay-single test 66a: AT: verify MDT service time adjusts with no early replies ========================================================== 18:20:12 (1786832412) [ 6912.560841] LustreError: 17699:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c743f1f800 x1873622472500608/t0(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:379/0 lens 576/0 e 0 to 0 dl 1786832454 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6912.600535] LustreError: 17699:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 5000ms [ 6917.623139] LustreError: 17699:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6930.315547] LustreError: 67185:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6930.336141] LustreError: 6521:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c877eb2680 x1873622472523264/t0(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:436/0 lens 664/0 e 0 to 0 dl 1786832511 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6930.356565] LustreError: 6521:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 120 previous similar messages [ 6951.638768] Lustre: DEBUG MARKER: == replay-single test 66b: AT: verify net latency adjusts ========================================================== 18:21:21 (1786832481) [ 7048.047406] Lustre: DEBUG MARKER: == replay-single test 67a: AT: verify slow request processing doesn't induce reconnects ========================================================== 18:22:58 (1786832578) [ 7078.698253] LustreError: 11771:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff99c875d4ce00 x1873622472583552/t0(0) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:546/0 lens 576/0 e 0 to 0 dl 1786832621 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 7078.729144] LustreError: 11771:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 98 previous similar messages [ 7078.744453] LustreError: 11771:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 400ms [ 7078.754776] LustreError: 11771:0:(service.c:2559:ptlrpc_server_handle_request()) Skipped 1 previous similar message [ 7079.175120] LustreError: 11771:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 7095.107095] LustreError: 6523:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 400ms [ 7095.114157] LustreError: 6523:0:(service.c:2559:ptlrpc_server_handle_request()) Skipped 37 previous similar messages [ 7095.543132] LustreError: 6523:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 7095.564076] LustreError: 6523:0:(service.c:2559:ptlrpc_server_handle_request()) Skipped 37 previous similar messages [ 7131.247602] Lustre: DEBUG MARKER: == replay-single test 67b: AT: verify instant slowdown doesn't induce reconnects ========================================================== 18:24:21 (1786832661) [ 7168.819834] Lustre: DEBUG MARKER: phase 2 [ 7184.208960] Lustre: DEBUG MARKER: == replay-single test 68: AT: verify slowing locks ======= 18:25:13 (1786832713) [ 7267.172936] Lustre: DEBUG MARKER: == replay-single test 70a: check multi client t-f ======== 18:26:37 (1786832797) [ 7268.784622] Lustre: DEBUG MARKER: SKIP: replay-single test_70a Need two or more clients, have 1 [ 7270.475721] Lustre: DEBUG MARKER: == replay-single test 70b: dbench 2mdts recovery; 1 clients ========================================================== 18:26:40 (1786832800) [ 7275.606720] Lustre: DEBUG MARKER: Started rundbench load pid=135449 ... [ 7285.491619] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 7288.358925] Lustre: DEBUG MARKER: test_70b fail mds1 1 times [ 7290.205982] Lustre: Failing over lustre-MDT0000 [ 7290.466440] Lustre: lustre-MDT0000: Not available for connect from 192.168.201.20@tcp (stopping) [ 7290.478433] Lustre: Skipped 3 previous similar messages [ 7290.750828] Lustre: server umount lustre-MDT0000 complete [ 7292.384973] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 7292.389189] LustreError: 67185:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 7292.399129] Lustre: Skipped 8 previous similar messages [ 7292.413165] LustreError: 67185:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 47 previous similar messages [ 7307.743578] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786832823/real 1786832823] req@ffff99c87992b800 x1873622494803072/t0(0) o400->MGC192.168.201.120@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786832839 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 7307.770924] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 1 previous similar message [ 7307.780117] LustreError: MGC192.168.201.120@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 7307.796203] LustreError: Skipped 4 previous similar messages [ 7313.447150] LDISKFS-fs (dm-0): 4 truncates cleaned up [ 7313.450397] LDISKFS-fs (dm-0): recovery complete [ 7313.556798] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7317.984351] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c87f53b480 x1873622494905856/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 7318.376676] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 7318.382068] Lustre: Skipped 7 previous similar messages [ 7318.441708] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 7318.453366] Lustre: Skipped 7 previous similar messages [ 7319.548333] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 7319.556787] Lustre: Skipped 1 previous similar message [ 7323.027355] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7323.624348] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 7323.645673] Lustre: Skipped 13 previous similar messages [ 7323.706922] Lustre: lustre-MDT0000: Recovery over after 0:04, of 2 clients 2 recovered and 0 were evicted. [ 7323.715065] Lustre: Skipped 1 previous similar message [ 7323.772463] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:5046 to 0x280000401:5089) [ 7323.778039] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:5011 to 0x2c0000401:5057) [ 7333.775846] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 7335.657620] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7346.320739] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 7349.223899] Lustre: DEBUG MARKER: test_70b fail mds2 2 times [ 7351.353060] Lustre: Failing over lustre-MDT0001 [ 7351.389380] Lustre: lustre-MDT0001: Not available for connect from 192.168.201.20@tcp (stopping) [ 7351.666140] Lustre: server umount lustre-MDT0001 complete [ 7354.339547] LustreError: lustre-MDT0001-osp-MDT0000: operation mds_statfs to node 0@lo failed: rc = -107 [ 7354.343604] LustreError: Skipped 1 previous similar message [ 7377.548255] LDISKFS-fs (dm-1): 8 truncates cleaned up [ 7377.550149] LDISKFS-fs (dm-1): recovery complete [ 7377.569804] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7382.983382] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7385.389223] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:558 to 0x2c0000400:577) [ 7385.390390] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:576 to 0x280000400:609) [ 7393.931480] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 7395.801793] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7427.748941] Lustre: DEBUG MARKER: == replay-single test 70c: tar 2mdts recovery ============ 18:29:17 (1786832957) [ 7558.885994] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 7571.673470] Lustre: DEBUG MARKER: test_70c fail mds1 1 times [ 7573.848883] Lustre: Failing over lustre-MDT0000 [ 7573.905568] Lustre: lustre-MDT0000: Not available for connect from 192.168.201.20@tcp (stopping) [ 7574.150589] Lustre: server umount lustre-MDT0000 complete [ 7600.686570] LDISKFS-fs (dm-0): 3 truncates cleaned up [ 7600.691677] LDISKFS-fs (dm-0): recovery complete [ 7600.706327] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7602.912082] LustreError: 178652:0:(import.c:339:ptlrpc_invalidate_import()) MGS: timeout waiting for callback (1 != 0) [ 7602.932373] LustreError: 178652:0:(import.c:363:ptlrpc_invalidate_import()) @@@ still on sending list req@ffff99c87f480700 x1873622496437760/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 1786833134 ref 1 fl Rpc:NQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 7602.977235] LustreError: 178652:0:(import.c:373:ptlrpc_invalidate_import()) MGS: Unregistering RPCs found (0). Network is sluggish? Waiting for them to error out. [ 7608.758488] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7611.716655] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:5305 to 0x2c0000401:5345) [ 7611.716953] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:5338 to 0x280000401:5377) [ 7619.818875] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 7621.902423] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7693.808289] Lustre: DEBUG MARKER: == replay-single test 70d: mkdir/rmdir striped dir 2mdts recovery ========================================================== 18:33:43 (1786833223) [ 7823.528437] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 7837.357729] Lustre: DEBUG MARKER: test_70d fail mds2 1 times [ 7840.089909] Lustre: Failing over lustre-MDT0001 [ 7840.169686] Lustre: lustre-MDT0001: Not available for connect from 192.168.201.20@tcp (stopping) [ 7840.668660] Lustre: server umount lustre-MDT0001 complete [ 7842.291205] LustreError: 6504:0:(ldlm_lockd.c:2564:ldlm_cancel_handler()) ldlm_cancel from 0@lo arrived at 1786833373 with bad export cookie 17311334267966605640 [ 7842.304619] LustreError: 6504:0:(ldlm_lockd.c:2564:ldlm_cancel_handler()) Skipped 1 previous similar message [ 7867.464706] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 7867.471306] LDISKFS-fs (dm-1): recovery complete [ 7867.491357] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7873.827590] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7875.114266] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:1052 to 0x2c0000400:1089) [ 7875.115516] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:1084 to 0x280000400:1121) [ 7884.637474] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 7887.161275] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7897.649822] Lustre: DEBUG MARKER: == replay-single test 70e: rename cross-MDT with random fails ========================================================== 18:37:07 (1786833427) [ 8030.343024] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 8042.816658] Lustre: DEBUG MARKER: test_70e fail mds2 1 times [ 8045.266605] Lustre: Failing over lustre-MDT0001 [ 8045.876905] Lustre: server umount lustre-MDT0001 complete [ 8047.073138] Lustre: lustre-MDT0001-lwp-OST0000: Connection to lustre-MDT0001 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 8047.087366] LustreError: 67185:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0001: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 8047.089748] Lustre: Skipped 11 previous similar messages [ 8047.111488] LustreError: 67185:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 106 previous similar messages [ 8070.056601] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 8070.065921] LDISKFS-fs (dm-1): recovery complete [ 8070.084032] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8070.661416] Lustre: lustre-MDT0001: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 8070.709545] Lustre: lustre-MDT0001: in recovery but waiting for the first client to connect [ 8070.723167] Lustre: Skipped 3 previous similar messages [ 8072.528951] Lustre: lustre-MDT0001: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 8072.534630] Lustre: Skipped 3 previous similar messages [ 8075.765987] Lustre: lustre-MDT0001-lwp-OST0001: Connection restored to 0@lo (at 0@lo) [ 8075.776473] Lustre: Skipped 13 previous similar messages [ 8075.829374] Lustre: lustre-MDT0001: Recovery over after 0:03, of 2 clients 2 recovered and 0 were evicted. [ 8075.845949] Lustre: Skipped 3 previous similar messages [ 8075.900829] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:1129 to 0x280000400:1153) [ 8075.901872] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:1097 to 0x2c0000400:1121) [ 8076.260368] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8088.148959] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 8090.168614] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8100.195661] Lustre: DEBUG MARKER: == replay-single test 70f: OSS O_DIRECT recovery with 1 clients ========================================================== 18:40:30 (1786833630) [ 8112.870379] Lustre: DEBUG MARKER: ost1 REPLAY BARRIER on lustre-OST0000 [ 8115.407127] Lustre: DEBUG MARKER: test_70f failing OST 1 times [ 8117.391268] Lustre: Failing over lustre-OST0000 [ 8117.500188] Lustre: server umount lustre-OST0000 complete [ 8118.752679] LustreError: lustre-OST0000-osc-MDT0000: operation ost_statfs to node 0@lo failed: rc = -107 [ 8118.757421] LustreError: Skipped 1 previous similar message [ 8141.379274] LDISKFS-fs (dm-2): 3 truncates cleaned up [ 8141.381081] LDISKFS-fs (dm-2): recovery complete [ 8141.398673] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 8141.673951] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 8149.211622] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8160.763434] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 8163.496885] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 8176.173924] Lustre: DEBUG MARKER: == replay-single test 71a: mkdir/rmdir striped dir with 2 mdts recovery ========================================================== 18:41:46 (1786833706) [ 8305.098886] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 8313.468246] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8326.886483] Lustre: DEBUG MARKER: fail mds2 mds1 1 times [ 8329.324585] Lustre: Failing over lustre-MDT0001 [ 8329.377340] Lustre: lustre-MDT0001: Not available for connect from 0@lo (stopping) [ 8329.599568] Lustre: server umount lustre-MDT0001 complete [ 8332.889716] Lustre: Failing over lustre-MDT0000 [ 8338.987780] Lustre: lustre-MDT0000: Not available for connect from 192.168.201.20@tcp (stopping) [ 8339.001091] Lustre: Skipped 4 previous similar messages [ 8341.288866] Lustre: server umount lustre-MDT0000 complete [ 8356.831259] Lustre: 3644:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786833872/real 1786833872] req@ffff99c742d4ed80 x1873622501510272/t0(0) o400->MGC192.168.201.120@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786833888 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 8356.859185] Lustre: 3644:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 1 previous similar message [ 8356.867437] LustreError: MGC192.168.201.120@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 8356.876720] LustreError: Skipped 1 previous similar message [ 8366.002465] LDISKFS-fs (dm-0): 3 truncates cleaned up [ 8366.008479] LDISKFS-fs (dm-0): recovery complete [ 8366.032090] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8366.311133] LDISKFS-fs (dm-1): 5 truncates cleaned up [ 8366.315175] LDISKFS-fs (dm-1): recovery complete [ 8366.328127] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8367.661276] Lustre: lustre-MDT0001: Imperative Recovery not enabled, recovery window 60-180 [ 8367.668615] Lustre: Skipped 3 previous similar messages [ 8372.321874] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8373.048073] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8374.688425] LustreError: 8446:0:(ldlm_lockd.c:915:ldlm_server_blocking_ast()) ### BUG 6063: lock collide during recovery ns: mdt-lustre-MDT0001_UUID lock: ffff99c877e59600/0xf03e36cc1e3f767d lrc: 3/0,0 mode: PW/PW res: [0x240000fba:0x467:0x0].0x0 bits 0x2/0x0 rrc: 3 type: IBT gid 0 flags: 0x40000000000020 nid: 0@lo remote: 0xf03e36cc1e3f7676 expref: 35 pid: 177448 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 [ 8389.088808] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:1129 to 0x280000400:1185) [ 8389.092342] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:1097 to 0x2c0000400:1153) [ 8394.526983] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:6090 to 0x280000401:6113) [ 8394.528440] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6057 to 0x2c0000401:6081) [ 8400.942716] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid,mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 8402.547762] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8404.013123] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8411.378454] Lustre: DEBUG MARKER: == replay-single test 73a: open(O_CREAT), unlink, replay, reconnect before open replay, close ========================================================== 18:45:41 (1786833941) [ 8418.098564] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8420.313600] Lustre: Failing over lustre-MDT0000 [ 8420.327750] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 8420.503309] Lustre: server umount lustre-MDT0000 complete [ 8443.307215] LDISKFS-fs (dm-0): 3 truncates cleaned up [ 8443.309645] LDISKFS-fs (dm-0): recovery complete [ 8443.317375] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8448.491143] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xf03e36cc1e3fde0a [ 8450.619857] Lustre: *** cfs_fail_loc=302, val=2147483648*** [ 8453.519211] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8467.011932] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnected, waiting for 2 clients in recovery for 0:53 [ 8467.176754] Lustre: 187958:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c881def800 x1873622484877824/t326417526413(326417526413) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:424/0 lens 520/3488 e 0 to 0 dl 1786834009 ref 1 fl Interpret:/606/0 rc 0/0 job:'lfs.0' uid:0 gid:0 projid:0 [ 8467.309264] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6057 to 0x2c0000401:6113) [ 8467.310727] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:6115 to 0x280000401:6145) [ 8473.776418] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8475.456064] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8485.535231] Lustre: DEBUG MARKER: == replay-single test 73b: open(O_CREAT), unlink, replay, reconnect at open_replay reply, close ========================================================== 18:46:55 (1786834015) [ 8494.420410] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8497.472028] Lustre: Failing over lustre-MDT0000 [ 8497.755484] Lustre: server umount lustre-MDT0000 complete [ 8523.717709] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 8523.721722] LDISKFS-fs (dm-0): recovery complete [ 8523.734336] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8526.047545] LustreError: 192405:0:(import.c:339:ptlrpc_invalidate_import()) MGS: timeout waiting for callback (1 != 0) [ 8526.059178] LustreError: 192405:0:(import.c:363:ptlrpc_invalidate_import()) @@@ still on sending list req@ffff99c746bb4e00 x1873622502085760/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 1786834057 ref 1 fl Rpc:NQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 8526.083971] LustreError: 192405:0:(import.c:373:ptlrpc_invalidate_import()) MGS: Unregistering RPCs found (0). Network is sluggish? Waiting for them to error out. [ 8528.468513] Lustre: *** cfs_fail_loc=157, val=2147483648*** [ 8528.473410] LustreError: 188957:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c881b57100 x1873622484877824/t326417526413(326417526413) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:485/0 lens 520/664 e 0 to 0 dl 1786834070 ref 1 fl Interpret:/604/0 rc 301/0 job:'lfs.0' uid:0 gid:0 projid:0 [ 8531.842983] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8544.850509] Lustre: lustre-MDT0000: Client 2ee7416c-2f87-4b00-a3c6-9e1355d088c2 (at 192.168.201.20@tcp) reconnected, waiting for 2 clients in recovery for 0:53 [ 8544.872224] Lustre: 187958:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff99c742d4e680 x1873622484877824/t326417526413(326417526413) o101->2ee7416c-2f87-4b00-a3c6-9e1355d088c2@192.168.201.20@tcp:502/0 lens 520/3488 e 0 to 0 dl 1786834087 ref 1 fl Interpret:/606/0 rc 0/0 job:'lfs.0' uid:0 gid:0 projid:0 [ 8545.031716] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:6147 to 0x280000401:6177) [ 8545.034939] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6057 to 0x2c0000401:6145) [ 8551.542794] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8553.460868] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8563.627799] Lustre: DEBUG MARKER: == replay-single test 74: Ensure applications don't fail waiting for OST recovery ========================================================== 18:48:13 (1786834093) [ 8568.018144] Lustre: Failing over lustre-OST0000 [ 8568.256675] Lustre: server umount lustre-OST0000 complete [ 8572.542923] Lustre: Failing over lustre-MDT0000 [ 8572.834841] Lustre: server umount lustre-MDT0000 complete [ 8592.285755] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8597.479900] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8598.034949] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6057 to 0x2c0000401:6177) [ 8606.865284] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 8608.923682] Lustre: lustre-OST0000: Denying connection for new client 7653632c-411c-434f-911f-6b6ffd81ad63 (at 192.168.201.20@tcp), waiting for 2 known clients (1 recovered, 0 in progress, and 0 evicted) to recover in 1:09 [ 8612.368514] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:6147 to 0x280000401:6209) [ 8614.937471] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8629.216217] Lustre: DEBUG MARKER: == replay-single test 80a: DNE: create remote dir, drop update rep from MDT0, fail MDT0 ========================================================== 18:49:19 (1786834159) [ 8630.274759] Lustre: *** cfs_fail_loc=1701, val=2147483648*** [ 8630.283358] LustreError: 190702:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c877262d80 x1873622502156800/t343597383687(0) o1000->lustre-MDT0001-mdtlov_UUID@0@lo:587/0 lens 1312/4320 e 0 to 0 dl 1786834172 ref 1 fl Interpret:/200/0 rc 0/0 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8637.372722] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8638.718885] Lustre: Failing over lustre-MDT0000 [ 8638.961842] Lustre: server umount lustre-MDT0000 complete [ 8649.188502] LustreError: 188743:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 8649.219804] LustreError: 188743:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 147 previous similar messages [ 8661.203651] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 8661.206946] LDISKFS-fs (dm-0): recovery complete [ 8661.227183] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8669.673408] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xf03e36cc1e3fee41 [ 8672.226422] Lustre: mdt00_000: service thread pid 187956 was inactive for 42.037 seconds. The thread might be hung, or it might only be slow and will resume later. Dumping the stack trace for debugging purposes: [ 8672.237874] task:mdt00_000 state:I stack:0 pid:187956 ppid:2 flags:0x80004000 [ 8672.242175] Call Trace: [ 8672.244112] __schedule+0x351/0xcb0 [ 8672.247323] schedule+0xc0/0x180 [ 8672.248587] top_trans_stop+0xbc5/0x1590 [ptlrpc] [ 8672.253306] ? woken_wake_function+0x30/0x30 [ 8672.255844] lod_trans_stop+0xe2/0x560 [lod] [ 8672.257969] mdd_trans_stop+0x29/0x1f0 [mdd] [ 8672.265313] mdd_create+0x1f9c/0x2590 [mdd] [ 8672.268723] ? mdd_links_rename+0x550/0x550 [mdd] [ 8672.282768] mdt_create+0xcc6/0x2110 [mdt] [ 8672.286468] mdt_reint_create+0x336/0x5d0 [mdt] [ 8672.289593] mdt_reint_rec+0x139/0x2b0 [mdt] [ 8672.292350] mdt_reint_internal+0x693/0xdc0 [mdt] [ 8672.295387] mdt_reint+0x163/0x190 [mdt] [ 8672.298346] tgt_handle_request0+0x137/0xaf0 [ptlrpc] [ 8672.301670] tgt_request_handle+0x575/0x1f70 [ptlrpc] [ 8672.305182] ptlrpc_server_handle_request+0x443/0x13b0 [ptlrpc] [ 8672.310680] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8672.313117] ptlrpc_main+0xce8/0x1400 [ptlrpc] [ 8672.317887] ? ptlrpc_wait_event+0x690/0x690 [ptlrpc] [ 8672.323463] kthread+0x1d1/0x200 [ 8672.331768] ? set_kthread_struct+0x70/0x70 [ 8672.336764] ret_from_fork+0x1f/0x30 [ 8673.789481] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8675.373408] Lustre: mdt00_000: service thread pid 187956 completed after 45.184s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 8675.383433] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:6211 to 0x280000401:6241) [ 8675.403490] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6057 to 0x2c0000401:6209) [ 8683.329228] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8685.209560] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8695.653581] Lustre: DEBUG MARKER: == replay-single test 80b: DNE: create remote dir, drop update rep from MDT0, fail MDT1 ========================================================== 18:50:26 (1786834226) [ 8696.815907] Lustre: *** cfs_fail_loc=1701, val=2147483648*** [ 8696.818570] LustreError: 187963:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c85e4b6450 x1873622502199296/t347892350988(0) o1000->lustre-MDT0001-mdtlov_UUID@0@lo:654/0 lens 488/4320 e 0 to 0 dl 1786834239 ref 1 fl Interpret:/200/0 rc 0/0 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8705.736906] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8713.185565] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 8713.200266] Lustre: Skipped 25 previous similar messages [ 8713.211159] Lustre: lustre-MDT0000: Received new MDS connection from 0@lo, keep former export from same NID [ 8713.226301] Lustre: lustre-MDT0000-osp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 8713.244161] Lustre: Skipped 29 previous similar messages [ 8713.596149] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 8715.265049] Lustre: Failing over lustre-MDT0001 [ 8715.967541] Lustre: server umount lustre-MDT0001 complete [ 8740.844658] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 8740.847466] LDISKFS-fs (dm-1): recovery complete [ 8740.856402] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8741.149220] Lustre: lustre-MDT0001: in recovery but waiting for the first client to connect [ 8741.155266] Lustre: Skipped 8 previous similar messages [ 8742.969222] Lustre: lustre-MDT0001: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 8742.978786] Lustre: Skipped 8 previous similar messages [ 8746.489687] Lustre: lustre-MDT0001: Recovery over after 0:04, of 2 clients 2 recovered and 0 were evicted. [ 8746.502257] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8746.503957] Lustre: Skipped 8 previous similar messages [ 8746.571188] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:1196 to 0x280000400:1217) [ 8746.574668] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:1164 to 0x2c0000400:1185) [ 8757.667689] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 8759.543331] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8769.303833] Lustre: DEBUG MARKER: == replay-single test 80c: DNE: create remote dir, drop update rep from MDT1, fail MDT[0,1] ========================================================== 18:51:39 (1786834299) [ 8770.906803] LustreError: 187963:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff99c85f14ad80 x1873622502245376/t347892351009(0) o1000->lustre-MDT0001-mdtlov_UUID@0@lo:728/0 lens 2520/4320 e 0 to 0 dl 1786834313 ref 1 fl Interpret:/200/0 rc 0/0 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8777.919477] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8784.196427] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 8785.731178] Lustre: Failing over lustre-MDT0000 [ 8785.977985] Lustre: server umount lustre-MDT0000 complete [ 8808.848840] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 8808.850976] LDISKFS-fs (dm-0): recovery complete [ 8808.858502] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8811.487226] Lustre: mdt00_003: service thread pid 188743 was inactive for 40.648 seconds. The thread might be hung, or it might only be slow and will resume later. Dumping the stack trace for debugging purposes: [ 8811.499409] task:mdt00_003 state:I stack:0 pid:188743 ppid:2 flags:0x80004000 [ 8811.509078] Call Trace: [ 8811.511860] __schedule+0x351/0xcb0 [ 8811.519297] schedule+0xc0/0x180 [ 8811.523215] top_trans_stop+0xbc5/0x1590 [ptlrpc] [ 8811.530309] ? woken_wake_function+0x30/0x30 [ 8811.534240] lod_trans_stop+0xe2/0x560 [lod] [ 8811.537224] mdd_trans_stop+0x29/0x1f0 [mdd] [ 8811.539891] mdd_create+0x1f9c/0x2590 [mdd] [ 8811.543273] ? mdd_links_rename+0x550/0x550 [mdd] [ 8811.547794] mdt_create+0xcc6/0x2110 [mdt] [ 8811.551703] mdt_reint_create+0x336/0x5d0 [mdt] [ 8811.557212] mdt_reint_rec+0x139/0x2b0 [mdt] [ 8811.562976] mdt_reint_internal+0x693/0xdc0 [mdt] [ 8811.571208] mdt_reint+0x163/0x190 [mdt] [ 8811.574544] tgt_handle_request0+0x137/0xaf0 [ptlrpc] [ 8811.578059] tgt_request_handle+0x575/0x1f70 [ptlrpc] [ 8811.584633] ptlrpc_server_handle_request+0x443/0x13b0 [ptlrpc] [ 8811.587984] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8811.591606] ptlrpc_main+0xce8/0x1400 [ptlrpc] [ 8811.594828] ? ptlrpc_wait_event+0x690/0x690 [ptlrpc] [ 8811.597769] kthread+0x1d1/0x200 [ 8811.599614] ? set_kthread_struct+0x70/0x70 [ 8811.602111] ret_from_fork+0x1f/0x30 [ 8813.543393] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff99c881e34700 x1873622502264960/t0(0) o250->MGC192.168.201.120@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 8813.579069] LustreError: 3643:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 2 previous similar messages [ 8818.760222] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8819.289384] Lustre: 201256:0:(out_handler.c:886:out_tx_end()) lustre-MDT0000-osd: error during execution of #0 from /home/green/git/lustre-release/lustre/ptlrpc/../target/out_handler.c:565: rc = -2 [ 8819.301251] LustreError: 201256:0:(tgt_lastrcvd.c:1516:tgt_last_rcvd_update()) lustre-MDT0000: replay transno 347892351004 failed: rc = -2 [ 8819.306278] LustreError: 3643:0:(client.c:3435:ptlrpc_replay_interpret()) @@@ status -2, old was 0 req@ffff99c743802300 x1873622502239744/t347892351004(347892351004) o1000->lustre-MDT0000-osp-MDT0001@0@lo:24/4 lens 264/4320 e 0 to 0 dl 1786834366 ref 2 fl Interpret:RQU/204/0 rc -2/-2 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8819.322356] Lustre: mdt00_003: service thread pid 188743 completed after 48.484s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 8819.451382] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6057 to 0x2c0000401:6241) [ 8819.460763] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000401:6211 to 0x280000401:6273) [ 8827.999835] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8829.726617] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8833.494628] Lustre: Failing over lustre-MDT0001 [ 8833.722794] Lustre: server umount lustre-MDT0001 complete [ 8834.544266] LustreError: lustre-MDT0001-osp-MDT0000: operation mds_statfs to node 0@lo failed: rc = -107 [ 8834.556714] LustreError: Skipped 8 previous similar messages [ 8856.454365] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 8856.460241] LDISKFS-fs (dm-1): recovery complete [ 8856.475761] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8860.683982] Lustre: DEBUG MARKER: oleg120-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8862.191534] Lustre: 202593:0:(update_recovery.c:1391:distribute_txn_replay_handle()) lustre-MDT0001-mdtlov: error during execution of #0 from /home/green/git/lustre-release/lustre/ptlrpc/../target/update_recovery.c:946: rc = -116 [ 8862.220340] LustreError: 202593:0:(update_trans.c:1080:top_trans_stop()) lustre-MDT0000-osp-MDT0001: stop trans failed: rc = -20 [ 8862.290187] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000400:1228 to 0x280000400:1249) [ 8862.294582] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:1196 to 0x2c0000400:1217) [ 8869.958728] Lustre: DEBUG MARKER: oleg120-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 8871.965941] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8875.344656] Lustre: DEBUG MARKER: replay-single test_80c: @@@@@@ FAIL: remote creation failed