[ 3593.724571] Lustre: Failing over lustre-MDT0000 [ 3593.863852] Lustre: server umount lustre-MDT0000 complete [ 3606.648296] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 3606.651351] LDISKFS-fs (dm-0): recovery complete [ 3606.658651] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3607.046882] Lustre: lustre-MDT0000: Aborting client recovery [ 3607.052862] LustreError: 107272:0:(ldlm_lib.c:3004:target_stop_recovery_thread()) lustre-MDT0000: Aborting recovery [ 3607.055582] Lustre: 107304:0:(ldlm_lib.c:2404:target_recovery_overseer()) recovery is aborted, evict exports in recovery [ 3607.063606] Lustre: 107304:0:(ldlm_lib.c:2404:target_recovery_overseer()) Skipped 2 previous similar messages [ 3607.068275] Lustre: 107304:0:(genops.c:1622:class_disconnect_stale_exports()) lustre-MDT0000: disconnect stale client lustre-MDT0001-mdtlov_UUID@ [ 3607.075282] Lustre: 107304:0:(genops.c:1622:class_disconnect_stale_exports()) Skipped 1 previous similar message [ 3607.080376] Lustre: lustre-MDT0000: disconnecting 2 stale clients [ 3607.086366] Lustre: lustre-MDT0000-osd: cancel update llog [0x200017b00:0x1:0x0] [ 3607.096334] Lustre: lustre-MDT0001-osp-MDT0000: cancel update llog [0x2400007ea:0x1:0x0] [ 3607.150243] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:1543 to 0x2c0000401:1633) [ 3607.157712] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1571 to 0x280000400:1633) [ 3612.069123] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3612.138138] LustreError: lustre-MDT0000-osp-MDT0001: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 3635.396671] Lustre: DEBUG MARKER: == replay-single test 38: test recovery from unlink llog (test llog_gen_rec) ========================================================== 08:12:56 (1786104776) [ 3669.710558] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 3671.484408] Lustre: Failing over lustre-MDT0000 [ 3671.796400] Lustre: server umount lustre-MDT0000 complete [ 3695.970185] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 3695.972497] LDISKFS-fs (dm-0): recovery complete [ 3695.978050] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3699.214399] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff930987c84a80 x1872862922839168/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 3699.237888] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 13566 previous similar messages [ 3699.491619] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 3699.504515] Lustre: Skipped 9 previous similar messages [ 3704.160981] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3705.126269] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2034 to 0x280000400:2049) [ 3705.129939] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2034 to 0x2c0000401:2049) [ 3713.980257] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 3715.952295] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 3740.062870] Lustre: DEBUG MARKER: == replay-single test 39: test recovery from unlink llog (test llog_gen_rec) ========================================================== 08:14:40 (1786104880) [ 3765.712364] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 3776.061224] Lustre: Failing over lustre-MDT0000 [ 3776.320962] Lustre: server umount lustre-MDT0000 complete [ 3798.845524] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 3798.848265] LDISKFS-fs (dm-0): recovery complete [ 3798.854125] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 3814.045553] Lustre: lustre-MDT0000-lwp-OST0000: Connection restored to 0@lo (at 0@lo) [ 3814.117969] Lustre: Skipped 41 previous similar messages [ 3814.364029] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3820.010171] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2450 to 0x280000400:2465) [ 3820.010419] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2450 to 0x2c0000401:2465) [ 3825.353203] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 3827.122889] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 3848.640861] Lustre: DEBUG MARKER: == replay-single test 41: read from a valid osc while other oscs are invalid ========================================================== 08:16:30 (1786104990) [ 3851.558410] Lustre: setting import lustre-OST0001_UUID INACTIVE by administrator request [ 3852.666230] Lustre: lustre-OST0001: Client lustre-MDT0000-mdtlov_UUID (at 0@lo) reconnecting [ 3852.715059] LustreError: lustre-OST0001-osc-MDT0000: This client was evicted by lustre-OST0001; in progress operations using this service will fail. [ 3858.910839] Lustre: DEBUG MARKER: == replay-single test 42: recovery after ost failure ===== 08:16:40 (1786105000) [ 3885.385682] Lustre: DEBUG MARKER: ost1 REPLAY BARRIER on lustre-OST0000 [ 3896.090856] Lustre: Failing over lustre-OST0000 [ 3896.290306] Lustre: lustre-OST0000-osc-MDT0001: Connection to lustre-OST0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 3896.304091] Lustre: Skipped 36 previous similar messages [ 3896.311280] Lustre: lustre-OST0000: Not available for connect from 0@lo (stopping) [ 3896.316826] Lustre: Skipped 1 previous similar message [ 3898.228160] Lustre: server umount lustre-OST0000 complete [ 3921.456157] LDISKFS-fs (dm-2): 3 truncates cleaned up [ 3921.459054] LDISKFS-fs (dm-2): recovery complete [ 3921.465278] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 3922.672343] Lustre: lustre-OST0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 3922.679893] Lustre: Skipped 3 previous similar messages [ 3923.786795] Lustre: lustre-OST0000: Recovery over after 0:01, of 3 clients 3 recovered and 0 were evicted. [ 3923.793874] Lustre: Skipped 3 previous similar messages [ 3928.936745] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 3989.134918] Lustre: DEBUG MARKER: == replay-single test 43: mds osc import failure during recovery; don't LBUG ========================================================== 08:18:50 (1786105130) [ 3996.274694] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 3998.874856] Lustre: Failing over lustre-MDT0000 [ 3999.339152] Lustre: server umount lustre-MDT0000 complete [ 3999.714433] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 3999.718520] LustreError: Skipped 5 previous similar messages [ 4009.711220] LustreError: 6527:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 192.168.203.38@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 4009.719747] LustreError: 6527:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 137 previous similar messages [ 4020.191145] Lustre: 3648:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786105147/real 1786105147] req@ffff930a8a7ed500 x1872862923352448/t0(0) o400->MGC192.168.203.138@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786105163 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4020.212570] Lustre: 3648:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 4020.219806] LustreError: MGC192.168.203.138@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 4020.230786] LustreError: Skipped 7 previous similar messages [ 4021.239615] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 4021.241852] LDISKFS-fs (dm-0): recovery complete [ 4021.246931] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4034.388661] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4036.151474] Lustre: *** cfs_fail_loc=204, val=2147483648*** [ 4036.154523] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2867 to 0x280000400:2913) [ 4042.938152] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4044.672979] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4052.612237] LustreError: 116720:0:(osp_precreate.c:974:osp_precreate_cleanup_orphans()) lustre-OST0001-osc-MDT0000: cannot cleanup orphans: rc = -11 [ 4052.620408] Lustre: lustre-OST0001: Client lustre-MDT0000-mdtlov_UUID (at 0@lo) reconnecting [ 4053.665270] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2866 to 0x2c0000401:2881) [ 4063.718814] Lustre: DEBUG MARKER: == replay-single test 44a: race in target handle connect ========================================================== 08:20:05 (1786105205) [ 4068.410535] LustreError: 9535:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4073.439130] LustreError: 9535:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4073.443507] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4073.531266] LustreError: 43916:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 waking [ 4075.477469] LustreError: 6528:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4080.607164] LustreError: 6528:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4080.618691] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4082.450096] LustreError: 11965:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4087.775171] LustreError: 11965:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4087.790528] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4090.070988] LustreError: 6527:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4095.457066] LustreError: 6527:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4097.478890] LustreError: 43903:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4102.629628] LustreError: 43903:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4102.640165] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4102.644141] Lustre: Skipped 1 previous similar message [ 4111.032271] LustreError: 6529:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4111.047944] LustreError: 6529:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 1 previous similar message [ 4116.447149] LustreError: 6529:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4116.465587] LustreError: 6529:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 1 previous similar message [ 4123.615839] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4123.624895] Lustre: Skipped 2 previous similar messages [ 4131.936400] LustreError: 43903:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_race id 701 sleeping [ 4131.942557] LustreError: 43903:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 2 previous similar messages [ 4137.524219] LustreError: 43903:0:(ldlm_lib.c:1178:target_handle_connect()) cfs_fail_race id 701 awake: rc=0 [ 4137.531593] LustreError: 43903:0:(ldlm_lib.c:1178:target_handle_connect()) Skipped 2 previous similar messages [ 4146.535294] Lustre: DEBUG MARKER: == replay-single test 44b: race in target handle connect ========================================================== 08:21:28 (1786105288) [ 4148.810495] LustreError: 9535:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4159.229256] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4160.264045] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4164.272181] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4169.516270] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4174.570903] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4174.585292] Lustre: Skipped 1 previous similar message [ 4184.818186] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4184.823666] Lustre: Skipped 1 previous similar message [ 4188.815485] LustreError: 9535:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4188.826227] Lustre: 9535:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff930ab2963b80 x1872862904118144/t0(0) o38->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:0/0 lens 520/416 e 0 to 0 dl 1786105312 ref 1 fl Complete:H/200/0 rc 0/0 job:'lctl.0' uid:0 gid:0 projid:4294967295 [ 4189.933172] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4189.941406] Lustre: Skipped 3 previous similar messages [ 4189.950433] LustreError: 6528:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4214.507563] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4230.047130] LustreError: 6528:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4230.062038] Lustre: 6528:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff930a97c4c380 x1872862904122752/t0(0) o38->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:0/0 lens 520/416 e 0 to 0 dl 1786105353 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4234.988851] LustreError: 6527:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4259.774326] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4259.779086] Lustre: Skipped 4 previous similar messages [ 4275.047300] LustreError: 6527:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4275.055199] Lustre: 6527:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff930ab9593480 x1872862904126336/t0(0) o38->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:0/0 lens 520/416 e 0 to 0 dl 1786105398 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4275.434536] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4275.441349] Lustre: Skipped 1 previous similar message [ 4275.443920] LustreError: 43903:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4315.543131] LustreError: 43903:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4315.549049] Lustre: 43903:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff930ab7e6b800 x1872862904129536/t0(0) o38->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:0/0 lens 520/416 e 0 to 0 dl 1786105438 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4316.906967] LustreError: 6528:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4341.548620] Lustre: lustre-MDT0000: Export ffff930ab199c000 already connecting from 192.168.203.38@tcp [ 4341.558295] Lustre: Skipped 7 previous similar messages [ 4356.927236] LustreError: 6528:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 awake [ 4356.938861] Lustre: 6528:0:(service.c:2582:ptlrpc_server_handle_request()) @@@ Request took longer than estimated (20/20s); client may timeout req@ffff930984cfdf80 x1872862904132992/t0(0) o38->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:0/0 lens 520/416 e 0 to 0 dl 1786105480 ref 1 fl Complete:H/200/0 rc 0/0 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4361.966306] LustreError: 6527:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout id 704 sleeping for 40000ms [ 4376.953791] LustreError: 6527:0:(ldlm_lib.c:1436:target_handle_connect()) cfs_fail_timeout interrupted [ 4382.293795] Lustre: DEBUG MARKER: == replay-single test 44c: race in target handle connect ========================================================== 08:25:23 (1786105523) [ 4391.099483] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 4395.658393] Lustre: Failing over lustre-MDT0000 [ 4395.980978] Lustre: server umount lustre-MDT0000 complete [ 4408.007094] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 4408.009694] LDISKFS-fs (dm-0): recovery complete [ 4408.021578] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4408.375920] Lustre: *** cfs_fail_loc=712, val=0*** [ 4408.385337] LustreError: 8443:0:(service.c:1394:ptlrpc_check_req()) @@@ Invalid replay without recovery req@ffff930ab92ac000 x1872862923558144/t0(0) o400->lustre-MDT0000-mdtlov_UUID@0@lo:0/0 lens 224/0 e 0 to 0 dl 0 ref 1 fl New:/2c0/ffffffff rc 0/-1 job:'ptlrpcd_rcv.0' uid:0 gid:0 projid:4294967295 [ 4408.505468] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 4408.510954] Lustre: Skipped 3 previous similar messages [ 4408.539475] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 4408.539974] Lustre: lustre-MDT0000: Aborting client recovery [ 4408.543141] Lustre: Skipped 15 previous similar messages [ 4408.544969] LustreError: 121572:0:(ldlm_lib.c:3004:target_stop_recovery_thread()) lustre-MDT0000: Aborting recovery [ 4408.552794] Lustre: 121605:0:(ldlm_lib.c:2404:target_recovery_overseer()) recovery is aborted, evict exports in recovery [ 4408.558659] Lustre: 121605:0:(ldlm_lib.c:2404:target_recovery_overseer()) Skipped 2 previous similar messages [ 4408.563418] Lustre: 121605:0:(genops.c:1622:class_disconnect_stale_exports()) lustre-MDT0000: disconnect stale client lustre-MDT0001-mdtlov_UUID@ [ 4408.569436] Lustre: 121605:0:(genops.c:1622:class_disconnect_stale_exports()) Skipped 1 previous similar message [ 4408.580839] Lustre: lustre-MDT0000: disconnecting 2 stale clients [ 4408.591690] Lustre: lustre-MDT0000-osd: cancel update llog [0x2000182d0:0x1:0x0] [ 4408.641605] Lustre: lustre-MDT0001-osp-MDT0000: cancel update llog [0x2400007eb:0x1:0x0] [ 4408.721515] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2866 to 0x2c0000401:2913) [ 4408.738210] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2867 to 0x280000400:2945) [ 4412.624883] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4413.932618] LustreError: lustre-MDT0000-osp-MDT0001: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 4428.311834] Lustre: Failing over lustre-MDT0000 [ 4428.528217] Lustre: lustre-MDT0000: Not available for connect from 192.168.203.38@tcp (stopping) [ 4428.551170] Lustre: Skipped 1 previous similar message [ 4428.718666] Lustre: server umount lustre-MDT0000 complete [ 4448.571233] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4454.898647] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff930ab7e6b100 x1872862923583744/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4454.910965] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 1 previous similar message [ 4460.081262] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4460.553547] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 4460.563164] Lustre: Skipped 15 previous similar messages [ 4460.659121] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2866 to 0x2c0000401:2945) [ 4460.659782] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2867 to 0x280000400:2977) [ 4469.971265] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4471.636841] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4479.733949] Lustre: DEBUG MARKER: == replay-single test 45: Handle failed close ============ 08:27:00 (1786105620) [ 4479.809589] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 4479.815213] Lustre: Skipped 2 previous similar messages [ 4487.660819] Lustre: DEBUG MARKER: == replay-single test 46: Don't leak file handle after open resend (3325) ========================================================== 08:27:08 (1786105628) [ 4488.726537] Lustre: *** cfs_fail_loc=122, val=2147483648*** [ 4488.736584] LustreError: 6537:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930ac0e8f480 x1872862904197376/t0(0) o700->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:633/0 lens 264/248 e 0 to 0 dl 1786105643 ref 1 fl Interpret:/200/0 rc 0/0 job:'touch.0' uid:0 gid:0 projid:4294967295 [ 4507.459810] Lustre: Failing over lustre-MDT0000 [ 4507.646343] Lustre: server umount lustre-MDT0000 complete [ 4511.721120] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 4511.734818] Lustre: Skipped 14 previous similar messages [ 4525.577564] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4527.292961] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 4527.303330] Lustre: Skipped 2 previous similar messages [ 4530.194804] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4531.215820] Lustre: lustre-MDT0000: Recovery over after 0:04, of 2 clients 2 recovered and 0 were evicted. [ 4531.226148] Lustre: Skipped 2 previous similar messages [ 4531.259739] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2947 to 0x2c0000401:2977) [ 4531.262183] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2979 to 0x280000400:3009) [ 4539.410255] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4541.193398] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4552.451871] Lustre: DEBUG MARKER: == replay-single test 47: MDS->OSC failure during precreate cleanup (2824) ========================================================== 08:28:13 (1786105693) [ 4554.719618] Lustre: Failing over lustre-OST0000 [ 4554.941831] Lustre: server umount lustre-OST0000 complete [ 4572.891139] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 4579.171453] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4588.715853] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 4590.190751] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 4661.402307] Lustre: DEBUG MARKER: == replay-single test 48: MDS->OSC failure during precreate cleanup (2824) ========================================================== 08:30:02 (1786105802) [ 4668.362874] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 4670.669053] Lustre: Failing over lustre-MDT0000 [ 4670.882370] Lustre: server umount lustre-MDT0000 complete [ 4671.232605] LustreError: 6529:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 192.168.203.38@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 4671.269818] LustreError: 6529:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 100 previous similar messages [ 4674.535251] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 4674.539238] LustreError: Skipped 2 previous similar messages [ 4690.895294] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786105818/real 1786105818] req@ffff930a893c0380 x1872862923710208/t0(0) o400->MGC192.168.203.138@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786105834 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4690.912470] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 4690.918794] LustreError: MGC192.168.203.138@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 4690.925924] LustreError: Skipped 3 previous similar messages [ 4693.041890] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 4693.046644] LDISKFS-fs (dm-0): recovery complete [ 4693.128556] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4707.171899] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4708.545704] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:2998 to 0x2c0000401:3041) [ 4708.548202] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3030 to 0x280000400:3073) [ 4781.013297] Lustre: DEBUG MARKER: == replay-single test 50: Double OSC recovery, don't LASSERT (3812) ========================================================== 08:32:02 (1786105922) [ 4782.762587] Lustre: lustre-OST0000: Client lustre-MDT0000-mdtlov_UUID (at 0@lo) reconnecting [ 4782.766548] Lustre: Skipped 2 previous similar messages [ 4796.189328] Lustre: DEBUG MARKER: == replay-single test 52: time out lock replay (3764) ==== 08:32:17 (1786105937) [ 4799.085749] Lustre: Failing over lustre-MDT0000 [ 4799.341467] Lustre: server umount lustre-MDT0000 complete [ 4816.883431] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4820.937638] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4822.517454] Lustre: *** cfs_fail_loc=157, val=2147483648*** [ 4822.523258] LustreError: 129323:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930ab9698000 x1872862904323456/t0(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:254/0 lens 328/344 e 0 to 0 dl 1786106019 ref 1 fl Complete:/240/0 rc 0/0 job:'ldlm_lock_repla.0' uid:0 gid:0 projid:4294967295 [ 4880.626717] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnected, waiting for 2 clients in recovery for 1:04 [ 4880.641906] Lustre: 129323:0:(ldlm_lib.c:2085:extend_recovery_timer()) lustre-MDT0000: extended recovery timer reached hard limit: 180, extend: 1 [ 4880.722600] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3053 to 0x2c0000401:3073) [ 4880.736992] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3084 to 0x280000400:3105) [ 4888.226805] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4889.870793] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4902.533567] Lustre: DEBUG MARKER: == replay-single test 53a: |X| close request while two MDC requests in flight ========================================================== 08:34:03 (1786106043) [ 4904.676666] Lustre: *** cfs_fail_loc=115, val=2147483648*** [ 4913.813750] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 4916.042510] Lustre: Failing over lustre-MDT0000 [ 4916.399628] Lustre: server umount lustre-MDT0000 complete [ 4938.746427] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 4938.751051] LDISKFS-fs (dm-0): recovery complete [ 4938.756239] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 4952.199935] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 4952.608539] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3053 to 0x2c0000401:3105) [ 4952.612216] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3107 to 0x280000400:3137) [ 4962.152879] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 4963.568108] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4974.677266] Lustre: DEBUG MARKER: == replay-single test 53b: |X| open request while two MDC requests in flight ========================================================== 08:35:15 (1786106115) [ 4975.768584] Lustre: *** cfs_fail_loc=107, val=2147483648*** [ 4985.740806] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 4987.932078] Lustre: Failing over lustre-MDT0000 [ 4988.124021] Lustre: server umount lustre-MDT0000 complete [ 5011.921121] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5011.923882] LDISKFS-fs (dm-0): recovery complete [ 5011.929722] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5014.051744] LustreError: 133447:0:(import.c:339:ptlrpc_invalidate_import()) MGS: timeout waiting for callback (1 != 0) [ 5014.061176] LustreError: 133447:0:(import.c:363:ptlrpc_invalidate_import()) @@@ still on sending list req@ffff930ab969bb80 x1872862923882368/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 1786106157 ref 1 fl Rpc:NQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5014.091175] LustreError: 133447:0:(import.c:373:ptlrpc_invalidate_import()) MGS: Unregistering RPCs found (0). Network is sluggish? Waiting for them to error out. [ 5015.475974] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 5015.506674] Lustre: Skipped 6 previous similar messages [ 5015.546855] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 5015.557020] Lustre: Skipped 8 previous similar messages [ 5020.569859] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5020.724191] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3107 to 0x2c0000401:3137) [ 5020.726406] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3107 to 0x280000400:3169) [ 5031.998508] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5033.555865] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5043.669571] Lustre: DEBUG MARKER: == replay-single test 53c: |X| open request and close request while two MDC requests in flight ========================================================== 08:36:25 (1786106185) [ 5044.945243] Lustre: *** cfs_fail_loc=107, val=2147483648*** [ 5055.172782] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5057.295179] Lustre: Failing over lustre-MDT0000 [ 5057.470989] Lustre: server umount lustre-MDT0000 complete [ 5080.121541] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5080.125514] LDISKFS-fs (dm-0): recovery complete [ 5080.132523] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5088.253618] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff930ab7ecbb80 x1872862923923456/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5093.514756] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5093.864590] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 5093.869590] Lustre: Skipped 27 previous similar messages [ 5093.924712] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3107 to 0x2c0000401:3169) [ 5093.926088] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3171 to 0x280000400:3201) [ 5107.180918] Lustre: DEBUG MARKER: == replay-single test 53d: close reply while two MDC requests in flight ========================================================== 08:37:28 (1786106248) [ 5109.245127] Lustre: *** cfs_fail_loc=13b, val=315*** [ 5109.255935] Lustre: *** cfs_fail_loc=13b, val=2147483648*** [ 5109.269436] LustreError: 43658:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930a97de5c00 x1872862904386048/t257698037777(0) o35->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:498/0 lens 392/456 e 0 to 0 dl 1786106263 ref 1 fl Interpret:/600/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5111.783864] Lustre: Failing over lustre-MDT0000 [ 5112.005973] Lustre: server umount lustre-MDT0000 complete [ 5114.340725] Lustre: lustre-MDT0000-lwp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 5114.354795] Lustre: Skipped 29 previous similar messages [ 5132.373290] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5140.961509] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xdb52b77d7c8ab331 [ 5140.967653] Lustre: Skipped 3 previous similar messages [ 5142.017531] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 5142.038373] Lustre: Skipped 6 previous similar messages [ 5145.691321] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5146.639444] Lustre: lustre-MDT0000: Recovery over after 0:04, of 2 clients 2 recovered and 0 were evicted. [ 5146.646231] Lustre: Skipped 6 previous similar messages [ 5146.681535] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3107 to 0x2c0000401:3201) [ 5146.682816] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3203 to 0x280000400:3233) [ 5146.748605] Lustre: 6531:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff930a893c1500 x1872862904386048/t257698037777(0) o35->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:536/0 lens 392/456 e 0 to 0 dl 1786106301 ref 1 fl Interpret:/602/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5155.512725] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5157.566560] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5168.227207] Lustre: DEBUG MARKER: == replay-single test 53e: |X| open reply while two MDC requests in flight ========================================================== 08:38:29 (1786106309) [ 5170.094690] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 5170.098627] LustreError: 9535:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930ac1044e00 x1872862904402304/t261993005072(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:603/0 lens 504/448 e 0 to 0 dl 1786106368 ref 1 fl Interpret:/200/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5179.145118] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5180.991452] Lustre: Failing over lustre-MDT0000 [ 5181.156913] Lustre: server umount lustre-MDT0000 complete [ 5203.964750] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5203.967757] LDISKFS-fs (dm-0): recovery complete [ 5203.973229] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5212.318523] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5214.272446] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3203 to 0x2c0000401:3233) [ 5214.285036] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3203 to 0x280000400:3265) [ 5214.324355] Lustre: 9535:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff93098346b800 x1872862904402304/t261993005072(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:647/0 lens 504/2880 e 0 to 0 dl 1786106412 ref 1 fl Interpret:/202/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5222.360956] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5223.672047] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5233.799904] Lustre: DEBUG MARKER: == replay-single test 53f: |X| open reply and close reply while two MDC requests in flight ========================================================== 08:39:35 (1786106375) [ 5235.245554] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 5235.253212] LustreError: 11965:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930ab7ec9c00 x1872862904418816/t266287972368(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:668/0 lens 504/448 e 0 to 0 dl 1786106433 ref 1 fl Interpret:/200/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5237.176083] Lustre: *** cfs_fail_loc=13b, val=315*** [ 5244.635426] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5246.472302] Lustre: Failing over lustre-MDT0000 [ 5246.622603] Lustre: server umount lustre-MDT0000 complete [ 5269.954575] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5269.957028] LDISKFS-fs (dm-0): recovery complete [ 5269.965899] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5271.276782] LustreError: 6529:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 192.168.203.38@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 5271.286432] LustreError: 6529:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 250 previous similar messages [ 5282.689776] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5283.428851] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3267 to 0x280000400:3297) [ 5283.439184] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3203 to 0x2c0000401:3265) [ 5283.461583] Lustre: 43903:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff930ab9287100 x1872862904418816/t266287972368(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:716/0 lens 504/2880 e 0 to 0 dl 1786106481 ref 1 fl Interpret:/202/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5283.520417] Lustre: 43903:0:(mdt_recovery.c:102:mdt_req_from_lrd()) Skipped 1 previous similar message [ 5295.937748] Lustre: DEBUG MARKER: == replay-single test 53g: |X| drop open reply and close request while close and open are both in flight ========================================================== 08:40:36 (1786106436) [ 5297.353835] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 5297.360506] Lustre: Skipped 1 previous similar message [ 5297.364834] LustreError: 11965:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930a97cef800 x1872862904434048/t270582939664(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:730/0 lens 504/448 e 0 to 0 dl 1786106495 ref 1 fl Interpret:/200/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5297.378049] LustreError: 11965:0:(ldlm_lib.c:3346:target_send_reply_msg()) Skipped 1 previous similar message [ 5299.108820] Lustre: *** cfs_fail_loc=115, val=2147483648*** [ 5299.222961] Lustre: Skipped 1 previous similar message [ 5308.160401] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5310.904379] Lustre: Failing over lustre-MDT0000 [ 5311.202295] Lustre: server umount lustre-MDT0000 complete [ 5314.023704] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 5314.043041] LustreError: Skipped 6 previous similar messages [ 5330.400586] Lustre: 3648:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786106457/real 1786106457] req@ffff930ac0ceea00 x1872862924061056/t0(0) o400->MGC192.168.203.138@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786106473 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5330.438602] Lustre: 3648:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 6 previous similar messages [ 5330.444604] LustreError: MGC192.168.203.138@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 5330.452617] LustreError: Skipped 7 previous similar messages [ 5334.877369] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5334.878975] LDISKFS-fs (dm-0): recovery complete [ 5334.883553] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5340.648749] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xdb52b77d7c8ac65c [ 5345.171311] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5346.341161] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3267 to 0x280000400:3329) [ 5346.347394] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3267 to 0x2c0000401:3297) [ 5346.371180] Lustre: 6528:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff930ac087f480 x1872862904434048/t270582939664(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:24/0 lens 504/2880 e 0 to 0 dl 1786106544 ref 1 fl Interpret:/202/0 rc 0/0 job:'mcreate.0' uid:0 gid:0 projid:4294967295 [ 5356.762598] Lustre: DEBUG MARKER: == replay-single test 53h: open request and close reply while two MDC requests in flight ========================================================== 08:41:37 (1786106497) [ 5357.758912] Lustre: *** cfs_fail_loc=107, val=2147483648*** [ 5359.377459] Lustre: *** cfs_fail_loc=13b, val=315*** [ 5359.383582] Lustre: *** cfs_fail_loc=13b, val=2147483648*** [ 5359.388443] LustreError: 43658:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930a97c4e680 x1872862904449280/t274877906960(0) o35->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:748/0 lens 392/456 e 0 to 0 dl 1786106513 ref 1 fl Interpret:/600/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5366.473477] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5368.302080] Lustre: Failing over lustre-MDT0000 [ 5368.473055] Lustre: server umount lustre-MDT0000 complete [ 5392.104721] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5392.107365] LDISKFS-fs (dm-0): recovery complete [ 5392.112931] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5403.952410] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5404.784345] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3331 to 0x280000400:3361) [ 5404.792706] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3267 to 0x2c0000401:3329) [ 5404.795756] Lustre: 6531:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff930ab92adf80 x1872862904449280/t274877906960(0) o35->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:39/0 lens 392/456 e 0 to 0 dl 1786106559 ref 1 fl Interpret:/602/0 rc 0/0 job:'multiop.0' uid:0 gid:0 projid:0 [ 5418.597433] Lustre: DEBUG MARKER: == replay-single test 55: let MDS_CHECK_RESENT return the original return code instead of 0 ========================================================== 08:42:39 (1786106559) [ 5419.284548] Lustre: *** cfs_fail_loc=12b, val=2147483991*** [ 5419.371722] LustreError: 6527:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930985a90e00 x1872862904461440/t279172874255(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:97/0 lens 664/608 e 0 to 0 dl 1786106617 ref 1 fl Interpret:/600/0 rc 301/0 job:'touch.0' uid:0 gid:0 projid:0 [ 5479.670489] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnecting [ 5479.690808] Lustre: Skipped 1 previous similar message [ 5479.708775] Lustre: 9535:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff930a8a5d1180 x1872862904461440/t279172874255(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:158/0 lens 664/3488 e 0 to 0 dl 1786106678 ref 1 fl Interpret:/602/0 rc 0/0 job:'touch.0' uid:0 gid:0 projid:0 [ 5488.088943] Lustre: DEBUG MARKER: == replay-single test 56: don't replay a symlink open request (3440) ========================================================== 08:43:49 (1786106629) [ 5495.204533] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5497.153933] Lustre: Failing over lustre-MDT0000 [ 5497.375709] Lustre: server umount lustre-MDT0000 complete [ 5518.291177] LDISKFS-fs (dm-0): 4 truncates cleaned up [ 5518.293716] LDISKFS-fs (dm-0): recovery complete [ 5518.300547] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5528.061471] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xdb52b77d7c8ad33d [ 5532.342676] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5533.839064] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3331 to 0x2c0000401:3361) [ 5533.843962] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3331 to 0x280000400:3393) [ 5542.538494] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5544.804600] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5564.752548] Lustre: DEBUG MARKER: == replay-single test 57: test recovery from llog for setattr op ========================================================== 08:45:05 (1786106705) [ 5573.236494] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5575.627820] Lustre: Failing over lustre-MDT0000 [ 5575.840767] Lustre: server umount lustre-MDT0000 complete [ 5599.783485] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5599.785589] LDISKFS-fs (dm-0): recovery complete [ 5599.790217] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5610.530315] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5611.743361] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:3331 to 0x2c0000401:3393) [ 5611.745390] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3395 to 0x280000400:3425) [ 5620.761758] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5622.361898] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5627.750241] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing _wait_recovery_complete *.lustre-MDT0000.recovery_status 1475 [ 5640.542449] Lustre: DEBUG MARKER: == replay-single test 58a: test recovery from llog for setattr op (test llog_gen_rec) ========================================================== 08:46:21 (1786106781) [ 5682.591456] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5684.222520] Lustre: Failing over lustre-MDT0000 [ 5684.553812] Lustre: server umount lustre-MDT0000 complete [ 5706.453992] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5706.456343] LDISKFS-fs (dm-0): recovery complete [ 5706.462459] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5714.911963] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff930ab9288e00 x1872862924272256/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5714.924193] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 24 previous similar messages [ 5715.129266] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 5715.132318] Lustre: Skipped 8 previous similar messages [ 5715.190219] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 5715.199393] Lustre: Skipped 8 previous similar messages [ 5720.546155] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5720.566144] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 5720.577947] Lustre: Skipped 35 previous similar messages [ 5721.259310] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4676 to 0x280000400:4705) [ 5721.271684] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4644 to 0x2c0000401:4673) [ 5731.604392] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5733.929251] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5797.413323] Lustre: DEBUG MARKER: == replay-single test 58b: test replay of setxattr op ==== 08:48:58 (1786106938) [ 5806.168380] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 5808.622714] Lustre: Failing over lustre-MDT0000 [ 5808.911945] Lustre: server umount lustre-MDT0000 complete [ 5812.704754] Lustre: lustre-MDT0000-lwp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 5812.717461] Lustre: Skipped 30 previous similar messages [ 5831.962036] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 5831.965076] LDISKFS-fs (dm-0): recovery complete [ 5831.972764] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 5840.730763] Lustre: lustre-MDT0000: Not available for connect from 192.168.203.38@tcp (not set up) [ 5840.747802] Lustre: Skipped 1 previous similar message [ 5842.235164] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 3 clients reconnect [ 5842.247783] Lustre: Skipped 7 previous similar messages [ 5845.683888] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 5846.031050] Lustre: lustre-MDT0000: Recovery over after 0:04, of 3 clients 3 recovered and 0 were evicted. [ 5846.036176] Lustre: Skipped 7 previous similar messages [ 5846.056468] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4676 to 0x280000400:4737) [ 5846.056572] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4675 to 0x2c0000401:4705) [ 5855.698665] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 5857.471859] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 5868.485930] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount FULL mgc.*.mgs_server_uuid 1475 0 [ 5870.190678] Lustre: DEBUG MARKER: mgc.*.mgs_server_uuid in FULL state after 0 sec [ 5879.195536] Lustre: DEBUG MARKER: == replay-single test 58c: resend/reconstruct setxattr op ========================================================== 08:50:20 (1786107020) [ 5886.898888] Lustre: *** cfs_fail_loc=123, val=2147483648*** [ 5948.688362] Lustre: *** cfs_fail_loc=119, val=2147483648*** [ 5948.698277] Lustre: Skipped 1 previous similar message [ 5948.705644] LustreError: 6527:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930a97fa9880 x1872862907275392/t296352743435(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:626/0 lens 66040/440 e 0 to 0 dl 1786107146 ref 1 fl Interpret:/600/0 rc 0/0 job:'setfattr.0' uid:0 gid:0 projid:0 [ 6009.081100] Lustre: 6529:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff930aaffeb850 x1872862907275392/t296352743435(0) o36->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:687/0 lens 66040/440 e 0 to 0 dl 1786107207 ref 1 fl Interpret:/602/0 rc 0/0 job:'setfattr.0' uid:0 gid:0 projid:0 [ 6019.674909] Lustre: DEBUG MARKER: SKIP: replay-single test_59 skipping ALWAYS excluded test 59 [ 6021.587442] Lustre: DEBUG MARKER: == replay-single test 60: test llog post recovery init vs llog unlink ========================================================== 08:52:42 (1786107162) [ 6034.819738] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 6037.385175] Lustre: Failing over lustre-MDT0000 [ 6037.658452] Lustre: server umount lustre-MDT0000 complete [ 6039.838296] LustreError: 11965:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 192.168.203.38@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 6039.847893] LustreError: 11965:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 228 previous similar messages [ 6040.551326] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 6040.567593] LustreError: Skipped 2 previous similar messages [ 6056.927319] Lustre: 3647:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786107183/real 1786107183] req@ffff930ac0abf100 x1872862925086464/t0(0) o400->MGC192.168.203.138@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786107199 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 6056.976374] Lustre: 3647:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 5 previous similar messages [ 6056.985138] LustreError: MGC192.168.203.138@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 6056.998978] LustreError: Skipped 5 previous similar messages [ 6060.296682] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 6060.298583] LDISKFS-fs (dm-0): recovery complete [ 6060.304138] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6072.851311] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6075.689637] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4806 to 0x2c0000401:4833) [ 6075.695971] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4839 to 0x280000400:4865) [ 6082.559642] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 6084.732628] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6094.633625] Lustre: DEBUG MARKER: == replay-single test 61a: test race llog recovery vs llog cleanup ========================================================== 08:53:55 (1786107235) [ 6117.830569] Lustre: DEBUG MARKER: ost1 REPLAY BARRIER on lustre-OST0000 [ 6131.160449] Lustre: Failing over lustre-OST0000 [ 6131.187921] Lustre: lustre-OST0000: Not available for connect from 0@lo (stopping) [ 6131.280639] Lustre: server umount lustre-OST0000 complete [ 6153.170456] LDISKFS-fs (dm-2): 3 truncates cleaned up [ 6153.173136] LDISKFS-fs (dm-2): recovery complete [ 6153.179669] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 6159.346639] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6174.441814] Lustre: Failing over lustre-OST0000 [ 6174.640895] Lustre: server umount lustre-OST0000 complete [ 6192.607904] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 6198.951898] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6208.382435] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 6209.955816] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 6250.223947] Lustre: DEBUG MARKER: == replay-single test 61b: test race mds llog sync vs llog cleanup ========================================================== 08:56:31 (1786107391) [ 6253.111397] Lustre: Failing over lustre-MDT0000 [ 6253.348719] Lustre: server umount lustre-MDT0000 complete [ 6272.867278] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6280.162277] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xdb52b77d7c9021f3 [ 6280.169638] Lustre: Skipped 1 previous similar message [ 6284.583968] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6285.864564] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4839 to 0x280000400:4897) [ 6285.865654] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4806 to 0x2c0000401:4865) [ 6299.481638] Lustre: Failing over lustre-MDT0000 [ 6299.723457] Lustre: server umount lustre-MDT0000 complete [ 6317.464988] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6326.815906] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff930ab2cf2680 x1872862925371008/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 6326.984939] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 6326.990225] Lustre: Skipped 5 previous similar messages [ 6327.017118] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 6327.029972] Lustre: Skipped 5 previous similar messages [ 6330.861477] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6332.413667] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 6332.420700] Lustre: Skipped 20 previous similar messages [ 6332.501491] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4806 to 0x2c0000401:4897) [ 6332.503544] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4839 to 0x280000400:4929) [ 6340.492711] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 6342.494545] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6352.889900] Lustre: DEBUG MARKER: == replay-single test 61c: test race mds llog sync vs llog cleanup ========================================================== 08:58:13 (1786107493) [ 6367.939926] Lustre: Failing over lustre-OST0000 [ 6368.030933] Lustre: server umount lustre-OST0000 complete [ 6387.957266] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 6394.646125] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6403.675252] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 6406.040632] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 6417.267254] Lustre: DEBUG MARKER: == replay-single test 61d: error in llog_setup should cleanup the llog context correctly ========================================================== 08:59:18 (1786107558) [ 6419.397630] Lustre: Failing over lustre-MDT0000 [ 6419.435866] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 6419.488554] Lustre: Skipped 20 previous similar messages [ 6419.500588] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 6419.762481] Lustre: server umount lustre-MDT0000 complete [ 6428.986613] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6429.042680] Lustre: *** cfs_fail_loc=605, val=0*** [ 6429.045000] LustreError: 164892:0:(llog_obd.c:192:llog_setup()) MGS: ctxt 0 lop_setup=ffffffffc0a34110 failed: rc = -95 [ 6429.050153] LustreError: 164892:0:(obd_config.c:845:class_setup()) setup MGS failed (-95) [ 6429.053590] LustreError: 164892:0:(obd_mount.c:259:lustre_start_simple()) MGS setup error -95 [ 6429.060503] LustreError: 164892:0:(tgt_mount.c:116:server_deregister_mount()) MGS not registered [ 6429.067180] LustreError: Failed to start MGS 'MGS' (-95). Is the 'mgs' module loaded? [ 6429.071418] LustreError: 164892:0:(tgt_mount.c:2129:server_put_super()) no obd lustre-MDT0000 [ 6429.080645] Lustre: server umount lustre-MDT0000 complete [ 6429.082667] LustreError: 164892:0:(super25.c:178:lustre_fill_super()) llite: Unable to mount : rc = -95 [ 6437.249434] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6441.882709] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6443.064691] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4899 to 0x2c0000401:4929) [ 6443.066817] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4931 to 0x280000400:4961) [ 6450.190341] Lustre: DEBUG MARKER: == replay-single test 62: don't mis-drop resent replay === 08:59:51 (1786107591) [ 6457.387788] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 6460.765833] Lustre: Failing over lustre-MDT0000 [ 6460.956670] Lustre: server umount lustre-MDT0000 complete [ 6485.710940] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 6485.713212] LDISKFS-fs (dm-0): recovery complete [ 6485.718292] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 6491.788032] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 6491.803169] Lustre: Skipped 7 previous similar messages [ 6491.941437] Lustre: *** cfs_fail_loc=707, val=0*** [ 6495.323144] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 6551.886676] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnected, waiting for 2 clients in recovery for 0:54 [ 6551.910390] Lustre: 167067:0:(ldlm_lib.c:2085:extend_recovery_timer()) lustre-MDT0000: extended recovery timer reached hard limit: 180, extend: 1 [ 6551.928215] Lustre: 167067:0:(ldlm_lib.c:2085:extend_recovery_timer()) Skipped 1 previous similar message [ 6552.205473] Lustre: lustre-MDT0000: Recovery over after 1:01, of 2 clients 2 recovered and 0 were evicted. [ 6552.212555] Lustre: Skipped 7 previous similar messages [ 6552.309501] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4974 to 0x280000400:4993) [ 6552.321266] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:4943 to 0x2c0000401:4961) [ 6558.376550] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 6560.393528] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6570.605989] Lustre: DEBUG MARKER: == replay-single test 65a: AT: verify early replies ====== 09:01:51 (1786107711) [ 6601.069891] LustreError: 6528:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff93098642b800 x1872862908182272/t0(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:480/0 lens 664/0 e 0 to 0 dl 1786107755 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6601.089158] LustreError: 6528:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 11000ms [ 6612.143209] LustreError: 6528:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6612.164918] LustreError: 6530:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff930985b9c700 x1872862908183552/t0(0) o35->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:491/0 lens 392/0 e 0 to 0 dl 1786107766 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6614.976505] LustreError: 6528:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff930982e52300 x1872862908189952/t0(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:533/0 lens 576/0 e 0 to 0 dl 1786107808 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'unlinkmany.0' uid:0 gid:0 projid:0 [ 6615.005844] LustreError: 6528:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 27 previous similar messages [ 6618.529560] LustreError: 43919:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff930ac0e82300 x1872862925529728/t0(0) o6->lustre-MDT0000-mdtlov_UUID@0@lo:497/0 lens 544/0 e 0 to 0 dl 1786107772 ref 1 fl Interpret:/200/ffffffff rc 0/-1 job:'osp-syn-1-0.0' uid:0 gid:0 projid:4294967295 [ 6618.544986] LustreError: 43919:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 20 previous similar messages [ 6624.225053] LustreError: 151898:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff930984d4ad80 x1872862925534336/t0(0) o6->lustre-MDT0000-mdtlov_UUID@0@lo:503/0 lens 544/0 e 0 to 0 dl 1786107778 ref 1 fl Interpret:/200/ffffffff rc 0/-1 job:'osp-syn-0-0.0' uid:0 gid:0 projid:4294967295 [ 6624.238278] LustreError: 151898:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 7 previous similar messages [ 6631.720802] Lustre: DEBUG MARKER: == replay-single test 65b: AT: verify early replies on packed reply / bulk ========================================================== 09:02:53 (1786107773) [ 6664.169276] LustreError: 8448:0:(tgt_handler.c:2830:tgt_brw_write()) cfs_fail_timeout id 224 sleeping for 11000ms [ 6675.215133] LustreError: 8448:0:(tgt_handler.c:2830:tgt_brw_write()) cfs_fail_timeout id 224 awake [ 6687.168578] Lustre: DEBUG MARKER: == replay-single test 66a: AT: verify MDT service time adjusts with no early replies ========================================================== 09:03:48 (1786107828) [ 6717.742655] LustreError: 9535:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff930ab8f29500 x1872862908211456/t0(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:597/0 lens 576/0 e 0 to 0 dl 1786107872 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6717.786780] LustreError: 9535:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 5000ms [ 6722.863306] LustreError: 9535:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6725.813187] LustreError: 11965:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 10000ms [ 6735.895462] LustreError: 11965:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6735.945995] LustreError: 43903:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff930984cfd880 x1872862908234112/t0(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:654/0 lens 664/0 e 0 to 0 dl 1786107929 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6735.965892] LustreError: 43903:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 120 previous similar messages [ 6757.407127] Lustre: DEBUG MARKER: == replay-single test 66b: AT: verify net latency adjusts ========================================================== 09:04:58 (1786107898) [ 6852.894671] Lustre: DEBUG MARKER: == replay-single test 67a: AT: verify slow request processing doesn't induce reconnects ========================================================== 09:06:33 (1786107993) [ 6883.303940] LustreError: 6528:0:(service.c:2558:ptlrpc_server_handle_request()) @@@ HIT req@ffff930ac0abf800 x1872862908294272/t0(0) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:7/0 lens 576/0 e 0 to 0 dl 1786108037 ref 1 fl Interpret:/600/ffffffff rc 0/-1 job:'createmany.0' uid:0 gid:0 projid:0 [ 6883.372658] LustreError: 6528:0:(service.c:2558:ptlrpc_server_handle_request()) Skipped 98 previous similar messages [ 6883.384556] LustreError: 6528:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 400ms [ 6883.815613] LustreError: 6528:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6900.007418] LustreError: 6529:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a awake [ 6900.019493] LustreError: 6529:0:(service.c:2559:ptlrpc_server_handle_request()) Skipped 35 previous similar messages [ 6915.642781] LustreError: 6529:0:(service.c:2559:ptlrpc_server_handle_request()) cfs_fail_timeout id 50a sleeping for 400ms [ 6915.647308] LustreError: 6529:0:(service.c:2559:ptlrpc_server_handle_request()) Skipped 74 previous similar messages [ 6937.047667] Lustre: DEBUG MARKER: == replay-single test 67b: AT: verify instant slowdown doesn't induce reconnects ========================================================== 09:07:58 (1786108078) [ 6973.049418] Lustre: DEBUG MARKER: phase 2 [ 6987.125498] Lustre: DEBUG MARKER: == replay-single test 68: AT: verify slowing locks ======= 09:08:47 (1786108127) [ 7073.548629] Lustre: DEBUG MARKER: == replay-single test 70a: check multi client t-f ======== 09:10:14 (1786108214) [ 7075.467755] Lustre: DEBUG MARKER: SKIP: replay-single test_70a Need two or more clients, have 1 [ 7077.094847] Lustre: DEBUG MARKER: == replay-single test 70b: dbench 2mdts recovery; 1 clients ========================================================== 09:10:18 (1786108218) [ 7082.827570] Lustre: DEBUG MARKER: Started rundbench load pid=135434 ... [ 7093.107513] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 7096.766874] Lustre: DEBUG MARKER: test_70b fail mds1 1 times [ 7099.016603] Lustre: Failing over lustre-MDT0000 [ 7099.108236] Lustre: lustre-MDT0000: Not available for connect from 192.168.203.38@tcp (stopping) [ 7099.727345] Lustre: server umount lustre-MDT0000 complete [ 7100.384338] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 7100.396354] LustreError: Skipped 15 previous similar messages [ 7100.404685] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 7100.412175] Lustre: Skipped 7 previous similar messages [ 7100.446709] LustreError: 7804:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 7100.459448] LustreError: 7804:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 183 previous similar messages [ 7119.327367] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786108246/real 1786108246] req@ffff93098642b480 x1872862925809792/t0(0) o400->MGC192.168.203.138@tcp@0@lo:26/25 lens 224/224 e 0 to 1 dl 1786108262 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 7119.337837] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 3 previous similar messages [ 7119.341723] LustreError: MGC192.168.203.138@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 7119.346581] LustreError: Skipped 4 previous similar messages [ 7123.157960] LDISKFS-fs (dm-0): 4 truncates cleaned up [ 7123.160432] LDISKFS-fs (dm-0): recovery complete [ 7123.169233] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7129.573287] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff9309825aea00 x1872862925819264/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 7129.854138] Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 [ 7129.858911] Lustre: Skipped 3 previous similar messages [ 7129.880423] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 7129.891671] Lustre: Skipped 3 previous similar messages [ 7131.380896] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 7134.498989] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7135.253480] Lustre: lustre-MDT0000-lwp-MDT0001: Connection restored to 0@lo (at 0@lo) [ 7135.282474] Lustre: Skipped 13 previous similar messages [ 7135.938896] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:5042 to 0x2c0000401:5057) [ 7135.939652] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:5078 to 0x280000400:5121) [ 7146.273620] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 7148.420698] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7159.439442] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 7162.095305] Lustre: DEBUG MARKER: test_70b fail mds2 2 times [ 7164.870101] Lustre: Failing over lustre-MDT0001 [ 7165.203712] Lustre: server umount lustre-MDT0001 complete [ 7192.361224] LDISKFS-fs (dm-1): 8 truncates cleaned up [ 7192.363755] LDISKFS-fs (dm-1): recovery complete [ 7192.374605] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7197.720898] Lustre: lustre-MDT0001: Recovery over after 0:05, of 2 clients 2 recovered and 0 were evicted. [ 7197.729872] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7197.734484] Lustre: Skipped 1 previous similar message [ 7197.801062] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000401:421 to 0x280000401:449) [ 7197.801216] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:403 to 0x2c0000400:449) [ 7209.629091] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 7211.547862] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7233.656748] Lustre: DEBUG MARKER: == replay-single test 70c: tar 2mdts recovery ============ 09:12:55 (1786108375) [ 7364.817444] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 7378.409875] Lustre: DEBUG MARKER: test_70c fail mds2 1 times [ 7380.366877] Lustre: Failing over lustre-MDT0001 [ 7380.451709] Lustre: lustre-MDT0001: Not available for connect from 192.168.203.38@tcp (stopping) [ 7380.736762] Lustre: server umount lustre-MDT0001 complete [ 7407.286860] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 7407.289436] LDISKFS-fs (dm-1): recovery complete [ 7407.298691] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7413.167030] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7416.524397] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:730 to 0x2c0000400:769) [ 7416.524722] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000401:729 to 0x280000401:769) [ 7425.253516] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 7427.325787] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7483.051707] Lustre: DEBUG MARKER: == replay-single test 70d: mkdir/rmdir striped dir 2mdts recovery ========================================================== 09:17:04 (1786108624) [ 7615.495822] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 7628.077711] Lustre: DEBUG MARKER: test_70d fail mds2 1 times [ 7631.197226] Lustre: Failing over lustre-MDT0001 [ 7631.281941] Lustre: lustre-MDT0001: Not available for connect from 192.168.203.38@tcp (stopping) [ 7631.646596] Lustre: server umount lustre-MDT0001 complete [ 7631.860078] LustreError: 9538:0:(ldlm_lockd.c:2564:ldlm_cancel_handler()) ldlm_cancel from 0@lo arrived at 1786108775 with bad export cookie 15803895791985479087 [ 7655.461402] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 7655.466712] LDISKFS-fs (dm-1): recovery complete [ 7655.478631] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7655.875722] Lustre: lustre-MDT0001: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 7660.782324] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7663.382133] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000401:885 to 0x280000401:929) [ 7663.386655] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:886 to 0x2c0000400:929) [ 7675.056086] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 7678.362294] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7688.992663] Lustre: DEBUG MARKER: == replay-single test 70e: rename cross-MDT with random fails ========================================================== 09:20:30 (1786108830) [ 7820.420835] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 7833.140383] Lustre: DEBUG MARKER: test_70e fail mds2 1 times [ 7835.583360] Lustre: Failing over lustre-MDT0001 [ 7835.818317] Lustre: server umount lustre-MDT0001 complete [ 7836.415935] LustreError: 6528:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0001: not available for connect from 192.168.203.38@tcp (no target). If you are running an HA pair check that the target is mounted on the other server. [ 7836.427196] LustreError: 6528:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 99 previous similar messages [ 7837.716172] LustreError: lustre-MDT0001-osp-MDT0000: operation mds_statfs to node 0@lo failed: rc = -107 [ 7837.722018] LustreError: Skipped 2 previous similar messages [ 7837.726886] Lustre: lustre-MDT0001-osp-MDT0000: Connection to lustre-MDT0001 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 7837.775441] Lustre: Skipped 12 previous similar messages [ 7862.709897] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 7862.712363] LDISKFS-fs (dm-1): recovery complete [ 7862.720196] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 7863.001526] Lustre: lustre-MDT0001: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 7863.063519] Lustre: lustre-MDT0001: in recovery but waiting for the first client to connect [ 7863.087427] Lustre: Skipped 3 previous similar messages [ 7864.513413] Lustre: lustre-MDT0001: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 7864.533344] Lustre: Skipped 3 previous similar messages [ 7868.066251] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7868.426071] Lustre: lustre-MDT0001-lwp-OST0001: Connection restored to 0@lo (at 0@lo) [ 7868.452247] Lustre: Skipped 12 previous similar messages [ 7868.544460] Lustre: lustre-MDT0001: Recovery over after 0:04, of 2 clients 2 recovered and 0 were evicted. [ 7868.550908] Lustre: Skipped 2 previous similar messages [ 7868.604883] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000401:936 to 0x280000401:961) [ 7868.610236] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:936 to 0x2c0000400:961) [ 7882.521664] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 7885.818839] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7896.520425] Lustre: DEBUG MARKER: == replay-single test 70f: OSS O_DIRECT recovery with 1 clients ========================================================== 09:23:57 (1786109037) [ 7909.752822] Lustre: DEBUG MARKER: ost1 REPLAY BARRIER on lustre-OST0000 [ 7912.429373] Lustre: DEBUG MARKER: test_70f failing OST 1 times [ 7914.635976] Lustre: Failing over lustre-OST0000 [ 7914.720283] Lustre: server umount lustre-OST0000 complete [ 7939.306842] LDISKFS-fs (dm-2): 3 truncates cleaned up [ 7939.308850] LDISKFS-fs (dm-2): recovery complete [ 7939.315207] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 7939.512112] Lustre: lustre-OST0000: Imperative Recovery enabled, recovery window shrunk from 60-180 down to 60-180 [ 7947.483873] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 7958.528139] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid 1475 0 [ 7961.475085] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 7973.055414] Lustre: DEBUG MARKER: == replay-single test 71a: mkdir/rmdir striped dir with 2 mdts recovery ========================================================== 09:25:14 (1786109114) [ 8100.906331] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8109.731212] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 8124.425747] Lustre: DEBUG MARKER: fail mds1 mds2 1 times [ 8126.385840] Lustre: Failing over lustre-MDT0000 [ 8126.419181] Lustre: lustre-MDT0000: Not available for connect from 192.168.203.38@tcp (stopping) [ 8126.625617] Lustre: server umount lustre-MDT0000 complete [ 8131.444931] LustreError: 9538:0:(ldlm_lockd.c:2564:ldlm_cancel_handler()) ldlm_cancel from 0@lo arrived at 1786109274 with bad export cookie 15803895791985139146 [ 8131.446435] Lustre: Failing over lustre-MDT0001 [ 8131.447401] LustreError: MGC192.168.203.138@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail [ 8131.455936] LustreError: 9538:0:(ldlm_lockd.c:2564:ldlm_cancel_handler()) Skipped 4 previous similar messages [ 8132.007846] Lustre: server umount lustre-MDT0001 complete [ 8150.496095] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1786109277/real 1786109277] req@ffff930986553b80 x1872862932669952/t0(0) o400->lustre-MDT0001-lwp-OST0001@0@lo:12/10 lens 224/224 e 0 to 1 dl 1786109293 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 8150.546025] Lustre: 3646:0:(client.c:2490:ptlrpc_expire_one_request()) Skipped 1 previous similar message [ 8165.515870] LDISKFS-fs (dm-0): 3 truncates cleaned up [ 8165.520808] LDISKFS-fs (dm-0): recovery complete [ 8165.531702] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8165.879583] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 8165.881853] LDISKFS-fs (dm-1): recovery complete [ 8165.891352] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8176.417278] Lustre: Evicted from MGS (at 0@lo) after server handle changed from 0x0 to 0xdb52b77d7ca3077f [ 8176.715398] Lustre: lustre-MDT0001: Imperative Recovery not enabled, recovery window 60-180 [ 8176.721258] Lustre: Skipped 2 previous similar messages [ 8180.930501] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000401:963 to 0x280000401:993) [ 8180.931666] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:963 to 0x2c0000400:993) [ 8185.542582] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8186.278277] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8194.879340] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6196 to 0x2c0000401:6241) [ 8194.882156] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:6261 to 0x280000400:6305) [ 8202.229339] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid,mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 8204.170233] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8207.267855] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8217.444788] Lustre: DEBUG MARKER: == replay-single test 73a: open(O_CREAT), unlink, replay, reconnect before open replay, close ========================================================== 09:29:18 (1786109358) [ 8226.788832] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8230.770353] Lustre: Failing over lustre-MDT0000 [ 8230.881693] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 8230.892659] Lustre: Skipped 3 previous similar messages [ 8231.206190] Lustre: server umount lustre-MDT0000 complete [ 8254.772451] LDISKFS-fs (dm-0): 3 truncates cleaned up [ 8254.775651] LDISKFS-fs (dm-0): recovery complete [ 8254.783880] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8263.939367] Lustre: *** cfs_fail_loc=302, val=2147483648*** [ 8266.780476] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8279.303788] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnected, waiting for 2 clients in recovery for 0:54 [ 8279.327395] Lustre: 187900:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff93098ba4d180 x1872862920179840/t322122566376(322122566376) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:648/0 lens 520/3488 e 0 to 0 dl 1786109433 ref 1 fl Interpret:/606/0 rc 0/0 job:'lfs.0' uid:0 gid:0 projid:0 [ 8279.412649] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6243 to 0x2c0000401:6273) [ 8279.415210] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:6261 to 0x280000400:6337) [ 8286.745166] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8288.501241] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8298.286866] Lustre: DEBUG MARKER: == replay-single test 73b: open(O_CREAT), unlink, replay, reconnect at open_replay reply, close ========================================================== 09:30:39 (1786109439) [ 8306.699644] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8309.383543] Lustre: Failing over lustre-MDT0000 [ 8309.641564] Lustre: server umount lustre-MDT0000 complete [ 8332.902342] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 8332.905455] LDISKFS-fs (dm-0): recovery complete [ 8332.911695] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8340.959441] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff930a9780dc00 x1872862933221376/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 8342.462779] Lustre: *** cfs_fail_loc=157, val=2147483648*** [ 8342.470579] LustreError: 187899:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930a9780ce00 x1872862920179840/t322122566376(322122566376) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:711/0 lens 520/664 e 0 to 0 dl 1786109496 ref 1 fl Interpret:/604/0 rc 301/0 job:'lfs.0' uid:0 gid:0 projid:0 [ 8346.272818] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8358.665953] Lustre: lustre-MDT0000: Client 397d331e-10fc-4558-be87-308c3622e8eb (at 192.168.203.38@tcp) reconnected, waiting for 2 clients in recovery for 0:53 [ 8358.682143] Lustre: 189813:0:(mdt_recovery.c:102:mdt_req_from_lrd()) @@@ restoring transno req@ffff930a9780bb80 x1872862920179840/t322122566376(322122566376) o101->397d331e-10fc-4558-be87-308c3622e8eb@192.168.203.38@tcp:728/0 lens 520/3488 e 0 to 0 dl 1786109513 ref 1 fl Interpret:/606/0 rc 0/0 job:'lfs.0' uid:0 gid:0 projid:0 [ 8358.738050] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:6261 to 0x280000400:6369) [ 8358.739318] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6275 to 0x2c0000401:6305) [ 8365.531643] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8367.496607] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8378.237391] Lustre: DEBUG MARKER: == replay-single test 74: Ensure applications don't fail waiting for OST recovery ========================================================== 09:31:59 (1786109519) [ 8381.897391] Lustre: Failing over lustre-OST0000 [ 8382.029368] Lustre: server umount lustre-OST0000 complete [ 8386.157625] Lustre: Failing over lustre-MDT0000 [ 8386.379460] Lustre: server umount lustre-MDT0000 complete [ 8405.317638] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8417.025338] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8418.826680] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6275 to 0x2c0000401:6337) [ 8424.684695] LDISKFS-fs (dm-2): mounted filesystem with ordered data mode. Opts: errors=remount-ro,no_mbcache,nodelalloc [ 8426.836135] Lustre: lustre-OST0000: Denying connection for new client fac74500-4523-4794-af20-dfdcdfae19c1 (at 192.168.203.38@tcp), waiting for 2 known clients (1 recovered, 0 in progress, and 0 evicted) to recover in 1:09 [ 8430.144486] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:6261 to 0x280000400:6401) [ 8430.783978] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8444.112207] Lustre: DEBUG MARKER: == replay-single test 80a: DNE: create remote dir, drop update rep from MDT0, fail MDT0 ========================================================== 09:33:05 (1786109585) [ 8445.007632] Lustre: *** cfs_fail_loc=1701, val=2147483648*** [ 8445.015205] LustreError: 8451:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff93098e358700 x1872862933289216/t339302416391(0) o1000->lustre-MDT0001-mdtlov_UUID@0@lo:59/0 lens 1312/4320 e 0 to 0 dl 1786109599 ref 1 fl Interpret:/200/0 rc 0/0 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8450.701228] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8452.703923] Lustre: Failing over lustre-MDT0000 [ 8453.007764] Lustre: server umount lustre-MDT0000 complete [ 8454.625222] LustreError: lustre-MDT0000-osp-MDT0001: operation mds_statfs to node 0@lo failed: rc = -107 [ 8454.635381] LustreError: Skipped 5 previous similar messages [ 8454.642162] Lustre: lustre-MDT0000-osp-MDT0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete [ 8454.665592] Lustre: Skipped 23 previous similar messages [ 8454.678590] LustreError: 177399:0:(ldlm_lib.c:1192:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. [ 8454.707910] LustreError: 177399:0:(ldlm_lib.c:1192:target_handle_connect()) Skipped 146 previous similar messages [ 8473.828347] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 8473.830576] LDISKFS-fs (dm-0): recovery complete [ 8473.836585] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8483.298768] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) @@@ invalidate in flight req@ffff930a8909b100 x1872862933303808/t0(0) o250->MGC192.168.203.138@tcp@0@lo:26/25 lens 520/544 e 0 to 0 dl 0 ref 1 fl Rpc:NQU/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 8483.337349] LustreError: 3645:0:(client.c:1391:ptlrpc_import_delay_req()) Skipped 1 previous similar message [ 8483.746857] Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect [ 8483.763243] Lustre: Skipped 7 previous similar messages [ 8483.995879] Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect [ 8484.011634] Lustre: Skipped 7 previous similar messages [ 8487.906852] Lustre: mdt00_000: service thread pid 187899 was inactive for 42.945 seconds. The thread might be hung, or it might only be slow and will resume later. Dumping the stack trace for debugging purposes: [ 8487.919799] task:mdt00_000 state:I stack:0 pid:187899 ppid:2 flags:0x80004000 [ 8487.923583] Call Trace: [ 8487.928168] __schedule+0x351/0xcb0 [ 8487.935433] schedule+0xc0/0x180 [ 8487.939564] top_trans_stop+0xbc5/0x1590 [ptlrpc] [ 8487.949319] ? woken_wake_function+0x30/0x30 [ 8487.959419] lod_trans_stop+0xe2/0x560 [lod] [ 8487.963510] mdd_trans_stop+0x29/0x1f0 [mdd] [ 8487.968436] mdd_create+0x1f9c/0x2590 [mdd] [ 8487.970566] ? mdd_links_rename+0x550/0x550 [mdd] [ 8487.980170] mdt_create+0xcc6/0x2110 [mdt] [ 8487.985996] mdt_reint_create+0x336/0x5d0 [mdt] [ 8487.995185] mdt_reint_rec+0x139/0x2b0 [mdt] [ 8488.000461] mdt_reint_internal+0x693/0xdc0 [mdt] [ 8488.002487] mdt_reint+0x163/0x190 [mdt] [ 8488.011418] tgt_handle_request0+0x137/0xaf0 [ptlrpc] [ 8488.014558] tgt_request_handle+0x575/0x1f70 [ptlrpc] [ 8488.018872] ptlrpc_server_handle_request+0x443/0x13b0 [ptlrpc] [ 8488.021813] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8488.024797] ptlrpc_main+0xce8/0x1400 [ptlrpc] [ 8488.031248] ? ptlrpc_wait_event+0x690/0x690 [ptlrpc] [ 8488.034361] kthread+0x1d1/0x200 [ 8488.035929] ? set_kthread_struct+0x70/0x70 [ 8488.038214] ret_from_fork+0x1f/0x30 [ 8488.910982] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8488.932504] Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 0@lo (at 0@lo) [ 8488.949458] Lustre: Skipped 24 previous similar messages [ 8488.984463] Lustre: lustre-MDT0000: Recovery over after 0:05, of 2 clients 2 recovered and 0 were evicted. [ 8488.995749] Lustre: Skipped 7 previous similar messages [ 8489.013408] Lustre: mdt00_000: service thread pid 187899 completed after 44.051s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 8489.041724] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:6403 to 0x280000400:6433) [ 8489.042935] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6275 to 0x2c0000401:6369) [ 8498.627439] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8500.210540] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8509.792504] Lustre: DEBUG MARKER: == replay-single test 80b: DNE: create remote dir, drop update rep from MDT0, fail MDT1 ========================================================== 09:34:11 (1786109651) [ 8510.999799] Lustre: *** cfs_fail_loc=1701, val=2147483648*** [ 8511.026393] LustreError: 179257:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930ac17bfb80 x1872862933333376/t343597383692(0) o1000->lustre-MDT0001-mdtlov_UUID@0@lo:125/0 lens 488/4320 e 0 to 0 dl 1786109665 ref 1 fl Interpret:/200/0 rc 0/0 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8519.176464] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8527.335795] Lustre: lustre-MDT0000: Received new MDS connection from 0@lo, keep former export from same NID [ 8527.827684] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 8529.686094] Lustre: Failing over lustre-MDT0001 [ 8529.888228] Lustre: lustre-MDT0001: Not available for connect from 0@lo (stopping) [ 8529.970288] Lustre: server umount lustre-MDT0001 complete [ 8551.980129] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 8551.982779] LDISKFS-fs (dm-1): recovery complete [ 8551.991929] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8557.117887] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8557.625397] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000401:1004 to 0x280000401:1025) [ 8557.625837] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:1004 to 0x2c0000400:1025) [ 8567.518843] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 8569.818702] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8580.105935] Lustre: DEBUG MARKER: == replay-single test 80c: DNE: create remote dir, drop update rep from MDT1, fail MDT[0,1] ========================================================== 09:35:21 (1786109721) [ 8581.736448] LustreError: 199402:0:(ldlm_lib.c:3346:target_send_reply_msg()) @@@ dropping reply req@ffff930a97fd7800 x1872862933377408/t343597383713(0) o1000->lustre-MDT0001-mdtlov_UUID@0@lo:196/0 lens 2520/4320 e 0 to 0 dl 1786109736 ref 1 fl Interpret:/200/0 rc 0/0 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8588.302735] Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 [ 8596.217626] Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 [ 8597.892168] Lustre: Failing over lustre-MDT0000 [ 8597.988288] Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) [ 8598.012786] Lustre: Skipped 3 previous similar messages [ 8598.196423] Lustre: server umount lustre-MDT0000 complete [ 8621.814650] LDISKFS-fs (dm-0): 2 truncates cleaned up [ 8621.816873] LDISKFS-fs (dm-0): recovery complete [ 8621.824333] LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8623.071141] Lustre: mdt00_005: service thread pid 196401 was inactive for 41.360 seconds. The thread might be hung, or it might only be slow and will resume later. Dumping the stack trace for debugging purposes: [ 8623.097553] task:mdt00_005 state:I stack:0 pid:196401 ppid:2 flags:0x80004000 [ 8623.108233] Call Trace: [ 8623.110785] __schedule+0x351/0xcb0 [ 8623.114739] schedule+0xc0/0x180 [ 8623.118863] top_trans_stop+0xbc5/0x1590 [ptlrpc] [ 8623.123673] ? woken_wake_function+0x30/0x30 [ 8623.127978] lod_trans_stop+0xe2/0x560 [lod] [ 8623.133568] mdd_trans_stop+0x29/0x1f0 [mdd] [ 8623.140292] mdd_create+0x1f9c/0x2590 [mdd] [ 8623.150723] ? mdd_links_rename+0x550/0x550 [mdd] [ 8623.160664] mdt_create+0xcc6/0x2110 [mdt] [ 8623.164501] mdt_reint_create+0x336/0x5d0 [mdt] [ 8623.169911] mdt_reint_rec+0x139/0x2b0 [mdt] [ 8623.172687] mdt_reint_internal+0x693/0xdc0 [mdt] [ 8623.175871] mdt_reint+0x163/0x190 [mdt] [ 8623.180554] tgt_handle_request0+0x137/0xaf0 [ptlrpc] [ 8623.187531] tgt_request_handle+0x575/0x1f70 [ptlrpc] [ 8623.198532] ptlrpc_server_handle_request+0x443/0x13b0 [ptlrpc] [ 8623.206423] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8623.210768] ptlrpc_main+0xce8/0x1400 [ptlrpc] [ 8623.216963] ? ptlrpc_wait_event+0x690/0x690 [ptlrpc] [ 8623.221543] kthread+0x1d1/0x200 [ 8623.224960] ? set_kthread_struct+0x70/0x70 [ 8623.229291] ret_from_fork+0x1f/0x30 [ 8633.325083] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8634.456185] Lustre: 201173:0:(out_handler.c:886:out_tx_end()) lustre-MDT0000-osd: error during execution of #0 from /home/green/git/lustre-release/lustre/ptlrpc/../target/out_handler.c:565: rc = -2 [ 8634.471612] LustreError: 201173:0:(tgt_lastrcvd.c:1516:tgt_last_rcvd_update()) lustre-MDT0000: replay transno 343597383708 failed: rc = -2 [ 8634.490305] LustreError: 3645:0:(client.c:3435:ptlrpc_replay_interpret()) @@@ status -2, old was 0 req@ffff930ab1c20000 x1872862933370496/t343597383708(343597383708) o1000->lustre-MDT0000-osp-MDT0001@0@lo:24/4 lens 264/4320 e 0 to 0 dl 1786109793 ref 2 fl Interpret:RQU/204/0 rc -2/-2 job:'osp_up0-1.0' uid:0 gid:0 projid:4294967295 [ 8634.535816] Lustre: mdt00_005: service thread pid 196401 completed after 52.824s. This likely indicates the system was overloaded (too many service threads, or not enough hardware resources). [ 8634.623524] LustreError: 3645:0:(ldlm_request.c:2758:replay_lock_interpret()) received replay ack for unknown local cookie 0xdb52b77d7ca3888f remote cookie 0xdb52b77d7ca38ce8 from server (efault) id 0-0@<0:0> [ 8634.635973] LustreError: 3645:0:(import.c:717:ptlrpc_connect_import_locked()) already connecting [ 8634.647597] Lustre: lustre-MDT0000: Received new MDS connection from 0@lo, keep former export from same NID [ 8634.705918] Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:6403 to 0x280000400:6465) [ 8634.708070] Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000401:6275 to 0x2c0000401:6401) [ 8642.339310] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 [ 8644.435346] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8647.844118] Lustre: Failing over lustre-MDT0001 [ 8648.013636] Lustre: server umount lustre-MDT0001 complete [ 8668.691651] LDISKFS-fs (dm-1): 6 truncates cleaned up [ 8668.694193] LDISKFS-fs (dm-1): recovery complete [ 8668.704650] LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc [ 8673.226757] Lustre: DEBUG MARKER: oleg338-server.virtnet: executing set_default_debug vfstrace rpctrace dlmtrace neterror ha config ioctl super lfsck all [ 8674.335612] Lustre: 202512:0:(update_recovery.c:1391:distribute_txn_replay_handle()) lustre-MDT0001-mdtlov: error during execution of #0 from /home/green/git/lustre-release/lustre/ptlrpc/../target/update_recovery.c:946: rc = -116 [ 8674.373512] LustreError: 202512:0:(update_trans.c:1080:top_trans_stop()) lustre-MDT0000-osp-MDT0001: stop trans failed: rc = -20 [ 8674.492920] Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x280000401:1036 to 0x280000401:1057) [ 8674.494863] Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000400:1036 to 0x2c0000400:1057) [ 8681.804546] Lustre: DEBUG MARKER: oleg338-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 [ 8683.078160] Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8686.576487] Lustre: DEBUG MARKER: replay-single test_80c: @@@@@@ FAIL: remote creation failed