[ 2155.851843] LustreError: lustre-OST0000-osc-ffff96fd057df000: operation ost_sync to node 192.168.201.103@tcp failed: rc = -107 [ 2155.922157] LustreError: lustre-OST0000-osc-ffff96fd057df000: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 2155.994965] Lustre: 2355:0:(llite_lib.c:4198:ll_dirty_page_discard_warn()) lustre: dirty page discard: 192.168.201.103@tcp:/lustre/fid: [0x200000404:0x35:0x0]// may get corrupted (rc -108) [ 2173.815675] Lustre: DEBUG MARKER: == recovery-small test 26a: evict dead exports =========== 17:51:07 (1781819467) [ 2179.261801] Lustre: DEBUG MARKER: SKIP: recovery-small test_26a msg and ost1 are at the same node [ 2184.408976] Lustre: DEBUG MARKER: == recovery-small test 26b: evict dead exports =========== 17:51:17 (1781819477) [ 2188.991963] Lustre: DEBUG MARKER: SKIP: recovery-small test_26b msg and ost1 are at the same node [ 2193.587116] Lustre: DEBUG MARKER: == recovery-small test 27: fail LOV while using OSC's ==== 17:51:27 (1781819487) [ 2200.743431] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -19 [ 2200.791667] Lustre: lustre-MDT0000-mdc-ffff96fd057df000: Connection to lustre-MDT0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 2200.805045] LustreError: Skipped 2 previous similar messages [ 2200.870794] Lustre: Skipped 10 previous similar messages [ 2219.487622] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 2244.080314] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f949c2e6 to 0xaa643b29f949d9b4 [ 2244.156107] Lustre: MGC192.168.201.103@tcp: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 2244.210794] Lustre: Skipped 9 previous similar messages [ 2362.458163] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -19 [ 2382.303947] Lustre: 2357:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781819663/real 1781819663] req@ffff96fd18a09180 x1868370982087296/t0(0) o400->MGC192.168.201.103@tcp@192.168.201.103@tcp:26/25 lens 224/224 e 0 to 1 dl 1781819679 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 2382.386551] Lustre: 2357:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 10 previous similar messages [ 2382.409651] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 2391.634314] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f949d9b4 to 0xaa643b29f94abd57 [ 2419.276815] Lustre: DEBUG MARKER: == recovery-small test 28: handle error adding new clients (bug 6086) ========================================================== 17:55:14 (1781819714) [ 2420.304588] Lustre: *** cfs_fail_loc=305, val=0*** [ 2420.309573] Lustre: Skipped 2 previous similar messages [ 2441.510555] LustreError: lustre-OST0000-osc-ffff96fd057df000: operation ost_connect to node 192.168.201.103@tcp failed: rc = -75 [ 2468.326208] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 2478.579308] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94abd57 to 0xaa643b29f94ad480 [ 2479.763258] Lustre: 15888:0:(mgc_request.c:1901:mgc_process_log()) MGC192.168.201.103@tcp: IR log lustre-cliir failed, not fatal: rc = -2 [ 2504.299667] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 2507.057591] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 2522.861705] Lustre: DEBUG MARKER: == recovery-small test 29a: error adding new clients doesn't cause LBUG (bug 22273) ========================================================== 17:56:57 (1781819817) [ 2545.148259] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 2545.187178] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94ad480 to 0xaa643b29f94ada99 [ 2545.307564] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 2582.219032] Lustre: DEBUG MARKER: == recovery-small test 29b: error adding new clients doesn't cause LBUG (bug 22273) ========================================================== 17:57:57 (1781819877) [ 2629.444415] Lustre: DEBUG MARKER: == recovery-small test 50: failover MDS under load ======= 17:58:44 (1781819924) [ 2644.225339] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -19 [ 2662.881752] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 2673.134855] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94ada99 to 0xaa643b29f94b154e [ 2696.078386] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 2699.716104] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 2768.612323] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: operation mds_reint to node 192.168.201.103@tcp failed: rc = -19 [ 2784.735903] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 2795.003989] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94b154e to 0xaa643b29f94c1d2c [ 2796.289123] Lustre: 15888:0:(mgc_request.c:1901:mgc_process_log()) MGC192.168.201.103@tcp: IR log lustre-cliir failed, not fatal: rc = -2 [ 2824.877720] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 2828.268469] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 2895.813303] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: operation mds_reint to node 192.168.201.103@tcp failed: rc = -19 [ 2895.815433] Lustre: lustre-MDT0000-mdc-ffff96fd057df000: Connection to lustre-MDT0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 2895.845958] LustreError: Skipped 1 previous similar message [ 2895.874021] Lustre: Skipped 5 previous similar messages [ 2916.895359] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 2927.236044] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94c1d2c to 0xaa643b29f94d5c4d [ 2927.285651] Lustre: MGC192.168.201.103@tcp: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 2927.302766] Lustre: Skipped 11 previous similar messages [ 2953.479074] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 2955.397762] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 2987.832719] Lustre: DEBUG MARKER: == recovery-small test 51: failover MDS during recovery == 18:04:43 (1781820283) [ 3012.639250] Lustre: 2354:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781820294/real 1781820294] req@ffff96fd086ec700 x1868370984861184/t0(0) o400->MGC192.168.201.103@tcp@192.168.201.103@tcp:26/25 lens 224/224 e 0 to 1 dl 1781820310 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 3012.707067] Lustre: 2354:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 12 previous similar messages [ 3012.734305] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 3023.365511] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94d5c4d to 0xaa643b29f94df7f8 [ 3024.145512] Lustre: DEBUG MARKER: test_51: failover in 1 sec [ 3074.304095] Lustre: DEBUG MARKER: test_51: failover in 5 sec [ 3081.909930] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -19 [ 3081.924194] LustreError: Skipped 2 previous similar messages [ 3124.828773] Lustre: DEBUG MARKER: test_51: failover in 10 sec [ 3160.550226] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 3160.586275] LustreError: Skipped 2 previous similar messages [ 3160.609606] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94e0c26 to 0xaa643b29f94e5583 [ 3160.623250] Lustre: Skipped 2 previous similar messages [ 3162.293341] Lustre: 15888:0:(mgc_request.c:1901:mgc_process_log()) MGC192.168.201.103@tcp: IR log lustre-cliir failed, not fatal: rc = -2 [ 3185.611262] Lustre: DEBUG MARKER: test_51: failover in 20 sec [ 3254.859970] Lustre: DEBUG MARKER: test_51: failover in 25 sec [ 3313.154291] Lustre: 15888:0:(mgc_request.c:1901:mgc_process_log()) MGC192.168.201.103@tcp: IR log lustre-cliir failed, not fatal: rc = -2 [ 3325.067285] Lustre: DEBUG MARKER: test_51: failover in 30 sec [ 3362.042487] LustreError: lustre-MDT0000-mdc-ffff96fd057df000: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -19 [ 3362.081526] LustreError: Skipped 5 previous similar messages [ 3431.408234] Lustre: DEBUG MARKER: == recovery-small test 52: failover OST under load ======= 18:12:05 (1781820725) [ 3503.764162] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 3506.611903] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 3780.770065] Lustre: lustre-OST0000-osc-ffff96fd057df000: Connection to lustre-OST0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 3780.850179] Lustre: Skipped 7 previous similar messages [ 3811.628050] Lustre: lustre-OST0000-osc-ffff96fd057df000: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 3811.687424] Lustre: Skipped 14 previous similar messages [ 3838.526858] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 3841.491747] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 4111.746396] LustreError: lustre-OST0000-osc-ffff96fd057df000: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -107 [ 4111.777722] LustreError: Skipped 2 previous similar messages [ 4170.159379] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 4173.688094] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 4408.813062] Lustre: DEBUG MARKER: == recovery-small test 53a: touch: drop rep ============== 18:28:24 (1781821704) [ 4427.231171] Lustre: 50333:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781821708/real 1781821708] req@ffff96fd09bb9f80 x1868370996504320/t0(0) o101->lustre-MDT0000-mdc-ffff96fd057df000@192.168.201.103@tcp:12/10 lens 576/1152 e 0 to 1 dl 1781821724 ref 2 fl Rpc:XQr/600/ffffffff rc 0/-1 job:'openfile.0' uid:0 gid:0 projid:0 [ 4427.269595] Lustre: 50333:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 6 previous similar messages [ 4427.274146] Lustre: lustre-MDT0000-mdc-ffff96fd057df000: Connection to lustre-MDT0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 4427.287595] Lustre: Skipped 1 previous similar message [ 4427.340893] Lustre: lustre-MDT0000-mdc-ffff96fd057df000: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 4427.358096] Lustre: Skipped 1 previous similar message [ 4441.296449] Lustre: DEBUG MARKER: == recovery-small test 53b: touch: drop rep ============== 18:28:56 (1781821736) [ 4472.535739] Lustre: DEBUG MARKER: == recovery-small test 53c: touch: drop rep ============== 18:29:27 (1781821767) [ 4502.008959] Lustre: DEBUG MARKER: == recovery-small test 54: back in time ================== 18:29:57 (1781821797) [ 4502.601488] Lustre: Mounted lustre-client [ 4534.239428] Lustre: 2355:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781821815/real 1781821815] req@ffff96fd0a229880 x1868370996525696/t0(0) o400->MGC192.168.201.103@tcp@192.168.201.103@tcp:26/25 lens 224/224 e 0 to 1 dl 1781821831 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 4534.276706] Lustre: 2355:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 4534.288431] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 4534.310253] LustreError: Skipped 3 previous similar messages [ 4544.522373] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f94fa41e to 0xaa643b29f95b917c [ 4544.533639] Lustre: Skipped 3 previous similar messages [ 4563.634780] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 4566.017803] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 4568.990920] Lustre: Unmounted lustre-client [ 4579.564201] Lustre: DEBUG MARKER: == recovery-small test 55: ost_brw_read/write drops timed-out read/write request ========================================================== 18:31:14 (1781821874) [ 4691.935253] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781821973/real 1781821973] req@ffff96fd08b90700 x1868370996562176/t0(0) o4->lustre-OST0000-osc-ffff96fd057df000@192.168.201.103@tcp:6/4 lens 488/448 e 0 to 1 dl 1781821989 ref 2 fl Rpc:XQr/602/ffffffff rc 0/-1 job:'dd.0' uid:0 gid:0 projid:0 [ 4691.991115] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 57 previous similar messages [ 4888.989591] Lustre: DEBUG MARKER: == recovery-small test 56: do not fail on getattr resend ========================================================== 18:36:24 (1781822184) [ 4942.541968] Lustre: DEBUG MARKER: == recovery-small test 57: read procfs entries causes kernel crash ========================================================== 18:37:17 (1781822237) [ 4947.682323] Lustre: Unmounted lustre-client [ 4978.049593] Lustre: Mounted lustre-client [ 4987.114400] Lustre: DEBUG MARKER: == recovery-small test 58: Eviction in the middle of open RPC reply processing ========================================================== 18:38:02 (1781822282) [ 4987.729180] LustreError: 56257:0:(mdc_locks.c:1335:mdc_finish_intent_lock()) cfs_fail_timeout id 801 sleeping for 20000ms [ 4988.801337] LustreError: 56257:0:(mdc_locks.c:1335:mdc_finish_intent_lock()) cfs_fail_timeout interrupted [ 4988.975754] Lustre: *** cfs_fail_loc=305, val=0*** [ 5013.712142] Lustre: DEBUG MARKER: == recovery-small test 59: Read cancel race on client eviction ========================================================== 18:38:29 (1781822309) [ 5014.172616] Lustre: Mounted lustre-client [ 5018.325187] Lustre: setting import lustre-MDT0000_UUID INACTIVE by administrator request [ 5028.665314] Lustre: Unmounted lustre-client [ 5041.318843] Lustre: DEBUG MARKER: == recovery-small test 60: Add Changelog entries during MDS failover ========================================================== 18:38:56 (1781822336) [ 5202.967372] LustreError: lustre-MDT0000-mdc-ffff96fd09ad9800: operation mds_reint to node 192.168.201.103@tcp failed: rc = -19 [ 5202.987550] Lustre: lustre-MDT0000-mdc-ffff96fd09ad9800: Connection to lustre-MDT0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 5203.016307] Lustre: Skipped 23 previous similar messages [ 5220.327563] Lustre: 2355:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781822501/real 1781822501] req@ffff96fd04682300 x1868370999233664/t0(0) o400->MGC192.168.201.103@tcp@192.168.201.103@tcp:26/25 lens 224/224 e 0 to 1 dl 1781822517 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 5220.390047] Lustre: 2355:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 93 previous similar messages [ 5220.412351] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 5229.632385] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f95bac79 to 0xaa643b29f95e5b28 [ 5229.660450] Lustre: MGC192.168.201.103@tcp: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 5229.695329] Lustre: Skipped 24 previous similar messages [ 5512.050290] Lustre: DEBUG MARKER: == recovery-small test 61: Verify to not reuse orphan objects - bug 17025 ========================================================== 18:46:47 (1781822807) [ 5519.248938] Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 [ 5541.365595] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 5541.397393] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f95e5b28 to 0xaa643b29f9633076 [ 5541.461450] LustreError: lustre-MDT0000-mdc-ffff96fd09ad9800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 5568.056161] Lustre: DEBUG MARKER: == recovery-small test 65: lock enqueue for destroyed export ========================================================== 18:47:42 (1781822862) [ 5569.116616] Lustre: Mounted lustre-client [ 5577.778616] LustreError: lustre-OST0000-osc-ffff96fd09ad9800: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 5591.104407] Lustre: Unmounted lustre-client [ 5601.095491] Lustre: DEBUG MARKER: == recovery-small test 66: lock enqueue re-send vs client eviction ========================================================== 18:48:16 (1781822896) [ 5609.078570] LustreError: lustre-MDT0000-mdc-ffff96fd09ad9800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 5609.118691] LustreError: 60689:0:(file.c:6088:ll_inode_revalidate_fini()) lustre: revalidate FID [0x200000007:0x1:0x0] error: rc = -5 [ 5622.508402] Lustre: DEBUG MARKER: == recovery-small test 67: connect vs import invalidate race ========================================================== 18:48:37 (1781822917) [ 5622.740955] LustreError: 61373:0:(recover.c:329:ptlrpc_recover_import()) cfs_race id 531 sleeping [ 5627.384150] LustreError: lustre-MDT0000-mdc-ffff96fd09ad9800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 5627.422463] LustreError: 61397:0:(import.c:293:ptlrpc_invalidate_import()) cfs_fail_race id 531 waking [ 5627.448427] LustreError: 61373:0:(recover.c:329:ptlrpc_recover_import()) cfs_fail_race id 531 awake: rc=307 [ 5627.488801] LustreError: 61373:0:(import.c:719:ptlrpc_connect_import_locked()) already connecting [ 5627.541581] LustreError: 61402:0:(file.c:6088:ll_inode_revalidate_fini()) lustre: revalidate FID [0x200000007:0x1:0x0] error: rc = -108 [ 5628.833846] LustreError: 61408:0:(file.c:6088:ll_inode_revalidate_fini()) lustre: revalidate FID [0x200000007:0x1:0x0] error: rc = -108 [ 5628.848700] LustreError: 61408:0:(file.c:6088:ll_inode_revalidate_fini()) Skipped 5 previous similar messages [ 5631.300531] LustreError: 61430:0:(file.c:6088:ll_inode_revalidate_fini()) lustre: revalidate FID [0x200000007:0x1:0x0] error: rc = -108 [ 5631.323101] LustreError: 61430:0:(file.c:6088:ll_inode_revalidate_fini()) Skipped 13 previous similar messages [ 5636.273967] LustreError: 61475:0:(file.c:6088:ll_inode_revalidate_fini()) lustre: revalidate FID [0x200000007:0x1:0x0] error: rc = -108 [ 5636.290307] LustreError: 61475:0:(file.c:6088:ll_inode_revalidate_fini()) Skipped 27 previous similar messages [ 5651.819796] Lustre: DEBUG MARKER: == recovery-small test 100: IR: Make sure normal recovery still works w/o IR ========================================================== 18:49:06 (1781822946) [ 5709.007542] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 5711.691500] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 5726.425712] Lustre: DEBUG MARKER: == recovery-small test 101a: IR: Make sure IR works w/o normal recovery ========================================================== 18:50:21 (1781823021) [ 5782.022913] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 5785.588845] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 5808.797207] Lustre: DEBUG MARKER: == recovery-small test 101b: IR: Make sure IR works w/o normal recovery and proceed EAGAIN ========================================================== 18:51:43 (1781823103) [ 5908.578193] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 5912.496121] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 5928.226894] Lustre: DEBUG MARKER: == recovery-small test 102: IR: New client gets updated nidtbl after MGS restart ========================================================== 18:53:41 (1781823221) [ 5999.883789] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 6002.388663] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 6012.871936] Lustre: Unmounted lustre-client [ 6114.633422] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 6119.381805] Lustre: Mounted lustre-client [ 6135.997808] Lustre: DEBUG MARKER: == recovery-small test 103: IR: MDS can start w/o MGS and get updated nidtbl later ========================================================== 18:57:09 (1781823429) [ 6141.240974] Lustre: DEBUG MARKER: SKIP: recovery-small test_103 needs separate mgs and mds [ 6145.590509] Lustre: DEBUG MARKER: == recovery-small test 104: IR: ost can disable IR voluntarily ========================================================== 18:57:19 (1781823439) [ 6218.947659] Lustre: DEBUG MARKER: == recovery-small test 105: IR: NON IR clients support === 18:58:33 (1781823513) [ 6221.629603] Lustre: DEBUG MARKER: SKIP: recovery-small test_105 Needs multiple clients [ 6226.326875] Lustre: DEBUG MARKER: == recovery-small test 106: lightweight connection support ========================================================== 18:58:39 (1781823519) [ 6229.130947] Lustre: *** cfs_fail_loc=805, val=0*** [ 6229.257297] Lustre: Mounted lustre-client [ 6242.920547] Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 [ 6266.336693] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781823547/real 1781823547] req@ffff96fd043b9f80 x1868371001320960/t0(0) o400->MGC192.168.201.103@tcp@192.168.201.103@tcp:26/25 lens 224/224 e 0 to 1 dl 1781823563 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 6266.394526] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 2 previous similar messages [ 6266.434811] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 6266.489919] Lustre: lustre-MDT0000-mdc-ffff96fd18907800: Connection to lustre-MDT0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 6266.576298] Lustre: Skipped 8 previous similar messages [ 6291.246118] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9633a7f to 0xaa643b29f9633ca8 [ 6291.314328] Lustre: MGC192.168.201.103@tcp: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 6291.345363] Lustre: Skipped 9 previous similar messages [ 6293.198806] Lustre: 68006:0:(mgc_request.c:1901:mgc_process_log()) MGC192.168.201.103@tcp: IR log lustre-cliir failed, not fatal: rc = -2 [ 6320.846184] Lustre: Unmounted lustre-client [ 6339.647765] Lustre: DEBUG MARKER: == recovery-small test 107: drop reint reply, then restart MDT ========================================================== 19:00:33 (1781823633) [ 6369.251557] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 6393.851398] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9633ca8 to 0xaa643b29f9634307 [ 6404.243593] LustreError: 2353:0:(client.c:3447:ptlrpc_replay_interpret()) @@@ status 301, old was 0 req@ffff96fd0a0e8380 x1868371001334912/t90194313220(90194313220) o101->lustre-MDT0000-mdc-ffff96fd09ab3800@192.168.201.103@tcp:12/10 lens 664/608 e 0 to 0 dl 1781823717 ref 2 fl Interpret:RPQU/604/0 rc 301/301 job:'multiop.0' uid:0 gid:0 projid:0 [ 6429.447722] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 6432.761571] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 6448.480429] Lustre: DEBUG MARKER: == recovery-small test 108: client eviction don't crash == 19:02:23 (1781823743) [ 6450.807907] LustreError: lustre-OST0000-osc-ffff96fd09ab3800: operation ost_write to node 192.168.201.103@tcp failed: rc = -107 [ 6450.833158] LustreError: Skipped 3 previous similar messages [ 6450.868273] LustreError: lustre-OST0000-osc-ffff96fd09ab3800: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 6450.937481] Lustre: 2354:0:(llite_lib.c:4198:ll_dirty_page_discard_warn()) lustre: dirty page discard: 192.168.201.103@tcp:/lustre/fid: [0x20000a042:0x6:0x0]// may get corrupted (rc -5) [ 6451.062193] LustreError: 72557:0:(ldlm_resource.c:1180:ldlm_resource_complain()) lustre-OST0000-osc-ffff96fd09ab3800: namespace resource [0x240000400:0x2b02:0x0].0x0 (ffff96fd05c9ae00) refcount nonzero (1) after lock cleanup; forcing cleanup. [ 6469.953746] Lustre: DEBUG MARKER: == recovery-small test 110a: create remote directory: drop client req ========================================================== 19:02:43 (1781823763) [ 6473.743434] Lustre: DEBUG MARKER: SKIP: recovery-small test_110a needs >= 2 MDTs [ 6478.180049] Lustre: DEBUG MARKER: == recovery-small test 110b: create remote directory: drop Master rep ========================================================== 19:02:52 (1781823772) [ 6481.880344] Lustre: DEBUG MARKER: SKIP: recovery-small test_110b needs >= 2 MDTs [ 6486.342805] Lustre: DEBUG MARKER: == recovery-small test 110c: create remote directory: drop update rep on slave MDT ========================================================== 19:03:00 (1781823780) [ 6490.219733] Lustre: DEBUG MARKER: SKIP: recovery-small test_110c needs >= 2 MDTs [ 6494.272791] Lustre: DEBUG MARKER: == recovery-small test 110d: remove remote directory: drop client req ========================================================== 19:03:08 (1781823788) [ 6497.876246] Lustre: DEBUG MARKER: SKIP: recovery-small test_110d needs >= 2 MDTs [ 6501.726231] Lustre: DEBUG MARKER: == recovery-small test 110e: remove remote directory: drop master rep ========================================================== 19:03:16 (1781823796) [ 6505.175771] Lustre: DEBUG MARKER: SKIP: recovery-small test_110e needs >= 2 MDTs [ 6509.000987] Lustre: DEBUG MARKER: == recovery-small test 110f: remove remote directory: drop slave rep ========================================================== 19:03:23 (1781823803) [ 6512.969738] Lustre: DEBUG MARKER: SKIP: recovery-small test_110f needs >= 2 MDTs [ 6518.204371] Lustre: DEBUG MARKER: == recovery-small test 110g: drop reply during migration ========================================================== 19:03:31 (1781823811) [ 6522.298259] Lustre: DEBUG MARKER: SKIP: recovery-small test_110g needs >= 2 MDTs [ 6526.461300] Lustre: DEBUG MARKER: == recovery-small test 110h: drop update reply during cross-MDT file rename ========================================================== 19:03:40 (1781823820) [ 6529.651401] Lustre: DEBUG MARKER: SKIP: recovery-small test_110h needs >= 2 MDTs [ 6533.528864] Lustre: DEBUG MARKER: == recovery-small test 110i: drop update reply during cross-MDT dir rename ========================================================== 19:03:47 (1781823827) [ 6537.010388] Lustre: DEBUG MARKER: SKIP: recovery-small test_110i needs >= 2 MDTs [ 6540.727579] Lustre: DEBUG MARKER: == recovery-small test 110j: drop update reply during cross-MDT ln ========================================================== 19:03:55 (1781823835) [ 6544.035126] Lustre: DEBUG MARKER: SKIP: recovery-small test_110j needs >= 2 MDTs [ 6548.333175] Lustre: DEBUG MARKER: == recovery-small test 110k: FID_QUERY failed during recovery ========================================================== 19:04:02 (1781823842) [ 6552.625906] Lustre: DEBUG MARKER: SKIP: recovery-small test_110k needs >= 2 MDTS [ 6556.430213] Lustre: DEBUG MARKER: == recovery-small test 110m: update resent vs original RPC race ========================================================== 19:04:10 (1781823850) [ 6562.767285] Lustre: DEBUG MARKER: SKIP: recovery-small test_110m needs at least 2 MDTs [ 6566.855613] Lustre: DEBUG MARKER: == recovery-small test 111: mdd setup fail should not cause umount oops ========================================================== 19:04:21 (1781823861) [ 6592.483195] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 6617.008197] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9634307 to 0xaa643b29f9634990 [ 6644.055270] Lustre: DEBUG MARKER: == recovery-small test 112a: bulk resend while orignal request is in progress ========================================================== 19:05:37 (1781823937) [ 6687.906210] Lustre: DEBUG MARKER: == recovery-small test 115a: read: late REQ MDunlink and no bulk ========================================================== 19:06:21 (1781823981) [ 6688.858912] Lustre: *** cfs_fail_loc=51b, val=3*** [ 6709.304365] Lustre: DEBUG MARKER: == recovery-small test 115b: write: late REQ MDunlink and no bulk ========================================================== 19:06:42 (1781824002) [ 6710.981139] Lustre: *** cfs_fail_loc=51b, val=4*** [ 6718.817582] Lustre: DEBUG MARKER: recovery-small test_115b: @@@@@@ FAIL: dd success [ 6772.378730] Lustre: DEBUG MARKER: == recovery-small test 115c: read: late Reply MDunlink and no bulk ========================================================== 19:07:45 (1781824065) [ 6773.224123] Lustre: *** cfs_fail_loc=50f, val=3*** [ 6793.883587] Lustre: DEBUG MARKER: == recovery-small test 115d: write: late Reply MDunlink and no bulk ========================================================== 19:08:08 (1781824088) [ 6794.548462] Lustre: *** cfs_fail_loc=50f, val=4*** [ 6814.859867] Lustre: DEBUG MARKER: == recovery-small test 115e: read: late Bulk MDunlink and no reply ========================================================== 19:08:28 (1781824108) [ 6815.547780] Lustre: *** cfs_fail_loc=510, val=3*** [ 6837.521189] Lustre: DEBUG MARKER: == recovery-small test 115f: read: late REQ MDunlink and no reply ========================================================== 19:08:50 (1781824130) [ 6838.629322] Lustre: *** cfs_fail_loc=51b, val=3*** [ 6897.695635] Lustre: 2355:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781824140/real 1781824140] req@ffff96fd041ced80 x1868371001395584/t0(0) o3->lustre-OST0000-osc-ffff96fd09ab3800@192.168.201.103@tcp:6/4 lens 488/4536 e 0 to 1 dl 1781824195 ref 2 fl Rpc:XQr/600/ffffffff rc 0/-1 job:'dd.0' uid:0 gid:0 projid:0 [ 6897.788247] Lustre: 2355:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 18 previous similar messages [ 6897.817856] Lustre: lustre-OST0000-osc-ffff96fd09ab3800: Connection to lustre-OST0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 6897.880193] Lustre: Skipped 4 previous similar messages [ 6916.220773] Lustre: DEBUG MARKER: == recovery-small test 115g: read: late REQ MDunlink and Reply MDunlink ========================================================== 19:10:09 (1781824209) [ 6917.345578] Lustre: *** cfs_fail_loc=51c, val=3*** [ 6996.643935] Lustre: DEBUG MARKER: == recovery-small test 120: flock race: completion vs. evict ========================================================== 19:11:30 (1781824290) [ 6997.516230] LustreError: 82937:0:(ldlm_flock.c:804:ldlm_flock_completion_ast()) cfs_fail_timeout id 320 sleeping for 4000ms [ 7001.575153] LustreError: 82937:0:(ldlm_flock.c:804:ldlm_flock_completion_ast()) cfs_fail_timeout id 320 awake [ 7003.148122] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -107 [ 7003.241178] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 7003.385908] Lustre: lustre-MDT0000-mdc-ffff96fd09ab3800: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 7003.437115] Lustre: Skipped 6 previous similar messages [ 7007.057826] LustreError: 82968:0:(ldlm_flock.c:804:ldlm_flock_completion_ast()) cfs_fail_timeout id 320 sleeping for 4000ms [ 7011.063176] LustreError: 82968:0:(ldlm_flock.c:804:ldlm_flock_completion_ast()) cfs_fail_timeout id 320 awake [ 7011.421556] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 7020.885996] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 7022.626581] LustreError: 83024:0:(ldlm_flock.c:809:ldlm_flock_completion_ast()) cfs_fail_timeout id 321 sleeping for 4000ms [ 7026.659043] LustreError: 83024:0:(ldlm_flock.c:809:ldlm_flock_completion_ast()) cfs_fail_timeout id 321 awake [ 7027.237707] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 7030.755355] LustreError: 83054:0:(ldlm_flock.c:809:ldlm_flock_completion_ast()) cfs_fail_timeout id 321 sleeping for 4000ms [ 7034.824256] LustreError: 83054:0:(ldlm_flock.c:809:ldlm_flock_completion_ast()) cfs_fail_timeout id 321 awake [ 7035.105899] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 7043.698417] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 7045.320472] LustreError: 83110:0:(ldlm_flock.c:858:ldlm_flock_completion_ast()) cfs_fail_timeout id 322 sleeping for 4000ms [ 7049.384610] LustreError: 83110:0:(ldlm_flock.c:858:ldlm_flock_completion_ast()) cfs_fail_timeout id 322 awake [ 7050.130374] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 7050.160133] LustreError: 83125:0:(ldlm_resource.c:1180:ldlm_resource_complain()) lustre-MDT0000-mdc-ffff96fd09ab3800: namespace resource [0x200000007:0x1:0x0].0x0 (ffff96fd05c9af00) refcount nonzero (1) after lock cleanup; forcing cleanup. [ 7055.270520] Lustre: DEBUG MARKER: recovery-small test_120: @@@@@@ FAIL: DEADLOCK race failed [ 7103.691443] Lustre: DEBUG MARKER: == recovery-small test 113: ldlm enqueue dropped reply should not cause deadlocks ========================================================== 19:13:16 (1781824396) [ 7130.616587] LustreError: 68003:0:(ldlm_lockd.c:2883:ldlm_bl_thread_blwi()) cfs_fail_timeout id 31f sleeping for 4000ms [ 7134.618620] LustreError: 68003:0:(ldlm_lockd.c:2883:ldlm_bl_thread_blwi()) cfs_fail_timeout id 31f awake [ 7154.428701] Lustre: DEBUG MARKER: == recovery-small test 130a: enqueue resend on not existing file ========================================================== 19:14:07 (1781824447) [ 7219.423954] Lustre: DEBUG MARKER: == recovery-small test 130b: enqueue resend on a stale inode ========================================================== 19:15:13 (1781824513) [ 7297.126406] Lustre: DEBUG MARKER: == recovery-small test 130c: layout intent resend on a stale inode ========================================================== 19:16:31 (1781824591) [ 7348.885727] Lustre: DEBUG MARKER: == recovery-small test 132: long punch =================== 19:17:23 (1781824643) [ 7350.071405] Lustre: Mounted lustre-client [ 7478.622251] Lustre: Unmounted lustre-client [ 7494.452705] Lustre: DEBUG MARKER: == recovery-small test 131: IO vs evict results to IO under staled lock ========================================================== 19:19:48 (1781824788) [ 7496.701456] LustreError: 87553:0:(osc_request.c:2989:osc_build_rpc()) cfs_fail_timeout id 414 sleeping for 10000ms [ 7504.242041] LustreError: lustre-OST0000-osc-ffff96fd09ab3800: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -107 [ 7504.274128] LustreError: Skipped 6 previous similar messages [ 7504.283674] Lustre: lustre-OST0000-osc-ffff96fd09ab3800: Connection to lustre-OST0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 7504.343067] Lustre: Skipped 11 previous similar messages [ 7504.397647] LustreError: lustre-OST0000-osc-ffff96fd09ab3800: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [ 7504.464249] LustreError: 87636:0:(ldlm_resource.c:1180:ldlm_resource_complain()) lustre-OST0000-osc-ffff96fd09ab3800: namespace resource [0x240000400:0x2b2c:0x0].0x0 (ffff96fd05c21100) refcount nonzero (1) after lock cleanup; forcing cleanup. [ 7505.399145] LustreError: 87553:0:(osc_request.c:2989:osc_build_rpc()) cfs_fail_timeout interrupted [ 7505.431234] Lustre: 2356:0:(llite_lib.c:4198:ll_dirty_page_discard_warn()) lustre: dirty page discard: 192.168.201.103@tcp:/lustre/fid: [0x20000bf87:0x9:0x0]// may get corrupted (rc -108) [ 7520.997881] Lustre: DEBUG MARKER: == recovery-small test 133: don't fail on flock resend === 19:20:15 (1781824815) [ 7580.641377] Lustre: 88224:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781824823/real 1781824823] req@ffff96fd05c94700 x1868371001525632/t0(0) o101->lustre-MDT0000-mdc-ffff96fd09ab3800@192.168.201.103@tcp:12/10 lens 328/344 e 0 to 1 dl 1781824878 ref 2 fl Rpc:XQr/600/ffffffff rc 0/-1 job:'multiop.0' uid:0 gid:0 projid:0 [ 7580.798134] Lustre: 88224:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 4 previous similar messages [ 7601.966774] Lustre: DEBUG MARKER: == recovery-small test 134: race between failover and search for reply data free slot ========================================================== 19:21:34 (1781824894) [ 7606.220896] Lustre: DEBUG MARKER: SKIP: recovery-small test_134 Need 2+ clients, have 1 [ 7612.051778] Lustre: DEBUG MARKER: == recovery-small test 135: DOM: open/create resend to return size ========================================================== 19:21:45 (1781824905) [ 7672.853291] Lustre: lustre-MDT0000-mdc-ffff96fd09ab3800: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 7672.905548] Lustre: Skipped 10 previous similar messages [ 7691.947196] Lustre: DEBUG MARKER: SKIP: recovery-small test_136 skipping excluded test 136 [ 7696.271410] Lustre: DEBUG MARKER: == recovery-small test 137: late resend must be skipped if already applied ========================================================== 19:23:10 (1781824990) [ 7732.346697] Lustre: DEBUG MARKER: == recovery-small test 138: Umount MDT during recovery === 19:23:46 (1781825026) [ 7736.724864] Lustre: DEBUG MARKER: SKIP: recovery-small test_138 needs >= 2 MDTs [ 7742.250346] Lustre: DEBUG MARKER: == recovery-small test 139: corrupted catid won't cause crash ========================================================== 19:23:55 (1781825035) [ 7746.021121] Lustre: DEBUG MARKER: SKIP: recovery-small test_139 needs >= 2 MDTs [ 7750.702124] Lustre: DEBUG MARKER: == recovery-small test 140a: local mount is flagged properly ========================================================== 19:24:04 (1781825044) [ 7837.103929] Lustre: DEBUG MARKER: == recovery-small test 140b: local mount is excluded from recovery ========================================================== 19:25:31 (1781825131) [ 7865.817493] Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 [ 7901.727650] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 7912.014427] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9634990 to 0xaa643b29f9635fa8 [ 7948.874477] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 7953.565381] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 7975.899604] Lustre: DEBUG MARKER: == recovery-small test 141: do not lose locks on MGS restart ========================================================== 19:27:50 (1781825270) [ 7982.471728] Lustre: DEBUG MARKER: SKIP: recovery-small test_141 cannot run in local mode or from build tree [ 7986.408123] Lustre: DEBUG MARKER: == recovery-small test 142: orphan name stub can be cleaned up in startup ========================================================== 19:28:00 (1781825280) [ 8013.826564] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 8013.872474] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9635fa8 to 0xaa643b29f96362f7 [ 8015.442977] Lustre: 68006:0:(mgc_request.c:1901:mgc_process_log()) MGC192.168.201.103@tcp: IR log lustre-cliir failed, not fatal: rc = -2 [ 8057.558290] Lustre: DEBUG MARKER: == recovery-small test 143: orphan cleanup thread shouldn't be blocked even delete failed ========================================================== 19:29:11 (1781825351) [ 8081.376122] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 8126.472764] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f96362f7 to 0xaa643b29f963676c [ 8168.258657] Lustre: DEBUG MARKER: == recovery-small test 144a: MDT failover should stop precreation threads ========================================================== 19:31:02 (1781825462) [ 8187.057169] LustreError: lustre-OST0000-osc-ffff96fd09ab3800: operation ost_setattr to node 192.168.201.103@tcp failed: rc = -107 [ 8187.076459] Lustre: lustre-OST0000-osc-ffff96fd09ab3800: Connection to lustre-OST0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 8187.137268] Lustre: Skipped 6 previous similar messages [ 8283.794392] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 8292.318826] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 8356.831152] INFO: task touch:95830 blocked for more than 120 seconds. [ 8356.844622] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8356.868472] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8356.897737] task:touch state:D stack:0 pid:95830 ppid:95562 flags:0x80000000 [ 8356.947233] Call Trace: [ 8356.958236] __schedule+0x351/0xcb0 [ 8356.969220] schedule+0xc0/0x180 [ 8356.971360] schedule_preempt_disabled+0x21/0x40 [ 8356.982331] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8356.989152] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8357.005724] down_write+0x80/0xd0 [ 8357.012605] do_last+0x2eb/0xfc0 [ 8357.026058] ? nd_jump_root+0xe5/0x160 [ 8357.035101] ? path_init+0x437/0x520 [ 8357.053320] path_openat+0xf7/0x500 [ 8357.058810] ? __mod_lruvec_state+0x5a/0x80 [ 8357.070488] do_filp_open+0x99/0x140 [ 8357.076801] ? getname_flags+0x6e/0x330 [ 8357.098793] ? __check_object_size+0xff/0x256 [ 8357.114174] ? do_raw_spin_unlock+0x75/0x190 [ 8357.116535] ? _raw_spin_unlock+0x12/0x30 [ 8357.139731] do_sys_openat2+0x2b4/0x410 [ 8357.154833] do_sys_open+0x73/0xa0 [ 8357.165266] __x64_sys_openat+0x24/0x30 [ 8357.174739] do_syscall_64+0xc1/0x440 [ 8357.180896] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8357.201218] RIP: 0033:0x7fe4fbd8b332 [ 8357.209793] Code: Unable to access opcode bytes at RIP 0x7fe4fbd8b308. [ 8357.232801] RSP: 002b:00007ffef29987d0 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8357.265091] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007fe4fbd8b332 [ 8357.283175] RDX: 0000000000000941 RSI: 00007ffef299ac57 RDI: 00000000ffffff9c [ 8357.305281] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8357.318389] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8357.343514] R13: 0000000000000001 R14: 00007ffef299ac57 R15: 00007fe4fc02e374 [ 8357.373272] INFO: task touch:95831 blocked for more than 120 seconds. [ 8357.382833] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8357.406310] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8357.439749] task:touch state:D stack:0 pid:95831 ppid:95562 flags:0x80000000 [ 8357.461464] Call Trace: [ 8357.484919] __schedule+0x351/0xcb0 [ 8357.500397] schedule+0xc0/0x180 [ 8357.514879] schedule_preempt_disabled+0x21/0x40 [ 8357.526185] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8357.533166] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8357.548222] down_write+0x80/0xd0 [ 8357.559964] do_last+0x2eb/0xfc0 [ 8357.563813] ? nd_jump_root+0xe5/0x160 [ 8357.573329] ? path_init+0x437/0x520 [ 8357.590545] path_openat+0xf7/0x500 [ 8357.598658] do_filp_open+0x99/0x140 [ 8357.615366] ? getname_flags+0x6e/0x330 [ 8357.624268] ? __check_object_size+0xff/0x256 [ 8357.645499] ? do_raw_spin_unlock+0x75/0x190 [ 8357.663130] ? _raw_spin_unlock+0x12/0x30 [ 8357.676784] do_sys_openat2+0x2b4/0x410 [ 8357.687564] do_sys_open+0x73/0xa0 [ 8357.700430] __x64_sys_openat+0x24/0x30 [ 8357.714847] do_syscall_64+0xc1/0x440 [ 8357.729177] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8357.740017] RIP: 0033:0x7f38e3dc8332 [ 8357.757369] Code: Unable to access opcode bytes at RIP 0x7f38e3dc8308. [ 8357.782127] RSP: 002b:00007fff7b67cdf0 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8357.806236] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007f38e3dc8332 [ 8357.830737] RDX: 0000000000000941 RSI: 00007fff7b67dc57 RDI: 00000000ffffff9c [ 8357.854293] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8357.871227] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8357.894766] R13: 0000000000000001 R14: 00007fff7b67dc57 R15: 00007f38e406b374 [ 8357.923710] INFO: task touch:95833 blocked for more than 120 seconds. [ 8357.929788] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8357.959145] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8357.981669] task:touch state:D stack:0 pid:95833 ppid:95562 flags:0x80000000 [ 8358.002839] Call Trace: [ 8358.010422] __schedule+0x351/0xcb0 [ 8358.021990] schedule+0xc0/0x180 [ 8358.032025] schedule_preempt_disabled+0x21/0x40 [ 8358.043114] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8358.050416] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8358.066441] down_write+0x80/0xd0 [ 8358.071565] do_last+0x2eb/0xfc0 [ 8358.079857] ? nd_jump_root+0xe5/0x160 [ 8358.090818] ? path_init+0x437/0x520 [ 8358.098804] path_openat+0xf7/0x500 [ 8358.101661] do_filp_open+0x99/0x140 [ 8358.118975] ? getname_flags+0x6e/0x330 [ 8358.134688] ? __check_object_size+0xff/0x256 [ 8358.145029] ? do_raw_spin_unlock+0x75/0x190 [ 8358.155657] ? _raw_spin_unlock+0x12/0x30 [ 8358.158033] do_sys_openat2+0x2b4/0x410 [ 8358.168921] do_sys_open+0x73/0xa0 [ 8358.188391] __x64_sys_openat+0x24/0x30 [ 8358.199306] do_syscall_64+0xc1/0x440 [ 8358.216042] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8358.233251] RIP: 0033:0x7f25bb572332 [ 8358.239423] Code: Unable to access opcode bytes at RIP 0x7f25bb572308. [ 8358.261514] RSP: 002b:00007ffd519ee510 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8358.289318] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007f25bb572332 [ 8358.313420] RDX: 0000000000000941 RSI: 00007ffd519f0c57 RDI: 00000000ffffff9c [ 8358.333193] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8358.362499] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8358.394698] R13: 0000000000000001 R14: 00007ffd519f0c57 R15: 00007f25bb815374 [ 8358.417309] INFO: task touch:95834 blocked for more than 120 seconds. [ 8358.434702] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8358.466704] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8358.487226] task:touch state:D stack:0 pid:95834 ppid:95562 flags:0x80000000 [ 8358.524461] Call Trace: [ 8358.535114] __schedule+0x351/0xcb0 [ 8358.560663] schedule+0xc0/0x180 [ 8358.572330] schedule_preempt_disabled+0x21/0x40 [ 8358.585474] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8358.598232] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8358.616834] down_write+0x80/0xd0 [ 8358.626143] do_last+0x2eb/0xfc0 [ 8358.630478] ? nd_jump_root+0xe5/0x160 [ 8358.635554] ? path_init+0x437/0x520 [ 8358.642562] path_openat+0xf7/0x500 [ 8358.656637] do_filp_open+0x99/0x140 [ 8358.658651] ? getname_flags+0x6e/0x330 [ 8358.666966] ? __check_object_size+0xff/0x256 [ 8358.682388] ? do_raw_spin_unlock+0x75/0x190 [ 8358.697047] ? _raw_spin_unlock+0x12/0x30 [ 8358.714804] do_sys_openat2+0x2b4/0x410 [ 8358.725625] do_sys_open+0x73/0xa0 [ 8358.731076] __x64_sys_openat+0x24/0x30 [ 8358.746149] do_syscall_64+0xc1/0x440 [ 8358.756403] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8358.776136] RIP: 0033:0x7fa369722332 [ 8358.783651] Code: Unable to access opcode bytes at RIP 0x7fa369722308. [ 8358.810284] RSP: 002b:00007ffe7d14b100 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8358.836039] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007fa369722332 [ 8358.864689] RDX: 0000000000000941 RSI: 00007ffe7d14cc57 RDI: 00000000ffffff9c [ 8358.884331] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8358.907928] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8358.916488] R13: 0000000000000001 R14: 00007ffe7d14cc57 R15: 00007fa3699c5374 [ 8358.937124] INFO: task touch:95835 blocked for more than 120 seconds. [ 8358.955271] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8358.981960] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8359.007964] task:touch state:D stack:0 pid:95835 ppid:95562 flags:0x80000000 [ 8359.038638] Call Trace: [ 8359.049480] __schedule+0x351/0xcb0 [ 8359.055143] schedule+0xc0/0x180 [ 8359.059673] schedule_preempt_disabled+0x21/0x40 [ 8359.066512] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8359.084467] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8359.101740] down_write+0x80/0xd0 [ 8359.120438] do_last+0x2eb/0xfc0 [ 8359.129115] ? nd_jump_root+0xe5/0x160 [ 8359.138786] ? path_init+0x437/0x520 [ 8359.155580] path_openat+0xf7/0x500 [ 8359.170271] do_filp_open+0x99/0x140 [ 8359.186396] ? getname_flags+0x6e/0x330 [ 8359.202246] ? __check_object_size+0xff/0x256 [ 8359.227602] ? do_raw_spin_unlock+0x75/0x190 [ 8359.242748] ? _raw_spin_unlock+0x12/0x30 [ 8359.253351] do_sys_openat2+0x2b4/0x410 [ 8359.257754] do_sys_open+0x73/0xa0 [ 8359.280648] __x64_sys_openat+0x24/0x30 [ 8359.296643] do_syscall_64+0xc1/0x440 [ 8359.308838] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8359.332091] RIP: 0033:0x7f0f8bcc8332 [ 8359.342264] Code: Unable to access opcode bytes at RIP 0x7f0f8bcc8308. [ 8359.380503] RSP: 002b:00007ffe5a9ac0c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8359.424762] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007f0f8bcc8332 [ 8359.464180] RDX: 0000000000000941 RSI: 00007ffe5a9acc57 RDI: 00000000ffffff9c [ 8359.479463] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8359.493288] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8359.509856] R13: 0000000000000001 R14: 00007ffe5a9acc57 R15: 00007f0f8bf6b374 [ 8359.529101] INFO: task touch:95836 blocked for more than 120 seconds. [ 8359.537894] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8359.565593] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8359.584413] task:touch state:D stack:0 pid:95836 ppid:95562 flags:0x80000004 [ 8359.615214] Call Trace: [ 8359.616607] __schedule+0x351/0xcb0 [ 8359.636415] schedule+0xc0/0x180 [ 8359.638412] schedule_preempt_disabled+0x21/0x40 [ 8359.658033] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8359.666060] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8359.702044] down_write+0x80/0xd0 [ 8359.714291] do_last+0x2eb/0xfc0 [ 8359.734177] ? nd_jump_root+0xe5/0x160 [ 8359.743625] ? path_init+0x437/0x520 [ 8359.759453] path_openat+0xf7/0x500 [ 8359.773283] ? __mod_lruvec_state+0x5a/0x80 [ 8359.775696] do_filp_open+0x99/0x140 [ 8359.787676] ? getname_flags+0x6e/0x330 [ 8359.795877] ? __check_object_size+0xff/0x256 [ 8359.809786] ? do_raw_spin_unlock+0x75/0x190 [ 8359.822721] ? _raw_spin_unlock+0x12/0x30 [ 8359.841496] do_sys_openat2+0x2b4/0x410 [ 8359.856326] do_sys_open+0x73/0xa0 [ 8359.872621] __x64_sys_openat+0x24/0x30 [ 8359.880754] do_syscall_64+0xc1/0x440 [ 8359.895900] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8359.905869] RIP: 0033:0x7fd38860c332 [ 8359.914485] Code: Unable to access opcode bytes at RIP 0x7fd38860c308. [ 8359.928742] RSP: 002b:00007fff1b2d4430 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8359.941812] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007fd38860c332 [ 8359.953911] RDX: 0000000000000941 RSI: 00007fff1b2d5c57 RDI: 00000000ffffff9c [ 8359.974230] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8359.990824] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8359.998156] R13: 0000000000000001 R14: 00007fff1b2d5c57 R15: 00007fd3888af374 [ 8360.029331] INFO: task touch:95837 blocked for more than 120 seconds. [ 8360.038424] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8360.060024] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8360.077977] task:touch state:D stack:0 pid:95837 ppid:95562 flags:0x80000004 [ 8360.108369] Call Trace: [ 8360.117324] __schedule+0x351/0xcb0 [ 8360.130387] schedule+0xc0/0x180 [ 8360.135941] schedule_preempt_disabled+0x21/0x40 [ 8360.147189] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8360.162026] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8360.174765] down_write+0x80/0xd0 [ 8360.183852] do_last+0x2eb/0xfc0 [ 8360.190657] ? nd_jump_root+0xe5/0x160 [ 8360.204513] ? path_init+0x437/0x520 [ 8360.222379] path_openat+0xf7/0x500 [ 8360.229214] ? __mod_lruvec_state+0x5a/0x80 [ 8360.240994] do_filp_open+0x99/0x140 [ 8360.247445] ? getname_flags+0x6e/0x330 [ 8360.255065] ? __check_object_size+0xff/0x256 [ 8360.260825] ? do_raw_spin_unlock+0x75/0x190 [ 8360.268693] ? _raw_spin_unlock+0x12/0x30 [ 8360.277061] do_sys_openat2+0x2b4/0x410 [ 8360.284170] do_sys_open+0x73/0xa0 [ 8360.293138] __x64_sys_openat+0x24/0x30 [ 8360.300859] do_syscall_64+0xc1/0x440 [ 8360.306589] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8360.312500] RIP: 0033:0x7f40f44df332 [ 8360.326154] Code: Unable to access opcode bytes at RIP 0x7f40f44df308. [ 8360.341936] RSP: 002b:00007ffd72c64db0 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8360.359023] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007f40f44df332 [ 8360.379731] RDX: 0000000000000941 RSI: 00007ffd72c66c57 RDI: 00000000ffffff9c [ 8360.409839] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8360.429475] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8360.452493] R13: 0000000000000001 R14: 00007ffd72c66c57 R15: 00007f40f4782374 [ 8360.572761] INFO: task touch:95838 blocked for more than 120 seconds. [ 8360.583232] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8360.594851] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8360.617482] task:touch state:D stack:0 pid:95838 ppid:95562 flags:0x80000004 [ 8360.633979] Call Trace: [ 8360.635480] __schedule+0x351/0xcb0 [ 8360.642160] schedule+0xc0/0x180 [ 8360.652599] schedule_preempt_disabled+0x21/0x40 [ 8360.660512] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8360.676297] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8360.689178] down_write+0x80/0xd0 [ 8360.702087] do_last+0x2eb/0xfc0 [ 8360.709150] ? nd_jump_root+0xe5/0x160 [ 8360.726648] ? path_init+0x437/0x520 [ 8360.737044] path_openat+0xf7/0x500 [ 8360.751573] do_filp_open+0x99/0x140 [ 8360.765527] ? getname_flags+0x6e/0x330 [ 8360.780247] ? __check_object_size+0xff/0x256 [ 8360.796918] ? do_raw_spin_unlock+0x75/0x190 [ 8360.813439] ? _raw_spin_unlock+0x12/0x30 [ 8360.822290] do_sys_openat2+0x2b4/0x410 [ 8360.829224] do_sys_open+0x73/0xa0 [ 8360.841186] __x64_sys_openat+0x24/0x30 [ 8360.844573] do_syscall_64+0xc1/0x440 [ 8360.848093] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8360.860575] RIP: 0033:0x7fdd94502332 [ 8360.863343] Code: Unable to access opcode bytes at RIP 0x7fdd94502308. [ 8360.873437] RSP: 002b:00007fff575579f0 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8360.887477] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007fdd94502332 [ 8360.915454] RDX: 0000000000000941 RSI: 00007fff57559c57 RDI: 00000000ffffff9c [ 8360.935158] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8360.963389] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8360.982160] R13: 0000000000000001 R14: 00007fff57559c57 R15: 00007fdd947a5374 [ 8390.623154] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781825672/real 1781825672] req@ffff96fd38566a00 x1868371007038848/t0(0) o400->MGC192.168.201.103@tcp@192.168.201.103@tcp:26/25 lens 224/224 e 0 to 1 dl 1781825688 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 8390.715297] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 13 previous similar messages [ 8390.757412] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 8416.459726] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f963676c to 0xaa643b29f96374fc [ 8416.493873] Lustre: MGC192.168.201.103@tcp: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 8416.535435] Lustre: Skipped 7 previous similar messages [ 8449.478610] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 8453.655115] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8483.295635] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 8493.609530] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f96374fc to 0xaa643b29f96377aa [ 8528.275317] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid [ 8532.732612] Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec [ 8839.619420] Lustre: DEBUG MARKER: == recovery-small test 144b: orphan cleanup shouldn't be blocked for no objects+failover situation ========================================================== 19:42:13 (1781826133) [ 8859.438751] LustreError: lustre-OST0000-osc-ffff96fd09ab3800: operation ost_setattr to node 192.168.201.103@tcp failed: rc = -19 [ 8859.445861] Lustre: lustre-OST0000-osc-ffff96fd09ab3800: Connection to lustre-OST0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 8859.516764] LustreError: Skipped 7 previous similar messages [ 8859.534481] Lustre: Skipped 2 previous similar messages [ 8964.391493] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [ 8970.386962] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [ 8975.327141] INFO: task touch:98970 blocked for more than 120 seconds. [ 8975.336353] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8975.370945] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8975.389878] task:touch state:D stack:0 pid:98970 ppid:98714 flags:0x80000000 [ 8975.430514] Call Trace: [ 8975.439709] __schedule+0x351/0xcb0 [ 8975.457805] schedule+0xc0/0x180 [ 8975.467940] schedule_preempt_disabled+0x21/0x40 [ 8975.481671] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8975.484948] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8975.513509] down_write+0x80/0xd0 [ 8975.532370] do_last+0x2eb/0xfc0 [ 8975.544167] ? nd_jump_root+0xe5/0x160 [ 8975.554368] ? path_init+0x437/0x520 [ 8975.576071] path_openat+0xf7/0x500 [ 8975.591941] do_filp_open+0x99/0x140 [ 8975.609900] ? getname_flags+0x6e/0x330 [ 8975.624188] ? __check_object_size+0xff/0x256 [ 8975.638097] ? do_raw_spin_unlock+0x75/0x190 [ 8975.666909] ? _raw_spin_unlock+0x12/0x30 [ 8975.691401] do_sys_openat2+0x2b4/0x410 [ 8975.700324] do_sys_open+0x73/0xa0 [ 8975.723107] __x64_sys_openat+0x24/0x30 [ 8975.728241] do_syscall_64+0xc1/0x440 [ 8975.744317] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8975.760765] RIP: 0033:0x7f6bd7139332 [ 8975.767874] Code: Unable to access opcode bytes at RIP 0x7f6bd7139308. [ 8975.778597] RSP: 002b:00007ffd5c251df0 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8975.793335] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007f6bd7139332 [ 8975.807487] RDX: 0000000000000941 RSI: 00007ffd5c253c57 RDI: 00000000ffffff9c [ 8975.827955] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8975.845387] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8975.859096] R13: 0000000000000001 R14: 00007ffd5c253c57 R15: 00007f6bd73dc374 [ 8975.893858] INFO: task touch:98971 blocked for more than 120 seconds. [ 8975.904236] Tainted: G O -------- - - 4.18.0rh8.10-debug #2 [ 8975.914321] "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. [ 8975.926893] task:touch state:D stack:0 pid:98971 ppid:98714 flags:0x80000000 [ 8975.936954] Call Trace: [ 8975.943384] __schedule+0x351/0xcb0 [ 8975.961902] schedule+0xc0/0x180 [ 8975.977132] schedule_preempt_disabled+0x21/0x40 [ 8975.986421] rwsem_down_write_slowpath+0x5d7/0xa40 [ 8975.997980] ? lprocfs_counter_add+0x15b/0x210 [obdclass] [ 8976.012892] down_write+0x80/0xd0 [ 8976.016026] do_last+0x2eb/0xfc0 [ 8976.019869] ? nd_jump_root+0xe5/0x160 [ 8976.024173] ? path_init+0x437/0x520 [ 8976.031969] path_openat+0xf7/0x500 [ 8976.044799] ? __mod_lruvec_state+0x5a/0x80 [ 8976.056898] do_filp_open+0x99/0x140 [ 8976.073414] ? getname_flags+0x6e/0x330 [ 8976.085781] ? __check_object_size+0xff/0x256 [ 8976.095304] ? do_raw_spin_unlock+0x75/0x190 [ 8976.105209] ? _raw_spin_unlock+0x12/0x30 [ 8976.114500] do_sys_openat2+0x2b4/0x410 [ 8976.144120] do_sys_open+0x73/0xa0 [ 8976.158432] __x64_sys_openat+0x24/0x30 [ 8976.171177] do_syscall_64+0xc1/0x440 [ 8976.182949] entry_SYSCALL_64_after_hwframe+0x49/0xae [ 8976.205354] RIP: 0033:0x7f58afeec332 [ 8976.220270] Code: Unable to access opcode bytes at RIP 0x7f58afeec308. [ 8976.243506] RSP: 002b:00007fff2a5350e0 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 [ 8976.265918] RAX: ffffffffffffffda RBX: 00000000ffffffff RCX: 00007f58afeec332 [ 8976.284865] RDX: 0000000000000941 RSI: 00007fff2a536c57 RDI: 00000000ffffff9c [ 8976.311879] RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001 [ 8976.324089] R10: 00000000000001b6 R11: 0000000000000246 R12: 0000000000000001 [ 8976.351805] R13: 0000000000000001 R14: 00007fff2a536c57 R15: 00007f58b018f374 [ 9421.086508] Lustre: DEBUG MARKER: == recovery-small test 144c: reconnection during orphan cleanup shouldn't lose LAST_ID synchronization ========================================================== 19:51:54 (1781826714) [ 9759.237390] Lustre: lustre-MDT0000-mdc-ffff96fd09ab3800: Connection to lustre-MDT0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [ 9779.679464] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781827061/real 1781827061] req@ffff96fd06ae3800 x1868371018274944/t0(0) o400->MGC192.168.201.103@tcp@192.168.201.103@tcp:26/25 lens 224/224 e 0 to 1 dl 1781827077 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [ 9779.753413] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 8 previous similar messages [ 9779.789579] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [ 9779.841943] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f96377aa to 0xaa643b29f9731a37 [ 9779.894774] Lustre: MGC192.168.201.103@tcp: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 9779.917428] Lustre: Skipped 3 previous similar messages [ 9781.274516] Lustre: 68006:0:(mgc_request.c:1901:mgc_process_log()) MGC192.168.201.103@tcp: IR log lustre-cliir failed, not fatal: rc = -2 [ 9853.261313] Lustre: DEBUG MARKER: == recovery-small test 145: connect mdtlovs and process update logs after recovery expire ========================================================== 19:59:07 (1781827147) [ 9856.681383] Lustre: DEBUG MARKER: SKIP: recovery-small test_145 needs >= 3 MDTs [ 9861.372694] Lustre: DEBUG MARKER: == recovery-small test 146: test eviction is counted properly ========================================================== 19:59:14 (1781827154) [ 9865.576817] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: operation ldlm_enqueue to node 192.168.201.103@tcp failed: rc = -107 [ 9865.610408] LustreError: Skipped 332 previous similar messages [ 9865.650954] LustreError: lustre-MDT0000-mdc-ffff96fd09ab3800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [ 9865.697194] Lustre: lustre-MDT0000-mdc-ffff96fd09ab3800: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [ 9865.711858] Lustre: Skipped 1 previous similar message [ 9882.515632] Lustre: DEBUG MARKER: == recovery-small test 147: Check client reconnect ======= 19:59:36 (1781827176) [10058.116697] Lustre: DEBUG MARKER: == recovery-small test 148: data corruption through resend ========================================================== 20:02:32 (1781827352) [10089.954646] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781827367/real 1781827367] req@ffff96fd30c8f800 x1868371018299392/t0(0) o4->lustre-OST0000-osc-ffff96fd09ab3800@192.168.201.103@tcp:6/4 lens 4584/448 e 0 to 1 dl 1781827387 ref 2 fl Rpc:XQr/600/ffffffff rc 0/-1 job:'dd.0' uid:0 gid:0 projid:0 [10121.402216] Lustre: DEBUG MARKER: == recovery-small test 149: skip orphan removal at umount ========================================================== 20:03:35 (1781827415) [10124.522634] Lustre: DEBUG MARKER: SKIP: recovery-small test_149 needs >= 2 MDTs [10127.720172] Lustre: DEBUG MARKER: == recovery-small test 150: statfs when MDT0 offline with lazystatfs option ========================================================== 20:03:42 (1781827422) [10131.000298] Lustre: DEBUG MARKER: SKIP: recovery-small test_150 needs >= 2 MDTs [10134.589416] Lustre: DEBUG MARKER: == recovery-small test 152: QoS object allocation could be awakened in case of OST failover ========================================================== 20:03:49 (1781827429) [10187.353841] Lustre: DEBUG MARKER: == recovery-small test 153: evict vs reconnect race ====== 20:04:41 (1781827481) [10239.456568] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [10248.695587] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9731a37 to 0xaa643b29f97341d6 [10248.752747] Lustre: MGC192.168.201.103@tcp: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [10279.392362] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781827520/real 1781827520] req@ffff96fd05d32a00 x1868371018456832/t0(0) o400->lustre-MDT0000-mdc-ffff96fd09ab3800@192.168.201.103@tcp:12/10 lens 224/224 e 0 to 1 dl 1781827576 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [10279.475904] Lustre: 2356:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 1 previous similar message [10293.036511] Lustre: DEBUG MARKER: == recovery-small test 154a: corruption update llog can be skipped ========================================================== 20:06:27 (1781827587) [10296.007969] Lustre: DEBUG MARKER: SKIP: recovery-small test_154a needs >= 2 MDTs [10300.546454] Lustre: DEBUG MARKER: == recovery-small test 154b: restore update llog after failed recovery ========================================================== 20:06:34 (1781827594) [10304.618870] Lustre: DEBUG MARKER: SKIP: recovery-small test_154b needs >= 2 MDTs [10308.071466] Lustre: DEBUG MARKER: == recovery-small test 155: failover after client remount ========================================================== 20:06:42 (1781827602) [10353.867530] Lustre: Unmounted lustre-client [10358.623882] Lustre: Mounted lustre-client [10365.507480] Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 [10390.559349] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [10400.754686] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9734c4f to 0xaa643b29f9734ca3 [10403.717977] Lustre: lustre-MDT0000-mdc-ffff96fd086a1800: Connection to lustre-MDT0000 (at 192.168.201.103@tcp) was lost; in progress operations using this service will wait for recovery to complete [10403.769089] Lustre: Skipped 3 previous similar messages [10437.915862] Lustre: DEBUG MARKER: == recovery-small test 156: tot_granted miscount after client eviction ========================================================== 20:08:51 (1781827731) [10455.206067] Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-OST0000 [10491.162411] LustreError: 2353:0:(client.c:3532:ptlrpc_replay_req()) cfs_fail_timeout id 536 sleeping for 45000ms [10536.193341] LustreError: 2353:0:(client.c:3532:ptlrpc_replay_req()) cfs_fail_timeout id 536 awake [10536.213801] LustreError: lustre-OST0000-osc-ffff96fd086a1800: operation ost_write to node 192.168.201.103@tcp failed: rc = -107 [10536.279122] LustreError: lustre-OST0000-osc-ffff96fd086a1800: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [10551.358563] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount (FULL|IDLE) osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [10554.523874] Lustre: DEBUG MARKER: osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid in FULL state after 0 sec [10578.952797] Lustre: DEBUG MARKER: == recovery-small test 157: eviction during mmaped i/o === 20:11:13 (1781827873) [10580.344751] LustreError: 109182:0:(vvp_io.c:1473:vvp_io_fault_start()) cfs_fail_timeout id 1432 sleeping for 3000ms [10583.406058] LustreError: 109182:0:(vvp_io.c:1473:vvp_io_fault_start()) cfs_fail_timeout id 1432 awake [10583.471671] LustreError: lustre-OST0000-osc-ffff96fd086a1800: This client was evicted by lustre-OST0000; in progress operations using this service will fail. [10596.711309] Lustre: DEBUG MARKER: == recovery-small test 158a: connect without access right ========================================================== 20:11:31 (1781827891) [10600.106750] Lustre: DEBUG MARKER: SKIP: recovery-small test_158a needs >= 2 MDTS [10604.959299] Lustre: DEBUG MARKER: == recovery-small test 160: MDT destroys are blocked by grouplocks ========================================================== 20:11:38 (1781827898) [10609.853764] LustreError: lustre-MDT0000-mdc-ffff96fd086a1800: This client was evicted by lustre-MDT0000; in progress operations using this service will fail. [10610.074452] Lustre: lustre-MDT0000-mdc-ffff96fd086a1800: Connection restored to 192.168.201.103@tcp (at 192.168.201.103@tcp) [10610.113826] Lustre: Skipped 3 previous similar messages [10670.426467] Lustre: DEBUG MARKER: == recovery-small test 161: evict osp by ping evictor ==== 20:12:45 (1781827965) [10673.933653] Lustre: DEBUG MARKER: SKIP: recovery-small test_161 needs >= 2 MDTs [10678.025978] Lustre: DEBUG MARKER: == recovery-small test 162: File attributes should be persisted after MDS failover ========================================================== 20:12:52 (1781827972) [10705.389471] LustreError: MGC192.168.201.103@tcp: Connection to MGS (at 192.168.201.103@tcp) was lost; in progress operations using this service will fail [10705.473690] Lustre: Evicted from MGS (at 192.168.201.103@tcp) after server handle changed from 0xaa643b29f9734ca3 to 0xaa643b29f9736221 [10707.577856] LustreError: 2353:0:(client.c:3447:ptlrpc_replay_interpret()) @@@ status 301, old was 0 req@ffff96fd05d33100 x1868371018615296/t133143986234(133143986234) o101->lustre-MDT0000-mdc-ffff96fd086a1800@192.168.201.103@tcp:12/10 lens 576/608 e 0 to 0 dl 1781828059 ref 2 fl Interpret:RPQU/604/0 rc 301/301 job:'chattr.0' uid:0 gid:0 projid:0 [10745.198616] Lustre: 2354:0:(client.c:2489:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1781827987/real 1781827987] req@ffff96fd305be680 x1868371018621056/t0(0) o400->lustre-MDT0000-mdc-ffff96fd086a1800@192.168.201.103@tcp:12/10 lens 224/224 e 0 to 1 dl 1781828042 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 [10745.327793] Lustre: 2354:0:(client.c:2489:ptlrpc_expire_one_request()) Skipped 14 previous similar messages [10748.950347] Lustre: DEBUG MARKER: == recovery-small test 163: changelog check for fail write and processing records ========================================================== 20:14:02 (1781828042) [10793.357750] Lustre: DEBUG MARKER: == recovery-small test 170: Reconnect after REPLAY_LOCKS hangs (LU-18154) ========================================================== 20:14:47 (1781828087) [10804.283402] Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-OST0000 [10876.961778] Lustre: DEBUG MARKER: oleg103-client.virtnet: executing wait_import_state_mount REPLAY_LOCKS osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid [12637.340396] Lustre: DEBUG MARKER: rpc test_170: @@@@@@ FAIL: can't put import for osc.lustre-OST0000-osc-[-0-9a-f]*.ost_server_uuid into REPLAY_LOCKS state after 1475 sec, have IDLE