# node crashed and i am unable to recover

**URL:** <https://forums.percona.com/t/node-crashed-and-i-am-unable-to-recover/4892>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [May 21, 2016, 1:25am UTC](https://forums.percona.com/t/node-crashed-and-i-am-unable-to-recover/4892 "2016-05-21T01:25:06Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![pavithra](https://avatars.discourse-cdn.com/v4/letter/p/e9a140/32.png) [@pavithra](https://forums.percona.com/u/pavithra)\
**Post date:** [May 21, 2016, 1:25am UTC](https://forums.percona.com/t/node-crashed-and-i-am-unable-to-recover/4892/1 "2016-05-21T01:25:06Z")

</div>

i have a 3 node PXC setup we i have added some performance variable to my config file in node1 and restarted i was unable to add the node1 to existing cluster . i have comments all the changes on tried to restart the node but this didn’t work out . please find below for error messages that were logged. Need help as it is production PXC setup .

160519 23:24:12 mysqld\_safe Skipping wsrep-recover for b4821329-1e39-11e6-8378-baa69bfe7dc9:4 pair  
160519 23:24:12 mysqld\_safe Assigning b4821329-1e39-11e6-8378-baa69bfe7dc9:4 to wsrep\_start\_position  
2016-05-19 23:24:12 0 [Warning] TIMESTAMP with implicit DEFAULT value is deprecated. Please use --explicit\_defaults\_for\_timestamp server option (see documentation for more details).  
2016-05-19 23:24:12 0 [Note] /usr/sbin/mysqld (mysqld 5.6.24-72.2-56) starting as process 2547 …  
2016-05-19 23:24:12 2547 [Note] WSREP: Read nil XID from storage engines, skipping position init  
2016-05-19 23:24:12 2547 [Note] WSREP: wsrep\_load(): loading provider library ‘/usr/lib64/libgalera\_smm.so’  
2016-05-19 23:24:12 2547 [Note] WSREP: wsrep\_load(): Galera 3.11(ra0189ab) by Codership Oy \<info@codership.com\> loaded successfully.  
2016-05-19 23:24:12 2547 [Note] WSREP: CRC-32C: using hardware acceleration.  
2016-05-19 23:24:12 2547 [Note] WSREP: Found saved state: b4821329-1e39-11e6-8378-baa69bfe7dc9:4  
2016-05-19 23:24:12 2547 [Note] WSREP: Passing config to GCS: base\_dir = /var/lib/mysql/; base\_host = 10.0.1.225; base\_port = 4567; cert.log\_conflicts = no; debug = no; evs.auto\_evict = 0; evs.delay\_margin = PT1S; evs.delayed\_keep\_period = PT30S; evs.inactive\_check\_period = PT0.5S; evs.inactive\_timeout = PT15S; evs.join\_retrans\_period = PT1S; evs.max\_install\_timeouts = 3; evs.send\_window = 4; evs.stats\_report\_period = PT1M; evs.suspect\_timeout = PT5S; evs.user\_send\_window = 2; evs.view\_forget\_timeout = PT24H; gcache.dir = /var/lib/mysql/; gcache.keep\_pages\_size = 0; gcache.mem\_size = 0; gcache.name = /var/lib/mysql//galera.cache; gcache.page\_size = 128M; gcache.size = 128M; gcs.fc\_debug = 0; gcs.fc\_factor = 1.0; gcs.fc\_limit = 16; gcs.fc\_master\_slave = no; gcs.max\_packet\_size = 64500; gcs.max\_throttle = 0.25; gcs.recv\_q\_hard\_limit = 9223372036854775807; gcs.recv\_q\_soft\_limit = 0.25; gcs.sync\_donor = no; gmcast.segment = 0; gmcast.version = 0; pc.announce\_timeout = PT3S; pc.checksum = false; pc.ignore\_quorum = false; pc.ignore\_sb = false; p  
2016-05-19 23:24:12 2547 [Note] WSREP: Service thread queue flushed.  
2016-05-19 23:24:12 2547 [Note] WSREP: Assign initial position for certification: 4, protocol version: -1  
2016-05-19 23:24:12 2547 [Note] WSREP: wsrep\_sst\_grab()  
2016-05-19 23:24:12 2547 [Note] WSREP: Start replication  
2016-05-19 23:24:12 2547 [Note] WSREP: Setting initial position to b4821329-1e39-11e6-8378-baa69bfe7dc9:4  
2016-05-19 23:24:12 2547 [Note] WSREP: protonet asio version 0  
2016-05-19 23:24:12 2547 [Note] WSREP: Using CRC-32C for message checksums.  
2016-05-19 23:24:12 2547 [Note] WSREP: backend: asio  
2016-05-19 23:24:12 2547 [Warning] WSREP: access file(/var/lib/mysql//gvwstate.dat) failed(No such file or directory)  
2016-05-19 23:24:12 2547 [Note] WSREP: restore pc from disk failed  
2016-05-19 23:24:12 2547 [Note] WSREP: GMCast version 0  
2016-05-19 23:24:12 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) listening at tcp://0.0.0.0:4567  
2016-05-19 23:24:12 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) multicast: , ttl: 1  
2016-05-19 23:24:12 2547 [Note] WSREP: EVS version 0  
2016-05-19 23:24:12 2547 [Note] WSREP: gcomm: connecting to group ‘my\_ubuntu\_cluster’, peer ‘10.0.1.225:,10.0.1.224:’  
2016-05-19 23:24:12 2547 [Warning] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) address ‘tcp://10.0.1.225:4567’ points to own listening address, blacklisting  
2016-05-19 23:24:12 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) address ‘tcp://10.0.1.225:4567’ pointing to uuid 529594ab is blacklisted, skipping  
2016-05-19 23:24:12 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) turning message relay requesting on, nonlive peers: tcp://10.0.1.226:4567  
2016-05-19 23:24:12 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) address ‘tcp://10.0.1.225:4567’ pointing to uuid 529594ab is blacklisted, skipping  
2016-05-19 23:24:12 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) address ‘tcp://10.0.1.225:4567’ pointing to uuid 529594ab is blacklisted, skipping  
2016-05-19 23:24:12 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) address ‘tcp://10.0.1.225:4567’ pointing to uuid 529594ab is blacklisted, skipping  
2016-05-19 23:24:12 2547 [Note] WSREP: declaring 56560110 at tcp://10.0.1.224:4567 stable  
2016-05-19 23:24:12 2547 [Note] WSREP: declaring 6e832bd2 at tcp://10.0.1.226:4567 stable  
2016-05-19 23:24:12 2547 [Note] WSREP: Node 56560110 state prim  
2016-05-19 23:24:12 2547 [Note] WSREP: view(view\_id(PRIM,529594ab,51) memb {  
529594ab,0  
56560110,0  
6e832bd2,0  
} joined {  
} left {  
} partitioned {  
})

---

<div class="post-metadata">

**Author:** ![pavithra](https://avatars.discourse-cdn.com/v4/letter/p/e9a140/32.png) [@pavithra](https://forums.percona.com/u/pavithra)\
**Post date:** [May 21, 2016, 1:26am UTC](https://forums.percona.com/t/node-crashed-and-i-am-unable-to-recover/4892/2 "2016-05-21T01:26:15Z")

</div>

2016-05-19 23:24:12 2547 [Note] WSREP: save pc into disk  
2016-05-19 23:24:13 2547 [Note] WSREP: gcomm: connected  
2016-05-19 23:24:13 2547 [Note] WSREP: Changing maximum packet size to 64500, resulting msg size: 32636  
2016-05-19 23:24:13 2547 [Note] WSREP: Shifting CLOSED → OPEN (TO: 0)  
2016-05-19 23:24:13 2547 [Note] WSREP: Opened channel ‘my\_ubuntu\_cluster’  
2016-05-19 23:24:13 2547 [Note] WSREP: Waiting for SST to complete.  
2016-05-19 23:24:13 2547 [Note] WSREP: New COMPONENT: primary = yes, bootstrap = no, my\_idx = 0, memb\_num = 3  
2016-05-19 23:24:13 2547 [Note] WSREP: STATE\_EXCHANGE: sent state UUID: 532ea367-1e3a-11e6-9dda-2a89c5480b29  
2016-05-19 23:24:13 2547 [Note] WSREP: STATE EXCHANGE: sent state msg: 532ea367-1e3a-11e6-9dda-2a89c5480b29  
2016-05-19 23:24:13 2547 [Note] WSREP: STATE EXCHANGE: got state msg: 532ea367-1e3a-11e6-9dda-2a89c5480b29 from 0 (ip-10-0-1-225.ap-southeast-1.compute.internal)  
2016-05-19 23:24:13 2547 [Note] WSREP: STATE EXCHANGE: got state msg: 532ea367-1e3a-11e6-9dda-2a89c5480b29 from 1 (ip-10-0-1-224.ap-southeast-1.compute.internal)  
2016-05-19 23:24:13 2547 [Note] WSREP: STATE EXCHANGE: got state msg: 532ea367-1e3a-11e6-9dda-2a89c5480b29 from 2 (ip-10-0-1-226.ap-southeast-1.compute.internal)  
2016-05-19 23:24:13 2547 [Note] WSREP: Quorum results:  
version = 3,  
component = PRIMARY,  
conf\_id = 50,  
members = 2/3 (joined/total),  
act\_id = 2183133,

protocols = 0/7/3 (gcs/repl/appl),  
group UUID = e8e2feed-1753-11e6-bfe3-972d646081e6  
2016-05-19 23:24:13 2547 [Note] WSREP: Flow-control interval: [28, 28]  
2016-05-19 23:24:13 2547 [Note] WSREP: Shifting OPEN → PRIMARY (TO: 2183133)  
2016-05-19 23:24:13 2547 [Note] WSREP: State transfer required:  
Group state: e8e2feed-1753-11e6-bfe3-972d646081e6:2183133  
Local state: b4821329-1e39-11e6-8378-baa69bfe7dc9:4  
2016-05-19 23:24:13 2547 [Note] WSREP: New cluster view: global state: e8e2feed-1753-11e6-bfe3-972d646081e6:2183133, view# 51: Primary, number of nodes: 3, my index: 0, protocol version 3  
2016-05-19 23:24:13 2547 [Warning] WSREP: Gap in state sequence. Need state transfer.  
2016-05-19 23:24:13 2547 [Note] WSREP: Running: ‘wsrep\_sst\_xtrabackup-v2 --role ‘joiner’ --address ‘10.0.1.225’ --auth ‘sstuser:s3cret’ --datadir ‘/var/lib/mysql/’ --defaults-file ‘/etc/my.cnf’ --defaults-group-suffix ‘’ --parent ‘2547’ ‘’ ’  
WSREP\_SST: [INFO] Streaming with xbstream (20160519 23:24:13.511)  
WSREP\_SST: [INFO] Using socat as streamer (20160519 23:24:13.512)  
WSREP\_SST: [INFO] Stale sst\_in\_progress file: /var/lib/mysql//sst\_in\_progress (20160519 23:24:13.515)  
WSREP\_SST: [INFO] Evaluating timeout -s9 100 socat -u TCP-LISTEN:4444,reuseaddr stdio | xbstream -x; RC=( ${PIPESTATUS[@]} ) (20160519 23:24:13.537)  
2016-05-19 23:24:13 2547 [Note] WSREP: Prepared SST request: xtrabackup-v2|10.0.1.225:4444/xtrabackup\_sst//1  
2016-05-19 23:24:13 2547 [Note] WSREP: wsrep\_notify\_cmd is not defined, skipping notification.  
2016-05-19 23:24:13 2547 [Note] WSREP: REPL Protocols: 7 (3, 2)  
2016-05-19 23:24:13 2547 [Note] WSREP: Service thread queue flushed.  
2016-05-19 23:24:13 2547 [Note] WSREP: Assign initial position for certification: 2183133, protocol version: 3  
2016-05-19 23:24:13 2547 [Note] WSREP: Service thread queue flushed.  
2016-05-19 23:24:13 2547 [Warning] WSREP: Failed to prepare for incremental state transfer: Local state UUID (b4821329-1e39-11e6-8378-baa69bfe7dc9) does not match group state UUID (e8e2feed-1753-11e6-bfe3-972d646081e6): 1 (Operation not permitted)  
at galera/src/replicator\_str.cpp:prepare\_for\_IST():463. IST will be unavailable.  
2016-05-19 23:24:13 2547 [Note] WSREP: Member 0.0 (ip-10-0-1-225.ap-southeast-1.compute.internal) requested state transfer from ‘_any_’. Selected 1.0 (ip-10-0-1-224.ap-southeast-1.compute.internal)(SYNCED) as donor.  
2016-05-19 23:24:13 2547 [Note] WSREP: Shifting PRIMARY → JOINER (TO: 2183133)  
2016-05-19 23:24:13 2547 [Note] WSREP: Requesting state transfer: success, donor: 1  
WSREP\_SST: [INFO] WARNING: Stale temporary SST directory: /var/lib/mysql//.sst from previous state transfer (20160519 23:24:13.932)  
WSREP\_SST: [INFO] Proceeding with SST (20160519 23:24:13.934)  
WSREP\_SST: [INFO] Evaluating socat -u TCP-LISTEN:4444,reuseaddr stdio | xbstream -x; RC=( ${PIPESTATUS[@]} ) (20160519 23:24:13.935)  
WSREP\_SST: [INFO] Cleaning the existing datadir and innodb-data/log directories (20160519 23:24:13.937)  
removed `/var/lib/mysql/auto.cnf' removed `/var/lib/mysql/ib\_logfile0’  
removed directory: `/var/lib/mysql/test' removed `/var/lib/mysql/performance\_schema/host\_cache.frm’

removed `/var/lib/mysql/performance_schema/setup_actors.frm' removed `/var/lib/mysql/performance\_schema/setup\_instruments.frm’  
removed directory: `/var/lib/mysql/performance_schema' removed `/var/lib/mysql/ib\_logfile1’

removed `/var/lib/mysql/mysql/proc.MYI' removed `/var/lib/mysql/mysql/slave\_worker\_info.frm’  
removed `/var/lib/mysql/mysql/time_zone.MYI' removed directory: `/var/lib/mysql/mysql’  
removed `/var/lib/mysql/ibdata1’  
WSREP\_SST: [INFO] Waiting for SST streaming to complete! (20160519 23:24:13.964)  
2016-05-19 23:24:15 2547 [Note] WSREP: (529594ab, ‘tcp://0.0.0.0:4567’) turning message relay requesting off  
WSREP\_SST: [ERROR] xtrabackup\_checkpoints missing, failed innobackupex/SST on donor (20160519 23:24:23.935)  
WSREP\_SST: [ERROR] Cleanup after exit with status:2 (20160519 23:24:23.936)  
2016-05-19 23:24:23 2547 [ERROR] WSREP: Process completed with error: wsrep\_sst\_xtrabackup-v2 --role ‘joiner’ --address ‘10.0.1.225’ --auth ‘sstuser:s3cret’ --datadir ‘/var/lib/mysql/’ --defaults-file ‘/etc/my.cnf’ --defaults-group-suffix ‘’ --parent ‘2547’ ‘’ : 2 (No such file or directory)  
2016-05-19 23:24:23 2547 [ERROR] WSREP: Failed to read uuid:seqno from joiner script.  
2016-05-19 23:24:23 2547 [ERROR] WSREP: SST failed: 2 (No such file or directory)  
2016-05-19 23:24:23 2547 [ERROR] Aborting

2016-05-19 23:24:23 2547 [Warning] WSREP: 1.0 (ip-10-0-1-224.ap-southeast-1.compute.internal): State transfer to 0.0 (ip-10-0-1-225.ap-southeast-1.compute.internal) failed: -22 (Invalid argument)  
2016-05-19 23:24:23 2547 [ERROR] WSREP: gcs/src/gcs\_group.cpp:int gcs\_group\_handle\_join\_msg(gcs\_group\_t\*, const gcs\_recv\_msg\_t\*)():731: Will never receive state. Need to abort.  
2016-05-19 23:24:23 2547 [Note] WSREP: gcomm: terminating thread  
2016-05-19 23:24:23 2547 [Note] WSREP: gcomm: joining thread  
2016-05-19 23:24:23 2547 [Note] WSREP: gcomm: closing backend  
2016-05-19 23:24:23 2547 [Note] WSREP: view(view\_id(NON\_PRIM,529594ab,51) memb {  
529594ab,0  
} joined {  
} left {  
} partitioned {  
56560110,0  
6e832bd2,0  
})  
2016-05-19 23:24:23 2547 [Note] WSREP: view((empty))  
2016-05-19 23:24:23 2547 [Note] WSREP: gcomm: closed  
2016-05-19 23:24:23 2547 [Note] WSREP: /usr/sbin/mysqld: Terminated.  
160519 23:24:23 mysqld\_safe mysqld from pid file /var/lib/mysql/ip-10-0-1-225.ap-southeast-1.compute.internal.pid ended

---

<div class="post-metadata">

**Author:** ![pavithra](https://avatars.discourse-cdn.com/v4/letter/p/e9a140/32.png) [@pavithra](https://forums.percona.com/u/pavithra)\
**Post date:** [May 21, 2016, 1:29am UTC](https://forums.percona.com/t/node-crashed-and-i-am-unable-to-recover/4892/3 "2016-05-21T01:29:35Z")

</div>

galera state on node2 and node3 are as below :

# GALERA saved state

version: 2.1  
uuid: e8e2feed-1753-11e6-bfe3-972d646081e6  
seqno: -1  
cert\_index:

# GALERA saved state

version: 2.1  
uuid: e8e2feed-1753-11e6-bfe3-972d646081e6  
seqno: -1  
cert\_index:  
~  
~

show global status like ‘wsrep\_local\_cached\_downto’;  
±--------------------------±--------+  
| Variable\_name | Value |  
±--------------------------±--------+  
| wsrep\_local\_cached\_downto | 2523303 |  
±--------------------------±--------+

show global status like ‘wsrep\_local\_cached\_downto’;  
±--------------------------±--------+  
| Variable\_name | Value |  
±--------------------------±--------+  
| wsrep\_local\_cached\_downto | 2523303 |  
±--------------------------±--------+
