# Problem when node goes down

**URL:** <https://forums.percona.com/t/problem-when-node-goes-down/16312>\
**Category:** Percona XtraDB Cluster 5.x\
**Tags:** mysql, percona\
**Created:** [June 27, 2022, 9:11am UTC](https://forums.percona.com/t/problem-when-node-goes-down/16312 "2022-06-27T09:11:09Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![hieu\_nguyen](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/hieu_nguyen/32/6616_2.png) [@hieu\_nguyen](https://forums.percona.com/u/hieu_nguyen)\
**Post date:** [June 27, 2022, 9:11am UTC](https://forums.percona.com/t/problem-when-node-goes-down/16312/1 "2022-06-27T09:11:09Z")

</div>

i was running 3 nodes , i shut down my node 1 , and when i restarted node 1 was down i shut down node2 and node 3 then turn on boostrap with node 3 , start again with node2 , but when starting node 1 again an error occurs

> 2022-06-27 15:54:19 12660 [Note] WSREP: Member 1.0 (localhost.localdomain) requested state transfer from ‘_any_’. Selected 0.0 (localhost.localdomain)(SYNCED) as donor.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Shifting PRIMARY → JOINER (TO: 4)  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Requesting state transfer: success, donor: 0  
> 2022-06-27 15:54:19 12660 [Warning] WSREP: 0.0 (localhost.localdomain): State transfer to 1.0 (localhost.localdomain) failed: -255 (Unknown error 255)  
> 2022-06-27 15:54:19 12660 [ERROR] WSREP: gcs/src/gcs\_group.cpp:gcs\_group\_handle\_join\_msg():736: Will never receive state. Need to abort.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: gcomm: terminating thread  
> 2022-06-27 15:54:19 12660 [Note] WSREP: gcomm: joining thread  
> 2022-06-27 15:54:19 12660 [Note] WSREP: gcomm: closing backend  
> 2022-06-27 15:54:20 12660 [Note] WSREP: view(view\_id(NON\_PRIM,aab8cec1,2) memb {  
> ba9c56fc,0  
> } joined {  
> } left {  
> } partitioned {  
> aab8cec1,0  
> })  
> 2022-06-27 15:54:20 12660 [Note] WSREP: view((empty))  
> 2022-06-27 15:54:20 12660 [Note] WSREP: gcomm: closed  
> 2022-06-27 15:54:20 12660 [Note] WSREP: /usr/sbin/mysqld: Terminated.  
> 220627 15:54:20 mysqld\_safe mysqld from pid file /var/lib/mysql/localhost.localdomain.pid ended  
> WSREP\_SST: [ERROR] Parent mysqld process (PID:12660) terminated unexpectedly. (20220627 15:54:21.423)  
> WSREP\_SST: [INFO] Joiner cleanup. rsync PID: 12701 (20220627 15:54:21.425)  
> WSREP\_SST: [INFO] Joiner cleanup done. (20220627 15:54:21.929)

---

<div class="post-metadata">

**Author:** ![hieu\_nguyen](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/hieu_nguyen/32/6616_2.png) [@hieu\_nguyen](https://forums.percona.com/u/hieu_nguyen)\
**Post date:** [June 27, 2022, 10:28am UTC](https://forums.percona.com/t/problem-when-node-goes-down/16312/2 "2022-06-27T10:28:07Z")

</div>

I want to provide more log

> 2022-06-27 15:54:18 0 [Warning] TIMESTAMP with implicit DEFAULT value is deprecated. Please use --explicit\_defaults\_for\_timestamp server option (see documentation for more details).  
> 2022-06-27 15:54:18 0 [Note] /usr/sbin/mysqld (mysqld 5.6.30-76.3-56) starting as process 12660 …  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Read nil XID from storage engines, skipping position init  
> 2022-06-27 15:54:18 12660 [Note] WSREP: wsrep\_load(): loading provider library ‘/usr/lib64/libgalera\_smm.so’  
> 2022-06-27 15:54:18 12660 [Note] WSREP: wsrep\_load(): Galera 3.16(r5c765eb) by Codership Oy [info@codership.com](mailto:info@codership.com) loaded successfully.  
> 2022-06-27 15:54:18 12660 [Note] WSREP: CRC-32C: using hardware acceleration.  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Found saved state: 00000000-0000-0000-0000-000000000000:-1  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Passing config to GCS: base\_dir = /var/lib/mysql/; base\_host = 192.168.254.128; base\_port = 4567; cert.log\_conflicts = no; debug = no; evs.auto\_evict = 0; evs.delay\_margin = PT1S; evs.delayed\_keep\_period = PT30S; evs.inactive\_check\_period = PT0.5S; evs.inactive\_timeout = PT15S; evs.join\_retrans\_period = PT1S; evs.max\_install\_timeouts = 3; evs.send\_window = 4; evs.stats\_report\_period = PT1M; evs.suspect\_timeout = PT5S; evs.user\_send\_window = 2; evs.view\_forget\_timeout = PT24H; gcache.dir = /var/lib/mysql/; gcache.keep\_pages\_count = 0; gcache.keep\_pages\_size = 0; gcache.mem\_size = 0; gcache.name = /var/lib/mysql//galera.cache; gcache.page\_size = 128M; gcache.size = 128M; gcomm.thread\_prio = ; gcs.fc\_debug = 0; gcs.fc\_factor = 1.0; gcs.fc\_limit = 16; gcs.fc\_master\_slave = no; gcs.max\_packet\_size = 64500; gcs.max\_throttle = 0.25; gcs.recv\_q\_hard\_limit = 9223372036854775807; gcs.recv\_q\_soft\_limit = 0.25; gcs.sync\_donor = no; gmcast.segment = 0; gmcast.version = 0; pc.announce\_timeout = PT3S; pc.checksum =  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Service thread queue flushed.  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Assign initial position for certification: -1, protocol version: -1  
> 2022-06-27 15:54:18 12660 [Note] WSREP: wsrep\_sst\_grab()  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Start replication  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Setting initial position to 00000000-0000-0000-0000-000000000000:-1  
> 2022-06-27 15:54:18 12660 [Note] WSREP: protonet asio version 0  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Using CRC-32C for message checksums.  
> 2022-06-27 15:54:18 12660 [Note] WSREP: backend: asio  
> 2022-06-27 15:54:18 12660 [Note] WSREP: gcomm thread scheduling priority set to other:0  
> 2022-06-27 15:54:18 12660 [Warning] WSREP: access file(/var/lib/mysql//gvwstate.dat) failed(No such file or directory)  
> 2022-06-27 15:54:18 12660 [Note] WSREP: restore pc from disk failed  
> 2022-06-27 15:54:18 12660 [Note] WSREP: GMCast version 0  
> 2022-06-27 15:54:18 12660 [Note] WSREP: (ba9c56fc, ‘tcp://0.0.0.0:4567’) listening at tcp://0.0.0.0:4567  
> 2022-06-27 15:54:18 12660 [Note] WSREP: (ba9c56fc, ‘tcp://0.0.0.0:4567’) multicast: , ttl: 1  
> 2022-06-27 15:54:18 12660 [Note] WSREP: EVS version 0  
> 2022-06-27 15:54:18 12660 [Note] WSREP: gcomm: connecting to group ‘my\_centos\_cluster’, peer ‘192.168.254.128:,192.168.254.102:,192.168.254.150:’  
> 2022-06-27 15:54:18 12660 [Warning] WSREP: (ba9c56fc, ‘tcp://0.0.0.0:4567’) address ‘tcp://192.168.254.128:4567’ points to own listening address, blacklisting  
> 2022-06-27 15:54:18 12660 [Note] WSREP: (ba9c56fc, ‘tcp://0.0.0.0:4567’) turning message relay requesting on, nonlive peers:  
> 2022-06-27 15:54:18 12660 [Note] WSREP: declaring aab8cec1 at tcp://192.168.254.150:4567 stable  
> 2022-06-27 15:54:18 12660 [Note] WSREP: Node aab8cec1 state prim  
> 2022-06-27 15:54:18 12660 [Note] WSREP: view(view\_id(PRIM,aab8cec1,2) memb {  
> aab8cec1,0  
> ba9c56fc,0  
> } joined {  
> } left {  
> } partitioned {  
> })  
> 2022-06-27 15:54:18 12660 [Note] WSREP: save pc into disk  
> 2022-06-27 15:54:18 12660 [Note] WSREP: discarding pending addr without UUID: tcp://192.168.254.102:4567  
> 2022-06-27 15:54:19 12660 [Note] WSREP: gcomm: connected  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Changing maximum packet size to 64500, resulting msg size: 32636  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Shifting CLOSED → OPEN (TO: 0)  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Opened channel ‘my\_centos\_cluster’  
> 2022-06-27 15:54:19 12660 [Note] WSREP: New COMPONENT: primary = yes, bootstrap = no, my\_idx = 1, memb\_num = 2  
> 2022-06-27 15:54:19 12660 [Note] WSREP: STATE EXCHANGE: Waiting for state UUID.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Waiting for SST to complete.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: STATE EXCHANGE: sent state msg: b9ac93c0-f5f6-11ec-b706-7ab2a73219b6  
> 2022-06-27 15:54:19 12660 [Note] WSREP: STATE EXCHANGE: got state msg: b9ac93c0-f5f6-11ec-b706-7ab2a73219b6 from 0 (localhost.localdomain)  
> 2022-06-27 15:54:19 12660 [Note] WSREP: STATE EXCHANGE: got state msg: b9ac93c0-f5f6-11ec-b706-7ab2a73219b6 from 1 (localhost.localdomain)  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Quorum results:  
> version = 4,  
> component = PRIMARY,  
> conf\_id = 1,  
> members = 1/2 (joined/total),  
> act\_id = 4,  
> last\_appl. = -1,  
> protocols = 0/7/3 (gcs/repl/appl),  
> group UUID = f6c8c6d1-e0b3-11ec-82be-cfbd0b61dbd0  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Flow-control interval: [23, 23]  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Shifting OPEN → PRIMARY (TO: 4)  
> 2022-06-27 15:54:19 12660 [Note] WSREP: State transfer required:  
> Group state: f6c8c6d1-e0b3-11ec-82be-cfbd0b61dbd0:4  
> Local state: 00000000-0000-0000-0000-000000000000:-1  
> 2022-06-27 15:54:19 12660 [Note] WSREP: New cluster view: global state: f6c8c6d1-e0b3-11ec-82be-cfbd0b61dbd0:4, view# 2: Primary, number of nodes: 2, my index: 1, protocol version 3  
> 2022-06-27 15:54:19 12660 [Warning] WSREP: Gap in state sequence. Need state transfer.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Running: 'wsrep\_sst\_rsync --role ‘joiner’ --address ‘192.168.254.128’ --datadir ‘/var/lib/mysql/’ --defaults-file ‘/etc/my.cnf’ --defaults-group-suffix ‘’ --parent ‘12660’ ‘’ ’  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Prepared SST request: rsync|192.168.254.128:4444/rsync\_sst  
> 2022-06-27 15:54:19 12660 [Note] WSREP: wsrep\_notify\_cmd is not defined, skipping notification.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: REPL Protocols: 7 (3, 2)  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Service thread queue flushed.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Assign initial position for certification: 4, protocol version: 3  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Service thread queue flushed.  
> 2022-06-27 15:54:19 12660 [Warning] WSREP: Failed to prepare for incremental state transfer: Local state UUID (00000000-0000-0000-0000-000000000000) does not match group state UUID (f6c8c6d1-e0b3-11ec-82be-cfbd0b61dbd0): 1 (Operation not permitted)  
> at galera/src/replicator\_str.cpp:prepare\_for\_IST():507. IST will be unavailable.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Member 1.0 (localhost.localdomain) requested state transfer from ‘_any_’. Selected 0.0 (localhost.localdomain)(SYNCED) as donor.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Shifting PRIMARY → JOINER (TO: 4)  
> 2022-06-27 15:54:19 12660 [Note] WSREP: Requesting state transfer: success, donor: 0  
> 2022-06-27 15:54:19 12660 [Warning] WSREP: 0.0 (localhost.localdomain): State transfer to 1.0 (localhost.localdomain) failed: -255 (Unknown error 255)  
> 2022-06-27 15:54:19 12660 [ERROR] WSREP: gcs/src/gcs\_group.cpp:gcs\_group\_handle\_join\_msg():736: Will never receive state. Need to abort.  
> 2022-06-27 15:54:19 12660 [Note] WSREP: gcomm: terminating thread  
> 2022-06-27 15:54:19 12660 [Note] WSREP: gcomm: joining thread  
> 2022-06-27 15:54:19 12660 [Note] WSREP: gcomm: closing backend  
> 2022-06-27 15:54:20 12660 [Note] WSREP: view(view\_id(NON\_PRIM,aab8cec1,2) memb {  
> ba9c56fc,0  
> } joined {  
> } left {  
> } partitioned {  
> aab8cec1,0  
> })  
> 2022-06-27 15:54:20 12660 [Note] WSREP: view((empty))  
> 2022-06-27 15:54:20 12660 [Note] WSREP: gcomm: closed  
> 2022-06-27 15:54:20 12660 [Note] WSREP: /usr/sbin/mysqld: Terminated.  
> 220627 15:54:20 mysqld\_safe mysqld from pid file /var/lib/mysql/localhost.localdomain.pid ended  
> WSREP\_SST: [ERROR] Parent mysqld process (PID:12660) terminated unexpectedly. (20220627 15:54:21.423)  
> WSREP\_SST: [INFO] Joiner cleanup. rsync PID: 12701 (20220627 15:54:21.425)  
> WSREP\_SST: [INFO] Joiner cleanup done. (20220627 15:54:21.929)

---

<div class="post-metadata">

**Author:** ![CTutte](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/ctutte/32/1341_2.png) [@CTutte](https://forums.percona.com/u/CTutte)\
**Post date:** [June 28, 2022, 4:35pm UTC](https://forums.percona.com/t/problem-when-node-goes-down/16312/3 "2022-06-28T16:35:44Z")

</div>

Hi hieu\_nguyen,

What versions are all the nodes ?  
Note that node failing to join is 5.6.30 which is ~6 years old and already reached end of life.  
Possible reasons for failing include 1) a bug, 2) cluster being bootstrapped by a node with version 5.7 thus not possible for 5.6 to join, 3) some of the communication ports are blocked

I suggest you double check versions and ports, and as a last resort remove grastate.dat file from 5.6 data dir. Doing the latter will force a clean SST from the node

Also keep in mind that all node versions should be same major version (or cluster bootstraped by the oldest version of the nodes). Ideally all nodes should have same version to avoid conflicts

Regards
