# my percona xtraDB cluster suddently dead and how to fix it

**URL:** <https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952>\
**Category:** Percona XtraDB Cluster 8.x\
**Tags:** community, mysql, percona\
**Created:** [August 26, 2020, 8:32am UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952 "2020-08-26T08:32:06Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 26, 2020, 8:32am UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/1 "2020-08-26T08:32:06Z")

</div>

hi,  
my PXC is down and this is the error message I got from error log, please see attached.  
  
any reason why and how to fix it ?

[PXC errorr message.docx](https://forums.percona.com/uploads/short-url/A5P1x90D2rxcQzlTlOtCGlzQE7g.docx) (19.2 KB)

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 26, 2020, 1:23pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/2 "2020-08-26T13:23:27Z")

</div>

```auto
2020-08-26T21:10:55.088728+08:00 0 [Note] [MY-000000] [Galera] (995a4f35, 'tcp://0.0.0.0:4567') connection to peer 4855f0e4 with addr tcp://&lt;IP address&gt;:4567 timed out, no messages seen in PT3S (gmcast.peer_timeout)<br><br><br><br><br>2020-08-26T21:10:55.089033+08:00 0 [Note] [MY-000000] [Galera] (995a4f35, 'tcp://0.0.0.0:4567') connection to peer f2266321 with addr tcp://&lt;IP address&gt;:4567 timed out, no messages seen in PT3S (gmcast.peer_timeout)

```

Your nodes lost connections to eachother. Network outage. If all nodes are offline, you need to stop them all and then re-bootstrap the whole cluster.

---

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 26, 2020, 9:21pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/3 "2020-08-26T21:21:58Z")

</div>

but I ping each other and it’s pingable!  
I read that line too but it seems not the case ! by your statement, you seems telling me that standalone boot by systemctl start mysql will be ok ?  
  
and the error message by saying nodes communicate with each other using port&nbsp;4567 ? mysql is 3306 right ?  
and only this cluster communicate using port 4567 … ? my other cluster do not have this kind of problem .

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 26, 2020, 9:49pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/4 "2020-08-26T21:49:35Z")

</div>

If you have other PXC clusters on this same network and none of those clusters are having issues, then I’d say there is an issue with the nodes of this cluster causing network timeout issues. Look at your metrics/monitoring for CPU saturation, disk IO saturation, etc. There could be something else causing the node to be unable to process network packets and thus miss heartbeats and be ejected from the cluster.

---

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 26, 2020, 9:50pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/5 "2020-08-26T21:50:44Z")

</div>

“be unable to process network packets and thus miss heartbeats and be ejected from the cluster.”  
  
so you are sure that MUST BE network problem ! and why port 4567 ? I never use it for mysql

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 26, 2020, 9:56pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/6 "2020-08-26T21:56:07Z")

</div>

No, I am not sure it is directly network related. The logs say “connection timed out” which means network issues. However, many other things can cause “network issues.”  
  
4444 is used by PXC/Galera for SST/IST. 4567 is used by PXC/Galera internal node-node communication. 3306 is used by MySQL.

---

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 26, 2020, 10:18pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/7 "2020-08-26T22:18:08Z")

</div>

“4444 is used by PXC/Galera for SST/IST. 4567 is used by PXC/Galera internal node-node communication. 3306 is used by MySQL.”  
  
good ! and can I just telnet \<nodes IP\> 4444 and&nbsp;telnet \<nodes IP\> 4567 to verify it ?&nbsp; I am thinking firewall block it.

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 26, 2020, 10:23pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/8 "2020-08-26T22:23:48Z")

</div>

Probably, yes, you can telnet to 4567 to see if you get a response from another node. 4444 only responds while an SST/IST is in progress.

---

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 26, 2020, 10:27pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/9 "2020-08-26T22:27:17Z")

</div>

"&nbsp;SST/IST is in progress"  
  
replication ?

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 26, 2020, 10:30pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/10 "2020-08-26T22:30:39Z")

</div>

No. You need to go learn PXC Basics 101 if you don’t know what IST/SST are. Fundamental to PXC/Galera operations.

---

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 26, 2020, 10:31pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/11 "2020-08-26T22:31:44Z")

</div>

yeah!I might forget it, it is for the start of replication and the stream replication, right?  
  
at this moment want to troubleshoot the cluster first. sorry

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 26, 2020, 10:32pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/12 "2020-08-26T22:32:55Z")

</div>

No. IST/SST have nothing to do with replication. SST is for when new nodes join. IST is for when nodes leave and then come back.

---

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 26, 2020, 10:36pm UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/13 "2020-08-26T22:36:24Z")

</div>

"&nbsp;SST is for when new nodes join. IST is for when nodes leave and then come back.＂  
good and tks. will have a look later.

---

<div class="post-metadata">

**Author:** ![DBA100](https://avatars.discourse-cdn.com/v4/letter/d/b3f665/32.png) [@DBA100](https://forums.percona.com/u/DBA100)\
**Post date:** [August 27, 2020, 1:06am UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/14 "2020-08-27T01:06:42Z")

</div>

hi,  
probably some network problem and today very funny that, the cluster up without any problem anymore without I restart /bootstrap again !&nbsp;  
amazing …  
one quetion is, if next time it happens again, just because of some network problem, the linkage between nodes&nbsp; broken again by some reason but it recover later, will the cluster also reform automatically ?&nbsp;  
  
and I found if situation like this happen again, really need to bootstrap again even it recover itself automatically.

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 27, 2020, 8:12am UTC](https://forums.percona.com/t/my-percona-xtradb-cluster-suddently-dead-and-how-to-fix-it/7952/15 "2020-08-27T08:12:52Z")

</div>

If all nodes are down, you must always bootstrap the first node.
