# Strange node hang

**URL:** <https://forums.percona.com/t/strange-node-hang/2940>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [September 16, 2013, 1:42am UTC](https://forums.percona.com/t/strange-node-hang/2940 "2013-09-16T01:42:32Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![marko\_s](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@marko\_s](https://forums.percona.com/u/marko_s)\
**Post date:** [September 16, 2013, 1:42am UTC](https://forums.percona.com/t/strange-node-hang/2940/1 "2013-09-16T01:42:32Z")

</div>

Hello,

We have a following setup:

- 4 Percona cluster nodes running 5.5.31-23.7.5.438. build x64
- 1 master r/w node with others standing as backup over HAProxy

The issue happens occasionally at random (sometimes two days in a row, sometimes after a few weeks).  
The active node just stops processing queries and process list grows. There’s absolutely nothing in the error log (we have warnings enabled as well). By nothing I mean  
there’s no usual warnings in log in that period as well (as in usual I mean “update was ineffective” sort of thing).

After that we issue a node restart and then it starts up normally. I have zero log trace that I can analyze (at least those I know about, such as mysqld error log, syslog etc.)  
It’s not an memory/CPU issue, I’m monitoring server 0/24 for performance and there’s no anomaly in graphs.

What’s even worse, clustercheck script sees the node as fully synced and operational. I can connect to the instance, issue a query and it gets added to process list, but waits indefinitely as do other queries issued. So we have backup nodes ready to kick in but HAProxy never detects the outage.

The my.cnf buffer and memory parameters are the same as we had with standalone Percona Server that never hung and this server has even extra 8G of RAM over the standalone instance (40GB total).

Any ideas where to start the troubleshooting?

---

<div class="post-metadata">

**Author:** ![marko\_s](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@marko\_s](https://forums.percona.com/u/marko_s)\
**Post date:** [September 19, 2013, 4:29am UTC](https://forums.percona.com/t/strange-node-hang/2940/2 "2013-09-19T04:29:42Z")

</div>

I have upped gcache.size from 128M to 4GB. Also, have modified wsrep slave threads from 1 to 32.  
What worries me is that I’m getting ‘wsrep\_cert\_deps\_distance’ no larger than 1.

---

<div class="post-metadata">

**Author:** ![marko\_s](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@marko\_s](https://forums.percona.com/u/marko_s)\
**Post date:** [September 20, 2013, 2:27am UTC](https://forums.percona.com/t/strange-node-hang/2940/3 "2013-09-20T02:27:17Z")

</div>

I have noticed (during the yesterday’s hangup) that the oldest query among those piling up was related with “REPLACE INTO” on a myisam table (we have two altogether).  
The replication queue was empty (0) and the cluster rep was not stalled.

Could this be the issue? What I’ve read mostly is that myisam replication is unreliable (regarding consistency) but nothing on the issue of hanging a node.  
I could switch off myisam replication in my.cnf but am unwilling to do so until I’m certain that this is causing lockups.

---

<div class="post-metadata">

**Author:** ![przemek](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/przemek/32/3_2.png) [@przemek](https://forums.percona.com/u/przemek)\
**Post date:** [September 24, 2013, 8:40am UTC](https://forums.percona.com/t/strange-node-hang/2940/4 "2013-09-24T08:40:31Z")

</div>

Can you paste a sample show processlist part from a moment a node is hung? What are the queries waiting for? Also, I wonder if the same issue happens after converting those tables to InnoDB?

---

<div class="post-metadata">

**Author:** ![marko\_s](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@marko\_s](https://forums.percona.com/u/marko_s)\
**Post date:** [September 25, 2013, 1:17am UTC](https://forums.percona.com/t/strange-node-hang/2940/5 "2013-09-25T01:17:52Z")

</div>

I will post the full processlist as soon as it happens again. We have upgraded nodes to 5.33, no hangups yet.

I don’t know what the queries are waiting for, they seem to be in various states of execution. There is also a certain number of ‘wsrep in pre-commit stage’ processes.  
You can execute show status/variables etc, but no DB query ever gets executed.

As for InnoDB conversion, we have to keep one table on MyISAM because of the full-text search. We could turn off MyISAM replication off on the cluster but I want to make sure it is the root cause of the hangups.

---

<div class="post-metadata">

**Author:** ![marko\_s](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@marko\_s](https://forums.percona.com/u/marko_s)\
**Post date:** [October 5, 2013, 12:33am UTC](https://forums.percona.com/t/strange-node-hang/2940/6 "2013-10-05T00:33:25Z")

</div>

We had two hangs yesterday. I haven’t been able to get a processlist since I haven’t been at PC (did restart via webmin script over smartphone), but I did find this on every node except master node (others are backup/standby nodes):

WSREP: (7569e389-2698-11e3-8693-139228b987aa, ‘tcp://0.0.  
0.0:4567’) address ‘tcp://xx.xx.xx.xx:4567’ pointing to uuid 7569e389-2698-11e3-  
8693-139228b987aa is blacklisted, skipping  
131004 10:48:57 [Note] WSREP: (7569e389-2698-11e3-8693-139228b987aa, ‘tcp://0.0.  
0.0:4567’) address ‘tcp://xx.xx.xx.xx:4567’ pointing to uuid 7569e389-2698-11e3-  
8693-139228b987aa is blacklisted, skipping  
131004 10:48:57 [Note] WSREP: (7569e389-2698-11e3-8693-139228b987aa, ‘tcp://0.0.  
0.0:4567’) address ‘tcp://xx.xx.xx.xx:4567’ pointing to uuid 7569e389-2698-11e3-  
8693-139228b987aa is blacklisted, skipping  
131004 10:48:57 [Note] WSREP: (7569e389-2698-11e3-8693-139228b987aa, ‘tcp://0.0.  
0.0:4567’) address ‘tcp://xx.xx.xx.xx:4567’ pointing to uuid 7569e389-2698-11e3-  
8693-139228b987aa is blacklisted, skipping  
131004 10:48:57 [Note] WSREP: (7569e389-2698-11e3-8693-139228b987aa, ‘tcp://0.0.  
0.0:4567’) address ‘tcp://xx.xx.xx.xx:4567’ pointing to uuid 7569e389-2698-11e3-  
8693-139228b987aa is blacklisted, skipping  
131004 10:48:57 [Note] WSREP: (7569e389-2698-11e3-8693-139228b987aa, ‘tcp://0.0.  
0.0:4567’) address ‘tcp://xx.xx.xx.xx:4567’ pointing to uuid 7569e389-2698-11e3-  
8693-139228b987aa is blacklisted, skipping  
131004 10:48:57 [Note] WSREP: (7569e389-2698-11e3-8693-139228b987aa, ‘tcp://0.0.  
0.0:4567’) address ‘tcp://xx.xx.xx.xx:4567’ pointing to uuid 7569e389-2698-11e3-  
8693-139228b987aa is blacklisted, skipping

The master node has absolutely nothing relevant logged until restart. I have the following entries on both hangs. What do these messages indicate? xx.xx.xx.xx is the IP of the master node.

---

<div class="post-metadata">

**Author:** ![marko\_s](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@marko\_s](https://forums.percona.com/u/marko_s)\
**Post date:** [October 10, 2013, 3:42am UTC](https://forums.percona.com/t/strange-node-hang/2940/7 "2013-10-10T03:42:02Z")

</div>

Hello,

You asked me for full processlist when cluster hangs…can I PM it to you? Don’t want to obfuscate the queries and I rather wouldn’t see DB queries shown on Google search. 😃  
We have some attacks now and then.

Thanks in advance.

Regards,

Marko

---

<div class="post-metadata">

**Author:** ![przemek](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/przemek/32/3_2.png) [@przemek](https://forums.percona.com/u/przemek)\
**Post date:** [October 26, 2013, 3:58am UTC](https://forums.percona.com/t/strange-node-hang/2940/8 "2013-10-26T03:58:56Z")

</div>

Sorry for long delay, I forgot about this thread.  
This message:

```auto
WSREP: (xxx, 'tcp://0.0.0.0:4567') address 'tcp://xx.xx.xx.xx:4567' pointing to uuid xxx is blacklisted, skipping

```

is usually seen on a nodes trying to connect to the cluster, but no working node is in primary state. So looks like the master node went into non-Primary state after this hangup.  
To allow them to re-join, you need to force the primary state on the surviving node like this:

```auto
SET GLOBAL wsrep_provider_options="pc.bootstrap=1";

```

and then starting mysql on the other nodes should succeed.

Yes, you can send me the processlist via priv.

---

<div class="post-metadata">

**Author:** ![marko\_s](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@marko\_s](https://forums.percona.com/u/marko_s)\
**Post date:** [October 28, 2013, 2:06am UTC](https://forums.percona.com/t/strange-node-hang/2940/9 "2013-10-28T02:06:13Z")

</div>

It seems you cannot receive private messages 🙂

At the time of the hangup, the primary node (read/write, others a re only for backup) becomes non responsive and I’m seeing “blacklisted” messages on other nodes while the primary node has nothing in error log and never recovers until restart.

When I restart the primary node only, the rest of the cluster quickly syncs with the primary.

This is the wsrep status dump at the time of the hangup (please note that you can connect to the primary node and exec variable and status queries, but no query on actual tables succeeds, only waits indefinitely until number of connections is saturated):

Variable\_name Value  
wsrep\_local\_state\_uuid 4c3aae36-ff25-11e2-b3f2-f2c8c18d67ea  
wsrep\_protocol\_version 4  
wsrep\_last\_committed 201895717  
wsrep\_replicated 17129710  
wsrep\_replicated\_bytes 17889140306  
wsrep\_received 35425  
wsrep\_received\_bytes 982833  
wsrep\_local\_commits 17127813  
wsrep\_local\_cert\_failures 0  
wsrep\_local\_bf\_aborts 0  
wsrep\_local\_replays 0  
wsrep\_local\_send\_queue 0  
wsrep\_local\_send\_queue\_avg 0.000000  
wsrep\_local\_recv\_queue 34  
wsrep\_local\_recv\_queue\_avg 0.000000  
wsrep\_flow\_control\_paused 1.000000  
wsrep\_flow\_control\_sent 0  
wsrep\_flow\_control\_recv 0  
wsrep\_cert\_deps\_distance 193.730000  
wsrep\_apply\_oooe 0.000000  
wsrep\_apply\_oool 0.000000  
wsrep\_apply\_window 0.000000  
wsrep\_commit\_oooe 0.000000  
wsrep\_commit\_oool 0.000000  
wsrep\_commit\_window 0.000000  
wsrep\_local\_state 4  
wsrep\_local\_state\_comment Synced  
wsrep\_cert\_index\_size 368  
wsrep\_causal\_reads 0  
wsrep\_incoming\_addresses 10.42.71.90:3306,10.42.71.69:3306,10.42.71.68:3306,10.42.71.91:3306  
wsrep\_cluster\_conf\_id 121  
wsrep\_cluster\_size 4  
wsrep\_cluster\_state\_uuid 4c3aae36-ff25-11e2-b3f2-f2c8c18d67ea  
wsrep\_cluster\_status Primary  
wsrep\_connected ON  
wsrep\_local\_index 2  
wsrep\_provider\_name Galera  
wsrep\_provider\_vendor Codership Oy \<info@codership.com\>  
wsrep\_provider\_version 2.7(r157)  
wsrep\_ready ON
