# two-node cluster hangs when primary node is gone (ignore-sb=true)

**URL:** <https://forums.percona.com/t/two-node-cluster-hangs-when-primary-node-is-gone-ignore-sb-true/3112>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [November 21, 2013, 5:36am UTC](https://forums.percona.com/t/two-node-cluster-hangs-when-primary-node-is-gone-ignore-sb-true/3112 "2013-11-21T05:36:55Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![ClusterJohn](https://avatars.discourse-cdn.com/v4/letter/c/c5a1d2/32.png) [@ClusterJohn](https://forums.percona.com/u/ClusterJohn)\
**Post date:** [November 21, 2013, 5:36am UTC](https://forums.percona.com/t/two-node-cluster-hangs-when-primary-node-is-gone-ignore-sb-true/3112/1 "2013-11-21T05:36:55Z")

</div>

Hi,  
I have a two node Percona XtraDB Cluster (v5.5.33) with Galera setup. Galera has been configured to ignore split-brain.

When I perform failover tests for these nodes then I see strange behaviour which I cant get a grip on.

The following is the scenario:  
Node2 was the cluster creator. It has wsrep\_cluster\_address of “gcomm://”, Node1 has address “gcomm://10.0.100.2” when I start this scenario.  
Node1 (10.0.100.1) and node2 (10.0.100.2) are both running and only node 1 is receiving data.  
When I reboot node1 then node2 detects this and it happily receives data. No problem here. Node1 comes back up and performs IST to rejoin the cluster.  
All still well.  
When the IST of node1 is finished and the node is ready then after several minutes I stop node2. As soon as node2 is busy with stopping Percona then node1 hangs all transactions.  
Doing a “show status like ‘wsrep%’;” shows me that node 1 ‘believes’ its still part of the cluster and does not seem to detect that the 2nd node is gone.

I’m using all innoDB tables and have a high-load on the server. Several TBs of data with 60GB configured innodb buffer pool size.

I also tried to do a "SET GLOBAL wsrep\_cluster\_address=‘gcomm://’; " to force node2 to be the cluster creator. But alas, it does not solve the issue described above.

Why is the node hanging? and more importantly how can I fix this?

Many Thanks!

My my.cnf (config of node1, node2 is the same apart from ofcourse IP-addresses) looks like :

# GENERAL

user = mysql  
default-storage-engine = innodb  
socket = /data/mysql/mysql.sock  
pid-file = /data/mysql/mysql.pid

slow-query-log = ON  
log-queries-not-using-indexes = ON  
innodb\_print\_all\_deadlocks = ON

max\_allowed\_packet = 120M  
max\_connect\_errors = 2000000000000  
skip-name-resolve

sysdate-is-now = 1  
innodb = FORCE  
innodb-strict-mode = 1

datadir = /data/mysql  
tmpdir = /data/mysql-tmp

log-bin = /data/mysql/mysql-bin  
expire-logs-days = 5  
sync-binlog = 1

log-slave-updates = 1  
relay-log = /data/mysql/relay-bin  
slave-net-timeout = 60  
sync-master-info = 1  
sync-relay-log = 1  
sync-relay-log-info = 1

tmp-table-size = 32M  
max-heap-table-size = 32M  
query-cache-type = 0  
query-cache-size = 0

max-connections = 1000  
thread-cache-size = 50  
open-files-limit = 65535  
table-definition-cache = 1024  
table-open-cache = 1000

innodb-flush-method = O\_DIRECT  
innodb-log-files-in-group = 2  
innodb-log-file-size = 512M  
innodb-flush-log-at-trx-commit = 1  
innodb-file-per-table = 1

innodb-buffer-pool-size = 60G

server-id = 1  
binlog\_format=ROW

innodb\_autoinc\_lock\_mode=2  
innodb\_locks\_unsafe\_for\_binlog=1  
bind-address=0.0.0.0

wsrep\_provider=“/usr/lib/libgalera\_smm.so”  
wsrep\_provider\_options=“pc.ignore\_sb = yes; evs.keepalive\_period = PT1S; evs.inactive\_check\_period = PT1S; evs.suspect\_timeout = PT5S; evs.inactive\_timeout = PT10S; evs.install\_timeout = PT10S; gcache.size=32G”

wsrep\_cluster\_name=“percona\_cluster”  
wsrep\_cluster\_address=gcomm://10.0.100.2

wsrep\_node\_name=node1  
wsrep\_node\_address=10.0.100.1

wsrep\_slave\_threads=16

wsrep\_certify\_nonPK=1  
wsrep\_max\_ws\_rows=131072  
wsrep\_max\_ws\_size=1073741824  
wsrep\_debug=0  
wsrep\_convert\_LOCK\_to\_trx=0  
wsrep\_retry\_autocommit=1  
wsrep\_auto\_increment\_control=1  
wsrep\_drupal\_282555\_workaround=0

wsrep\_causal\_reads=0  
wsrep\_notify\_cmd=

wsrep\_sst\_method=xtrabackup  
wsrep\_sst\_auth=mysql\_sst:\*\*\*\*\*\*\*\*\*

# Desired SST donor name.

#wsrep\_sst\_donor=

# Reject client queries when donating SST (false)

#wsrep\_sst\_donor\_rejects\_queries=0

# Protocol version to use

# wsrep\_protocol\_version=

---

<div class="post-metadata">

**Author:** ![madhusudan](https://avatars.discourse-cdn.com/v4/letter/m/76d3ee/32.png) [@madhusudan](https://forums.percona.com/u/madhusudan)\
**Post date:** [November 21, 2013, 7:39am UTC](https://forums.percona.com/t/two-node-cluster-hangs-when-primary-node-is-gone-ignore-sb-true/3112/2 "2013-11-21T07:39:56Z")

</div>

If you had bootstrapped node2 then try starting again first node2(let it start completely) and then node1,  
how did you come to conclusion that it was hanging!?, you say its high load, and with innodb you can expect slowness.  
What the mysql error log says…?
