# 2 Nodes are going out of sync in 3 Nodes PXC set up at the time of re syncing

**URL:** <https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [December 27, 2013, 6:45am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165 "2013-12-27T06:45:47Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Poorna\_PC](https://avatars.discourse-cdn.com/v4/letter/p/e19b73/32.png) [@Poorna\_PC](https://forums.percona.com/u/Poorna_PC)\
**Post date:** [December 27, 2013, 6:45am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/1 "2013-12-27T06:45:47Z")

</div>

2 Nodes are going out of sync in 3 Nodes Percona XtraDB cluster set up at the time of re syncing:

## Requesting you Guys in helping to re sync all 3 nodes. Thanks you..

Used RPMS:  
Percona-XtraDB-Cluster-shared-5.5.27-23.6.356.rhel6.x86\_64  
Percona-XtraDB-Cluster-server-5.5.27-23.6.356.rhel6.x86\_64  
percona-release-0.0-1.x86\_64  
Percona-XtraDB-Cluster-client-5.5.27-23.6.356.rhel6.x86\_64  
Percona-XtraDB-Cluster-galera-2.0-1.114.rhel6.x86\_64  
percona-xtrabackup-2.0.3-470.rhel6.x86\_64

## OS: CentOS release 6.3 (Final) Environment: Virtual Systems.

Here is the mysql-error log from all 3 nodes:  
Node 2: which is up  
WSREP: FK key len exceeded 0 4294967295 3500  
131227 2:58:46 [ERROR] WSREP: FK key set failed: 11  
WSREP: FK key append failed

Node 3: is down  
131227 5:00:11 [Note] WSREP: sst\_donor\_thread signaled with 0  
131227 5:00:11 [Note] WSREP: Flushing tables for SST…  
131227 5:00:11 [Note] WSREP: Provider paused at cf67b4da-6ea7-11e3-0800-7176739bc3d8:261  
131227 5:00:11 [Note] WSREP: Tables flushed.  
InnoDB: Warning: a long semaphore wait:  
–Thread 139738020943616 has waited at trx0rseg.ic line 46 for 241.00 seconds the semaphore:  
X-lock (wait\_ex) on RW-latch at 0x7f177f07a6b8 ‘&block-\>lock’  
a writer (thread id 139738020943616) has reserved it in mode wait exclusive  
number of readers 1, waiters flag 0, lock\_word: ffffffffffffffff  
Last time read locked in file buf0flu.c line 1319  
Last time write locked in file /home/jenkins/workspace/percona-xtradb-cluster-rpms/label\_exp/centos6-64/target/BUILD/Percona-XtraDB-Cluster-5.5.27/Percona-XtraDB-Cluster-5.5.27/storage/innobase/include/trx0rseg.ic line 46  
InnoDB: ###### Starts InnoDB Monitor for 30 secs to print diagnostic info:

* * *

## SEMAPHORES

OS WAIT ARRAY INFO: reservation count 46, signal count 44  
–Thread 139738020943616 has waited at trx0rseg.ic line 46 for 271.00 seconds the semaphore:  
X-lock (wait\_ex) on RW-latch at 0x7f177f07a6b8 ‘&block-\>lock’  
a writer (thread id 139738020943616) has reserved it in mode wait exclusive  
number of readers 1, waiters flag 0, lock\_word: ffffffffffffffff  
Last time read locked in file buf0flu.c line 1319  
Last time write locked in file /home/jenkins/workspace/percona-xtradb-cluster-rpms/label\_exp/centos6-64/target/BUILD/Percona-XtraDB-Cluster-5.5.27/Percona-XtraDB-Cluster-5.5.27/storage/innobase/include/trx0rseg.ic line 46  
Mutex spin waits 38, rounds 925, OS waits 30  
RW-shared spins 15, rounds 432, OS waits 14  
RW-excl spins 1, rounds 60, OS waits 2  
Spin rounds per wait: 24.34 mutex, 28.80 RW-shared, 60.00 RW-excl

## RANSACTIONS

## Trx id counter A0E406071 Purge done for trx’s n:o \< A0E40606E undo n:o \< 0 History list length 618 LIST OF TRANSACTIONS FOR EACH SESSION: —TRANSACTION A0E40606E, not started MySQL thread id 3, OS thread handle 0x7f174a757700, query id 2974 committed 260 —TRANSACTION A0E406070, not started MySQL thread id 1, OS thread handle 0x7f1b16edb700, query id 2976 committed 261

# END OF INNODB MONITOR OUTPUT

InnoDB: ###### Diagnostic info printed to the standard error stream  
InnoDB: Warning: a long semaphore wait:  
–Thread 139738020943616 has waited at trx0rseg.ic line 46 for 303.00 seconds the semaphore:  
X-lock (wait\_ex) on RW-latch at 0x7f177f07a6b8 ‘&block-\>lock’  
a writer (thread id 139738020943616) has reserved it in mode wait exclusive  
number of readers 1, waiters flag 0, lock\_word: ffffffffffffffff  
Last time read locked in file buf0flu.c line 1319  
Last time write locked in file /home/jenkins/workspace/percona-xtradb-cluster-rpms/label\_exp/centos6-64/target/BUILD/Percona-XtraDB-Cluster-5.5.27/Percona-XtraDB-Cluster-5.5.27/storage/innoba  
se/include/trx0rseg.ic line 46  
InnoDB: ###### Starts InnoDB Monitor for 30 secs to print diagnostic info:  
InnoDB: Pending preads 0, pwrites 0

Node 1: down  
131227 4:49:46 [Note] WSREP: 1 (Node3): State transfer from 0 (Node1) complete.  
131227 4:49:46 [Note] WSREP: Member 1 (Node3) synced with group.  
05:00:03 UTC - mysqld got signal 11 ;  
This could be because you hit a bug. It is also possible that this binary  
or one of the libraries it was linked against is corrupt, improperly built,  
or misconfigured. This error can also be caused by malfunctioning hardware.  
We will try our best to scrape up some info that will hopefully help  
diagnose the problem, but since we have already crashed,  
something is definitely wrong and this may fail.  
Please help us make Percona Server better by reporting any  
bugs at [http://bugs.percona.com/](http://bugs.percona.com/)

131227 5:11:34 [Note] WSREP: New COMPONENT: primary = yes, bootstrap = no, my\_idx = 0, memb\_num = 2  
131227 5:11:34 [Note] WSREP: forgetting 49cd72df-6eb2-11e3-0800-3db8fd926ddb (tcp://XXX.XXX.XXX.53-Node3:4567)  
131227 5:11:34 [Note] WSREP: (bf5de37d-6eb3-11e3-0800-1b8b698cefc9, ‘tcp://0.0.0.0:4567’) turning message relay requesting off  
131227 5:11:34 [Note] WSREP: STATE\_EXCHANGE: sent state UUID: 5ac327af-6eb5-11e3-0800-8a7f196d2532  
131227 5:11:34 [Note] WSREP: STATE EXCHANGE: sent state msg: 5ac327af-6eb5-11e3-0800-8a7f196d2532  
131227 5:11:34 [Note] WSREP: STATE EXCHANGE: got state msg: 5ac327af-6eb5-11e3-0800-8a7f196d2532 from 0 (Node1)  
131227 5:11:34 [Note] WSREP: STATE EXCHANGE: got state msg: 5ac327af-6eb5-11e3-0800-8a7f196d2532 from 1 (Node2)  
131227 5:11:34 [Note] WSREP: Quorum results:  
version = 2,  
component = PRIMARY,  
conf\_id = 4,  
members = 1/2 (joined/total),  
act\_id = 864,  
last\_appl. = 835,  
protocols = 0/4/2 (gcs/repl/appl),  
group UUID = cf67b4da-6ea7-11e3-0800-7176739bc3d8  
131227 5:11:34 [Warning] WSREP: Donor 49cd72df-6eb2-11e3-0800-3db8fd926ddb is no longer in the group. State transfer cannot be completed, need to abort. Aborting…  
131227 5:11:34 [Note] WSREP: /usr/sbin/mysqld: Terminated.  
131227 05:11:34 mysqld\_safe mysqld from pid file /mnt/data//Node1.pid ended

131227 5:24:10 [Note] WSREP: Assign initial position for certification: 960, protocol version: 2  
131227 5:24:10 [Warning] WSREP: Failed to prepare for incremental state transfer: Local state UUID (00000000-0000-0000-0000-000000000000) does not match group state UUID (cf67b4da-6ea7-11e3-

## 131227 5:25:53 [Note] WSREP: Quorum results: version = 2, component = NON-PRIMARY, conf\_id = -1, members = 1/1 (joined/total), act\_id = -1, last\_appl. = -1, protocols = -1/-1/-1 (gcs/repl/appl), group UUID = 00000000-0000-0000-0000-000000000000 131227 5:25:53 [Note] WSREP: Flow-control interval: [8, 16] 131227 5:25:53 [Note] WSREP: Received NON-PRIMARY. 131227 5:25:53 [Note] WSREP: Shifting JOINER → OPEN (TO: 961) 131227 5:25:59 [Note] WSREP: cleaning up f9d65922-6eb6-11e3-0800-4de8ca27dd9e (tcp://XXX.XXX.XXX.52-Node2:4567)

---

<div class="post-metadata">

**Author:** ![Poorna\_PC](https://avatars.discourse-cdn.com/v4/letter/p/e19b73/32.png) [@Poorna\_PC](https://forums.percona.com/u/Poorna_PC)\
**Post date:** [December 27, 2013, 6:46am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/2 "2013-12-27T06:46:30Z")

</div>

* * *

Node 2 my.cnf: mysqld section: which is up  
[mysqld]

# GENERAL

user = mysql  
default\_storage\_engine = InnoDB

server\_id=1  
wsrep\_cluster\_address=gcomm://  
wsrep\_provider=/usr/lib64/libgalera\_smm.so  
wsrep\_slave\_threads=2  
wsrep\_cluster\_name= ecomm  
wsrep\_sst\_method=rsync  
wsrep\_node\_name=Node2  
wsrep\_sst\_receive\_address=XXX.XXX.XXX.52-Node2

# MyISAM

key\_buffer\_size = 32M  
myisam\_recover = FORCE,BACKUP

# SAFETY

max\_allowed\_packet = 64M  
max\_connect\_errors = 1000000  
skip\_name\_resolve  
sql\_mode = STRICT\_TRANS\_TABLES,ERROR\_FOR\_DIVISION\_BY\_ZERO,NO\_AUTO\_CREATE\_USER,NO\_AUTO\_VALUE\_  
ON\_ZERO,NO\_ENGINE\_SUBSTITUTION,NO\_ZERO\_DATE,NO\_ZERO\_IN\_DATE,ONLY\_FULL\_GROUP\_BY  
sysdate\_is\_now = 1  
innodb = FORCE  
innodb\_strict\_mode = 1

# DATA STORAGE

datadir = /mnt/data/

# BINARY LOGGING

log\_bin = /mnt/data/mysql-bin  
expire\_logs\_days = 14  
sync\_binlog = 1  
binlog\_format = ROW

# CACHES AND LIMITS

tmp\_table\_size = 128M  
max\_heap\_table\_size = 128M  
query\_cache\_type = 0  
query\_cache\_size = 8  
max\_connections = 2010  
thread\_cache\_size = 50  
open\_files\_limit = 65535  
table\_definition\_cache = 4096  
table\_open\_cache = 12000

# INNODB

innodb\_flush\_method = O\_DIRECT  
innodb\_log\_files\_in\_group = 2  
innodb\_log\_file\_size = 256M  
innodb\_flush\_log\_at\_trx\_commit = 1  
innodb\_file\_per\_table = 1  
innodb\_buffer\_pool\_size = 14G  
innodb\_locks\_unsafe\_for\_binlog = 1  
innodb\_autoinc\_lock\_mode = 2  
wait\_timeout = 1500  
interactive\_timeout = 1500

* * *

Node 1 my.cnf: mysqld section: which is down  
[mysqld]

# GENERAL

user = mysql  
default\_storage\_engine = InnoDB

server\_id=2  
wsrep\_cluster\_address=gcomm://XXX.XXX.XXX.52-Node2  
wsrep\_provider=/usr/lib64/libgalera\_smm.so  
wsrep\_slave\_threads=2  
wsrep\_cluster\_name= ecomm  
wsrep\_sst\_method=rsync  
wsrep\_node\_name=Node1  
wsrep\_sst\_receive\_address=XXX.XXX.XXX.51-Node1

# MyISAM

key\_buffer\_size = 32M  
myisam\_recover = FORCE,BACKUP

# SAFETY

max\_allowed\_packet = 64M  
max\_connect\_errors = 1000000  
skip\_name\_resolve  
sql\_mode = STRICT\_TRANS\_TABLES,ERROR\_FOR\_DIVISION\_BY\_ZERO,NO\_AUTO\_CREATE\_USER,NO\_AUTO\_VALUE\_  
ON\_ZERO,NO\_ENGINE\_SUBSTITUTION,NO\_ZERO\_DATE,NO\_ZERO\_IN\_DATE,ONLY\_FULL\_GROUP\_BY  
sysdate\_is\_now = 1  
innodb = FORCE  
innodb\_strict\_mode = 1

# DATA STORAGE

datadir = /mnt/data/

# BINARY LOGGING

log\_bin = /mnt/data/mysql-bin  
expire\_logs\_days = 14  
sync\_binlog = 1  
binlog\_format = ROW

# CACHES AND LIMITS

tmp\_table\_size = 128M  
max\_heap\_table\_size = 128M  
query\_cache\_type = 0  
query\_cache\_size = 8  
max\_connections = 2010  
thread\_cache\_size = 50  
open\_files\_limit = 65535  
table\_definition\_cache = 4096  
table\_open\_cache = 12000

# INNODB

innodb\_flush\_method = O\_DIRECT  
innodb\_log\_files\_in\_group = 2  
innodb\_log\_file\_size = 256M  
innodb\_flush\_log\_at\_trx\_commit = 1  
innodb\_file\_per\_table = 1  
innodb\_buffer\_pool\_size = 14G  
innodb\_locks\_unsafe\_for\_binlog = 1  
innodb\_autoinc\_lock\_mode = 2  
wait\_timeout = 1500  
interactive\_timeout = 1500

* * *

---

<div class="post-metadata">

**Author:** ![przemek](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/przemek/32/3_2.png) [@przemek](https://forums.percona.com/u/przemek)\
**Post date:** [December 27, 2013, 8:56am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/3 "2013-12-27T08:56:05Z")

</div>

Try changing two things: PXC never version then 5.5.27 (latest if possible), and wsrep\_sst\_method=xtrabackup (much less locking then rsync).

---

<div class="post-metadata">

**Author:** ![Yoganand](https://avatars.discourse-cdn.com/v4/letter/y/f17d59/32.png) [@Yoganand](https://forums.percona.com/u/Yoganand)\
**Post date:** [December 28, 2013, 12:18am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/4 "2013-12-28T00:18:06Z")

</div>

Hi Przemek, Thanks for the suggestion.  
Actually all these days, all the 3 nodes are working quite fine. Now and then nodes went out of sync and we used to take down time and resync the nodes. After this cluster used to come back to normalcy.  
But in the latest scenario, 2 nodes went out of sync. Actions done are below.

1. Took downtime 2) resynced the nodes from surviving Node 2 , resync successfully completed 3) created a blank schema in one node and verified the same in other nodes. blank schema synchronized in other nodes also. i.e, OK. 4) After 15 minutes, nodes went out of sync again. ERRORS are posted above.  
From the errors, errors are different from each node which are out of sync.(Node 1 & Node 3)  
Could u pl check these errors and suggest what can be done to fix and bring nodes to sync and stay up and running. Our problem is nodes goes out of sync very frequently.

As a long term action, we can upgrade the PXC and galera to latest version.  
But for immediate action, any suggestions , so that 3 node cluster comes back to working status as it was working previously.

---

<div class="post-metadata">

**Author:** ![Yoganand](https://avatars.discourse-cdn.com/v4/letter/y/f17d59/32.png) [@Yoganand](https://forums.percona.com/u/Yoganand)\
**Post date:** [December 28, 2013, 12:28am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/5 "2013-12-28T00:28:45Z")

</div>

Also to give info on the cluster type, this is a multi master 3 node cluster ( all are masters).  
Please suggest…

---

<div class="post-metadata">

**Author:** ![Poorna\_PC](https://avatars.discourse-cdn.com/v4/letter/p/e19b73/32.png) [@Poorna\_PC](https://forums.percona.com/u/Poorna_PC)\
**Post date:** [December 28, 2013, 4:04am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/6 "2013-12-28T04:04:52Z")

</div>

Thanks Przemek.

Just to check, Galera 2.0 does not support [COLOR=#252C2F]wsrep\_sst\_method=xtrabackup (pls correct me, If i’m wrong).

Also we are using 5 HAProxy clients in 5 App system for Application connection to cluster database.

Also write always happens to 1 node only from all 5 APP(HAProxy), to avoid deadlock.

---

<div class="post-metadata">

**Author:** ![mixa](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/mixa/32/16881_2.png) [@mixa](https://forums.percona.com/u/mixa)\
**Post date:** [January 1, 2014, 11:31am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/7 "2014-01-01T11:31:35Z")

</div>

Poorna PC,

[COLOR=#252C2F]wsrep\_sst\_method=xtrabackup - it’s ok for galera 2.0  
you can read about pros and cons on codership site:  
[http://www.codership.com/wiki/doku.php?id=sst\_mysql](http://www.codership.com/wiki/doku.php?id=sst_mysql)

the one of errors is similar to error described in bug:  
[https://bugs.launchpad.net/codership-mysql/+bug/1057910](https://bugs.launchpad.net/codership-mysql/+bug/1057910)

> [@](#):
>
> This issue has not been reproduced so far. Code analysis shows that foreign key check will fail if one of the parts in the key has NULL value.

The bug fixed in 5.5.28 version.  
So I’d suggest to upgrade.

---

<div class="post-metadata">

**Author:** ![Poorna\_PC](https://avatars.discourse-cdn.com/v4/letter/p/e19b73/32.png) [@Poorna\_PC](https://forums.percona.com/u/Poorna_PC)\
**Post date:** [January 5, 2014, 7:33am UTC](https://forums.percona.com/t/2-nodes-are-going-out-of-sync-in-3-nodes-pxc-set-up-at-the-time-of-re-syncing/3165/8 "2014-01-05T07:33:19Z")

</div>

Thanks Mixa…
