# Randomly IST fail

**URL:** <https://forums.percona.com/t/randomly-ist-fail/4235>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [May 20, 2015, 11:21am UTC](https://forums.percona.com/t/randomly-ist-fail/4235 "2015-05-20T11:21:49Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fernando\_Mattera](https://avatars.discourse-cdn.com/v4/letter/f/bcef8e/32.png) [@Fernando\_Mattera](https://forums.percona.com/u/Fernando_Mattera)\
**Post date:** [May 20, 2015, 11:21am UTC](https://forums.percona.com/t/randomly-ist-fail/4235/1 "2015-05-20T11:21:49Z")

</div>

Hello,

first, thank you for your excellent products like xtrabackup, server and cluster (and so on…)

Very well, we have 5-node cluster, I graceful stopped 1st node ad the others still writing, and when I start the stopped node sometimes I have this message:

2015-05-20 12:45:38 29587 [Note] WSREP: New cluster view: global state: f7dc1bbb-f8b8-11e4-8172-3f0c0cf99383:165132, view# 59: Primary, number of nodes: 5, my index: 1, protocol version 3  
2015-05-20 12:45:38 29587 [Warning] WSREP: Gap in state sequence. Need state transfer.  
2015-05-20 12:45:38 29587 [Note] WSREP: Running: 'wsrep\_sst\_xtrabackup-v2 --role ‘joiner’ --address ‘10.1.3.221’ --auth ‘sstuser:despegar#mysql’ --datadir ‘/mysql/data/’ --defaults-file ‘/etc/my.cnf’ --parent ‘29587’ ‘’ ’  
.WSREP\_SST: [INFO] Streaming with xbstream (20150520 12:45:39.638)  
WSREP\_SST: [INFO] Using socat as streamer (20150520 12:45:39.644)  
2015-05-20 12:45:39 29587 [Note] WSREP: Prepared SST request: xtrabackup-v2|10.1.3.221:4444/xtrabackup\_sst//1  
WSREP\_SST: [INFO] Evaluating timeout -s9 100 socat -u TCP-LISTEN:4444,reuseaddr stdio | xbstream -x; RC=( ${PIPESTATUS[@]} ) (20150520 12:45:39.731)  
2015/05/20 12:45:39 socat[29858] E bind(3, {AF=2 0.0.0.0:4444}, 16): Address already in use  
WSREP\_SST: [ERROR] Error while getting data from donor node: exit codes: 1 0 (20150520 12:45:39.748)  
WSREP\_SST: [ERROR] Cleanup after exit with status:32 (20150520 12:45:39.754)  
2015-05-20 12:45:39 29587 [ERROR] WSREP: Process completed with error: wsrep\_sst\_xtrabackup-v2 --role ‘joiner’ --address ‘10.1.3.221’ --auth ‘sstuser:despegar#mysql’ --datadir ‘/mysql/data/’ --defaults-file ‘/etc/my.cnf’ --parent ‘29587’ ‘’ : 32 (Broken pipe)  
2015-05-20 12:45:39 29587 [ERROR] WSREP: Failed to read uuid:seqno from joiner script.  
2015-05-20 12:45:39 29587 [ERROR] WSREP: SST failed: 32 (Broken pipe)  
2015-05-20 12:45:39 29587 [ERROR] Aborting

2015-05-20 12:45:39 29587 [Note] WSREP: REPL Protocols: 7 (3, 2)  
2015-05-20 12:45:39 29587 [Note] WSREP: Service thread queue flushed.  
2015-05-20 12:45:39 29587 [Note] WSREP: Assign initial position for certification: 165132, protocol version: 3  
2015-05-20 12:45:39 29587 [Note] WSREP: Service thread queue flushed.  
2015-05-20 12:45:39 29587 [Note] WSREP: Prepared IST receiver, listening at: tcp://10.1.3.221:4568  
2015-05-20 12:45:39 29587 [Note] WSREP: Member 1.0 (xtradb-cloudia-01) requested state transfer from ‘_any_’. Selected 4.0 (xtradb-cloudia-04)(SYNCED) as donor.  
2015-05-20 12:45:39 29587 [Note] WSREP: Shifting PRIMARY → JOINER (TO: 165132)  
2015-05-20 12:45:39 29587 [Note] WSREP: Requesting state transfer: success, donor: 4  
State transfer in progress, setting sleep higher  
.xbstream: Can’t create/write to file ‘./xtrabackup\_galera\_info’ (Errcode: 2 - No such file or directory)  
xbstream: failed to create file.

Any idea?  
Need any other data?  
Thank you so much

---

<div class="post-metadata">

**Author:** ![Fernando\_Mattera](https://avatars.discourse-cdn.com/v4/letter/f/bcef8e/32.png) [@Fernando\_Mattera](https://forums.percona.com/u/Fernando_Mattera)\
**Post date:** [May 20, 2015, 12:03pm UTC](https://forums.percona.com/t/randomly-ist-fail/4235/2 "2015-05-20T12:03:33Z")

</div>

Interest thing, when I retry to start this node, PXC do a SST.  
And they’re not a lot of writes in order to do that.

---

<div class="post-metadata">

**Author:** ![Fernando\_Mattera](https://avatars.discourse-cdn.com/v4/letter/f/bcef8e/32.png) [@Fernando\_Mattera](https://forums.percona.com/u/Fernando_Mattera)\
**Post date:** [May 21, 2015, 10:38am UTC](https://forums.percona.com/t/randomly-ist-fail/4235/3 "2015-05-21T10:38:28Z")

</div>

It’s seems to be socat hang or take port 4444 after successful IST and when you try to stop/start MySQL again, port is taken.  
A workaround is “killall -9 socat”  
But I don’t like it.  
This is a Galera or Percona bug?  
Thanks.

---

<div class="post-metadata">

**Author:** ![jrivera](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/jrivera/32/13_2.png) [@jrivera](https://forums.percona.com/u/jrivera)\
**Post date:** [May 21, 2015, 3:52pm UTC](https://forums.percona.com/t/randomly-ist-fail/4235/4 "2015-05-21T15:52:59Z")

</div>

Post your my.cnf, and if you still have it, post contents of innobackup.\* logs on donor and joiner node

---

<div class="post-metadata">

**Author:** ![Fernando\_Mattera](https://avatars.discourse-cdn.com/v4/letter/f/bcef8e/32.png) [@Fernando\_Mattera](https://forums.percona.com/u/Fernando_Mattera)\
**Post date:** [May 22, 2015, 9:34am UTC](https://forums.percona.com/t/randomly-ist-fail/4235/5 "2015-05-22T09:34:31Z")

</div>

Here is my.cnf

[mysqld\_safe]  
pid-file = /home/despegar/mysql/mysql.pid

[mysqld]  
user = despegar  
port = 33033  
socket = /home/despegar/mysql/mysql.sock  
back\_log = 50  
datadir = /home/despegar/mysql/data  
log-error = /home/despegar/mysql/mysql.err

max\_connections = 10000  
max\_connect\_errors = 99999  
max\_allowed\_packet = 16M  
skip-host-cache  
skip-name-resolve

explicit\_defaults\_for\_timestamp = 1  
performance\_schema = off

transaction\_isolation = READ-COMMITTED  
log-bin = /home/despegar/mysql/binlog/mysql-bin  
slow\_query\_log = 1  
log\_output = TABLE  
long\_query\_time = 5  
event\_scheduler = ON

# \*\*\* INNODB Specific options \*\*\*

innodb\_buffer\_pool\_size = 3G  
innodb\_data\_file\_path = ibdata1:256M:autoextend  
innodb\_thread\_concurrency = 16  
innodb\_log\_file\_size = 256M  
innodb\_log\_files\_in\_group = 3  
innodb\_file\_per\_table = ON  
innodb\_undo\_tablespaces = 3  
innodb\_undo\_directory = /home/despegar/mysql/undo

# \*\*\* Galera Cluster Settings \*\*\*

binlog\_format = ROW  
default-storage-engine = innodb  
innodb\_autoinc\_lock\_mode = 2  
log\_slave\_updates = ON  
query\_cache\_size = 0  
query\_cache\_type = 0

# wsrep Provider Settings

wsrep\_provider = /usr/lib64/galera3/libgalera\_smm.so  
wsrep\_provider\_options = “gcache.size=128M;gcache.page\_size=128M”  
wsrep\_cluster\_address = “gcomm://xtradblab00,xtradblab01,xtradblab02,xtradblab03”  
wsrep\_cluster\_name = “xtradb-lab”  
wsrep\_node\_address = “xtradblab00”  
wsrep\_node\_name = “xtradblab00”  
wsrep\_sst\_method = xtrabackup-v2  
wsrep\_sst\_auth = “sstuser:despegar#mysql”  
wsrep\_node\_incoming\_address = “10.70.128.203:33033”  
wsrep\_sst\_receive\_address = “10.70.128.203”  
wsrep\_slave\_threads = 8

innobackup.log not been created just for this reason (port in use)
