# xtradb cluster crashing problem

**URL:** <https://forums.percona.com/t/xtradb-cluster-crashing-problem/2527>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [January 8, 2013, 8:59am UTC](https://forums.percona.com/t/xtradb-cluster-crashing-problem/2527 "2013-01-08T08:59:30Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![jasonwolf](https://avatars.discourse-cdn.com/v4/letter/j/bc8723/32.png) [@jasonwolf](https://forums.percona.com/u/jasonwolf)\
**Post date:** [January 8, 2013, 8:59am UTC](https://forums.percona.com/t/xtradb-cluster-crashing-problem/2527/1 "2013-01-08T08:59:30Z")

</div>

We converted an existing single point failure mysql to a percona xtradb cluster last wednesday and have been having some issues.

Description:  
During normal load everything works completely fine.

On saturday during nightly maintenance jobs the nodes started to become unresponsive, loaded, and/or failed.  
Running unclustered the node can easily handle all incoming requests and sits about 0.1 load most of the time.  
The nightly maintenance does spike cpu and IO load from 3:00AM to 3:40AM-ish  
However it doesnt appear that the load should be soo excessive as to bring down both nodes of the cluster.  
I am struggling to understand why a single unclustered mysql handles it fine, but a clustered mysql crashes 2 nodes completely.  
I can completely understand if it uses more resources and takes longer to do its tasks, but not crash.

Question: When xtradb cluster starts generating “gcache.page” files does that mean it isn’t keeping up with demand?

Information:  
The cluster has around 250GB of actual mysql data.

Each of the 2 nodes is a dual quad core xeon E5520, 16GB ram, 74GB RAID1 SAS for OS, 600GB RAID10 15K SAS just for mysql data.

Both nodes running Ubuntu 12.04, with percona xtradb cluster 5.5.28-23.7-369.precise

Yes, I know I should add a 3rd node and it has been ordered but hasnt arrived yet. Yes, I should have just waited until I received the 3rd node to convert the existing mysql to a percona cluster.

Config:

# Generated by Percona Configuration Wizard version REL5-20120208

# Configuration name mysqlc03 generated at 2012-12-27 14:04:51

[mysql]

# CLIENT

port = 3306  
socket = /var/run/mysqld/mysqld.sock

[mysqld\_safe]  
wsrep\_urls=[gcomm://192.168.252.107:4567,gcomm://](https://gcomm)

[mysqld]

# GENERAL

user = mysql  
default\_storage\_engine = InnoDB  
socket = /var/run/mysqld/mysqld.sock  
pid\_file = /var/run/mysqld/mysqld.pid

# MyISAM

key\_buffer\_size = 32M  
myisam\_recover = FORCE,BACKUP

# SAFETY

max\_allowed\_packet = 16M  
max\_connect\_errors = 1000000  
skip\_name\_resolve  
#sql\_mode = STRICT\_TRANS\_TABLES,ERROR\_FOR\_DIVISION\_BY\_ZERO,NO\_AUTO\_CREAT E\_USER,NO\_AUTO\_VALUE\_ON\_ZERO,NO\_ENGINE\_SUBSTITUTION,NO\_ZERO\_ DATE,NO\_ZERO\_IN\_DATE,ONLY\_FULL\_GROUP\_BY  
#sysdate\_is\_now = 1  
#innodb = FORCE  
#innodb\_strict\_mode = 1

# DATA STORAGE

datadir = /var/lib/mysql/

# BINARY LOGGING

log\_bin = /var/lib/mysql/mysql-bin  
expire\_logs\_days = 14  
sync\_binlog = 1

# CACHES AND LIMITS

tmp\_table\_size = 32M  
max\_heap\_table\_size = 32M  
query\_cache\_type = 0  
query\_cache\_size = 0  
max\_connections = 3000  
thread\_cache\_size = 100  
open\_files\_limit = 65535  
table\_definition\_cache = 1024  
table\_open\_cache = 2048

# INNODB

innodb\_flush\_method = O\_DIRECT  
innodb\_log\_files\_in\_group = 2  
innodb\_log\_file\_size = 256M  
innodb\_flush\_log\_at\_trx\_commit = 1  
innodb\_file\_per\_table = 1  
innodb\_buffer\_pool\_size = 14G # total system memory - ~2G for system processes if system is dedicated to mysql

# LOGGING

log\_error = /var/log/mysql/mysql-error.log  
log\_queries\_not\_using\_indexes = 1  
slow\_query\_log = 1  
slow\_query\_log\_file = /var/log/mysql/mysql-slow.log

# XTRADB CLUSTERING

#wsrep\_cluster\_address=gcomm:// # on first node  
#wsrep\_cluster\_address=[gcomm://192.168.252.107,192.168.252.109,](https://gcomm) # on each additional node  
wsrep\_provider=/usr/lib/libgalera\_smm.so  
wsrep\_provider\_options = “gmcast.listen\_addr=[tcp://192.168.252.108](https://tcp);”  
wsrep\_sst\_receive\_address=192.168.252.108  
wsrep\_node\_incoming\_address=192.168.252.108  
wsrep\_node\_name=node2 # name each node unique  
wsrep\_slave\_threads=16 # set to the number of cpu cores  
wsrep\_sst\_method=xtrabackup  
wsrep\_sst\_auth=  
wsrep\_cluster\_name=mysqlc03  
binlog\_format=ROW  
default\_storage\_engine=InnoDB  
innodb\_autoinc\_lock\_mode=2  
innodb\_locks\_unsafe\_for\_binlog=1

Nightly Procedure log:  
130104 3:01:52 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000000 of size 162693128 bytes  
130104 3:01:59 [Note] WSREP: Deleted page /var/lib/mysql/gcache.page.000000  
… \< Every day between 3AM and 3:40ish it generates these gcache.page files during nightly stored procedure jobs \>  
130106 3:29:33 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000011 of size 372497251 bytes  
130106 3:29:45 [Note] WSREP: Deleted page /var/lib/mysql/gcache.page.000011

Sunday crash log:  
Node1:  
130106 3:29:33 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000011 of size 372497251 bytes  
130106 3:29:45 [Note] WSREP: Deleted page /var/lib/mysql/gcache.page.000011  
130106 3:40:54 [Warning] Too many connections  
130106 3:40:56 [Warning] Too many connections  
…  
130106 4:06:16 [Warning] Too many connections  
130106 4:06:17 [Note] WSREP: Read nil XID from storage engines, skipping position init  
130106 4:06:17 [Note] WSREP: wsrep\_load(): loading provider library ‘/usr/lib/libgalera\_smm.so’  
130106 4:06:17 [Note] WSREP: wsrep\_load(): Galera 2.1(r113) by Codership Oy \<[info&#64;codership.com](mailto:info&#64;codership.com)\> loaded succesfully.  
130106 4:06:17 [Note] WSREP: Found saved state: 510f28eb-5582-11e2-0800-414351a60eb6:-1  
130106 4:06:17 [Note] WSREP: Reusing existing ‘/var/lib/mysql//galera.cache’.  
\<Node 1 appears to start a second instance of mysqld for some reason\>  
130106 4:06:18 InnoDB: Initializing buffer pool, size = 6.0G  
130106 4:06:18 [Warning] Too many connections  
130106 4:06:18 InnoDB: Completed initialization of buffer pool  
InnoDB: Unable to lock ./ibdata1, error: 11  
InnoDB: Check that you do not already have another mysqld process  
InnoDB: using the same InnoDB data or log files.  
130106 4:06:18 InnoDB: Retrying to lock the first data file  
130106 4:06:18 [Warning] Too many connections  
\< it generates repeated “unable to lock ./ibdata1” errors at the same time as “Too many connections” from the first mysql instance  
However it is unresponsive to queries. \>

Node2:  
130106 3:43:40 [Note] WSREP: Deleted page /var/lib/mysql/gcache.page.000011  
130106 3:43:41 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000012 of size 312412590 bytes  
130106 3:58:15 [Warning] Too many connections  
130106 3:58:17 [Warning] Too many connections  
…  
130106 4:11:10 [Warning] Too many connections  
130106 4:11:11 [Warning] Too many connections  
130106 4:11:11 [Note] /usr/sbin/mysqld: Normal shutdown  
\<At this point I shut down this node as it is unresponsive to all queries and just dumping “Too many connections” to the logs  
The load is incoming insert writes from a web application on the order of 15 inserts per second.  
The maintenance procedures are purging old entries and generating a few reports.\>  
\<I kill both mysql processes and restart both cluster nodes and sync them by deleting  
/var/lib/mysql/grastate.dat on node2 to trigger an SST\>

Monday crash log:  
Node1:  
130107 3:04:58 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000000 of size 419565861 bytes  
130107 3:05:18 [Note] WSREP: Deleted page /var/lib/mysql/gcache.page.000000  
…  
130107 3:43:12 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000005 of size 564372059 bytes  
130107 3:43:37 [Note] WSREP: Deleted page /var/lib/mysql/gcache.page.000005  
130107 3:45:07 [Note] WSREP: (19068069-57ef-11e2-0800-1a8ff159f030, ‘[tcp://192.168.252.107:4567](https://tcp)’) turning message relay requesting on, nonlive peers: [tcp://192.168.252.108:4567](https://tcp)  
130107 3:45:08 [Note] WSREP: (19068069-57ef-11e2-0800-1a8ff159f030, ‘[tcp://192.168.252.107:4567](https://tcp)’) reconnecting to 724621c5-57ea-11e2-0800-3eac77a6d452 ([tcp://192.168.252.108:4567](https://tcp)), attempt 0  
130107 3:45:09 [Note] WSREP: evs::proto(19068069-57ef-11e2-0800-1a8ff159f030, OPERATIONAL, view\_id(REG,19068069-57ef-11e2-0800-1a8ff159f030,2)) suspecting node: 724621c5-57ea-11e2-0800-3eac77a6d452

Node2:  
130107 3:04:58 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000000 of size 419565861 bytes  
130107 3:11:42 [Note] WSREP: Deleted page /var/lib/mysql/gcache.page.000000  
…  
130107 3:43:12 [Note] WSREP: Created page /var/lib/mysql/gcache.page.000005 of size 564372059 bytes  
Killed  
130107 03:45:07 mysqld\_safe Number of processes running now: 0  
130107 03:45:07 mysqld\_safe WSREP: not restarting wsrep node automatically  
130107 03:45:07 mysqld\_safe mysqld from pid file /var/run/mysqld/mysqld.pid ended

---

<div class="post-metadata">

**Author:** ![jasonwolf](https://avatars.discourse-cdn.com/v4/letter/j/bc8723/32.png) [@jasonwolf](https://forums.percona.com/u/jasonwolf)\
**Post date:** [January 8, 2013, 9:15am UTC](https://forums.percona.com/t/xtradb-cluster-crashing-problem/2527/2 "2013-01-08T09:15:51Z")

</div>

Reading through the Percona Xtradb Cluster Limitations page again I see,

“DELETE operation is unsupported on tables without primary key. Also rows in tables without primary key may appear in different order on different nodes. As a result SELECT…LIMIT… may return slightly different sets.”

I went and looked and quite a few large tables do not have keys or indexes. Do you think this could actually cause the behavior I was seeing?
