# Galera inconsistent data?

**URL:** <https://forums.percona.com/t/galera-inconsistent-data/4794>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [March 20, 2016, 2:18pm UTC](https://forums.percona.com/t/galera-inconsistent-data/4794 "2016-03-20T14:18:22Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![bradh352](https://avatars.discourse-cdn.com/v4/letter/b/ec9cab/32.png) [@bradh352](https://forums.percona.com/u/bradh352)\
**Post date:** [March 20, 2016, 2:18pm UTC](https://forums.percona.com/t/galera-inconsistent-data/4794/1 "2016-03-20T14:18:22Z")

</div>

I’ve been doing some testing of my application against Percona-XtraDB-Cluster-server-56-5.6.28-25.14.1.el7.x86\_64,  
and have run across something I cannot explain, and it appears to only occur under load.

I’ve got a load balancer sitting in front of 3 nodes directing traffic in a round robin fashion (Linux IPVS/LVS using server load  
balancing features in keepalived). My application spawns 25 DB connections, resulting in 8-9 connections per DB node.

Essentially my application is doing something like:

SET TRANSACTION ISOLATION LEVEL SERIALIZABLE;  
BEGIN;  
SELECT var FROM foo WHERE id = ? FOR UPDATE;

# Do some math on ‘var’

UPDATE foo SET var=? WHERE id=?;  
COMMIT;

“id” is the primary key in this case, so basically doing the same thing as this blog post:  
[http://galeracluster.com/2015/09/sup…alera-cluster/](http://galeracluster.com/2015/09/support-for-mysql-transaction-isolation-levels-in-galera-cluster/)  
I understand “SERIALIZABLE” isn’t really supported, but I’d hope Galera wouldn’t downgrade it  
to worse than REPEATABLE READ or something.

What I am observing is “N” concurrent committers get the same select result (expected with galera), do their math,  
update with math result which succeeds (expected with galera), then commit and all succeed (_unexpected_).  
I would have expected all but one committer to deadlock (1213) and I’d need to retry the transaction on the  
deadlocked nodes. I do get some occasional deadlocks reported, but not as many as I should. This leads to  
inconsistent results since the math performed is now off … needless to say this causes major issues.

If I attempt to reproduce this with the mysql command line tool, I cannot, all but one case results in a deadlock as  
expected.

A few other notes:

1. The math operation performed may or may not be the same for each transaction, it depends on the operation,  
but some may indeed appear to return the _same_ result (Does this matter?)
2. The number of committers (“N”) here could be as high as the number of connections, in theory, but in  
practice it is probably 1-3 … but this statement during a load test is executed a few dozen times per  
_second_.
3. This works flawlessly if all traffic is directed to a single node in the cluster, so it is definitely something introduced  
by using Galera.
4. The value of wsrep\_sync\_wait appears to have zero effect (values 0, 1, 3 tried)
5. I could attempt to change the “UPDATE” statement to include the original value of ‘foo’ that is expected  
to try to catch this condition, but I’m not actually sure if it would depending on how the record updates  
are transmitted. I have not tried that, but wouldn’t think it should be necessary.
6. I’m 90% sure we tested this same load scenario a year or so ago and this worked fine, could there be  
a regression in newer versions? I haven’t yet tried to roll back to something older.
7. I haven’t yet tried to create a reduced test case, I wanted to first ask to see if anyone else had seen  
what I’m seeing.

And finally, I’m sure everyone wants my DB configs …

/etc/my.cnf:  
[mysqld]  
datadir = /var/lib/mysql

# move tmpdir due to /tmp being a memory backed tmpfs filesystem, mysql uses this for on disk sorting

tmpdir = /var/lib/mysql/tmp

[mysqld\_safe]  
pid-file = /run/mysqld/mysql.pid  
syslog  
!includedir /etc/my.cnf.d

/etc/my.cnf.d/base.cnf:  
[mysqld]  
bind-address = 0.0.0.0  
key\_buffer = 256M  
max\_allowed\_packet = 16M  
max\_connections = 256

# Some optimizations

thread\_concurrency = 10  
sort\_buffer\_size = 2M  
query\_cache\_limit = 100M  
query\_cache\_size = 256M  
log\_bin  
binlog\_format = ROW  
gtid\_mode = ON  
log\_slave\_updates  
enforce\_gtid\_consistency = 1  
group\_concat\_max\_len = 102400  
innodb\_buffer\_pool\_size = 10G  
innodb\_log\_file\_size = 64M  
innodb\_file\_per\_table = 1  
innodb\_file\_format = barracuda  
default\_storage\_engine = innodb

# SSD Tuning

innodb\_flush\_neighbors = 0  
innodb\_io\_capacity = 6000

/etc/my.cnf.d/cluster.cnf:

# Galera cluster

[mysqld]  
wsrep\_provider = /usr/lib64/libgalera\_smm.so  
wsrep\_sst\_method = xtrabackup-v2  
wsrep\_sst\_auth = “sstuser:s3cretPass”  
wsrep\_cluster\_name = cluster  
wsrep\_slave\_threads = 32  
wsrep\_max\_ws\_size = 2G  
wsrep\_provider\_options = “gcache.size = 5G; pc.recovery = true”  
wsrep\_cluster\_address = gcomm://[10.30.30.11](http://10.30.30.11),10.30.30.12,10.30.30.13  
wsrep\_sync\_wait = 0  
innodb\_autoinc\_lock\_mode = 2  
innodb\_locks\_unsafe\_for\_binlog = 1  
innodb\_flush\_log\_at\_trx\_commit = 0  
sync\_binlog = 0  
innodb\_support\_xa = 0  
innodb\_flush\_method = ALL\_O\_DIRECT

[sst]  
progress = 1  
time = 1  
streamfmt = xbstream

Thanks!  
-Brad

---

<div class="post-metadata">

**Author:** ![bradh352](https://avatars.discourse-cdn.com/v4/letter/b/ec9cab/32.png) [@bradh352](https://forums.percona.com/u/bradh352)\
**Post date:** [March 21, 2016, 2:02pm UTC](https://forums.percona.com/t/galera-inconsistent-data/4794/2 "2016-03-21T14:02:54Z")

</div>

I’ve written a test case for this, though I’m having problems attaching it to here, its not clear what the rules are for attachments, its just a .c file.

Anyhow, one thing I neglected to mention previously is I’m actually getting a deadlock on a couple of connections as well that never return results. If I restart the mysql daemon itself, it unlocks the client side so it can error out.

We’re in the process of trying to roll back to older versions with this test case we created to see if this is a recently introduced regression or not.

---

<div class="post-metadata">

**Author:** ![bradh352](https://avatars.discourse-cdn.com/v4/letter/b/ec9cab/32.png) [@bradh352](https://forums.percona.com/u/bradh352)\
**Post date:** [March 21, 2016, 3:05pm UTC](https://forums.percona.com/t/galera-inconsistent-data/4794/3 "2016-03-21T15:05:04Z")

</div>

I just filed a bug here with my test case: [url][https://bugs.launchpad.net/percona-xtradb-cluster/+bug/1560206[/url]](https://bugs.launchpad.net/percona-xtradb-cluster/+bug/1560206%5B/url%5D)
