# Is my cluster back in sync?

**URL:** <https://forums.percona.com/t/is-my-cluster-back-in-sync/4324>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [July 28, 2015, 2:03am UTC](https://forums.percona.com/t/is-my-cluster-back-in-sync/4324 "2015-07-28T02:03:31Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![nightfly](https://avatars.discourse-cdn.com/v4/letter/n/e79b87/32.png) [@nightfly](https://forums.percona.com/u/nightfly)\
**Post date:** [July 28, 2015, 2:03am UTC](https://forums.percona.com/t/is-my-cluster-back-in-sync/4324/1 "2015-07-28T02:03:31Z")

</div>

Hello,

I have a 2 node percona-xtradb-cluster-56 in multi-master mode. The server was crashed (both nodes) then bootstrapped, maybe in the wrong order because the users noticed that some data wasn’t up to date. It’s not relevant, what I would like to know if it is back to normal synchronized mode right now or it’s still in some flaky unpredictable state.

Here are the variables for NODE1:

±-----------------------------±--------------------------------------------------+  
| Variable\_name | Value |  
±-----------------------------±--------------------------------------------------+  
| wsrep\_local\_state\_uuid | 53ef5e93-33de-11e5-adb5-1b37c6b643bd |  
| wsrep\_protocol\_version | 6 |  
| wsrep\_last\_committed | 912789 |  
| wsrep\_replicated | 0 |  
| wsrep\_replicated\_bytes | 0 |  
| wsrep\_repl\_keys | 0 |  
| wsrep\_repl\_keys\_bytes | 0 |  
| wsrep\_repl\_data\_bytes | 0 |  
| wsrep\_repl\_other\_bytes | 0 |  
| wsrep\_received | 776 |  
| wsrep\_received\_bytes | 56548807 |  
| wsrep\_local\_commits | 0 |  
| wsrep\_local\_cert\_failures | 0 |  
| wsrep\_local\_replays | 0 |  
| wsrep\_local\_send\_queue | 0 |  
| wsrep\_local\_send\_queue\_max | 2 |  
| wsrep\_local\_send\_queue\_min | 0 |  
| wsrep\_local\_send\_queue\_avg | 0.111111 |  
| wsrep\_local\_recv\_queue | 0 |  
| wsrep\_local\_recv\_queue\_max | 151 |  
| wsrep\_local\_recv\_queue\_min | 0 |  
| wsrep\_local\_recv\_queue\_avg | 15.442010 |  
| wsrep\_local\_cached\_downto | 912023 |  
| wsrep\_flow\_control\_paused\_ns | 458 |  
| wsrep\_flow\_control\_paused | 0.000000 |  
| wsrep\_flow\_control\_sent | 1 |  
| wsrep\_flow\_control\_recv | 1 |  
| wsrep\_cert\_deps\_distance | 24.422425 |  
| wsrep\_apply\_oooe | 0.000000 |  
| wsrep\_apply\_oool | 0.000000 |  
| wsrep\_apply\_window | 1.000000 |  
| wsrep\_commit\_oooe | 0.000000 |  
| wsrep\_commit\_oool | 0.000000 |  
| wsrep\_commit\_window | 1.000000 |  
| wsrep\_local\_state | 4 |  
| wsrep\_local\_state\_comment | Synced |  
| wsrep\_cert\_index\_size | 7 |  
| wsrep\_causal\_reads | 0 |  
| wsrep\_cert\_interval | 0.234681 |  
| wsrep\_incoming\_addresses | |  
| wsrep\_evs\_delayed | |  
| wsrep\_evs\_evict\_list | |  
| wsrep\_evs\_repl\_latency | 0.000237055/0.000330571/0.00051501/8.63108e-05/13 |  
| wsrep\_evs\_state | OPERATIONAL |  
| wsrep\_gcomm\_uuid | fa66fe2c-34f5-11e5-8c84-035250e696ad |  
| wsrep\_cluster\_conf\_id | 4 |  
| wsrep\_cluster\_size | 2 |  
| wsrep\_cluster\_state\_uuid | 53ef5e93-33de-11e5-adb5-1b37c6b643bd |  
| wsrep\_cluster\_status | Primary |  
| wsrep\_connected | ON |  
| wsrep\_local\_bf\_aborts | 0 |  
| wsrep\_local\_index | 1 |  
| wsrep\_provider\_name | Galera |  
| wsrep\_provider\_vendor | Codership Oy \<info@codership.com\> |  
| wsrep\_provider\_version | 3.8(rf6147dd) |  
| wsrep\_ready | ON |  
±-----------------------------±--------------------------------------------------+

---

<div class="post-metadata">

**Author:** ![nightfly](https://avatars.discourse-cdn.com/v4/letter/n/e79b87/32.png) [@nightfly](https://forums.percona.com/u/nightfly)\
**Post date:** [July 28, 2015, 2:05am UTC](https://forums.percona.com/t/is-my-cluster-back-in-sync/4324/2 "2015-07-28T02:05:12Z")

</div>

Here are the variables for NODE2:

±-----------------------------±---------------------------------------------------+  
| Variable\_name | Value |  
±-----------------------------±---------------------------------------------------+  
| wsrep\_local\_state\_uuid | 53ef5e93-33de-11e5-adb5-1b37c6b643bd |  
| wsrep\_protocol\_version | 6 |  
| wsrep\_last\_committed | 922781 |  
| wsrep\_replicated | 922780 |  
| wsrep\_replicated\_bytes | 23855221085 |  
| wsrep\_repl\_keys | 3071856 |  
| wsrep\_repl\_keys\_bytes | 45801377 |  
| wsrep\_repl\_data\_bytes | 6849937076 |  
| wsrep\_repl\_other\_bytes | 0 |  
| wsrep\_received | 7261 |  
| wsrep\_received\_bytes | 60045 |  
| wsrep\_local\_commits | 922777 |  
| wsrep\_local\_cert\_failures | 0 |  
| wsrep\_local\_replays | 0 |  
| wsrep\_local\_send\_queue | 0 |  
| wsrep\_local\_send\_queue\_max | 10 |  
| wsrep\_local\_send\_queue\_min | 0 |  
| wsrep\_local\_send\_queue\_avg | 0.002493 |  
| wsrep\_local\_recv\_queue | 0 |  
| wsrep\_local\_recv\_queue\_max | 2 |  
| wsrep\_local\_recv\_queue\_min | 0 |  
| wsrep\_local\_recv\_queue\_avg | 0.003443 |  
| wsrep\_local\_cached\_downto | 921845 |  
| wsrep\_flow\_control\_paused\_ns | 43461847996 |  
| wsrep\_flow\_control\_paused | 0.000357 |  
| wsrep\_flow\_control\_sent | 0 |  
| wsrep\_flow\_control\_recv | 72 |  
| wsrep\_cert\_deps\_distance | 19.291316 |  
| wsrep\_apply\_oooe | 0.057334 |  
| wsrep\_apply\_oool | 0.000002 |  
| wsrep\_apply\_window | 1.086623 |  
| wsrep\_commit\_oooe | 0.000000 |  
| wsrep\_commit\_oool | 0.000000 |  
| wsrep\_commit\_window | 1.029616 |  
| wsrep\_local\_state | 4 |  
| wsrep\_local\_state\_comment | Synced |  
| wsrep\_cert\_index\_size | 29 |  
| wsrep\_causal\_reads | 0 |  
| wsrep\_cert\_interval | 0.093278 |  
| wsrep\_incoming\_addresses | |  
| wsrep\_evs\_delayed | |  
| wsrep\_evs\_evict\_list | |  
| wsrep\_evs\_repl\_latency | 0.000275245/0.00125926/0.00935097/0.000899017/1104 |  
| wsrep\_evs\_state | OPERATIONAL |  
| wsrep\_gcomm\_uuid | 53eef647-33de-11e5-9145-eade2fa688ff |  
| wsrep\_cluster\_conf\_id | 4 |  
| wsrep\_cluster\_size | 2 |  
| wsrep\_cluster\_state\_uuid | 53ef5e93-33de-11e5-adb5-1b37c6b643bd |  
| wsrep\_cluster\_status | Primary |  
| wsrep\_connected | ON |  
| wsrep\_local\_bf\_aborts | 0 |  
| wsrep\_local\_index | 0 |  
| wsrep\_provider\_name | Galera |  
| wsrep\_provider\_vendor | Codership Oy \<info@codership.com\> |  
| wsrep\_provider\_version | 3.8(rf6147dd) |  
| wsrep\_ready | ON |  
±-----------------------------±---------------------------------------------------+

The wsrep\_local\_state\_comment say it is Synced but it even say that if I shut one node down. When I create database and load data in on any side that gets replicated all right to the other side. What worries me is this missing data the users said and what I have found in the documentation “How to recover a PXC cluster Scenario6”.

If I look into this grastate.dat on the nodes:

# GALERA saved state

version: 2.1  
uuid: 53ef5e93-33de-11e5-adb5-1b37c6b643bd  
seqno: -1  
cert\_index:

The seqno is -1 instead of the last valid sequence number, now I don’t know if this is normal.

Also the logs were full of warnings such as (on node2):

2015-07-28 06:25:02 30682 [Warning] InnoDB: Cannot open table mydb/field\_revision\_field\_decision\_nl\_mydb\_one from the internal data dictionary of InnoDB though the .frm file for the table exists. See [http://dev.mysql.com/doc/refman/5.6/en/innodb-troubleshooting.html](http://dev.mysql.com/doc/refman/5.6/en/innodb-troubleshooting.html) for how you can resolve the problem.

I would just really like to know if this is back to normal state so we can continue the work or not.

Can someone help?

Thanks

---

<div class="post-metadata">

**Author:** ![jrivera](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/jrivera/32/13_2.png) [@jrivera](https://forums.percona.com/u/jrivera)\
**Post date:** [July 28, 2015, 10:23pm UTC](https://forums.percona.com/t/is-my-cluster-back-in-sync/4324/3 "2015-07-28T22:23:14Z")

</div>

You can clean up node2’s datadir and then restart node2 to SST from node1 to make sure you have consistent data between the nodes. Check to make sure you do not use MyISAM tables, if so convert them to InnoDB. Add a Garbd node to avoid split brain when one node can’t communicate to the other node.
