# PXC IST Recovery Time After Heavy Write Load (11+ Hours of Continuous Writes)

**URL:** <https://forums.percona.com/t/pxc-ist-recovery-time-after-heavy-write-load-11-hours-of-continuous-writes/41187>\
**Category:** Percona XtraDB Cluster 5.x\
**Created:** [August 20, 2026, 9:46am UTC](https://forums.percona.com/t/pxc-ist-recovery-time-after-heavy-write-load-11-hours-of-continuous-writes/41187 "2026-08-20T09:46:38Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![vagrawal](https://avatars.discourse-cdn.com/v4/letter/v/6a8cbe/32.png) [@vagrawal](https://forums.percona.com/u/vagrawal)\
**Post date:** [August 20, 2026, 9:46am UTC](https://forums.percona.com/t/pxc-ist-recovery-time-after-heavy-write-load-11-hours-of-continuous-writes/41187/1 "2026-08-20T09:46:38Z")

</div>

Hello Percona Team,

We are testing **Percona XtraDB Cluster (PXC) 8.4.x** with the objective of understanding the **recovery time after a node has been offline during a heavy write workload**.

### Test Environment

- Percona XtraDB Cluster: 3 nodes
- PXC Version: 8.4.x
- Kubernetes (RKE2)
- Continuous workload generated with 1000TPS using Sysbench
- Approximately 11 hours of continuous write activity

### Test Scenario

1. Started a continuous Sysbench write workload.
2. Shut down one PXC node.
3. Kept the workload running for approximately 11 hours.
4. Brought the node back online.
5. The node recovered using **IST (Incremental State Transfer)** instead of SST.
6. The synchronization completed much faster than we expected.

### Our Understanding

Our understanding is that:

- PXC uses **Galera replication**.
- IST uses the **gcache** to transfer only the missing write-sets.
- It does **not** replay MySQL binary logs.
- Therefore, even after a large number of writes, recovery can still be very fast as long as the required write-sets are still available in the donor’s gcache.

### Current Cluster Information

Current `wsrep` status shows:

- `wsrep_local_state_comment = Synced`

- `wsrep_cluster_size = 3`

- `wsrep_cluster_status = Primary`

- `wsrep_provider_version = 4.24`

- `wsrep_gcache_pool_size = 327056696` (approximately 312 MB)

- `wsrep_local_cached_downto = 140433700`

- `wsrep_last_committed = 140547929`  
`gcache size - 10GB`

### Our Questions

1. Is our understanding correct?
2. Is it expected that a node can recover very quickly after approximately 11 hours of continuous writes if IST is used?
3. Does the recovery time depend primarily on the amount of missing write-set data rather than the duration for which the node was offline?
4. Am I missing anything in my understanding of how IST recovery works?
5. Are there any recommended methods for accurately testing recovery time after a very large write workload?

We would appreciate any clarification or best practices from the community.

Thank you!

---

<div class="post-metadata">

**Author:** ![vagrawal](https://avatars.discourse-cdn.com/v4/letter/v/6a8cbe/32.png) [@vagrawal](https://forums.percona.com/u/vagrawal)\
**Post date:** [August 20, 2026, 9:47am UTC](https://forums.percona.com/t/pxc-ist-recovery-time-after-heavy-write-load-11-hours-of-continuous-writes/41187/2 "2026-08-20T09:47:20Z")

</div>

Here are the logs-

2026-08-19T05:17:19.554264Z 0 [Note] [MY-000000] [Galera] wsrep\_load(): loading provider library ‘/usr/lib64/galera4/libgalera\_smm.so’  
2026-08-19T05:17:19.554975Z 0 [Note] [MY-000000] [Galera] wsrep\_load(): Galera 4.24(a430f07) by Codership Oy [info@codership.com](mailto:info@codership.com) (modified by Percona [https://percona.com/](https://percona.com/)) loaded successfully.  
2026-08-19T05:17:19.554993Z 0 [Note] [MY-000000] [Galera] Resolved symbol ‘wsrep\_node\_isolation\_mode\_set\_v1’  
2026-08-19T05:17:19.554998Z 0 [Note] [MY-000000] [Galera] Resolved symbol ‘wsrep\_certify\_v1’  
2026-08-19T05:17:19.555002Z 0 [Note] [MY-000000] [Galera] Initializing config service v2  
2026-08-19T05:17:19.555242Z 0 [Note] [MY-000000] [Galera] Deinitializing config service v2  
2026-08-19T05:17:19.555268Z 0 [Note] [MY-000000] [Galera] CRC-32C: using 64-bit x86 acceleration.  
2026-08-19T05:17:19.555398Z 0 [Note] [MY-000000] [Galera] not using SSL compression  
2026-08-19T05:17:19.555895Z 0 [Note] [MY-000000] [Galera] Found saved state: 7cb3e8b0-7149-11f1-b8df-26bb63d99319:140547925, safe\_to\_bootstrap: 0  
2026-08-19T05:17:19.556987Z 0 [Note] [MY-000000] [Galera] GCache DEBUG: opened preamble:  
Version: 2  
UUID: 7cb3e8b0-7149-11f1-b8df-26bb63d99319  
Seqno: 140433700 - 140547925  
Offset: 121950336  
Synced: 1  
EncVersion: 1  
Encrypted: 0  
MasterKeyConst UUID: 7cb05436-7149-11f1-91a3-72f3100dbcc8  
MasterKey UUID: 00000000-0000-0000-0000-000000000000  
MasterKey ID: 0  
2026-08-19T05:17:19.557004Z 0 [Note] [MY-000000] [Galera] Recovering GCache ring buffer: version: 2, UUID: 7cb3e8b0-7149-11f1-b8df-26bb63d99319, offset: 121950336  
2026-08-19T05:17:19.557053Z 0 [Note] [MY-000000] [Galera] GCache::RingBuffer initial scan… 0.0% (0/10737418264 bytes) complete.  
2026-08-19T05:17:19.984982Z 0 [Note] [MY-000000] [Galera] GCache::RingBuffer initial scan… 100.0% (10737418264/10737418264 bytes) complete.  
2026-08-19T05:17:19.986400Z 0 [Note] [MY-000000] [Galera] Recovering GCache ring buffer: found gapless sequence 140433700-140547925  
2026-08-19T05:17:19.986450Z 0 [Note] [MY-000000] [Galera] GCache::RingBuffer unused buffers scan… 0.0% (0/205103168 bytes) complete.  
2026-08-19T05:17:19.989060Z 0 [Note] [MY-000000] [Galera] GCache::RingBuffer unused buffers scan… 100.0% (205103168/205103168 bytes) complete.  
2026-08-19T05:17:19.989075Z 0 [Note] [MY-000000] [Galera] Recovering GCache ring buffer: found 2/114228 locked buffers  
2026-08-19T05:17:19.989079Z 0 [Note] [MY-000000] [Galera] Recovering GCache ring buffer: free space: 10532315504/10737418240  
2026-08-19T05:17:19.990004Z 0 [Note] [MY-000000] [Galera] Passing config to GCS: allocator.disk\_pages\_encryption = no; allocator.encryption\_cache\_page\_size = 32K; allocator.encryption\_cache\_size = 16777216; base\_dir = /var/lib/mysql/; base\_host = 10.42.209.202; base\_port = 4567; cert.log\_conflicts = no; cert.optimistic\_pa = no; debug = no; evs.auto\_evict = 0; evs.causal\_keepalive\_period = PT1S; evs.debug\_log\_mask = 0x1; evs.delay\_margin = PT1S; evs.delayed\_keep\_period = PT30S; evs.inactive\_check\_period = PT0.5S; evs.inactive\_timeout = PT15S; evs.info\_log\_mask = 0; evs.join\_retrans\_period = PT1S; evs.keepalive\_period = PT1S; evs.max\_install\_timeouts = 3; evs.send\_window = 10; evs.stats\_report\_period = PT1M; evs.suspect\_timeout = PT5S; evs.use\_aggregate = true; evs.user\_send\_window = 4; evs.version = 1; evs.view\_forget\_timeout = PT24H; gcache.dir = /var/lib/mysql/; gcache.encryption = no; gcache.encryption\_cache\_page\_size = 32K; gcache.encryption\_cache\_size = 16777216; gcache.freeze\_purge\_at\_seqno = -1; gcache.keep\_pages\_count = 0; gcache.keep\_pages\_size = 0; gcache.mem\_size = 0; gcache.name = galera.cache; gcache.page\_size = 128M; gcache.recover = yes; gcache.size = 10G; gcomm.thread\_prio = ; gcs.check\_appl\_proto = 1; gcs.fc\_auto\_evict\_threshold = 0.75; gcs.fc\_auto\_evict\_window = 0; gcs.fc\_debug = 0; gcs.fc\_factor = 1.0; gcs.fc\_limit = 100; gcs.fc\_master\_slave = no; gcs.fc\_single\_primary = no; gcs.max\_packet\_size = 64500; gcs.max\_throttle = 0.25; gcs.recv\_q\_hard\_limit = 9223372036854775807; gcs.recv\_q\_soft\_limit = 0.25; gcs.sync\_donor = no; gmcast.mcast\_ttl = 1; gmcast.peer\_timeout = PT3S; gmcast.segment = 0; gmcast.time\_wait = PT5S; gmcast.version = 0; pc.announce\_timeout = PT3S; pc.checksum = false; pc.ignore\_quorum = false; pc.ignore\_sb = false; pc.linger = PT20S; pc.npvo = false; pc.recovery = true; pc.version = 0; pc.wait\_prim = true; pc.wait\_prim\_timeout = PT30S; pc.wait\_restored\_prim\_timeout = PT0S; pc.weight = 1; protonet.backend = asio; protonet.version = 0; repl.causal\_read\_timeout = PT30S; repl.commit\_order = 3; repl.key\_format = FLAT8; repl.max\_ws\_size = 2147483647; repl.proto\_max = 11; socket.checksum = 2; socket.recv\_buf\_size = auto; socket.send\_buf\_size = auto; socket.ssl = YES; socket.ssl\_ca = /etc/mysql/ssl-internal/ca.crt; socket.ssl\_cert = /etc/mysql/ssl-internal/tls.crt; socket.ssl\_cipher = ; socket.ssl\_key = /etc/mysql/ssl-internal/tls.key; socket.ssl\_reload = 1;  
2026-08-19T05:17:20.004419Z 0 [Note] [MY-000000] [Galera] Service thread queue flushed.  
2026-08-19T05:17:20.004482Z 0 [Note] [MY-000000] [Galera] ####### Assign initial position for certification: 7cb3e8b0-7149-11f1-b8df-26bb63d99319:140547925, protocol version: -1  
2026-08-19T05:17:20.004576Z 0 [Note] [MY-000000] [WSREP] Starting replication  
2026-08-19T05:17:20.004592Z 0 [Note] [MY-000000] [Galera] Connecting with bootstrap option: 0  
2026-08-19T05:17:20.004598Z 0 [Note] [MY-000000] [Galera] Setting GCS initial position t

---

<div class="post-metadata">

**Author:** ![matthewb](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/matthewb/32/34_2.png) [@matthewb](https://forums.percona.com/u/matthewb)\
**Post date:** [August 20, 2026, 8:41pm UTC](https://forums.percona.com/t/pxc-ist-recovery-time-after-heavy-write-load-11-hours-of-continuous-writes/41187/3 "2026-08-20T20:41:17Z")

</div>

> 1. Is our understanding correct?

Yes, all 4 points are correct.

> 1. Is it expected that a node can recover very quickly after approximately 11 hours of continuous writes if IST is used?

Yes. Even days of writes.

> 1. Does the recovery time depend primarily on the amount of missing write-set data rather than the duration for which the node was offline?

It depends 100% on the amount of missing write-set data, and has absolutely nothing to do with time/duration.

> 1. Am I missing anything in my understanding of how IST recovery works?

Nope. All good.

> 1. Are there any recommended methods for accurately testing recovery time after a very large write workload?

Your method is pretty spot on.

Check out gcache.[freeze\_purge\_at\_seqno](https://docs.percona.com/percona-xtradb-cluster/8.4/wsrep-provider-index.html#gcachefreeze_purge_at_seqno) which gives the ability for gcache to grow indefinitely if needed in certain situations.

---

<div class="post-metadata">

**Author:** ![vagrawal](https://avatars.discourse-cdn.com/v4/letter/v/6a8cbe/32.png) [@vagrawal](https://forums.percona.com/u/vagrawal)\
**Post date:** [August 21, 2026, 10:24am UTC](https://forums.percona.com/t/pxc-ist-recovery-time-after-heavy-write-load-11-hours-of-continuous-writes/41187/4 "2026-08-21T10:24:39Z")

</div>

Thank you so much Sir!!
