Percona XtraDB 84 - Clone SST joiner aborts with “Operation not permitted” on Ubuntu 24.04: mv -n exits 1 under errexit when datadir already has *.pem

Project: PXC
Component: SST (clone)
Affects version: 8.4.10-10 (code path also present on 8.0 and trunk branches)
Severity: a node that previously belonged to the cluster cannot rejoin via clone SST

Summary

On Ubuntu 24.04, a node that has already been part of the cluster cannot rejoin through wsrep_sst_method=clone. The joiner script exits with status 1, and mysqld reports it as SST script aborted with error 1 (Operation not permitted). The “Operation not permitted” text is only strerror(1); no permission problem is involved.

The cause is in scripts/wsrep_sst_clone.sh:

  1. The script runs with set -o nounset -o errexit (line 97 on the 8.4 branch).
  2. The joiner cleans the datadir with find ... -regex $cpat -prune -o -exec rm -rfv {}. The default cpat contains .*\.pem$, so *.pem files in the datadir are preserved (lines 1023-1024).
  3. mysqld --initialize-insecure runs into a temporary datadir. It generates private_key.pem and public_key.pem, the RSA key pair used by caching_sha2_password.
  4. The temporary datadir is moved into place with mv -n "$tmp_datadir"/* "$WSREP_SST_OPT_DATA/" (line 1068).

In GNU coreutils 9.2 through 9.4, mv -n exits with a nonzero status when it skips an existing destination file, and 9.3 added the mv: not replacing '...' message. Coreutils 9.5 reverted to exiting 0. Ubuntu 24.04 LTS ships coreutils 9.4, so step 4 returns 1 and errexit aborts the SST.

A node’s first join succeeds because its datadir is empty, so no key files collide. Every later SST of the same node fails because its old key pair survives step 2.

Environment

  • OS: Ubuntu 24.04 LTS (noble), coreutils 9.4-3ubuntu6.x
  • Packages:
    percona-xtradb-cluster          1:8.4.10-10-1.noble
    percona-xtradb-cluster-server   1:8.4.10-10-1.noble
    percona-xtradb-cluster-common   1:8.4.10-10-1.noble
    percona-xtradb-cluster-client   1:8.4.10-10-1.noble
    
  • Topology: 3 PXC nodes + garbd, Galera traffic over SSL
  • Relevant configuration:
    [mysqld]
    wsrep_provider            = /usr/lib/galera4/libgalera_smm.so
    wsrep_cluster_address     = gcomm://172.18.0.2,172.18.0.3,172.18.0.5
    wsrep_sst_method          = clone
    wsrep_sst_allowed_methods = clone,xtrabackup-v2
    datadir                   = /var/lib/mysql/data
    [sst]
    ssl_ca   = /etc/<certs>/ca.pem
    ssl_cert = /etc/<certs>/server-cert.pem
    ssl_key  = /etc/<certs>/server-key.pem
    
    The TLS certificates are outside the datadir. The only *.pem files in the datadir are the auto-generated private_key.pem and public_key.pem.

Steps to reproduce

  1. Bootstrap a PXC 8.4 cluster on Ubuntu 24.04 with wsrep_sst_method=clone, and join a second node via clone SST. This first join succeeds.
  2. On the joined node, stop mysqld.
  3. Force a full SST without deleting the key files. For example, remove grastate.dat, or wipe the datadir except private_key.pem and public_key.pem. This matches what a failed or crashed node looks like.
  4. Start mysqld on that node.

Expected: the clone SST completes and the node reaches SYNCED.

Actual: the joiner script aborts right after initializing the temporary datadir, and mysqld terminates.

Joiner error log (excerpt)

[Note] [WSREP-SST] DONOR SAY SST (yes|SST@61804391-b1ae-11f1-aad9-5e14c80868d9:54<EOF>)
[Note] [WSREP-SST] Cleaning Data directory /var/lib/mysql/data/
[Note] [WSREP-SST] Initializing data directory at /var/lib/mysql/data/tmp.fsKAPjWhRO
[Note] [WSREP-SST] mv: not replacing '/var/lib/mysql/data/private_key.pem'
[Note] [WSREP-SST] mv: not replacing '/var/lib/mysql/data/public_key.pem'
[Note] [WSREP-SST] Joiner cleanup. SST daemon PID:
[Note] [WSREP-SST] Joiner cleanup done.
[ERROR] [WSREP] Process completed with error: wsrep_sst_clone --role 'joiner' --address '172.18.0.5' --datadir '/var/lib/mysql/data/' ... : 1 (Operation not permitted)
[ERROR] [WSREP] Failed to read uuid:seqno from joiner script.
[ERROR] [WSREP] SST script aborted with error 1 (Operation not permitted)
[ERROR] [Galera] Application received wrong state:
        Received: 00000000-0000-0000-0000-000000000000
        Required: 61804391-b1ae-11f1-aad9-5e14c80868d9
[ERROR] [Galera] Application state transfer failed. This is unrecoverable condition, restart required.

Proposed fix

Replace mv -n with an explicit skip-if-exists loop. This keeps the current semantics (existing datadir files win) and does not depend on coreutils version:

     # Move initialized data directory structure to real datadir and cleanup
-    mv -n "$tmp_datadir"/* "$WSREP_SST_OPT_DATA/"
+    for f in "$tmp_datadir"/*; do [ -e "$WSREP_SST_OPT_DATA/${f##*/}" ] || mv "$f" "$WSREP_SST_OPT_DATA/"; done

I tested this with the reproducer against coreutils 8.32, 9.4 and 9.5, and with a real mysqld --initialize-insecure datadir on 9.4.

Alternatives considered:

  • mv --update=none: exits 0 on 9.4, but the --update= option only exists since coreutils 9.3, so it fails on RHEL 8/9 (coreutils 8.x).
  • mv -n ... || true: portable, but it would also hide real move failures such as a full disk or wrong permissions.

The same mv -n "$tmp_datadir"/* line is present on the 8.0 and trunk branches. On the 8.4 branch, wsrep_sst_xtrabackup-v2.sh contains no mv -n or cp -n.

A separate suggestion: the joiner could log the name of the failing command before exiting. Right now the only error-level message is error 1 (Operation not permitted), which points administrators toward permissions, AppArmor or SELinux instead of the actual cause.

Workarounds

Any one of these lets the node rejoin:

  1. Delete private_key.pem and public_key.pem from the datadir before starting the joiner. New keys are generated by the initialization step, as on a first join.
  2. Override cpat in the [sst] group with the default value minus .*\.pem$. This is only safe when the TLS certificates are stored outside the datadir.
  3. Use wsrep_sst_method=xtrabackup-v2.

References

I’ve opened the Jira Issue: https://perconadev.atlassian.net/browse/PXC-5330

Hi @moro15011 ,

Thanks, creating a Jira ticket is the correct approach in this case.