Pt-online-schema-change 3.7.1 exhausts max_user_connections on MySQL 8.4 replica lag checks

Environment:

  • pt-online-schema-change 3.7.1-4 (percona-toolkit)
  • Primary: Percona Server 8.0.36-28
  • Replicas: MySQL 8.4
  • OS: Linux (x86_64)
  • Replica discovery: --recursion-method dsn=“D=,t=dsns”

  • Primary: Percona Server 8.0.36-28
  • Replicas: MySQL 8.4
  • OS: Linux (x86_64)
  • Replica discovery: --recursion-method dsn=“D=,t=dsns”

We upgraded from an earlier version of percona-toolkit to 3.7.1 specifically to support MySQL 8.4 replicas, which removed SHOW SLAVE STATUS in favour of SHOW REPLICA STATUS. The older version could not check lag on 8.4 replicas at all. The previous version was smoothly working but after upgrade we have encountered so many issues specifically related to replica status/lag checking using dsns table. I applied some workarounds but most recent one is most annoying.

Issue:
We are seeing max_user_connections exhaustion during execution against a table with 6 replicas in the DSN table:

Error while waiting for replica lag: Cannot connect to MySQL: DBI connect(‘…host=<replica_ip>;port=3306…’,‘schemachange’,…)
failed: User schemachange already has more than ‘max_user_connections’ active connections at /usr/bin/pt-online-schema-change line 2351.

We also observe:

  • Can’t create TCP/IP socket (24) — file descriptor exhaustion at default ulimit -n 1024
  • OS-level alert: “Too many TCP sockets” on the primary after raising ulimit
  • DNS resolution failures (EAI_AGAIN) under load when using hostnames in the DSN table

What we believe is happening?

It appears that 3.7.1 opens a new connection per replica per lag check cycle rather than reusing existing connections. With 6 replicas and frequent polling during a large table copy, the connection count grows rapidly and exhausts both the OS file descriptor limit and the MySQL max_user_connections per-user cap.

Questions:

  1. Was connection reuse behaviour changed in 3.7.1 for the SHOW REPLICA STATUS support? If so, is this intentional?
  2. Is there a known fix or configuration flag to limit connections opened per lag check cycle?
  3. Is a patch release planned, or is there a recommended workaround beyond removing the max_user_connections cap and trimming the DSN table to a single replica?

Any guidance appreciated.

Hello Aneel,

It seems you are hitting this bug,

https://perconadev.atlassian.net/browse/PT-2531

Its planned to fix in another version 3.7.2

So we suggest you to wait for the fix soon.

Regards,

Yunus Shaikh