Percona backup getting timeout for Mongodb

Hi We are using percona backup version as 2.15 and mongodb version as 8.0 and while taking backup through docker we are getting error as follows

docker run --rm -v /etc/mongo/ssl/ca-chainv4.cert.pem:/ca-chainv4.cert.pem:ro percona/percona-backup-mongodb:2.15.0 pbm backup --mongodb-uri=“mongodb://pbmuser:UEK1eYDLPB6fsV5YpxQ3u5vy@cfgsvr1.snb.internal:27019,cfgsvr2.snb.internal:27019,cfgsvr3.snb.internal:27019/?authSource=admin&replicaSet=cfgRS&tls=true&tlsCAFile=/ca-chainv4.cert.pem”

  • ‘[’ pbm = pbm-agent ‘]’

  • exec pbm backup ‘–mongodb-uri=mongodb://pbmuser:password@cfgsvr1.snb.internal:27019,cfgsvr2.snb.internal:27019,cfgsvr3.snb.internal:27019/?authSource=admin&replicaSet=cfgRS&tls=true&tlsCAFile=/ca-chainv4.cert.pem’
    Starting backup “2026-08-03T13:10:19Z”…
    Error: wait for backup status: couldn’t get response from all shards: convergeClusterWithTimeout: 33s: reached converge timeout

  • Backup on replicaset “cfgRS” in state: error: couldn’t get response from all shards: convergeClusterWithTimeout: 33s: reached converge timeout

We have currently 3300+ collections and getting timeout error as follows - journalctl -u pbm-agent --since “2026-08-03 08:07:30” --until “2026-08-03 08:08:20” -o cat | grep -A2 “timed out”
2026-08-03T08:08:11.000+0000 I [backup/2026-08-03T08:07:39Z] mark RS as error get namespaces size: collStats "audit.elements.versions.autoapitest0812": timed out while checking out a connection from connection pool: context deadline exceeded; total connections: 200, maxPoolSize: 200, idle connections: 0, wait duration: 29.999841437s:
2026-08-03T08:08:11.000+0000 D [backup/2026-08-03T08:07:39Z] set balancer on
2026-08-03T08:08:11.000+0000 E [backup/2026-08-03T08:07:39Z] backup: get namespaces size: collStats “audit.elements.versions.autoapitest0812”: timed out while checking out a connection from connection pool: context deadline exceeded; total connections: 200, maxPoolSize: 200, idle connections: 0, wait duration: 29.999841437s

2026-08-03T08:08:11.000+0000 D [backup/2026-08-03T08:07:39Z] releasing lock

We have following configuration set where we have maxPoolSize=200 to resolve the timeout error

cat /etc/sysconfig/pbm-agent
PBM_MONGODB_URI=“mongodb://pbmuser:password@shardsvr2.snb.internal:27018/?authSource=admin&tls=true&tlsCAFile=/data/mongo/config/ca-chainv4.cert.pem&maxPoolSize=200”

cat /usr/local/bin/pbm-wrapper
#!/bin/bash
export PBM_MONGODB_URI=“mongodb://pbmuser:password@cfgsvr1.snb.internal:27019,cfgsvr2.snb.internal:27019,cfgsvr3.snb.internal:27019/admin?replicaSet=cfgRS&tls=true&tlsCAFile=/data/mongo/config/ca-chainv4.cert.pem”
exec /usr/bin/pbm “$@”

We have disable pitr for now and here is following configuration

pbm-wrapper config
storage:
type: s3
s3:
region: us-east-1
forcePathStyle: true
bucket: snb-dev-int-s3-backups
prefix: mongo-shard-config
credentials: {}
maxUploadParts: 10000
storageClass: STANDARD
insecureSkipTLSVerify: false
pitr:
enabled: false
oplogSpanMin: 60
compression: zstd
backup:
oplogSpanMin: 0
compression: zstd
numParallelCollections: 1
restore:
numDownloadWorkers: 4
maxDownloadBufferMb: 256
downloadChunkMb: 64

I have increased timeout as well to 180 seconds in /etc/pbm-config.yaml but still the issue is not resolved…failure is always during collStats on one of the audit.elements.versions.* collections.

could you help us resolve this issue and what more information we can provide

Hello and thank you for describing your case.

Unfortunately, you are affected by the bug that has been present for longer period within PBM, including the latest release (v2.15): https://perconadev.atlassian.net/browse/PBM-1700

The better news is that it is already fixed in dev branch and will be released as part of v2.16 (approx. second half of September, 2026).

There’s no possible workaround, but you can try the following:

  • increase cpu/memory for the running mongodb cluster so that it can serve the number of namespaces on your server during backup preparation phase.
  • use selective backup and do the backup just for a smaller number of databases/collections, which are the most important to you (probably not suitable for your use case, but just to be aware of).
  • do a custom build from the source from the dev branch.

Even after upgrading to 2.16 version and as you told to wait till Sep 2026 end ..issue is still there and looks like the bug is closed and added in 2.16 version which is https://perconadev.atlassian.net/browse/PBM-1700 but I am still getting above error and after increasing timeout to 2m also same issue..could you check if the bug is added or not.

{"t":{"$date":"2026-09-30T04:13:07.850+00:00"},"s":"I", "c":"-", "id":4333222, "svc":"-", "ctx":"ReplicaSetMonitor-TaskExecutor","msg":"RSM received error response","attr":{"host":"shardsvr2.snb.internal:27018","error":"HostUnreachable: Timed out connecting to shardsvr2.snb.internal:27018 after 20000ms","replicaSet":"shardRS","response":{}}}
{"t":{"$date":"2026-09-30T04:13:07.850+00:00"},"s":"I", "c":"NETWORK", "id":4712102, "svc":"-", "ctx":"ReplicaSetMonitor-TaskExecutor","msg":"Host failed in replica set","attr":{"replicaSet":"shardRS","host":"shardsvr2.snb.internal:27018","error":{"code":6,"codeName":"HostUnreachable","errmsg":"Timed out connecting to shardsvr2.snb.internal:27018 after 20000ms"},"action":{"dropConnections":true,"requestImmediateCheck":true}}}
{"t":{"$date":"2026-09-30T04:13:18.350+00:00"},"s":"I", "c":"ASIO", "id":6496500, "svc":"-", "ctx":"ReplicaSetMonitor-TaskExecutor","msg":"Operation timed out while waiting to acquire connection","attr":{"requestId":186167,"durationMillis":10000}}
{"t":{"$date":"2026-09-30T04:13:19.230+00:00"},"s":"I", "c":"ASIO", "id":6496500, "svc":"-", "ctx":"ReplNetwork","msg":"Operation timed out while waiting to acquire connection","attr":{"requestId":186181,"durationMillis":2000}}
{"t":{"$date":"2026-09-30T04:13:28.350+00:00"},"s":"I", "c":"-", "id":4333222, "svc":"-", "ctx":"ReplicaSetMonitor-TaskExecutor","msg":"RSM received error response","attr":{"host":"shardsvr2.snb.internal:27018","error":"HostUnreachable: Timed out connecting to shardsvr2.snb.internal:27018 after 20000ms","replicaSet":"shardRS","response":{}}}
{"t":{"$date":"2026-09-30T04:13:28.350+00:00"},"s":"I", "c":"NETWORK", "id":4712102, "svc":"-", "ctx":"ReplicaSetMonitor-TaskExecutor","msg":"Host failed in replica set","attr":{"replicaSet":"shardRS","host":"shardsvr2.snb.internal:27018","error":{"code":6,"codeName":"HostUnreachable","errmsg":"Timed out connecting to shardsvr2.snb.internal:27018 after 20000ms"},"action":{"dropConnections":true,"requestImmediateCheck":false,"outcome":{"host":"shardsvr2.snb.internal:27018","success":false,"errorMessage":"HostUnreachable: Timed out connecting to shardsvr2.snb.internal:27018 after 20000ms"}}}}
{"t":{"$date":"2026-09-30T04:13:31.230+00:00"},"s":"I", "c":"ASIO", "id":6496500, "svc":"-", "ctx":"ReplNetwork","msg":"Operation timed out while waiting to acquire connection","attr":{"requestId":186194,"durationMillis":10000}}

This the status of percona backup

 pbm-wrapper status
Cluster:
========
shardRS:
  - shardsvr1.snb.internal:27018 [S]: pbm-agent [v2.16.0] OK
  - shardsvr2.snb.internal:27018 [S]: pbm-agent [v2.16.0] OK
  - shardsvr3.snb.internal:27018 [P]: pbm-agent [v2.16.0] OK
cfgRS:
  - cfgsvr1.snb.internal:27019 [S]: pbm-agent [v2.16.0] OK
  - cfgsvr2.snb.internal:27019 [P]: pbm-agent [v2.16.0] OK
  - cfgsvr3.snb.internal:27019 [S]: pbm-agent [v2.16.0] OK


PITR incremental backup:
========================
Status [ON]
! ERROR while running PITR backup: 2026-10-02T14:29:26.000+0000 E [shardRS/shardsvr2.snb.internal:27018] [pitr] init: catchup: get last backup: no backup found. full backup is required to start PITR; 2026-10-02T14:29:27.000+0000 E [cfgRS/cfgsvr3.snb.internal:27019] [pitr] init: catchup: get last backup: no backup found. full backup is required to start PITR

Currently running:
==================
(none)

Backups:
========
Main storage:
  Type:       S3
  Region:     us-east-1
  Path:       s3://snb-dev-int-s3-backups/mongo-shard-config
Snapshots:
  NAME                      SIZE        TYPE          PROFILE               SEL    BASE  RESTORE TIME         DURATION    STATUS
  ------------------------------------------------------------------------------------------------------------------------------
  2026-10-02T14:29:18Z      0.00B       logical                             no     no    2026-10-02T14:31:25  2m6s        error: couldn't get respons...
  2026-10-02T13:38:17Z      0.00B       logical                             no     no    2026-10-02T13:38:52  35s         error: couldn't get respons...
  2026-10-02T13:18:44Z      0.00B       logical                             no     no    2026-10-02T13:19:22  37s         error: couldn't get respons...

rpm -qa | grep percona

percona-backup-mongodb-2.16.0-1.amzn2023.x86_64

Hi, can you please attach PBM’s logs, e.g.

pbm logs -e backup/2026-10-02T14:29:18Z -s D -t 0

or even better, all logs from the agents that failed.

Here are the logs

First looks default value of timeout was 20000ms and it got failed first time on one shard node and we are getting error which I have pasted above and then we increase it to 2m just to check if it works or not but same issue

 pbm logs -e backup/2026-10-02T14:29:18Z -s D -t 0
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] init backup meta
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination list for cfgRS: [[cfgsvr3.snb.internal:27019 cfgsvr1.snb.internal:27019] [cfgsvr2.snb.internal:27019]]
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination list for shardRS: [[shardsvr1.snb.internal:27018 shardsvr2.snb.internal:27018] [shardsvr3.snb.internal:27018]]
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination shardRS, set candidates [shardsvr1.snb.internal:27018 shardsvr2.snb.internal:27018]
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination cfgRS, set candidates [cfgsvr3.snb.internal:27019 cfgsvr1.snb.internal:27019]
2026-10-02T14:29:19Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] backup started
2026-10-02T14:29:19Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] stopping balancer
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:19Z I [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] backup started
2026-10-02T14:29:19Z D [shardRS/shardsvr2.snb.internal:27018] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:19Z D [shardRS/shardsvr3.snb.internal:27018] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:20Z D [cfgRS/cfgsvr3.snb.internal:27019] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:23Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] waiting for balancer off
2026-10-02T14:29:24Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] balancer is disabled
2026-10-02T14:29:24Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] bcp nomination: cfgRS won by cfgsvr1.snb.internal:27019
2026-10-02T14:29:24Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] bcp nomination: shardRS won by shardsvr1.snb.internal:27018
2026-10-02T14:29:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] got sizes of 49 namespaces in 166ms
2026-10-02T14:31:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] mark RS as error `couldn't get response from all shards: convergeClusterWithTimeout: 2m0s: reached converge timeout`: <nil>
2026-10-02T14:31:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] mark backup as error `couldn't get response from all shards: convergeClusterWithTimeout: 2m0s: reached converge timeout`: <nil>
2026-10-02T14:31:25Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] set balancer on
2026-10-02T14:31:25Z E [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] backup: couldn't get response from all shards: convergeClusterWithTimeout: 2m0s: reached converge timeout
2026-10-02T14:31:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] backup: 2026-10-02T14:29:18Z, start: 2026-10-02T14:29:18Z, finish: 2026-10-02T14:31:25Z, duration: 2m7s
2026-10-02T14:31:25Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] releasing lock
2026-10-02T14:37:07Z I [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] got sizes of 7932 namespaces in 7m46.459s
2026-10-02T14:37:08Z I [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] mark RS as error `waiting for running: backup stuck, last beat ts: 1790951484`: <nil>
2026-10-02T14:37:08Z D [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] set balancer on
2026-10-02T14:37:08Z E [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] backup: waiting for running: backup stuck, last beat ts: 1790951484
2026-10-02T14:37:08Z D [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] releasing lock

this is error on cfg node

pbm logs -e backup/2026-10-02T14:29:18Z -s D -t 0
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] init backup meta
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination list for cfgRS: [[cfgsvr3.snb.internal:27019 cfgsvr1.snb.internal:27019] [cfgsvr2.snb.internal:27019]]
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination list for shardRS: [[shardsvr1.snb.internal:27018 shardsvr2.snb.internal:27018] [shardsvr3.snb.internal:27018]]
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination shardRS, set candidates [shardsvr1.snb.internal:27018 shardsvr2.snb.internal:27018]
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] nomination cfgRS, set candidates [cfgsvr3.snb.internal:27019 cfgsvr1.snb.internal:27019]
2026-10-02T14:29:19Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] backup started
2026-10-02T14:29:19Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] stopping balancer
2026-10-02T14:29:19Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:19Z I [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] backup started
2026-10-02T14:29:19Z D [shardRS/shardsvr2.snb.internal:27018] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:19Z D [shardRS/shardsvr3.snb.internal:27018] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:20Z D [cfgRS/cfgsvr3.snb.internal:27019] [backup/2026-10-02T14:29:18Z] skip after nomination, probably started by another node
2026-10-02T14:29:23Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] waiting for balancer off
2026-10-02T14:29:24Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] balancer is disabled
2026-10-02T14:29:24Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] bcp nomination: cfgRS won by cfgsvr1.snb.internal:27019
2026-10-02T14:29:24Z D [cfgRS/cfgsvr2.snb.internal:27019] [backup/2026-10-02T14:29:18Z] bcp nomination: shardRS won by shardsvr1.snb.internal:27018
2026-10-02T14:29:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] got sizes of 49 namespaces in 166ms
2026-10-02T14:31:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] mark RS as error `couldn't get response from all shards: convergeClusterWithTimeout: 2m0s: reached converge timeout`: <nil>
2026-10-02T14:31:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] mark backup as error `couldn't get response from all shards: convergeClusterWithTimeout: 2m0s: reached converge timeout`: <nil>
2026-10-02T14:31:25Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] set balancer on
2026-10-02T14:31:25Z E [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] backup: couldn't get response from all shards: convergeClusterWithTimeout: 2m0s: reached converge timeout
2026-10-02T14:31:25Z I [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] backup: 2026-10-02T14:29:18Z, start: 2026-10-02T14:29:18Z, finish: 2026-10-02T14:31:25Z, duration: 2m7s
2026-10-02T14:31:25Z D [cfgRS/cfgsvr1.snb.internal:27019] [backup/2026-10-02T14:29:18Z] releasing lock
2026-10-02T14:37:07Z I [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] got sizes of 7932 namespaces in 7m46.459s
2026-10-02T14:37:08Z I [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] mark RS as error `waiting for running: backup stuck, last beat ts: 1790951484`: <nil>
2026-10-02T14:37:08Z D [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] set balancer on
2026-10-02T14:37:08Z E [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] backup: waiting for running: backup stuck, last beat ts: 1790951484
2026-10-02T14:37:08Z D [shardRS/shardsvr1.snb.internal:27018] [backup/2026-10-02T14:29:18Z] releasing lock

please help us resolve this issue as soon as possible as we are stuck with backup and we cannot move ahead for deploying mongo sharded cluster in production.

Hi, the new issue you were getting is different from the original one.

shardsvr1 needed ~8 minutes to resolve namespaces sizes:

[backup/2026-10-02T14:29:18Z] got sizes of 7932 namespaces in 7m46.459s

You should increase backup.timeouts.startingStatus to something above that (e.g. 10 minutes). Please see in docs here: