# Mongod-failure-after-restore

**URL:** <https://forums.percona.com/t/mongod-failure-after-restore/39588>\
**Category:** Percona Backup for MongoDB\
**Created:** [October 21, 2025, 8:24am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588 "2025-10-21T08:24:43Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [October 21, 2025, 8:24am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/1 "2025-10-21T08:24:43Z")

</div>

Hello Ivan,

Just a follow-up on this ticket. [Mongod-failure-after-restore - #2 by Ivan\_Groenewold](https://forums.percona.com/t/mongod-failure-after-restore/39528/2)

I attempted to perform a restore on the secondary instance despite the Mongo service failure. The PBM status has been mentioned above.

Could you please share the documentation or Linux commands for performing the restore? It would be very helpful.

Thank you.

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [October 21, 2025, 8:39am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/2 "2025-10-21T08:39:45Z")

</div>

I believe you need to follow the initial sync procedure as @Ivan_Groenewold mentioned. Here’s the guide:

> **[Resync a Member of a Self-Managed Replica Set - Database Manual - MongoDB Docs](https://www.mongodb.com/docs/manual/tutorial/resync-replica-set-member/#std-label-resync-replica-member)**
>
> Resync a stale replica set member by removing its data and performing an initial sync or by copying data files from another member.

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [October 24, 2025, 5:39am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/3 "2025-10-24T05:39:46Z")

</div>

Hi [radoslaw.szulgo](https://forums.percona.com/t/mongod-failure-after-restore/39588/2),

we have tried initial sync procedure also given output here – \> [Mongod-failure-after-restore](https://forums.percona.com/t/mongod-failure-after-restore/39528)

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [October 24, 2025, 7:05am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/4 "2025-10-24T07:05:18Z")

</div>

Hi **[radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)**

We have tried the initial sync procedure and provided the output [Mongod-failure-after-restore](https://forums.percona.com/t/mongod-failure-after-restore/39528) . We are also using PBM for backup and restore operations.

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [October 27, 2025, 8:54am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/5 "2025-10-27T08:54:23Z")

</div>

can you please provide some more information on what doesn’t work for you? Any error messages? Without that it’s really hard to help in your case.

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [October 27, 2025, 9:06am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/6 "2025-10-27T09:06:05Z")

</div>

We are using a cloud infrastructure running on Ubuntu with three servers: one primary and two secondary. We are running Percona Backup for MongoDB (PBM) version v2.10. For backups, we use Amazon S3 as the storage destination.

However, when we attempt to restore, the mongod service fails even though the restoration process is reported as successful. After this, all three mongod services fail. I am trying to restore the data only on a secondary server.

Command used for backup:  
pbm backup --type=physical

Command used for restore:  
pbm restore --time=“2025-09-24T13:30:00Z”

# PBM status after backup:

pbm status  
Cluster:

poc:

- private\_ip:27018 [S]: pbm-agent [v2.10.0] OK

- private\_ip:27018 [S]: pbm-agent [v2.10.0] OK

- private\_ip:27018 [P]: pbm-agent [v2.10.0] OK

# PITR incremental backup:

Status [ON]  
Running members: poc/private\_ip:27018

# Currently running:

(none)

# Backups:

S3 ap-south-1 s3://mongo-percona-bk  
Snapshots:  
2025-10-01T12:13:21Z 33.59GB success [restore\_to\_time: 2025-10-01T12:13:24]  
PITR chunks [11.10MB]:  
2025-10-01T12:13:25 - 2025-10-02T02:53:23

# PBM status after restore:

pbm status  
Cluster:

poc:

- private\_ip:27018 : pbm-agent [NOT FOUND]

- private\_ip:27018 : pbm-agent [NOT FOUND]

- private\_ip:27018 : pbm-agent [NOT FOUND]

# PITR incremental backup:

Status [OFF]

# Currently running:

(none)

# Backups:

S3 ap-south-1 s3://mongo-percona-bk  
(none)

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [October 27, 2025, 9:44am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/7 "2025-10-27T09:44:14Z")

</div>

> [@hemanth](#):
>
> However, when we attempt to restore, the mongod service fails even though the restoration process is reported as successful. After this, all three mongod services fail. I am trying to restore the data only on a secondary server.

Thanks for the explanation and more context - this definitely helps. Why do you want to use PBM instead of initial sync to “rebuild/recover” 1/3 nodes? As @Ivan_Groenewold already wrote in the other topic - PBM is used to recover all nodes rather than a specific one.

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [October 27, 2025, 10:16am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/8 "2025-10-27T10:16:18Z")

</div>

We are taking incremental backups through S3. When we initiate the backup from the secondary server, it works successfully. However, when we try to restore the data to the secondary server, all mongod processes on both the primary and secondary servers fail. Additionally, we are unable to perform this backup and restore activity using the Initial Sync process.

Through the Initial Sync process, we cannot take or restore database backups via S3, as it is a MongoDB replication-level rebuild mechanism, not a PBM-based backup/restore method.

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [October 27, 2025, 10:21am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/9 "2025-10-27T10:21:22Z")

</div>

Right, you cannot do “backup” for initial-sync… but why would you?

Can we take one step back and clarify what’s the use case you’re trying to implement? Why do you want to restore just 1 out of 3 nodes? Something is broken or that’s a “recovery” procedure for any node? Something else?

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [October 27, 2025, 11:23am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/10 "2025-10-27T11:23:10Z")

</div>

Right now, we are performing a POC with the goal of taking a full data backup using PBM. After that, we plan to enable incremental PITR to take hourly backups. During the POC, we successfully pushed the data to S3 through PBM. However, when we tried to restore the same data after enabling PITR, the data restoration completed successfully, but all mongod processes on both the primary and secondary servers failed. We have also configured a replica set between the primary and secondary servers to ensure data synchronization.

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [October 29, 2025, 9:36am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/11 "2025-10-29T09:36:04Z")

</div>

Can you provide logs from at least 1 server? What error is shown there?

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [October 30, 2025, 12:12pm UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/12 "2025-10-30T12:12:59Z")

</div>

The procedure I followed for the restore was: I removed the MongoDB data directory, but after doing so the mongod service failed to start. I started it again and then executed the restore command pbm restore 2025-10-30T07:01:38Z. A few minutes later, the mongod service went inactive. I’ve also attached the last few logs.

2025-10-30T05:51:17.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] download stat: buf 536870912, arena 268435456, span 33554432, spanNum 8, cc 2, [{2 0} {2 0}]  
2025-10-30T05:51:17.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] preparing data  
2025-10-30T05:51:35.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] oplogTruncateAfterPoint: {1761737952 1}  
2025-10-30T05:51:37.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] recovering oplog as standalone  
2025-10-30T05:51:54.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] clean-up and reset replicaset config  
2025-10-30T05:52:06.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] uploading “.pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/node.10.200.10.98:27018.hb” [size hint: 10 (10.00B); part size: 10485760 (10.00MB)]  
2025-10-30T05:52:06.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] uploading “.pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/rs.hb” [size hint: 10 (10.00B); part size: 10485760 (10.00MB)]  
2025-10-30T05:52:06.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] uploading “.pbm.restore/2025-10-30T05:38:05.927063022Z/cluster.hb” [size hint: 10 (10.00B); part size: 10485760 (10.00MB)]  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmAgents’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmBackups’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmRestores’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmCmd’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmPITRChunks’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmPITR’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmOpLog’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmLockOp’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmLock’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmLock’  
2025-10-30T05:52:12.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] dropping ‘admin.pbmLog’  
2025-10-30T05:52:14.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] restore on node succeed  
2025-10-30T05:52:14.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] moving to state done  
2025-10-30T05:52:14.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] uploading “.pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/node.10.200.10.98:27018.done” [size hint: 10 (10.00B); part size: 10485760 (10.00MB)]  
2025-10-30T05:52:14.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] waiting for `done` status in rs map[.pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/node.10.200.10.104:27018:{} .pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/node.10.200.10.94:27018:{} .pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/node.10.200.10.98:27018:{}]  
2025-10-30T05:52:19.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] uploading “.pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/rs.done” [size hint: 10 (10.00B); part size: 10485760 (10.00MB)]  
2025-10-30T05:52:19.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] waiting for shards map[.pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/rs:{}]  
2025-10-30T05:52:24.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] uploading “.pbm.restore/2025-10-30T05:38:05.927063022Z/cluster.done” [size hint: 10 (10.00B); part size: 10485760 (10.00MB)]  
2025-10-30T05:52:24.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] waiting for cluster  
2025-10-30T05:52:29.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] converged to state done  
2025-10-30T05:52:29.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] uploading “.pbm.restore/2025-10-30T05:38:05.927063022Z/rs.poc/stat.10.200.10.98:27018” [size hint: 73 (73.00B); part size: 10485760 (10.00MB)]  
2025-10-30T05:52:29.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] writing restore meta  
2025-10-30T05:52:29.000+0000 W [restore/2025-10-30T05:38:05.927063022Z] meta `.pbm.restore/2025-10-30T05:38:05.927063022Z.json` already exists, trying write done status with ‘’  
2025-10-30T05:52:29.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] rm tmp conf  
2025-10-30T05:52:29.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] wait for cluster status  
2025-10-30T05:52:34.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] no cleanup strategy to apply  
2025-10-30T05:52:34.000+0000 I [restore/2025-10-30T05:38:05.927063022Z] recovery successfully finished  
2025-10-30T05:52:34.000+0000 I change stream was closed  
2025-10-30T05:52:34.000+0000 D [restore/2025-10-30T05:38:05.927063022Z] hearbeats stopped  
2025-10-30T05:52:34.000+0000 I Exit:  
2025-10-30T05:52:34.000+0000 D [agentCheckup] deleting agent status  
2025-10-30T05:53:05.000+0000 E Exit: connect to PBM: create mongo connection: ping: server selection error: server selection timeout, current topology: { Type: Unknown, Servers: [{ Addr: 127.0.0.1:27018, Type: Unknown, Last error: dial tcp 127.0.0.1:27018: connect: connection refused },] }  
2025-10-30T05:53:35.000+0000 E Exit: connect to PBM: create mongo connection: ping: server selection error: server selection timeout, current topology: { Type: Unknown, Servers: [{ Addr: 127.0.0.1:27018, Type: Unknown, Last error: dial tcp 127.0.0.1:27018: connect: connection refused },] }

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [October 30, 2025, 12:18pm UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/13 "2025-10-30T12:18:18Z")

</div>

This states that the PBM agent connection to the PSMDB was unexpectedly interrupted. Can you please provide logs from the server so we can see what happened on the server side at that time?

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [November 7, 2025, 1:02pm UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/14 "2025-11-07T13:02:44Z")

</div>

which logs require syslog?

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [November 12, 2025, 11:49am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/15 "2025-11-12T11:49:27Z")

</div>

I don’t understand the question. Sorry…

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [November 12, 2025, 11:54am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/16 "2025-11-12T11:54:37Z")

</div>

You have asked for server logs we have many logs, so could you please specify which ones are required?

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [November 12, 2025, 12:58pm UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/17 "2025-11-12T12:58:06Z")

</div>

Ah, that’s clear now. I asked precisely about Percona Server for MongoDB logs:

By default under: /var/log/mongodb/server1.log\*

---

<div class="post-metadata">

**Author:** ![hemanth](https://avatars.discourse-cdn.com/v4/letter/h/7993a0/32.png) [@hemanth](https://forums.percona.com/u/hemanth)\
**Post date:** [November 13, 2025, 4:34am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/18 "2025-11-13T04:34:49Z")

</div>

Below are the MongoDB logs taken during the restoration.

[mongod.log](https://forums.percona.com/uploads/short-url/qKnZOUvpZxqkdzmM6NJJfq8S1vH.log) (60.9 KB)

---

<div class="post-metadata">

**Author:** ![radoslaw.szulgo](https://sea1.discourse-cdn.com/flex019/user_avatar/forums.percona.com/radoslaw.szulgo/32/17315_2.png) [@radoslaw.szulgo](https://forums.percona.com/u/radoslaw.szulgo)\
**Post date:** [November 20, 2025, 9:01am UTC](https://forums.percona.com/t/mongod-failure-after-restore/39588/19 "2025-11-20T09:01:07Z")

</div>

It’s not crystal clear from the logs, but most likely you try to run PBM restore for the cluster (destination) that is not healthy. Please prepare a new cluster to restore to and try again.
