Replication Slot WAL Retention

Inspect WAL retained by slots and consumer state. Continuously monitor disk risk to identify abandoned replication slots.

2026-09-23 · 2 min read

Why this problem matters

Inactive or lagging physical and logical slots can cause required WAL files to be retained. An unused slot may create a quietly growing risk until disk capacity is exhausted.

How to diagnose it

Review slot type, active state, restart_lsn, and safe WAL size information together. Inactivity alone does not prove abandonment because a consumer may be temporarily offline.

The query reads current slot state and WAL safety fields available in modern supported versions. safe_wal_size may be null depending on configuration and slot state.

SELECT slot_name, slot_type, database, active, active_pid,
       restart_lsn, confirmed_flush_lsn, wal_status, safe_wal_size
FROM pg_replication_slots
ORDER BY active, slot_name;

A safe solution approach

Verify each slot's owner and consumer through inventory, then repair connectivity or consumption. Dropping a slot requires separate approval because of data and resynchronization consequences.

Why continuous monitoring matters

moon helps continuously monitor slot retention and disk trends with low overhead. Slack, PagerDuty, or webhook alerts expose declining safe capacity.

Frequently asked questions

Should every inactive replication slot be dropped?

No, a consumer under maintenance or temporarily offline may still require its slot. Make no change before verifying ownership and resynchronization impact.

How does moon help with this problem?

moon continuously observes SQL Server, PostgreSQL, and MongoDB signals, helping teams evaluate the problem as a trend instead of relying on a one-time check.