Est.

Postgres Logical Replication Slots and CDC Reliability

Logical replication slots can fill your disk if consumers fall behind or go idle.

Staff Writer · · 14 min read
Cover illustration for “Postgres Logical Replication Slots and CDC Reliability”
Change Data Capture · October 10, 2026 · 14 min read · 3,226 words

A logical replication slot is the durability boundary that makes Postgres change data capture possible. It is a durable object that lives on the primary database, and it carries two positions in the write-ahead log (WAL) that together tell Postgres how much history it has to preserve for a consumer that has not yet finished reading it.

The first position is restart_lsn, the oldest WAL the slot requires. Postgres will not recycle any WAL segment older than this point, no matter how much disk pressure builds up elsewhere on the system. The second is confirmed_flush_lsn, the most recent position the subscriber has acknowledged processing. WAL older than confirmed_flush_lsn is no longer eligible to be re-sent to that consumer, but restart_lsn governs when WAL gets reclaimed on disk. The slot's entire function reduces to one instruction given to the database engine: do not recycle WAL segments past this point, because a consumer still needs them.

Three things have to be true before a slot can exist. The server needs wal_level set to logical, because replica, the default, doesn't retain enough for logical decoding, and changing it takes a restart. A publication has to declare which tables' changes should flow through the replication stream, which matters for the publication and subscription model Postgres uses natively, though it is not a requirement of the slot object itself. And the slot has to be created, either explicitly through a function call like pg_create_logical_replication_slot(), or implicitly when a subscriber connects.

What the slot actually does is decode WAL into row-level logical changes, filterable by table, rather than the raw binary stream that physical replication copies byte-for-byte to build an identical copy of the cluster. This decoding layer is the surface that every CDC tool, regardless of vendor or architecture, ultimately reads from. The consumer's obligation in that relationship is simple to state and easy to get wrong in practice: read from the slot, process the changes, and periodically confirm how far it has gotten. That confirmation step, more than any other single operational detail in Postgres CDC, is responsible for the majority of production incidents tied to logical replication.

The slot's durability guarantee as its primary failure mode

The same property that makes a replication slot trustworthy is the property that makes it dangerous. Because the slot refuses to let WAL go until the consumer confirms its position, a slot that goes inactive, or one attached to a consumer that has fallen behind, accumulates WAL with no ceiling. Postgres does not clean this up automatically by default. PostgreSQL 18 introduced an opt-in setting, idle_replication_slot_timeout, that allows automatic invalidation of slots that have sat idle too long, but outside of that explicit configuration, the accumulation simply continues.

That accumulation does not scale with anything convenient. WAL grows at the rate of the database's own write activity, completely independent of whether a consumer is reading any of it. A busy primary with an inactive slot can fill disk in a matter of days or less, regardless of how light the actual CDC workload was meant to be. The terminal outcome is stark: if WAL fills the disk, Postgres can refuse new writes, halting the production database that the CDC pipeline was built to read from.

This failure mode is the most common way teams damage their own production Postgres instance in the course of standing up CDC, not a theoretical edge case reserved for unusual configurations. A proof-of-concept consumer gets spun up, tested, and torn down, but the slot it created is never dropped. WAL quietly piles up behind that orphaned slot for days until someone notices the disk is nearly full.

One incident illustrates how quickly this escalates. A Kafka Connect cluster had a configuration error that kept it from starting its consumer tasks. The cluster was running, but it was not consuming anything. The replication slot it had created stayed open, so WAL began piling up behind it, the primary's disk filled, and an entire e-commerce platform went offline. What makes slot bloat especially dangerous is that the failure appears on the database side, not the CDC side: a monitoring system watching only the health of the connector would have reported everything as green while the primary quietly ran out of disk.

A related trap occurs when a host runs several logical databases that see different traffic patterns. WAL is shared across the entire Postgres instance, but replication slots are scoped to a single database. A slot attached to a quiet or low-traffic database cannot advance on its own, and it will sit there retaining growing amounts of WAL even while the rest of the instance looks healthy.

Configuration choices can make this worse without anyone intending it. REPLICA IDENTITY FULL exists as a last resort for tables that lack a primary key or a suitable unique index, but it forces Postgres to write the entire old row on every update rather than just the changed columns. On a high-volume table, that setting alone can multiply WAL volume several times over at the same write rate, so an already fragile slot situation turns into a much faster path to disk exhaustion.

How long-running transactions stall a slot through logical decoding

A second accumulation pathway operates independently of anything the CDC consumer does. Postgres logical decoding cannot skip past an open transaction. If any transaction on the database remains open for an extended period, decoding stalls at the exact LSN where that transaction began, and it cannot move forward until the transaction closes.

Long-running transactions occur in production for mundane reasons. A schema migration on a large table can hold a transaction open for a long stretch. A batch job may open a transaction and hold it while it works through a large volume of records. An abandoned BEGIN left open in a forgotten psql session can do the same damage with none of the apparent justification.

The consequence is counterintuitive: WAL accumulates even though the CDC consumer itself may be reading perfectly well and keeping pace with everything the slot has sent it. The slot simply cannot advance past the starting point of that open transaction, so from the database's perspective, the consumer is falling behind no matter how fast it processes what it receives.

Schema changes introduce a related complication. DDL is not replicated through native Postgres logical replication, so changes like ALTER TABLE never flow through the slot. A team has to apply those changes manually on both the publisher and the subscriber, and the Postgres documentation on logical replication is specific about the order: add the column on the subscriber first, so that incoming changes referencing the new column do not error out mid-stream on the side that hasn't caught up yet. If a CDC connector hits a schema change it wasn't built to expect, it typically stops consuming altogether, and the slot keeps accumulating WAL behind it whether or not anyone has noticed the connector has stalled.

Certain tables sit outside what a slot can see. Unlogged tables bypass WAL entirely, so CDC has no mechanism to capture changes to them. Large objects do generate WAL, but logical replication does not support them, so they remain invisible to CDC as well. Connectors need to handle TOAST-related changes carefully, so they don't silently miss updates to large column values.

Why failover breaks the slot, the HA coupling problem

A logical replication slot's state only advances when a consumer connects and acknowledges data, and that state has never traveled through WAL itself. That means standbys in a high-availability cluster had no authoritative copy of a slot's progress to work from if the primary failed. For a standby to be eligible to carry a slot after a promotion, three conditions all have to hold at once: the slot has to be marked as synced on that standby, the slot's WAL position has to be consistent with the standby's own position (not too far ahead, not too far behind), and the slot has to be persistent rather than invalidated, meaning temporary is false and invalidation_reason is null.

That requirement creates a practical trap. If the CDC client hasn't connected in several hours, any standby that was recently added or recently restarted simply will not qualify. So you can't cleanly execute a controlled, planned promotion of the primary without breaking the CDC stream that depends on that slot.

The failure scenarios that follow from this coupling are specific. During a quiet period on the CDC side, logical slots on standbys can remain stuck in temporary status because their positions haven't reconciled with the primary's. If a forced failover happens while that's the case, those temporary slots are not failover-ready, the CDC stream breaks outright, and recovery requires reinitializing the connector and reloading a snapshot from scratch. A similar problem appears when a team replaces old standbys with freshly built ones using pg_basebackup: each new standby starts synchronizing slot metadata from a conservative starting point and will not be considered synchronized until the subscriber has genuinely advanced past it. The overall effect is that the slowest slot in the cluster sets the pace for how far the whole system can move without someone stepping in manually.

Postgres 17 introduced logical replication failover slots specifically to address this, allowing slot state to synchronize to promotion candidates ahead of time. But eligibility still deliberately waits for the subscriber to advance the slot, precisely to avoid handing a promoted standby a broken stream. A subscriber that is slow or briefly offline still leaves the cluster with no eligible replica to promote. On managed Postgres platforms, production failover continuity depends on setting both sync_replication_slots and hot_standby_feedback to on, and on explicitly listing the relevant replication slot names in the failover configuration before any switchover takes place.

The contrast with MySQL's replication model matters as structural context, not as an argument for abandoning Postgres. MySQL's binlog, built around global transaction identifiers (GTIDs), does not create this same coupling. A CDC connector can point at any suitable replica and resume reading from its last committed GTID set, and a consumer that has fallen behind cannot stall a switchover the way it can in Postgres. Success there depends on binlog retention settings, not on how recently a connector happened to poll. Postgres teams that understand this difference can configure around the coupling deliberately, rather than discovering it during an actual failover event.

What "exactly-once" means across the full CDC chain

Many teams assume that if they read from a Postgres replication slot, they get exactly-once delivery by default, but they do not. Reliable Postgres CDC depends on four independent durability boundaries, and a guarantee at one of them does not substitute for a guarantee at any other: Postgres has to retain the source history, which is the slot's job; the connector has to resume from a valid offset; the transport has to actually deliver the records it carries; and the sink has to apply replayed records without repeating any effects that would cause harm.

The checkpoint mechanism inside Postgres is where this gets concrete. A logical slot normally emits each change exactly once during regular operation, but its position is only persisted to disk at checkpoint time. If the database crashes between checkpoints, the slot can revert to an earlier position on restart, and recent changes get sent to the consumer a second time. Preventing that duplication from causing real damage is the client's responsibility, not something the slot itself solves. Postgres 17's failover slots do not change this either: an event can already be committed to the sink while the connector's recorded offset is still behind it, and that event can be replayed again during recovery regardless of how well the failover slot synchronized.

The consequences of replay vary enormously depending on what the sink actually does with a duplicate. An upsert into a warehouse table is typically harmless when it runs twice, since the second write just overwrites the first with the same values. Sending a second email, charging a second payment, firing a second webhook, or reserving inventory a second time each produces a real-world effect, making it a far more serious problem than a database no-op.

The practical target for a production CDC pipeline is at-least-once delivery paired with a sink that is idempotent or capable of deduplicating incoming records, backed by a recovery runbook that has actually been tested. Tolerating occasional duplicates and handling them deliberately at the sink is a safer design than advancing a checkpoint before the target has durably accepted the payload.

The configuration and monitoring practices that keep slots safe in production

Every default in a fresh Postgres installation is unsafe for a slot running in production, so running one well requires active configuration choices, continuous monitoring, and disciplined cleanup.

The single most important parameter is max_slot_wal_keep_size: it defaults to -1, so a slot can retain an unlimited amount of WAL, which directly enables the disk-exhaustion failure mode described earlier. Setting an actual limit caps how much WAL a slot can hold onto at checkpoint time. If a slot falls too far behind that cap, the WAL it needs gets removed and the slot is invalidated, breaking replication but protecting the database from running out of disk. Production guidance recommends starting above 4 GB and tuning that figure against replication lag, the database's change rate, and available disk space, and leaving the parameter at its unlimited default is not an acceptable production configuration.

A handful of supporting parameters round out a safe setup. max_replication_slots should be set to twice the number of replicas or subscribers, with one slot per replica and the rest held in reserve for failover operations. max_wal_senders needs to be at least max_replication_slots plus the number of physical replicas, and setting it lower breaks replication. sync_replication_slots should be on to support synchronization to standbys for failover continuity, and hot_standby_feedback should be on to prevent query conflicts during replication. logical_slot_sync_timeout bounds how long failover will wait for slot synchronization, defaulting to 300 seconds, and should be tuned against the failover window the team is actually willing to tolerate.

The multi-database host fix is direct: writing a periodic record into a heartbeat table on the otherwise idle database forces its slot to advance even when there is no genuine application traffic to drive it forward.

Lifecycle hygiene matters as much as configuration. Unused slots should be dropped deliberately and immediately, because a forgotten proof-of-concept slot is still the single most common cause of runaway WAL growth in the wild. If you query pg_replication_slots and join pg_wal_lsn_diff against the current WAL position and each slot's restart_lsn, you can see how much WAL every slot is retaining, and pg_drop_replication_slot removes a slot cleanly once you no longer need it. Retained WAL per slot deserves a permanent place on an operational dashboard, not just a one-time check during initial setup.

Alerting thresholds should be concrete: flag any slot that has been inactive for more than 30 minutes, and flag any slot retaining more than 10 to 20 GB of WAL. Monitoring needs to track several things independently rather than treating CDC health as a single metric: source replication slots, connector offsets, transport lag, sink lag, retained WAL size, and the rate of duplicate events reaching the sink.

On managed platforms like Amazon RDS, Cloud SQL, and Azure Database for PostgreSQL, wal_level and slot-related limits are usually set through a parameter group rather than edited directly in postgresql.conf, and some managed tiers cap how many slots can be active concurrently. Those limits vary by provider and by service tier, so they need to be checked against the specific deployment before launch. On the schema side, adding new columns on the subscriber before applying the corresponding ALTER TABLE on the source remains the practice that keeps incoming changes from erroring out mid-stream.

CDC tool architecture and operational burden

Every failure mode covered so far, WAL accumulation, long-running transaction stalls, failover coupling, schema evolution handling, lands differently depending on which CDC architecture a team chooses, because the architecture determines how much of that operational surface the team has to own directly.

Using pg_recvlogical and building custom consumers directly against the WAL stream carries zero third-party dependencies, but it means the team has to hand-write state management, failover slot handling, retry logic, and deduplication from scratch. That approach fits specialized environments where running a JVM or any additional runtime is explicitly off the table.

AWS DMS relies on replication slots to capture WAL changes from Postgres sources, so the operational risk tied to slot management still sits with whoever runs the primary database, not with DMS itself. DMS Serverless specifically does not support setting custom CDC start points, so any recovery scenario that requires replaying from a specific LSN has to use the provisioned instance instead. When a DMS pipeline falls behind, WAL still accumulates on the source just as described earlier, because the underlying slot mechanics haven't changed.

Apache Flink CDC fits teams that need CDC combined with genuine stateful stream processing, transformations, joins, or parallel incremental snapshots. Running it means operating a Flink job and its checkpoints directly, and keeping connector and runtime versions compatible across upgrades.

A managed CDC service earns its value by absorbing exactly the failure modes covered in every section above: slot lifecycle, WAL accumulation, failover reconnection, and schema evolution handling all become the provider's responsibility, leaving the data team responsible for configuring the service rather than managing the internals of Postgres replication directly. So this tradeoff matters most for teams powering real-time analytics, AI agents, or customer-facing dashboards, where you can't afford stale data and building slot management in-house doesn't pay off for what the use case needs. The right question for any team isn't which architecture is technically superior in the abstract, but which of these failure modes the team can realistically operate, monitor, and recover from given its own size and expertise.

Slot reliability for real-time AI and analytics pipelines

As AI agents, real-time dashboards, and A/B testing frameworks draw increasingly on data flowing out of warehouse destinations, tolerance for pipeline unreliability keeps shrinking. If a slot accumulates WAL for hours before anyone notices, an operations problem becomes a product problem, because the systems downstream are making decisions on what that pipeline delivers.

AI agents and automated decision systems produce unreliable outputs when they act on stale data, and CDC replication latency sets the gap between when an operational event actually occurs and when an AI system can act on it. Time-sensitive use cases like fraud detection, inventory management, and customer-facing recommendation systems need pipelines that recover from slot failures within minutes rather than hours, so the monitoring thresholds and recovery runbooks described earlier become non-negotiable requirements, not aspirational best practices.

A replayed event that triggers a downstream action, a recommendation sent twice, an alert fired twice, a transaction processed twice, is a visible error that a user or customer experiences directly. Compliance environments add another layer of consequence: in HIPAA-regulated healthcare AI systems or SOC 2 environments, replication infrastructure has to function as a passthrough that never stores customer data on its own, and any pipeline gap caused by slot failure creates audit exposure on top of the underlying data quality problem.

Mastering slot configuration, monitoring, and recovery is what lets a CDC pipeline serve as a reliable foundation for the systems built on top of it.

Sources

  1. Logical replication and Change Data Capture (CDC) - PlanetScale

More in Change Data Capture