September 30, 2026

Deep Dive into PostgreSQL Logical Replication: Mastering max_sync_workers_per_subscription for Efficient Migrations

deep-dive-into-postgresql-logical-replication-mastering-max_sync_workers_per_subscription-for-efficient-migrations

deep-dive-into-postgresql-logical-replication-mastering-max_sync_workers_per_subscription-for-efficient-migrations

Database administrators and systems architects managing PostgreSQL environments are often tempted to tweak configuration parameters to accelerate initial data synchronization during logical replication migrations. Among these parameters, max_sync_workers_per_subscription is frequently misunderstood.

Far from being a speed booster for individual large tables, this setting controls table-level parallelism. Understanding its underlying mechanics, resource costs, failure loops, and operational traps is essential for anyone planning a smooth zero-downtime database migration or managing complex multi-table replication clusters.


Main Facts

At its core, max_sync_workers_per_subscription determines the maximum number of tables a single logical replication subscription can copy simultaneously during its initial synchronization phase.

The configuration parameter operates under specific technical parameters:

  • Default Value: 2
  • Configuration Context: sighup (meaning it can be reloaded dynamically without restarting the database server)
  • Valid Range: 0 to 262143
  • Origin: Introduced alongside logical replication in PostgreSQL 10.

Crucially, the parameter does not increase the speed of copying a single table. Each table synchronization worker is strictly dedicated to copying exactly one table from start to finish. Setting this parameter to 2, 8, or 200 will not alter the time it takes for your single largest table to copy.

Instead, it introduces parallelism across different tables. A subscription’s initial copy phase cannot officially finish until the largest, slowest-moving table synchronization worker completes its job.

The workers themselves are provisioned from a global pool managed by max_logical_replication_workers. Furthermore, this concurrency limit is evaluated on a per-subscription basis. If three distinct subscriptions attempt to initialize simultaneously, they will demand three times the allocated worker slots, heightening the risk of running out of logical replication worker slots if global limits are misconfigured.


Chronology and Dynamic Behavior

The lifecycle of the max_sync_workers_per_subscription parameter and its operational behavior during a migration follow a precise architectural timeline.

1. Initialization and The Apply Worker Loop

When a subscription is created, the primary apply worker inspects the system catalogs—specifically walking pg_subscription_rel in physical order. It evaluates tables marked as not ready. As long as free worker slots are available within the per-subscription and global pools, the apply worker launches individual table synchronization workers.

Because the traversal follows physical storage order (determined by the initial row insertion order during CREATE SUBSCRIPTION), synchronization does not inherently prioritize tables by name or file size.

2. Mid-Migration Reloads

One of the most powerful administrative features of this parameter is its sighup context. "Reload really does mean reload." The main loop of the apply worker continuously re-reads the configuration file.

If a database administrator initiates a migration with max_sync_workers_per_subscription = 1 and subsequently issues a pg_reload_conf() after raising the value to 3, the apply worker immediately responds. In testing environments (such as PostgreSQL 18.6), ramping up the parameter dynamically spawns additional synchronization workers mid-stream without requiring any intervention or restart on the subscription itself.

All Your GUCs in a Row: max_sync_workers_per_subscription

3. Error Handling and the Retry Trap

If a table synchronization worker encounters an error—such as a primary key constraint violation caused by pre-existing conflicting data on the subscriber—the worker exits with an error code.

The apply worker handles this by waiting for the duration specified by wal_retrieve_retry_interval and then restarting the entire table copy from scratch. Every retry drops the old replication slot on the publisher, creates a new one, and re-executes the entire COPY command.

Historically, without safety configurations like disable_on_error = true (available in PostgreSQL 15 and later), a single failing table traps the system in an endless retry loop. This halves the available parallelism for the duration of the migration, leaving administrators blind to the issue unless they actively monitor pg_stat_subscription_stats.sync_error_count and system log files.


Supporting Data and Resource Costs

Running table synchronization workers is not free. Every worker consumes distinct computational and storage resources on both the publisher and the subscriber nodes.

Resource Footprint on the Subscriber

On the receiving end (subscriber), each active synchronization worker consumes:

  • One worker slot from the global max_logical_replication_workers pool.
  • One replication origin slot (tracked via max_active_replication_origins in PostgreSQL 18+, or max_replication_slots in earlier versions).

Resource Footprint on the Publisher

The resource burden is even heavier on the source (publisher) database. For every active table synchronization worker, the publisher must allocate:

  1. A dedicated WAL sender process, which counts against the max_wal_senders limit.
  2. A permanent replication slot named following the convention pg_<subscription oid>_sync_<table oid>_<system identifier>.
  3. An open REPEATABLE READ transaction that stays active for the entire duration of the COPY operation.

The presence of this long-lived transaction introduces a critical maintenance hazard via the transaction’s backend_xmin horizon. For as long as the synchronization worker’s COPY is running, the publisher’s background VACUUM process cannot remove any dead rows created after that transaction’s snapshot was established—affecting every table in the database, not just the table currently being copied.

-- Checking active WAL senders and their transaction horizons on the publisher
SELECT application_name, backend_xmin, left(query, 40) AS query
FROM pg_stat_activity WHERE backend_type = 'walsender';

Furthermore, once a table’s initial COPY phase concludes, the synchronization worker enters a catch-up phase. The leader apply worker pauses—entering the LogicalSyncStateChange wait event—and waits passively while the sync worker decodes the WAL stream to catch the table up to the current LSN. On busy production systems with high write volumes, this post-copy catch-up can trigger unexpected, multi-second production stalls.


Official Recommendations and Future-Proofing

Database engineers must carefully weigh the performance benefits of parallelism against the hardware constraints of their underlying infrastructure.

The Danger of Extremes

  • Setting the Value to Zero (0): This is effectively a configuration trap. While CREATE SUBSCRIPTION will succeed, the apply worker will never request synchronization workers. Consequently, every table will remain indefinitely stuck in the 'i' (init) state, incoming changes will be silently discarded, and no error logs will be generated because technically, nothing is failing.
  • Over-Allocating Workers: Setting max_sync_workers_per_subscription to excessively high numbers (e.g., 20 or 50) on modest hardware leads to diminishing returns. Each parallel copy executes a heavy COPY ... TO STDOUT on the publisher core while simultaneously driving a resource-intensive COPY FROM statement with heavy index maintenance on the subscriber core. When CPU cores and disk I/O channels are saturated, high parallelism only introduces lock contention and context-switching overhead without reducing total migration time.

Sequence Synchronization in PostgreSQL 19

PostgreSQL 19 introduces native sequence synchronization, managed by a dedicated worker per subscription that coordinates all sequence states. This worker also draws from the max_sync_workers_per_subscription pool.

While official documentation notes that "one additional worker is also needed for sequence synchronization," testing shows that this worker simply rotates through the existing per-subscription worker cap without requiring structural increases to the parameter value.

Best Practices for Database Administrators

  1. Production Steady-State: Keep max_sync_workers_per_subscription at the default value of 2 for normal, long-running operational replicas. This prevents sudden administrative actions (such as running ALTER SUBSCRIPTION ... REFRESH PUBLICATION during peak hours) from unexpectedly flooding production with dozens of parallel heavy table copies.
  2. Migration Tuning: During major database migrations or initial bulk loads, temporarily scale the parameter up. Align the value with the actual capacity of your CPU cores and storage subsystems across both the publisher and subscriber.
  3. Post-Migration Cleanup: Always reset the parameter to safe defaults once the heavy initial synchronization phase completes.
  4. Enable Error Protections: Ensure disable_on_error = true is enabled on critical subscriptions to prevent silent, infinite retry loops from consuming system resources when schema or data anomalies occur.