Skip to content
back to writing
5 min readpostgresql · distributed-systems · backend-architecture

Why Postgres LISTEN/NOTIFY is a Scaling Trap for High-Throughput Systems

The simplicity of Postgres LISTEN/NOTIFY masks severe architectural risks. Discover why global locks and delivery guarantees make it unsuitable for high-volume eventing.

RG
Rahul Gupta
Senior Software Engineer
share

Postgres LISTEN/NOTIFY looks like a magic bullet when you are building your first real-time feature.

You have a table. You add a trigger. You call NOTIFY channel, 'payload'. Suddenly, your backend services are reacting to database changes in real-time without you having to manage a single external dependency like Kafka or RabbitMQ. It feels elegant. It feels lightweight.

Then you hit 500 events per second. Then you try to scale your consumer fleet to ten instances. Then you experience a network partition.

That is when the elegance turns into a production incident.

Most teams that adopt LISTEN/NOTIFY for core business logic discover that they haven't built an eventing system; they have built a distributed bottleneck that is tightly coupled to their primary database's availability and performance.

1. The "Global" problem in a local mechanism

The first thing you need to understand is that NOTIFY is not a distributed queue. It is a signaling mechanism built into the Postgres backend process.

When you execute a NOTIFY command, the payload is sent to the postmaster, which then broadcasts it to all connected backends that have issued a LISTEN on that channel. This sounds fine until you look at the internals.

In current stable versions like Postgres 16 or 17, the delivery of notifications is tied to the transaction lifecycle. A notification is only actually sent once the transaction that issued the NOTIFY commits. This is actually a good thing—it prevents "phantom notifications" where a consumer reacts to data that was eventually rolled back.

However, the mechanism for managing these pending notifications lives within the shared memory of the database engine. As your throughput increases, the overhead of managing the notification queue and the context switching required to signal all listening backends starts to bite.

You aren't just paying the cost of your SQL queries anymore; you are paying a tax on every single transaction that triggers a notification. In high-concurrency BFSI workloads where transaction latency is measured in microseconds, this tax becomes visible.

2. The reliability gap: No persistence, no retries

This is where the "footgun" becomes a landmine.

LISTEN/NOTIFY is fire-and-forget. It is a transient mechanism. If a consumer service is down, or if the network connection between your service and Postgres blips for even a second, those notifications are gone. Forever.

If you are using this for something trivial—like invalidating a local cache or updating a UI via WebSockets—you might get away with it. If you lose a message, the system eventually converges.

But if you are using LISTEN/NOTIFY to trigger downstream side effects—like sending a payment confirmation email, updating a ledger, or starting a fulfillment workflow—you have a massive reliability gap.

Compare this to a proper WAL-based approach or a dedicated broker:

FeaturePostgres LISTEN/NOTIFYTransactional Outbox (Polling)Kafka / Redpanda
PersistenceNone (Transient)High (Stored in Table)High (Distributed Log)
Delivery GuaranteeAt-most-onceAt-least-onceAt-least-once / Exactly-once
Consumer ScalingHard (Broadcast to all)Easy (Row-level locking)Native (Consumer Groups)
BackpressureNon-existentNatural (Table size)Native (Log offsets)

If your consumer fails to process an event, there is no built-in way to "replay" the notifications that were sent while it was offline. You end up writing custom, complex logic to "re-sync" the state, which usually involves scanning the entire table to find what you missed. At that point, you are just building a poorly optimized version of a Change Data Capture (CDC) system.

3. The Fan-out death spiral

In a multi-tenant SaaS environment, you might think: "I'll just use different channels for different tenants."

This doesn't solve the fundamental problem of fan-out. When you issue a NOTIFY, Postgres has to iterate through the list of all active listeners for that channel.

If you have 50 microservice instances all listening to a tenant_events channel, every single notification triggers a broadcast to all 50 instances. If you have 1,000 events per second, you aren't handling 1,000 events; you are handling 50,000 delivery attempts across your fleet.

This creates a massive amount of interrupt driven work for the Postgres backend processes. You will see CPU spikes on your database primary that don't correlate with your query volume. You will see increased latency on unrelated transactions because the postmaster is busy managing the notification broadcast.

This is why you cannot easily scale your consumer tier. In a proper pub/sub system, adding more consumers usually helps distribute the load. In Postgres LISTEN/NOTIFY, adding more consumers increases the load on the database.

4. What I actually do: The Transactional Outbox

If you need guaranteed delivery and you are already using Postgres, do not use NOTIFY. Use the Transactional Outbox pattern.

It is slightly more "heavyweight" in terms of storage, but it is infinitely more predictable. Instead of sending a signal, you write the event into a dedicated outbox table within the same transaction as your business logic.

SQL
-- The atomic operation
BEGIN;
  INSERT INTO orders (id, user_id, amount) VALUES ('ord_123', 'user_456', 100.00);
  INSERT INTO outbox (event_type, payload, created_at) 
  VALUES ('order_created', '{"id": "ord_123", "amount": 100.00}', NOW());
COMMIT;

Now, how do you get that event out of the database? You have two real options:

Option A: The Polling Publisher (Simple)

A background worker polls the outbox table every 100ms-500ms.

SQL
-- Fetch and lock rows to prevent multiple workers from picking up the same event
UPDATE outbox
SET processed = true
WHERE id IN (
    SELECT id FROM outbox
    WHERE processed = false
    ORDER BY created_at ASC
    LIMIT 100
    FOR UPDATE SKIP LOCKED
)
RETURNING *;

Using FOR UPDATE SKIP LOCKED is the key here. It allows you to run multiple instances of your worker without them stepping on each other, and without causing massive lock contention. This scales linearly with your worker count.

Option B: CDC (The Gold Standard)

Use a tool like Debezium to tail the Postgres Write-Ahead Log (WAL).

This is the most performant way to do eventing in Postgres. Debezium reads the WAL directly. It doesn't query the tables, and it doesn't interfere with your transaction performance. It turns your database changes into a stream of events that can be pushed into Kafka or Pulsar.

This gives you the best of both worlds: the ACID guarantees of your local transaction and the massive scale of a dedicated streaming platform.

The Verdict

LISTEN/NOTIFY is a tool for low-stakes, high-frequency, transient signaling.

Use it if:

  • You are building a real-time dashboard where missing an update is fine.
  • You are implementing local cache invalidation in a single-node setup.
  • Your throughput is low and predictable.

Avoid it if:

  • You are handling financial transactions or critical state changes.
  • You need to guarantee that an event is processed at least once.
  • You plan on scaling your consumer fleet to more than a handful of instances.
  • You want to avoid "hidden" CPU load on your primary database.

If you are building a production-grade eventing system for a fintech or a high-growth SaaS, skip the NOTIFY command. Implement a Transactional Outbox with SKIP LOCKED for medium scale, or move to CDC with Debezium for high scale. It is more work upfront, but it won't wake you up at 3 AM when your event lag starts killing your database.

— Rahul Gupta
share