Change Data Capture Patterns for Graph Synchronization
Log-based CDC and Kafka durability eliminate polling's blind spots in graph synchronization.

Picture a team running a graph database as a derived query layer sitting on top of PostgreSQL, using a scheduled poll to keep it current. For a few weeks it looks fine. Then someone runs a traversal query and gets a result that doesn't match what the source tables say, and nobody can pinpoint when the two systems stopped agreeing. Polling becomes a structural liability at that moment, because a graph can't reconstruct its own state the way a relational replica can.
A relational target can survive sloppy synchronization because a full-table comparison tells you almost everything you need to know: which rows exist, which changed, which vanished. A graph doesn't have that luxury. Relationships in a graph are first-class objects, stored as edges with no timestamp column of their own, so there's nothing to poll for on the relationship side. You can detect that a customer row changed. You can't detect, through polling, that the foreign key inside that row now points somewhere else, unless you re-derive the entire edge structure from scratch every cycle, which defeats the purpose of polling in the first place.
Three failure modes appear once a graph sits downstream of a polling loop, and each one is worse than it sounds at first pass. The first is missed deletes. Standard queries can't detect a deleted record unless the source implements soft deletes with timestamp columns, and even then, a deleted node that still has edges attached leaves the graph in a structurally broken state, not just a stale one. Dangling edges don't throw errors. They just sit there, waiting for a traversal to hit them.
The second failure mode is a race condition inside the polling window itself. Concurrent updates that land between two poll cycles can produce a graph edge pointing to a node version that was never actually stable in the source at any point in time. The graph ends up representing a state that never existed anywhere else, which is a strange kind of corruption because it looks internally consistent while being factually wrong.
The third is relationship drift. A row change that touches a foreign key changes an edge in the graph, but a poll-based model has no mechanism to detect that at the relationship level. It only sees the row. Over time the graph's topology quietly diverges from the schema it's supposed to mirror, and because nothing throws an alert, the divergence compounds. Warehouses and caches can tolerate a few hours of staleness because the semantic unit is a row. Graphs can't, because the semantic unit is a relationship, and relationships have no polling-friendly signal of their own.
Capturing changes with log-based CDC
Change Data Capture captures row-level inserts, updates, and deletes as they happen, delivering each one with its before-state, after-state, operation type, and source position metadata. The before-state matters more for graph synchronization than it does almost anywhere else, because knowing what a row looked like before a change is how a consumer knows which relationships to retract before it creates the updated ones.
Three capture methods exist, and they land very differently once a graph is the target. Query-based polling, the timestamp approach, inherits every problem described above: it can't reliably catch deletes, it depends on timestamp columns existing in the first place, and it's exposed to race conditions whenever updates land mid-cycle. It's disqualified for graph targets for the same structural reasons that make it fragile everywhere else, only more so.
Trigger-based capture fires on every write, which adds performance overhead directly on the source database and requires schema changes to install. It can also be disabled by anyone with the right permissions, which makes it fragile in ways that have nothing to do with the data itself. It works for low-write-volume sources, but trigger maintenance gets harder as table count grows, and graph targets typically pull from many source tables at once, so the maintenance burden scales in the wrong direction.
Log-based capture reads the database's transaction log directly. It adds minimal load to the source, captures every operation including deletes, needs no schema modifications, and preserves the exact order in which operations occurred. That ordering guarantee makes log-based capture the right foundation for graph work, not just a nice-to-have. A node has to exist before an edge referencing it can be created, and if a consumer applies an edge-creation event ahead of its corresponding node-creation event, the graph becomes structurally invalid. Log-based CDC's ordering guarantee removes that entire class of error, since events arrive in the same sequence they occurred in the source.
There's a second benefit that matters for teams running production databases under real load: log-based CDC reads from logs the database already writes for its own recovery purposes, so the source does no extra work beyond what it's doing anyway. On PostgreSQL, the standard foundation is the write-ahead log combined with logical replication, decoded through the native pgoutput plugin, which has shipped since PostgreSQL 10 and is what Debezium reads to produce structured change events. Every property that graph synchronization needs, order, completeness, delete visibility, no schema disruption, comes from that one mechanism. What's left is building the pipeline that carries those events somewhere useful.
Debezium and Kafka as the operational layer between relational source and graph target
The combination that's become the default relay layer pairs Debezium's log reading with Kafka's durable, replayable event log, so a graph consumer can process changes on its own schedule without losing anything if the graph write path goes down temporarily. Debezium reads the PostgreSQL WAL and turns each change into a structured event, before-state, after-state, operation type, transaction ID, and source position, published to named Kafka topics. It's the most widely used open-source CDC platform, and it ships both as Kafka Connect source connectors and as a standalone Debezium Server for teams that want to skip Kafka as an intermediary layer entirely.
A typical PostgreSQL connector configuration looks something like this in practice:
{
"connector.class": "io.debezium.connector.postgresql.PostgresConnector",
"database.hostname": "prod-pg.internal",
"database.dbname": "orders_db",
"plugin.name": "pgoutput",
"table.include.list": "public.customers,public.orders,public.order_items",
"topic.prefix": "orders_cdc",
"key.converter": "io.confluent.connect.avro.AvroConverter",
"value.converter": "io.confluent.connect.avro.AvroConverter",
"value.converter.schema.registry.url": "
}
That block shows the three decisions that matter most for a graph target. First, plugin.name is set to pgoutput, the native decoding mechanism built into PostgreSQL since version 10. Second, table.include.list routes each source table to its own Kafka topic under a shared prefix, topic-per-table routing, which becomes the seam that later sections build translation logic around. Third, the Avro converters point at a Schema Registry, where type safety enters the picture.
Kafka's contribution to graph synchronization comes down to three things. It provides durability, since events sit on disk and a graph write failure doesn't lose data, the consumer just resumes from its last committed offset. It provides replayability, letting a consumer reprocess the event log from any point, which matters when a graph schema migration requires rebuilding relationship structure from historical events. And it enables fan-out, so the same CDC stream can feed a data warehouse and a graph target at the same time without querying the source database twice.
Schema Registry and Avro serialization matter more for graph targets than for most other consumers, because a source column type change, an integer foreign key becoming a UUID, has to be caught before it silently corrupts edge identifiers in the graph. And Dead Letter Queue handling isn't optional here either. DLQ patterns are the standard way CDC connectors deal with schema incompatibilities, serialization errors, network failures, or bad data, and for a graph target, a silently dropped node-creation event means every subsequent edge-creation event referencing that node produces a dangling edge. The relay layer has to move events and make sure none of them go missing without anyone noticing.
The schema translation problem: mapping relational change events to graph writes
This is the part of graph synchronization that has no shortcut. A relational change event describes a single row mutation. A graph write has to express that mutation as some combination of node property updates, node creation or deletion, edge creation or deletion, and edge property updates, and a single row change can require several of those operations at once.
The structural mismatch runs deep. Mapping a relational database into a graph typically means running a query against the relational side, tagging result columns, and building a typed, directed property graph from the output. Going the other direction, from a graph transaction back into relational storage, means updating a mapping layer with a surrogate describing the change, and designing that mapping well requires real attention to constraint satisfaction and information integrity on both sides. None of this happens automatically. It has to be built.
Four translation cases come up constantly, and each one has a wrong answer that looks plausible until it corrupts the graph. A foreign key update in the source, say a customer's assigned sales rep changes, has to produce an edge deletion for the old relationship and an edge creation for the new one. It is not a property update, and treating it as one leaves a phantom edge in the graph. Getting this right depends entirely on the before-state carried in the CDC event, so that field isn't optional for graph targets.
Junction tables are the second case. A many-to-many join table row maps to a graph edge with no corresponding node of its own, so CDC events on that table have to be translated into pure edge operations, never node operations. Treating a junction row like any other row causes the translation layer to create a node that shouldn't exist.
The third case is a NULL foreign key. Setting a foreign key to NULL should produce an edge deletion in the graph without touching either endpoint node, and a naive translation rule that deletes "the row's graph object" gets this wrong immediately, deleting a node that's still valid.
The fourth case, column rename or type change, doesn't even produce a row-level event. Schema evolution in the source has to be signaled through the Schema Registry so the translation layer updates its mapping before the next data event arrives, otherwise the mapping silently misreads the incoming payload.
Apache Flink or a dedicated stream processor is the appropriate layer for this translation: it can maintain stateful mappings, join CDC events across multiple source tables to construct a single graph operation, and enforce ordering guarantees between node and edge writes. Get this layer wrong and the failure mode is silent. The graph drifts instead, a little more with every mistranslated event, until someone runs a query that exposes the gap.
Event routing strategies for multi-entity graph writes
Most meaningful graph operations aren't single-table affairs. Creating a subgraph that represents "customer placed order containing product" pulls from at least three source tables, customers, orders, and order_items, and those tables' CDC events can land on different Kafka topics at different times, with no guarantee they arrive together.
Three routing patterns handle this, and each fits a different shape of problem. Topic-per-table with an ordered consumer is the simplest: one Kafka topic per source table, one consumer applying events in offset order. It breaks down when a graph edge requires data from two tables that cannot be joined at the event level, for instance an edge between an order and a product that needs quantity data living in order_items.
Topic aggregation with a stream join covers that gap. Flink or a comparable stream processor joins CDC events from orders and order_items on a shared key like order_id within a time window, then emits one composite event to the graph-write consumer. Relationship drift and the missing polling signal
The outbox pattern takes a different approach. Rather than reconstructing a graph operation from separate table events after the fact, the source application writes a pre-composed graph operation payload into an outbox table inside the same transaction as its relational writes. CDC then captures that outbox row and routes it straight to the graph writer, skipping the join and translation problem. The tradeoff is coupling: the source application now has to know something about the graph schema it's feeding, which isn't always an acceptable design constraint depending on how the application team operates.
Relationship-aware transformation for entity resolution across data silos
The cross-silo case is where graph synchronization starts paying for itself, because it demands more than a technical exercise to get right. When the same real-world customer, or product, or account, exists across two or more source systems that were never designed to talk to each other, CDC transformation has to do more than translate one table's events. It has to recognize that two change events from different systems describe the same entity.
Entity resolution inside a CDC pipeline works by clustering pairwise matches into entity graphs, which limits the errors that transitive closure would otherwise introduce, chaining together matches that shouldn't actually be linked. Running this well inside a streaming pipeline demands three properties at once: incremental matching of the CDC stream, idempotent updates, and safety under event replay.
Idempotency isn't a nice-to-have here. If a CDC event replays, because a consumer restarted or a Kafka offset got reset, the graph write has to produce the exact same result rather than a duplicate node or a duplicate edge. That requirement rules out blind insert semantics entirely and pushes toward merge semantics instead, MERGE in Cypher or an equivalent upsert pattern in other graph APIs.
Incremental matching is the second constraint, and it's arguably the harder one. A full entity-resolution run against a static snapshot isn't something a CDC pipeline can do, since data never stops moving long enough for a batch run to make sense. Resolution logic has to evaluate each incoming event against the current graph state on the fly, without re-scanning every entity in the graph every time a new event lands.
This is the exact point where a graph target justifies its cost. Joining across system boundaries in a purely relational world means coordinating ETL jobs across multiple source databases, a brittle and slow process. A graph kept current through CDC serves that same join at query time, without touching any of the source systems directly. The business case for building the graph in the first place lives right here, in the query that used to take a multi-system ETL run and now takes a single traversal.
Neo4j's native CDC API and the current state of graph-side capture support
Everything covered so far assumes changes originate in the relational source and flow outward into the graph. Graph-side CDC support flips that assumption. Rather than treating the graph purely as a passive write target, a graph database with native change capture becomes a change-emitting source in its own right, which opens the door to bidirectional synchronization and to downstream consumers reacting directly to mutations happening inside the graph.
Since version 5.13, a graph database's enterprise edition can ship a built-in Change Data Capture API that exposes a change stream through Cypher procedures. That capability matters for architectures where the graph isn't just a derived read layer anymore but an active participant, feeding changes back into other systems, triggering downstream workflows, or synchronizing with a second graph instance in a different region.
Graph-side CDC support is still maturing relative to the decades of tooling built around relational transaction logs. Teams evaluating it should treat it as a useful complement to source-side capture for the translation and routing work described in the sections above. The two problems, getting changes out of a relational source correctly and letting a graph emit its own changes reliably, are related but distinct, and both need to be solved for a graph layer to stay trustworthy over time.


