Articles /
CQRS: when the user saves and cannot see their own data
Separating reads from writes solves a pile of problems and creates a new one, which shows up on the user's screen and arrives as a support ticket.
I have lost count of how many times I have seen this ticket. The user edits the delivery address, saves, the screen confirms it went through. They reload the page to check and the old address is still sitting there. They open a ticket saying the system did not save it, support escalates it as a bug, and someone on the team replies that it is working as expected.
And it is working as expected. That does not help the user at all, who is now unsure whether their order is going to show up at the right place.
This post is about that specific moment, but first I need to walk through the path that leads to it.
Separating reads from writes
CQRS stands for Command Query Responsibility Segregation. The name is long and the idea behind it fits in one sentence, which is to use one model to change data and another one to read it.
In most systems we write, the same model does both jobs. The same Order entity, with the same relationships and the same validations, is used to create an order and to build the listing that shows up on screen. And in something like 95% of cases that is enough, without any pattern with a fancy name.
The story changes when the query starts asking for something the write model was never built to give. A report that joins order, item, customer, payment and delivery, groups by month and by region, and scans a full year of history to close a total. The normalized model is great for writing, with short transactions and few rows touched, and it is the worst possible shape for answering that kind of question.
That has a name, it is an OLAP workload running inside an OLTP database. The two profiles want opposite things. Writing wants short transactions, selective indexes and the fewest rows touched. Analysis wants to scan a lot of rows, join tables and aggregate. When both share the same instance they fight over buffer, IO and CPU, and the damage usually lands on the wrong side, someone opens the monthly report and checkout latency goes up along with it.
Then comes the part that usually gets misunderstood, because there are two very different steps of CQRS and people normally picture the second one.
First step, same database
The cheapest step is to separate in code only. We keep a single database, with the same schema, and change the way we reach it.
The write goes through the full path, loads the aggregate, validates, applies the rule and saves. The read ignores all of that and runs a direct query, returning exactly the fields the screen needs, instead of loading a whole aggregate with every relationship attached to render four columns.
// write, with business rules in the middle
func (s *OrderService) Cancel(ctx context.Context, id string) error {
order, err := s.repo.FindByID(ctx, id)
if err != nil {
return err
}
if err := order.Cancel(); err != nil {
return err
}
return s.repo.Save(ctx, order)
}
// read, skipping the aggregate
func (q *OrderQuery) List(ctx context.Context, customerID string) ([]OrderSummary, error) {
return q.db.Select(ctx, `
SELECT o.number, o.created_at, c.name, o.total
FROM orders o
JOIN customers c ON c.id = o.customer_id
WHERE o.customer_id = $1
ORDER BY o.created_at DESC`, customerID)
}
No events, no queue, no synchronization. Not even eventual consistency, because the data is the same. And in my experience this step covers the vast majority of cases where someone says they need CQRS.
A note on the repository is worth it here, because it is the thing that makes that second snippet look like a hack. The community knows it as the Repository Pattern, and for a lot of people it became a synonym for clean architecture. Its job is to translate a domain aggregate into a database row, and a database row into a domain aggregate. It exists so the domain does not have to know there is SQL on the other side, and the price is that it always hands over the whole aggregate, hydrated, with the relationships the business rule might end up using.
On the read side none of that is needed. The order.Cancel() above is a domain operation, encapsulated inside a rich entity, the way good old DDD wants it, and nobody is going to call that from a listing. So you can skip the repository and go straight to the database, returning a read model that exists only to feed the screen and dies right there. In some cases you can push it a bit further and leave the query ready inside the database itself, with a view, or with a materialized view when it is heavy enough that recomputing it on every access stops making sense.
And notice that the materialized view already carries the same trade-off as the rest of this post in miniature, because somebody has to refresh it, and between one refresh and the next it serves stale data. If you go down that road, REFRESH MATERIALIZED VIEW CONCURRENTLY is worth the trouble, since it does not block reads while updating, at the cost of requiring a unique index on the view.
Second step, separate databases
The second step is when the two sides end up in different stores. The write stays in Postgres, with the normalized model, and the read moves somewhere else, which can be a replica, a Redis, an Elasticsearch, a denormalized table with everything ready for the screen.
Somebody has to keep the two in sync, and usually an event does that. The service records the change, publishes an event, and on the other side a consumer updates the read side.
Worth saying that publishing that event has a trap of its own, which is writing to the database and publishing to the broker with no shared transaction between them. I have written about it in the outbox and saga post, so I am not going to repeat it here. It is worth reading before you build this pipeline, because the consumer that updates the read model has exactly the same problems as any other consumer, including repeated delivery and out of order arrival.
Another common confusion is thinking this second step requires event sourcing, and it does not. You can split the stores and still keep the current state on both sides, with no event log, no replay, no aggregate rebuilt from history. Event sourcing pairs well with CQRS, which is why the two show up together so often, but neither one depends on the other.
When it is worth it
It is worth it when reads and writes have very different loads and reads win by a wide margin. A product catalog read thousands of times per second and updated three times a day is the classic case.
It is worth it when the query has to pull data from several services. If the order detail screen shows data living in four different databases, assembling that on demand costs four calls and a latency nobody wants.
It is worth it when the shape of the data on the read side is far from the shape on the write side. If rendering one screen takes six joins and an aggregation, keeping that precomputed in a read table takes weight off everyone.
And it is worth it when two parts of the team can work in parallel, one on the complicated business rules of the write side and another on the screen and the queries.
Now, the important one, which is when it is not worth it. If the domain is simple and the screen is a CRUD, CQRS delivers complexity and nothing in return. If the team is small, somebody is going to maintain both sides alone, plus the synchronization between them. And if the data has to be visible the moment it is written, the second step will fight that requirement all day long, which is exactly what the rest of this post is about.
The bill that arrives later
When the read lives somewhere else, there is a gap between writing and being able to read what was written. That gap is eventual consistency, and it shows up in every CQRS reference as an acceptable downside.
The problem is that it almost never shows up as a number, and almost never shows up as a decision about a screen.
From the point of view of whoever designed the system, the steps happened in exactly the expected order. From the point of view of whoever was in front of the screen, the system lost what they had just typed.
And notice the gap can be small on average and horrible in the tail. A projection that applies in 40ms at the median and 4 seconds at the 99th percentile is going to generate tickets, because 4 seconds is plenty of time for the user to reload the page.
Reading your own writes
The discussion around this usually gets stuck on strong consistency against eventual consistency, as if those were the only two options. There is a band in the middle that rarely gets mentioned, and it is the one that solves the support ticket.
Read-your-own-writes guarantees that a session always sees its own writes. It says nothing about what other sessions see. If two users edit different things, each one sees their own change right away, and sees the other one whenever the data arrives.
Monotonic reads guarantees that a session never goes back in time. Without it, two reloads in a row can land on different replicas, and the new address appears, disappears, and appears again. If you think seeing stale data is bad, seeing it flicker is much worse.
Both are session guarantees. They cost far less than making the whole system consistent, and they are exactly what that user in the ticket was asking for. They were not worried about whether the rest of the world could already see their new address.
What you can do about it
Before picking a solution, it is worth knowing what the screen in question actually has to promise.
Accept it
Worth starting here, because this is the answer more often than it looks.
Feed, search, report, dashboard, history listing. Nobody just wrote that data, and nobody notices two seconds of lag. If the screen is not the immediate result of an action by the user themselves, the lag can simply exist.
Read from the write model right after the command
The confirmation screen, or the redirect that happens right after saving, can go straight to the write database instead of going through the read side.
// GET /addresses/:id?after_write=1
if req.URL.Query().Get("after_write") == "1" {
return h.write.FindAddress(ctx, id)
}
return h.read.FindAddress(ctx, id)
It is ugly, and it solves a good chunk of the tickets for almost nothing. The care here is about discipline, this has to stay restricted to the route that comes immediately after the command. The day after_write=1 starts showing up in the listing, the separation became decoration and it is worth asking whether it should exist at all.
Render with what the client already has
The browser just sent the payload, so it knows the outcome without having to ask. You can render the screen with what was sent, no extra query.
Instant and cheap, with one consequence that needs handling. The command can fail after the 200. An asynchronous validation rejects it, a saga compensates, a consumer refuses the event. The screen showed a state that stopped existing, and somebody has to undo that in front of the user.
I use this without worrying for low risk changes, like editing a display name. I would not use it for anything involving money.
Carry the version along
This is the complete solution, and it is the one that takes the most work.
The command returns the version it produced, and the next read requires at least that version. On the write side, one column per aggregate handles it:
UPDATE addresses
SET street = $2, version = version + 1
WHERE id = $1
RETURNING version;
The number goes back in the command response:
HTTP/1.1 200 OK
x-resource-version: 42
And the projection stores the source version alongside the row it writes:
INSERT INTO address_view (address_id, street, source_version)
VALUES ($1, $2, $3)
ON CONFLICT (address_id) DO UPDATE
SET street = excluded.street,
source_version = excluded.source_version
WHERE address_view.source_version < excluded.source_version;
That WHERE at the end of the DO UPDATE does one extra job besides enabling the comparison, it keeps an older event, redelivered after a newer one, from overwriting the row with stale data. As I said earlier, a projection is a consumer like any other.
At query time, with the minimum version in hand, there are three ways out:
view, err := h.read.FindAddress(ctx, id)
if view.SourceVersion < minVersion {
// 1. wait a bit and try again
// 2. answer 409 and let the client decide
// 3. fall back to the write model
}
All three are defensible, and the choice is more about product than about architecture. The short wait makes the problem disappear, at the price of holding the request. The 409 hands the decision to the client, which can show a decent notice instead of wrong data. The fallback always works, and it puts load on the write database exactly when the projection is behind, which tends to be exactly when the system is already struggling.
Two things to know before choosing this path. The first is that every command starts returning a version, every read starts carrying one, and the client has to hold on to it for the session. Threading that number across services is real work.
The second is that this works for detail screens. Listings, counters and aggregations have no single version to require, because each row depends on several different aggregates. For those screens, accepting the lag is usually the answer.
When the read side is a replica
If the read model is a replica of the same Postgres, rather than a projection with its own schema, the database already hands you this control. The primary knows the position of the commit:
SELECT pg_current_wal_lsn();
And the replica knows how far it has replayed:
SELECT pg_last_wal_replay_lsn() >= $1;
If the comparison is false, the read goes to the primary or to another replica. If it is true, it answers right there. Same idea as the version, with the database doing the work instead of the application.
How to measure this
Broker consumer lag tends to be the metric already on the dashboard, and it answers a smaller question than we need. It tells you how far behind the end of the topic the consumer is, and it ignores the time between commit and publish. Worse, it stays at zero while the projection consumes events and fails to apply them, because as far as the broker is concerned those were delivered.
The metric that answers questions during an incident is the end to end one, from the moment the fact happened until the row becomes visible on the read side. A checkpoint table covers it:
CREATE TABLE projection_checkpoint (
projection TEXT PRIMARY KEY,
last_event UUID NOT NULL,
occurred_at TIMESTAMPTZ NOT NULL,
applied_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
Two readings worth alerting on come out of it. applied_at - occurred_at is the real projection lag, and it is the number that should be driving the design of the screen. now() - applied_at is how long it has been since anything was applied, and it is what exposes a stuck projection on a quiet night, when everything else on the dashboard looks healthy.
Closing
The cost of CQRS shows up in talks as the second database and the projection to maintain, and that part is honestly the easy one. What tends to hurt is the promise the interface makes, and that promise is born on the screen, far away from the architecture diagram.
That is why I think the conversation has to happen screen by screen, and before splitting the models. Which screens can live with lag, which ones have to show the write immediately, and what the interface does while the data is still on its way. Answer that first and the rest of the design gets a lot easier to decide.