Friday, October 2, 2026

Rethinking Real-Time Data Architecture on AWS

When Compute and Storage Stop Scaling Together

Why the Cluster Became the Scaling Unit

Real-time data platforms can become expensive when one kind of growth forces a team to buy capacity for another. Longer retention may require more nodes, even though their processors have spare capacity. A burst of readers may compete with the machines accepting new events. On AWS, separating storage from compute can reduce that coupling, but it does not by itself give ingestion and reads independent capacity.

Keeping data on local disks, close to the processors using it, can make low latency easier to achieve. When a node also provides the storage needed to retain history, however, capacity planning has to accommodate both requirements. Adding nodes can bring more processors, memory and network bandwidth than the immediate problem needs.

Attached storage already offers some flexibility: Amazon MSK Standard brokers, for example, allow EBS volumes to grow without adding brokers. A shared durable storage layer can also let replacement compute access retained data without copying the full history onto its own disks. Separating write and read workloads is another architectural decision, as Figure 1 shows.

Figure 1  Capacity can be separated at different boundaries. Shared bottlenecks may remain.
Figure 1 Capacity can be separated at different boundaries. Shared bottlenecks may remain.

What Storage Separation Changes on AWS

Real time describes how quickly information becomes usable; it does not prescribe where the full history lives. A payment event may need immediate processing while older events are rarely accessed. Systems built around Amazon S3 can keep durable data outside their compute nodes and cache frequently accessed records nearby. The benefit depends on how much data needs fast access and how often that working set changes.

Amazon MSK tiered storage illustrates a narrower, concrete version of this approach. On supported Standard-broker configurations, it copies completed log segments to a remote tier. Local retention can be shorter than overall retention, so a topic can retain more history without increasing broker count merely to hold that history. For example, a team could retain two days locally and thirty days overall, provided its throughput and local-storage needs still fit the existing brokers.

Local retention needs deliberate configuration. If it inherits the overall retention setting, the same data remains in both tiers until it expires. Tiering then provides no shorter local history. Nor does it create an independent read-serving fleet: consumers still fetch through brokers. Replaying older data adds network traffic, and AWS warns that excessive remote reads can affect current produce or consume traffic. Storage capacity has been separated more than serving capacity.

Uneven Growth Exposes the Cost of Coupling

Uneven demand can make the cost of coupling easier to see. Extending retention with stable traffic differs from handling a two-hour ingestion spike. The spike also creates more data, but it need not justify keeping extra processors after traffic subsides. More readers present a different problem again, especially when they scan history rather than follow newly arriving events.

A useful cost comparison includes the resources that remain idle and the work introduced by separation. On AWS, that means accounting for requests and retrieval, applicable data-transfer charges, cache capacity and the compute needed to move or process data. A lower storage price alone does not establish a lower operating cost. Frequent historical scans or repeated cache warm-ups can erode the expected saving.

Test
Figure 2 Test uneven growth against agreed freshness and latency limits.

Disaggregation Has Practical Limits

Remote storage changes the work required to retrieve a record. A cache miss may require a network request and a larger read than the application actually needs. AWS documents higher initial read latency when MSK retrieves data from its remote tier.  A replay job may tolerate that delay; an application with a tight response-time budget may need the relevant data cached or materialised in a serving database.

Independent compute pools also need explicit limits on shared resources. Readers and writers can still compete for network bandwidth, metadata services or storage throughput. Admission control and workload quotas help contain that competition, while separate serving resources may be needed to protect ingestion. These measures have their own cost and complexity. For a stable workload that already meets its latency and budget targets, keeping some functions together can be a reasonable choice.

A Capacity Test for AWS Architectures

The tests in Figure 2 reveal how much independence an architecture actually provides. Longer retention should not require more processing capacity simply to hold the additional history. More readers should not push ingestion beyond its agreed latency limits. Where those changes still affect other workloads, the system retains dependencies that a diagram of separate components can obscure.

The transition matters as much as the final capacity. New compute may need time to restore state or warm its caches before it can serve traffic effectively. A platform that handles a peak comfortably may also be difficult to shrink afterwards, leaving the business paying for capacity it no longer needs. Both behaviours belong in the cost assessment.

For an AWS architect, the useful question is how much of the platform must change when one demand grows. Object storage can make longer retention affordable without a matching increase in compute. Separating ingestion from serving requires further design work. The value of each step depends on the capacity it saves and the performance it preserves.

Sources

1  AWS — Scaling Standard broker storage
2  AWS — Tiered storage for Standard brokers
3  AWS — Tiered storage retention settings
4  AWS — Production practices for MSK tiered storage

Latest