Can One Tenant Overload Your Shared Cache?
Evaluate shared-cache tenant isolation through eviction, command cost, refill load and admission evidence. Compare bounded sharing with dedicated capacity.
A tenant prefix does not reserve capacity
One tenant can affect another tenant's service when they share cache capacity. A distinct key prefix helps identify ownership, but it does not reserve memory, execution time, connections or downstream refill capacity. Decide whether the shared design is acceptable by testing the other tenant's business function while an allowed workload consumes those resources.
A SaaS platform might cache each customer's report results under a separate namespace. During a large report import, one customer creates many new entries. Another customer's usual entries disappear, cache misses rise, and the application reads more often from the database. Nobody accessed the wrong customer's keys. The availability boundary still failed.
That scenario is hypothetical. This article proposes a review method for platform engineers operating a shared cache, with Redis Open Source examples where the documented behavior applies. It does not establish that every managed cache has the same isolation controls or that a particular instance has a defect. Record the actual product, version, topology and policy before drawing a conclusion.
Separate access permissions from resource protection
Redis's ACL documentation describes permissions for commands and key patterns. It also warns that key patterns do not limit commands that operate on an entire database without named key arguments. Command restrictions therefore need their own review. Our capacity recommendation is separate: do not treat permission to access only one namespace as evidence of a tenant memory or execution quota.
Trace the identity used by the application. If every customer request reaches Redis through one service identity, a tenant-specific key convention may depend entirely on application code. Check how the application constructs tenant identity and rejects missing or conflicting context. A user-supplied prefix must not decide whose cached data a request may read. Authorization tests and load tests answer different questions and both need evidence.
Review privileged maintenance paths too. A support tool, batch worker or cleanup task may have broader access than ordinary requests. Its command permissions and capacity limits can affect all tenants even when the customer-facing API is carefully restricted. Preserve the scope of the test identity in the evidence record. Testing a restricted client does not prove that an unrestricted worker cannot create a wider failure.
Avoid collecting raw cache keys or values merely to make an ownership dashboard. Keys can contain personal or commercially sensitive identifiers. Prefer an internal tenant reference, operation class and aggregated measurement. Define who can inspect detailed diagnostic samples, why they are needed and when they expire. Observability should not introduce a new cross-tenant disclosure path while investigating an availability issue.
Follow eviction into the application and database
Redis's eviction guidance describes a memory limit and policies that choose eligible keys when that limit is exceeded. An all-keys policy can select across the eligible keyspace; a volatile policy considers expiring keys. With no eligible expiring keys, volatile policies behave like noeviction. Under noeviction, writes that need additional memory can fail while reads of existing data remain possible. Replication and persistence buffers also need memory beyond the eviction comparison.
These mechanisms produce different application failure conditions. A report cache might rebuild an evicted result safely but slowly. A session, deduplication record or coordination key can have a different consequence. First classify whether each key is a disposable copy, a security-relevant decision or state needed for business correctness. Do not move all three classes into one eviction pool because their storage technology happens to match.
Check the miss path end to end. Can concurrent requests refill the same missing object? Does the application bound that work, coalesce duplicate requests and stop within the caller's deadline? Does a cache write rejection leave the caller with a valid database result, or cause another round of reads? A fallback that appears harmless for one request can exceed database capacity when many tenants lose their working sets together.
Keep recovery behavior in scope. Once the large writer stops, other tenants may still need to repopulate their data. A test that ends when memory usage falls misses that period. Observe database connections, application latency, error rates and refill backlog until the chosen business function is usable again. Record which operation remained unavailable and which owner accepted the recovery result.
Inspect placement and expensive operations
Count more than keys. A tenant with few large values, a costly command pattern or many simultaneous connections can affect shared capacity differently from a tenant with many small entries. Attribute demand at the application boundary by bytes admitted, operation class, request concurrency and response size. Measure the actual backend as well; an application counter alone cannot explain every cache resource cost.
Redis's cluster specification explains that hash tags place matching keys in the same hash slot. Our placement recommendation is to inspect whether a tenant's naming scheme concentrates its work. A convenient multi-key operation does not prove balanced resource consumption across the cluster.
Review burst shape as well as average volume. A scheduled import, cache warming job or AI-generated report workflow may create a short period of expensive reads and writes that an hourly average conceals. The fact that an AI agent initiated the work does not exempt it from the platform's admission limits. Bind work to the authenticated tenant and operation budget before the cache request is issued.
Do not assume that moving a tenant to another shard isolates every dependency. The application may retain a shared connection pool, worker queue or database. Draw the dependency boundary in the review notes and name which resource moved. Keep any remaining shared failure path explicit rather than calling the whole service isolated after one infrastructure change.
Compare three capacity boundaries
The following comparison is an engineering review framework, not a Redis feature list or a promise that each hosting product supports these options. Confirm available controls and their scope with the actual operator. Choose the smallest boundary that can meet the accepted service requirement under the permitted workload.
| Design choice | Evidence needed | Remaining failure path | | --- | --- | --- | | Shared cache with application admission limits | Every caller uses the limit; bursts and refill work stay bounded | Bypassed callers or shared backend capacity can affect other tenants | | Separate cache pool for a workload class | Routing, capacity and eviction behavior are independently tested | Shared application workers or database refill remain coupled | | Dedicated tenant cache capacity | Identity, routing, operation and recovery ownership are verified | Shared upstream dependencies can still affect the tenant |
A request-rate limit alone may permit large values or expensive operations. A byte limit alone may permit many repeated commands. Define separate budgets for the demand that caused the observed problem and explain how the application responds when a budget is exhausted. Rejection, delayed execution and a bounded degraded response have different effects on the customer. Avoid silently transferring all rejected cache work to the database.
More separation also creates operational work. An additional pool requires routing rules, credentials, monitoring, maintenance and a recovery owner. A dedicated tenant may need a migration that avoids serving stale data from both old and new locations. Compare those costs with the measured consequence of sharing. Neither the lowest infrastructure bill nor the largest number of separate instances establishes the correct design.
Run a bounded two-tenant rehearsal
Use synthetic tenants and an approved isolated environment. Give one tenant a representative steady business function, such as opening a cached report. Give the other tenant a bounded burst through the same application path used in production. Define the allowed value sizes, concurrency, duration and command classes before running it. Set stop conditions for both cache pressure and database fallback load; do not discover the safety limit by exhausting a live shared service.
Collect a baseline for the steady tenant before the burst, then compare its latency, errors, misses and downstream work during the burst and through recovery. Add an independent control operation that does not depend on the cache if the environment permits it. That can help distinguish cache-related pressure from an unrelated application bottleneck, though a correlation still needs investigation before a causal claim.
Global cache hit and eviction counters are useful signals, but they do not identify which tenant lost useful entries. Retain tenant-attributed application observations alongside node-level measurements. Separate ordinary expiry, explicit invalidation, failed writes and suspected eviction instead of labelling every miss as another tenant's fault. Record instrumentation gaps and sample uncertainty.
Test the proposed control under the same workload contract. Include an attempt from a batch caller or alternate code path to check whether it bypasses admission. Check the behavior when the budget service is unavailable, without inventing a fail-open default. A passing rehearsal covers its tested function, workload and dependency state. It cannot establish isolation for every future operation or prove a security boundary through performance measurements alone.
Keep a decision record the next engineer can use
Write down the tested tenant identities, cache product and policy, placement, key classes, workload envelope, admission point, baseline, burst observations, recovery observations and unresolved dependencies. Link the measurements without embedding secrets. Name the owner who accepts the steady tenant's result and the changes that require another test, such as a larger value type, new batch caller or different eviction policy.
The next useful step is one two-tenant rehearsal with an explicit acceptance condition for the quieter tenant. If it fails, trace the first exhausted resource and compare a targeted admission control with a separate workload pool. Bring the failed observation and proposed boundary to a reliability review. The PostgreSQL recovery-evidence paper covers a different dependency: proving recovered database state and application acceptance when fallback or recovery work reaches the database. Reading either resource does not require submitting contact information.