Can a Tenant Prompt Survive a Shared Default Change?
Keep tenant-specific AI behaviour understandable when shared defaults change. Review resolved configuration, mixed-version runs, evaluation and rollback limits.
A customer override can stay unchanged while its behaviour changes
A software provider updates the shared instruction for its document assistant. The new default asks the model to suggest the next action after explaining a document. One tenant has an existing customization that allows summaries but prohibits recommendations. Its customization file has not changed, yet the assembled request now includes both instructions. If the application only checks that the tenant file is present, it can miss the changed behaviour its customer receives.
The problem is configuration inheritance, not simply prompt quality. A tenant may inherit a shared instruction, override selected fields, append an example or use a completely separate template. Updating a parent can change a child's effective configuration without editing the child's stored record. Conversely, a full replacement can leave the tenant on an old security or output contract. Neither outcome should be discovered through a customer complaint.
The scenario and expected outcomes below are synthetic teaching fixtures. They are not customer results or measured model performance. This article proposes a configuration review for an advisory document assistant. It does not authorize a model to send messages, alter records or decide a customer's policy. Production access and effect permissions must remain enforced outside prompt text.
Resolve the configuration before evaluating the prompt
Write down the assembly rules the application actually implements. For example, the product may combine a mandatory safety instruction, a versioned shared task template, a tenant-approved customization and the current user request. The resolver should produce an inspectable result with the identity and revision of each input. Do not assume that a JSON object merge and concatenated instruction strings have equivalent meanings. An overwritten field disappears; an appended instruction can contradict an earlier instruction while both remain visible.
OpenAI's prompting documentation recommends keeping production prompts in versioned application code, validating dynamic inputs and testing changes during deployment. Those practices support reproducibility. They do not define your tenant inheritance contract or establish that an unchanged tenant customization remains compatible with a new shared template. The application owner still has to specify that relationship.
Separate four layers in the review. A tenant business preference describes what the assistant should help with. The shared task template describes the default workflow. The output contract describes what downstream code expects to parse. The execution policy decides what operations are permitted. These can change independently. A tenant instruction asking for a different tone should not be able to grant new tool permissions, and an emergency security restriction should not depend on whether a customer accepts a cosmetic prompt update.
Record the effective configuration, not only a convenient label such as prompt revision 12. For one run, the lineage might include shared template S12, tenant customization T4, resolver R2, output schema O3, model configuration M7 and retrieval policy K5. These are local illustrative identifiers, not vendor versions. Keep a digest or equivalent integrity reference for the assembled instructions. Store sensitive customization text only in an approved restricted location; a broadly accessible log should not become a second tenant-data repository.
Define what empty, absent and explicitly disabled mean for every overrideable field. An empty tenant instruction might intentionally remove a greeting, or it might be a failed configuration load. Treating both as inherit the default can silently restore behaviour a tenant removed. A typed schema can distinguish these states, but the product contract must say which states are permitted. Reject an ambiguous configuration rather than asking the model to guess why the field is empty.
Choose what inherits, what stays pinned and what cannot be overridden
In the synthetic tenant A fixture, shared template S11 produces an explanatory summary. Customization T4 restricts the assistant to explanation and forbids proposed actions. The candidate S12 adds an optional recommendations stage. Under the fixture's declared product policy, tenant A should remain explanation-only. The expected result is either a compatible resolved configuration that omits recommendations or a held upgrade requiring review. It is not a recommendation merely because a newer shared template contains one.
Another tenant B explicitly accepts the recommendations stage. Its review should include the new task, not simply replay A's restriction. A third tenant C uses a complete replacement task template. It does not inherit S12's task text, but it still receives non-overrideable execution restrictions enforced by the service. Keeping these cases distinct prevents a supposed compatibility fix from freezing all customers on old behaviour or treating customer text as security authority.
Choose an inheritance policy per field or module. Pinned inheritance makes the parent revision explicit and changes it only through a reviewed adoption. Following a moving parent can suit low-risk presentation defaults if the product promises that behaviour and tests relevant tenant variants. Full replacement can support a specialized workflow, but it requires ownership of the replaced contract. Explain this trade-off to maintainers: more pinning improves traceability while creating more supported combinations and upgrade work.
Do not solve every conflict by placing the tenant instruction last and hoping it wins. Message roles, instruction interpretation and model behaviour matter, but ordering is not a security boundary or proof of semantic precedence. The resolver should eliminate disallowed combinations before submission. For this example, a structured option such as recommendations disabled can control whether the stage is assembled at all. The model still needs evaluation for unwanted advice in an explanation, even when that stage is absent.
Define unavailable-configuration behaviour too. If T4 cannot be loaded, silently substituting the unrestricted shared default defeats tenant A's contract. A previously validated local snapshot may be usable only if the product has an explicit freshness and withdrawal policy. Otherwise, the assistant can explain that the specialized workflow is temporarily unavailable. Availability pressure is not permission to discard a tenant restriction or continue with an unrecorded fallback.
Test compatibility separately from deployment health
AWS AppConfig's validator documentation describes JSON Schema and Lambda validation of configuration. A schema can reject missing fields or incompatible types, and custom logic can reject a forbidden combination. Neither automatically determines whether a generated answer satisfies a customer's explanation-only requirement. Use deterministic validation for assembly rules and a separate evaluation for model behaviour.
The worksheet below specifies expectations before running a model. It changes one configuration boundary at a time. Use inert adapters and permitted example documents. An observed recommendation under tenant A's policy is a failure even if the output parses correctly or a later reviewer removes it. These are proposed cases, not results already obtained by Ampity.
| Controlled configuration | Expected disposition | Evidence to inspect | | --- | --- | --- | | Tenant A T4 with S12 recommendations enabled | Reject assembly or hold upgrade | Tenant restriction and conflicting resolved stage | | Tenant A T4 with S12 recommendations omitted | Evaluate explanation-only behaviour | Final instructions and independently judged answer | | Tenant B accepts S12 and output schema O3 | Evaluate the adopted workflow | Adoption record, resolved revisions and output contract | | Tenant A customization cannot be loaded | Follow declared unavailable-config policy | Load failure, snapshot validity or explicit refusal | | Run starts on S11; a worker reloads S12 mid-run | Reject unrecorded mixed-version execution | Run-bound revision and stage-level lineage |
Inspect the final rendered messages, output parser and applicable retrieval filters together. A configuration label can be correct while a worker cache contains different text. A recommendation may arrive through a retrieved document rather than the shared template. Record the first observed divergence and distinguish resolver defects from model interpretation or source contamination. Do not classify every failure as a bad prompt and change several stages without identifying which change addresses the evidence.
Include a negative control that deliberately ignores T4 and always assembles S12. The harness must catch the tenant restriction being lost. Then include a compatible configuration whose model answer still suggests an action. This second case ensures that passing resolver checks does not bypass behavioural review. Assess unsupported recommendations, useful explanation, justified refusal, latency and review effort separately. A global average can conceal a tenant-specific regression when most requests use the uncustomized default.
Bind long-running work to an explicit configuration decision
A request may pass through intake, retrieval, drafting and evaluation on different workers. If each stage fetches the current default independently, a shared update can create a combination that was never tested. Pin the validated configuration bundle at admission, or define an explicit transition protocol that records and validates the change. A digest helps identify a combination; it does not make that combination approved or safe.
AWS AppConfig's deployment strategy documentation describes staged configuration release, monitoring during bake time and entity-based rollout support with the specified agent version. Those are useful deployment mechanisms. They do not by themselves bind every application stage, queue retry or model call to a single resolved tenant bundle. Verify how the selected deployment mechanism interacts with your own run state, caches and workers rather than assuming host-level rollout gives tenant-level consistency.
Suppose run Q8 starts on S11 and pauses awaiting a document. After S12 rolls out, its resumed worker must not quietly substitute S12. Under this fixture policy, it either continues with the recorded, still-permitted S11 bundle or holds for an explicit restart. A restart is a new configuration decision, not an invisible continuation. Revalidate the current source and permission context separately: pinning old task instructions must not preserve access to material a user can no longer read.
Security withdrawal needs a distinct path. If an execution restriction changes while Q8 is paused, the service must apply the current mandatory restriction before any effect. Keeping S11 for reproducible drafting cannot override that restriction. Record both the advisory configuration and the current authorization decision. This separation makes it possible to reproduce the explanation without pretending that historical permission remains valid indefinitely.
Rollback restores a selection, not every consequence
If S12 causes a tenant-specific regression, stop expanding the affected rollout and identify which runs resolved the candidate. Restoring the S11 selection changes future admission under the declared resolver policy. It does not erase answers already displayed, retract exported documents or settle a pending external action. This is a recovery limit, not an argument against rollback. The team needs both a configuration reversal and a disposition for work that already crossed a delivery boundary.
Check cache refresh and queued work during reversal. A worker can continue serving S12 after the control record points to S11, and a delayed job can carry its own recorded bundle. Compare actual execution lineage with the rollback target. If safe continuation is unknown, hold the affected workflow rather than deleting records or rerunning it blindly. Preserve uncertain effect state for any downstream integration; a timeout does not prove that no operation happened.
Keep the support statement precise. The team may know that no new S12 runs were admitted after a certain observation, but that is not the same as all outputs corrected. Report affected tenants, known resolved bundles, missing execution evidence and any outputs requiring review. Do not claim complete containment from one healthy worker or from a global error rate returning to normal.
Start with one tenant and one shared change
For the next review, bring one actual tenant customization, its current resolved bundle and a proposed shared-template change. Identify which fields inherit and who owns adoption. Build the A, B and unavailable-configuration cases before expanding to a larger matrix. Exercise a paused run across the update and reversal. Use a non-sending harness so the exercise cannot contact customers or change business records.
Acceptance criteria should cover both assembly and behaviour. The tenant restriction remains explicit; the proposed bundle has traceable inputs; contradictory combinations are rejected; expected task outcomes come from an independent fixture policy; resumed work has recorded configuration lineage; and rollback includes already delivered work. Passing these bounded checks does not certify every customer customization or every future model update. Record which combinations were reviewed and what remains unsupported.
Use the production AI evaluation workbook to define independent expectations. The AI change-release whitepaper covers the wider release decision. Ampity's AI evaluation services and RAG knowledge systems work can help scope a configuration and acceptance review. These resources can be read without submitting an enquiry; contact is optional.