AI Software Supply Chains: Evidence Before Production Admission

Separate artifact identity, build provenance, model loading, behavior tests and tool permissions. Design production admission and withdrawal without treating a...

audience="AI platform owners, release engineers and security reviewers deciding what may run with production data or credentials." decision="What evidence admits a particular model artifact, dependency and tool runner into a specified production workflow, and what withdraws that admission." position="Evaluate origin, executable content, behavior and authority separately. Bind the decision to exact artifacts and operating conditions; do not turn a verified signature into unrestricted permission." scope="A proposed admission framework with a hypothetical document-processing assistant. Not a security certification, a complete threat model or an Ampity customer implementation." outputs={['A release evidence contract', 'An artifact and permission inventory', 'A comparison of admission options', 'A change invalidation map', 'A withdrawal and recovery record', 'A production decision checklist']} />

Executive summary

An AI application can load a model, install a dependency, execute a tool runner and transmit information to a remote service within one workflow. These actions depend on different kinds of trust. Knowing who published a package does not establish that its code is acceptable. Knowing which container bytes will run does not establish that the container should receive a production credential. A passing answer-quality test does not establish that a document parser has no dangerous installation behavior. The release decision needs those questions answered separately, then brought together for the actual business use.

Thesis: production admission should bind exact inputs, evidence, permission scope and an accountable decision, rather than assign a permanent trusted label to a vendor, repository or model family. Evidence must describe what was examined and what remains unknown. When a relevant input changes, the system should identify which evidence needs renewal. When trust is withdrawn, operators need a way to prevent new use and investigate existing effects without destroying the evidence they need for recovery.

This paper proposes an evidence contract for a hypothetical document assistant that extracts information and prepares a customer-record update for approval. It considers application dependencies, model weights and loaders, tool implementations, external endpoints and configuration. It recommends separating acquisition, evaluation and production operation because each stage needs different access. It also compares lighter admission designs for small teams with stronger controls for high-consequence workflows. The goal is an explainable release decision, not an ever-growing security checklist that nobody can operate.

Scope, assumptions and the actual decision

The illustrative assistant reads uploaded business documents, extracts candidate fields and asks an authorized employee to approve a proposed update. It uses an application image, a document parser, a model or model endpoint and a runner for the approved write operation. The organization owns the application and the business decision. Some software comes from external maintainers. The model may be downloaded or accessed as a service. These are assumptions for discussing the design, not evidence that Ampity or a named customer runs this architecture.

The first question is what crosses into production. A downloaded weight file, its associated configuration and the code used to load it are separate inventory entries. A remote endpoint has no application-owned weight digest; it instead needs a record of the selected service, available version controls, request handling assumptions and observed behavior. A tool description is not its implementation. A prompt is not a dependency manifest. An inventory that conflates these objects cannot explain which change caused a new risk or which release must be withdrawn.

Define the consequence before the control. A public-information assistant that returns links has different exposure from a runner that updates financial records. Record the data categories, permitted actions, external recipients and recoverability of mistakes. State who may approve exceptions and which uncertainty makes release unacceptable. This framework assumes the team can constrain deployment and credential issuance. Where an unmanaged desktop agent downloads and runs arbitrary software outside those controls, the admission design below does not cover the real execution path. That path needs its own ownership and restrictions first.

Four questions that cannot substitute for one another

Use four headings in the release record: origin, content, behavior and authority. Origin asks whether the object comes from the expected publisher and build context. Content asks what executable material, dependencies and loading mechanisms are present. Behavior asks what the proposed system does under defined tests and failure conditions. Authority asks which data and actions the running component can access. A positive answer under one heading is not evidence under another. Keeping the headings separate makes incomplete decisions visible without pretending that uncertainty has been eliminated.

For example, a signed tool runner can originate from the expected team and still implement an overly broad update endpoint. A known image digest can identify the bytes while a secret injected at runtime grants more access than the test environment provided. A model can pass a field-extraction evaluation while the surrounding loader imports unreviewed code. Conversely, an unsigned internal artifact is not automatically malicious; it has missing origin evidence that the release owner must resolve through the approved process. Avoid converting missing evidence into either a guaranteed compromise or a guaranteed safe exception.

The proposed admission record should therefore carry separate results, explicit limitations and a scoped final decision. It should not calculate a single reassuring score by averaging away a failed permission boundary. Some conditions are hard stops, such as an unidentified artifact where identity is mandatory or an uncontrolled write credential in a read-only workflow. Other findings may have compensating controls. Name the person who accepts those controls, the evidence supporting them and the condition that expires the exception. An unnamed team acceptance is difficult to challenge during an incident.

Identify the objects, not just their names

Build an inventory that can distinguish the reviewed object from the deployed one. A package name or image tag is a discovery reference, not sufficient release identity. Docker documents that an image digest identifies content while a tag can be reused or changed. Use the resolved digest in the release record and deployment configuration where supported. Record the platform-specific selection when a multi-platform image is involved. A digest provides an exact reference; its existence does not independently establish an acceptable publisher or safe behavior. See Docker's image digest documentation.

For the hypothetical assistant, inventory the application image, parser dependency set, model artifact or endpoint selection, loader code, tokenizer and configuration, tool runner image, prompt revision and permission policy. Include referenced artifacts downloaded at startup. If the application fetches a fresh component after deployment, reviewing only the original image is incomplete. Either make that fetched component part of the admission contract or prohibit the fetch in production. Include build-time generators where they can influence output, but distinguish build dependencies from runtime dependencies so response teams know where each object was used.

An inventory also needs an owner and an observation path. Explain how the running service reports its release identity, how deployment records preserve the selected artifact and how an operator compares them. Avoid logging full prompts or customer documents just to establish version identity. Keep identifiers and evidence references where those are enough. For a hosted model, record what the provider actually exposes and mark unavailable identity details as unknown. Do not invent a digest for a service you cannot inspect. An honest limitation is more useful than a false precision that misleads incident scoping.

Verify origin against an expected identity

Signature verification needs a trust policy, not merely a valid cryptographic operation. Decide which signer, issuer or key is expected for the selected artifact. Sigstore's verification documentation shows identity-based checks that specify both certificate identity and OIDC issuer. It also distinguishes signature verification from checking claims in the payload. The design implication is to store the expected identity and the checked artifact relation, not just a command exit code. A valid signature from an unrelated publisher should fail the organization's admission policy. See Sigstore's verification documentation.

Treat provenance as evidence about a build, then evaluate it against expectations. The application team should decide which source repository, revision, workflow and builder are acceptable for its own releases. External artifacts may offer different evidence. Document the actual coverage rather than assuming every supplier uses the same system. If the record proves publication but not the intended build input, do not label it a complete build verification. A verifier that accepts any available attestation makes the evidence look stronger than the rule it enforces.

Origin verification also has operational dependencies. An unavailable evidence service must not silently turn mandatory verification into acceptance. Decide whether locally retained verification material is sufficient for the relevant case and test that path. Separate an evidence outage from a failed identity check in alerts, while keeping the release result fail-closed when required. Name the mechanism for approved exceptions and prohibit ad hoc disabling of claim checks to make a pipeline green. The goal is not to stop every release during every outage; it is to ensure that the fallback has a reviewed contract instead of a hidden bypass.

Provenance is not a content review

npm explicitly notes that provenance can link a package to source and build instructions without guaranteeing that it contains no malicious code. That distinction matters when an AI coding assistant suggests a plausible dependency or an engineer adds a library under delivery pressure. The package still needs selection and review appropriate to its role. Provenance can make that review more traceable, but it cannot decide whether the functionality, maintenance risk or executable installation behavior is acceptable. See npm's provenance documentation.

For the proposed assistant, assess the parser's relevant dependency graph and its behavior during installation and execution. Record the lockfile used for the build, resolved sources and changes from the prior release. A manifest alone may not identify transitive versions. A dependency inventory alone may not identify how a build script modifies the resulting image. Ask which steps execute code, which steps need network access and which credentials are present at those steps. Acquisition should not happen in an environment that already holds the customer-record write credential simply because that environment is convenient.

Review findings in relation to the actual use. A scanner result can identify a known issue, but the release decision also needs a reasoned assessment and a disposition. Avoid declaring a zero-finding report proof of safety. Record scanner coverage and the artifact examined. If a vulnerable parser is replaced, confirm the resulting image and behavior rather than only changing the manifest. If a finding is accepted temporarily, assign a corrective action, expiry and deployment scope. Evidence from a different artifact is not a substitute even when the dependency names look the same.

Model loading creates its own execution boundary

Model artifacts are not all inert data. Hugging Face's pickle guidance describes arbitrary code execution risks during loading and explains that import scanning is not foolproof. The production decision must examine how the selected format is loaded, not only which organization hosts it. A signed origin can establish who supplied the file without establishing that loading it is safe. Avoid copying the older guidance's implementation details into a claim about every current library version; inspect the exact loader and configuration in use. See Hugging Face's pickle security guidance.

Safetensors offers a tensor-storage format designed as an alternative to pickle. Choosing such a format can address a particular loading concern, but it does not review the surrounding application code, prove weight quality or authorize data access. A model bundle can still require configuration, tokenization and executable libraries. The proposed rule is to examine the complete loading path and minimize executable material, not to stamp the whole system safe based on a file extension. See the Safetensors documentation.

Custom model code needs explicit treatment as code. Transformers documents that loading a custom model may require trust_remote_code=True. That setting is an execution decision, not a harmless compatibility toggle. In the proposed design, review the required code and bind its revision before approving use. Test loading in an isolated environment without production secrets and with controlled network access. Record any downloads and generated files. See Transformers' custom model documentation. An inability to operate the model without uncontrolled startup downloads is an unresolved admission condition, not a detail to hide in deployment instructions.

Compare admission options before building a control plane

Teams do not all need a new central service. A small application with a limited artifact set may use a reviewed manifest, controlled build job and deployment checks owned by the release engineer. This can be effective when the inputs are observable and the procedure is consistently followed. Its weakness is dependence on manual evidence gathering and the risk that one team silently changes a shared rule. The design should still preserve exact identities, explicit exceptions and a withdrawal procedure. Simplicity is valuable when it does not hide missing decisions.

| Option | Useful when | Main operational trade-off | Evidence needed before choosing | | --- | --- | --- | --- | | Release-owned manifest and checks | Few applications and predictable inputs | Less infrastructure, more disciplined manual review | A rehearsed release and withdrawal record | | Shared admission policy and artifact catalog | Several teams reuse components | Consistent decisions, additional platform dependency | Policy ownership, failure behavior and support capacity | | Restricted execution with limited supplier evidence | A necessary component has incomplete provenance | Smaller exposure, unresolved supplier uncertainty | Isolation tests, data limits and a scoped exception | | Decline the component or retain a non-AI path | Consequence exceeds available evidence | Lost automation benefit, clearer risk boundary | Business impact and a viable alternative process |

The third option must not become a blanket excuse for unknown software. Restrictions need to be observable: no production write credential during loading, permitted destinations enforced outside the component, and a clear boundary on input data. Test whether those restrictions survive deployment changes and exception handling. The fourth option is a valid engineering decision. A faster extraction workflow is not worth accepting an unidentified executable with unrestricted access. When choosing between these options, compare the cost of control ownership and recovery as well as the expected benefit of the feature.

Bind the release decision to operating conditions

The evidence record should describe a release tuple: exact artifacts, evaluated configuration, permission policy, intended data use and supported environment. This is a proposed design vocabulary, not a standard that every vendor implements. The tuple helps answer whether the current deployment is still the one that was reviewed. If the prompt or allowed tool set changes behavior materially, the tuple needs a new decision even when the application image is unchanged. If a remote service can change without a stable revision contract, preserve that limitation and add monitoring or reduced authority appropriate to the workflow.

Use a structured decision record with fields for object identities, evidence references, test results, unresolved findings, scope, approver and invalidation conditions. Do not embed secrets or full sensitive test documents in it. Link to controlled evidence stores with the appropriate retention and access rules. Distinguish admission from publication: a component may be acceptable in a development sandbox and unacceptable in a production workflow. An approval to experiment does not authorize processing all customer data. Capture the environment and permitted data class explicitly so operators are not left interpreting a generic approved label.

Record the verification policy version as well as the evidence. A result from an older rule can remain valid only if the current policy permits that relationship. Define how policy changes trigger reevaluation and which changes require immediate restriction. Avoid expiring every decision on an arbitrary short interval if that creates unmanageable reapproval work without improving evidence. Instead combine scheduled review with event-driven invalidation: changed artifact, changed signer policy, newly relevant finding, changed data exposure or expanded tool authority. Test the event mapping with a real release record before automating it across the platform.

Production permissions are a separate admission condition

The tool runner should receive only the authority required for its approved action. For the illustrative assistant, extracting candidate fields is not the same as approving a customer-record update. The business service should enforce the write rules using the caller's identity, target record and current approval. A prompt that says not to update without permission cannot replace that check. The release record should specify whether the runner prepares a proposal, reads authoritative state or performs a committed change. Those distinctions determine which credentials and recovery controls are appropriate.

Keep acquisition and loading away from production credentials. A download job may need access to an artifact source but not the customer database. A behavior test may use synthetic records without the production write token. A production runner may need an action-specific identity while being unable to fetch arbitrary new executables. These separations are proposed controls that must be enforced by the actual platform. Drawing three boxes on a diagram is not isolation evidence. Observe network access, filesystem access, secret injection and identity behavior under the deployed configuration.

Review what happens when the tool fails or is manipulated. Unexpected arguments should be rejected by the authoritative service. Retried writes need a business-level recovery contract so that uncertainty does not create duplicate effects. Returned text from a tool or document should not change the caller's permission scope. Include rejected operations in the acceptance evidence, but avoid logging their sensitive payloads unnecessarily. The operational and security owners should agree on which observations establish that the runner stayed within its authority and which alarms require suspending it before further requests are admitted.

Changes invalidate different parts of the evidence

A changed parser dependency can affect executable content and output behavior. A changed model weight can affect extraction results without changing the tool implementation. A changed tool permission policy can expand consequence without changing either artifact. A changed build signer policy can invalidate origin acceptance even if the bytes stay the same. Create a change-to-evidence map so each release renews the relevant checks. Do not require a generic full review that nobody can explain, and do not assume that passing unchanged tests covers a new exposure those tests never exercised.

In the proposed map below, the evidence families are intentionally independent. A model change triggers loading-path and behavior assessment; a permission change triggers authority checks and may require new behavior cases; a signer-policy change triggers origin reevaluation. Actual dependencies can be broader. If a loader change also modifies network destinations, it affects more than content. The release owner should document these relationships for the application, rather than treat the diagram as a universal dependency list. The useful outcome is an explainable scope for the next review.

The map also supports incident response. A newly identified unsafe loader should reveal which active releases contain it and which environments executed it. A withdrawn signer should reveal decisions that relied on that identity. That lookup requires retained inventories and an observed deployment identity. A directory of PDF approvals without machine-readable relationships will be slow to reconcile under pressure. Start with a small accurate catalog and prove the lookup. A large catalog that omits startup downloads or exception deployments gives an impressive count and incomplete incident scope.

Withdraw trust without losing the recovery record

Withdrawal should stop new admissions and address already running instances separately. Blocking a deployment pipeline does not necessarily terminate existing workers or invalidate credentials. Define what the application can safely suspend, drain or disable when an artifact is rejected after release. Preserve the identifiers needed to determine which work used it. For an extraction-only component, hold affected proposals for review. For a runner with write capability, assess completed and unresolved effects against authoritative business state. Do not assume rollback of code reverses a committed record update.

The hypothetical assistant illustrates this difference. If a parser is found to transform a particular field incorrectly, replacing it prevents future proposals from using that version. Existing proposals still need identification and reconsideration. If a compromised runner may have made changes, the investigation must examine actual record effects and access, not merely the output messages retained by the agent. Freeze further execution where appropriate, retain evidence under the incident procedure and assign recovery decisions to the business owner. Recovery may require compensating actions rather than replaying every request through the new runner.

Test emergency withdrawal before relying on it. Demonstrate that an affected identity can be located, new work is denied, existing jobs are handled according to policy and the on-call engineer can determine the resulting state. Include a case where the primary evidence store is unavailable. The recovery record should name unresolved items instead of closing the incident when the deployment is green. If a new replacement release cannot be admitted safely, retain the non-AI workflow or manual review path. Continuity does not justify granting a replacement broader authority than the removed component had.

Evidence that is useful to the next operator

Evidence should allow another engineer to reproduce the reasoning, not just admire a collection of green icons. Keep the selected artifact, expected identity, verification policy, observed result and decision together. For behavior tests, record the evaluated configuration, input class, expected outcome and failures. For permissions, preserve the relevant policy identity and observed denied operations. Where evidence is unavailable, state that fact and the consequence for admission. A record that says scanner passed without identifying its object or coverage leaves the next reviewer guessing.

Minimize sensitive material. A release evidence store does not need complete customer documents when controlled synthetic fixtures can exercise the same rule. If a case genuinely requires a restricted real example, use the approved access and retention procedure and keep that evidence separate from a broadly visible catalog. Be careful with public transparency records and logs. Before publishing attestations or annotations, inspect whether they disclose private repository names, internal paths or other information the organization did not intend to expose. Evidence generation is itself a data-handling workflow with an owner and an access boundary.

Measure control usefulness rather than report only coverage percentages. Useful observations include releases with unresolved identity, deployments that differ from their approved tuple, exceptions past their expiry and the time needed to identify affected running instances during a drill. These are suggested indicators, not benchmark targets. A high proportion of signed artifacts can coexist with overly broad permissions. A low number of scanner findings can coexist with missing artifact inventory. Pair each indicator with a question an operator can answer and a defined action when the result changes.

Limitations and cases that need a different design

This framework does not prove the absence of malicious code, establish complete model training provenance or guarantee that every future input produces acceptable behavior. It organizes admission evidence and authority around the information the application owner can actually control. Hosted services may expose limited build or version detail. The appropriate response is to record that uncertainty and constrain the use, seek stronger evidence or choose another process. Do not fill the gap with a certificate-like claim that the provider did not make and the application owner cannot verify.

Not every application needs the same gate strength. A low-consequence internal prototype with synthetic data can have a lighter, explicit sandbox decision. A production workflow using sensitive information and irreversible writes needs stronger evidence and enforcement. The distinction is the operating scope, not whether someone calls the system a pilot. A pilot that quietly receives production secrets has production exposure. Revisit the decision when data, users or authority expand. The paper also does not replace specialist assessment of licensing, contractual obligations, privacy rules or regulated deployment requirements.

Finally, supplier trust and runtime trust remain partly independent. A carefully reviewed component can receive adversarial input. An exact artifact can run in an incorrectly configured environment. A remote service can be unavailable without being compromised. Keep runtime validation, monitoring and recovery work connected to admission without claiming admission replaces them. Use the AI change release governance paper for behavior-change review and the AI action recovery paper for uncertain writes. This paper's contribution is explaining why artifact acceptance and permitted execution are different decisions.

Decision checklist and the first implementation slice

Start with one workflow and one release. Build its inventory from observed build and startup behavior rather than only the declared manifest. Select the consequence and record the permission scope. Gather evidence under the four headings, note the missing evidence and make a scoped decision. Then rehearse a relevant change and a withdrawal. The first implementation slice is complete when another engineer can identify the running artifacts, explain their admission and stop an affected component without guessing which work it touched. Do not roll out a platform-wide trusted label before that drill works.

  • Identify every executable component and startup download in the selected workflow.
  • Bind the deployed artifacts or available service revision controls to the record.
  • State the expected signer, issuer or source evidence where origin verification is required.
  • Review installation and loading behavior separately from provenance.
  • Name the behavior cases and configuration actually evaluated.
  • Enforce data access and write authority outside model instructions.
  • Give each unresolved finding an owner, disposition and review condition.
  • Map relevant changes to the evidence that needs reconsideration.
  • Prove that a withdrawn component is identifiable in running deployments.
  • Reconcile affected proposals and business effects under the recovery procedure.

The next useful artifact is one completed admission record and the withdrawal-drill result. If they expose an uncontrolled download path or an oversized credential, fix that boundary before buying another scanning tool. For dependency-selection steps, use the AI-generated dependency verification guide. For a tool-server review, use the MCP trust review playbook. These companion resources address narrower tasks; none substitutes for deciding whether the full workflow may operate within its proposed production authority.