Engineering Partner Evaluation Playbook

A practical partner-evaluation method based on a bounded delivery proof, secure development evidence, production behavior, recovery, continuity, commercial clarity,...

trigger="A company needs an external engineering partner for a material product, platform, cloud, data, AI, modernization, or production-support outcome and wants evidence beyond proposals, interviews, certifications, or hourly rates." owner="One executive or engineering sponsor accountable for the intended outcome, paired with one technical evaluator who can inspect delivery and production evidence. Security, procurement, legal, finance, product, and operations retain decisions in their remit." timebox="Two weeks for evidence review and a representative working session, followed by a two-to-four-week bounded proof when the decision warrants it. Do not begin unrestricted production work during evaluation." participants={["Executive sponsor", "Technical evaluator", "Product owner", "Platform or operations owner", "Security", "Procurement or commercial owner", "Legal or privacy owner when relevant", "Proposed partner delivery lead", "Engineers who will do the work"]} prerequisites={[ "One named outcome or production problem, current constraints, retained client responsibilities, and material risks.", "A representative slice that can test discovery, engineering, review, security, release, support, and handover without exposing the full estate.", "Approved access boundaries, test data, evaluation environment, evidence-retention rules, and a stop decision." ]} outputs={[ "A work-boundary and responsibility record covering scope, assumptions, dependencies, decisions, acceptance, change, and support.", "A partner evidence pack covering people, delivery path, secure development, production access, artifacts, operations, recovery, and continuity.", "A bounded-proof result with accepted evidence, returned work, incidents, decisions, and observed collaboration costs.", "A risk and dependency register with client owner, partner owner, fallback, expiry, and exit treatment.", "An expand, correct, retain-specialist, or stop decision with reasons and next actions." ]} doneWhen={[ "The buyer has observed the proposed team complete representative work through the intended delivery controls.", "Security and production claims link to inspectable artifacts or observed behavior.", "The client can access, build, deploy, operate, recover, and continue the bounded slice under the agreed model.", "Commercial terms align with measurable scope, client dependencies, change treatment, acceptance, support, and exit.", "Open risks and gaps have owners and dates, and the decision does not depend on unsupported claims." ]} />

Define the outcome before comparing partners

Start with the production or business condition that must change. “Add developers” is an input. “Reduce failed releases for the registration journey while modernizing its deployment path” is closer to an outcome that can be evaluated.

Describe current behavior, affected users, system boundary, constraints, non-negotiable controls, desired evidence, and client dependencies. State what the client retains: product priority, data decisions, access approval, acceptance authority, legal judgement, and any platform or operational ownership that will not transfer.

Separate known delivery from discovery. A partner can commit to producing an architecture decision, dependency map, tested slice, migration plan, or risk baseline when implementation scope is still uncertain. Forcing a fixed build promise before evidence exists rewards optimism.

Write the stop conditions. Examples include inability to segregate production access, no reproducible build, unclear intellectual-property rights, refusal to retain change evidence, critical skill substitution without approval, or a design that cannot meet a named recovery constraint.

Choose a representative proof

Select a small slice that exposes the way the real engagement must work. It should require clarification, an architectural judgement, code or configuration, review, security, tests, release evidence, observability, and handover. Avoid a polished isolated prototype that bypasses the client’s systems and controls.

The proof could be one API integration, deployment path, production defect class, cloud-cost change, data pipeline, AI evaluation gate, tenant-isolation test, or thin user journey. Use non-production access unless production observation is necessary and explicitly approved.

Define its outputs and done-when evidence before starting. Set a timebox, access scope, data class, budget, rollback, repository ownership, and treatment of artifacts if the relationship does not continue.

Do not make unpaid speculative work the default evaluation method. A meaningful proof consumes engineering effort and produces reusable artifacts. Use a paid bounded engagement where appropriate and judge the work, not willingness to absorb sales cost.

Evaluate the people who will do the work

Meet the proposed delivery lead and representative engineers, not only executives or solution architects. Ask them to explain a relevant decision, production failure, trade-off, or recovery in concrete terms while protecting prior-client confidentiality.

Assess whether the team asks about users, constraints, dependencies, data, operations, acceptance, and failure before prescribing technology. Strong engineers can say what they do not know, how they will find out, and what would change their recommendation.

Record role, expected allocation, time zone overlap, communication pattern, replacement process, escalation, and who reviews work. Do not infer quality from senior titles alone. Clarify which named people are committed and which are illustrative.

Test collaboration through the proof: response to ambiguous input, written decisions, review quality, ability to challenge unsafe requests, follow-through on dependencies, and respect for client context. Presentation fluency is not delivery evidence.

Inspect the delivery system

Follow one change from request to accepted production evidence. Inspect how the team records outcome, scope, assumptions, client dependencies, decision owner, design trade-offs, implementation, tests, review, security checks, release, rollback, telemetry, and acceptance.

Ask for representative sanitized artifacts or observe the process during the proof. Useful evidence includes an architecture decision, pull request, test failure, security finding, deployment record, incident timeline, runbook, reconciliation result, and acceptance record.

Check batch size and feedback time. A partner that reports progress through activity can still reveal risk late. Working slices, visible blockers, and explicit returned items make the state inspectable.

Examine how scope change is handled. The team should distinguish clarified requirement, corrected defect, changed assumption, new outcome, and client delay. Every change does not need a commercial dispute, but invisible absorption creates schedule and quality problems.

Review secure development evidence

Use a recognized framework as a vocabulary, not a badge. NIST SP 800-218 describes Secure Software Development Framework practices that purchasers can use in supplier conversations. Ask which relevant practices operate for this work and inspect their evidence.

Review development identity, repository access, branch controls, peer review, secrets handling, dependency management, vulnerability response, build isolation, artifact identity, test integrity, environment segregation, logging, and incident response. Tailor depth to the workload and data.

The SLSA build specification provides a useful model for artifact provenance and build integrity. Do not demand a level without understanding the current platform. Ask whether the delivered artifact can be tied to reviewed source and a controlled build, and what remains outside that chain.

Inspect how the partner handles a finding. Evidence includes triage, reachability and consequence analysis, remediation, test, release, disclosure, and root-cause prevention. A scan with unresolved findings is an input, not proof of secure delivery.

Assess assurance maturity without scoring theatre

Use maturity models to expose missing practices and inconsistent execution, not to manufacture a single reassuring grade. OWASP SAMM groups software assurance work across governance, design, implementation, verification, and operations. That breadth is useful because a partner can have strong scanners and weak threat modelling, or careful design review and no reliable vulnerability response.

Select the practices that match the proof's data, threat, availability, and change risk. Ask for one recent artifact and one observed behavior for each selected practice. A written policy shows intent. A repository rule, reviewed threat model, dependency decision, failed security test, remediation record, or incident exercise shows how the practice operates. Record whether the evidence applies to the proposed team and delivery path, rather than assuming an organization-wide policy reaches every project.

Do not give every practice equal weight. Missing identity control for production access can be a hard gate. An immature reporting dashboard may be a correctable gap. Record the consequence, current control, required treatment, owner, due date, and evidence that will close each gap.

Turn the assessment into a delivery roadmap. The first bounded outcome should include the assurance changes required to make that outcome safe. Broader improvements can be sequenced with explicit risk acceptance and review dates. A maturity label without owners, dates, and acceptance evidence is not an operating control.

Reassess the selected practices when the workload, data class, production authority, or delivery team changes.

Bound production access

List every partner identity, environment, resource, data class, action, access path, approval, expiry, and evidence source. Prefer named identities, least privilege, time-bounded elevation, client-controlled trust, and attributable sessions.

Separate build access, deployment authority, runtime support, data access, security administration, and emergency access. A developer who can change code does not automatically need standing database or cloud-administrator permission.

Test removal and emergency use. Confirm that the client can revoke access without partner cooperation, see privileged activity, rotate affected credentials, and continue operation. Avoid shared accounts and secrets sent through informal channels.

CISA’s Secure by Demand guide encourages buyers to evaluate product security rather than relying only on enterprise compliance. Apply the same distinction to engineering delivery: organizational controls matter, but the delivered system and its defaults need direct evidence.

Evaluate architecture and engineering judgement

Give the team a real constraint and ask for options, not a predetermined stack. The decision record should explain current evidence, alternatives, trade-offs, rejected options, reversibility, operational consequences, security, cost, and validation plan.

Check whether diagrams show authority, trust boundaries, data movement, failure, recovery, and ownership rather than boxes connected for appearance. Ask what is intentionally omitted and which decision the diagram supports.

Prefer the simplest design that meets the outcome and constraints. Unnecessary services, custom frameworks, premature distribution, or fashionable AI can create dependence without customer value. Conversely, a simplistic design that ignores known scale, recovery, or isolation is not efficient.

Observe response to challenge. A credible partner updates a recommendation when evidence changes and retains the reasoning. Defending every initial choice creates hidden risk.

Test production and recovery behavior

Release one reversible slice through the intended path when the proof permits it. Observe pre-deployment checks, infrastructure preview, migration, exposure, business and service telemetry, stop signals, rollback or roll-forward, and final acceptance.

Run a failure exercise: dependency unavailable, timeout after an external effect, failed deployment, delayed event, stale permission, corrupted configuration, or restore. Ask the team to detect, contain, communicate, recover, reconcile, and document the result.

Assess on-call and support boundaries. Name who watches, who responds, who can act, how severity is set, when the client is contacted, and which unresolved obligations remain after service restoration.

Do not accept a scripted demo as production proof. Retain timestamps, version, telemetry, action records, recovery time, discrepancies, and follow-up owners.

Inspect quality and acceptance evidence

Define acceptance before implementation. Evidence should correspond to the outcome and constraints: journey completion, tests, performance distribution, security controls, data reconciliation, tenant isolation, recovery, accessibility, cost, and operating readiness as relevant.

Review failed tests and returned work, not only the final green result. They show whether the process catches meaningful defects. Check whether tests are tied to requirements and failure modes rather than raw coverage percentage.

Ask the client reviewer to reproduce the decision from the evidence pack. If acceptance depends on the partner verbally explaining unrecorded context, handover is incomplete.

Distinguish a completed deliverable from an achieved business outcome. Some outcomes mature after release or depend on client adoption. State the observation window and how responsibility is divided.

Measure collaboration cost

Track client time spent clarifying work, resolving blockers, reviewing, coordinating environments, correcting misunderstandings, and operating the result. External cost alone hides retained management work.

Measure waiting and decision latency as well as coding activity. A partner can appear productive while work queues behind unclear ownership, unavailable access, or large review batches.

Record rework cause. New evidence, changed business decision, partner defect, client dependency, and discovered legacy constraint require different responses. Avoid using a single velocity measure to assign blame.

Compare the proof with the current arrangement where possible: elapsed time, accepted output, incidents, review load, operating burden, and complete cost. Small samples do not prove long-term performance, so preserve uncertainty.

Verify continuity and exit

Confirm where repositories, infrastructure definitions, build configuration, artifacts, documentation, diagrams, test data, secrets references, runbooks, dashboards, tickets, and decision records live. The client should have authorized access during delivery, not only after termination.

Test a handover. Ask a client engineer to build, deploy to an approved environment, find telemetry, follow a runbook, and explain one architecture decision. Record gaps and correct them during the proof.

List partner-managed accounts, licenses, domains, signing keys, certificates, data, and third-party relationships. Define transfer, replacement, export, deletion, and evidence. Repository ownership alone does not establish continuity.

Protect intellectual property and confidential information through appropriate legal review. This playbook does not determine ownership or applicable law. It identifies the technical artifacts and dependencies counsel and commercial owners need to examine.

Align the commercial model

Compare the commercial unit with the work boundary. Time and materials, dedicated team, fixed scope, milestone, managed service, and outcome-linked delivery distribute uncertainty and responsibility differently. None is automatically best.

Write included work, exclusions, assumptions, client dependencies, acceptance, change treatment, pause, invoice basis, support, incident work, travel or tooling, third-party cost, replacement, and exit. Avoid unrestricted claims that payment is guaranteed by satisfaction or that the partner absorbs every change.

For outcome-linked work, define the bounded outcome, accepted evidence, timing, assumptions, dependencies, and change process. Payment can be linked to accepted delivery without transferring product priority, legal judgement, unlimited scope, or client-caused delay.

Test one scenario: a critical client dependency is late, a provider limitation invalidates an assumption, or production evidence fails. The agreement and operating process should produce the same next action.

Score evidence, not adjectives

Use a decision record rather than a universal weighted score. For each criterion record required evidence, observed evidence, gap, consequence, confidence, owner, and treatment. A numerical score can summarize, but it should not let excellent presentation offset a missing production-access control.

Use hard gates for legal eligibility, confidentiality, unacceptable security exposure, irreproducible artifacts, prohibited data handling, or lack of continuity. Use comparative judgement for collaboration, domain learning, design quality, and operating fit.

Separate unknown from fail. Missing evidence may require a proof, contract condition, reduced scope, or rejection depending on consequence and timing.

Document dissent and conflicts. Procurement, engineering, security, product, and finance can value different evidence. The executive sponsor owns the final bounded decision.

Decide how to proceed

The decision options are not only select or reject:

  • expand the partner into the next bounded outcome;
  • correct the work, access, evidence, or commercial boundary before expansion;
  • retain the partner for a specialist scope while the client owns integration and operation;
  • run a second proof for a distinct critical uncertainty;
  • stop, revoke access, retain artifacts, reconcile obligations, and complete handback.

Set the next review trigger. Team composition, scope, production authority, data class, commercial model, major incident, acquisition, or new regulated requirement can invalidate prior evidence.

Final checklist

"One outcome, baseline, constraints, retained client responsibilities, evidence, dependencies, and stop conditions are explicit.", "The proof uses representative work and the people expected to deliver it.", "The buyer observed discovery, decisions, engineering, review, security, release, operations, and handover.", "Secure-development claims link to relevant artifacts or observed controls.", "Partner production access is named, least-privileged, time-bound where feasible, attributable, and removable by the client.", "Architecture decisions explain alternatives, trade-offs, failure, recovery, security, cost, and reversibility.", "One release or failure exercise produced inspectable production and recovery evidence.", "Acceptance evidence is tied to the outcome and can be understood without oral context.", "Client collaboration time, review load, waiting, rework cause, operating burden, and complete cost are visible.", "Repositories, builds, infrastructure, accounts, runbooks, telemetry, decisions, and knowledge remain usable by the client.", "Commercial terms align with scope, acceptance, dependencies, change, support, failure, and exit.", "The final decision records gaps, owners, conditions, dissent, review trigger, and handback tasks." ]} />

Failure and handback

Stop the proof if access exceeds the approved boundary, protected data is mishandled, artifacts cannot be retained, a hard gate fails without containment, or the partner substitutes the evaluated team or delivery path without agreement.

Revoke access, rotate affected credentials, preserve client-owned artifacts and evidence, reconcile open changes and external effects, return or delete data under the approved process, assign unresolved incidents, and confirm account and repository control.

Do not hide an unsuccessful proof. Record the conditions and evidence. The result can prevent a much larger delivery failure and improve the next work boundary.

The next action is to choose one representative slice and write the evidence the client must retain after the proof. That single exercise reveals whether the evaluation tests engineering delivery or only sales confidence.