Choose an AI consulting partner by testing whether the proposed team can turn your specific workflow into a measurable, secure, operable capability that your organization can own. Do not select primarily on a framework logo, generic demo, or claims about model access.
The evaluation should produce evidence before a large commitment: a clear problem contract, a representative test set, an architecture and risk boundary, a working vertical slice, and an exit or transfer plan.
Prepare before contacting vendors
Write a one-page buyer brief:
user and workflowcurrent process and baselinedesired outcome and unit of valuevolume, latency, and availability needscost of incorrect output or actiondata sources, owners, sensitivity, and access constraintssystems that must integratelegal, security, and regional constraintsinternal owner and available subject-matter expertsdecision date and initial budget boundary
Do not prescribe “build a multi-agent RAG platform” unless the architecture is already justified. The UK government's AI procurement guidance recommends describing the challenge and remaining open to alternative solutions, assessing data before procurement, and using iterative proof-of-concept mechanisms.
Use knockout questions first
Reject or pause a vendor that cannot answer:
- Who exactly will deliver the work, and what similar system decisions can they demonstrate?
- How will success and critical failure be measured before launch?
- Where will our data, prompts, traces, and derived artifacts be stored and processed?
- How are identity, authorization, tenant isolation, deletion, and incident response handled?
- What remains ours if the engagement ends?
- Which components can be replaced without rewriting the product?
- What does the team need from us, and who owns operation after handoff?
Ask for evidence, not assurances. “Enterprise-grade security” should resolve into controls, boundaries, test results, ownership, and contract terms.
Buyer scorecard
Score each category from 0 to 3 and multiply by weight:
0 — absent or generic claim1 — plausible approach, no relevant artifact2 — relevant evidence with limitations3 — demonstrated on our representative task or data
| Category | Weight | Evidence to request |
|---|---|---|
| Problem and product judgment | 15 | Baseline, non-AI alternative, task contract, failure cost |
| Delivery team | 10 | Named people, roles, availability, comparable decision artifacts |
| Evaluation | 15 | Dataset schema, metrics, critical gates, regression process |
| Data/security/privacy | 15 | Data flow, threat model, access/deletion tests, incident process |
| Architecture/operations | 10 | Boundaries, objectives, traces, fallback, rollback, cost model |
| Proof of capability | 15 | Working vertical slice on a controlled representative set |
| Knowledge transfer | 10 | Documentation, pairing, runbooks, training, ownership milestones |
| Commercial/exit terms | 10 | Deliverables, IP, third parties, portability, support, termination |
Set minimum scores for security, evaluation, and exit terms. Do not allow a polished demo to compensate for a zero in a critical category.
Evaluate the actual team
Sales and leadership credentials do not prove who will write, test, deploy, or operate your system. Interview the proposed delivery lead and key engineers. Ask them to walk through one failure they found, how they measured it, and which permanent test or control followed.
Verify:
- named allocation and substitution rules;
- experience with your integration and regulatory constraints;
- software, data, evaluation, security, and product skills;
- who makes architecture and risk decisions;
- who joins incidents and for how long;
- which work is subcontracted.
Require evaluation before architecture theater
A credible proposal defines a representative evaluation set early. For an internal support agent, it might include answerable, unanswerable, permission-sensitive, stale-document, tool-error, and adversarial cases. Agree on task success, critical failures, latency, cost, and manual-review effort.
The partner should explain how model, prompt, retrieval, and tool changes pass a release gate. If quality is described only as “the model looked good in our demo,” the project has no acceptance contract.
Use the LLM evaluation dataset guide to assess the proposed test plan and LLM observability guidance to check the operating plan.
Run a bounded proof of capability
The proof should reduce the largest uncertainties, not imitate a complete product. Give shortlisted teams the same controlled problem, context, constraints, and scoring method where procurement rules permit.
Require these outputs:
- task contract and non-goals;
- data flow and trust-boundary diagram;
- baseline and one working vertical slice;
- evaluation result with failed examples;
- latency and cost measurement;
- known risks and next experiment;
- production and transfer plan;
- code, configuration, and documentation deliverables defined by contract.
Protect sensitive data during selection. Use synthetic or minimized representative data when possible and apply the same access requirements you expect in delivery.
Inspect data and security boundaries
Ask the partner to trace one request end to end:
user identity -> application -> model gateway -> retrieval/tools-> logs/traces -> evaluation store -> support access -> deletion
For each boundary, identify processor, region, encryption, retention, access, subprocessor, training use, deletion verification, and incident owner. The UK NCSC recommends applying existing supply-chain security expectations to AI components and documenting data, models, prompts, dependencies, and technical debt.
Security review should cover prompt injection, excessive agency, secret handling, cross-tenant access, unsafe tool output, dependency changes, model-provider outages, and auditability. A policy in a system prompt is not an authorization control.
Prevent avoidable lock-in
Not all managed services are bad, and portability is not free. Decide where lock-in is acceptable and where it threatens continuity or bargaining power.
Contract for access to:
- source code and deployment configuration;
- prompts, schemas, evaluation data, and test results;
- source/index/model/version records needed for reproduction;
- operational dashboards, runbooks, and incident history;
- data export and verified deletion;
- third-party licenses and recurring costs;
- transition support and termination assistance.
Prefer explicit interfaces around model providers, search, tools, identity, and traces when replacement is a realistic requirement.
Structure delivery around evidence gates
Use staged commitments:
| Stage | Exit evidence |
|---|---|
| Discovery | Task contract, baseline, data/risk assessment, prioritized uncertainty |
| Prototype | Working vertical slice and initial evaluation on controlled data |
| Pilot | Real integration, users, security review, operating measures, rollback |
| Production | Release gates passed, ownership/runbooks established, support active |
| Transfer | Team can deploy, evaluate, operate, and change the system independently |
Define what happens if a gate fails. Time-and-materials work still needs decision checkpoints and transparent artifacts.
Reference and contract questions
Ask references about the system six months after launch:
- Did the named team remain on the engagement?
- Which assumptions were wrong, and how were they surfaced?
- Could your staff reproduce evaluation results and deploy changes?
- How did the partner respond to an incident or provider change?
- Were operating costs and manual review close to the forecast?
- Could you continue without the partner?
Have counsel review liability, data protection, IP, confidentiality, warranties, security obligations, subprocessors, service levels, and termination terms for your jurisdiction and risk profile. This article is an engineering buyer framework, not legal advice.
Final selection rule
Choose the partner with the strongest relevant evidence and clearest limits—not the most ambitious promise. A good team may recommend a smaller automation, a search interface, or deterministic workflow instead of an agent. That is product judgment.
Before signing, compare the proposal with the first production LLM application checklist and use the AI prototype-to-production plan to make milestones and ownership explicit. If you need a scoped technical partner, review JoinAI consulting against the same scorecard.
When you are ready to invite structured proposals, copy the AI automation project RFP template so every supplier responds to the same outcome, evidence, security, transfer, and exit requirements.




