An AI automation RFP should describe the business problem, evidence required for acceptance, data and authority boundaries, operating constraints, transfer expectations, and commercial exit—not prescribe fashionable architecture. Give suppliers enough context to propose a smaller or safer approach when it meets the outcome.
Download the copyable Markdown RFP template. Replace every bracketed field and have procurement, security, privacy, legal, data, engineering, product, and affected operational owners review the final document.
This is an engineering procurement template, not legal advice.
1. Procurement summary
Organization:Procurement owner and contact:Technical owner and contact:Question deadline:Proposal deadline:Decision and target start dates:Contract form:Budget range or boundary:Required proposal validity period:
State the procurement rules, communication channel, confidentiality process, and whether suppliers will be compensated for proof work. Apply the rules and evaluation information consistently.
2. Problem, users, and baseline
Describe:
- who performs or receives the workflow;
- current steps, systems, volume, cycle time, and manual effort;
- current quality/error baseline and how it was measured;
- user harm or business cost when the workflow fails;
- desired outcome and unit of value;
- why AI may be relevant;
- non-AI alternatives already considered.
The UK government's AI procurement guidance recommends output-based requirements backed by user needs and required performance, while staying open to alternative solutions. Do not ask for “an agentic platform” when the actual requirement is reducing invoice exception handling time without unauthorized payments.
3. Scope, authority, and non-goals
In-scope users, teams, regions, and languages:In-scope systems and data sources:Decisions the system may recommend:Actions it may draft:Actions it may execute:Required human approvals:Prohibited actions:Explicit non-goals:
Define autonomy at each action, not once for the whole system. A tool that reads an order and a tool that issues a refund need different authority and evidence.
4. Success and acceptance
Require suppliers to propose and run a versioned evaluation. Provide representative minimized cases where possible.
Specify:
| Area | Buyer requirement |
|---|---|
| Task outcome | Success definition and current baseline |
| Critical failure | Security, privacy, safety, legal, or financial events that fail the gate |
| Offline evaluation | Dataset slices, metrics, thresholds, review, uncertainty |
| Pilot outcome | User/task result, adoption, override, manual effort, incident rate |
| Operations | p95 latency, availability, throughput, recovery, support |
| Economics | Cost per successful task and usage assumptions |
| Decision gates | Evidence to proceed, extend, redesign, or stop |
Do not accept a single “accuracy” number. Separate task, retrieval, tool, policy, and critical-boundary measures.
5. Data and lifecycle requirements
Provide a source inventory or require discovery to create one:
source and ownerdata categories and sensitivitypurpose and permitted usequality and known limitationsaccess/identity modelprocessing and storage regionsfreshness objectiveretention and deletionderived artifacts: chunks, embeddings, prompts, traces, labels, fine-tunesexport and termination behaviorprovider training-use restrictions
Ask suppliers to identify missing data and assumptions before committing to model performance. NIST's AI RMF emphasizes documenting test sets, tools, deployment-like conditions, and production monitoring; the same evidence belongs in acceptance and operation.
6. Security, privacy, and safety
Require:
- end-to-end data flow and trust-boundary diagram;
- identity propagation and server-side authorization;
- tenant isolation and negative tests;
- encryption, secrets management, and privileged access controls;
- logging minimization and trace access/retention;
- prompt-injection and unsafe-tool-output testing;
- approval immediately before consequential actions;
- model, dependency, and supply-chain vulnerability management;
- incident classification, notification, response, and evidence retention;
- subprocessors, hosting regions, and deletion verification;
- applicable privacy, impact, and security assessments.
Ask for control evidence at the proof and pilot stages. Policy prose is not a test result.
7. Architecture and integration
Describe existing identity, APIs, databases, events, networks, deployment standards, observability, and support constraints. Ask the supplier to propose:
- the smallest architecture that meets the task;
- boundaries for models, retrieval, tools, state, policy, and evaluation;
- typed interfaces and error behavior;
- fallback, degraded mode, rollback, and kill switch;
- model/provider replacement and index migration approach;
- capacity, latency, and cost model;
- build-versus-buy decisions and third-party dependencies.
Require assumptions and rejected alternatives. A diagram without contracts is not architecture evidence.
8. Evaluation and monitoring
The response should include:
- representative dataset design and ownership;
- deterministic and human evaluation methods;
- calibration of model judges against reviewed labels;
- slice-level and critical-failure gates;
- baseline/candidate comparison in CI;
- production trace schema and sampling;
- feedback, override, appeal, and incident handling where relevant;
- drift, provider-change, and regression process;
- reproducible evaluation reports and retention controls.
Use the Python LLM evaluation harness as a minimum inspectable example, not as a mandated implementation.
9. Delivery and governance
Ask for a staged plan:
| Stage | Required exit evidence |
|---|---|
| Discovery | Task contract, baseline, data/risk assessment, architecture options |
| Proof | Working vertical slice, evaluation result, failed cases, cost/latency |
| Pilot | Real integration, users, security review, operating data, rollback |
| Production | Release gates, runbooks, ownership, on-call/support, acceptance |
| Transfer | Internal team can deploy, evaluate, operate, and modify independently |
Require named accountable roles, allocation, substitution terms, decision/risk logs, reporting cadence, change control, and what happens when a gate fails.
10. Deliverables and knowledge transfer
Define required ownership and formats for:
- source code and infrastructure configuration;
- prompts, schemas, policies, and tool contracts;
- datasets, graders, evaluation reports, and incident regressions;
- architecture, data flow, threat model, and decision records;
- dashboards, alerts, runbooks, recovery tests, and cost model;
- administrator, operator, reviewer, and developer training;
- pairing, documentation acceptance, and internal ownership milestones.
“Documentation included” is too vague. State which roles must perform which task without supplier intervention.
11. Commercial, IP, and exit terms
Ask suppliers to itemize:
- discovery, implementation, pilot, production, and support fees;
- model, hosting, data, tool, observability, and license costs;
- volume assumptions and cost-change triggers;
- subcontractors and third-party terms;
- rights to code, configuration, prompts, data, labels, and derived artifacts;
- warranties, liability, indemnities, and insurance for legal review;
- service levels, incident support, and change response;
- export formats, transition assistance, termination, and verified deletion.
Decide which managed lock-in is acceptable. Require replaceable interfaces only where portability has material value.
12. Supplier response format
Require every response in the same order:
- proposed outcome and non-AI alternative;
- assumptions, exclusions, and dependencies;
- architecture and integration;
- data, security, privacy, and safety;
- evaluation and acceptance plan;
- staged delivery and governance;
- named team and relevant evidence;
- knowledge transfer and exit;
- implementation and recurring cost;
- risks, deviations, and completed scorecard.
This makes comparison easier and exposes missing sections.
13. Scoring example
| Category | Weight |
|---|---|
| Problem and product judgment | 10% |
| Delivery team | 10% |
| Evaluation and acceptance | 15% |
| Data, security, and privacy | 15% |
| Architecture and operations | 10% |
| Proof of capability | 15% |
| Delivery and governance | 10% |
| Knowledge transfer and exit | 5% |
| Commercial response | 10% |
Set critical minimums. A serious data or authorization gap should not be offset by a low price. Use the detailed AI consulting partner scorecard during shortlisting.
14. Proof-of-capability brief
Give shortlisted suppliers the same controlled task, representative data, allowed services, time boundary, required artifacts, and scoring method where your procurement process permits.
Require:
- a working vertical slice;
- architecture and data flow;
- test dataset and result;
- failed-example analysis;
- security boundary and negative tests;
- measured latency and cost;
- next-stage risks and transfer plan.
The proof should reduce the largest uncertainty, not reward the most polished frontend.
Final checklist
- Requirements describe user outcomes rather than a preferred buzzword stack.
- Scope includes prohibited actions and approval boundaries.
- Data owners, limits, access, retention, and deletion are visible.
- Critical failures and slice-level acceptance gates are defined.
- Security requirements request evidence and negative tests.
- Delivery is staged with explicit stop decisions.
- Named team and substitution terms are requested.
- Deliverables allow internal evaluation and operation.
- Recurring costs and usage assumptions are itemized.
- IP, portability, support, termination, and deletion receive legal review.
Download the Markdown template, then align its milestones with the AI prototype-to-production plan.




