Steps
- 1
Define the job, owner, pass/fail criteria, test budget, and stop conditions.
- 2
Create a dated source log from current primary sources; separate vendor claims, verified facts, and missing evidence.
- 3
Verify access, pricing, limits, integrations, security, privacy, and data-use terms for the intended account and region.
- 4
Run representative non-sensitive cases and compare quality, errors, time, and cost with the current method or alternative.
- 5
Have a human reviewer check consequential outputs, policy fit, rights, data boundaries, and operational ownership.
- 6
Record one decision: test with safeguards, pause for named evidence, or reject with reasons. Do not invent a fit score.
What you should get
- Dated source and evidence log
- Trial record against acceptance criteria
- Test / pause / reject decision memo
- Owner, safeguards, monitoring, and stop conditions
Prompt
Run the AI Tool Evaluation Workflow for the goal below.
Goal: Decide whether to run a bounded test, pause for evidence, or reject an AI tool for a defined job.
Inputs (replace each placeholder with verified information):
- Job, target user, and pass/fail criteria: {{add verified input}}
- Official product, pricing, security, and privacy sources: {{add verified input}}
- Trial account, region, and access limits: {{add verified input}}
- Representative non-sensitive test cases: {{add verified input}}
- Current method or alternative: {{add verified input}}
- Test owner, budget, and data boundary: {{add verified input}}
Method:
1. Define the job, owner, pass/fail criteria, test budget, and stop conditions.
2. Create a dated source log from current primary sources; separate vendor claims, verified facts, and missing evidence.
3. Verify access, pricing, limits, integrations, security, privacy, and data-use terms for the intended account and region.
4. Run representative non-sensitive cases and compare quality, errors, time, and cost with the current method or alternative.
5. Have a human reviewer check consequential outputs, policy fit, rights, data boundaries, and operational ownership.
6. Record one decision: test with safeguards, pause for named evidence, or reject with reasons. Do not invent a fit score.
Return exactly these reviewable outputs:
- Dated source and evidence log
- Trial record against acceptance criteria
- Test / pause / reject decision memo
- Owner, safeguards, monitoring, and stop conditions
Human review boundary: This workflow does not endorse a tool. Do not connect sensitive data, purchase, publish, or recommend until an authorized human approves the evidence and bounded next step.