Workflow

Stuck on: Decide whether to run a bounded test, pause for evidence, or reject an AI tool for a defined job.

AI Tool Evaluation Workflow

Teams evaluating a tool before paying, connecting data, recommending it, or adding it to a catalog.

Follow the sequence

Steps

  1. 1

    Define the job, owner, pass/fail criteria, test budget, and stop conditions.

  2. 2

    Create a dated source log from current primary sources; separate vendor claims, verified facts, and missing evidence.

  3. 3

    Verify access, pricing, limits, integrations, security, privacy, and data-use terms for the intended account and region.

  4. 4

    Run representative non-sensitive cases and compare quality, errors, time, and cost with the current method or alternative.

  5. 5

    Have a human reviewer check consequential outputs, policy fit, rights, data boundaries, and operational ownership.

  6. 6

    Record one decision: test with safeguards, pause for named evidence, or reject with reasons. Do not invent a fit score.

Output contract

What you should get

  • Dated source and evidence log
  • Trial record against acceptance criteria
  • Test / pause / reject decision memo
  • Owner, safeguards, monitoring, and stop conditions
Copy and adapt

Prompt

Run the AI Tool Evaluation Workflow for the goal below.

Goal: Decide whether to run a bounded test, pause for evidence, or reject an AI tool for a defined job.

Inputs (replace each placeholder with verified information):
- Job, target user, and pass/fail criteria: {{add verified input}}
- Official product, pricing, security, and privacy sources: {{add verified input}}
- Trial account, region, and access limits: {{add verified input}}
- Representative non-sensitive test cases: {{add verified input}}
- Current method or alternative: {{add verified input}}
- Test owner, budget, and data boundary: {{add verified input}}

Method:
1. Define the job, owner, pass/fail criteria, test budget, and stop conditions.
2. Create a dated source log from current primary sources; separate vendor claims, verified facts, and missing evidence.
3. Verify access, pricing, limits, integrations, security, privacy, and data-use terms for the intended account and region.
4. Run representative non-sensitive cases and compare quality, errors, time, and cost with the current method or alternative.
5. Have a human reviewer check consequential outputs, policy fit, rights, data boundaries, and operational ownership.
6. Record one decision: test with safeguards, pause for named evidence, or reject with reasons. Do not invent a fit score.

Return exactly these reviewable outputs:
- Dated source and evidence log
- Trial record against acceptance criteria
- Test / pause / reject decision memo
- Owner, safeguards, monitoring, and stop conditions

Human review boundary: This workflow does not endorse a tool. Do not connect sensitive data, purchase, publish, or recommend until an authorized human approves the evidence and bounded next step.
Human review boundaryThis workflow does not endorse a tool. Do not connect sensitive data, purchase, publish, or recommend until an authorized human approves the evidence and bounded next step.