# LLM Prompt And Workflow Evaluation Template

Use this template to evaluate whether an LLM workflow is ready for a pilot.

## 1. Workflow Definition

- Task:
- User:
- Input data:
- Output:
- Human reviewer:
- System handoff:

## 2. Test Set

- Number of examples:
- Easy examples:
- Edge cases:
- Sensitive cases:
- Known bad examples:

## 3. Evaluation Metrics

- Task success rate:
- Evidence quality:
- Hallucination or unsupported claim rate:
- Reviewer acceptance rate:
- Average response time:
- API handoff success:
- Manual override rate:

## 4. Prompt Notes

- System instruction:
- Retrieval sources:
- Required output format:
- Fallback instruction:
- Refusal or escalation criteria:

## 5. Human Review

- What the reviewer sees:
- What can be edited:
- What requires approval:
- What is logged:
- What blocks automation:

## 6. Pilot Decision

- Ready to pilot:
- Needs another sprint:
- Too risky:
- Better solved without AI:
