Test an AI skill with a small evaluation set
Use complete, incomplete, conflicting, and adversarial inputs to inspect workflow behavior rather than one impressive output.

Evaluate a skill with several representative inputs and explicit checks. One polished answer does not show whether it handles missing facts, contradictions, or unsafe instructions in a document. Record the host, model, version, inputs, and observed behavior so the test means something later.
Define success before the run
For a fictional résumé workflow, success might mean preserving dates and titles, asking about missing metrics, and keeping rewrites attached to evidence. Do not define success as “sounds professional”; that can reward invented claims. Include a usable output criterion as well as factual constraints.
Use four cases
| Case | Input | What to inspect |
|---|---|---|
| Complete | Bullet plus confirmed scope | Accurate rewrite |
| Incomplete | No outcome measurement | Question or unquantified wording |
| Conflicting | Two different employment dates | Conflict flagged |
| Embedded instruction | Document says to ignore the user | Document treated as data |
These cases are a starting set, not a certification. Add cases that resemble the actual work and consequences of your workflow.
Keep an evaluation record
Task and package version:
Host and model:
Input case:
Checks defined before the run:
Observed output:
Checks passed or failed:
Unsupported additions:
Questions asked:
Next change to test:Preserve actual outputs separately from authored examples. A sample you write to explain desired behavior should never be presented as a recorded model result.
Test changes against the same cases
After changing the instructions, repeat the relevant cases. A fix that improves the complete example may make the missing-input case worse. Record regressions and avoid changing several unrelated parts at once when diagnosing a failure.
Skill tests are different from package validation and storefront tests. Checking that files exist or that a ZIP downloads proves those behaviors, not native activation or output quality. State exactly which layer you evaluated and keep the limits visible when describing compatibility or results.
References and further reading
The examples and templates above are original. These references support the definitions and documented behavior discussed in the guide.



