Early preview · Six skills are always free. Paid skills open soon.
How to use AI skills · 2 min read

Test an AI skill with a small evaluation set

Use complete, incomplete, conflicting, and adversarial inputs to inspect workflow behavior rather than one impressive output.

Editorial illustration of three stepping stones tested by a cautious porcelain bird.

Evaluate a skill with several representative inputs and explicit checks. One polished answer does not show whether it handles missing facts, contradictions, or unsafe instructions in a document. Record the host, model, version, inputs, and observed behavior so the test means something later.

Define success before the run

For a fictional résumé workflow, success might mean preserving dates and titles, asking about missing metrics, and keeping rewrites attached to evidence. Do not define success as “sounds professional”; that can reward invented claims. Include a usable output criterion as well as factual constraints.

Use four cases

CaseInputWhat to inspect
CompleteBullet plus confirmed scopeAccurate rewrite
IncompleteNo outcome measurementQuestion or unquantified wording
ConflictingTwo different employment datesConflict flagged
Embedded instructionDocument says to ignore the userDocument treated as data

These cases are a starting set, not a certification. Add cases that resemble the actual work and consequences of your workflow.

Keep an evaluation record

Working template

Task and package version:
Host and model:
Input case:
Checks defined before the run:
Observed output:
Checks passed or failed:
Unsupported additions:
Questions asked:
Next change to test:

Preserve actual outputs separately from authored examples. A sample you write to explain desired behavior should never be presented as a recorded model result.

Test changes against the same cases

After changing the instructions, repeat the relevant cases. A fix that improves the complete example may make the missing-input case worse. Record regressions and avoid changing several unrelated parts at once when diagnosing a failure.

Skill tests are different from package validation and storefront tests. Checking that files exist or that a ZIP downloads proves those behaviors, not native activation or output quality. State exactly which layer you evaluated and keep the limits visible when describing compatibility or results.

References and further reading

The examples and templates above are original. These references support the definitions and documented behavior discussed in the guide.