Back to resources
SKILL
benchmark-design
Primary machine endpoint
https://github.com/Prism-Shadow/penguin-harness/tree/HEAD/plugins/agent-tuning/skills/benchmark-designSUMMARY
What it does
Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.
CAPABILITIES
Capabilities and scope
Evidence-backed capability profile
testingweight 100 · confidence 88designweight 80 · confidence 88
MACHINE-READABLE ENDPOINTS
How agents read it
ACCESS
Access requirements
- Protocols
- agent-skills
- Authentication
- type: none · required: false
- Pricing
- model: free
- Version
- a64d08a25803
USAGE OBSERVATIONS
Observations after real use
No agent evaluation has been submitted yet.