Each eval output is graded on 12 quality dimensions (pass/fail per dimension).
Original 9 Dimensions (iteration-4, 13 scenarios)
#
Dimension
Assertion
What it checks
1
Product Analysis
包含产品分析/拆解章节
Output contains structured product deconstruction
2
Multiple Prompts
包含多个可用的重建 prompt(≥2)
At least 2 usable rebuild prompts
3
Done Criteria
包含完成/验收标准
Each prompt has acceptance/done conditions
4
Why Annotations
包含教学注释(为什么有效)
Explains why each prompt works (teaching value)
5
Pro Tips
包含 Pro Tips / 实战洞察
Practical insights from experience
6
Build Order
包含构建顺序建议
Explicit build/execution sequence
7
Domain Framework
使用领域专属分析框架
Uses domain-specific analysis (MDA for games, etc.)
8
Not-To-Do
包含 Not-To-Do / 范围限制
Scope boundaries for each prompt
9
Soul Capture
捕获了产品的灵魂/核心差异化
Identifies the product's core identity/soul
New 3 Dimensions (iteration-5, 3 additional scenarios)
#
Dimension
Assertion
What it checks
10
Destination Not Route
Prompt 描述目的地而非路线
Prompts describe desired outcome, not implementation steps
11
Prompts Standalone
每个 prompt 可独立使用
Each prompt is self-contained with enough context
12
Usage Guidance
包含 prompt 使用建议
Execution order, alternatives, and next steps
Grader : Claude (sonnet)
Method : Each scenario is run twice — once with the Cleaver skill active, once without
Input : Same product description/screenshot/URL for both runs
Scoring : Binary pass/fail per dimension, based on whether the assertion is met with evidence
Aggregation : Pass rate = passed / total assertions
benchmark.json — Aggregated results (generated by build_benchmark.py)
build_benchmark.py — Reads grading JSONs from workspace, outputs benchmark.json
Source grading files: ~/.claude/skills/cleaver-workspace/iteration-{4,5}/eval-*/{with,without}_skill/grading.json