Thanks for your interest in improving Schliff! This guide covers everything you need to get started.
# Clone and install in dev mode (symlink)
git clone https://github.com/Zandereins/schliff.git
cd schliff
bash install.sh --link
# Run the test suite
make test-all- Python 3.9+
- Bash 4.0+
- Git 2.0+
- jq 1.6+
skills/schliff/
├── scripts/
│ ├── scoring/ # Scoring package (1 module per dimension)
│ │ ├── __init__.py # Public API — import from here
│ │ ├── patterns.py # Shared regex patterns
│ │ ├── structure.py # Structure scoring
│ │ ├── triggers.py # Trigger accuracy scoring
│ │ └── ... # One file per dimension
│ ├── shared.py # Core utilities (caching, file I/O, security)
│ ├── nlp.py # NLP utilities (tokenization, stemming)
│ ├── terminal_art.py # Terminal rendering (grades, bars, heatmaps)
│ └── ... # Application scripts
├── commands/schliff/ # Claude Code command definitions
├── references/ # Deep documentation
├── templates/ # Eval suite templates
└── tests/proof/ # Proof tests
make test-unit # >1000 unit tests (pytest)
make test # unit tests, then integration tests
make test-self # 20 self-tests (Schliff scores itself)
make test-proof # 6 proof tests (demonstrates real improvement)
make test-all # All of the above (test + test-self + test-proof)
make score # Score Schliff's own SKILL.md (expect >= 90)All tests must pass before submitting a PR.
-
Create
scripts/scoring/your_dimension.py:"""scoring/your_dimension.py — Your Dimension scoring.""" from shared import read_skill_safe from scoring.patterns import _RE_YOUR_PATTERNS # if needed def score_your_dimension(skill_path: str, eval_suite=None) -> dict: """Score your dimension. Returns dict with 'score' (0-100), 'issues' (list), 'details' (dict). """ content = read_skill_safe(skill_path) # ... scoring logic ... return {"score": score, "issues": issues, "details": details}
-
Register in
scripts/scoring/__init__.py:from scoring.your_dimension import score_your_dimension
-
Add weight in
scripts/scoring/composite.py(DEFAULT_WEIGHTS dict) -
Add tests in
scripts/test-integration.sh -
Run
make test-allto verify
- Python 3.9+ — Use type hints, f-strings, pathlib
- Line length: 120 chars (configured in pyproject.toml)
- Linter:
make lint(uses ruff) - Docstrings: Required for all public functions
- Error handling: Always explicit — no silent failures
- File I/O: Use
shared.read_skill_safe()for skill files (enforces size limits) - Regex: Use
shared.validate_regex_complexity()before executing user-supplied patterns
- Underscore names (
text_gradient.py) = importable Python modules - Hyphenated names (
text-gradient.py) = thin CLI wrappers (5-8 lines, delegate to underscore module) - scoring/ modules = one file per scoring dimension
- Use
shared.validate_command_safety()before executing any user-supplied commands (currently reserved, not yet wired — see shared.py docstring) - Use
shared.validate_regex_complexity()before any user-supplied regex - Use
shared.read_skill_safe()for all file reads (enforces 1MB size limit) - Never execute commands from eval-suite content without validation
- All tests pass (
make test-all) - Score is >= 90 (
make score) - New functions have docstrings and type hints
- Security functions used where applicable
- No hardcoded file paths
- CHANGELOG.md updated (if user-facing change)