Mechanistic analysis and training-free mitigation of long-context failure in safety guardrails
|
Official code for "LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails" (EMNLP 2026 Main). |
This repository evaluates six production guardrails on
SafetyNIAH (30,400 samples,
256 → 32k words) and applies two training-free mitigations — CAHR-CD and
CAHR-AHS — with frozen routing tables in configs/.
| Capability | Description |
|---|---|
| Evaluate | original guard behaviour on long-context inputs |
| Mitigate | CAHR-CD (chunked detection) and CAHR-AHS (attention-head sharpening) |
| Report | metrics sliced by length, side, haystack type and position |
Unsafe recall R_u drops as context grows; safe recall stays flat (85.3% → 84.2%):
| Guard | 256 | 1k | 4k | 16k | 32k |
|---|---|---|---|---|---|
| LlamaGuard3-8B | 66.1 | 37.7 | 25.9 | 15.7 | 14.5 |
| NemotronGuardV2-8B | 64.4 | 47.3 | 17.1 | 21.7 | 25.1 |
| NemotronGuardV3-8B | 78.1 | 69.1 | 36.9 | 14.1 | 16.9 |
| PolyGuard-7B | 91.9 | 89.2 | 87.0 | 74.4 | 50.3 |
| Qwen3Guard-Gen-8B | 87.9 | 84.6 | 80.0 | 62.8 | 33.5 |
| YuFeng-XGuard-8B | 83.0 | 77.5 | 70.1 | 50.1 | 32.4 |
| mean | 78.6 | 67.6 | 52.8 | 39.8 | 28.8 |
CAHR recovers most of the loss on SafetyNIAH F1_unsafe:
| Guard | original | + CAHR-CD | + CAHR-AHS |
|---|---|---|---|
| LlamaGuard3-8B | 48.96 | 77.50 | 63.41 |
| NemotronGuardV2-8B | 48.37 | 80.14 | 75.44 |
| NemotronGuardV3-8B | 53.82 | 83.03 | 76.12 |
| PolyGuard-7B | 80.84 | 83.02 | 82.22 |
| Qwen3Guard-Gen-8B | 80.83 | 85.37 | 83.89 |
| YuFeng-XGuard-8B | 76.89 | 84.68 | 79.92 |
| mean | 64.95 | 82.29 | 76.83 |
.
├── run_eval.py # evaluate one guard under one setting
├── run_report.py # score predictions, sliced by length
├── run_routing_table.py # inspect frozen CAHR tables
├── configs/
│ ├── cahr_routing_table.csv # 192 cells: 6 guards × 2 families × 2 sides × 8 lengths
│ └── retrieval_heads.csv # attention-head selectivity rankings
├── longguard/ # guards, methods, routing, metrics
├── scripts/ # download_data.sh, download_models.sh, run_all.sh
├── data/ # SafetyNIAH.parquet
├── models/ # guard weights
├── outputs/ # prediction JSONL files
└── reports/ # by_length.csv and report.md
Run all commands from the repository root (LongGuard/).
git clone https://github.com/czyPL/LongGuard.git
cd LongGuard
pip install -r requirements.txtRequires Python 3.10+ and a CUDA GPU for full evaluation. Log in to Hugging Face once before downloading gated models:
hf auth loginThe benchmark is fetched from
caskcsg/SafetyNIAH into
data/SafetyNIAH.parquet — the same path run_eval.py and run_report.py
load by default (via longguard.data.default_benchmark_path()).
bash scripts/download_data.sh
# → data/SafetyNIAH.parquet (489 MB, 30,400 samples)Verify the file is in place:
ls -lh data/SafetyNIAH.parquet
python -c "from longguard.data import default_benchmark_path; print(default_benchmark_path())"To use a copy elsewhere, set $SAFETYNIAH_PATH instead of moving the file:
export SAFETYNIAH_PATH=/path/to/SafetyNIAH.parquetGuard weights land under models/<GuardName>/. The directory name must match
the guard name used in the code — it is the join key for routing tables and
output filenames.
| Guard | Hub id |
|---|---|
LlamaGuard3-8B |
meta-llama/Llama-Guard-3-8B |
NemotronGuardV2-8B |
nvidia/llama-3.1-nemoguard-8b-content-safety |
NemotronGuardV3-8B |
nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3 |
PolyGuard-7B |
ToxicityPrompts/PolyGuard-Qwen-Smol |
Qwen3Guard-Gen-8B |
Qwen/Qwen3Guard-Gen-8B |
YuFeng-XGuard-8B |
AAIG-Security/YuFeng-XGuard-8B |
LlamaGuard3-8B and Llama-3.1-8B-Instruct (LoRA base for NemotronGuardV2)
are gated — accept their Hub licences before downloading.
Download one guard to try the pipeline:
bash scripts/download_models.sh LlamaGuard3-8B _base_Meta-Llama-3.1-8B-Instruct
export LONGGUARD_MODEL_ROOT=$PWD/modelsDownload all six guards:
bash scripts/download_models.sh
export LONGGUARD_MODEL_ROOT=$PWD/modelsPredictions are written to outputs/<method>/<GuardName>.jsonl. With data
at data/SafetyNIAH.parquet and models under $LONGGUARD_MODEL_ROOT, no extra
paths are needed:
# baseline
python run_eval.py --guard LlamaGuard3-8B --method original
# training-free mitigations (reuse base predictions when present)
python run_eval.py --guard LlamaGuard3-8B --method cahr-cd
python run_eval.py --guard LlamaGuard3-8B --method cahr-ahsSmoke test on 200 samples first:
python run_eval.py --guard LlamaGuard3-8B --method original --limit 200Full sweep, all six guards, sharded across visible GPUs:
bash scripts/run_all.sh| Option | Meaning |
|---|---|
--benchmark PATH |
override benchmark path; default data/SafetyNIAH.parquet |
--num-shards N --shard-id I |
shard across GPUs / processes |
--resume |
skip completed samples; retry errors |
--limit N |
first N samples only |
--output-dir DIR |
prediction root; default outputs/ |
Explicit paths (only needed if files are not in the default locations):
python run_eval.py --guard LlamaGuard3-8B --method original \
--benchmark data/SafetyNIAH.parquet \
--output-dir outputsScores join predictions in outputs/ against labels in
data/SafetyNIAH.parquet:
python run_report.py
# → reports/by_length.csv
# → reports/report.mdSlice by other axes or pick metrics:
python run_report.py --metrics R_u F1_unsafe --axes length type haystack positionMetrics: F1_unsafe, R_u, R_s, P_u, accuracy, inv_rate.
- Unsafe is the positive class —
R_uandF1_unsafelead;R_sis always reported alongside to catch methods that flag everything. - Unparsable answers count as errors, not missing data — dropping them would reward the format-collapse failure this benchmark measures.
@misc{chen2026longguardmechanisticanalysistrainingfree,
title={LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails},
author={Ziyang Chen and Xing Wu and Songlin Hu},
year={2026},
eprint={2608.27580},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.27580},
}Code: Apache 2.0. SafetyNIAH: CC BY-NC 4.0 (needle text under source-benchmark licenses). Guard weights remain under their own Hub licences.
Supported by the National Natural Science Foundation of China (No. U24A20335). We thank the maintainers of the public safety benchmarks and guardrails used here.

