Skip to content
caskcsgPublic

About

Official code for 《LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails》 (EMNLP 2026 Main).

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

🛡️ LongGuard

Mechanistic analysis and training-free mitigation of long-context failure in safety guardrails

Paper Dataset License

Official code for "LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails" (EMNLP 2026 Main).

A guardrail catches an unsafe request in a short context and misses the same request once it is embedded in a long one

📖 Overview

This repository evaluates six production guardrails on SafetyNIAH (30,400 samples, 256 → 32k words) and applies two training-free mitigations — CAHR-CD and CAHR-AHS — with frozen routing tables in configs/.

Capability Description
Evaluate original guard behaviour on long-context inputs
Mitigate CAHR-CD (chunked detection) and CAHR-AHS (attention-head sharpening)
Report metrics sliced by length, side, haystack type and position
Evaluation with SafetyNIAH, analysis of attention dilution, and mitigation by context-adaptive hyperparameter routing
Mean metric against relative processed-token budget: original at 1.00x, CAHR-AHS at 1.53x, CAHR-CD at 2.24x

📊 Key results

Unsafe recall R_u drops as context grows; safe recall stays flat (85.3% → 84.2%):

Guard 256 1k 4k 16k 32k
LlamaGuard3-8B 66.1 37.7 25.9 15.7 14.5
NemotronGuardV2-8B 64.4 47.3 17.1 21.7 25.1
NemotronGuardV3-8B 78.1 69.1 36.9 14.1 16.9
PolyGuard-7B 91.9 89.2 87.0 74.4 50.3
Qwen3Guard-Gen-8B 87.9 84.6 80.0 62.8 33.5
YuFeng-XGuard-8B 83.0 77.5 70.1 50.1 32.4
mean 78.6 67.6 52.8 39.8 28.8

CAHR recovers most of the loss on SafetyNIAH F1_unsafe:

Guard original + CAHR-CD + CAHR-AHS
LlamaGuard3-8B 48.96 77.50 63.41
NemotronGuardV2-8B 48.37 80.14 75.44
NemotronGuardV3-8B 53.82 83.03 76.12
PolyGuard-7B 80.84 83.02 82.22
Qwen3Guard-Gen-8B 80.83 85.37 83.89
YuFeng-XGuard-8B 76.89 84.68 79.92
mean 64.95 82.29 76.83

📁 Repository layout

.
├── run_eval.py                  # evaluate one guard under one setting
├── run_report.py                # score predictions, sliced by length
├── run_routing_table.py         # inspect frozen CAHR tables
├── configs/
│   ├── cahr_routing_table.csv   # 192 cells: 6 guards × 2 families × 2 sides × 8 lengths
│   └── retrieval_heads.csv      # attention-head selectivity rankings
├── longguard/                   # guards, methods, routing, metrics
├── scripts/                     # download_data.sh, download_models.sh, run_all.sh
├── data/                        # SafetyNIAH.parquet
├── models/                      # guard weights
├── outputs/                     # prediction JSONL files
└── reports/                     # by_length.csv and report.md

🚀 Getting started

Run all commands from the repository root (LongGuard/).

📦 1. Environment setup

git clone https://github.com/czyPL/LongGuard.git
cd LongGuard
pip install -r requirements.txt

Requires Python 3.10+ and a CUDA GPU for full evaluation. Log in to Hugging Face once before downloading gated models:

hf auth login

📥 2. Download data

The benchmark is fetched from caskcsg/SafetyNIAH into data/SafetyNIAH.parquet — the same path run_eval.py and run_report.py load by default (via longguard.data.default_benchmark_path()).

bash scripts/download_data.sh
# → data/SafetyNIAH.parquet  (489 MB, 30,400 samples)

Verify the file is in place:

ls -lh data/SafetyNIAH.parquet
python -c "from longguard.data import default_benchmark_path; print(default_benchmark_path())"

To use a copy elsewhere, set $SAFETYNIAH_PATH instead of moving the file:

export SAFETYNIAH_PATH=/path/to/SafetyNIAH.parquet

🤖 3. Download models

Guard weights land under models/<GuardName>/. The directory name must match the guard name used in the code — it is the join key for routing tables and output filenames.

Guard Hub id
LlamaGuard3-8B meta-llama/Llama-Guard-3-8B
NemotronGuardV2-8B nvidia/llama-3.1-nemoguard-8b-content-safety
NemotronGuardV3-8B nvidia/Llama-3.1-Nemotron-Safety-Guard-8B-v3
PolyGuard-7B ToxicityPrompts/PolyGuard-Qwen-Smol
Qwen3Guard-Gen-8B Qwen/Qwen3Guard-Gen-8B
YuFeng-XGuard-8B AAIG-Security/YuFeng-XGuard-8B

LlamaGuard3-8B and Llama-3.1-8B-Instruct (LoRA base for NemotronGuardV2) are gated — accept their Hub licences before downloading.

Download one guard to try the pipeline:

bash scripts/download_models.sh LlamaGuard3-8B _base_Meta-Llama-3.1-8B-Instruct
export LONGGUARD_MODEL_ROOT=$PWD/models

Download all six guards:

bash scripts/download_models.sh
export LONGGUARD_MODEL_ROOT=$PWD/models

🔬 4. Evaluation

Predictions are written to outputs/<method>/<GuardName>.jsonl. With data at data/SafetyNIAH.parquet and models under $LONGGUARD_MODEL_ROOT, no extra paths are needed:

# baseline
python run_eval.py --guard LlamaGuard3-8B --method original

# training-free mitigations (reuse base predictions when present)
python run_eval.py --guard LlamaGuard3-8B --method cahr-cd
python run_eval.py --guard LlamaGuard3-8B --method cahr-ahs

Smoke test on 200 samples first:

python run_eval.py --guard LlamaGuard3-8B --method original --limit 200

Full sweep, all six guards, sharded across visible GPUs:

bash scripts/run_all.sh
Option Meaning
--benchmark PATH override benchmark path; default data/SafetyNIAH.parquet
--num-shards N --shard-id I shard across GPUs / processes
--resume skip completed samples; retry errors
--limit N first N samples only
--output-dir DIR prediction root; default outputs/

Explicit paths (only needed if files are not in the default locations):

python run_eval.py --guard LlamaGuard3-8B --method original \
  --benchmark data/SafetyNIAH.parquet \
  --output-dir outputs

📊 5. Reporting

Scores join predictions in outputs/ against labels in data/SafetyNIAH.parquet:

python run_report.py
# → reports/by_length.csv
# → reports/report.md

Slice by other axes or pick metrics:

python run_report.py --metrics R_u F1_unsafe --axes length type haystack position

Metrics: F1_unsafe, R_u, R_s, P_u, accuracy, inv_rate.

📏 Scoring conventions

  • Unsafe is the positive class — R_u and F1_unsafe lead; R_s is always reported alongside to catch methods that flag everything.
  • Unparsable answers count as errors, not missing data — dropping them would reward the format-collapse failure this benchmark measures.

📄 Citation

@misc{chen2026longguardmechanisticanalysistrainingfree,
  title={LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails},
  author={Ziyang Chen and Xing Wu and Songlin Hu},
  year={2026},
  eprint={2608.27580},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2608.27580},
}

⚖️ License

Code: Apache 2.0. SafetyNIAH: CC BY-NC 4.0 (needle text under source-benchmark licenses). Guard weights remain under their own Hub licences.

🙏 Acknowledgements

Supported by the National Natural Science Foundation of China (No. U24A20335). We thank the maintainers of the public safety benchmarks and guardrails used here.

About

Official code for 《LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails》 (EMNLP 2026 Main).

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages