arXiv Paper • Models • Data & Results
About • Documents • Citation • Acknowledgement
Accepted at NeurIPS 2026 and AdvML-Frontiers × CoTMA: From Model Security to Compositional Threats in Multi-Agent AI Systems @ COLM 2026 (Oral)
CoT-Guard is a 4B-parameter monitor trained to detect hidden objectives in code-generation tasks using chain-of-thought and generated code. This repository contains the training, evaluation, and data curation code, with SFT and RL model checkpoints available on Hugging Face.
Read CoT-Guard: Small Models for Strong Monitoring for the method and evaluation details.
From the project root, set the data and temporary directories and add the project to PYTHONPATH:
export COT_GLOBAL_DIRECTORY="/path/to/data_and_results"
export COT_TEMP_DIRECTORY="/path/to/temp_for_vllm_logs"
export OPENAI_API_KEY="your-openai-api-key"
export GOOGLE_API_KEY="your-gemini-api-key"
export PYTHONPATH=$PYTHONPATH:$(pwd)- Attack: create tasks and generate attack responses.
- Verify: test responses for main-task correctness and side-task success.
- Monitor: create monitor tasks from the responses and run monitors.
- Score: evaluate monitor predictions.
a) Create prompts containing the main task (bigcodebench), side task (varname), and attack policy (baseline):
python3 task_creation/for_attack.py --main_task=bigcodebench --side_tasks=varname --attack_policies=baselineb) Run the tasks with Qwen/Qwen3-8B on GPU 0 to generate one attack response per prompt:
python3 gen/attack.py single --main_task=bigcodebench --side_task=varname --attack_policy=baseline --model=Qwen/Qwen3-14B --devices=0 --tp=1 --num_rollouts=5With Docker running:
python3 eval/verify_task.py --main_task=bigcodebench --side_task=varname --attack_policy=baseline --attack_model=Qwen/Qwen3-14B --num_rollouts=5Output: $COT_GLOBAL_DIRECTORY/results/attack/bigcodebench.varname.baseline/Qwen3-14B.5.tested.jsonl
a) Create monitor prompts from the attacks that were successful:
python3 task_creation/for_monitor.py --main_task=bigcodebench --side_tasks=varname --attack_policies=baseline --model=Qwen/Qwen3-14B --attack_rollouts=5 --filter_mode=side-all --monitor_policy=main_aware --monitor_type=cot_actionb) Run the monitor with Qwen/Qwen3-8B:
python3 gen/monitor.py single --main_task=bigcodebench --side_task=varname --attack_policy=baseline --attack_model=Qwen/Qwen3-14B --attack_rollouts=5 --filter_mode=side-all --monitor_policy=main_aware --monitor_model=Qwen/Qwen3-8B --monitor_type=cot_action --devices=0 --tp=1 --num_rollouts=1Display one setting:
python3 eval/display.py --main_task=bigcodebench --evaluated_model=Qwen/Qwen3-14B --rollout_num=5 --attack_policy=baseline --monitor_policy=main_aware --monitor_objective=cot_action --filter_mode=side-all --side_tasks=varname --monitor_models=Qwen/Qwen3-8BDisplay the result tables:
python3 eval/display_table.py --main_task=bigcodebench --evaluated_model=Qwen/Qwen3-14B --rollout_num=5 --attack_policy=baseline --monitor_policy=main_aware --monitor_objective=cot_action --filter_mode=side-all --side_tasks=varname --monitor_models=Qwen/Qwen3-8BAttack policy values include --attack_policy=baseline, no_side_task, stealth, and monitor_notice.
--tp=1 shards the model across one GPU. --num_rollouts=1 generates one response per prompt.
Side task values include --side_task=varname, defaultval, logio, iodelete, iodownload, permissions, ioexfil, edgecase, complexity, longlines, exit, vulnimport, typosquat, benign_comment_watermark, benign_print_watermark.
You can also use comma-separated values, like --side_tasks=varname,exit.
For parallel generation across 2 GPUs --
python3 gen/attack.py parallel --main_task=bigcodebench --side_tasks=varname,exit --attack_policies=baseline --model=Qwen/Qwen3-14B --devices=0,1 --tp=1 --num_rollouts=5tp per N devices --> tensor parallel for vllm
- Paper: method, threat model, and evaluation.
- Inference pipeline: generate attacks, verify tasks, run monitors, and score results.
- Training: prepare data and train with SFT and RL / GRPO.
- Stored evaluation and results: download datasets and reproduce result tables.
- Model checkpoints: SFT and RL monitors.
The paper uses I/O Delete, I/O Download,
and I/O Exfiltrate Secret for SFT, then Default Value, Complexity, and Variable
Name for RL. For each group, prepare verified baseline and
no_side_task monitor outputs from Qwen3-14B attacks and a Qwen3-32B monitor,
with all three scopes (--monitor_type=all).
Install LlamaFactory, then collect the I/O task examples:
python3 train/sft/collect.py --main-tasks kodcode \
--side-tasks iodelete,ioexfil,iodownload --attack-rollouts 5The collector registers three datasets in dataset_info.json. Edit
train/sft/config/qwen3_full_sft_three.yaml before training: replace the
example deepspeed, dataset_dir, and output_dir paths with absolute paths.
$COT_GLOBAL_DIRECTORY in YAML is literal text and will not expand. Check the
dataset and model names if you use different sources. Then run:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 llamafactory-cli train train/sft/config/qwen3_full_sft_three.yamlCollect training prompts from the three subtle tasks:
python3 train/rl/collect.py --main-tasks bigcodebench \
--side-tasks defaultval,complexity,varname --set-type trainInstall veRL in a separate training
environment and activate it. Before running, edit DATA_ROOT, TRAIN_FILE,
MODEL_PATH, CHECKPOINT_DIR, GPU_IDS, and EXPERIMENT at the top
of train/rl/run_grpo.sh. Use your collected JSONL paths and set MODEL_PATH
to the same checkpoint as output_dir in the edited SFT YAML. The launcher
disables validation. Then run:
bash train/rl/run_grpo.shAll our results and datasets available for download from here.
After downloading and extracting the files, place datasets under
$COT_GLOBAL_DIRECTORY/datasets/ and evaluation outputs under
$COT_GLOBAL_DIRECTORY/results/. The display script reads outputs from
results/monitor/; preserve the downloaded subdirectory structure.
From the repository root, with the project environment activated, run the
following commands. Replace the example paths below; COT_DOWNLOAD_DIRECTORY
should point to the extracted folder containing datasets/ and results/.
export COT_GLOBAL_DIRECTORY="/path/to/data_and_results"
export COT_DOWNLOAD_DIRECTORY="/path/to/extracted_download"
export PYTHONPATH="$PWD${PYTHONPATH:+:$PYTHONPATH}"
mkdir -p "$COT_GLOBAL_DIRECTORY/datasets" "$COT_GLOBAL_DIRECTORY/results"
cp -a "$COT_DOWNLOAD_DIRECTORY/datasets/." "$COT_GLOBAL_DIRECTORY/datasets/"
cp -a "$COT_DOWNLOAD_DIRECTORY/results/." "$COT_GLOBAL_DIRECTORY/results/"
# Display tables for the default evaluation settings.
python3 eval/display_table.py
# Include all configured side tasks for those settings.
python3 eval/display_table.py --side_tasks=NoneUse --monitor_models, --attack_policy, and --monitor_objective to select
other downloaded evaluation settings; the command does not automatically
discover every experiment. See the results commands for an example.
If you use CoT-Guard in your research, please cite our paper:
@inproceedings{diwan2026cotguard,
title = {CoT-Guard: Small Models for Strong Monitoring},
author = {Nirav Diwan and Han Wang and Berkcan Kapusuzoglu and Ramin Moradi and Supriyo Chakraborty and Giri Iyengar and Sambit Sahu and Huan Zhang and Gang Wang},
booktitle = {The Fortieth Annual Conference on Neural Information Processing Systems},
year = {2026},
url = {https://openreview.net/forum?id=mURtPVmTjF}
}We thank the developers and maintainers of BigCodeBench, Inspect AI, vLLM, LlamaFactory, and veRL for the open-source tools used in this project.