Dataset release

Adversarial Agent Intent Safety Analysis 240K

Intent safety and clarification / 242,454 safety analyses / Open RAIL-D

A 242,454-row adversarial safety corpus for command-and-control models, guardrail classifiers, and red-team agents. Each record separates a request's plausible surface interpretation from its deeper capability footprint, then produces an intent audit and authorization-focused clarifying questions.

Adversarial AI safety agent reviewing intent-risk signals
242,454adversarial safety records
149.7M+exact cl100k_base tokens
480.6Mgeneration and context tokens processed
496atomic adversarial objectives
126critical risk vectors
808.5Mpossible threat permutations
509average words per record
0pairs above 80% cosine similarity in the audit

Technical overview

What this release is built to train.

Every row follows a tripartite safety structure: a naive surface interpretation, a forensic intent analysis that maps the requested specifications to their capability footprint, and targeted clarifying questions that test authorization, compliance, and ethical boundaries.

The generation matrix combines 496 atomic adversarial objectives, 247 disguise personas, 275 legitimate cover contexts, target-output formats, four sophistication tiers, and High or Critical risk levels across 12 broad threat domains.

The dataset contains about 123.6 million words and 149.7 million cl100k_base tokens. A 50,000-prompt diversity audit covered 1.25 billion pairwise comparisons and found 6.34% average TF-IDF cosine similarity, 5.80% median similarity, and no pair above 80%.

Access is gated for safety research and responsible model development. The Open RAIL-D license permits research and commercial training subject to restrictions against weaponization targeting people, AI systems, or critical infrastructure.

A three-part intent audit instead of keyword refusal

The dataset trains models to distinguish what a prompt claims to request from what the requested technical artifact could enable. Surface Interpretation records the immediate benign reading. Intent Analysis cross-references specifications, cover story, and dual-use consequences to uncover the deeper capability. Clarifying Questions introduce targeted friction around authorization, legal basis, institutional controls, scope, and intended deployment.

This makes the release useful for agents that act before execution: system-level intent routers, command-and-control safeguards, guardrail classifiers, and red-team evaluators. The desired behavior is not indiscriminate refusal; it is technically informed skepticism that can identify incongruencies and ask the right questions before a high-impact tool call proceeds.

Deterministic adversarial breadth and measured diversity

The matrix combines 496 literal payload objectives with 247 authoritative personas, 275 legitimate cover scenarios, target artifact formats, four obfuscation sophistication levels, and two risk ceilings. Those dimensions form more than 808.5 million valid threat permutations; 242,454 complete analyses were retained for the release.

Average density is approximately 509 words per row: 166 words of adversarial prompt, 37 of surface interpretation, 129 of intent analysis, and 178 of defense-in-depth questions. On a random 50,000-prompt subset, average TF-IDF cosine similarity was 0.0634 and average Jaccard overlap was 0.0791, with zero highly similar pairs across 1,249,975,000 comparisons.

Training applications and responsible access

The eight-column Parquet schema records batch index, operating mode, sophistication, risk level, adversarial prompt, surface interpretation, intent analysis, and clarifying questions. That supports supervised guardrail training, risk classification, frontier-model red teaming, System 2 intent routing, defense-in-depth clarification policies, and research into dual-use abstraction.

The release is access-gated and uses the Open RAIL-D license. It allows responsible research and commercial model training while imposing restrictions against using the data to weaponize systems against individuals, AI systems, or critical infrastructure. Users should review the complete repository license and access terms before training or redistribution.