Dataset release

White Hat Security Agent Prompts 600K

Defensive security reasoning / 596,295 prompts / CC BY 4.0

A practitioner-perspective corpus of 596K contextual security prompts that teaches models to reason from inside live defensive operations: incident response, red-team simulation, threat intelligence, post-mortems, CISO review, and AI safety.

White-hat security agent monitoring threat intelligence dashboards
596,295security prompts
131named threat categories
76.8M+threat-scenario search space
5impact severity tiers
211average prompt words
100%schema density target

Technical overview

What this release is built to train.

Prompts are framed from active operational roles rather than textbook Q&A, including SOC analysts, CISOs, threat hunters, red-teamers, and trust-and-safety operators.

The taxonomy spans network, malware, web, social engineering, cloud, supply chain, IoT/OT, DeFi, insider threat, IAM, critical infrastructure, telecom, AI safety, quantum, synthetic biology, autonomous systems, and APT operations.

Each scenario combines threat, attack vector, practitioner role, defensive system, target sector, and impact level to train threat-aware instruction following.

Impact levels range from low nuisance events to catastrophic scenarios involving existential, national-security, or loss-of-life stakes.

Security reasoning from the defender's operating position

The 596,295 prompts are written as live professional requests rather than sanitized textbook questions. They place the model inside incident response, authorized red-team simulation, skeptical CISO review, forensic post-mortems, and threat-intelligence briefings where urgency, context, prioritization, and operational language matter.

Each record asks the model to reason like a practitioner responsible for stopping a real event. The prompt combines a named threat, attack vector, practitioner role, defensive system, target sector, and impact tier, enabling both instruction-following and supervised threat classification from the accompanying metadata.

131 threats across conventional and frontier security

The taxonomy covers network, malware, web, social engineering, cloud, supply chain, IoT and operational technology, DeFi, insider threat, privacy, identity and access management, mobile, physical security, critical infrastructure, and telecom. It extends into adversarial machine learning, malicious-intent detection, model alignment, quantum cryptography, synthetic biology, autonomous systems, and nation-state APT operations.

The combinatorial base exceeds 76.8 million potential scenarios. Five approximately uniform impact tiers range from low nuisance events through business disruption, enterprise damage, national-security risk, and catastrophic loss-of-life or existential scenarios. Average prompt length is about 211 words.

Schema and responsible training applications

The compact Parquet schema contains a reproducible batch index, the complete practitioner-framed prompt, one of 131 threat labels, and the impact classification. That supports security-specialized language-model fine-tuning, SOC assistants, threat-aware response calibration, multi-domain classifiers, red-team scenario research, and focused AI-safety subsets.

This is defensive training material, not an authorization layer or a substitute for qualified incident response. The corpus is released under CC BY 4.0 and requires attribution when reused or redistributed.