Technical overview
What this release is built to train.
Each task is written as a first-person, technical execution command directed at an AI co-pilot embedded inside a specialized creative application.
The dataset covers frontend/UI infrastructure, node-based CGI, nonlinear editing, color grading, DAWs, vector design, raster design, packaging, motion graphics, and more.
The generation engine uses 67 to 210 native tools per archetype, producing 4,211,138 possible unique tool intersections before sampling roughly 1.07M operations.
A 76-parameter behavioral injection system creates lexical and emotional variance while preserving deterministic generation and avoiding sequential monotony.
Training contextual actuation across professional software
The 1,070,917 records are not generic questions. Each is a first-person technical command to an AI copilot embedded inside a professional application. The model must bridge subjective intent and literal actuation by inferring a sequence of GUI operations, node connections, timeline changes, parameter edits, routing decisions, or code-level modifications from dense domain language.
The 36 professional archetypes span 17 macro-categories across frontend and UI infrastructure, 3D visualization and node-based CGI, nonlinear editing and color science, digital audio workstations, raster and vector design, branding, packaging, photography, motion graphics, and related engineering environments.
Deterministic diversity at million-row scale
Every archetype has a curated taxonomy of 67 to 210 native tools. The generator selects deterministic two- or three-tool intersections from 4,211,138 possible combinations, then samples roughly one quarter of that space. Batch IDs provide reproducible seeding, while a 76-state behavioral array changes tone and lexical stance without introducing sequential correlation.
Prompts cluster around 175 words with a standard deviation of 15 and are bounded between 80 and 400 words. The full schema totals approximately 243.2 million tokens, including 231.6 million prompt tokens; generation consumed approximately 527.2 million tokens. Data is distributed as batch-ordered, Zstandard-compressed Parquet for high-throughput loading.
Coverage and intended model behavior
The dataset is designed for interface-aware autonomy: visual action agents, application copilots, API orchestration policies, multimodal problem solvers, and agents trained inside simulated professional environments. It can also seed environment generation, where a separate system constructs an isolated NLE, compositor, DAW, design canvas, or frontend workspace and the agent learns from execution traces.
The five-column schema records batch and local indices, professional identity, macro-group, and raw task payload. The release is MIT licensed for commercial, academic, and personal research into creative software orchestration and high-fidelity multimodal interaction.
