Verified agentic reasoning data for frontier AI labs.
Aurevus is building a proprietary, model-agnostic Cognitive OS that solves problems today's models can't — on ARC-AGI-3, a frontier benchmark — and turns the graded, step-by-step solving process into training data for agentic and reasoning models. The system is proven; the first datasets are being compiled now.
"record_type": "outcome",
"benchmark": "ARC-AGI-3",
"joined_by": "problem + step",
"grade": "pass",
"layers": ["decision", "outcome", "debate"]
}
Training data from real problem-solving, not scraped behavior.
Most agentic data is scraped text or unverified logs. Aurevus captures the complete problem-solving process — situation, reasoning, move, and graded result — as its Cognitive OS works scored ARC-AGI-3 problems.
Decision records
Board state, available actions, and a learned value estimate for every possible move at each step.
Outcome records
Ground-truth pass/fail grades on state transitions — the rare, RL-usable layer that makes the data trainable.
Debate records
The premium layer: multiple models reason adversarially — claim, counterexample, falsification test, verdict — before the chosen action.
It's not the model. It's the Cognitive OS on top of it.
Aurevus runs Orion — a proprietary, model-agnostic Cognitive OS above frontier LLMs. The underlying models are swappable components; Orion is the system that orchestrates, corrects, evaluates, and records the work, leaving a tamper-evident receipt at every step.
It's built for generality: every hard task exposes capability gaps, produces graded telemetry, and adds to the dataset. ARC-AGI-3 is the proving forge — not the product. The data is what we sell today; Orion, the architecture behind it, is the deeper asset.
Scored problem
ARC-AGI-3 provides hard, novel, multi-step tasks with objective grades.
Cognitive OS
A proprietary reasoning layer coordinates, corrects, and evaluates the problem-solving.
Verified telemetry
Every step is captured as structured trajectory data tied to outcomes.
Lab-ready feed
Sanitized JSONL with schema, dataset card, and linked record types.
The data an agentic model can actually learn from.
Offered non-exclusively as a data feed, evaluation pack, or strategic partnership — for teams training and evaluating agentic and reasoning models. Early access is opening now.
Verified Outcome Trajectories
Step-by-step solves, ground-truth graded by the environment.
Capability-Uplift Data
Base attempt → Cognitive OS correction → improved, graded result.
Failure-Mode Telemetry
Where models fail, how failures persist, and what changes the outcome.
Private Evaluation Packs
Scored, held-out probes for lab-side agentic capability evaluation.
ARC-AGI-3 is the forge, not the product.
Aurevus uses scored ARC-AGI-3 problem-solving to generate and verify the data. The product is the verified exhaust: decisions, outcomes, and falsification-driven reasoning, each tied to an objective grade.
For model labs, agent teams, and frontier data buyers.
Request the dataset card, review sanitized evidence, and discuss early, non-exclusive access to the trajectory pipeline as it comes online.
