Verified agentic reasoning data for frontier AI labs.
Aurevus is building a proprietary, model-agnostic Cognitive OS that solves problems today's models can't — on ARC-AGI-3, a frontier benchmark — and turns the graded, step-by-step solving process into training data for agentic and reasoning models. The system is proven; the first datasets are being compiled now.
"record_type": "outcome",
"benchmark": "ARC-AGI-3",
"joined_by": "problem + step",
"grade": "pass",
"layers": ["decision", "outcome", "debate"]
}
We ran a 60 GB model in ~11 GB — without quantizing a single weight. Then we edited what it knew.
MNEMOS is our patent-pending architecture that partitions a frozen model by function: run a model far larger than device memory with exact, never-quantized weights (bit-identical output at the fitting tier), edit its knowledge in seconds without retraining, and attach new capability across knowledge, skills, and reasoning — all on one frozen host, measured on a 24 GB laptop.
Training data from real problem-solving, not scraped behavior.
Most agentic data is scraped text or unverified logs. Aurevus captures the complete problem-solving process — situation, reasoning, move, and graded result — as its Cognitive OS works scored ARC-AGI-3 problems.
Decision records
Board state, available actions, and a learned value estimate for every possible move at each step.
Outcome records
Ground-truth pass/fail grades on state transitions — the rare, RL-usable layer that makes the data trainable.
Debate records
The premium layer: multiple models reason adversarially — claim, counterexample, falsification test, verdict — before the chosen action.
It's not the model. It's the Cognitive OS on top of it.
Aurevus runs Orion — a proprietary, model-agnostic Cognitive OS above frontier LLMs. The underlying models are swappable components; Orion is the system that orchestrates the multi-agent work — correcting, evaluating, and recording it — leaving a tamper-evident receipt at every step.
It's built for generality: every hard task exposes capability gaps, produces graded telemetry, and adds to the dataset. ARC-AGI-3 is the proving forge — not the product. The data is what we sell today; Orion, the architecture behind it, is the deeper asset.
Scored problem
ARC-AGI-3 provides hard, novel, multi-step tasks with objective grades.
Cognitive OS
A proprietary reasoning layer coordinates, corrects, and evaluates the problem-solving.
Verified telemetry
Every step is captured as structured trajectory data tied to outcomes.
Lab-ready feed
Sanitized JSONL with schema, dataset card, and linked record types.
The data an agentic model can actually learn from.
Offered non-exclusively as a data feed, evaluation pack, or strategic partnership — for teams training and evaluating agentic and reasoning models. Early access is opening now.
Verified Outcome Trajectories
Step-by-step solves, ground-truth graded by the environment.
Capability-Uplift Data
Base attempt → Cognitive OS correction → improved, graded result.
Failure-Mode Telemetry
Where models fail, how failures persist, and what changes the outcome.
Private Evaluation Packs
Scored, held-out probes for lab-side agentic capability evaluation.
Railspike: governed AI for the built world.
Railspike is Aurevus's governed AI operating layer for vertically integrated property developers, builders, and managers. It connects the systems a company already runs, carries a building's operational context across its lifecycle, and places consequential actions inside the customer's own governance.
Yardi/RentCafe · Microsoft 365/Entra · construction systems · accounting · HR · investor systems. Railspike connects the operating stack instead of asking the customer to rebuild it.
For model labs, agent teams, and frontier data buyers.
Contact Aurevus to discuss our work and potential collaboration.
