New closed-ended problems, not recycled exam text.
The system creates original expert tasks with known, auditable solutions while preserving the subject breadth, multimodal demands, and answer discipline that make frontier academic evaluations useful.
Each problem is authored or generated inside a controlled provenance workflow, independently solved, normalized into an exact or multiple-choice contract, and checked for ambiguity before entering an RL environment. Dense reward can combine correctness, rationale coverage, output format, and confidence calibration while strict accuracy remains the authoritative headline.
One schema, many fields of expertise.
Exact derivation
Algebra, probability, topology, numerical analysis, coding theory, and related mathematical domains.
Cross-disciplinary precision
Physics, chemistry, biology, medicine, and scientific inference with unit and convention checks.
Structured reasoning
Computer science, engineering, algorithms, control, and distributed systems, including image-backed prompts.
Knowledge plus interpretation
Humanities, social science, linguistics, formal semantics, policy, and other expert knowledge areas.
The judged GPT-5.6 Sol Pro run was broad, multimodal, and frozen before grading.
| Slice | Tasks | Accuracy | Mean reward |
|---|---|---|---|
| Multimodal | 7 | 100.00% | 0.947593 |
| Text-only | 43 | 95.35% | 0.915367 |
| Multiple choice | 20 | 95.00% | 0.904615 |
| Short answer | 30 | 96.67% | 0.930055 |
The candidate files were SHA-256 locked before the verifier and remained byte-identical afterward. The bundle retains per-item scores and concise derivations while excluding private hidden chain-of-thought.
Misses become targeted post-training material.
- DFA minimization: the submitted reachable-state interpretation produced 3 while the packaged verifier expected 6; the audit preserves the disagreement rather than silently changing the key.
- Normal-load reliability: the numerical calculation reached 6.889986…, but the selected multiple-choice option did not match the keyed rounded representation 6.8900.
- Training views: derive an ambiguity audit for the first and an answer-selection/format repair for the second.
This separation matters. One result is a potential task-quality or convention issue; the other is a clean agent error. They should not produce the same label or training response.
What an HLE-shaped program includes.
- Original text and multimodal problems with internally authored solutions, field labels, difficulty metadata, and answer contracts.
- Independent solve, ambiguity, duplication, leakage, image-rights, and formatting review before release.
- Exact, numeric-tolerance, multiple-choice, or structured graders plus optional rationale and calibration dimensions.
- Blind-run response files, concise derivations, confidence values, grader receipts, per-item labels, critiques, and repairs.
- Private holdouts designed around underrepresented fields and model-specific failure clusters.
Benchmark relationship and source boundary.
This is an independent clean-room task family inspired by the broad expert and multimodal capability tested by Humanity's Last Exam. It is not official HLE data, an affiliation, or a leaderboard submission.
Primary reference: Humanity's Last Exam by the Center for AI Safety and Scale AI. Evidence summary: JSON.