A renewable professional-work task factory.
The capability target is not document generation in isolation. It is the full chain from source inspection to a coherent, editable, decision-ready work product.
We build original briefs around occupations, artifact types, and operational decisions. A task can require a risk model and board report, a capacity plan and executive summary, or a portfolio recommendation backed by closed reference files. The data system varies the role, source pack, scenario, constraints, deliverable schema, and verifier while preserving a stable outer contract.
How the RL environment works.
Closed source pack
Instructions, CSV/JSON inputs, policies, PDFs, and artifact requirements are mounted inside a reproducible Harbor task.
Artifact production
The agent creates exact filenames and editable deliverables such as XLSX, PDF, DOCX, or PPTX plus a concise decision record.
Criterion-level scoring
Structural checks, formula tests, traceability, cross-file consistency, professional quality, and safety controls produce dense reward.
Failure-specific variants
Observed semantic mismatches become new task families, adversarial source packs, critiques, repairs, and targeted curriculum slices.
How GPT-5.6 Sol Pro performed across three judged blind runs.
| Suite | Visible input | Required output | Observed signal |
|---|---|---|---|
| v0.1 | Risk register + assumptions | Scenario workbook + PDF | Strong artifact execution across all occupations |
| v2 | Role-specific operational sources | Workbook + decision brief | Recurrent definition mismatch exposed by hidden metrics |
| v3 | Contingency exposures + thresholds | Stress workbook + 3-page PDF | 100% file integrity; semantic value-basis mismatch |
In the v3 bundle, all 88 required files opened successfully, instruction following and artifact integrity averaged 1.0, and all 132 rendered PDF pages passed visual review. The principal blind failure was semantic: the rollout used contribution value while the hidden verifier expected total revenue. That is exactly the kind of narrow, consequential mismatch an RL environment should make trainable.
Trajectories preserve the work, not hidden chain-of-thought.
- Observe: load the brief, permitted reference files, required filenames, and protected evaluation boundary.
- Calculate: record disclosed formulas, selected source rows, derived metrics, and consistency checks.
- Build: produce the workbook and report, then run structural and visual preflight checks.
- Freeze: hash the exact submission before the hidden verifier runs.
- Score: attach criterion-level reward and diagnose the earliest consequential mismatch without rewriting the frozen artifact.
Production trajectory releases can include observations, tool events, concise rationales, artifact manifests, grader receipts, failure labels, critiques, and repaired variants. They do not need private hidden reasoning to be useful for post-training.
What a client receives.
- Original train, development, and sealed evaluation tasks across selected occupations and work-product families.
- Harbor-compatible environments with immutable source packs, exact deliverable contracts, and isolated verifier mounts.
- Executable structural checks plus criterion-level rubrics for correctness, consistency, traceability, safety, and professional quality.
- Frozen blind-run submissions, safe trajectories, per-criterion scores, rendered visual audits, and integrity manifests.
- Hillclimb batches targeted at model-specific misses rather than prompt paraphrases.
Benchmark relationship and source boundary.
This is an independent, clean-room task system inspired by the capability shape of GDPval and GDPval-AA. It is not an official benchmark release, replica, endorsement, or leaderboard result. Original task materials and hidden evaluation assets remain separate from benchmark source data.
Primary references: OpenAI GDPval and Artificial Analysis GDPval-AA. Page evidence is summarized in machine-readable JSON.