GPT-5.6 Sol Pro · judged blind runs

Professional work that ends in usable artifacts.

Original, source-grounded tasks across real occupations. The agent must inspect closed reference files, make defensible decisions, and deliver editable workbooks, reports, models, briefs, or other professional artifacts.

A renewable professional-work task factory.

The capability target is not document generation in isolation. It is the full chain from source inspection to a coherent, editable, decision-ready work product.

We build original briefs around occupations, artifact types, and operational decisions. A task can require a risk model and board report, a capacity plan and executive summary, or a portfolio recommendation backed by closed reference files. The data system varies the role, source pack, scenario, constraints, deliverable schema, and verifier while preserving a stable outer contract.

Evaluated model · GPT-5.6 Sol Pro

GPT-5.6 Sol Pro produced the three frozen 44-task blind-run bundles. Artifact verifiers and criterion-level graders then judged every submitted workbook, document, and PDF. These are Yotta/Ulam synthetic-suite results, not official GDPval or GDPval-AA leaderboard scores.

How the RL environment works.

01 / Observation

Closed source pack

Instructions, CSV/JSON inputs, policies, PDFs, and artifact requirements are mounted inside a reproducible Harbor task.

02 / Action

Artifact production

The agent creates exact filenames and editable deliverables such as XLSX, PDF, DOCX, or PPTX plus a concise decision record.

03 / Reward

Criterion-level scoring

Structural checks, formula tests, traceability, cross-file consistency, professional quality, and safety controls produce dense reward.

04 / Hillclimb

Failure-specific variants

Observed semantic mismatches become new task families, adversarial source packs, critiques, repairs, and targeted curriculum slices.

How GPT-5.6 Sol Pro performed across three judged blind runs.

v0.1 risk scenarios0.909732mean reward · 44 tasks
v2 work programs0.904354mean reward · 44 tasks
v3 contingencies0.830111mean reward · 44 tasks
SuiteVisible inputRequired outputObserved signal
v0.1Risk register + assumptionsScenario workbook + PDFStrong artifact execution across all occupations
v2Role-specific operational sourcesWorkbook + decision briefRecurrent definition mismatch exposed by hidden metrics
v3Contingency exposures + thresholdsStress workbook + 3-page PDF100% file integrity; semantic value-basis mismatch

In the v3 bundle, all 88 required files opened successfully, instruction following and artifact integrity averaged 1.0, and all 132 rendered PDF pages passed visual review. The principal blind failure was semantic: the rollout used contribution value while the hidden verifier expected total revenue. That is exactly the kind of narrow, consequential mismatch an RL environment should make trainable.

Trajectories preserve the work, not hidden chain-of-thought.

Representative record shapeJSON / JSONL
  1. Observe: load the brief, permitted reference files, required filenames, and protected evaluation boundary.
  2. Calculate: record disclosed formulas, selected source rows, derived metrics, and consistency checks.
  3. Build: produce the workbook and report, then run structural and visual preflight checks.
  4. Freeze: hash the exact submission before the hidden verifier runs.
  5. Score: attach criterion-level reward and diagnose the earliest consequential mismatch without rewriting the frozen artifact.

Production trajectory releases can include observations, tool events, concise rationales, artifact manifests, grader receipts, failure labels, critiques, and repaired variants. They do not need private hidden reasoning to be useful for post-training.

What a client receives.

  • Original train, development, and sealed evaluation tasks across selected occupations and work-product families.
  • Harbor-compatible environments with immutable source packs, exact deliverable contracts, and isolated verifier mounts.
  • Executable structural checks plus criterion-level rubrics for correctness, consistency, traceability, safety, and professional quality.
  • Frozen blind-run submissions, safe trajectories, per-criterion scores, rendered visual audits, and integrity manifests.
  • Hillclimb batches targeted at model-specific misses rather than prompt paraphrases.

Benchmark relationship and source boundary.

This is an independent, clean-room task system inspired by the capability shape of GDPval and GDPval-AA. It is not an official benchmark release, replica, endorsement, or leaderboard result. Original task materials and hidden evaluation assets remain separate from benchmark source data.

Primary references: OpenAI GDPval and Artificial Analysis GDPval-AA. Page evidence is summarized in machine-readable JSON.

Professional work environments

Build the artifact tasks your model still fails.

Start a program