GPT-5.6 Sol Pro · judged blind run

Graduate science built for useful wrong answers.

Original biology, chemistry, and physics questions with difficult distractors, exact choices, concise scientific derivations, checkpoint claims, elimination evidence, and confidence-aware reward.

New questions with adversarially plausible distractors.

The target is scientific reasoning under uncertainty: retrieve the right theory, execute the calculation, reject answers that encode common conceptual errors, and commit to a calibrated choice.

Our system generates original question families around a scientific mechanism rather than templating one prompt. Subject experts and programmatic checks validate units, constants, conventions, answer uniqueness, numerical tolerances, and distractor provenance. Each distractor is labeled by the misconception or calculation failure it represents.

Evaluated model · GPT-5.6 Sol Pro

GPT-5.6 Sol Pro produced 33 frozen blind responses. The exact-choice verifier and criterion graders then judged answers, checkpoints, rationales, distractor eliminations, and calibration. This is a synthetic-suite result, not an official GPQA Diamond benchmark score.

Exact accuracy stays authoritative; dense reward makes it trainable.

Choice

Exact correctness

The selected option receives the benchmark-style pass/fail outcome.

Checkpoints

Intermediate claims

Numerical or conceptual checkpoints show whether the derivation reached the right scientific state.

Elimination

Distractor diagnosis

Credits correct rejection of options without letting eloquent text override a wrong final choice.

Calibration

Confidence discipline

Rewards confidence that matches correctness and creates useful abstention or review policies.

GPT-5.6 Sol Pro solved every exact choice in the judged v3 run.

Biology11 / 11mean reward 0.968177
Chemistry11 / 11mean reward 0.946207
Physics11 / 11mean reward 0.961359
DimensionMeanInterpretation
Exact accuracy1.000000All 33 selected choices correct
Checkpoint score1.000000All keyed intermediate claims matched
Rationale score0.704545Hidden rationale groups create remaining training signal
Elimination score0.762626Distractor-specific explanation is incomplete in places
Calibration score0.999900Confidence 0.99 on each correct task

No answer was revised after verifier output. All response hashes remained unchanged through scoring, and the run used no web, network, model API, solution file, or hidden specification during solving.

Correct answers still contain hillclimb signal.

Safe science trajectoryPer-task JSON
  1. Observation: question, four choices, allowed local tools, and confidence schema.
  2. Derivation: concise equations, scientific checkpoints, assumptions, and unit conversions.
  3. Decision: selected option plus a short reason each distractor fails.
  4. Freeze: hash the answer artifact before any hidden grading.
  5. Reward: attach exact choice, checkpoint, rationale, elimination, calibration, and integrity dimensions.

A 100% exact score does not end the data loop. Missing rationale groups and weak distractor diagnoses can seed new variants, contrastive explanations, critiques, and targeted preference pairs.

What a GPQA-shaped program includes.

  • Original graduate-level question families across biology, chemistry, physics, or client-selected scientific domains.
  • Gold choices, exact derivations, independent solves, difficulty review, and misconception-mapped distractors.
  • Exact-choice reward plus optional checkpoint, rationale, elimination, calibration, and artifact-integrity dimensions.
  • Frozen blind responses, concise trajectory records, per-dimension scores, critiques, and repaired responses.
  • Private eval families separated by mechanism and source lineage to reduce near-duplicate leakage.

Benchmark relationship and source boundary.

This is an independent clean-room task system inspired by the graduate science capability shape of GPQA Diamond. It does not reproduce official questions and is not affiliated with the benchmark authors.

Primary reference: GPQA: A Graduate-Level Google-Proof Q&A Benchmark. Evidence summary: JSON.

Scientific reasoning data

Turn distractors into a diagnostic instrument.

Start a program