Benchmark what the system proves—and what it refuses.

A useful industrial benchmark varies identity, geometry, pose and evidence quality, then measures successful realization and calibrated refusal separately.

Concept illustration of six controlled engineering test fixtures producing pass, review and fail-closed outcomes
Concept illustration — no performance percentage is implied; published results require a frozen fixture set and auditable run records.

01 / VARIATION MATRIX

A replay of one happy path is not a benchmark.

Cases are grouped by the engineering fact that changed so reviewers can distinguish lookup, geometric reasoning, generalization and safe refusal. These six cards define benchmark dimensions; they are not article links.

Browse published research
IDENTITY

Unseen approved part number

Tests whether reasoning survives beyond a previously seen identifier.

REPRESENTATION

Face ordering change

Detects hidden dependence on brittle hardcoded face indices.

POSE

Orientation variation

Tests interface reasoning under transformed component poses.

GEOMETRY

Split or merged surfaces

Challenges formal interface evidence across B-Rep representation changes.

AMBIGUITY

Competing interfaces

Measures whether uncertainty is surfaced for engineering review.

MISSING TRUTH

Insufficient evidence

Requires a fail-closed or cannot-verify result instead of a plausible guess.

Outcome model

Success and refusal are different metrics.

A system that passes many valid cases but invents answers on invalid cases is not safe enough for engineering. Reporting must keep both sides visible.

VALID CASECorrect native realizationArtifact + reopened evidence + verdict
REVIEW CASECalibrated escalationReason and missing decision are explicit
INVALID CASECorrect refusalNo unsupported artifact is released

02 / PUBLICATION CONTRACT

Freeze the fixture before publishing the number.

  1. Declare the scope

    Part families, drawing conventions, Rulepack, CAD version and unsupported conditions are named.

  2. Freeze inputs and expected obligations

    Hashes, expected outcomes and review rules are fixed before execution.

  3. Preserve every run record

    Task identity, artifact hash, read-back facts, refusal reason and reviewer decision remain auditable.

  4. Report error classes, not only an average

    False pass, false refusal, geometry failure, execution failure and unverifiable outcome remain separate.

TECHNICAL DILIGENCE

Ask for the fixture, the failure classes and the raw verdicts.

Request an evidence review