Evidence & Coverage
TorchCTS measures which backend-relevant PyTorch behavior has a test strategy, which cases executed, and which work remains pending or excluded.
Coverage makes the focused-suite tradeoff reviewable. TorchCTS can remain much smaller than a framework-wide test sweep because each backend-relevant dispatcher surface receives an explicit coverage disposition.
PyTorch Support
Supports PyTorch 2.7.0-2.12.1.
TorchCTS 0.4.1 includes versioned CPU dtype operator contracts collected for 2.7.0, 2.7.1, 2.8.0, 2.9.0, 2.9.1, 2.10.0, 2.11.0, 2.12.0, and 2.12.1. Package metadata bounds PyTorch to torch>=2.7.0,<2.12.2.
PyTorch defines the behavior. TorchCTS records version-specific testing expectations for the PyTorch releases it supports.
Operator Contract Flow
TorchCTS reads the installed PyTorch version, dispatcher inventory, and matching packaged operator contract data.
The manifest selects the device, dtypes, capabilities, and semantic levels included in the run.
Unconfigured, out-of-contract, or out-of-scope cases are removed from pytest execution where possible and recorded with a reason.
Each included case uses PyTorch reference behavior, a reviewed TorchCTS reference, a legal-output check, or a direct property check.
The case runs on the selected backend. Runtime unsupported errors, wrong results, and crashes remain backend results.
Result artifacts, filtered accounting, crash records, reports, and coverage files are written for review.
Operator Contract Evidence
The packaged operator contract file records dtype behavior by dispatcher entry and PyTorch version range.
The packaged runtime contract file is compact. Source evidence is tracked in the TorchCTS repository under evidence/pytorch/dtype-contracts/ for audit and regeneration. Wheels and sdists do not ship that source evidence.
Expected Result Sources
PyTorch defines the operator semantics. TorchCTS uses four comparison methods to test those semantics on a backend.
Most value tests use PyTorch operator metadata, sample inputs, and CPU execution to establish the expected result.
Some cases use an independently implemented calculation. This avoids deriving the expected answer from the same operation being tested and supports cases that need higher-precision or operator-specific handling.
Some operators can return more than one valid representation. TorchCTS checks the required mathematical or structural properties instead of requiring exact agreement with one representation.
Aliasing, mutation, out= identity, storage relationships, RNG behavior, and allocator behavior are checked directly.
Expected-result routing is specific to the operator, dtype, input condition, and arguments. A matched TorchCTS reference does not silently fall back to another comparison if reference construction fails.
Oracle Quality Assurance
TorchCTS-owned references have a separate development and release QA suite in the TorchCTS source repository.
- Expected values come from reviewed fixed records.
- The QA suite does not generate an expected value by calling the TorchCTS reference being tested.
- It does not use the same PyTorch CPU operation that the reference replaces as its independent expected-value source.
- Checks cover fixed values, legal-output rules, backward formulas, routing, provenance, schema behavior, and backend-specific execution rules.
- Release checks verify the accepted case inventory and evidence manifests.
This QA suite is not part of an installed backend run. Its cases, generators, and raw evidence stay in the source repository, and its results do not contribute to backend conformance totals.
See the oracle QA guide and oracle authoring rules in the TorchCTS repository.
Evidence Storage Model
TorchCTS keeps runtime data, source evidence, backend evidence, and user run results in separate places.
Compact operator contracts, operator metadata, the path-shape corpus, templates, and known crash rules that TorchCTS needs while installed.
evidence/pytorch/dtype-contracts/Version-scoped PyTorch dtype source evidence used to audit and regenerate the compact runtime operator contract file.
evidence/backends/Canonical backend evidence records. These records are path-free and avoid host identity, full audit snapshots, and transport archive metadata.
Reviewed cases, source records, generators, and raw evidence used during TorchCTS development and release validation. These remain in the source repository.
results/User run output. Completed runs use canonical history artifacts and a latest reference file that TorchCTS tools resolve automatically.
sample-results/Committed examples of real result artifacts, reports, runlogs, filtered accounting, and crash evidence.
How Operator Contracts Are Used
During collection, TorchCTS checks whether a dtype/operator case is inside the PyTorch CPU operator contract for the installed PyTorch version.
- If the case is supported and included by the backend configuration, TorchCTS runs it.
- If the case is unsupported, pending, or unknown in the operator contract, TorchCTS filters it and records the reason.
- If the installed PyTorch version is outside the validated matrix, TorchCTS warns because the packaged operator contract is not validated for that version.
Dispatcher Coverage Snapshot
| Metric | Value |
|---|---|
| ATen overloads inventoried | 3,225 |
| 3,214 | |
| Covered backend-relevant overloads | 3,062 |
| Dispatcher coverage | 95.3% |
| 0 | |
| 96 | |
| 56 |
Dispatcher coverage is one measurement of suite breadth. It does not replace semantic tests for values, dtypes, layouts, aliasing, mutation, autograd, device APIs, memory behavior, workloads, or crashes.
Coverage Kinds
Internal key: opinfo. Coverage mapped from PyTorch's OpInfo operator metadata and sample inputs.
handwrittenExact coverage from TorchCTS-authored tests.
generatedMaterialized generated dispatcher cases.
TorchCTS-owned value oracle.
propertyChecks for aliasing, mutation, out= identity, metadata, RNG, or allocator behavior.
backend_packBackend-gated coverage for vendor or build-family surfaces.
Reviewed proof that a public entrypoint reaches an exact dispatcher surface.
excludedReviewed surface that does not count as covered.
Expected-result methods and coverage kinds answer different questions. For example, a PyTorch OpInfo-derived test can use a legal-output check while its coverage kind remains opinfo.
Generated Coverage
The generated suite currently has 4,258 pytest nodes.
Coverage tracks 1,896 generated coverage surfaces and 1,907 required generated dispatcher cases. There are no optional generated dispatcher cases in the current stats artifact.
Targeted Path-Shape Coverage
TorchCTS 0.4.1 includes a curated path-shape corpus. It is not a cartesian product. Each row is a reviewed shape, dtype, layout, stride, cost, and resource-tier case built to hit real backend branch points.
The corpus contains 1,320 total rows. The default selection contains 850 rows. The heavy tier contains 470 rows.
| Family | Total Rows | Default Rows |
|---|---|---|
| matmul | 210 | 140 |
| attention | 150 | 80 |
| convolution | 140 | 90 |
| reduction | 120 | 85 |
| indexing | 110 | 80 |
| spatial | 100 | 65 |
| model_patterns | 95 | 55 |
| normalization | 90 | 65 |
| fft | 85 | 45 |
| sorting | 80 | 55 |
| broadcasting | 75 | 55 |
| linear_algebra | 65 | 35 |
The default collection includes 850 path-shape nodes. It selects standard-tier rows from levels 5 and 7. Heavy tier rows include additional level 5, 7, and 8 cases for targeted stress runs.
Backend Packs
Backend packs cover surfaces that require a specific backend, vendor library, or build family.
A backend-pack runner that only skips on the current host does not demonstrate coverage. Backend-private coverage requires canonical backend evidence from a build that can execute the target dispatcher path.
promote_nowAccepted contract, runner, and evidence already support promotion.
candidate_onlyA real runner exists, but matching backend evidence has not promoted it.
blocked_contractNo accepted source-derived contract.
blocked_schemaNo safe exact dispatcher invocation path.
blocked_hardwareRequires unavailable backend or build hardware.
blocked_runtimeKnown runtime blocker such as OOM or unsupported layout.
torchcts coverage collect-backend-evidence --store evidence/backends --device cuda --backend-gate cuda
torchcts coverage collect-backend-evidence --store evidence/backends --device cuda --backend-gate cuda --run-pending-candidates --require-oracle-results --fail-on-oracle-failure
torchcts coverage collect-backend-evidence --store evidence/backends --device cuda --surface aten::_fused_dropout --surface aten::_fused_dropout.out
Package Artifact Policy
Wheels and sdists include compact runtime contracts, operator metadata, the path-shape corpus, runtime reference implementations, templates, and known crash rules. They do not ship PyTorch source evidence, backend evidence, oracle QA cases, oracle generators, raw oracle evidence, sample results, local run output, or release scratch files.
Coverage Release Gate
torchcts coverage audit
torchcts coverage check --fail-on-unknown
torchcts coverage report
The release target is zero unknown tensor-touching backend-relevant surfaces. Pending and excluded rows remain visible so a smaller focused suite does not create invisible coverage gaps.
A backend team needs to know which dispatcher surfaces are tested, pending, excluded, or still unclassified before the package ships.