Conformance
TorchCTS turns backend configuration into repeatable tests and structured results. Versioned operator contracts define the applicable PyTorch behavior, the manifest selects the backend support surface, and result artifacts record what ran and what did not.
Configuration To Result Pipeline
TorchCTS follows a fixed pipeline. The manifest configures support. Versioned operator contracts define what applies to the installed PyTorch version. Filtering records non-executed cases. Pytest executes the included cases. Artifacts explain the outcome.
Device, import path, dtype support, capabilities, semantic level, and resource limits.
PyTorch-version-specific operator and dtype expectations define what is fair to test.
Unconfigured, out-of-contract, or out-of-scope cases are recorded as filtered before pytest runs them.
Included cases execute through the normal pytest harness with crash isolation when needed.
Resolved result artifacts, markdown reports, runlogs, coverage files, and crash evidence explain what happened.
PyTorch Operators
PyTorch user code eventually reaches operators. An operator is a named behavior such as add, matmul, reshape, or scaled_dot_product_attention.
TorchCTS works at the level where possible. One public Python call can reach different dispatcher overloads based on:
- dtype
- shape
- layout
- arguments
- PyTorch version
Backends
A backend is the software layer that makes PyTorch operations run on a target. In-tree backend families include:
- CPU
- CUDA
- ROCm
- MPS
- XPU
lets an out-of-tree package register a custom device name and dispatch behavior.
Successful backend import, device registration, and tensor creation are the first integration milestones. TorchCTS continues from there into operator values, dtypes, shapes, layouts, aliasing, mutation, errors, device behavior, and workloads.
PyTorch Is The Source Of Truth
TorchCTS builds on PyTorch operator schemas, dispatcher behavior, OpInfo metadata, CPU reference behavior, and testing conventions.
TorchCTS operator contracts are versioned testing artifacts. They help the suite select the behavior applicable to an installed PyTorch release. They do not replace PyTorch documentation, schemas, implementations, or upstream tests.
Operator Contracts
An operator contract is the expected behavior for a PyTorch operator surface.
For TorchCTS, an operator contract can cover:
- operator name and overload
- dtype behavior
- valid input conditions
- output values
- output shape, dtype, and layout
- aliasing and mutation behavior
- RNG behavior
- allocator behavior (memory reuse, contiguity, and storage semantics)
- error behavior
TorchCTS 0.4.1 includes versioned CPU dtype operator contracts for PyTorch 2.7.0 through 2.12.1.
How Expected Results Are Determined
PyTorch defines the operator semantics. TorchCTS selects the comparison method that fits those semantics and the test case.
- PyTorch reference behavior. Most value tests compare backend output with behavior derived from PyTorch operator metadata and CPU execution.
- Reviewed TorchCTS references. Some edge cases need an independent calculation so the test does not derive its expected result from the operation being checked.
- Legal-output checks. When an operator permits more than one valid representation, TorchCTS checks the required mathematical or structural properties instead of requiring one exact representation.
- Property checks. Aliasing, mutation,
out=identity, storage, RNG, and allocator behavior are checked directly.
See Evidence & Coverage for the runtime comparison and project QA model.
PyTorch Version Contract
Each PyTorch release can change the operator surface TorchCTS tests against:
- New operators. New ATen operators require new dtype contracts, test cases, and expected-behavior definitions before TorchCTS can cover them.
- Changed behavior on existing operators. Bug fixes, dtype promotion rule changes, stricter input validation, and changed defaults can all alter which inputs are accepted and what outputs are produced. Contracts collected under the previous version become wrong and must be re-collected.
- Expanded operator surface. New optional arguments, keyword parameters, or broader dtype support widen what an existing operator can do, requiring new test cases and updated contracts.
Adding support for a new PyTorch version requires a fresh operator-contract matrix, a diff against the previous version, validation of every change, and updated packaged contracts. TorchCTS only claims compatibility with versions that have gone through that process. Uncollected versions are treated as unvalidated and contract lookup .
Backend Manifests
The manifest configures the backend support surface. See Configure Backend for the workflow and Manifest Reference for exact fields.
It tells TorchCTS:
- which backend to import
- which device to test
- which dtypes are included
- which capabilities are included
- run depth
- resource limits
During development, the manifest controls the focused implementation surface. In release results, it also records the support represented by the run.
True includes matching in-contract cases in backend execution. False records the support as unconfigured and filters matching cases where applicable. A runtime unsupported result from an included case remains a backend failure or error.
When Configured Support Is Unavailable
A manifest can configure torch.float16 globally or for a regex-matched operator set. Matching in-contract cases are then included in the run.
If div.float16 is in contract for the installed PyTorch version and the configured backend path reports "unsupported dtype" at runtime, the manifest and runtime behavior are out of alignment. TorchCTS records a failure or error rather than filtering the case after execution begins.
What Is Fair To Test
TorchCTS only asks a backend to run cases that apply to the installed PyTorch version and the configured support surface.
Operator contracts keep unsupported PyTorch dtype and operator combinations out of backend execution while preserving the reason in structured accounting. Configured in-contract cases remain part of backend execution.
Filtered Is Not Pytest Skipped
TorchCTS tracks filtered cases in its own run accounting.
These filters are removed from pytest execution where possible and recorded as filtered cases in TorchCTS accounting:
- manifest filters
- dtype filters
- CPU operator contract filters
- semantic-level filters
Some tests are still pytest skip-marked for runtime conditions. Those include:
- missing capabilities
- device count
- CPU-only applicability
- resource limits
Pytest skip-marked cases are reported separately.
Semantic Levels
A backend's manifest declares its intended end-state support. Semantic levels let engineers run progressively deeper slices of that goal without changing the manifest.
- Level 1: core operator dispatch and basic dtype support.
- Levels 2-4: normal operator correctness, generated variants, mutation, aliasing, RNG, metadata, and broad production behavior.
- Levels 5-6: advanced numeric behavior, layout behavior, sparse and nested coverage, compiler paths, device API behavior, allocator behavior, and backend integration details.
- Levels 7-8: heavy workloads, multi-device behavior, release-depth stress, and adversarial cases.
This means a team can write the manifest once with their target support, then work through levels in order, concentrating effort on the cases that matter most before moving to finer-grained coverage.
The detailed level inventory, including node counts and what each level contains, is in Test Inventory.
Diagnostic Probes
TorchCTS can run setup probes for these areas:
- declared dtypes
- declared capabilities
- compiler behavior
Probe failures are diagnostic only. They are recorded in result metadata and probe artifacts, but they do not rewrite the manifest, skip tests, xfail tests, or change conformance outcomes.
Normal backend runs print probe output as a compact table with these columns:
- kind
- probe
- status
- note
These modes suppress the table to keep output readable:
- validation
- collect-only
- show-skips
- known-segfault audit
- xdist worker
- subprocess child
Result Meanings
The backend matched the expected operator behavior for that case.
The test ran and produced the wrong behavior. That can include:
- value
- shape
- dtype
- layout
- aliasing behavior
- mutation behavior
- error behavior
- property result
The test could not complete because the backend or runtime raised an unexpected error.
The case was included by the manifest and operator contract, but the backend reported unsupported or not implemented. This is an implementation result, not a filter, and remains a failure or error.
TorchCTS did not run the case and recorded why. Common reasons include:
- manifest
- dtype
- operator contract
- semantic level
- suite selection
- another structured reason
The process crashed, hung, timed out, or exited by signal. Isolation can contain the blast radius, but the crash still counts.