# Testing and quality contracts

Status: M1 core plus M2 layout/element and M3 scene foundation tests; current headless tests run in CI.

The current test coverage exercises the headless `primitives` value layer,
core app/entity/scheduler lifecycle, deterministic flex-line and recursive
flex-tree layout, the flat `element` tree, a declarative Render/IntoElement and
request-layout/prepaint/paint path, ordered `scene` commands, and the
implemented versioned `SceneSnapshot` v1 subset. It includes deterministic
value and lifecycle cases, two seeded rectangle properties, a seeded entity
lifecycle reference-model property, seeded flex-line and flex-tree geometry/order
properties, and exact element lifecycle, hit-test/dispatch/focus/scene-order tests.
Mutation coverage now includes independently scoped `capability/` and `mcp/`
observations with reviewed baselines; their original final-head stability/CI
merge requirements are historical provenance recorded below. Full-schema scene golden fixtures,
text/path/image rendering, mutation coverage of other critical packages, and
broader native integration assertions remain unimplemented. Bounded recursive auto container sizing is now included in the
active layout test surface.
Event/focus PBT exercises route reversal, global
stop-propagation, and focus normalization after subtree removal; scene PBT
exercises strict clip-stack nesting and rejects unclosed stacks before snapshots
are emitted. Deterministic callback tests cover capture/bubble execution and
stop-propagation. Scene tests cover both the provisional `CommandSnapshot` and
the versioned v1 subset, including canonical serialization, negative-zero
normalization, envelope validation, and deterministic clip-chain reuse. This document
separates active checks from future budgets. Release status is tracked in
[release gates](release-gates.json) and the [release policy](release.md).

## Test layers

| Layer | Purpose | Required evidence |
| --- | --- | --- |
| L0 compile/static | Formatting, warnings, package checks, target builds | Tool versions and complete command output |
| L1 deterministic examples | Exact API contracts and regressions | Test result plus case identifier |
| L2 property-based | Geometry, layout, entity, event, focus, and scene invariants | Seed, generated-case count, shrink trace, final counterexample |
| L3 model-based | Compare lifecycle command sequences with an independent reference model | Model version, command sequence, seed, expected and actual observations |
| L4 mutation | Measure whether tests detect meaningful source changes | Per-package/operator report, baseline revision, survivor and timeout list |
| L5 visual | Compare normalized scene data and pinned raster output | Fixture manifest, baseline/current/diff images, rendering environment |
| L6 artifact review | Optional external image and accessibility-oriented review | Input artifact checksum and review result; never a native-driver pass |
| L7 native integration | Exercise windows, input, focus, text services, accessibility, and recovery | Platform, runner, test result, logs, and attached artifacts |

The active M1 property suites are `geometry.rect.invariants` with seed
`12725326473013217044` (`0xb09971016ad3c714`) and
`geometry.rect.translations` with seed `13790154211114911481`
(`0xbf6078310f5c86f9`). Each currently runs 256 cases with `max_size=64` and
`max_shrinks=500`. QuickCheck comes from the pinned MoonBit core library and is
imported as `moonbitlang/core/quickcheck` only with `for "test"`; the dependency
audit rejects it in runtime imports. The `core-model` entity lifecycle
reference-model property uses seed `584118800423729943`
(`0x081b34f445e23b17`), 256 cases, `max_size=32`, and `max_shrinks=500`.
The portable text editing property `text.utf16-replacement.invariants` uses seed
`17836885540012499121` (`0xf78958e50fbf20b1`), 256 cases, `max_size=32`, and
`max_shrinks=500`; its reference strings are independently assembled from
generated Unicode scalar fragments. The portable composition property
`text.utf16-composition.invariants` uses seed `4169804048700389933`
(`0x39de1eb091c5ba2d`), 256 cases, `max_size=32`, and `max_shrinks=500`. It
generates three previews followed by either commit or cancel, independently
assembles prefix/preview/suffix/final strings, and checks that prior immutable
snapshots remain unchanged. The strict text offset bridge property
`text.utf16-utf8-bridge.invariants` uses seed `4815306127957175914`
(`0x42d3676929a62e6a`), 256 cases, `max_size=32`, and `max_shrinks=500`. It
builds independent UTF-16 and UTF-8 boundary tables from generated scalar
fragments, then checks every boundary, round trip, surrogate interior, and
UTF-8 byte interior. These exact counts and seeds are active.
The command-palette property `command-palette.reference-filter-selection-v1`
uses derived seed `8367235218704275144` (`0x741e6419930196c8`), 256 cases,
`max_size=32`, and `max_shrinks=500`. Generated enabled-row arrays and operation
traces are checked after every filter/navigation step against an independent
index/viewport model. QuickCheck is imported only with `for "wbtest"`.

The larger nightly and stress budgets below are policy targets and are not
scheduled jobs yet.

Headless tests are the default for core semantics. Pixel comparisons are
platform-specific unless the same deterministic rasterizer and assets are used.
Semantic scene changes and raster changes have separate oracles. The optional
vlmkit workflow may inspect PNG output, but it does not drive native windows and
does not become a runtime dependency.

Linux text layout has a separate headless native test gate. It installs
PangoFT2/Fontconfig only on Ubuntu Linux, installs the DejaVu and Noto test
fixtures, and runs without `DISPLAY` or `WAYLAND_DISPLAY`; it does not use the
Wayland compositor or test raster output. Resolved system-library/font versions
are logged because exact metrics vary by distro and font revision. See the
[Linux text guide](linux-text.md) and
[`ubuntu-native.yml`](../.github/workflows/ubuntu-native.yml).

## Determinism and workload budgets

Except for the explicitly declared bounded regression below, every property or
generated model test uses a recorded unsigned 64-bit seed.
The fixed PR seed root is `0x47505549` (`GPUI`). A suite seed is the first 64
bits of `SHA-256("gpui.mbt:pbt:v1:<root-hex>:<suite-id>")`, interpreted as an
unsigned big-endian integer. Current QuickCheck properties pass this suite seed
directly to the core library. For generated model cases, derive a replay seed
from the suite seed and stable case identifier; a model runner must pass and
report it. If the selected API cannot accept the required seed, that suite is
not eligible for a deterministic gate until an adapter is provided.

The fixed-sequence ASCII history regression in
[`history_reference_test.mbt`](../controls/text_field/history_reference_test.mbt)
is an explicit bounded deterministic exception: it records literal `Int` seeds
`1`, `1777`, and `4242`, runs 256 commands per seed with at most 12 ASCII
characters, and checks an independent journal plus retained snapshot branches.
It does not derive unsigned-64-bit suite/case seeds or shrink failures. Replay
uses the recorded literal seed and source revision; this regression is not the
general generated-model runner or a nightly/stress qualification. The fixed-sequence palette viewport regression in
[`command_palette_pbt_wbtest.mbt`](../controls/command_palette/command_palette_pbt_wbtest.mbt)
is another explicit bounded deterministic exception: literal `Int` seeds `1`,
`1777`, `4242`, and `65535`, six catalog sizes (0, 1, 8, 9, 32, 128), and 256
operations per pair yield 6,144 reproducible intermediate states. Failures show
catalog size, literal seed, operation index, and step. This boundary regression
does not shrink or claim general generated-model qualification; the separate
QuickCheck suite above follows the derived seed policy. The seed
policy above remains the requirement for new general property/model suites.

| Job | Seeds | Cases per property | Max model commands | Shrink candidates | Wall-clock cap |
| --- | --- | ---: | ---: | ---: | ---: |
| Pull request | Fixed seed root, stable suite seeds | 256 | 32 | 500 | 5 minutes total for PBT |
| Nightly expanded | Fixed root plus 10 recorded rotating roots | 2,000 per seed | 256 | 5,000 per failure | 30 minutes total for PBT |
| Scheduled stress | 100 recorded rotating roots | 10,000 per root, divided into bounded batches | 2,048 | 20,000 per failure | 2 hours per platform job |

These are initial policy caps, not measured throughput claims. The current PR
workflow enforces 256 cases per active geometry property; total PBT wall-clock
and model-command caps are not yet enforced. Nightly and stress rows are future
jobs. Before a budget is made blocking, record runner class, MoonBit toolchain,
target, elapsed time, and case count. A budget change requires a reviewed update
to this table and the CI configuration. A timeout is a failure with its seed and
last completed case; rerun-to-pass does not erase it.

For every failed generated case, retain the original seed, derived suite/case
seeds where applicable (or the declared literal regression seed), property or
model identifier, generator version, attempted case count,
shrink trace, minimized counterexample, toolchain, target, and reproducing
command. A failure should be reproducible from the recorded seed and fixture
revision alone.

## Reference-model strategy

The model is deliberately small and independent of implementation internals.
It specifies only public observations and valid command preconditions. The
first model covers entity identity and lifecycle with create, read, update,
observe, subscribe, unsubscribe/drop, notify, and destroy commands. The runner
compares observable state after every command, not only at sequence end.

Generators produce valid commands from the model state. Shrinking removes
commands or simplifies their values while preserving command validity. The
model records notification count/order, subscription liveness, entity liveness,
and externally visible identity. Later models add focus/event trees, window
lifecycle, task cancellation, and platform services. Each model has a versioned
contract identifier so that a failing sequence remains interpretable after a
generator changes.

Core property families include:

- geometry round trips, containment, intersection symmetry, clamping, and
  transform composition within documented precision;
- layout determinism, finite/non-negative geometry where required, containment,
  constraint satisfaction, stable order, and idempotent relayout;
- entity liveness, identity, deterministic observer order, documented
  notification behavior, and subscription drop behavior;
- event propagation order, stop-propagation, live-tree focus, and valid focus
  after destruction;
- stable scene ordering, clip nesting, normalized serialization, and equivalent
  transforms yielding equivalent normalized output.

## Fixture and visual-artifact contract

Every committed fixture has a stable ID and a manifest entry containing its
kind, source/provenance, input checksum, expected semantic result, target scope,
and any required environment. Fixtures are original unless an explicit
compatible license and attribution are recorded. Generated fixtures include
the generator version and seed. Text fixtures state script coverage, font files
and checksums, locale, scale factor, and expected shaping/layout observations.

Raster goldens pin the runner image, renderer, fonts, locale, device scale,
color profile when applicable, and animation/time state. A visual failure
retains baseline, current output, and a difference image. Font or renderer
updates require an explicit baseline review. Exact pixels are not compared
across different native text/rendering stacks.

CI retains these artifacts for diagnosis:

| Artifact | Minimum contents | Retention target |
| --- | --- | ---: |
| PBT failure | Seed manifest, command, shrink trace, counterexample | 14 days for PR; 30 days for nightly |
| Mutation result | turtles JSON, baseline revision, survivor/timeout diffs | 14 days for PR; 30 days for nightly |
| Scene result | Normalized expected/current scene and semantic diff | 14 days for PR; 30 days for nightly |
| Raster result | Fixture manifest, baseline/current/diff PNGs | 14 days for PR; 30 days for nightly |
| Native failure | Platform/test identity, crash or native logs, relevant screenshots | 14 days for PR; 30 days for nightly |

Each artifact bundle has a manifest with repository revision, workflow/run ID,
target, toolchain, suite/fixture IDs, and checksums. A red gate must identify
the failed case without requiring a local reproduction first.

## Mutation ratchet

Use turtles as development/CI tooling, never as a public runtime dependency.
Mutation reports are split by package and mutation operator. A surviving or
timed-out viable mutant is an actionable test gap. A mutant can be classified
as equivalent or redundant only with a short reviewed rationale linked from
the report.

The ratchet is:

1. Record a non-blocking deterministic baseline by package and operator.
2. Make the gate blocking only after a rerun proves the baseline is stable.
3. Require each package/operator score to stay at or above its reviewed
   baseline; do not hide a subsystem regression in one global percentage.
4. Raise thresholds as APIs stabilize. Production-critical core packages target
   100% of viable mutants killed, with only documented equivalent/redundant
   exceptions.
5. At release, allow zero unexplained `SURVIVED` or `TIMEOUT` mutants in
   critical core packages. Report property-only kills separately and convert
   useful shrunk witnesses into permanent regression examples.

The repository pins `turtles` 0.3.0 in
[`mutation-primitives.yml`](../.github/workflows/mutation-primitives.yml) and
selects the stable primitives core in [`turtles.toml`](../turtles.toml). The
job runs the normal deterministic and fixed-seed property tests in turtles'
isolated copy, retains schema-2 JSON plus survivor/timeout diffs, and audits the
report before applying the checked-in
[`primitives` ratchet baseline](../mutation-baselines/primitives.json).

The accepted initial baseline comes from hosted PR #14 run `37202065139` at
head `0dac8f5e9141a960cf0d85e23da81c30e04f26dd`: 140 of 145 viable mutants
were killed (96.552%), with five survivors and zero timeouts. The operator
floors are arithmetic 20/20, boolean 25/28, comparison 37/37, condition 54/56,
and literal 4/4. CI compares exact killed/viable fractions using integer
cross-products rather than rounded percentages. The job fails if the overall
fraction regresses, any operator fraction regresses, or timeout count rises
above the recorded baseline. Scope, turtles version, target, and test-scope
metadata must also match, so a toolchain/scope change cannot silently inherit
the old threshold.

This completes the first blocking `primitives/` mutation ratchet on Moon's
default test target. It does not complete the broader release mutation gate:
the five current survivors still need either stronger tests or reviewed
equivalent/redundant classifications, and all-target mutation plus additional
critical packages remain open. The release ledger therefore remains pending.

### Capability and MCP semantic ratchets

The separate [`mutation-capability-mcp.yml`](../.github/workflows/mutation-capability-mcp.yml)
matrix runs turtles 0.3.0 over exactly `capability/` or `mcp/`, using full-module
tests on Moon's default target. The audited PR #17
[observation run](https://github.com/gpui-mbt/gpui.mbt/actions/runs/37241489153)
killed 271/283 capability mutants and 262/271 MCP mutants. Its 12 and nine
respective survivors have independently reviewed equivalent/redundant
rationales; neither scope had timeouts, unviable mutants, or in-scope parser
skips. The observations led to 41 deterministic regression tests and the
checked-in [capability](../mutation-baselines/capability.json) and
[MCP](../mutation-baselines/mcp.json) baselines.

The semantic gate compares exact overall and per-operator killed/viable
fractions, permits no timeouts, rejects any unreviewed survivor or identity/diff
swap, and checks the Moon driver build identity. Compiler/core pins remain in
the workflow; the report does not independently attest those two hashes. CI
retains observation JSON, ratchet JSON when enforcement succeeds, raw reports,
console logs, and survivor diffs for 14 days on PRs and 30 days otherwise.

The source observation deliberately failed only because the reviewed baseline
files were not yet present. For the original PR #17 acceptance,
a successful hosted baseline stability rerun and all required final-candidate
checks were explicit merge gates; exact successful run links belonged in that
PR's merge evidence. This historical provenance does not impose a hosted wait
on current Linux development. The source
observation alone is not a green stability result. See
[semantic mutation evidence and review caveats](semantic-mutation.md) for exact
operator floors and provenance. Closure of
[0014](../issues/closed/0014-gui-api-mcp-equivalence-conformance.md) takes effect
only with that verified merge,
while [0016](../issues/open/0016-mcp-endpoint-lifecycle-conformance.md) retains
asynchronous cancellation/disconnect and native/browser endpoint topology.
The broader mutation and production release gates stay pending.

## Local Linux acceptance

Current affected Linux development uses the pinned local
[actrun gate](../infra/linux-desktop/ACTRUN.md), without waiting for GitHub
Actions. Default fast mode runs contract/Python, infrastructure/recipe,
formatting, Linux text C/native and affected native package checks. Explicit
acceptance mode adds all Linux-runnable targets and a freshly built,
manifest-bound private-display native keyboard replay. Legacy headless mode
remains a labeled text-only subset.

Green applies only to the selected local scope. Every run records required
passed/failed/not-run steps and skipped coverage. macOS/Windows native, real
browser proof, mutation, MCP stdio interoperability and the separate Ubuntu
scale/readback/lifecycle GPU suite remain separate scopes;
Linux success makes no claim for them. Current catalog skips for adaptive held
repeat and unsupported GPUI IME remain visible. A native IPC denial fails wider
acceptance rather than turning unexecuted native cases Green. Broader
production release gates remain pending and are unchanged by this development
policy.

## CI split

Future expanded/nightly jobs add larger PBT budgets, full turtles, all supported
target builds, raster goldens, and native E2E. Future scheduled stress rotates
and records seeds for long command sequences, resource churn, repeated window
lifecycle, large text/layout fixtures, and renderer recovery.

The contracts workflow runs the Python contract-validator unit tests,
document/ledger validation, `moon fmt --check`, and warning-denied all-target
MoonBit checks and tests. Independent mutation workflows run the selected
primitives scope and the capability/MCP scope matrix, retaining JSON and
survivor diffs; the semantic baseline stability merge gate is recorded above.
The Ubuntu native workflow also retains an opt-in recovery-to-first-frame timing report. These
jobs do not cover expanded PBT, renderer goldens, broad stress workloads, or a
comparable Tier 1 performance baseline and cannot satisfy production release
gates. See [`contracts.yml`](../.github/workflows/contracts.yml),
[`mutation-primitives.yml`](../.github/workflows/mutation-primitives.yml),
[`mutation-capability-mcp.yml`](../.github/workflows/mutation-capability-mcp.yml),
and [`ubuntu-native.yml`](../.github/workflows/ubuntu-native.yml) for exact commands.


## Event/focus and clip invariant implementation update — 2026-10-03

The active M2 property surface now includes stable seeded suites for pointer-route
reversal, capture stop-propagation, and focus validity after immutable subtree
removal. `ElementTree::without_subtree` removes the complete contiguous subtree
and clears focus only when the focused node is removed.

The active M3 property surface now validates strict LIFO clip nesting. A pop with
no open clip, a mismatched clip ID, or a frame that ends with an open clip is a
typed `SceneError`; `command_snapshot` refuses to freeze such a scene. These
checks strengthen the provisional command model only and do not consume the
reserved `SceneSnapshot` v1 schema.


## Recursive layout and SceneSnapshot v1 subset update — 2026-10-03

The recursive layout suite adds deterministic nested row/column cases, preorder
validation, missing-container rejection, and a fixed-seed 256-case property that
checks deterministic absolute geometry and descendant ordering. The versioned
scene suite checks schema version 1 metadata, owned resource/clip/item arrays,
clip-reference validation, finite opacity/transform inputs, canonical JSON, and
reuse of identical active rectangle clip chains.

PR CI runs `moon fmt --check`, `moon check --target all --deny-warn`, and
`moon test --target all --deny-warn`. At this slice the MoonBit suite passes
62/62 tests on wasm, wasm-gc, js, and native. This is headless target evidence,
not native GUI/platform evidence.


## Bounded field history and hosted compositor observation

The bounded field history tests cover original before/after selections,
independent entry/payload caps, whole-group eviction, failed restore/clipboard
transactions, Busy/undo rollback and immutable branching. Fixed-seed reference
coverage uses three 256-operation ASCII histories with documents of at most
12 characters; it is not general Unicode segmentation or arbitrary-layout
coverage. Native real-sans tests separately cover mixed fallback and bearings.
Undo/redo GPU fixtures replay actual-control scenes, while actual MoonBit
Host.present tests exercise serialization; neither injects compositor keyboard
input. PR32's [PR run37402479619](https://github.com/gpui-mbt/gpui.mbt/actions/runs/37402479619)
and [merged-main run37403927714](https://github.com/gpui-mbt/gpui.mbt/actions/runs/37403927714)
passed on first attempts at scales 1/2: 16/16 field `Host.present` tests per
scale and 11 accepted GPU scenes from 13 fixtures. The other two fixtures are
headless rejected-edit identity checks. Reviewed undo/redo decoded pixels
match between the PR and exact merged main36bcb245/tree5441e254.

The merged origin baseline's
[main Ubuntu run37393518087](https://github.com/gpui-mbt/gpui.mbt/actions/runs/37393518087)
recorded compositor exit139/reset in attempt1 and passed an unchanged-source
failed-job retry in attempt2. First-failure and successful artifacts remain
separate. Cause is unknown; a successful retry is not proof of production
stability or live keyboard/IME qualification. See the
[Ubuntu stability/recovery limits](ubuntu.md#known-hosted-compositor-observation).

## Bounded direct keyboard repeat evidence

Private Ubuntu direct-repeat checks are deterministic headless evidence.
Mocked Wayland callbacks and controlled clock/queue observations cover the
v4 `repeat_info` policy, lower-version/rate-zero/missing-policy suppression,
repeatable key/text caching, empty-queue dispatch admission, one atomic group
per dispatch, no overdue catch-up, 1 ms rate saturation, Compose isolation and
focus/epoch/mode/modifier/layout/device/close/fatal cancellation. They are not
compositor-delivered typing or a measurement of a desktop repeat timer.

MoonBit ABI2 tests separately check physical/synthetic press flags and typed
`InvalidInput` for NaN, infinity, fractional, negative or out-of-range flags,
including repeat flags on release/text. Pointer slot 7 remains a y coordinate.
Portable field tests and native real-font controller tests inject repeated
`KeyPressed` and committed `TextInput` records. They check insertion exactly
once, separate content history groups, navigation without history, current
repeated undo/redo and clipboard/submit semantics, stale paste rejection after
repeated editing, and pending/rollback history consistency.

The existing Weston/llvmpipe rendering tier and replayed `Host.present` scenes
do not exercise `wl_keyboard` delivery or qualify repeat timing. Separately,
the frozen integrated source eb6c164f/tree d1b53388 passed isolated real
basic/Shift and Ctrl+A input, plus held-repeat release/focus-loss/refocus
through authenticated private Xvfb and Weston 14. The held gates disabled
upstream X11 autorepeat, read actual 40 Hz/400 ms policy, observed one real
physical press, and checked growth and post-release/leave stability with
completed frames and retained pixels. The earlier startup click-delivery
failure remains retained; R2 changed readiness ordering without relaxing
acceptance predicates. See the [qualification manifest](native-input-qualification.json)
and [field evidence tiers](linux-text-field.md#evidence-and-remaining-gates).
The later [packaged R4 qualification](native-input-qualification-r4.json)
records a fresh locked-prefix build of frozen `eb6c164f` and seven runnable
native keyboard cases with exact case/driver/binary identities. Held-repeat
and GPUI IME are two explicit skips; `all_cases_executed` remains false. The
unchanged basic golden passed after the official missing font was restored;
the earlier mismatch artifacts and pre-native snapshot stay preserved. A later
publication assembly is built and checked separately, and inherits this native
evidence only as byte-equivalent application source, not as the exact executed
commit/binary.

Those passes are exact-source/profile evidence, not general desktop timing
accuracy or a pass for every catalog case. GPUI IME remains unimplemented;
GTK Japanese conversion is environment-baseline evidence only. This bounded
opt-in route adds no public `TextInput`/IME capability, typing coalescing, or
one-shot shortcut filtering. See the
[field contract and remaining gates](linux-text-field.md#evidence-and-remaining-gates).

The earlier repeat-only Linux verification on 2026-10-06 passed native
380/380, JavaScript/Wasm/Wasm-GC 318/318 each, Python 101/101, full MoonBit
formatting and the contract-document checker. The C direct-input/repeat and
clipboard fixtures pass with `-Wall -Wextra -Werror`; repeat also passes
ASan+UBSan with leak detection disabled. Headless Pango raster, field-admission
and origin-raster consumers pass. These results use MoonBit
`0.10.14+7d59c7ec9`, Wayland 1.23.1, XKBCommon 1.7.0, PangoFT2 1.56.3 and
Fontconfig 2.15.0. Generated native MoonBit C retains existing compiler warnings;
MoonBit `--deny-warn` and the handwritten C strict-warning checks pass.
The native `GPUI_FIELD_E2E` compositor gate was skipped, not passed.

Final integrated source verification at eb6c164f0277a656125804154a3d8ad9a8abb78d
passed 402/402 aggregate native and 100/100 focused native tests, 325/325
JavaScript/Wasm/Wasm-GC tests each, 101/101 Python tests, formatting, contracts,
strict handwritten C, and ASan+UBSan with leak detection disabled. The native
aggregate still skipped `GPUI_FIELD_E2E`; the separately retained native input
gates do not turn that skip into a pass. Generated native MoonBit C retains
existing warnings. The earlier per-feature and pre-native evidence remains
unchanged; new source/runtime iterations need their own build and provenance.
