piker/plans/opencode/human-facing-e2e-coverage.md

58 KiB
Raw Permalink Blame History

Human-facing end-to-end coverage plan

Purpose

Build durable coverage for Piker as a user encounters it through installed console scripts, supported Python APIs, and the Qt chart. The highest-value result is a small set of true black-box system journeys backed by a deterministic, protocol-faithful offline provider. Real QtBot input journeys must cover the same critical chart behavior in process because QtBot cannot drive widgets owned by another process.

This plan incorporates the chart-specific findings in plans/opencode/pytest-qt-chart-ui-e2e.md and broadens them to all human-facing interfaces. It is based on source and tests at 09ddcf50 on wkt/human_iface_e2e, audited on 2026-08-30.

The first implementation patch is intentionally only a harness-foundation patch. It must not claim application or system E2E coverage, add a replay backend, rewrite the chart runtime, or start a broad widget audit.

Coverage vocabulary

The suite must use these labels consistently. A test does not become E2E just because it is expensive, asynchronous, uses a real Qt object, or starts a child actor.

Tier Name Required boundary Examples E2E credit
T0 Contract/unit One function, schema, or pure state transition FeedInit validation, gap-spec conversion, accounting math No
T1 Component integration Real component and its immediate framework, with narrow collaborators substituted Real PlotWidget, ViewBox, scene index, or CompleterView No
T2 Application integration Production window/widget composition, Trio guest run, and protocol-faithful offline services in one pytest process QtBot searches for a symbol and submits a paper order No; call it an application journey
T3 System E2E Fresh process through an installed console script or a fresh Python interpreter using a public API, with real process, Tractor, IPC, and SHM boundaries piker chart ..., pikerd, or a subprocess importing piker.open_feed Yes
T4 Qualification T3 under a real compositor or against an explicitly selected live provider Wayland/X11 focus, DPI, or read-only live feed qualification Yes, but opt-in and non-default

Two complementary tests are required for a critical Qt journey:

  1. A T2 application journey uses real QtBot key and mouse APIs against the production widget tree and proves user-visible state.
  2. A T3 system journey launches the installed command and proves process, service, IPC, SHM, startup, failure, and shutdown behavior.

Neither test may be described as proving the other boundary. A single true black-box UI journey that also drives native window chrome belongs in T4 and requires real compositor/accessibility tooling, not pytest-qt.

Non-negotiable test policy

  1. Reserve the terms E2E and system for T3 and T4.
  2. Drive claimed user interactions with qtbot.keyClick(), qtbot.keyPress(), qtbot.keyRelease(), qtbot.keyClicks(), qtbot.mouseMove(), and qtbot.mouseClick() against the displayed production widget or graphics viewport.
  3. Do not count direct handler calls, direct model mutation, direct command callbacks, CliRunner, manually constructed QKeyEvent or QMouseEvent instances, or QApplication.sendEvent() as user interaction. Such calls remain valid only in clearly named T0 or T1 tests.
  4. Mocked user interactions are banned in T2 through T4. If QtBot cannot express an interaction, either move that assertion to T1 without journey credit or add a T4 compositor driver. Do not disguise direct dispatch as a user event.
  5. Positive-duration time.sleep(), trio.sleep(), asyncio.sleep(), qtbot.wait(), and shell sleep are banned as readiness or ordering mechanisms in all new harness, application, and E2E tests.
  6. Synchronize on Qt signals, visible postconditions, typed Trio events, Tractor context startup messages, replay control acknowledgements, process exit, or bounded protocol probes. A timeout is a failure bound, not a pacing mechanism.
  7. trio.sleep(0) is allowed only as an explicit scheduler checkpoint. It must not be the evidence that startup, rendering, cancellation, or cleanup completed.
  8. A deterministic offline service implementation is allowed and preferred when it implements the same provider protocol, runs through the same service/actor boundary, and emits production message types. It is not a mock merely because its data is fixed.
  9. T3 tests must not monkeypatch names imported by the program under test. Configuration, a replay scenario, and a typed replay control protocol are valid inputs. Replacing open_feed, open_ems, a stream, or a widget with a spy is not.
  10. Default gates are credential-free, network-free, and write only beneath test-owned temporary roots.

Baseline evidence

Declared interfaces

pyproject.toml:183-187 installs exactly three scripts:

  • piker, the Click command tree;
  • pikerd, the standalone root service actor;
  • ledger, the Typer accounting command.

The piker command dynamically registers broker, UI, watchlist, storage, and accounting command modules at import time (piker/cli/__init__.py:367-387). Its human-facing surface includes services, broker commands, watchlists, store, and chart. The README promises both implicit daemon startup from piker chart and attachment to a separately started pikerd (README.rst:129-191). Neither journey currently has a system test.

The top-level Python package explicitly exports only open_piker_runtime and open_feed (piker/__init__.py:21-27). Additional supported subsystem APIs include piker.data.open_feed, piker.service.open_pikerd, piker.service.maybe_open_pikerd, piker.clearing.open_ems, and accounting context managers. Existing tests call several of these directly, but none executes the top-level API contract from a fresh installed-package process with an offline provider.

Current suite topology

A static scan finds 152 test functions across 20 test modules. Parametrization changes collected case counts, and two EMS functions contain only ellipsis, so the source count is not a coverage claim.

Area Current modules Current boundary and gap
Storage and time series test_storage_audit.py, test_storage_nativedb.py, test_store_cli.py, test_ldshm.py, test_tsp_analysis.py, test_history_backfill.py, test_backfill_audit_snippet.py Strong deterministic function and component coverage. CliRunner calls in storage tests bypass the installed piker script and are not E2E.
Broker adapters test_ib_history.py, test_ib_method_proxy.py, test_questrade.py IB tests use focused fake clients. Questrade is credentialed, obsolete, and module-skipped. No provider-neutral offline conformance suite exists.
Runtime, feeds, and clearing test_services.py, test_feeds.py, test_ems.py, test_shm_cleanup.py Real Tractor and SHM coverage exists, but meaningful feed and paper-EMS paths still reach live Kraken/Binance data. Two named EMS tests have empty bodies at tests/test_ems.py:397-412.
Accounting and watchlists test_accounting.py, test_watchlists.py Mostly direct Python APIs and local files. Some accounting cases can use configured state or write tracked fixtures.
CLI test_cli.py, plus direct storage command tests All 11 legacy subprocess CLI tests are disabled by one unconditional module mark at tests/test_cli.py:13-16, including offline watchlist cases. There is no active installed-script gate.
Qt test_dpi_font.py, test_gap_overlays.py DPI uses a MockScreen. Gap coverage has real Qt/PyQtGraph objects, five real-Qt component cases, one direct key-event relay case, three schema/state cases, and one Tractor endpoint case. There is no composed MainWindow, GodWidget, application journey, or chart process E2E.
Containers test_docker_services.py Both tests are skipped, but optional imports occur before the skip marker and may fail collection. This is qualification coverage, not a default E2E prerequisite.

There is no tests/ui/, tests/app/, tests/e2e/, or tests/qualification/ topology today. Protocol tests, real-Qt component tests, actor integration, and live tests are mixed in top-level modules.

Existing evidence worth preserving

The chart-local gap work is a useful component-integration baseline:

  • Real PlotWidget, PlotItem, ViewBox, and QGraphicsScene ownership assertions catch stale PyQtGraph registries (tests/test_gap_overlays.py:293-409).
  • Scene spatial-index assertions catch geometry updates performed in the wrong order (tests/test_gap_overlays.py:412-482).
  • Duplicate FQMEs prove that chart-local identity cannot be reduced to an actor-global FQME key (tests/test_gap_overlays.py:485-540).
  • Focused-chart behavior is exercised with two independent chart pairs (tests/test_gap_overlays.py:543-648).
  • The current Ctrl-G case traverses EventRelay, but manually constructs and sends a QKeyEvent (tests/test_gap_overlays.py:651-735). It is T1, not a QtBot user journey.
  • A real child actor proves typed request correlation and endpoint cleanup (tests/test_gap_overlays.py:817-1025). Rendering is replaced in that actor, so it is protocol integration, not chart E2E.

The root autouse SHM tracker records exact segments created by the current pytest process and restores Tractors token cache (tests/conftest.py:97-239). tests/test_shm_cleanup.py:18-109 proves that it does not unlink an attached segment owned by somebody else. This exact-owner discipline must be extended to subprocess and child-actor SHM rather than replaced by /dev/shm scans or broad cleanup.

Baseline gaps

  1. There is no active installed-script test for any of the three declared scripts.
  2. There is no deterministic provider that can run real datad, history, sampling, feed, paper EMS, search, and chart paths without network access.
  3. There is no nonblocking production seam for pytest-qt to own the Qt loop while Piker owns Trio guest mode.
  4. There is no observable guest-run completion or error result.
  5. There is no safe way for pytest-qt teardown to close a production MainWindow.
  6. There is no stable selector/accessibility policy for finding human-facing Qt controls.
  7. There is no child-process ownership ledger covering PIDs, actors, registry sockets, streams, and child-created SHM.
  8. Existing live actor tests use timing sleeps and network availability as implicit readiness.
  9. The tracked CI does not provision the current project. It uses Python 3.10, setup.py, missing requirements files, and an uncontrolled full suite (.github/workflows/ci.yml:14-63).

Risk-ranked user journey matrix

Risk combines impact, concurrency/lifecycle complexity, and present coverage. P0 is the first coverage target, not an assertion that lower-ranked behavior is unimportant.

Risk Surface Human journey and observable contract Current evidence Target evidence
P0 piker chart From a clean config, start one replay FQME, see history and realtime chart state become ready, then close without survivors README promise only; component chart fragments Paired T2 QtBot application journey and T3 installed-script process journey
P0 Qt search Press Ctrl-L, type a query, select B, switch A to B to A, and retain the right visible chart/focus/cache state Direct search and load methods are untested; known ownership race T2 QtBot journey using real search model and replay search service
P0 Qt overlapping startup Select B while uncached A is blocked, release A in both completion and cancellation orders, and keep each display attached to its own feed/widgets No coverage; shared mutable GodWidget fields make the race credible T2 typed-barrier application journey plus focused ownership tests
P0 Qt paper order Stage with keyboard, submit with a viewport mouse click, receive open/fill/cancel state, and see position/status/line transitions EMS API tests bypass UI; some paths are live T2 QtBot journey with real paper EMS and replay quotes, paired with T3 open_ems probe
P0 Runtime failure Provider, Trio guest task, or child actor fails; the window/process terminates, error is observable, and exit is nonzero Guest callback prints and quits T2 outcome test and T3 subprocess failure propagation test
P0 Public Python API A fresh interpreter opens piker.open_piker_runtime and piker.open_feed, receives deterministic history/quote state, pauses/resumes, and exits cleanly Direct live/in-process uses only T3 subprocess probe against replay provider
P0 pikerd plus client Start standalone daemon, observe protocol readiness, attach and detach a chart/API client, keep daemon alive, then stop it cleanly In-process fixture boot only T3 installed pikerd and piker process journey
P1 Script discovery piker, pikerd, and ledger help, invalid arguments, and documented subcommand discovery have stable exit/output semantics test_cli.py is skipped; store uses direct Typer runner T3 subprocess contract tests for all installed scripts
P1 piker watchlists Create, show, merge, and remove entries under an isolated config through the installed command Direct functions and skipped subprocess cases T3 subprocess filesystem journey; retain T0 function tests
P1 piker store Discover commands and perform read-only series, ls, audit, and SHM diagnostics without mutation Good CliRunner component coverage Selected T3 installed-script read-only journeys; destructive commands stay excluded
P1 Feed/service control Open multiple replay feeds, verify real datad/samplerd/SHM, pause hidden feeds, resume focused feeds, and close all streams Live feed tests and actor fixtures T3 public API probe and T2 chart application journey
P1 Clearing API Open paper EMS, submit/cancel/fill orders, persist isolated ledger/position state, and reconnect Partial real EMS coverage; live symbology; two empty tests T3 replay-backed open_ems probe with exact config assertions
P1 Qt navigation Focus chart/search, zoom, pan, reset, toggle gaps, and return with Escape using user events One manually sent key event; direct helpers T2 QtBot journeys, with geometry details kept in T1
P2 Python accounting Open ledger/account contexts, calculate positions, close/reopen under a temporary root Direct tests, some configured or tracked-state risk T3 fresh-interpreter smoke for public contexts; detailed math stays T0
P2 Graphics composition Linked realtime/history plots, cursor, labels, overlays, region movement, and SHM update cycle remain coherent Gap items only T1 real PyQtGraph composition suite, not mislabeled E2E
P2 Real compositor Window activation, native focus, DPI, font selection, titlebar close, and screenshot diagnostics work on Wayland and X11 Mock screen only Opt-in T4 compositor matrix
P3 Live provider Read-only symbol search, history, and quotes conform for a selected backend Ad hoc live tests with mixed CI rules Protected, opt-in T4 provider qualification
P3 Live account Account discovery and explicitly non-transmitting order validation where a backend supports it Configuration-dependent tests Manual protected qualification only; never a default gate

Stable contract policy

E2E tests are expensive to maintain when they freeze implementation details. They must preferentially assert interfaces expected to survive medium-term refactors.

Surface Stable enough for T2/T3 assertions Volatile; keep out of E2E assertions
Console Installed script names, documented option meaning, exit status, stdout/stderr role, JSON field schema, and noninteractive behavior Click/Typer callback objects, full help whitespace/color rendering, import order, log wording
Public Python Names exported from package APIs, async context lifetime, yielded production types, stream semantics, cancellation/error type Private attributes, helper call counts, module-local caches, exact task layout
Provider protocol MktPair, FeedInit, OHLCV dtype, quote/tick schema, context startup payload, pause/resume semantics, typed replay controls Backend helper classes, logger output, actor PID/UUID, exact scheduling order
Qt Production objectName/accessible names added as selectors, enabled/visible/focus state, semantic model roles/text, window title fields, scene attachment Child index traversal, private widget dictionaries, pixel coordinates unrelated to mapped data, exact size, style sheet text, render timing
Chart data FQME, timeframe, bar/tick identity, selected symbol, visible semantic labels, exact replay sequence number Generated SHM names, array allocation size, FPS, downsampling internals
Lifecycle Typed ready/completed outcomes, bounded shutdown, no owned survivors, expected cancellation reason Incidental exception-group nesting, log ordering, port number, cleanup callback order not exposed by contract
Visual Widget visibility, geometry relationships, scene hit testing, failure screenshots Pixel goldens, font rasterization, antialiasing, colors unless they convey required state

Policy consequences:

  1. Add a minimal production objectName or accessible name when a user control has no stable selector. Do not teach E2E tests to traverse private layouts by position.
  2. Give runtime readiness and completion typed handles. Do not parse logs to infer that a chart, actor, feed, or guest task is ready.
  3. Keep white-box registry and geometry assertions in T1 because they are valuable diagnostics, but pair them with a stable visible-state assertion when a journey depends on the behavior.
  4. A stable contract change requires an intentional test update and a short compatibility note. A private refactor should not require T2/T3 changes.
  5. Do not expose a test-only alternate application implementation. Test seams must be production lifecycle interfaces used by the blocking CLI adapter.

Harness architecture

Proposed layout

tests/
  conftest.py
  replay/
    conftest.py
    test_provider.py
  ui/
    conftest.py
    test_harness.py
    test_widgets.py
    test_graphics_items.py
    test_chart_interaction.py
    test_chart_composition.py
  app/
    conftest.py
    test_lifecycle.py
    test_symbol_journeys.py
    test_order_journeys.py
  e2e/
    conftest.py
    probes/
      open_feed.py
      open_ems.py
    test_console_scripts.py
    test_public_api.py
    test_chart_process.py
    test_failure_paths.py
  qualification/
    test_compositor.py
    test_live_provider.py
  _inputs/
    replay/
      basic-v1.json
      overlap-v1.json
      failure-v1.json

Existing fast tests stay in their current focused modules unless moving one is necessary to separate live collection or clarify a Qt tier. A directory move alone provides no coverage and should not dominate a patch.

Root ownership fixtures

tests/conftest.py remains the single owner of process-global test state:

  • set PYTEST_QT_API=pyqt6, choose the default QPA, and create an isolated XDG_CONFIG_HOME before importing Piker or Qt;
  • snapshot and restore Piker config globals, imported CLI path caches, Tractor runtime variables, and test-owned environment keys;
  • allocate unique actor names and registry addresses;
  • retain exact current-process SHM creation tracking;
  • expose an ownership ledger to child-process and Qt fixtures;
  • aggregate cleanup errors rather than hiding them behind the primary test failure.

The fixture must not scan for arbitrary tractor processes, kill by name, or unlink unknown SHM. Every cleanup action needs an identity recorded by the current test.

Qt component harness

tests/ui/conftest.py provides non-autouse typed factories around pytest-qt:

  • use pytest-qts qapp; never define another QApplication fixture;
  • register a widget with qtbot.addWidget() immediately after construction;
  • show and focus interaction targets, then use waitExposed(), waitActive(), waitSignal(), or waitUntil();
  • map data/scene coordinates through the real ChartPlotWidget viewport for mouse input;
  • create deterministic OHLCV arrays and, only where production drawing requires it, real disposable ShmArray instances under a local Tractor root runtime;
  • compare top-level widgets, event filters, signals, PyQtGraph registries, QSettings, and globals to a pre-test snapshot.

Small dataclasses and protocols should replace new SimpleNamespace graphs. The existing namespaces do not need a cosmetic rewrite in the foundation patch.

Application harness

tests/app/conftest.py composes the production application without calling the blocking QApplication.exec() adapter. It owns:

  • one QtractorSession returned by the production startup seam;
  • the real MainWindow, GodWidget, search widget, chart hierarchy, Trio guest run, and replay-backed services;
  • typed handles for readiness, shutdown request, and guest outcome;
  • QtBot input helpers that operate only on real displayed widgets/viewports;
  • a replay controller that advances provider state through typed IPC;
  • final assertions that the application, guest run, actors, SHM, and Qt state all returned to their baseline.

This is application integration, not system E2E, because pytest and the application share a process and pytest receives object handles.

Process harness

tests/e2e/conftest.py launches installed scripts and fresh Python probes with subprocess.Popen(..., start_new_session=True). It records:

  • executable path and resolved worktree package;
  • PID, process start identity, process group, and attributable descendants;
  • command, environment, temporary config root, registry address, and replay scenario;
  • stdout and stderr with bounded readers;
  • typed service/replay readiness and owned SHM identities;
  • graceful termination request, exit status, and whether escalation was required.

Readiness must come from a protocol response or a child-owned readiness FD, never a log substring. On teardown, request normal shutdown, wait for the bounded process result, and inspect only recorded descendants. A forced signal may be used to contain a failed test, but forced cleanup itself fails the test. Broad pkill, process-name matching, and global port/SHM cleanup are banned.

CliRunner remains useful for T1 command callback tests. T3 always resolves and invokes the installed executable or starts a fresh interpreter.

Assertion and artifact model

Every journey asserts one semantic user outcome and one ownership outcome. Examples are selected FQME plus no display-task leak, filled position plus no open EMS dialog, or nonzero process exit plus no actor/SHM survivor.

On failure retain:

  • stdout/stderr and pytest output;
  • replay input and an ordered typed replay transcript;
  • actor/service identity snapshot;
  • exact SHM ownership manifest;
  • Qt warnings and virtual-method exceptions;
  • qtbot.screenshot() for displayed widgets.

Screenshots are diagnostics. They are not initial golden assertions.

Deterministic replay provider

Role

Add an explicitly selected piker.brokers.replay provider. It is a real offline backend implementation, not a monkeypatch layer. piker -b replay chart alpha.replay and public open_feed() calls must traverse normal provider loading, datad spawn, Tractor contexts, history management, SHM, sampling, search, and paper EMS code.

Do not add replay to the default live broker list initially. Tests and users select it explicitly, avoiding any change to normal provider discovery or startup cost.

Protocol surface

The provider implements current production endpoint shapes:

  • get_mkt_info() returns a real MktPair;
  • open_history_client() returns deterministic OHLCV chunks and explicit history completion boundaries;
  • stream_quotes() starts with a real FeedInit, emits production quote and tick dictionaries, and honors feed-live startup;
  • open_symbol_search() is a real Tractor context and stream;
  • paper trading uses Pikers real paper engine and feed; the replay provider does not synthesize EMS status messages in the UI test;
  • a separate replay-control context accepts typed commands and returns typed acknowledgements.

If provider endpoint typing is formalized later, replay is the first conformance implementation. Do not invent a parallel test-only feed API.

Scenario format

Versioned JSON fixtures under tests/_inputs/replay/ contain only portable data and behavior declarations:

  • scenario version and deterministic seed;
  • markets and complete MktPair inputs;
  • sorted history chunks for 1-second and 60-second paths;
  • quote/tick events with stable sequence IDs and timestamps;
  • search query/result mappings;
  • optional startup barriers and named failure points;
  • expected terminal sequence, not expected widget internals.

The provider validates a scenario before starting any actor or SHM. Invalid fixtures fail with a typed configuration error.

Control and synchronization

The control protocol supports commands such as AwaitState, Advance, ReleaseBarrier, FailAt, and Snapshot. Exact message names may follow project msgspec conventions, but each request and response carries a scenario ID and monotonic sequence ID.

The provider emits no event based on elapsed wall-clock time. A quote advances only after a typed command or after an explicitly declared startup event. Pause/resume tests assert subscriber state through the control snapshot before advancing another quote. Failure tests inject at a protocol boundary and wait for the corresponding acknowledgement before asserting propagation.

This design permits deterministic offline behavior while preserving real serialization, IPC, actor scheduling, SHM, feed, and EMS boundaries.

Synchronization rules

  1. Process environment and Qt binding are fixed before importing Piker, Qt, or PyQtGraph.
  2. pytest-qt owns the one QApplication in T1 and T2.
  3. Every test-owned widget is registered before it can raise or return.
  4. User input follows show, exposure, activation, and focus assertions.
  5. QtBot targets the actual input widget or graphics viewport, not a backing ViewBox that is not a QWidget.
  6. A wait names a postcondition and has a bounded failure timeout.
  7. Tractor startup uses ctx.started(), task-status startup, or a typed service probe.
  8. Replay ordering uses protocol sequence IDs and barriers.
  9. Process readiness uses a protocol/FD handshake; logs remain diagnostics.
  10. No test infers success from one event-loop turn, one quote delay, or a process remaining alive for a guessed interval.
  11. Cancellation tests first prove the target operation reached a named barrier, then request cancellation, then await its typed terminal state.
  12. Geometry tests retain both internal bounds and black-box scene hit tests, but only the latter contributes to visible behavior evidence.

Cleanup invariants

Cleanup is part of every acceptance gate, including failure and cancellation paths. Cleanup errors fail the test even when the primary assertion passed.

Required unwind order

  1. Stop issuing QtBot and replay commands.
  2. Request application shutdown through the production session interface.
  3. Await and classify the Trio guest outcome.
  4. Close EMS/feed/search/replay streams and Tractor contexts.
  5. Cancel and await service tasks and child actors through their recorded portals/nurseries.
  6. Release chart, Viz, graphics-item, and SHM attachments.
  7. Unlink only exact SHM segments created by the test-owned actor tree and restore token caches.
  8. Disconnect Qt signals, remove event filters, detach PyQtGraph items, close widgets, schedule deletion, and wait for the top-level widget baseline.
  9. Await subprocess exit and pipe-reader completion; verify recorded descendants and sockets are gone.
  10. Restore config, QSettings, environment, runtime globals, and temporary roots.

Process and subprocess invariants

  • Every child belongs to a recorded process session and has a start identity, so PID reuse cannot authorize cleanup.
  • Normal shutdown gets one bounded opportunity. Escalation is exact-PID or exact-process-group containment and marks the test failed.
  • stdout/stderr readers always terminate and file descriptors close.
  • No attributable descendant remains after the parent exits.
  • No production registry address is used. Parallel tests receive distinct addresses and actor names.
  • No test reaps an unrelated local pikerd, datad, brokerd, or chart.

Trio and Tractor invariants

  • Every nursery/context entered by the harness exits before fixture return.
  • Guest completion is one of explicit normal completion, requested cancellation, or error. A missing outcome is a leak/failure.
  • Unexpected Error outcomes and remote actor errors fail the owning test; printing them is insufficient.
  • Context streams close in both directions and no replay request remains unacknowledged.
  • Actor registries contain none of the tests recorded actor IDs after teardown.
  • Repeated same-process application sessions do not accumulate enabled modules, runtime kwargs, callbacks, or service tasks.

SHM invariants

  • Record creator identity and all data/index segment names at allocation.
  • Attachments are closed but never unlinked by non-owners.
  • Creators unlink exact surviving segments and remove exact token entries.
  • Child-created SHM is reported through the ownership manifest before actor teardown, so the parent can verify absence without scanning globally.
  • A leak is cleaned only for containment and still fails the test.

Qt and PyQtGraph invariants

  • QApplication.topLevelWidgets() returns to the pre-test baseline.
  • _window._qt_win, GodWidget globals/caches, search globals, remote-control globals, fonts, and QSettings are restored.
  • The global zoom filter and all local EventRelay filters are removed.
  • Every harness-owned signal connection is disconnected.
  • PlotItem.items, ViewBox.addedItems, overlay registries, and scene membership contain no detached test item.
  • Widgets are closed only after guest tasks stop using them. deleteLater() completion is observed through a bounded Qt postcondition, not arbitrary processEvents() loops.
  • Qt warnings, virtual-method exceptions, and callbacks into closed channels fail the test.

Known blockers and defects

Item Evidence and impact Required disposition
pytest.ini shadowing The comments-only pytest.ini:1-3 wins config discovery, so testpaths and -p no:xonsh in pyproject.toml:169-180 are inactive. Delete pytest.ini and keep one active pytest configuration source. Assert the selected config in the foundation gate.
Missing pytest-qt pyproject.toml:150-152 and uv.lock include pytest 9.0.2 but no pytest-qt. Add pytest-qt to testing, resolve compatibility through uv.lock, and set qt_api = "pyqt6".
Config leakage Session confdir does nothing without --confdir (tests/conftest.py:53-67); tmpconfdir assigns config._config_dir without restoration (tests/conftest.py:299-327). Config is initialized at import (piker/config.py:125-156), UI fonts touch conf.toml at import (piker/ui/_style.py:54-77, piker/ui/_style.py:236-238), and several CLI modules cache app-dir paths at import. Establish process-start XDG isolation, snapshot/restore globals, and make subprocesses inherit only test-owned paths. Add a no-user-config-write regression.
MainWindow SIGINT close MainWindow.closeEvent() writes QSettings and sends SIGINT to its own process (piker/ui/_window.py:331-348). qtbot.addWidget() teardown could interrupt pytest. Separate close request from process signalling. The CLI adapter owns any signal policy; T2 requests and awaits session cancellation without signalling pytest.
Guest outcome is lost run_qtractor() prints non-keyboard errors and calls app.quit() (piker/ui/_exec.py:138-150), returns no handle, and blocks in app.exec_() (piker/ui/_exec.py:193-211). Add a nonblocking production session with observable typed completion. The blocking CLI wrapper must convert guest error to a nonzero process result.
Symbol-switch ownership race GodWidget.load_symbols() assigns shared self.rt_linked/self.hist_linked, then starts display_symbol_data() (piker/ui/_widget.py:195-217). The async task later rereads those mutable fields (piker/ui/_display.py:1409-1416). Overlapping uncached loads can bind As task to Bs widgets; search then rewrites cache state from current globals (piker/ui/_search.py:745-789). Give each load an explicit owned chart pair/session and pass it into display startup. Publish/focus only the current generation. Test completion and cancellation orderings with typed barriers.
Skipped CLI One unconditional module mark skips all 11 tests (tests/test_cli.py:13-16), even local watchlist subprocess cases. Active storage CLI tests invoke the Typer object directly (tests/test_store_cli.py:26-180). Split deterministic installed-script journeys from live legacy quote/API cases. Remove blanket skipping; retain direct callback tests as T1.
Signal/filter teardown gaps open_signal_handler() connects a proxy but has no matching disconnect (piker/ui/_event.py:195-219). run_qtractor() connects focus change and installs a global zoom filter without an owned teardown path (piker/ui/_exec.py:174-207). Make each connection/filter session-owned and prove repeated sessions do not accumulate callbacks.
Mutable runtime defaults run_qtractor() mutates a default dict (piker/ui/_exec.py:86-93, piker/ui/_exec.py:180-183); open_piker_runtime() has mutable list defaults (piker/service/_actor_runtime.py:58-65), and open_pikerd() extends enabled modules (piker/service/_actor_runtime.py:189-197). Remove mutable defaults when the repeated-session lifecycle test is introduced. Assert the second session has the same module/runtime state as the first.
Timing-based actor tests Positive sleeps remain in live/provider and actor tests, including delayed gap correlation and service readiness. Do not copy these patterns. Replace sleeps in migrated journeys with typed receipts, events, or service probes. Existing focused timing tests can be handled separately when touched.
CI cannot enforce the plan Current workflow targets obsolete Python/install inputs and runs an uncontrolled suite. Add current frozen-uv headless jobs only after the foundation command is proven locally; keep live/compositor jobs opt-in.

Phased implementation

Each phase is independently reviewable. Production fixes stay with the first test that demonstrates the defect. Later phases do not expand the first patch.

Phase 1: pytest, Qt, config, and leak foundation

This is the entire first implementation patch.

Exact targets:

  • pyproject.toml
  • uv.lock
  • pytest.ini (remove it)
  • tests/conftest.py
  • tests/ui/conftest.py
  • tests/ui/test_harness.py
  • tests/test_gap_overlays.py
  • .claude/skills/run-tests/test-harness-reference.md

Planned tests:

  • tests/ui/test_harness.py::test_qt_process_state_is_isolated
  • tests/ui/test_harness.py::test_qapplication_is_reused_without_widget_leaks
  • the five existing real-Qt gap nodes at tests/test_gap_overlays.py:293-735
  • tests/test_gap_overlays.py::test_remote_gap_dialog_real_actor

Implementation:

  1. Add pytest-qt beside pytest and regenerate the frozen lock against pytest 9.0.2.
  2. Delete the shadowing pytest.ini; set qt_api = "pyqt6" in the one active config and retain -p no:xonsh.
  3. Establish PyQt6/QPA/XDG isolation before Piker or Qt imports, with an explicit override path for later compositor tests.
  4. Make config and QSettings restoration exact and function-scoped while the process-level XDG root remains session-owned.
  5. Add typed OHLCV/PlotWidget helpers and top-level-widget/PyQtGraph leak checks without creating a parallel qapp.
  6. Migrate the module-local gap qapp fixture to qtbot. Register every widget and replace the manually sent Ctrl-G key event with a real QtBot key interaction after show/focus.
  7. Replace the positive delay in the actor correlation case with an explicit typed barrier if that node is retained in the same patch.
  8. Update only the test-harness reference facts made stale by this patch.

Acceptance gate:

  • pytest reports pyproject.toml as its config source and loads pytest-qt for PyQt6;
  • collection works without a separate -p no:xonsh argument;
  • the migrated gap tests preserve ownership and scene-index assertions;
  • at least two Qt tests run in one process using the same QApplication and leave the baseline unchanged after each test;
  • a test-owned config/QSettings sentinel is removed while a pre-existing non-test sentinel remains untouched;
  • no user config path, tracked fixture, actor, socket, or SHM is changed;
  • no arbitrary sleep, screenshot golden, MainWindow, GodWidget, replay provider, new widget audit, or system E2E claim is added.

Phase 2: protocol-faithful replay provider

Exact targets:

  • piker/brokers/replay.py
  • tests/_inputs/replay/basic-v1.json
  • tests/_inputs/replay/failure-v1.json
  • tests/replay/conftest.py
  • tests/replay/test_provider.py

Planned tests:

  • tests/replay/test_provider.py::test_replay_backend_satisfies_feed_contract
  • tests/replay/test_provider.py::test_replay_history_is_repeatable
  • tests/replay/test_provider.py::test_replay_control_advances_exact_sequence
  • tests/replay/test_provider.py::test_replay_pause_resume_has_no_hidden_ticks
  • tests/replay/test_provider.py::test_replay_failure_is_acknowledged_before_raise

Implementation:

  1. Implement the normal provider endpoint signatures and production message types in one importable backend module.
  2. Validate versioned scenarios before opening runtime resources.
  3. Add a typed Tractor control endpoint with monotonic sequence IDs and named barriers.
  4. Drive real open_feed(), datad, samplerd, history SHM, quote SHM, and pause/resume paths. Tests may inspect replay controls but may not replace feed internals.
  5. Keep real paper EMS support implicit through the normal feed interface; order behavior is covered in a later slice.

Acceptance gate:

  • the provider suite passes with outbound network denied;
  • two identical scenario runs produce identical protocol transcripts and history/tick values;
  • validate_backend() accepts the real FeedInit path without deprecated payload fallback;
  • pause/resume and failure ordering are proven by acknowledgements, with no positive-duration sleep or log parsing;
  • all actors, streams, and exact SHM names are gone after success and failure.

Phase 3: installed scripts and public API system harness

Exact targets:

  • tests/e2e/conftest.py
  • tests/e2e/probes/open_feed.py
  • tests/e2e/test_console_scripts.py
  • tests/e2e/test_public_api.py
  • tests/test_cli.py
  • piker/cli/__init__.py only for defects proven by these tests
  • piker/watchlists/cli.py only for defects proven by these tests
  • piker/storage/cli.py only for defects proven by these tests
  • piker/config.py only for config-contract defects proven by these tests

Planned tests:

  • tests/e2e/test_console_scripts.py::test_installed_scripts_expose_help
  • tests/e2e/test_console_scripts.py::test_piker_rejects_invalid_chart_symbol
  • tests/e2e/test_console_scripts.py::test_watchlist_round_trip_is_isolated
  • tests/e2e/test_console_scripts.py::test_store_read_only_commands_use_isolated_config
  • tests/e2e/test_console_scripts.py::test_standalone_pikerd_starts_and_stops
  • tests/e2e/test_public_api.py::test_open_runtime_and_feed_from_fresh_python
  • tests/e2e/test_public_api.py::test_feed_pause_resume_and_shutdown

Implementation:

  1. Resolve all three executables from the provisioned environment and assert their package resolves inside the active worktree/install.
  2. Add exact-process ownership, bounded pipe readers, protocol readiness, and graceful shutdown to the subprocess fixture.
  3. Move deterministic watchlist subprocess behavior out of the blanket skip. Move legacy live quote/API cases to later qualification or retire them with explicit rationale; do not silently turn live expectations into fixtures.
  4. Run a fresh Python probe using only piker.open_piker_runtime and piker.open_feed plus replay configuration.
  5. Keep existing CliRunner storage tests as T1 for fast diagnostics.

Acceptance gate:

  • each declared script is invoked as an installed executable, not a Click or Typer object;
  • public API probes run in a fresh interpreter and receive real replay-backed Tractor/SHM data;
  • stdout, stderr, exit status, config files, and process ownership are asserted through stable contracts;
  • no test contacts a live provider or writes outside its temporary root;
  • normal teardown needs no forced signal and leaves no descendant, actor, socket, stream, or SHM;
  • tests/test_cli.py no longer uses one unconditional module skip.

Phase 4: nonblocking Qt/Trio lifecycle seam

Exact targets:

  • piker/ui/_exec.py
  • piker/ui/_window.py
  • piker/ui/_event.py
  • piker/ui/_app.py
  • tests/app/conftest.py
  • tests/app/test_lifecycle.py

Planned tests:

  • tests/app/test_lifecycle.py::test_window_close_requests_guest_shutdown
  • tests/app/test_lifecycle.py::test_guest_error_is_observable
  • tests/app/test_lifecycle.py::test_cli_adapter_maps_guest_error_to_failure
  • tests/app/test_lifecycle.py::test_repeated_sessions_restore_qt_and_runtime_state
  • tests/app/test_lifecycle.py::test_signal_and_filter_ownership_is_bounded

Implementation:

  1. Extract a production start_qtractor() seam that configures an existing QApplication, creates the production window/widget, installs Trio guest mode, and returns a QtractorSession without calling app.exec_().
  2. Give QtractorSession stable handles for window, main widget, readiness, shutdown request, completion signal, and outcome.Outcome.
  3. Keep run_qtractor() as the blocking console adapter around that seam. It owns app.exec_() and maps unexpected guest errors to process failure.
  4. Make MainWindow.closeEvent() emit/request shutdown instead of sending SIGINT. Keep geometry persistence, but under the active isolated settings root.
  5. Own and remove focus signals, zoom filters, event relays, and signal proxies. Remove mutable default state that breaks a second session.

Acceptance gate:

  • pytest-qt starts and stops the production window without SIGINT;
  • normal close, requested cancellation, and guest error each produce one observable terminal outcome;
  • an unexpected guest error fails both T2 and blocking-adapter tests;
  • two complete sessions in one pytest process have identical initial globals and no accumulated callback, widget, actor, or module state;
  • run_qtractor() remains the only blocking CLI path; there is no test-only event loop.

Phase 5: chart boot and symbol ownership journeys

Exact targets:

  • piker/ui/_widget.py
  • piker/ui/_display.py
  • piker/ui/_search.py
  • piker/ui/_chart.py only for stable selectors required by the journeys
  • tests/_inputs/replay/overlap-v1.json
  • tests/app/test_symbol_journeys.py
  • tests/e2e/test_chart_process.py

Planned tests:

  • tests/app/test_symbol_journeys.py::test_initial_symbol_renders_from_replay
  • tests/app/test_symbol_journeys.py::test_qtbot_search_switches_a_b_a
  • tests/app/test_symbol_journeys.py::test_overlapping_load_keeps_each_display_owner
  • tests/app/test_symbol_journeys.py::test_superseded_load_cannot_take_focus
  • tests/e2e/test_chart_process.py::test_chart_cold_start_and_graceful_close
  • tests/e2e/test_chart_process.py::test_chart_attaches_to_standalone_pikerd

Implementation:

  1. Represent each load with an owned group key, linked chart pair, generation, readiness, and cancellation/completion handle.
  2. Pass the owned pair into display_symbol_data(); do not let an older task rediscover its widgets through current GodWidget fields.
  3. Return the load handle to search so cache/focus updates use the completed request, not whichever globals are current when an await resumes.
  4. Add only the production object/accessibility names needed to select search, results, realtime chart viewport, and semantic current-symbol state.
  5. Use QtBot to press Ctrl-L, type, navigate, select, and switch. Use replay barriers to force both overlapping completion orders.
  6. Pair the T2 interaction tests with installed piker -b replay chart ... process startup/shutdown tests. The process test proves system boundaries; the QtBot test proves input/render behavior.

Acceptance gate:

  • all claimed interactions are QtBot events against shown/focused production widgets;
  • A to B to A reuses the intended cached pair without a duplicate display task or feed consumer;
  • blocked A completing after B cannot steal Bs focus, sidepane, title, order mode, or cache identity;
  • both linked pairs retain their own feed, Viz, scene items, and cancellation owner;
  • the installed chart process reaches typed replay/feed readiness and exits cleanly with implicit and explicit pikerd;
  • no private layout position or arbitrary timing assertion enters T2/T3.

Phase 6: paper-order human journey

Exact targets:

  • piker/brokers/replay.py
  • piker/ui/_interaction.py only for behavior proven by the journey
  • piker/ui/order_mode.py only for behavior proven by the journey
  • piker/ui/_lines.py only for behavior proven by the journey
  • tests/e2e/probes/open_ems.py
  • tests/e2e/test_public_api.py
  • tests/app/test_order_journeys.py

Planned tests:

  • tests/app/test_order_journeys.py::test_qtbot_submits_and_fills_paper_buy
  • tests/app/test_order_journeys.py::test_qtbot_cancels_open_order
  • tests/app/test_order_journeys.py::test_symbol_switch_preserves_order_owner
  • tests/e2e/test_public_api.py::test_open_ems_paper_fill_from_fresh_python
  • tests/e2e/test_public_api.py::test_open_ems_cancel_and_reconnect

Implementation:

  1. Use real keyboard staging and a QtBot mouse click mapped through the chart viewport to submit.
  2. Let real OrderClient, EMS, paper engine, ledger, position table, and replay feed produce state. Do not inject fabricated Status messages into the UI.
  3. Advance a replay quote through the typed control endpoint to cross or fill the order only after the open-order state is acknowledged.
  4. Assert semantic status/position/line state and isolated persisted files.
  5. Exercise the same core order flow from a fresh open_ems probe to cover process and public API boundaries.

Acceptance gate:

  • mouse and keyboard actions traverse production event relays and order mode;
  • open, fill, cancel, position, and reconnect transitions come from real EMS protocol messages;
  • no direct handler invocation, fake user event, fake EMS stream, live broker, credential, or arbitrary sleep is present;
  • orders and positions remain attached to the initiating symbol/session under a switch;
  • all dialogs, lines, actors, streams, ledgers, and SHM clean up exactly.

Phase 7: component depth for interaction and graphics

Exact targets:

  • tests/ui/test_widgets.py
  • tests/ui/test_graphics_items.py
  • tests/ui/test_chart_interaction.py
  • tests/ui/test_chart_composition.py
  • piker/ui/_search.py, piker/ui/_window.py, piker/ui/_annotate.py, piker/ui/_editors.py, piker/ui/_interaction.py, piker/ui/_overlay.py, and piker/ui/_cursor.py only when a focused test demonstrates a defect

Planned coverage:

  • completer sections, selection, status groups, and focus;
  • annotation/selection attachment, geometry, repositioning, and removal;
  • keyboard focus, cursor, drag, zoom, overlay x-linking, and independent axes;
  • real LinkedSplits/Viz/SHM update cycles and realtime/history region movement.

Acceptance gate:

  • every user-like claim uses QtBot; direct dispatch cases are named as relay or geometry component tests;
  • real Qt/PyQtGraph objects and real SHM are used where their ownership contracts matter;
  • no T1 test is reported as E2E;
  • same-process cleanup holds across the complete tests/ui/ run;
  • production edits remain narrow and defect-driven.

Phase 8: system failure, cancellation, and teardown matrix

Exact targets:

  • tests/e2e/test_chart_process.py
  • tests/e2e/test_failure_paths.py
  • tests/e2e/conftest.py
  • tests/_inputs/replay/failure-v1.json
  • piker/ui/_exec.py, piker/service/_actor_runtime.py, and piker/data/_sharedmem.py only for failures demonstrated by these tests

Planned tests:

  • tests/e2e/test_failure_paths.py::test_provider_failure_exits_nonzero
  • tests/e2e/test_failure_paths.py::test_guest_failure_exits_nonzero
  • tests/e2e/test_failure_paths.py::test_child_actor_failure_has_no_survivors
  • tests/e2e/test_failure_paths.py::test_sigint_unwinds_chart_tree_once
  • tests/e2e/test_failure_paths.py::test_second_system_run_reuses_no_state
  • tests/e2e/test_chart_process.py::test_client_exit_leaves_standalone_pikerd_healthy

Acceptance gate:

  • each fault is acknowledged at its injection boundary before propagation;
  • expected exit status and error class are observable without log matching;
  • success, cancellation, and every injected failure satisfy all cleanup invariants;
  • containment never touches an unrecorded process or SHM;
  • running the full offline E2E set twice yields no fixed-port conflict, inherited actor state, config drift, or forced teardown.

Phase 9: CI and opt-in qualification

Exact targets:

  • .github/workflows/ci.yml
  • .github/workflows/qualification.yml
  • tests/qualification/test_compositor.py
  • tests/qualification/test_live_provider.py
  • .claude/skills/run-tests/test-harness-reference.md
  • README.rst only if replay becomes a supported user-facing mode

Implementation:

  1. Replace obsolete CI installation with frozen uv and supported Python.
  2. Add deterministic headless jobs by tier and keep commands path-explicit.
  3. Upload process, replay, actor, SHM, Qt log, and screenshot artifacts only on failure.
  4. Put real Wayland/X11 and live-provider jobs in an opt-in workflow with protected environments and explicit inputs.
  5. Keep live account/order transmission outside automation unless a separate human-approved qualification protocol is written.

Acceptance gate:

  • required PR jobs run without credentials, network provider access, Docker, or a real desktop session;
  • qualification jobs cannot run accidentally from an ordinary pull request;
  • the workflow commands match the repo-local harness reference exactly;
  • CI labels and summaries preserve T1/T2/T3 terminology and do not advertise component tests as E2E.

CI matrix

The matrix is introduced incrementally as each target directory exists.

Job Trigger Python/platform Selection Blocking policy
Deterministic core Pull request and push 3.12 and 3.13, Linux Existing deterministic non-live modules Required
Qt component headless Pull request and push 3.13, QT_QPA_PLATFORM=offscreen, PyQt6 tests/ui/ plus migrated gap/DPI nodes Required after Phase 1 stabilizes
Offline application Pull request and push 3.13, offscreen tests/app/ with replay provider Required after lifecycle and replay phases
Offline system E2E Pull request and push 3.13, offscreen, subprocess/Tractor/SHM enabled tests/e2e/ Required after process teardown passes repeated runs
Python 3.12 Qt smoke Pull request, initially nonblocking 3.12, offscreen Harness and one application smoke Promote to required after reproducible Qt provisioning is proven
Wayland compositor Manual, scheduled, or protected self-hosted 3.13, real Wayland session tests/qualification/test_compositor.py Opt-in qualification
X11 compositor Manual, scheduled, or protected self-hosted 3.13, real X11 session tests/qualification/test_compositor.py Opt-in qualification
Live provider Manual protected environment with named backend/FQME 3.13, real network tests/qualification/test_live_provider.py Opt-in, read-only, never a PR gate

All headless jobs set these before Python starts:

PYTEST_QT_API=pyqt6
QT_QPA_PLATFORM=offscreen
XDG_CONFIG_HOME=<job-owned-empty-directory>
UV_PROJECT_ENVIRONMENT=py313

Use the frozen project environment and explicit test paths. Do not use develop.nix, add a second Nix pytest-qt package, or depend on ambient pytest plugins. The current flakes Python 3.13/Qt 6 paths are the reference environment (flake.nix:23-57, flake.nix:63-96); uv owns Python test dependencies.

The real-compositor jobs unset QT_QPA_PLATFORM=offscreen, validate the actual platform plugin and display identity, and run focus/exposure/DPI tests. The live job requires an explicit backend and read-only FQME input. Absence of credentials or network is a skip only in that opt-in workflow, never in the deterministic jobs.

Minimal adjacent-document synchronization

  1. Treat plans/opencode/pytest-qt-chart-ui-e2e.md as historical input. Do not edit it as implementation advances.
  2. Update .claude/skills/run-tests/test-harness-reference.md in the same patch only when executable commands, config-source behavior, test topology, environment requirements, or known outcomes actually change.
  3. Update README.rst only if a public command/API contract changes or replay is intentionally supported for users. Internal fixtures, selectors, and CI details do not warrant README churn.
  4. Keep provider scenario documentation beside the scenario schema/provider once that code exists; do not add a speculative standalone guide first.
  5. Do not add broad marker documentation unless the suite actually adopts and consistently enforces those markers. Directory selection is the initial tier contract.
  6. Do not change task/checklist state in this or adjacent planning artifacts as a side effect of implementation. Report gates in review/PR text and leave acceptance state to the human owner.

Change-to-test mapping

Run the narrowest stable layer first, then its paired journey where one exists.

Changed area First deterministic test Required higher-boundary test
pyproject.toml, uv.lock, pytest config tests/ui/test_harness.py collection/config checks tests/e2e/test_console_scripts.py::test_installed_scripts_expose_help
piker/config.py, CLI config paths existing config/storage/watchlist tests test_watchlist_round_trip_is_isolated, test_store_read_only_commands_use_isolated_config, and fresh API probe
piker/cli/__init__.py focused Click callback tests where present all installed-script help plus affected command journey
piker/ui/cli.py, chart arguments argument validation component test test_piker_rejects_invalid_chart_symbol and chart cold-start system test
piker/accounting/cli.py accounting parser/context tests installed ledger help and selected isolated read-only journey
piker/brokers/replay.py complete tests/replay/test_provider.py public feed probe; chart/order journey if affected
piker/data/validate.py, feed.py, flows.py provider conformance plus focused data tests fresh open_feed probe and chart cold start
piker/service/_actor_runtime.py, registry, service manager tests/test_services.py::test_runtime_boot and replay runtime tests standalone pikerd, attach/detach, and system failure matrix
piker/data/_sharedmem.py tests/test_shm_cleanup.py and affected history tests public feed probe plus repeated system cleanup
piker/ui/_exec.py, _window.py, _event.py tests/app/test_lifecycle.py chart process shutdown and failure propagation
piker/ui/_widget.py, _display.py, _search.py focused T1 ownership/search tests all symbol T2 journeys and chart process cold start
piker/ui/_interaction.py tests/ui/test_chart_interaction.py affected QtBot symbol/order journey
piker/ui/order_mode.py, _lines.py, clearing client focused order state tests and replay EMS contract QtBot paper-order journey plus fresh open_ems probe
piker/ui/_annotate.py, _overlay.py, _cursor.py, _editors.py gap/graphics component suite only the affected semantic QtBot journey; no automatic system run if behavior is internal
piker/storage/cli.py tests/test_store_cli.py, tests/test_storage_audit.py, tests/test_ldshm.py selected installed piker store read-only journey
piker/accounting/ persistence focused accounting tests under copied inputs fresh public accounting/EMS probe; never mutate tracked fixtures
workflow/environment files dependency import, collection, harness tests one offline system smoke in the provisioned job

Program-level exit criteria

The coverage program is mature when all of the following are demonstrable, without changing the definition of E2E:

  1. Every installed script has an active fresh-process contract test.
  2. The top-level open_piker_runtime and open_feed APIs have a replay-backed fresh-interpreter journey.
  3. Chart cold start, search switching, paper order flow, failure propagation, and graceful close each have paired T2 QtBot and T3 system evidence where applicable.
  4. Offline provider behavior traverses real protocols, actors, streams, SHM, sampling, and paper EMS without network or credential access.
  5. No T2/T3 test uses mocked user interaction, arbitrary sleep, log readiness, broad process cleanup, or unowned SHM cleanup.
  6. Stable user/protocol selectors carry E2E assertions; volatile internals remain in focused component tests.
  7. Repeated success, cancellation, and failure runs leave no process, subprocess, Trio task, Tractor actor/context/stream, registry socket, SHM, Qt widget/filter/signal, PyQtGraph item, config, or QSettings residue.
  8. Required headless CI is deterministic and credential-free; compositor and live qualifications remain explicit opt-ins.

Until Phase 1 passes, the only honest implementation claim is that the pytest-qt and ownership foundation is being established. Full human-facing E2E remains subsequent work.