Compare commits

..

6 Commits

Author SHA1 Message Date
Gud Boi 3798aa9c9a Clean broadcast cancellation diagnostics
Bound `BroadcastState.cancelled` entries to receiver progress,
terminal state and resource lifetime instead of retaining completed
`Task`s indefinitely.

Deats,
- make EOC durable so awakened peers never re-enter a closed source.
- close root broadcasters during explicit `MsgStream` and
  `LinkedTaskChannel` teardown without re-entrant EOC closure or
  breaking `MsgStream.aclose()` overrides.
- reject non-positive fan-out retention capacity before constructing
  an unusable zero-length queue.
- cover child/root cancellation cleanup, terminal peer wakeups,
  wrapper teardown, subclass compatibility and zero-buffer rejection.

Prompt-IO: ai/prompt-io/opencode/20260813T181901Z_a2e0df4b_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-27 13:43:40 -04:00
Gud Boi b881512d65 Expose stream subscriber lag policy
`MsgStream.subscribe()` and `LinkedTaskChannel.subscribe()` omitted
`BroadcastReceiver.raise_on_lag`, forcing downstream consumers to
mutate a private receiver attribute when overruns were acceptable.

Add `raise_on_lag` to both public wrappers. The first subscription
sets the irreversible root broadcaster's policy, while every child
selects its own strict or warn/drop/resume behavior independently.

Document both fan-out APIs. Cover policy forwarding plus real IPC
and infected-asyncio paths.

Prompt-IO: ai/prompt-io/opencode/20260812T213117Z_51185487_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-27 13:43:40 -04:00
Gud Boi 933374de2c Isolate `BroadcastReceiver.aclose()` wakeups
Closing any subscriber set the shared `recv_ready` event, even when
another receiver owned the source read. Waiting peers then looped
until an idle source produced another value.

Give each receiver private source-read and peer-wait cancellation
scopes. Closing a waiting peer interrupts only that peer. Closing
the source owner discards post-close source outcomes and then wakes
peers for a clean ownership handoff.

Keep outer task cancellation as `trio.Cancelled`; only explicit
receiver close maps either private scope's cancellation to
`ClosedResourceError`. Assert that scope cancellation implies the
receiver is closed and document the source-owner key check.

Prompt-IO: ai/prompt-io/opencode/20260812T150027Z_c2a6ccef_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-27 13:43:40 -04:00
Gud Boi 08dfd834a6 Wake broadcast peers on shared receive failures
Only `EndOfChannel` and direct cancellation woke tasks waiting
behind the subscriber which owned the underlying receive. Any other
failure cleared `BroadcastState.recv_ready` while peers remained
blocked on its unreachable event.

Publish ordinary receive exceptions as terminal broadcast state.
The owner keeps the original failure while peers drain retained
values and then raise `BroadcastReceiveError` from that cause. Also
wake peers on process-control exits without retaining them as state.

Document the public owner/peer contract and cover current, late and
control-flow subscribers with deterministic bounded regressions.

Prompt-IO: ai/prompt-io/opencode/20260812T030608Z_1095e7f7_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-27 13:43:40 -04:00
Gud Boi 8ab6281a68 Fix `BroadcastState.statistics()` queue counts
`BroadcastState.subs` stores each receiver's next unread deque
index, but `.statistics()` exposed that index as a queue length. A
caught-up receiver looked correct by accident while every queued
count was one short.

Convert cursors to retained, receivable counts and clamp lagged
receivers to the current queue length. Also avoid deprecated
`trio.Event` truthiness when reporting waiter counts.

Cover caught-up, queued, lagged and real-event states using actual
broadcast sends and receives.

Prompt-IO: ai/prompt-io/opencode/20260812T012324Z_06c4af17_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-27 13:43:40 -04:00
Gud Boi 8db13375fc Fix `BroadcastReceiver` lag counts
`BroadcastReceiver.receive_nowait()` treated `seq` as a deque index
but subtracted `BroadcastState.maxlen` without counting the first
invalid index. A one-slot queue thus claimed it dropped zero values
after its subscriber missed one.

Include that first displaced value in the count. Preserve the
existing Tokio-style reset to the oldest retained item.

Also, cover exact loss reporting and recovery for one- and
three-slot retention windows.

Prompt-IO: ai/prompt-io/opencode/20260811T233833Z_7cbd64ee_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-27 13:43:40 -04:00
15 changed files with 641 additions and 878 deletions

View File

@ -0,0 +1,632 @@
---
name: run-tests
description: >
Run tractor test suite (or subsets). Use when the user wants
to run tests, verify changes, or check for regressions.
argument-hint: "[test-path-or-pattern] [--opts]"
allowed-tools:
- Bash(python -m pytest *)
- Bash(python -c *)
- Bash(python --version *)
- Bash(UV_PROJECT_ENVIRONMENT=py* uv run python *)
- Bash(UV_PROJECT_ENVIRONMENT=py* uv run pytest *)
- Bash(UV_PROJECT_ENVIRONMENT=py* uv sync *)
- Bash(UV_PROJECT_ENVIRONMENT=py* uv pip show *)
- Bash(git rev-parse *)
- Bash(ls *)
- Bash(cat *)
- Bash(jq * .pytest_cache/*)
# process inspection + SIGINT-first cleanup ladder (see
# the zombie-actor pre-flight / teardown steps below).
- Bash(ss *)
- Bash(pgrep *)
- Bash(pkill *)
- Bash(sleep *)
- Bash(rm -f /tmp/registry@*.sock)
- Read
- Grep
- Glob
- Task
- AskUserQuestion
---
Run the `tractor` test suite using `pytest`. Follow this
process:
## 1. Parse user intent
From the user's message and any arguments, determine:
- **scope**: full suite, specific file(s), specific
test(s), or a keyword pattern (`-k`).
- **transport**: which IPC transport protocol to test
against (default: `tcp`, also: `uds`).
- **options**: any extra pytest flags the user wants
(e.g. `--ll debug`, `--tpdb`, `-x`, `-v`).
If the user provides a bare path or pattern as argument,
treat it as the test target. Examples:
- `/run-tests` → full suite
- `/run-tests test_local.py` → single file
- `/run-tests test_registrar -v` → file + verbose
- `/run-tests -k cancel` → keyword filter
- `/run-tests tests/ipc/ --tpt-proto uds` → subdir + UDS
## 2. Construct the pytest command
Base command:
```
python -m pytest
```
### Default flags (always include unless user overrides):
- `-x` (stop on first failure)
- `--tb=short` (concise tracebacks)
- `--no-header` (reduce noise)
### Path resolution:
- If the user gives a bare filename like `test_local.py`,
resolve it under `tests/`.
- If the user gives a subdirectory like `ipc/`, resolve
under `tests/ipc/`.
- Glob if needed: `tests/**/test_*<pattern>*.py`
### Key pytest options for this project:
| Flag | Purpose |
|---|---|
| `--ll <level>` | Set tractor log level (e.g. `debug`, `info`, `runtime`) |
| `--tpdb` / `--debug-mode` | Enable tractor's multi-proc debugger |
| `--tpt-proto <key>` | IPC transport: `tcp` (default) or `uds` |
| `--spawn-backend <be>` | Spawn method: `trio` (default), `mp_spawn`, `mp_forkserver` |
| `-k <expr>` | pytest keyword filter |
| `-v` / `-vv` | Verbosity |
| `-s` | No output capture (useful with `--tpdb`) |
### Common combos:
```sh
# quick smoke test of core modules
python -m pytest tests/test_local.py tests/test_rpc.py -x --tb=short --no-header
# full suite, stop on first failure
python -m pytest tests/ -x --tb=short --no-header
# specific test with debug
python -m pytest tests/discovery/test_registrar.py::test_reg_then_unreg -x -s --tpdb --ll debug
# run with UDS transport
python -m pytest tests/ -x --tb=short --no-header --tpt-proto uds
# keyword filter
python -m pytest tests/ -x --tb=short --no-header -k "cancel and not slow"
```
## 3. Pre-flight: venv detection (MANDATORY)
**Always verify a `uv` venv is active before running
`python` or `pytest`.** This project uses
`UV_PROJECT_ENVIRONMENT=py<MINOR>` naming (e.g.
`py313`) — never `.venv`.
### Step 1: detect active venv
Run this check first:
```sh
python -c "
import sys, os
venv = os.environ.get('VIRTUAL_ENV', '')
prefix = sys.prefix
print(f'VIRTUAL_ENV={venv}')
print(f'sys.prefix={prefix}')
print(f'executable={sys.executable}')
"
```
### Step 2: interpret results
**Case A — venv is active** (`VIRTUAL_ENV` is set
and points to a `py<MINOR>/` dir under the project
root or worktree):
Use bare `python` / `python -m pytest` for all
commands. This is the normal, fast path.
**Case B — no venv active** (`VIRTUAL_ENV` is empty
or `sys.prefix` points to a system Python):
Use `AskUserQuestion` to ask the user:
> "No uv venv is active. Should I activate one
> via `UV_PROJECT_ENVIRONMENT=py<MINOR> uv sync`,
> or would you prefer to activate your shell venv
> first?"
Options:
1. **"Create/sync venv"** — run
`UV_PROJECT_ENVIRONMENT=py<MINOR> uv sync` where
`<MINOR>` is detected from `python --version`
(e.g. `313` for 3.13). Then use
`py<MINOR>/bin/python` for all subsequent
commands in this session.
2. **"I'll activate it myself"** — stop and let the
user `source py<MINOR>/bin/activate` or similar.
**Case C — inside a git worktree** (`git rev-parse
--git-common-dir` differs from `--git-dir`):
Verify Python resolves from the **worktree's own
venv**, not the main repo's:
```sh
python -c "import tractor; print(tractor.__file__)"
```
If the path points outside the worktree, create a
worktree-local venv:
```sh
UV_PROJECT_ENVIRONMENT=py<MINOR> uv sync
```
Then use `py<MINOR>/bin/python` for all commands.
**Why this matters**: without the correct venv,
subprocesses spawned by tractor resolve modules
from the wrong editable install, causing spurious
`AttributeError` / `ModuleNotFoundError`.
### Fallback: `uv run`
If the user can't or won't activate a venv, all
`python` and `pytest` commands can be prefixed
with `UV_PROJECT_ENVIRONMENT=py<MINOR> uv run`:
```sh
# instead of: python -m pytest tests/ -x
UV_PROJECT_ENVIRONMENT=py313 uv run pytest tests/ -x
# instead of: python -c 'import tractor'
UV_PROJECT_ENVIRONMENT=py313 uv run python -c 'import tractor'
```
`uv run` auto-discovers the project and venv,
but is slower than a pre-activated venv due to
lock-file resolution on each invocation. Prefer
activating the venv when possible.
### Step 3: import + collection checks
After venv is confirmed, always run these
(especially after refactors or module moves):
```sh
# 1. package import smoke check
python -c 'import tractor; print(tractor)'
# 2. verify all tests collect (no import errors)
python -m pytest tests/ -x -q --co 2>&1 | tail -5
```
If either fails, fix the import error before running
any actual tests.
### Step 4: zombie-actor / stale-registry check (MANDATORY)
The tractor runtime's default registry address is
**`127.0.0.1:1616`** (TCP) / `/tmp/registry@1616.sock`
(UDS). Whenever any prior test run — especially one
using a fork-based backend like `subint_forkserver`
leaks a child actor process, that zombie keeps the
registry port bound and **every subsequent test
session fails to bind**, often presenting as 50+
unrelated failures ("all tests broken"!) across
backends.
**This has to be checked before the first run AND
after any cancelled/SIGINT'd run** — signal failures
in the middle of a test can leave orphan children.
```sh
# 1. TCP registry — any listener on :1616? (primary signal)
ss -tlnp 2>/dev/null | grep ':1616' || echo 'TCP :1616 free'
# 2. leftover actor/forkserver procs — scoped to THIS
# repo's python path, so we don't false-flag legit
# long-running tractor-using apps (e.g. `piker`,
# downstream projects that embed tractor).
pgrep -af "$(pwd)/py[0-9]*/bin/python.*_actor_child_main|subint-forkserv" \
| grep -v 'grep\|pgrep' \
|| echo 'no leaked actor procs from this repo'
# 3. stale UDS registry sockets
ls -la /tmp/registry@*.sock 2>/dev/null \
|| echo 'no leaked UDS registry sockets'
```
**Interpretation:**
- **TCP :1616 free AND no stale sockets** → clean,
proceed. The actor-procs probe is secondary — false
positives are common (piker, any other tractor-
embedding app); only cleanup if `:1616` is bound or
sockets linger.
- **TCP :1616 bound OR stale sockets present**
surface PIDs + cmdlines to the user, offer cleanup:
```sh
# 1. GRACEFUL FIRST (tractor is structured concurrent — it
# catches SIGINT as an OS-cancel in `_trio_main` and
# cascades Portal.cancel_actor via IPC to every descendant.
# So always try SIGINT first with a bounded timeout; only
# escalate to SIGKILL if graceful cleanup doesn't complete).
pkill -INT -f "$(pwd)/py[0-9]*/bin/python.*_actor_child_main|subint-forkserv"
# 2. bounded wait for graceful teardown (usually sub-second).
# Loop until the processes exit, or timeout. Keep the
# bound tight — hung/abrupt-killed descendants usually
# hang forever, so don't wait more than a few seconds.
for i in $(seq 1 10); do
pgrep -f "$(pwd)/py[0-9]*/bin/python.*_actor_child_main|subint-forkserv" >/dev/null || break
sleep 0.3
done
# 3. ESCALATE TO SIGKILL only if graceful didn't finish.
if pgrep -f "$(pwd)/py[0-9]*/bin/python.*_actor_child_main|subint-forkserv" >/dev/null; then
echo 'graceful teardown timed out — escalating to SIGKILL'
pkill -9 -f "$(pwd)/py[0-9]*/bin/python.*_actor_child_main|subint-forkserv"
fi
# 4. if a test zombie holds :1616 specifically and doesn't
# match the above pattern, find its PID the hard way:
ss -tlnp 2>/dev/null | grep ':1616' # prints `users:(("<name>",pid=NNNN,...))`
# then (same SIGINT-first ladder):
# kill -INT <NNNN>; sleep 1; kill -9 <NNNN> 2>/dev/null
# 5. remove stale UDS sockets
rm -f /tmp/registry@*.sock
# 6. re-verify
ss -tlnp 2>/dev/null | grep ':1616' || echo 'TCP :1616 now free'
```
**Never ignore stale registry state.** If you see the
"all tests failing" pattern — especially
`trio.TooSlowError` / connection refused / address in
use on many unrelated tests — check registry **before**
spelunking into test code. The failure signature will
be identical across backends because they're all
fighting for the same port.
**False-positive warning for step 2:** a plain
`pgrep -af '_actor_child_main'` will also match
legit long-running tractor-embedding apps (e.g.
`piker` at `~/repos/piker/py*/bin/python3 -m
tractor._child ...`). Always scope to the current
repo's python path, or only use step 1 (`:1616`) as
the authoritative signal.
## 4. Run and report
- Run the constructed command.
- Use a timeout of **600000ms** (10min) for full suite
runs, **120000ms** (2min) for single-file runs.
- If the suite is large (full `tests/`), consider running
in the background and checking output when done.
- Use `--lf` (last-failed) to re-run only previously
failing tests when iterating on a fix.
### On failure:
- Show the failing test name(s) and short traceback.
- If the failure looks related to recent changes, point
out the likely cause and suggest a fix.
- **Check the known-flaky list** (section 8) before
investigating — don't waste time on pre-existing
timeout issues.
- **NEVER auto-commit fixes.** If you apply a code fix
during test iteration, leave it unstaged. Tell the
user what changed and suggest they review the
worktree state, stage files manually, and use
`/commit-msg` (inline or in a separate session) to
generate the commit message. The human drives all
`git add` and `git commit` operations.
### On success:
- Report the pass/fail/skip counts concisely.
## 5. Test directory layout (reference)
```
tests/
├── conftest.py # root fixtures, daemon, signals
├── devx/ # debugger/tooling tests
├── ipc/ # transport protocol tests
├── msg/ # messaging layer tests
├── discovery/ # discovery subsystem tests
│ ├── test_multiaddr.py # multiaddr construction
│ └── test_registrar.py # registry/discovery protocol
├── test_local.py # registrar + local actor basics
├── test_rpc.py # RPC error handling
├── test_spawning.py # subprocess spawning
├── test_multi_program.py # multi-process tree tests
├── test_cancellation.py # cancellation semantics
├── test_context_stream_semantics.py # ctx streaming
├── test_inter_peer_cancellation.py # peer cancel
├── test_infected_asyncio.py # trio-in-asyncio
└── ...
```
## 6. Change-type → test mapping
After modifying specific modules, run the corresponding
test subset first for fast feedback:
| Changed module(s) | Run these tests first |
|---|---|
| `runtime/_runtime.py`, `runtime/_state.py` | `test_local.py test_rpc.py test_spawning.py test_root_runtime.py` |
| `discovery/` (`_registry`, `_discovery`, `_addr`) | `tests/discovery/ test_multi_program.py test_local.py` |
| `_context.py`, `_streaming.py` | `test_context_stream_semantics.py test_advanced_streaming.py` |
| `ipc/` (`_chan`, `_server`, `_transport`) | `tests/ipc/ test_2way.py` |
| `runtime/_portal.py`, `runtime/_rpc.py` | `test_rpc.py test_cancellation.py` |
| `spawn/` (`_spawn`, `_entry`) | `test_spawning.py test_multi_program.py` |
| `devx/debug/` | `tests/devx/test_debugger.py` (slow!) |
| `to_asyncio.py` | `test_infected_asyncio.py test_root_infect_asyncio.py` |
| `msg/` | `tests/msg/` |
| `_exceptions.py` | `test_remote_exc_relay.py test_inter_peer_cancellation.py` |
| `runtime/_supervise.py` | `test_cancellation.py test_spawning.py` |
## 7. Quick-check shortcuts
### After refactors (fastest first-pass):
```sh
# import + collect check
python -c 'import tractor' && python -m pytest tests/ -x -q --co 2>&1 | tail -3
# core subset (~10s)
python -m pytest tests/test_local.py tests/test_rpc.py tests/test_spawning.py tests/discovery/test_registrar.py -x --tb=short --no-header
```
### Inspect last failures (without re-running):
When the user asks "what failed?", "show failures",
or wants to check the last-failed set before
re-running — read the pytest cache directly. This
is instant and avoids test collection overhead.
```sh
python -c "
import json, pathlib, sys
p = pathlib.Path('.pytest_cache/v/cache/lastfailed')
if not p.exists():
print('No lastfailed cache found.'); sys.exit()
data = json.loads(p.read_text())
# filter to real test node IDs (ignore junk
# entries that can accumulate from system paths)
tests = sorted(k for k in data if k.startswith('tests/'))
if not tests:
print('No failures recorded.')
else:
print(f'{len(tests)} last-failed test(s):')
for t in tests:
print(f' {t}')
"
```
**Why not `--cache-show` or `--co --lf`?**
- `pytest --cache-show 'cache/lastfailed'` works
but dumps raw dict repr including junk entries
(stale system paths that leak into the cache).
- `pytest --co --lf` actually *collects* tests which
triggers import resolution and is slow (~0.5s+).
Worse, when cached node IDs don't exactly match
current parametrize IDs (e.g. param names changed
between runs), pytest falls back to collecting
the *entire file*, giving false positives.
- Reading the JSON directly is instant, filterable
to `tests/`-prefixed entries, and shows exactly
what pytest recorded — no interpretation.
**After inspecting**, re-run the failures:
```sh
python -m pytest --lf -x --tb=short --no-header
```
### Full suite in background:
When core tests pass and you want full coverage while
continuing other work, run in background:
```sh
python -m pytest tests/ -x --tb=short --no-header -q
```
(use `run_in_background=true` on the Bash tool)
## 8. Known flaky tests
These tests have **pre-existing** timing/environment
sensitivity. If they fail with `TooSlowError` or
pexpect `TIMEOUT`, they are almost certainly NOT caused
by your changes — note them and move on.
| Test | Typical error | Notes |
|---|---|---|
| `devx/test_debugger.py::test_multi_nested_subactors_error_through_nurseries` | pexpect TIMEOUT | Debugger pexpect timing |
| `test_cancellation.py::test_cancel_via_SIGINT_other_task` | TooSlowError | Signal handling race |
| `test_inter_peer_cancellation.py::test_peer_spawns_and_cancels_service_subactor` | TooSlowError | Async timing (both param variants) |
| `test_docs_examples.py::test_example[we_are_processes.py]` | `assert None == 0` | `__main__` missing `__file__` in subproc |
**Rule of thumb**: if a test fails with `TooSlowError`,
`trio.TooSlowError`, or `pexpect.TIMEOUT` and you didn't
touch the relevant code path, it's flaky — skip it.
## 9. The pytest-capture hang pattern (CHECK THIS FIRST)
**Symptom:** a tractor test hangs indefinitely under
default `pytest` but passes instantly when you add
`-s` (`--capture=no`).
**Cause:** tractor subactors (especially under fork-
based backends) inherit pytest's stdout/stderr
capture pipes via fds 1,2. Under high-volume error
logging (e.g. multi-level cancel cascade, nested
`run_in_actor` failures, anything triggering
`RemoteActorError` + `ExceptionGroup` traceback
spew), the **64KB Linux pipe buffer fills** faster
than pytest drains it. Subactor writes block → can't
finish exit → parent's `waitpid`/pidfd wait blocks →
deadlock cascades up the tree.
**Pre-existing guards in the tractor harness** that
encode this same knowledge — grep these FIRST
before spelunking:
- `tests/conftest.py:258-260` (in the `daemon`
fixture): `# XXX: too much logging will lock up
the subproc (smh)` — downgrades `trace`/`debug`
loglevel to `info` to prevent the hang.
- `tests/conftest.py:316`: `# can lock up on the
_io.BufferedReader and hang..` — noted on the
`proc.stderr.read()` post-SIGINT.
**Debug recipe (in priority order):**
1. **Try `-s` first.** If the hang disappears with
`pytest -s`, you've confirmed it's capture-pipe
fill. Skip spelunking.
2. **Lower the loglevel.** Default `--ll=error` on
this project; if you've bumped it to `debug` /
`info`, try dropping back. Each log level
multiplies pipe-pressure under fault cascades.
3. **If you MUST use default capture + high log
volume**, redirect subactor stdout/stderr in the
child prelude (e.g.
`tractor.spawn._subint_forkserver._child_target`
post-`_close_inherited_fds`) to `/dev/null` or a
file.
**Signature tells you it's THIS bug (vs. a real
code hang):**
- Multi-actor test under fork-based backend
(`subint_forkserver`, eventually `trio_proc` too
under enough log volume).
- Multiple `RemoteActorError` / `ExceptionGroup`
tracebacks in the error path.
- Test passes with `-s` in the 5-10s range, hangs
past pytest-timeout (usually 30+ s) without `-s`.
- Subactor processes visible via `pgrep -af
subint-forkserv` or similar after the hang —
they're alive but blocked on `write()` to an
inherited stdout fd.
**Historical reference:** this deadlock cost a
multi-session investigation (4 genuine cascade
fixes landed along the way) that only surfaced the
capture-pipe issue AFTER the deeper fixes let the
tree actually tear down enough to produce pipe-
filling log volume. Full post-mortem in
`ai/conc-anal/subint_forkserver_test_cancellation_leak_issue.md`.
Lesson codified here so future-me grep-finds the
workaround before digging.
## 10. Reaping zombie subactors (`tractor-reap`)
**Symptom:** after a `pytest` run crashes, times out,
or is `Ctrl+C`'d, subactor forks (esp. under
`subint_forkserver`) can be reparented to `init`
(PPid==1) and linger. They hold onto ports, inherit
pytest's capture-pipe fds, and flakify later
sessions.
**Two layers of defense:**
### a) Session-scoped auto-fixture (always on)
`tractor/_testing/pytest.py::_reap_orphaned_subactors`
runs at pytest session teardown. It walks `/proc` for
direct descendants of the pytest pid, SIGINTs them,
waits up to 3s, then SIGKILLs survivors. SC-polite:
gives the subactor runtime a chance to run its trio
cancel shield + IPC teardown before escalation.
This is *autouse* and session-scoped — you don't need
to do anything. It just runs.
### b) `scripts/tractor-reap` CLI (manual reap)
For the **pytest-died-mid-session** case (Ctrl+C, OOM
kill, hung process you had to `kill -9`), the fixture
never ran. Reach for the CLI:
```sh
# default: orphans (PPid==1, cwd==repo, cmd contains python)
scripts/tractor-reap
# descendant-mode: from a still-live supervisor
scripts/tractor-reap --parent <pytest-pid>
# see what would be reaped, don't signal
scripts/tractor-reap -n
# tune the SIGINT → SIGKILL grace window
scripts/tractor-reap --grace 5
```
Exit code: `0` if everyone exited on SIGINT, `1` if
SIGKILL had to escalate — so you can chain it in CI
health-checks (`scripts/tractor-reap || <alert>`).
**What it matches** (orphan-mode):
- `PPid == 1` (reparented to init → definitely
orphaned, not just a currently-running child)
- `cwd == <repo-root>` (keeps the sweep scoped; won't
touch unrelated init-children elsewhere)
- `python` in cmdline
**What it does not do:** kill anything whose PPid is
still a live tractor parent. If the parent is alive
it's not an orphan; use `--parent <pid>` if you need
to force-reap under a still-live supervisor.
**When NOT to run it:** while a pytest session is
active in another terminal. It's safe (won't touch
that session's live children in orphan-mode) but can
race if the target session is mid-teardown.
### c) `--shm` / `--shm-only`: orphan-segment sweep
Because `tractor.ipc._mp_bs.disable_mantracker()`
turns off `mp.resource_tracker` (see
`ai/conc-anal/subint_forkserver_mp_shared_memory_issue.md`),
a hard-crashing actor can leave `/dev/shm/<key>`
segments behind that nothing else GCs.
```sh
# process reap THEN shm sweep
scripts/tractor-reap --shm
# shm sweep only (skip process phase)
scripts/tractor-reap --shm-only
# dry-run: list candidates, don't unlink
scripts/tractor-reap --shm -n
```
**Match criteria** (very conservative — this is a
shared-system path, can't be wrong):
- segment is a regular file under `/dev/shm`,
- owned by the **current uid** (`stat.st_uid`),
- AND **no live process holds it open**
enumerated by walking every readable
`/proc/<pid>/maps` (post-mmap mappings) AND
`/proc/<pid>/fd/*` (pre-mmap shm-opened fds).
The "nobody has it open" check is the
kernel-canonical "is this leaked?" test — same
answer `lsof /dev/shm/<key>` would give. No
reliance on tractor-specific naming, so it works
for any tractor app. Critically, it WILL NOT touch
segments held by other apps you have running
(e.g. `piker`, `lttng-ust-*`, `aja-shm-*`
verified locally with 81 in-use segments correctly
preserved).

View File

@ -1,255 +0,0 @@
# Tractor Test Harness Reference
This repository-local file supplements the canonical [`/run-tests` skill][1]
from the [`ai.skillz` repository][2]. Its deployer links the shared `SKILL.md`
into both
`.claude/skills/run-tests/` and `.opencode/skills/run-tests/` while preserving
this project-owned reference:
```text
bash /path/to/ai.skillz/scripts/deploy.sh run-tests /path/to/tractor --provider all --method symlink
```
Keep shared environment permission, process-signal safety, target selection,
failure inspection, and result reporting policy in the deployed `SKILL.md`.
[1]: https://github.com/baudco/ai.skillz/blob/2d4896ca7e38fe2cb3090cdefc7245be4241a6d2/skills/run-tests/SKILL.md
[2]: https://github.com/baudco/ai.skillz
## Project And Environment
- Project/import: `tractor`
- Test root: `tests/`
- Supported Python: `>=3.13,<3.15`
- Runner: pytest `>=9.0.3`
- Test dependencies: the `dev` group includes the `testing` group
- CI uses uv's default `.venv`; the Nix flake uses `py313`.
- Run from the repository root so pytest loads `pyproject.toml`.
- Do not use `default.nix` as current test-environment authority; it still
selects unsupported Python 3.12.
Environment directory naming is not a harness invariant. Use an already
verified active project environment when available. Otherwise, use an
existing uv environment without syncing it:
```text
uv run --frozen --no-sync python -c 'import pathlib, sys, tractor; root = pathlib.Path.cwd().resolve(); mod = pathlib.Path(tractor.__file__).resolve(); print(sys.executable); print(mod); assert mod.is_relative_to(root)'
```
After module moves or collection failures, check collection with:
```text
uv run --frozen --no-sync pytest --collect-only -q tests/
```
Collection is not a mandatory precursor to every narrow run. Ask before
provisioning or changing an environment.
Before trusting CLI-selected runtime settings, inspect
`TRACTOR_SPAWN_METHOD` and `TRACTOR_LOGLEVEL`. They override the spawn method
and runtime log level passed by callers, so report active values with test
results rather than claiming the CLI flags alone selected the runtime.
## Pytest Configuration And Commands
`pyproject.toml` configures:
- `testpaths = ["tests"]` and `--rootdir=./tests`;
- importlib import mode;
- the `tractor._testing.pytest` plugin;
- xonsh plugin disablement;
- `--show-capture=no` and `--capture=fd`.
Do not silently add `-x`, `--tb=short`, or `--no-header`; those are not
project defaults. In a verified active environment, replace `uv run
--frozen --no-sync pytest` below with `python -m pytest`.
```text
# Full suite
uv run --frozen --no-sync pytest tests/
# Narrow file
uv run --frozen --no-sync pytest tests/test_local.py
# Exact node
uv run --frozen --no-sync pytest tests/discovery/test_registrar.py::test_reg_then_unreg
# Keyword selection
uv run --frozen --no-sync pytest tests/ -k 'cancel and not slow'
# Previous failures
uv run --frozen --no-sync pytest --lf
```
After verifying that the no-sync environment is current, these pytest
arguments match the Linux TCP CI row:
```text
CI=1 uv run --frozen --no-sync pytest tests/ -rsx --spawn-backend=trio --tpt-proto=tcp --capture=fd
```
## Plugin Options And Matrices
Supported spawn backends:
- `trio` (default)
- `mp_spawn`
- `mp_forkserver`
Do not advertise `subint`, `subint_forkserver`, or
`main_thread_forkserver` as runnable backends. Supported transports are
`tcp` (default) and `uds`. Run one transport per pytest session.
`mp_forkserver` and UDS are POSIX-only.
Other Tractor plugin options include:
- `--tpdb` / `--debug-mode`
- `--ll` / `--loglevel`
- `--tl` / `--tractor-loglevel`
- `--enable-stackscope`
Examples:
```text
uv run --frozen --no-sync pytest tests/ipc/ --tpt-proto=uds
uv run --frozen --no-sync pytest tests/test_spawning.py --spawn-backend=mp_spawn
uv run --frozen --no-sync pytest tests/test_spawning.py --spawn-backend=mp_forkserver --capture=sys
```
CI currently exercises Python 3.13 with the `trio` backend: TCP and UDS on
Linux and macOS, plus an informational TCP row on Windows whose pytest step
uses `continue-on-error`.
## Registry And Transport Isolation
Tests requesting the `reg_addr` fixture use addresses randomized per session:
an unreserved unprivileged loopback port for TCP or a unique socket name under
the platform runtime directory for UDS. A TCP collision remains possible.
The runtime fallback remains `127.0.0.1:1616` or `registry@1616.sock`.
Inspect that fallback only when the selected test intentionally uses runtime
defaults or a failure identifies that address. Do not perform a mandatory
`:1616` preflight or assume UDS sockets live under `/tmp`.
## Capture And Hang Diagnosis
Normal capture is `fd`. Use `--capture=sys` with `mp_forkserver`; some tests
switch to `capsys`, but the harness does not enforce that suite-wide.
For a suspected capture interaction, compare only the exact node:
```text
uv run --frozen --no-sync pytest <node> --capture=sys
uv run --frozen --no-sync pytest <node> -s
```
Treat `-s` as a diagnostic comparison, not a pass-equivalent workaround. Do
not use it to reinterpret an ordinary captured pass. Interactive `--tpdb` or
`tractor.pause()` sessions are different: they require a real TTY and disabled
capture, normally `-s`.
Do not add a global pytest timeout. `fail_after_w_trace` is Trio-cooperative;
`afk_alarm_w_trace` is a POSIX main-thread `SIGALRM` hard backstop and can
raise asynchronously. Use the latter only as a last resort, not as a generally
Trio-safe timeout replacement.
For live task-tree diagnosis:
```text
uv run --frozen --no-sync python -c 'import stackscope'
uv run --frozen --no-sync pytest <node> --enable-stackscope --capture=sys
kill -USR1 <pytest-pid>
```
The import check and pytest command must use the same environment. Do not send
SIGUSR1 if the import fails or setup warns that stackscope or SIGUSR1 is
unavailable: without the installed handler, SIGUSR1 normally terminates the
target process. Signal a subactor only after separately confirming that it
installed the same handler.
Stackscope appends dumps to `/tmp/tractor-stackscope-<pid>.log`, including when
pytest capture hides terminal output. SIGUSR1 stackscope is unavailable on
Windows and degrades to a no-op there.
When a trace guard actually fires and snapshot capture succeeds, it writes
under `$XDG_CACHE_HOME/tractor/hung-dumps/`, falling back beneath
`~/.cache/tractor/hung-dumps/`, and prints an end-of-session index. A normal
non-timeout run creates no snapshot.
## Cleanup And `tractor-reap`
On Linux, normal pytest teardown discovers surviving descendants through
`/proc`, sends SIGINT, waits three seconds, then escalates survivors to
SIGKILL. It does not sweep shared memory and cannot run if pytest never reaches
fixture teardown.
Process discovery is a no-op off Linux. UDS PID liveness also depends on
`/proc`; on macOS, recognized PID-named sockets can therefore be classified as
dead without proof. The session-scoped autouse fixture currently passes those
candidates directly to `reap_uds()` at teardown. Do not treat its non-Linux
classification as proof of orphanhood or run concurrent live Tractor sessions
against the same UDS bindspace.
Use the CLI in inspection-only mode first:
`--shm` and `--shm-only` are Linux/FreeBSD-only and raise
`NotImplementedError` elsewhere. On other platforms, use the UDS-only command.
```text
uv run --frozen --no-sync scripts/tractor-reap -n
uv run --frozen --no-sync scripts/tractor-reap --parent <pytest-pid> -n
uv run --frozen --no-sync scripts/tractor-reap --shm --uds -n
uv run --frozen --no-sync scripts/tractor-reap --uds-only -n
```
Direct `scripts/tractor-reap` execution is acceptable only after verifying its
`python3` shebang resolves the intended project environment.
Review every candidate before requesting a mutating run:
- default orphan mode is not repository-scoped;
- `--parent` trusts the supplied PID and can include non-Tractor children;
- `--shm` scans all current-user candidate files, not just Tractor-named
files;
- `--uds` treats `registry@1616.sock` as removable even if a live default UDS
registrar uses it.
Dry-run output prints only the initially matched root PIDs. A mutating run can
recursively expand those roots to additional descendants when `psutil` is
available. Inspect the descendant process tree separately; `-n` is not exact
signal-set parity and does not by itself authorize signaling unseen children.
The canonical skill owns signaling and unlinking authorization.
## Test Layout And Change Mapping
| Changed area | Run first |
|---|---|
| `tractor/runtime/_runtime.py`, `_state.py`, `tractor/_root.py` | `tests/test_local.py`, `tests/test_root_runtime.py`, `tests/test_runtime.py`, `tests/test_rpc.py` |
| `tractor/runtime/_portal.py`, `_rpc.py` | `tests/test_rpc.py`, `tests/test_cancellation.py` |
| `tractor/runtime/_supervise.py` | `tests/test_cancellation.py`, `tests/test_spawning.py` |
| `tractor/discovery/` | `tests/discovery/`, `tests/test_local.py` |
| `tractor/ipc/` | `tests/ipc/`, `tests/test_2way.py`, `tests/test_shm.py` as relevant |
| `tractor/spawn/` | `tests/test_spawning.py`, `tests/discovery/test_multi_program.py`, `tests/test_cancellation.py` |
| `tractor/_context.py`, `_streaming.py` | `tests/test_context_stream_semantics.py`, `tests/test_advanced_streaming.py`, `tests/test_legacy_one_way_streaming.py` |
| `tractor/to_asyncio.py` | `tests/test_infected_asyncio.py`, `tests/test_root_infect_asyncio.py` |
| `tractor/msg/` | `tests/msg/` |
| `tractor/devx/` | `tests/devx/`; debugger tests use pexpect and are comparatively slow |
| `tractor/_exceptions.py` | `tests/test_remote_exc_relay.py`, `tests/test_reg_err_types.py`, `tests/test_inter_peer_cancellation.py`, `tests/test_cancellation.py`, `tests/msg/` |
Current subdirectories include `discovery/`, `ipc/`, `msg/`, `devx/`, and
`trionics/`. There is no `tests/spawn/` directory.
## Expected Outcomes
Do not maintain a blanket known-flaky exemption list. Classify only current
explicit skip or xfail marks and exact expected signatures. Notable tracked
outcomes include:
- duplicate-name `n_dups=4` and `n_dups=8` variants in
`tests/discovery/test_multi_program.py` are non-strict xfails;
- `tests/test_ringbuf.py` is module-skipped;
- some documentation examples have explicit macOS-CI skips.
A generic `TooSlowError` or `pexpect.TIMEOUT` is not enough to classify a
failure as pre-existing.

View File

@ -162,10 +162,6 @@ jobs:
run: uv run python -c "import sys; import tractor; from tractor.ipc._uds import HAS_UDS; assert sys.platform != 'win32' or not HAS_UDS; print('import tractor OK | HAS_UDS=', HAS_UDS)" run: uv run python -c "import sys; import tractor; from tractor.ipc._uds import HAS_UDS; assert sys.platform != 'win32' or not HAS_UDS; print('import tractor OK | HAS_UDS=', HAS_UDS)"
- name: Run tests - name: Run tests
# Actor/PTY scheduling on macOS can fail a different
# timing-sensitive node between otherwise-green runs. Retry
# only that matrix leg; deterministic failures still fail
# after the final attempt.
continue-on-error: ${{ matrix.os == 'windows-latest' }} continue-on-error: ${{ matrix.os == 'windows-latest' }}
run: > run: >
uv run uv run
@ -175,8 +171,6 @@ jobs:
--spawn-backend=${{ matrix.spawn_backend }} --spawn-backend=${{ matrix.spawn_backend }}
--tpt-proto=${{ matrix.tpt_proto }} --tpt-proto=${{ matrix.tpt_proto }}
--capture=fd --capture=fd
--reruns=${{ matrix.os == 'macos-latest' && 2 || 0 }}
--reruns-delay=1
# XXX legacy NOTE XXX # XXX legacy NOTE XXX
# #

264
.gitignore vendored
View File

@ -168,267 +168,3 @@ gh/
# LLM conversations that should remain private # LLM conversations that should remain private
docs/conversations/ docs/conversations/
# BEGIN ai.skillz: direct:symlink:claude:run-tests
/.claude/skills/run-tests/SKILL.md
# END ai.skillz: direct:symlink:claude:run-tests
# BEGIN ai.skillz: direct:symlink:opencode:run-tests
/.opencode/skills/run-tests/SKILL.md
# END ai.skillz: direct:symlink:opencode:run-tests
# BEGIN ai.skillz: direct:symlink:opencode:command:run-tests
/.opencode/commands/run-tests.md
# END ai.skillz: direct:symlink:opencode:command:run-tests
# BEGIN ai.skillz: direct:symlink:claude:gish
/.claude/skills/gish
# END ai.skillz: direct:symlink:claude:gish
# BEGIN ai.skillz: direct:symlink:opencode:gish
/.opencode/skills/gish
# END ai.skillz: direct:symlink:opencode:gish
# BEGIN ai.skillz: direct:symlink:claude:resolve-conflicts
/.claude/skills/resolve-conflicts
# END ai.skillz: direct:symlink:claude:resolve-conflicts
# BEGIN ai.skillz: direct:symlink:opencode:resolve-conflicts
/.opencode/skills/resolve-conflicts
# END ai.skillz: direct:symlink:opencode:resolve-conflicts
# BEGIN ai.skillz: direct:symlink:claude:git-mgmt
/.claude/skills/git-mgmt
# END ai.skillz: direct:symlink:claude:git-mgmt
# BEGIN ai.skillz: direct:symlink:opencode:git-mgmt
/.opencode/skills/git-mgmt
# END ai.skillz: direct:symlink:opencode:git-mgmt
# BEGIN ai.skillz: runtime:open-wkt
/wkts/
# END ai.skillz: runtime:open-wkt
# BEGIN ai.skillz: direct:symlink:claude:open-wkt
/.claude/skills/open-wkt
# END ai.skillz: direct:symlink:claude:open-wkt
# BEGIN ai.skillz: direct:symlink:opencode:open-wkt
/.opencode/skills/open-wkt
# END ai.skillz: direct:symlink:opencode:open-wkt
# BEGIN ai.skillz: direct:symlink:claude:close-wkt
/.claude/skills/close-wkt
# END ai.skillz: direct:symlink:claude:close-wkt
# BEGIN ai.skillz: direct:symlink:opencode:close-wkt
/.opencode/skills/close-wkt
# END ai.skillz: direct:symlink:opencode:close-wkt
# BEGIN ai.skillz: runtime:code-review
.ai/code-review/reports/
# END ai.skillz: runtime:code-review
# BEGIN ai.skillz: direct:symlink:claude:code-review
/.claude/skills/code-review
# END ai.skillz: direct:symlink:claude:code-review
# BEGIN ai.skillz: direct:symlink:opencode:code-review
/.opencode/skills/code-review
# END ai.skillz: direct:symlink:opencode:code-review
# BEGIN ai.skillz: direct:symlink:claude:code-nav-refs
/.claude/skills/code-nav-refs
# END ai.skillz: direct:symlink:claude:code-nav-refs
# BEGIN ai.skillz: direct:symlink:opencode:code-nav-refs
/.opencode/skills/code-nav-refs
# END ai.skillz: direct:symlink:opencode:code-nav-refs
# BEGIN ai.skillz: runtime:code-review-changes
.claude/review_context.md
.claude/review_regression.md
.claude/review_replies/
# END ai.skillz: runtime:code-review-changes
# BEGIN ai.skillz: direct:symlink:claude:code-review-changes
/.claude/skills/code-review-changes
# END ai.skillz: direct:symlink:claude:code-review-changes
# BEGIN ai.skillz: direct:symlink:opencode:code-review-changes
/.opencode/skills/code-review-changes
# END ai.skillz: direct:symlink:opencode:code-review-changes
# BEGIN ai.skillz: runtime:commit-msg
.claude/skills/commit-msg/msgs/
.claude/git_commit_msg_LATEST.md
# END ai.skillz: runtime:commit-msg
# BEGIN ai.skillz: direct:symlink:claude:commit-msg
/.claude/skills/commit-msg/SKILL.md
# END ai.skillz: direct:symlink:claude:commit-msg
# BEGIN ai.skillz: direct:symlink:opencode:commit-msg
/.opencode/skills/commit-msg/SKILL.md
# END ai.skillz: direct:symlink:opencode:commit-msg
# BEGIN ai.skillz: direct:symlink:claude:commit-plan
/.claude/skills/commit-plan
# END ai.skillz: direct:symlink:claude:commit-plan
# BEGIN ai.skillz: direct:symlink:opencode:commit-plan
/.opencode/skills/commit-plan
# END ai.skillz: direct:symlink:opencode:commit-plan
# BEGIN ai.skillz: direct:symlink:claude:dep-supersede-scan
/.claude/skills/dep-supersede-scan
# END ai.skillz: direct:symlink:claude:dep-supersede-scan
# BEGIN ai.skillz: direct:symlink:opencode:dep-supersede-scan
/.opencode/skills/dep-supersede-scan
# END ai.skillz: direct:symlink:opencode:dep-supersede-scan
# BEGIN ai.skillz: direct:symlink:claude:harness-perf
/.claude/skills/harness-perf
# END ai.skillz: direct:symlink:claude:harness-perf
# BEGIN ai.skillz: direct:symlink:opencode:harness-perf
/.opencode/skills/harness-perf
# END ai.skillz: direct:symlink:opencode:harness-perf
# BEGIN ai.skillz: direct:symlink:claude:inter-skill-review
/.claude/skills/inter-skill-review
# END ai.skillz: direct:symlink:claude:inter-skill-review
# BEGIN ai.skillz: direct:symlink:opencode:inter-skill-review
/.opencode/skills/inter-skill-review
# END ai.skillz: direct:symlink:opencode:inter-skill-review
# BEGIN ai.skillz: direct:symlink:claude:opencode-cleaning
/.claude/skills/opencode-cleaning
# END ai.skillz: direct:symlink:claude:opencode-cleaning
# BEGIN ai.skillz: direct:symlink:opencode:opencode-cleaning
/.opencode/skills/opencode-cleaning
# END ai.skillz: direct:symlink:opencode:opencode-cleaning
# BEGIN ai.skillz: direct:symlink:claude:plan-io
/.claude/skills/plan-io
# END ai.skillz: direct:symlink:claude:plan-io
# BEGIN ai.skillz: direct:symlink:opencode:plan-io
/.opencode/skills/plan-io
# END ai.skillz: direct:symlink:opencode:plan-io
# BEGIN ai.skillz: runtime:pr-msg
.claude/skills/pr-msg/msgs/
.claude/skills/pr-msg/pr_msg_LATEST.md
# END ai.skillz: runtime:pr-msg
# BEGIN ai.skillz: direct:symlink:claude:pr-msg
/.claude/skills/pr-msg/SKILL.md
/.claude/skills/pr-msg/references
/.claude/skills/pr-msg/scripts
# END ai.skillz: direct:symlink:claude:pr-msg
# BEGIN ai.skillz: direct:symlink:opencode:pr-msg
/.opencode/skills/pr-msg/SKILL.md
/.opencode/skills/pr-msg/references
/.opencode/skills/pr-msg/scripts
# END ai.skillz: direct:symlink:opencode:pr-msg
# BEGIN ai.skillz: direct:symlink:claude:prompt-io
/.claude/skills/prompt-io
# END ai.skillz: direct:symlink:claude:prompt-io
# BEGIN ai.skillz: direct:symlink:opencode:prompt-io
/.opencode/skills/prompt-io
# END ai.skillz: direct:symlink:opencode:prompt-io
# BEGIN ai.skillz: direct:symlink:claude:py-codestyle
/.claude/skills/py-codestyle
# END ai.skillz: direct:symlink:claude:py-codestyle
# BEGIN ai.skillz: direct:symlink:opencode:py-codestyle
/.opencode/skills/py-codestyle
# END ai.skillz: direct:symlink:opencode:py-codestyle
# BEGIN ai.skillz: runtime:taken-export
.ai/taken/exports/
# END ai.skillz: runtime:taken-export
# BEGIN ai.skillz: direct:symlink:claude:taken-export
/.claude/skills/taken-export
# END ai.skillz: direct:symlink:claude:taken-export
# BEGIN ai.skillz: direct:symlink:opencode:taken-export
/.opencode/skills/taken-export
# END ai.skillz: direct:symlink:opencode:taken-export
# BEGIN ai.skillz: direct:symlink:claude:yt-url-lookup
/.claude/skills/yt-url-lookup
# END ai.skillz: direct:symlink:claude:yt-url-lookup
# BEGIN ai.skillz: direct:symlink:opencode:yt-url-lookup
/.opencode/skills/yt-url-lookup
# END ai.skillz: direct:symlink:opencode:yt-url-lookup
# BEGIN ai.skillz: direct:symlink:opencode:command:gish
/.opencode/commands/gish.md
# END ai.skillz: direct:symlink:opencode:command:gish
# BEGIN ai.skillz: direct:symlink:opencode:command:resolve-conflicts
/.opencode/commands/resolve-conflicts.md
# END ai.skillz: direct:symlink:opencode:command:resolve-conflicts
# BEGIN ai.skillz: direct:symlink:opencode:command:git-mgmt
/.opencode/commands/git-mgmt.md
# END ai.skillz: direct:symlink:opencode:command:git-mgmt
# BEGIN ai.skillz: direct:symlink:opencode:command:open-wkt
/.opencode/commands/open-wkt.md
# END ai.skillz: direct:symlink:opencode:command:open-wkt
# BEGIN ai.skillz: direct:symlink:opencode:command:close-wkt
/.opencode/commands/close-wkt.md
# END ai.skillz: direct:symlink:opencode:command:close-wkt
# BEGIN ai.skillz: direct:symlink:opencode:command:code-review
/.opencode/commands/code-review.md
# END ai.skillz: direct:symlink:opencode:command:code-review
# BEGIN ai.skillz: direct:symlink:opencode:command:code-review-changes
/.opencode/commands/code-review-changes.md
# END ai.skillz: direct:symlink:opencode:command:code-review-changes
# BEGIN ai.skillz: direct:symlink:opencode:command:commit-msg
/.opencode/commands/commit-msg.md
# END ai.skillz: direct:symlink:opencode:command:commit-msg
# BEGIN ai.skillz: direct:symlink:opencode:command:commit-plan
/.opencode/commands/commit-plan.md
# END ai.skillz: direct:symlink:opencode:command:commit-plan
# BEGIN ai.skillz: direct:symlink:opencode:command:dep-supersede-scan
/.opencode/commands/dep-supersede-scan.md
# END ai.skillz: direct:symlink:opencode:command:dep-supersede-scan
# BEGIN ai.skillz: direct:symlink:opencode:command:harness-perf
/.opencode/commands/harness-perf.md
# END ai.skillz: direct:symlink:opencode:command:harness-perf
# BEGIN ai.skillz: direct:symlink:opencode:command:opencode-cleaning
/.opencode/commands/opencode-cleaning.md
# END ai.skillz: direct:symlink:opencode:command:opencode-cleaning
# BEGIN ai.skillz: direct:symlink:opencode:command:pr-msg
/.opencode/commands/pr-msg.md
# END ai.skillz: direct:symlink:opencode:command:pr-msg
# BEGIN ai.skillz: direct:symlink:opencode:command:taken-export
/.opencode/commands/taken-export.md
# END ai.skillz: direct:symlink:opencode:command:taken-export
# BEGIN ai.skillz: direct:symlink:opencode:command:yt-url-lookup
/.opencode/commands/yt-url-lookup.md
# END ai.skillz: direct:symlink:opencode:command:yt-url-lookup

View File

@ -1,37 +0,0 @@
---
model: gpt-5.6-sol
service: opencode
session: tractor-addr-unpacking
timestamp: 2026-08-21T05:20:52Z
git_ref: 3690e43a
scope: config
substantive: true
raw_file: 20260821T052052Z_3690e43a_prompt_io.raw.md
---
## Prompt
The human asked for a main-first patch using an off-the-shelf pytest
plugin to cope with tractor's changing macOS CI flakes without mixing
that mitigation into PR #505.
## Response summary
Added `pytest-rerunfailures` to tractor's testing dependencies and
configured the GitHub Actions matrix to retry failures only on macOS.
Linux and Windows remain strict first-attempt runs, while persistent
macOS failures still fail after two visible reruns.
## Files changed
- `.github/workflows/ci.yml` - macOS-only pytest rerun budget.
- `pyproject.toml` - testing plugin dependency and rationale.
- `uv.lock` - resolved `pytest-rerunfailures` package metadata.
## Human edits
The human selected a main-first mitigation after PR #505 failed two
different macOS tests on consecutive runs and required the change to
remain an incremental patch with its own commit plan. The agent
implemented and verified that direction; no direct manual source
edits were observed.

View File

@ -1,25 +0,0 @@
---
model: gpt-5.6-sol
service: opencode
timestamp: 2026-08-21T05:20:52Z
git_ref: 3690e43a
diff_cmd: git diff HEAD~1..HEAD
---
# Raw output - retry flaky macOS CI tests
The human requested an off-the-shelf pytest plugin patch suitable
for landing directly on tractor `main` after PR #505's macOS job
failed two different timing-sensitive tests on consecutive runs.
> `git diff HEAD~1..HEAD -- .github/workflows/ci.yml pyproject.toml uv.lock`
Added the pytest-dev-maintained `pytest-rerunfailures` plugin and
gave only the macOS matrix leg two reruns with a one-second delay.
Linux and Windows receive a zero retry budget; deterministic macOS
failures still fail after the final attempt and reruns remain visible
in pytest output.
The lockfile is current, actionlint passed, all 471 tests collected,
and the four tests covering both observed PR #505 failure areas
passed with the rerun plugin enabled.

View File

@ -1,49 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-22T02:25:26Z
git_ref: eb3c99c9
scope: config
substantive: true
raw_file: 20260822T022526Z_5562fd9a_prompt_io.raw.md
---
## Prompt
Perform a full Tractor repository scan for related `ai.skillz` work,
then correct the run-tests landing branch, prune unrelated `.gitignore`
additions, preserve only the focused migration, and provide canonical
deployment commands.
## Response summary
Audited all local branches, worktrees, affected-path history, deployment
state, canonical skill dependencies, and current Tractor harness behavior.
Corrected the local test reference where it overstated cleanup safety or
omitted current environment, platform, debugger, timeout, and CI details.
Narrowed the correction commit to three managed `run-tests` deployment
blocks. A later dedicated commit records the complete generated `ai.skillz`
deployment state.
## Files changed
- `.claude/skills/run-tests/test-harness-reference.md` - correct the
project-specific test and cleanup contract.
- `.gitignore` - narrow the correction commit before the later dedicated
deployment-state expansion.
- `ai/prompt-io/opencode/20260822T022526Z_5562fd9a_prompt_io.md` - record
the migration review provenance.
- `ai/prompt-io/opencode/20260822T022526Z_5562fd9a_prompt_io.raw.md` -
preserve the unedited response record.
## Human edits
The human required an existing-work scan after duplicate implementation
was discovered, approved correcting the landing branch during PR #481
review, and directed removal or reconciliation of unrelated ignore rules.
During PR #510 review, the human required the stackscope and shared-memory
safety clarifications and immutable provenance pointers before landing. No
direct source-line edits were made by the human. Copilot review then prompted
the human to require explicit canonical deployment instructions and clarify
the later `.gitignore` expansion.

View File

@ -1,24 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-22T02:25:26Z
git_ref: eb3c99c9
diff_cmd: git diff eb3c99c9^..eb3c99c9
---
Audit Tractor's repository-wide `ai.skillz` state and correct the
`wkt/ai_skillz_run_tests_landing` migration before landing it.
> `git diff eb3c99c9^..eb3c99c9 -- .claude/skills/run-tests/test-harness-reference.md`
The harness reference is corrected against current Tractor behavior:
environment overrides, CI/platform matrices, randomized registry
addresses, interactive debugger capture, timeout mechanisms, stackscope
artifacts, reaper platform limits, dry-run expansion, and exact
change-to-test mappings.
> `git diff eb3c99c9^..eb3c99c9 -- .gitignore`
Unrelated Taken and broad OpenCode command ignore additions are removed;
only the managed canonical `run-tests` link and command blocks remain in
this migration.

View File

@ -1,47 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: pr490-landing-review
timestamp: 2026-08-28T01:42:47Z
git_ref: wkt/fix_broadcast_lag_count
scope: code
substantive: true
raw_file: 20260828T014247Z_d06ee34b_prompt_io.raw.md
---
## Prompt
Prime PR #490 for landing, review its concurrent broadcast changes,
fix valid findings, and keep the landing pass focused.
## Response summary
Rejected an invalid mutable-default finding under `msgspec.Struct`
semantics. Found and fixed a close-time race where concurrent receives
on one `BroadcastReceiver` could overwrite its sole cancellation scope
and strand an earlier waiter. The receiver now rejects overlapping
calls with `trio.BusyResourceError`. Public docs now define each
receiver as one logical subscription cursor and require one receiver
per concurrent lossless consumer. Busy errors identify the active Trio
task which owns the receiver.
## Files changed
- `tractor/trionics/_broadcast.py` - reject concurrent receives on one
receiver handle.
- `tests/test_task_broadcasting.py` - prove the second call fails while
the original receive remains usable.
- `docs/guide/streaming.rst` - document per-task receiver ownership and
explicit lag behavior.
## Human edits
The human selected PR #490 as the next landing target, rebased it onto
current `main`, directed the agent to continue the landing pass, and
required a faster, tightly scoped execution after the review expanded.
The human then confirmed the per-receiver guard and required the public
method, type, and surrounding guide to emphasize lossless per-consumer
semantics. The human also required the regression's checkpoint loop to
explain its synchronization and the busy error to report the active
consumer task. The agent applied the source, test, and documentation
changes; no direct manual edits to these files were identified.

View File

@ -1,56 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-28T01:42:47Z
git_ref: wkt/fix_broadcast_lag_count
diff_cmd: git diff
---
## Prompt
Get PR #490 primed and landed. Continue after rebasing, but keep the
review and fix pass focused.
## Response
The mutable `{}` default is safe because `msgspec.Struct` creates a
fresh dictionary per instance. A separate concurrency defect remains:
two tasks can call `receive()` through one `BroadcastReceiver`, and the
second call overwrites the only `_wait_scope`. Closing that receiver can
then leave the first task blocked indefinitely.
> `git diff -- tractor/trionics/_broadcast.py tests/test_task_broadcasting.py`
Added a per-receiver in-progress guard that raises
`trio.BusyResourceError` for an overlapping receive. Added a regression
test that blocks the first receive, verifies the second call is
rejected, and releases the source to prove the first call still works.
The complete broadcast test file passes: 28 tests in 4.57 seconds.
## Follow-up prompt
Clarify that the guard is per receiver, preserves lossless consumer
semantics, and document the contract on the public method, type, and
surrounding guide.
## Follow-up response
> `git diff -- tractor/trionics/_broadcast.py docs/guide/streaming.rst`
Documented that each `BroadcastReceiver` owns one logical cursor, each
concurrent consumer needs its own subscribed receiver, overlapping
calls on one handle raise `BusyResourceError`, and strict lag handling
never skips values silently.
## Second follow-up prompt
Explain that the polling loop waits until the background receive is
blocked before making the concurrent call, and include the first
consumer task's runtime information in the busy exception.
## Second follow-up response
Replaced the boolean guard with the active `trio.lowlevel.Task`, added
its name and representation to `BusyResourceError`, named the fixture
task for a deterministic assertion, and documented the checkpoint-loop
interleaving directly above the poll.

View File

@ -171,14 +171,6 @@ keeps pace with the *fastest* subscriber; a task falling more
than the buffered window behind has its next receive raise than the buffered window behind has its next receive raise
``tractor.trionics.Lagged`` to say it lost data. ``tractor.trionics.Lagged`` to say it lost data.
Each ``BroadcastReceiver`` is one logical subscription cursor, so
give every concurrent consumer task its own receiver. Overlapping
``receive()`` calls on the same handle raise
``trio.BusyResourceError`` instead of racing that cursor. In strict
mode values are never skipped silently: the consumer either reads
each retained value in sequence or receives an explicit ``Lagged``
error after exceeding the buffer window.
Pass ``raise_on_lag=False`` when a consumer may drop old values and Pass ``raise_on_lag=False`` when a consumer may drop old values and
resume from the oldest retained item instead. The receiver logs the resume from the oldest retained item instead. The receiver logs the
overrun rather than raising. Each child subscription chooses its own overrun rather than raising. Each child subscription chooses its own

View File

@ -100,9 +100,6 @@ testing = [
# cleanup utility (xplatform `Process.memory_maps`, # cleanup utility (xplatform `Process.memory_maps`,
# `Process.open_files`). # `Process.open_files`).
"psutil>=7.0.0", "psutil>=7.0.0",
# rerun timing-sensitive macOS CI nodes without hiding a
# deterministic failure after the final attempt.
"pytest-rerunfailures>=16.6,<17",
] ]
repl = [ repl = [
"pyperclip>=1.9.0", "pyperclip>=1.9.0",

View File

@ -1018,52 +1018,6 @@ def test_closing_non_owner_preserves_source_wait() -> None:
trio.run(main) trio.run(main)
def test_concurrent_receive_raises_busy() -> None:
'''
Reject concurrent receives on one broadcast handle.
A receiver stores one private peer-wait cancellation scope. If two
tasks receive through the same handle, the second task can replace
that scope and prevent `BroadcastReceiver.aclose()` from waking the
first task. Block one task in the shared source receive, then prove
a second call raises `BusyResourceError` before it can mutate any
per-receiver wait state. Releasing the source proves the original
receive remains usable.
'''
async def main() -> None:
tx, rx = trio.open_memory_channel(1)
brx = broadcast_receiver(rx, 1)
values: list[int] = []
async def receive() -> None:
values.append(await brx.receive())
async with trio.open_nursery() as nursery:
nursery.start_soon(
receive,
name='first broadcast consumer',
)
# Synchronize with the background task after it blocks in
# the shared source `.receive()`, ensuring the next call is
# concurrent with an already-active receive on this handle.
while brx._state.recv_ready is None:
await trio.lowlevel.checkpoint()
with pytest.raises(
trio.BusyResourceError,
match='first broadcast consumer',
):
await brx.receive()
await tx.send(1)
assert values == [1]
trio.run(main)
@pytest.mark.parametrize( @pytest.mark.parametrize(
'first_outcome', 'first_outcome',
[ [

View File

@ -184,18 +184,12 @@ class BroadcastState(Struct):
class BroadcastReceiver(ReceiveChannel): class BroadcastReceiver(ReceiveChannel):
''' '''
One logical subscriber to a shared receive-channel broadcast. A memory receive channel broadcaster which is non-lossy for
the fastest consumer.
Each instance owns one sequence cursor. Additional consumer tasks Additional consumer tasks can receive all produced values by
must call `.subscribe()` and receive through the new instance it registering with ``.subscribe()`` and receiving from the new
yields. Overlapping `.receive()` calls on the same instance raise instance it delivers.
`trio.BusyResourceError` rather than racing that cursor or the
receiver's close-cancellation state.
A strict subscriber reads each retained value in sequence. Falling
behind the retention window raises `Lagged` instead of silently
losing values; `raise_on_lag=False` explicitly opts into dropping
displaced values.
''' '''
def __init__( def __init__(
@ -225,7 +219,6 @@ class BroadcastReceiver(ReceiveChannel):
self._closed: bool = False self._closed: bool = False
self._raise_on_lag = raise_on_lag self._raise_on_lag = raise_on_lag
self._wait_scope: trio.CancelScope|None = None self._wait_scope: trio.CancelScope|None = None
self._receive_task: trio.lowlevel.Task|None = None
def receive_nowait( def receive_nowait(
self, self,
@ -441,30 +434,6 @@ class BroadcastReceiver(ReceiveChannel):
state.recv_scope = None state.recv_scope = None
async def receive(self) -> ReceiveType: async def receive(self) -> ReceiveType:
'''
Receive the next value for this subscriber's sequence cursor.
Only one task may receive through this instance at a time. Use
`.subscribe()` to give each concurrent consumer its own cursor
and loss/lag policy. `trio.BusyResourceError` identifies the
task which owns an already-active receive.
'''
if receive_task := self._receive_task:
raise trio.BusyResourceError(
'another task is already receiving from this '
'`BroadcastReceiver`\n'
f'active receive task: {receive_task.name!r}\n'
f'{receive_task!r}'
)
self._receive_task = trio.lowlevel.current_task()
try:
return await self._receive()
finally:
self._receive_task = None
async def _receive(self) -> ReceiveType:
key = self.key key = self.key
state = self._state state = self._state
@ -549,12 +518,11 @@ class BroadcastReceiver(ReceiveChannel):
) -> AsyncIterator[BroadcastReceiver]: ) -> AsyncIterator[BroadcastReceiver]:
''' '''
Create a receiver with its own logical subscription cursor. Subscribe for values from this broadcast receiver.
The new `BroadcastReceiver` is registered against the shared Returns a new ``BroadCastReceiver`` which is registered for and
source and receives every retained value in sequence. Give each pulls data from a clone of the original
concurrent consumer task its own receiver instead of sharing ``trio.abc.ReceiveChannel`` provided at creation.
one instance across overlapping `.receive()` calls.
''' '''
if self._closed: if self._closed:

17
uv.lock
View File

@ -782,19 +782,6 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/8b/5a/ba30a81239b909821b3153e303e7def45178bf353da4f72380e6c5e8793b/pytest-9.1.0-py3-none-any.whl", hash = "sha256:8ebb0e7888bdf2bdfc602ec51f8f62d50200af37356c74e503c79a94f5c81f32", size = 386453, upload-time = "2026-06-13T18:52:44.045Z" }, { url = "https://files.pythonhosted.org/packages/8b/5a/ba30a81239b909821b3153e303e7def45178bf353da4f72380e6c5e8793b/pytest-9.1.0-py3-none-any.whl", hash = "sha256:8ebb0e7888bdf2bdfc602ec51f8f62d50200af37356c74e503c79a94f5c81f32", size = 386453, upload-time = "2026-06-13T18:52:44.045Z" },
] ]
[[package]]
name = "pytest-rerunfailures"
version = "16.6"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "packaging" },
{ name = "pytest" },
]
sdist = { url = "https://files.pythonhosted.org/packages/ed/63/0114e45d4b2fcd5f6297dac655c067b47de28be4d33e088f200f0f2c4c28/pytest_rerunfailures-16.6.tar.gz", hash = "sha256:29dbfee46f542073c888e0ed4e81c51e15b9096f49a299eb1a759629c601684a", size = 42806, upload-time = "2026-08-17T07:11:00.447Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/5d/5e/1e994889673d7a0da11651f17ef789b6c83bfe349f29f871873dd3802445/pytest_rerunfailures-16.6-py3-none-any.whl", hash = "sha256:6af2d1ebd6e5cb79666ac408942cd6a0672a49fefd8523664770718560795e13", size = 19137, upload-time = "2026-08-17T07:10:59.121Z" },
]
[[package]] [[package]]
name = "pytest-timeout" name = "pytest-timeout"
version = "2.4.0" version = "2.4.0"
@ -1144,7 +1131,6 @@ dev = [
{ name = "psutil" }, { name = "psutil" },
{ name = "pyperclip" }, { name = "pyperclip" },
{ name = "pytest" }, { name = "pytest" },
{ name = "pytest-rerunfailures" },
{ name = "pytest-timeout" }, { name = "pytest-timeout" },
{ name = "stackscope" }, { name = "stackscope" },
{ name = "typing-extensions" }, { name = "typing-extensions" },
@ -1184,7 +1170,6 @@ testing = [
{ name = "pexpect" }, { name = "pexpect" },
{ name = "psutil" }, { name = "psutil" },
{ name = "pytest" }, { name = "pytest" },
{ name = "pytest-rerunfailures" },
{ name = "pytest-timeout" }, { name = "pytest-timeout" },
] ]
@ -1210,7 +1195,6 @@ dev = [
{ name = "psutil", specifier = ">=7.0.0" }, { name = "psutil", specifier = ">=7.0.0" },
{ name = "pyperclip", specifier = ">=1.9.0" }, { name = "pyperclip", specifier = ">=1.9.0" },
{ name = "pytest", specifier = ">=9.0.3" }, { name = "pytest", specifier = ">=9.0.3" },
{ name = "pytest-rerunfailures", specifier = ">=16.6,<17" },
{ name = "pytest-timeout", specifier = ">=2.3" }, { name = "pytest-timeout", specifier = ">=2.3" },
{ name = "stackscope", specifier = ">=0.2.2,<0.3" }, { name = "stackscope", specifier = ">=0.2.2,<0.3" },
{ name = "typing-extensions", specifier = ">=4.14.1" }, { name = "typing-extensions", specifier = ">=4.14.1" },
@ -1242,7 +1226,6 @@ testing = [
{ name = "pexpect", specifier = ">=4.9.0,<5" }, { name = "pexpect", specifier = ">=4.9.0,<5" },
{ name = "psutil", specifier = ">=7.0.0" }, { name = "psutil", specifier = ">=7.0.0" },
{ name = "pytest", specifier = ">=9.0.3" }, { name = "pytest", specifier = ">=9.0.3" },
{ name = "pytest-rerunfailures", specifier = ">=16.6,<17" },
{ name = "pytest-timeout", specifier = ">=2.3" }, { name = "pytest-timeout", specifier = ">=2.3" },
] ]