Compare commits

..

6 Commits

Author SHA1 Message Date
Gud Boi d1edecbf60 Clean broadcast cancellation diagnostics
Bound `BroadcastState.cancelled` entries to receiver progress,
terminal state and resource lifetime instead of retaining completed
`Task`s indefinitely.

Deats,
- make EOC durable so awakened peers never re-enter a closed source.
- close root broadcasters during explicit `MsgStream` and
  `LinkedTaskChannel` teardown without re-entrant EOC closure or
  breaking `MsgStream.aclose()` overrides.
- reject non-positive fan-out retention capacity before constructing
  an unusable zero-length queue.
- cover child/root cancellation cleanup, terminal peer wakeups,
  wrapper teardown, subclass compatibility and zero-buffer rejection.

Prompt-IO: ai/prompt-io/opencode/20260813T181901Z_a2e0df4b_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-13 15:49:05 -04:00
Gud Boi a2e0df4ba1 Expose stream subscriber lag policy
`MsgStream.subscribe()` and `LinkedTaskChannel.subscribe()` omitted
`BroadcastReceiver.raise_on_lag`, forcing downstream consumers to
mutate a private receiver attribute when overruns were acceptable.

Add `raise_on_lag` to both public wrappers. The first subscription
sets the irreversible root broadcaster's policy, while every child
selects its own strict or warn/drop/resume behavior independently.

Document both fan-out APIs. Cover policy forwarding plus real IPC
and infected-asyncio paths.

Prompt-IO: ai/prompt-io/opencode/20260812T213117Z_51185487_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-13 00:37:07 -04:00
Gud Boi 511854870b Isolate `BroadcastReceiver.aclose()` wakeups
Closing any subscriber set the shared `recv_ready` event, even when
another receiver owned the source read. Waiting peers then looped
until an idle source produced another value.

Give each receiver private source-read and peer-wait cancellation
scopes. Closing a waiting peer interrupts only that peer. Closing
the source owner discards post-close source outcomes and then wakes
peers for a clean ownership handoff.

Keep outer task cancellation as `trio.Cancelled`; only explicit
receiver close maps either private scope's cancellation to
`ClosedResourceError`. Assert that scope cancellation implies the
receiver is closed and document the source-owner key check.

Prompt-IO: ai/prompt-io/opencode/20260812T150027Z_c2a6ccef_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-12 17:22:59 -04:00
Gud Boi c2a6ccefd0 Wake broadcast peers on shared receive failures
Only `EndOfChannel` and direct cancellation woke tasks waiting
behind the subscriber which owned the underlying receive. Any other
failure cleared `BroadcastState.recv_ready` while peers remained
blocked on its unreachable event.

Publish ordinary receive exceptions as terminal broadcast state.
The owner keeps the original failure while peers drain retained
values and then raise `BroadcastReceiveError` from that cause. Also
wake peers on process-control exits without retaining them as state.

Document the public owner/peer contract and cover current, late and
control-flow subscribers with deterministic bounded regressions.

Prompt-IO: ai/prompt-io/opencode/20260812T030608Z_1095e7f7_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-12 00:51:22 -04:00
Gud Boi 1095e7f710 Fix `BroadcastState.statistics()` queue counts
`BroadcastState.subs` stores each receiver's next unread deque
index, but `.statistics()` exposed that index as a queue length. A
caught-up receiver looked correct by accident while every queued
count was one short.

Convert cursors to retained, receivable counts and clamp lagged
receivers to the current queue length. Also avoid deprecated
`trio.Event` truthiness when reporting waiter counts.

Cover caught-up, queued, lagged and real-event states using actual
broadcast sends and receives.

Prompt-IO: ai/prompt-io/opencode/20260812T012324Z_06c4af17_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-11 22:49:23 -04:00
Gud Boi 06c4af17e4 Fix `BroadcastReceiver` lag counts
`BroadcastReceiver.receive_nowait()` treated `seq` as a deque index
but subtracted `BroadcastState.maxlen` without counting the first
invalid index. A one-slot queue thus claimed it dropped zero values
after its subscriber missed one.

Include that first displaced value in the count. Preserve the
existing Tokio-style reset to the oldest retained item.

Also, cover exact loss reporting and recovery for one- and
three-slot retention windows.

Prompt-IO: ai/prompt-io/opencode/20260811T233833Z_7cbd64ee_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-11 19:45:53 -04:00
151 changed files with 952 additions and 8536 deletions

View File

@ -91,10 +91,6 @@ jobs:
name: '${{ matrix.os }} Python${{ matrix.python-version }} spawn_backend=${{ matrix.spawn_backend }} tpt_proto=${{ matrix.tpt_proto }}' name: '${{ matrix.os }} Python${{ matrix.python-version }} spawn_backend=${{ matrix.spawn_backend }} tpt_proto=${{ matrix.tpt_proto }}'
timeout-minutes: 16 timeout-minutes: 16
runs-on: ${{ matrix.os }} runs-on: ${{ matrix.os }}
# Windows support is nascent: its full test suite remains
# informational, while setup and the `import tractor` smoke below
# are hard signals. Promote the test step to required once the
# suite is green.
strategy: strategy:
fail-fast: false fail-fast: false
@ -102,7 +98,6 @@ jobs:
os: [ os: [
ubuntu-latest, ubuntu-latest,
macos-latest, macos-latest,
windows-latest,
] ]
python-version: [ python-version: [
'3.13', '3.13',
@ -123,10 +118,10 @@ jobs:
'tcp', 'tcp',
'uds', 'uds',
] ]
# https://github.com/orgs/community/discussions/26253#discussioncomment-3250989
exclude: exclude:
# UDS is POSIX-only; Windows has no `AF_UNIX` so the # don't do UDS run on macOS (for now)
# backend is intentionally unavailable there. - os: macos-latest
- os: windows-latest
tpt_proto: 'uds' tpt_proto: 'uds'
steps: steps:
@ -155,14 +150,7 @@ jobs:
- name: List deps tree - name: List deps tree
run: uv tree run: uv tree
# hard signal for the Windows import-safety fix: `import
# tractor` must succeed everywhere, and `HAS_UDS` reflects
# platform capability (False on Windows, True on POSIX).
- name: 'Smoke: import tractor'
run: uv run python -c "import sys; import tractor; from tractor.ipc._uds import HAS_UDS; assert sys.platform != 'win32' or not HAS_UDS; print('import tractor OK | HAS_UDS=', HAS_UDS)"
- name: Run tests - name: Run tests
continue-on-error: ${{ matrix.os == 'windows-latest' }}
run: > run: >
uv run uv run
pytest pytest

View File

@ -1,79 +0,0 @@
---
model: claude-fable-5
service: claude
session: f6c84722-471a-4458-9a80-e453fea9029f
timestamp: 2026-07-02T15:42:55Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260702T154255Z_65bf9df5_prompt_io.raw.md
---
## Prompt
Driver prompt file `ai/prompt-io/prompts/issue_477.md`:
> attempt to resolve
> https://github.com/goodboy/tractor/issues/477
> do it with /open-wkt.
(plus a hard stop-for-human-review deadline of 12:50PM
EST the same day)
Issue #477 asks to factor `ActorNursery.run_in_actor()`
(and possibly `Portal.run()`) out of the nursery
internals into a new `tractor.to_actor` wrapper
subpackage of "higher level one shot" single-remote-task
APIs, adopting the `trio.to_thread`/`anyio.to_process`
parlance, so that error collection/propagation moves up
into the caller's local `trio` scope and the nursery's
spawn machinery can eventually drop the
`._ria_nursery` coupling.
## Response summary
First-cut `tractor.to_actor` subpkg delivering the
one-shot API composed purely from the existing
daemon-spawn + portal primitives (`start_actor()` +
`Portal.run()` + `Portal.cancel_actor()`), leaving the
legacy `.run_in_actor()` machinery untouched (formal
deprecation deferred until in-repo usage migrates):
- `to_actor.run(fn, **fn_kwargs) -> Any`: spawn a
subactor, schedule `fn` as its lone remote task, wait
on and return its result, ALWAYS reaping the subactor
(shield-safe `finally`). Remote errors raise in the
caller's task as boxed `RemoteActorError`s.
- placement variants: `portal=` reuses a running actor
(no spawn/reap), `an=` spawns from a caller-managed
actor-nursery, neither opens a call-scoped private
`open_nursery()` (implicitly booting the runtime,
configurable via `runtime_kwargs`).
- fail-fast validation before any spawn: non-streaming
async fn required; `portal=`/`an=` mutually
exclusive; `runtime_kwargs` rejected alongside any
placement opt.
- `run_in_actor()` TODO/docstring now cross-reference
the successor API.
## Files changed
- `tractor/to_actor/__init__.py` — new subpkg,
re-exports `run`
- `tractor/to_actor/_api.py``run()` +
`_invoke_in_subactor()` + `_validate_one_shot_fn()`
- `tractor/__init__.py` — top-level `to_actor`
re-export
- `tractor/runtime/_supervise.py` — comment/docstring
pointers from `run_in_actor()` to the successor
- `tests/test_to_actor.py` — 11-test suite covering
all placement variants, error relay, the concurrent
worker-pool-ish pattern and arg validation
- `examples/parallelism/to_actor_one_shots.py`
runnable demo (auto-collected by
`test_docs_examples.py`)
## Human edits
None yet — pending human review (work paused before the
12:50PM EST deadline per the driver prompt).

View File

@ -1,100 +0,0 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T15:42:55Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/to_actor_subpkg
---
# Raw AI output (diff-ref mode)
All generated code is committed on the
`wkt/to_actor_subpkg` branch; per diff-ref mode each
file's verbatim content is reachable via the pointers
below rather than duplicated here.
## Generated files
> `git diff main..wkt/to_actor_subpkg -- tractor/to_actor/__init__.py`
New subpackage init: module docstring establishing the
`trio.to_thread`/`anyio.to_process` "run it over there"
parlance for actors, plus the single public re-export
`run as run` from `._api`.
> `git diff main..wkt/to_actor_subpkg -- tractor/to_actor/_api.py`
The one-shot invocation impl, composed entirely from the
lower level daemon-spawn + portal primitives as
prescribed by issue #477:
- `_validate_one_shot_fn()`: the `Portal.run()`
non-streaming-async-fn constraint checked up-front,
before any subactor is spawned.
- `_invoke_in_subactor()`: `an.start_actor()` ->
`Portal.run()` -> always-reap via
`Portal.cancel_actor()` in a `finally` (the cancel
req's bounded wait is internally shielded so the reap
also runs under caller-scope cancellation).
- `run()`: the public API. Placement options:
`portal=` (reuse a running actor, no spawn/reap),
`an=` (spawn from a caller-managed nursery), or
neither (private `open_nursery()` scoped to the call,
implicitly booting the runtime when needed, tunable
via pass-through `runtime_kwargs`). Spawn opts mirror
`ActorNursery.start_actor()`; `**fn_kwargs` are
relayed to the remote task. Errors raise in the
caller's task as boxed `RemoteActorError`s.
`runtime_kwargs` alongside any placement opt is a
hard `ValueError`, never silently ignored.
> `git diff main..wkt/to_actor_subpkg -- tractor/__init__.py`
Top-level `from . import to_actor as to_actor`
re-export.
> `git diff main..wkt/to_actor_subpkg -- tractor/runtime/_supervise.py`
Comment/docstring-only: the `run_in_actor()` deprecation
TODO now points at the implemented `.to_actor.run()`
successor (checkbox ticked) and the method docstring
gains a NOTE steering users to the new API; remaining
TODO items are the `DeprecationWarning` emission +
in-repo usage migration.
> `git diff main..wkt/to_actor_subpkg -- tests/test_to_actor.py`
11-test suite: private-nursery one-shot, implicit
runtime boot via `runtime_kwargs`, remote-error relay to
the caller's task (bare + caller-managed nursery),
caller-nursery spawn, portal reuse w/o implicit reap,
the concurrent worker-pool-ish pattern (local `trio`
nursery x shared `an`), and the four validation
rejections (sync fn, async-gen fn, `portal`+`an`
combo, `runtime_kwargs`+placement combo).
> `git diff main..wkt/to_actor_subpkg -- examples/parallelism/to_actor_one_shots.py`
Runnable example (auto-collected by
`test_docs_examples.py`): the fully-implicit one-shot
plus the concurrent worker-pool-ish prime-check pattern
against a shared caller-managed actor-nursery.
## Test runs (verbatim)
```
tests/test_to_actor.py .......... [100%]
============= 10 passed in 4.29s =============
```
Regression subset for touched modules
(`test_local.py test_rpc.py test_spawning.py
test_cancellation.py`):
```
38 passed, 1 xfailed, 24 warnings in 80.71s (0:01:20)
```
(warnings are pre-existing stdlib `os.fork()`
DeprecationWarnings from the mp spawn backends, not
introduced by this change)

View File

@ -1,77 +0,0 @@
---
model: claude-fable-5
service: claude
session: b8ec2ffe-1e47-4db5-8bc3-5bebc5feaaf6
timestamp: 2026-07-02T15:50:06Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260702T155006Z_65bf9df5_prompt_io.raw.md
---
## Prompt
From `ai/prompt-io/prompts/issue_473.md`:
> NOTE: you MUST pause this work at 12:50PM EST (BEFORE
> your weekly limit reset) for review by a human!
>
> resolve GH #473: https://github.com/goodboy/tractor/issues/474
> do it with /open-wkt.
(The prompt's issue-number/URL mismatch was resolved in
favor of #473 — file name, session name and prompt text
all say #473; the `/474` URL is a typo.)
GH #473: "Root-cause UDS-on-macOS —
`uds_transport_actor_tree` example fails on CI".
## Response summary
Worked all four #473 next-step checkboxes in wkt
`uds_macos_473`:
1. **stderr surfacing**: `test_docs_examples.py` now
raises with the full subproc stderr (+stdout) on ANY
non-zero exit. The old last-line-'Error' check could
never fire for `tractor` crashes since root-actor
stderr always ends with the strict-EG collapse note
`( ^^^ this exc was collapsed from a group ^^^ )`
proven against the real PR #460 macOS CI log (bare
`assert 1 == 0`, no traceback).
2. **root-cause (linux-provable layer)**: macOS-only
addr corruption in
`MsgpackUDSStream.get_stream_addrs()` — no
`SO_PASSCRED`/autobind on darwin means the accept
side's `getpeername()` is `''`, and the
`(str(), str())` arm took `peername` unconditionally →
`Path('')` garbage addrs on every accepted conn.
Proven + fixed via linux no-autobind simulation.
Possibly not the final macOS crasher (non-fatal on
linux-sim); the diagnostic patch guarantees the next
macOS CI run shows any remaining layer.
3. **CI matrix**: removed the `macos-latest`+`uds`
exclude.
4. **un-skip**: dropped the macOS+CI skip of the example.
Also: `start_listener()` bindspace mkdir hardened
(`parents=True, exist_ok=True`), example docstring
peer-pid mechanism corrected for macOS.
## Files changed
- `tests/test_docs_examples.py` — surface full stderr on
non-zero exit; remove macOS skip of the UDS example
- `tractor/ipc/_uds.py` — fix no-autobind
`get_stream_addrs()` arm; harden bindspace mkdir;
document autobind semantics
- `.github/workflows/ci.yml` — add macOS UDS leg (remove
matrix exclude)
- `examples/uds_transport_actor_tree.py` — platform-
correct peer-pid docs
## Human edits
None yet — pending human review (hard-stop review
checkpoint at 12:50PM EST per prompt); commit staging and
any edits are the human's.

View File

@ -1,101 +0,0 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T15:50:06Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/uds_macos_473
---
# Raw output — GH #473 UDS-on-macOS root-cause session
NOTE: code output is diff-referenced (not copied) per
prompt-io diff-ref mode; all generated code is on branch
`wkt/uds_macos_473` relative to `main` (65bf9df5).
## Diagnostic narrative (verbatim)
Root-cause work proceeded by linux-side simulation since no
macOS host is available:
1. Pulled the actual macOS CI failure log from PR #460 run
28343878933 (job 84115107992): the failure is a bare
`assert 1 == 0` on `proc.returncode` with NO stderr
surfaced — the harness only re-raised when the LAST
stderr line contained 'Error'.
2. Reproduced the stderr-swallow mechanism on linux: any
`tractor` root-actor crash ends stderr with the
strict-EG collapse note
`( ^^^ this exc was collapsed from a group ^^^ )` which
never matches 'Error' — so EVERY possible crash was
swallowed. (Verified by sabotaging the runtime dir via
an over-long `XDG_RUNTIME_DIR` → `OSError: AF_UNIX path
too long` → rc=1 + swallowed.)
3. Found + proved a macOS-only addr-corruption bug in
`MsgpackUDSStream.get_stream_addrs()`: the
`(str(), str())` match-arm unconditionally took
`peername`, but on no-autobind platforms (macOS lacks
linux's `SO_PASSCRED`-triggered autobind) the accept
side's `getpeername()` is `''``Path('')` garbage
laddr/raddr on EVERY accepted UDS conn. Simulated on
linux by nulling `SO_PASSCRED` (no autobind → same `''`
shape): pre-fix the example printed
`listener sock file: .`; post-fix it prints the real
registry sockpath. Non-fatal on linux-sim (rc=0), so
possibly not the final macOS crasher — the diagnostic
patch guarantees the next macOS CI run reveals any
remaining layer.
4. Falsified the missing-parent-dir theory:
`get_rt_dir()` already `mkdir(parents=True,
exist_ok=True)`s at import (and macOS TCP CI passes),
so `~/Library/Caches/TemporaryItems` absence cannot be
the crasher. Hardened `start_listener()`'s bindspace
mkdir anyway (custom `filedir` case + racing actors).
## Generated changes (diff pointers)
> `git diff main..wkt/uds_macos_473 -- tests/test_docs_examples.py`
- always raise with FULL subproc stderr (+stdout) on any
non-zero example exit; keep legacy last-line 'Error'
check for zero-rc cases; drop the macOS+CI skip of
`uds_transport_actor_tree.py` (GH #473 next-step).
> `git diff main..wkt/uds_macos_473 -- tractor/ipc/_uds.py`
- `get_stream_addrs()`: document the autobind semantics
(bytes = linux abstract-ns autobind artifact), add
no-autobind `(str, str)` arm picking the non-empty name
(`peername` connect-side, `sockname` accept-side) with
an empty-pair `ValueError` guard.
- `start_listener()`: `bs.mkdir(parents=True,
exist_ok=True)`.
> `git diff main..wkt/uds_macos_473 -- .github/workflows/ci.yml`
- remove the `macos-latest`+`uds` matrix exclude so
UDS-on-macOS is exercised by CI (GH #473 next-step).
> `git diff main..wkt/uds_macos_473 -- examples/uds_transport_actor_tree.py`
- docs nit: peer-pid mechanism is `SO_PEERCRED` on linux,
`LOCAL_PEERPID` on macOS.
## Verification (verbatim summary)
- macOS-shape sim (no autobind) + fix: example rc=0 with
correct listener sockpath.
- native linux post-fix: example rc=0, autobind arms
unchanged.
- sabotage smoke-test: diagnostic patch surfaces the full
EG traceback incl. collapse-note line.
- `pytest tests/ipc/ tests/test_2way.py tests/discovery/
--tpt-proto uds`: 61 passed, 1 `TooSlowError` flake
(test_simple_rpc) that passes solo in 2.2s — attributed
to a concurrent full-suite run (other session) loading
the host.
- full `tests/test_docs_examples.py` run pending at
raw-file write time (queued behind the concurrent
session's suite).

View File

@ -1,83 +0,0 @@
---
model: claude-fable-5
service: claude
session: b6b42e23-0454-4d9b-be1c-0a89d65aed1b
timestamp: 2026-07-02T15:56:26Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260702T155626Z_65bf9df5_prompt_io.raw.md
---
## Prompt
From `ai/prompt-io/prompts/issue_470.md`:
> attempt to resolve
> https://github.com/goodboy/tractor/issues/470
> do it with /open-wkt, ensure you /run-tests on all
> changes
(plus a hard pause-for-human-review deadline of
12:50PM EST.)
Issue #470 asks to trim the ~0.42s `import tractor`
cost — which dominates per-actor spawn latency on the
`trio` backend — by lazy-importing heavy/optional deps
(`pdbp`, `stackscope`, `multiaddr`, + audit of
`colorlog`/`bidict`/`wrapt`/`setproctitle`).
## Response summary
Profiling showed the issue's dep-list only accounted
for ~20ms; the dominant cost (~244ms) was
`log.get_logger()`'s `get_caller_mod()` calling
`inspect.stack()` at module level in ~39 modules —
each call walks every stack frame (deep during nested
imports) and scans `sys.modules` per frame via
`inspect.getmodule()`.
Changes, in impact order:
1. `get_caller_mod()` -> `sys._getframe()` +
`f_globals['__name__']` `sys.modules` lookup
(~240ms saved).
2. Issue's lazy-import checklist: `bidict`,
`multiaddr`, `colorlog`, `wrapt` moved to
`TYPE_CHECKING`/function-local imports;
`platformdirs` function-local; `asyncio` +
`.to_asyncio` deferred out of the `devx.debug` +
`spawn._entry` eager paths (~15ms saved).
3. PEP 562 `__getattr__` on `tractor/__init__.py`
preserving public `tractor.to_asyncio` attr access.
Results: `import tractor` 0.42s -> ~0.145s (~65%);
sequential `start_actor` latency 0.40-0.44s ->
~0.179s/actor. `pdbp` (needs `_repl.py` class-base
restructure) + `platformdirs` (needs
`UDSAddress.def_bindspace` protocol rework) documented
as follow-ups.
## Files changed
- `tractor/log.py``get_caller_mod()` perf fix +
lazy `colorlog`
- `tractor/__init__.py` — PEP 562 lazy `to_asyncio`
- `tractor/discovery/_addr.py``bidict` ->
`TYPE_CHECKING`
- `tractor/discovery/_multiaddr.py` — lazy `multiaddr`
- `tractor/ipc/_tcp.py`, `tractor/ipc/_uds.py`
`Multiaddr` -> `TYPE_CHECKING`
- `tractor/runtime/_state.py` — lazy `platformdirs`
- `tractor/devx/_frame_stack.py` — lazy `pdbp` +
`wrapt`
- `tractor/devx/debug/_trace.py`,
`tractor/devx/debug/_tty_lock.py` — lazy `asyncio` +
`.to_asyncio`
- `tractor/spawn/_entry.py` — lazy
`run_as_asyncio_guest`
## Human edits
None yet — pending user review at the 12:50PM EST
pause gate (test-suite results reported in-session).

View File

@ -1,122 +0,0 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T15:56:26Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/boot_latency_470
---
# Raw AI output — gh #470 `import tractor` latency trim
All generated code is committed on the
`wkt/boot_latency_470` branch; per diff-ref mode each
file's content is referenced via its diff instead of
copied verbatim.
## Profiling findings (verbatim analysis output)
Baseline: `import tractor` ~0.39-0.42s wall.
`python -X importtime` + `cProfile` traced the cost NOT
primarily to third-party deps (the issue's hypothesis)
but to `tractor/log.py:get_logger()` calling
`get_caller_mod()` -> `inspect.stack()` at module level
in ~39 tractor modules:
- `inspect.stack()` builds `FrameInfo` (incl. src-file
and line-context resolution) for EVERY frame on the
stack; during nested imports the stack is dozens of
importlib frames deep.
- each `FrameInfo` resolution calls
`inspect.getmodule()` which scans all of
`sys.modules` per frame (1.4M `ismodule()` calls in
one profiled import).
- aggregate: ~244ms of tractor-own module "self" time
vs ~20ms for ALL the issue-listed third-party deps
(`pdbp` ~10ms, `bidict` ~4.5ms, `multiaddr` ~3.5ms,
`wrapt`/`colorlog` ~1ms each); `trio` itself is
~70-100ms and unavoidable.
## Generated changes
> `git diff main..wkt/boot_latency_470 -- tractor/log.py`
`get_caller_mod()` rewritten from `inspect.stack()` +
`inspect.getmodule()` to `sys._getframe(frames_up)` +
`frame.f_globals['__name__']` -> `sys.modules` lookup
(O(1) vs O(stack x sys.modules)). Unused `inspect`
imports dropped; `FrameType` imported from `types`.
Also `colorlog` lazy-imported inside
`get_console_log()`.
> `git diff main..wkt/boot_latency_470 -- tractor/discovery/_addr.py`
`bidict` import moved under `TYPE_CHECKING`
(annotation-only use; `_address_types` is a plain dict
literal).
> `git diff main..wkt/boot_latency_470 -- tractor/discovery/_multiaddr.py`
`from __future__ import annotations` added; `multiaddr`
import moved under `TYPE_CHECKING` + function-local
imports in `mk_maddr()`/`parse_maddr()`.
> `git diff main..wkt/boot_latency_470 -- tractor/ipc/_tcp.py tractor/ipc/_uds.py`
`Multiaddr` imports moved under `TYPE_CHECKING`
(annotation-only in both transports).
> `git diff main..wkt/boot_latency_470 -- tractor/runtime/_state.py`
`platformdirs` lazy-imported inside `get_rt_dir()`
(NOTE: still imported eagerly via
`UDSAddress.def_bindspace` class-var eval; see
follow-ups).
> `git diff main..wkt/boot_latency_470 -- tractor/devx/_frame_stack.py`
`pdbp` + `wrapt` lazy-imported inside
`hide_runtime_frames()` / `api_frame()` respectively.
> `git diff main..wkt/boot_latency_470 -- tractor/devx/debug/_trace.py tractor/devx/debug/_tty_lock.py`
`asyncio` moved to `TYPE_CHECKING` + call-site local
imports (`asyncio.current_task()` sites);
`tractor.to_asyncio.run_trio_task_in_future` imports
moved into the infected-aio runtime branches.
> `git diff main..wkt/boot_latency_470 -- tractor/spawn/_entry.py`
`run_as_asyncio_guest` import moved into the
`infect_asyncio=True` branches of `_mp_main()` /
`_trio_main()`.
> `git diff main..wkt/boot_latency_470 -- tractor/__init__.py`
PEP 562 module `__getattr__` added so
`tractor.to_asyncio` attr-access still works (required
by `tests/test_child_manages_service_nursery.py` and
any downstream user) while keeping `asyncio` off the
eager import path.
## Measured results (verbatim)
- `import tractor`: 0.39-0.42s -> ~0.145s (~65% cut)
- `start_actor` spawn+boot+reg+cancel: ~0.40-0.44s ->
~0.179s/actor (n=5 sequential, warm parent)
- post-change eager-module check: only `pdbp` +
`platformdirs` of the issue's list remain eager.
## Known follow-ups (not implemented, deadline-bound)
- `pdbp` (~10ms): still eager via
`devx/debug/_repl.py` class bases
(`class PdbREPL(pdbp.Pdb)`) + `_tty_lock.py`
module-level `@pdbp.hideframe`; needs `_repl`
restructure + PEP 562 in `devx.debug.__init__`.
- `platformdirs` (~1.5ms): eager via
`UDSAddress.def_bindspace: ClassVar = get_rt_dir()`
class-body call; needs `Address`-protocol rework of
`def_bindspace` to a lazy accessor.
- `stackscope` + `setproctitle`: already lazy/absent —
no change needed.

View File

@ -1,80 +0,0 @@
---
model: claude-fable-5
service: claude
session: f4da6654-63f9-4641-b567-f200da46d697
timestamp: 2026-07-02T16:16:24Z
git_ref: 65bf9df5
scope: tests
substantive: true
raw_file: 20260702T161624Z_65bf9df5_prompt_io.raw.md
---
## Prompt
> attempt to resolve
> https://github.com/goodboy/tractor/issues/474
> do it with /open-wkt.
Issue #474: "We need test suite for
`.trionics.start_or_cancel()`" — dig out the use case
from `modden` as a unit test (and provide motivation
for upstream `trio` to possibly provide better
porcelain around this case). The wrapper landed via
PR #464 with no tests.
## Response summary
New `tests/trionics/test_taskc.py` (9 tests) covering
`start_or_cancel()`, authored in wkt
`start_or_cancel_tests_474`.
Behaviour was first probed empirically (trio 0.29):
the lossy `RuntimeError('child exited without calling
task_status.started()')` only fires when the child
exits pre-`.started()` WITHOUT propagating the ambient
`Cancelled` — i.e. when the child (or lib code it
calls) absorbs the cancel in a graceful-teardown
pattern; a well-behaved child surfaces `Cancelled`
straight out of `.start()`. The `modden`
`progman.open_wks()` use case was reconstructed from
`modden/runtime/progman.py` accordingly.
Tests (each `use_start_or_cancel` parametrization also
pins upstream trio's current lossy behaviour as
wart-documentation):
- `test_sibling_err_not_masked_by_startup_rte` — the
`modden` case: sibling error OOB-cancels the shared
nursery scope; with the wrapper ONLY the root-cause
`ValueError` escapes; bare `.start()` adds the lossy
RTE alongside.
- `test_pure_oob_cancel_not_morphed_to_rte` — plain
ancestor `cs.cancel()`: wrapper → clean exit; bare
→ eg-wrapped RTE.
- `test_genuine_startup_rte_still_raised` — no
cancellation → protocol-bug RTE re-raised same as
bare.
- `test_childs_own_rte_never_demoted_to_cancel` — a
child's own `RuntimeError('never got started!')` /
`RuntimeError(1234)` under ambient cancel is never
demoted to `Cancelled` (exact-msg-match + str-guard
regression cover).
- `test_started_value_and_args_passthru` — happy path:
positional args, `name=`, `.started()` value.
Verified: 9/9 pass; 0 flakes across 50 hammer runs;
two impl mutations (checkpoint removed; guard relaxed
to substring match) each caught by exactly the
targeted tests; `tests/trionics/` +
`tests/test_trioisms.py` subset green (23 passed,
5 xfailed); ruff clean; 69-col style.
## Files changed
- `tests/trionics/test_taskc.py` — new
`start_or_cancel()` unit-test suite (gh #474).
## Human edits
Pending review — session paused pre-commit per user
deadline; nothing committed as of this entry.

View File

@ -1,107 +0,0 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T16:16:24Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/start_or_cancel_tests_474
---
# Raw output — gh #474 `start_or_cancel()` test suite
## Generated test code
> `git diff main..wkt/start_or_cancel_tests_474 -- tests/trionics/test_taskc.py`
Prose summary of the generated module
(`tests/trionics/test_taskc.py`):
- module docstring framing the `trio.Nursery.start()`
startup-cancellation wart, the wrapper's repair, and
the intent that `use_start_or_cancel=False` params
double as upstream-trio wart-documentation (break on
a trio upgrade → upstream may have shipped porcelain,
re-audit the wrapper); cites gh #474 / PR #464 and
`modden`'s `progman.open_wks()` as the source use
case.
- shared children: `absorbs_cancel_pre_started()` (the
graceful-teardown cancel-absorber which triggers the
lossy RTE path) + `raise_val_err()` (fast-erroring
sibling).
- `test_sibling_err_not_masked_by_startup_rte`
(parametrized `use_start_or_cancel`): asserts eg
contains exactly one `ValueError` and, wrapper-case,
NO residual RTE (`eg.split(ValueError)` remainder is
`None`); bare-case, the residual RTE carries trio's
exact "child exited without calling" wording.
- `test_pure_oob_cancel_not_morphed_to_rte`
(parametrized): wrapper-case runs clean and asserts
`cs.cancelled_caught`; bare-case asserts the
eg-wrapped RTE.
- `test_genuine_startup_rte_still_raised`
(parametrized): no-cancel protocol bug → RTE with
trio's wording from both call forms.
- `test_childs_own_rte_never_demoted_to_cancel`
(parametrized `rte_arg` in `'never got started!'`,
`1234`): child cancels the ambient scope then raises
its own RTE synchronously (no checkpoint between →
deterministically under-cancellation at catch time);
asserts the RTE survives with `args[0]` intact.
- `test_started_value_and_args_passthru`: `.started()`
value, positional args and the `name=` kwarg (via
`trio.lowlevel.current_task().name`) all forward.
## Non-code output (verbatim highlights)
Behaviour probe (trio 0.29, scratchpad scripts) — the
decision basis for the test shapes:
```
== B-sibling-err use_soc=False
start raised: RuntimeError('child exited without
calling task_status.started()')
top-level: ExceptionGroup([ValueError('sibling blew
up!'), RuntimeError('child exited without calling
task_status.started()')])
== B-cs-cancel use_soc=False
top-level: ExceptionGroup([RuntimeError('child
exited without calling task_status.started()')])
== B-sibling-err use_soc=True
start raised: Cancelled()
top-level: ExceptionGroup([ValueError('sibling blew
up!')])
== B-cs-cancel use_soc=True
start raised: Cancelled()
top-level: clean return
== own-rte-under-cancel (both) -> RTE('never got
started!') propagates unchanged
```
Key finding: with a WELL-BEHAVED (non-absorbing) child
an OOB ancestor cancel surfaces `Cancelled` directly
from `.start()` on trio 0.29 — the lossy RTE requires
the child to absorb its cancel pre-`.started()`, which
is what `modden`'s `open_from_wks` teardown did. Trio's
nursery-exit wait defers cancel delivery to children,
so all tested shapes are deterministic (0 flakes / 50
runs).
Mutation verification:
```
mutation 1 (checkpoint_if_cancelled removed):
FAILED test_sibling_err_not_masked_by_startup_rte[True]
FAILED test_pure_oob_cancel_not_morphed_to_rte[True]
mutation 2 (guard relaxed to 'started' substring,
isinstance dropped):
FAILED test_childs_own_rte_never_demoted_to_cancel[never got started!]
FAILED test_childs_own_rte_never_demoted_to_cancel[1234]
```
Final runs:
```
tests/trionics/test_taskc.py: 9 passed in 0.03s
hammer: 0/50 runs failed
tests/trionics/ + tests/test_trioisms.py:
23 passed, 5 xfailed in 3.02s
```

View File

@ -1,39 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: pr475-review-fixes-20260817
timestamp: 2026-08-17T23:18:25Z
git_ref: 359fe75c
scope: code
substantive: true
raw_file: 20260817T231825Z_359fe75c_prompt_io.raw.md
---
## Prompt
Continue the `/code-review-changes` pass for PR #475 in its isolated
worktree. Address the seven accepted manual-review findings in
`tractor/ipc/_types.py` and `tractor/ipc/_uds.py`, preserve the existing
Windows capability behavior, verify the result, and prepare the work for
human-controlled commit and review-reply steps. Do not publish replies,
stage, commit, or push without the required explicit authorization.
## Response summary
Restored project quote, docstring, multiline-expression, and
`match/case` conventions while retaining the Windows-safe UDS guard.
Removed unnecessary structural and comment churn, then verified the
focused transport, discovery, and lazy-import paths plus the missing
`AF_UNIX` behavior.
## Files changed
- `tractor/ipc/_types.py` - restore project style and guarded
socket-family dispatch.
- `tractor/ipc/_uds.py` - format the UDS capability gate
consistently.
## Human edits
None - the generated patch remains uncommitted and awaits human
review.

View File

@ -1,45 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-17T23:18:25Z
git_ref: 359fe75c
diff_cmd: git diff HEAD~1..HEAD
---
Applied the seven accepted manual-review fixes for PR #475 while
preserving the Windows transport capability behavior.
> `git diff HEAD~1..HEAD -- tractor/ipc/_types.py`
The generated changes restore the project's single-quote docstring and
string conventions, remove the unnecessary helper divider, simplify the
transport-registry comments, and restore `match/case` socket-family
dispatch. The UDS case retains a `HAS_UDS` guard that short-circuits
before `socket.AF_UNIX` is evaluated on unsupported hosts. Nearby error
messages are wrapped without changing their content.
> `git diff HEAD~1..HEAD -- tractor/ipc/_uds.py`
The generated change reformats the `HAS_UDS` conjunction according to
the project's multiline boolean-expression convention and simplifies
the adjacent capability comment.
Verification:
`/home/goodboy/repos/tractor/py313/bin/pytest -q tests/test_lazy_imports.py tests/discovery tests/ipc/test_server.py`
Result: `66 passed, 2 xpassed in 60.62s`.
`ruff check --no-cache --output-format=json tractor/ipc/_types.py tractor/ipc/_uds.py`
Result: no findings.
`git diff --check`
Result: no whitespace errors.
An explicit missing-`AF_UNIX` probe set `HAS_UDS = False`, removed the
socket constant, and exercised an unsupported socket family. It raised
the expected `NotImplementedError` instead of `AttributeError`.
No review replies, commits, or pushes were published.

View File

@ -1,43 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: pr481-review-fixes-p1-20260818
timestamp: 2026-08-18T03:15:32Z
git_ref: 4151b956
scope: code
substantive: true
raw_file: 20260818T031532Z_4151b956_prompt_io.raw.md
---
## Prompt
Address the approved review findings on PR #481, but work
iteratively: implement and verify one finding at a time, prepare a
separate `/commit-plan` after each fix, and stop for the human commit
before starting the next finding. Begin with the P1 per-child
lifecycle issue. Also publish the already-approved review findings
against the reviewed PR head before editing.
## Response summary
Published the approved non-approving review at head `4151b956`, then
implemented only the P1 lifecycle fix. Owned one-shot actors now use
a child-specific cancellation and process-reap handshake, including
hard escalation for unacknowledged cancellation and deterministic
bookkeeping removal before `to_actor.run()` returns.
## Files changed
- `tractor/runtime/_supervise.py` - coordinate child-specific cancel
and reap.
- `tractor/spawn/_trio.py` - wait on the Trio child's reap request.
- `tractor/spawn/_mp.py` - wait on the multiprocessing child's reap
request.
- `tractor/spawn/_spawn.py` - publish monitor completion centrally.
- `tractor/to_actor/_api.py` - await owned-child process reaping.
- `tests/test_to_actor.py` - cover cleanup, escalation, and startup
ordering.
## Human edits
None - the generated P1 patch remains uncommitted for human review.

View File

@ -1,74 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-18T03:15:32Z
git_ref: 4151b956
diff_cmd: git diff HEAD~1..HEAD
---
Implemented only the P1 lifecycle finding from the approved PR #481
review, preserving the requested one-fix-at-a-time commit boundary.
> `git diff HEAD~1..HEAD -- tractor/runtime/_supervise.py`
Added per-child reap request/completion events to `ActorNursery`, a
shielded child-specific cancel-and-reap operation, late-registration
latching for nursery teardown, and cancellation escalation that waits
for debugger release before using non-ignorable process termination.
The nursery-wide cancellation path snapshots child records before
checkpointing so concurrent one-shot cleanup cannot invalidate its
iteration.
> `git diff HEAD~1..HEAD -- tractor/spawn/_trio.py`
Changed Trio child monitors to wait on their per-child reap requests.
> `git diff HEAD~1..HEAD -- tractor/spawn/_mp.py`
Changed multiprocessing child monitors to wait on their per-child reap
requests.
> `git diff HEAD~1..HEAD -- tractor/spawn/_spawn.py`
Ensured every backend publishes child-reap completion after its process
monitor exits.
> `git diff HEAD~1..HEAD -- tractor/to_actor/_api.py`
Changed owned one-shot cleanup to await child-specific process joining
and bookkeeping removal instead of treating the cancel RPC as reaping.
> `git diff HEAD~1..HEAD -- tests/test_to_actor.py`
Added regressions for immediate caller-managed nursery cleanup, failed
cancel acknowledgement escalation, and child registration after a
latched nursery-wide teardown request.
Verification:
`pytest -q tests/test_to_actor.py tests/test_cancellation.py tests/test_spawning.py tests/discovery/test_multi_program.py`
Result: `46 passed, 1 xfailed, 3 xpassed`.
`pytest -q tests/test_to_actor.py --tpt-proto uds`
Result: `13 passed`.
`pytest -q tests/test_to_actor.py --spawn-backend mp_spawn --tpt-proto tcp`
Result: `13 passed`.
One broad verification run was mistakenly launched in parallel with
the UDS and `mp_spawn` actor suites. It timed out
`test_remote_error_from_caller_nursery`; the node passed immediately
in isolation and the complete broad selection then passed serially.
The failure was classified as concurrent test-session interference,
not accepted as a passing boundary result.
Ruff, Python compilation, and `git diff --check` passed for the changed
boundary. Ruff's existing `_trio.py` F401 finding was reproduced at the
unmodified PR head and excluded from attribution to this patch.
No source files were staged, committed, pushed, or used for review
replies. The previously approved top-level review was published before
the fix at reviewed head `4151b956`.

View File

@ -1,36 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: ses_3e4c90d3eafeqHEtRYSIHgHhpA
timestamp: 2026-08-18T19:30:01Z
git_ref: bf06b4f8
scope: code
substantive: true
raw_file: 20260818T193001Z_bf06b4f8_prompt_io.raw.md
---
## Prompt
Fix cancellation during an indeterminate transport write without
allowing a second RPC msg to corrupt the stream. Keep the transport
fix separate from the higher-level `to_actor` remediation and explain
the length-prefixed protocol rather than using unexplained "framed
send" terminology.
## Response summary
Close a msgpack transport stream when cancellation interrupts its
length-prefixed `send_all()` operation. The stream can no longer be
safely reused because Trio cannot report how many bytes were written.
## Files changed
- `tractor/ipc/_transport.py` - close an interrupted send stream.
- `tests/ipc/test_each_tpt.py` - cover cancellation during the write.
## Human edits
The human required this transport edge-case fix to land as its own
behavioral commit with a detailed message. During staged review, the
human also rejected the unexplained "framed send" wording and asked
for terminology tied directly to the actual transport operation.

View File

@ -1,19 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-18T19:30:01Z
git_ref: bf06b4f8
diff_cmd: git diff HEAD~1..HEAD
---
Prospective review found that cancellation can interrupt
`MsgpackTransport.send()` after `send_all()` writes only part of its
length-prefixed msg. Sending a cancellation request afterward can
append another msg to the indeterminate stream and desynchronize the
peer decoder.
> `git diff HEAD~1..HEAD -- tractor/ipc/_transport.py tests/ipc/test_each_tpt.py`
Close the stream under a cancellation shield when `send_all()` is
cancelled. Cover the behavior with a fake stream that checkpoints
inside the write and records forced closure.

View File

@ -1,37 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: ses_3e4c90d3eafeqHEtRYSIHgHhpA
timestamp: 2026-08-18T19:30:02Z
git_ref: bf06b4f8
scope: code
substantive: true
raw_file: 20260818T193002Z_bf06b4f8_prompt_io.raw.md
---
## Prompt
Distill repeated `Actor._contexts.pop()` machinery into a wrapper like
the RPC-task registration helper so future teardown sites do not keep
reconstructing the context-registry key independently. Preserve the
existing lifecycle-specific cleanup behavior.
## Response summary
Add idempotent `Actor._drop_context()` registry removal keyed from the
context's own channel and CID. Use it for caller context teardown and
the strict callee-side RPC deregistration path.
## Files changed
- `tractor/runtime/_runtime.py` - own context-registry removal.
- `tractor/runtime/_rpc.py` - use the helper for callee teardown.
- `tractor/_context.py` - use the helper after caller teardown.
## Human edits
The human identified the repeated registry-pop code and requested a
central primitive analogous to `_register_rpc_task()`. The agent first
suggested an async helper that also closed receive channels; the final
design was narrowed to registry removal only so each lifecycle owner
retains its existing closure, debugger, shielding, and error policy.

View File

@ -1,18 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-18T19:30:02Z
git_ref: bf06b4f8
diff_cmd: git diff HEAD~1..HEAD
---
Repeated teardown sites reconstruct the `Actor._contexts` registry
key from a portal channel and context ID before popping it. Add an
idempotent actor-owned helper deriving the key from the context itself,
then route caller and callee context teardown through that helper.
> `git diff HEAD~1..HEAD -- tractor/runtime/_runtime.py tractor/runtime/_rpc.py tractor/_context.py`
Keep receive-channel closure and cancellation shielding in each
lifecycle owner so the helper centralizes registry machinery without
changing their teardown ordering.

View File

@ -1,38 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: ses_3e4c90d3eafeqHEtRYSIHgHhpA
timestamp: 2026-08-18T19:30:03Z
git_ref: bf06b4f8
scope: code
substantive: true
raw_file: 20260818T193003Z_bf06b4f8_prompt_io.raw.md
---
## Prompt
Cancel a remote task when its caller is cancelled after `Start`
publication but before startup acknowledgement. Keep cancellation
bounded, prevent its private `_cancel_task` RPC from recursively
cancelling itself and preserve public target kwargs unchanged.
## Response summary
Add private portal startup policy, use it for non-recursive context
cancellation and clean caller-side startup state under a shield.
## Files changed
- `tractor/runtime/_portal.py` - separate private startup policy.
- `tractor/_context.py` - disable recursion for cancellation RPCs.
- `tractor/runtime/_runtime.py` - clean cancelled task startup.
- `tests/test_context_stream_semantics.py` - control cancellation
between `Start` publication and acknowledgement.
## Human edits
The human required this cancellation behavior to remain a distinct
commit from general startup failures and from the public `to_actor`
API. The human also requested that its runtime comment describe the
actual length-prefixed transport guarantee and concrete `_cancel_task`
operation rather than referring to an unnamed wrapper.

View File

@ -1,20 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-18T19:30:03Z
git_ref: bf06b4f8
diff_cmd: git diff HEAD~1..HEAD
---
Cancellation while `Actor.start_remote_task()` waits for `StartAck`
can strand its caller-side context and leave the remote task running.
Make one bounded cleanup request, remove local startup state and close
its receive channel.
> `git diff HEAD~1..HEAD -- tractor/runtime/_runtime.py tractor/runtime/_portal.py tractor/_context.py tests/test_context_stream_semantics.py`
Separate private startup-cancellation policy from public target kwargs
using `Portal._run_from_ns()`. Have `Context.cancel()` disable recursive
startup cancellation for its own `_cancel_task` RPC. Exercise
cancellation after `Start` publication and prove the caller-owned actor
remains reusable without leaked contexts.

View File

@ -1,36 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: ses_3e4c90d3eafeqHEtRYSIHgHhpA
timestamp: 2026-08-18T19:30:04Z
git_ref: bf06b4f8
scope: code
substantive: true
raw_file: 20260818T193004Z_bf06b4f8_prompt_io.raw.md
---
## Prompt
Release caller-side context state for every remote-task startup failure,
not only local cancellation. Preserve the remote error, avoid unsafe
follow-up sends and prove pre-publication serialization failures leave
a reused portal healthy.
## Response summary
Extend remote-task startup cleanup across send, acknowledgement and
validation errors. Track completed publication, perform only safe
best-effort cancellation and deterministically remove local state.
## Files changed
- `tractor/runtime/_runtime.py` - clean every startup failure path.
- `tests/test_context_stream_semantics.py` - cover authorization and
serialization failures before context entry.
## Human edits
The human accepted the discovered edge-case fixes but required general
startup cleanup to land separately from cancellation cleanup, transport
integrity and the public API. This boundary preserves that behavioral
distinction and its dedicated commit-message rationale.

View File

@ -1,19 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-18T19:30:04Z
git_ref: bf06b4f8
diff_cmd: git diff HEAD~1..HEAD
---
`Actor.start_remote_task()` inserts a context before sending `Start`,
but startup errors other than cancellation escape without removing or
closing that caller state. Serialization errors, acknowledgement
timeouts, malformed acknowledgements and remote authorization errors
can therefore leak context-registry entries.
> `git diff HEAD~1..HEAD -- tractor/runtime/_runtime.py tests/test_context_stream_semantics.py`
Cover the complete send, acknowledgement and validation phase with
exceptional cleanup. Attempt remote cancellation only when publication
is known complete or protocol-safe, and always release local state.

View File

@ -1,52 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: ses_3e4c90d3eafeqHEtRYSIHgHhpA
timestamp: 2026-08-18T19:30:05Z
git_ref: bf06b4f8
scope: code
substantive: true
raw_file: 20260818T193005Z_bf06b4f8_prompt_io.raw.md
---
## Prompt
Replace abandoned `Portal.run()` one-shots with a static linked-context
endpoint. Follow Trio positional-call semantics, use partials for target
keywords, preserve Python 3.14 Placeholder behavior, keep target lookup
behind the RPC allowlist and support private, nursery and portal
placement.
## Response summary
Use `Portal.open_context()` and `Context.wait_for_result()` for one-shot
tasks. Normalize every partial layer, validate signatures locally and
send target namespace/function components separately to the authorized
remote resolver. Retain the client-side function in its `NamespacePath`
so `to_tuple()` does not re-import it. Owned actors enable the declaring
`_api.__name__` directly; caller-owned portals opt in through the public
`to_actor.MODULE` alias.
## Files changed
- `tractor/to_actor/_api.py` - implement linked one-shot calls.
- `tractor/to_actor/__init__.py` - export `MODULE`.
- `tractor/msg/ptr.py` - retain refs created by `from_ref()`.
- `tests/test_to_actor.py` - cover the public API and authorization.
- `examples/parallelism/to_actor_one_shots.py` - use positional inputs.
## Human edits
The human rejected nested target-kwargs configuration and selected
Trio-style positional inputs plus `functools.partial()`. During staged
review the human required a Python 3.14 compatibility comment rather
than removing Placeholder support, requested separate namespace and
function inputs, preserved `_get_rpc_func(ns: str, funcname: str)`
authorization, renamed `RPC_MODULE` to `MODULE`, rejected global module
exposure and deferred speculative nursery/module-list helpers to the
`open_taskman()` design line. The human also required this public API
to land only after its lower-level safety dependencies. In final staged
review, the human required `_invoke_from_portal()` to use
`NamespacePath.to_tuple()` with the already-held function ref and
required internal actor setup to use `_api.__name__` directly, keeping
`to_actor.MODULE` solely as the public importer-facing alias.

View File

@ -1,24 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-18T19:30:05Z
git_ref: bf06b4f8
diff_cmd: git diff HEAD~1..HEAD
---
Implement `to_actor.run()` with Trio-style positional target arguments,
`functools.partial` keyword and Python 3.14 Placeholder binding, and a
static context endpoint that links remote results, errors and caller
cancellation.
> `git diff HEAD~1..HEAD -- tractor/to_actor/_api.py tractor/to_actor/__init__.py`
Resolve target functions through `Actor._get_rpc_func()` so module
authorization remains authoritative. Automatically expose the helper
module for actors owned by `to_actor.run()` and document explicit
exposure for a caller-owned portal.
> `git diff HEAD~1..HEAD -- tests/test_to_actor.py examples/parallelism/to_actor_one_shots.py`
Cover placement modes, argument binding, nested partials, caller-linked
cancellation, remote errors and module authorization.

View File

@ -1,34 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: ses_3e4c90d3eafeqHEtRYSIHgHhpA
timestamp: 2026-08-19T02:07:57Z
git_ref: b38efed7
scope: code
substantive: true
raw_file: 20260819T020757Z_b38efed7_prompt_io.raw.md
---
## Prompt
Resolve the remaining P3 review finding before landing PR #481:
`runtime_kwargs={}` must not be silently accepted alongside either
`an=` or `portal=` merely because the dict is falsey. Keep this as its
own final review-remediation commit.
## Response summary
Treat any non-`None` `runtime_kwargs` value as provided when validating
placement. Cover both placement APIs with empty and configured dicts,
proving the error is raised locally before actor startup.
## Files changed
- `tractor/to_actor/_api.py` - validate option presence explicitly.
- `tests/test_to_actor.py` - cover four invalid option combinations.
## Human edits
No direct line edits. The human accepted the P3 finding, required it to
remain separate from the five P2 behavioral commits and prioritized it
before the final PR #484 integration rebase and PR #481 landing steps.

View File

@ -1,25 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-19T02:07:57Z
git_ref: b38efed7
diff_cmd: git diff HEAD~1..HEAD
---
Fix the final PR #481 review finding: `runtime_kwargs` is mutually
exclusive with both caller placement options whenever it is provided,
including an empty dict.
> `git diff HEAD~1..HEAD -- tractor/to_actor/_api.py tests/test_to_actor.py`
Use an explicit `is not None` check rather than dict truthiness. Expand
the validation regression across `an=` and `portal=`, each with empty
and configured runtime kwargs, so every invalid combination fails
before actor runtime startup.
Verification:
- Trio/TCP: `23 passed`
- Trio/UDS: `23 passed`
- `mp_spawn`/TCP: `23 passed`
- Ruff and `git diff --check`: clean

View File

@ -1,37 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-19T23:48:23Z
git_ref: 557065d8
scope: tests
substantive: true
raw_file: 20260819T234823Z_557065d8_prompt_io.raw.md
---
## Prompt
Investigate PR #481's red CI run, explain the missing T-800 and
debugger-output failures, and proceed with fixes in the PR #481
worktree.
## Response summary
Updated stale teardown assertions to match #481's direct hard-reap
path and observable process-lifetime invariants. Made nested debugger
checks consume the complete pexpect transcript rather than only the
last prompt latch.
## Files changed
- `tests/devx/test_debugger.py` - assert EOF/dead-process teardown and
accumulate nested debugger output across prompt boundaries.
- `tests/devx/test_tooling.py` - assert cancel-timeout hard-reap
escalation instead of the bypassed T-800 backend marker.
## Human edits
The human reported the still-red PR #481 CI, supplied a failing job URL,
required work in `/wkts/pr481_review_fixes` and directed the agent to
continue immediately. No direct source-line edits were made by the
human.

View File

@ -1,26 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-19T23:48:23Z
git_ref: 557065d8
diff_cmd: git diff HEAD~1..HEAD
---
Diagnose and fix the stale debugger and reaper assertions failing PR
#481's Unix CI jobs.
> `git diff HEAD~1..HEAD -- tests/devx/test_debugger.py tests/devx/test_tooling.py`
Replace the old T-800 backend-log requirement with the new bounded
cancel-ack escalation evidence. Prove debugger teardown with EOF and a
dead child process instead of requiring optional `KeyboardInterrupt`
text. Accumulate all pexpect prompt chunks for nested error propagation
so expected tracebacks are not lost when `child.before` advances.
Verification:
- exact failed debugger/reaper nodes: `4 passed`
- debugger/tooling TCP: `39 passed, 6 skipped`
- debugger/tooling UDS: `39 passed, 6 skipped`
- full TCP suite: `478 passed, 9 skipped, 7 xfailed, 3 xpassed`
- full UDS rerun: `476 passed, 11 skipped, 8 xfailed, 2 xpassed`

View File

@ -1,43 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-19T23:48:24Z
git_ref: 557065d8
scope: code
substantive: true
raw_file: 20260819T234824Z_557065d8_prompt_io.raw.md
---
## Prompt
Investigate and fix PR #481's macOS TCP clustering and stream-overrun
failures without sacrificing IPC frame integrity or structured
concurrency.
## Response summary
Changed cancellation during `send_all()` from actor-wide stream closure
to shielded complete-frame publication followed by immediate pending
cancellation. Prevented failed overrun error shipment from promoting a
secondary transport closure over the context-local primary condition.
## Files changed
- `tractor/ipc/_transport.py` - complete in-flight frames before
delivering sender cancellation.
- `tractor/_context.py` - absorb transport closure while reporting an
overrun on an already-closing channel.
- `tests/ipc/test_each_tpt.py` - prove complete framing, cancellation
delivery and channel reuse.
- `tests/test_context_stream_semantics.py` - prove overrun reporting
tolerates a closed transport.
## Human edits
The human reported PR #481's red CI, asked for diagnosis and directed
the agent to proceed in the dedicated PR #481 worktree. During final
review, the human required preservation of the original far-end
cancellation rationale and fuller documentation of frame shielding,
shared-channel ownership and cancellation-delay tradeoffs. These were
human-directed agent edits; the human made no direct source-line edits.

View File

@ -1,31 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-19T23:48:24Z
git_ref: 557065d8
diff_cmd: git diff HEAD~1..HEAD
---
Fix the macOS TCP regressions where cancellation during a framed send
closed the actor-wide channel and replaced primary stream errors with
secondary `TransportClosed` failures.
> `git diff HEAD~1..HEAD -- tractor/ipc/_transport.py tractor/_context.py tests/ipc/test_each_tpt.py tests/test_context_stream_semantics.py`
Shield complete frame publication, then deliver pending cancellation
immediately after leaving the shield. Preserve channel reuse instead of
closing the multiplexed socket from a context-local sender. Treat
`TransportClosed` while shipping `StreamOverrun` as failed delivery so
the secondary error can not crash the actor-wide RPC loop.
Add deterministic unit regressions for cancellation in the middle of a
frame and overrun reporting after transport closure.
Verification:
- transport/context unit regressions: `3 passed`
- exact TCP and UDS CI-node batches: `11 passed, 1 skipped`
- transport/context/clustering/RPC TCP: `88 passed`
- transport/context/clustering/RPC UDS: `86 passed, 2 skipped`
- full TCP suite: `478 passed, 9 skipped, 7 xfailed, 3 xpassed`
- full UDS rerun: `476 passed, 11 skipped, 8 xfailed, 2 xpassed`

View File

@ -1,32 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-20T02:30:04Z
git_ref: 88a23449
scope: tests
substantive: true
raw_file: 20260820T023004Z_88a23449_prompt_io.raw.md
---
## Prompt
Inspect the two failed macOS jobs in PR #481's new CI run and continue
toward a green landing candidate.
## Response summary
Confirmed both jobs fail only the known nested crash-REPL scenario from
issue #320, while Ubuntu TCP/UDS and Windows pass. Added a targeted
macOS-CI skip without reducing Linux coverage.
## Files changed
- `tests/devx/test_debugger.py` - skip the issue #320 nested
crash-REPL node on Darwin CI.
## Human edits
The human monitored the new CI run, reported both macOS jobs dead and
directed the agent to continue diagnosis. No direct source-line edits
were made by the human.

View File

@ -1,23 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-20T02:30:04Z
git_ref: 88a23449
diff_cmd: git diff HEAD~1..HEAD
---
Diagnose the remaining macOS PR #481 CI failures after the Linux
debugger and transport fixes passed.
> `git diff HEAD~1..HEAD -- tests/devx/test_debugger.py`
Both macOS transports failed the same deeply nested crash-REPL test
already tracked by issue #320: TCP omitted one actor-specific traceback
record and UDS timed out waiting for a nested prompt. Apply an explicit
Darwin-CI skip to this one node while retaining Linux TCP/UDS coverage.
Verification:
- debugger/tooling TCP: `39 passed, 6 skipped`
- debugger/tooling UDS: `39 passed, 6 skipped`
- Ruff, compilation and `git diff --check`: clean

View File

@ -1,40 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-20T02:30:05Z
git_ref: 88a23449
scope: docs
substantive: true
raw_file: 20260820T023005Z_88a23449_prompt_io.raw.md
---
## Prompt
Audit all documentation and executable examples once more, replacing
prescriptive `run_in_actor()` usage with `to_actor.run()` or explicit
actor/context lifetime APIs before PR #481 lands.
## Response summary
Rewrote one-shot documentation around direct blocking result delivery,
linked context execution and per-call reaping. Migrated all runnable
examples, using daemon actors where reciprocal dialogs require longer
lifetimes. Added API/guide cross-links and retained only three explicit
legacy references.
## Files changed
- `docs/` - update API, quickstart and subsystem guides to showcase
`tractor.to_actor.run()` and link its underlying core APIs.
- `examples/` - migrate one-shot calls and preserve explicit daemon
lifetimes for reciprocal or long-lived actor dialogs.
## Human edits
The human requested a final docs pass covering every place that should
showcase `to_actor` over `.run_in_actor()`. Earlier review also required
named target arguments to remain visible through `functools.partial()`
and core API references to link to local guides/reference pages. These
were human-directed agent edits; the human made no direct source-line
edits.

View File

@ -1,27 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-20T02:30:05Z
git_ref: 88a23449
diff_cmd: git diff HEAD~1..HEAD
---
Perform a final rendered-documentation and executable-example pass so
PR #481 showcases `tractor.to_actor.run()` instead of the legacy
`ActorNursery.run_in_actor()` API.
> `git diff HEAD~1..HEAD -- docs examples`
Migrate one-shot guides and examples to direct result delivery through
`to_actor.run()`, preserving named target inputs with target partials.
Use daemon actors and concurrent portal calls where reciprocal actor
lifetimes require both peers to coexist. Add API and guide cross-links,
and retain only explicit legacy/removal notes.
Verification:
- executable docs examples: `23 passed`
- debugger/tooling TCP: `39 passed, 6 skipped`
- debugger/tooling UDS: `39 passed, 6 skipped`
- Ruff, compilation and `git diff --check`: clean
- local Sphinx build unavailable because Sphinx is not installed

View File

@ -1,35 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-20T13:51:25Z
git_ref: 9f99043b
scope: tests
substantive: true
raw_file: 20260820T135125Z_9f99043b_prompt_io.raw.md
---
## Prompt
Continue preparing PR #481 for landing after the prior test and
documentation commits were pushed. Follow CI and proceed with clear next
steps without merging or changing remote content unasked.
## Response summary
Followed CI through completion and found both macOS jobs failed because the
new `skipif` expression returned the `CI=true` environment string. Corrected
the condition to pass pytest a boolean before evaluating the marker. A
simulated Darwin-CI run now skips cleanly, and the sequential TCP and UDS
debugger/tooling suites each pass with 39 passed and 6 skipped.
## Files changed
- `tests/devx/test_debugger.py` - coerce the Darwin-CI skip condition to a
boolean.
## Human edits
The human pushed the preceding commits, directed the agent to continue, and
approved recording this test-only follow-up in Prompt-IO. No direct
source-line edits were made by the human.

View File

@ -1,22 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-20T13:51:25Z
git_ref: 9f99043b
diff_cmd: git diff HEAD~1..HEAD
---
Continue preparing PR #481 for landing after the test and documentation
commits were pushed. Follow the new CI run to completion and diagnose any
failures.
> `git diff HEAD~1..HEAD -- tests/devx/test_debugger.py`
Both macOS jobs failed while evaluating the new `skipif` marker. The
expression returned the `CI=true` environment string instead of a boolean,
so pytest evaluated `true` as Python source and raised `NameError` during
test setup. Coerce `_ci_env` to `bool` so pytest receives a boolean marker
condition on Darwin CI.
Verification should exercise the condition with `CI=true` and a simulated
Darwin platform, then rerun the debugger/tooling TCP and UDS suites.

View File

@ -1,43 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-20T14:38:50Z
git_ref: 559fd0f1
scope: code
substantive: true
raw_file: 20260820T143845Z_559fd0f1_prompt_io.raw.md
---
## Prompt
Continue preparing PR #481 after the latest fix was pushed. Follow CI and
proceed with clear next steps toward a green landing candidate.
## Response summary
Traced the remaining macOS UDS failure to cancellation racing transport
teardown inside the shielded framed-send path. Preserve pending cancellation
over a transport error caused by concurrent teardown, and add a deterministic
regression for that ordering. A follow-up A/B run showed the corrected
cancellation precedence changes which nested debugger intermediary is
rendered as the immediate source versus relay, so retain coverage for both
actor levels without pinning those racy roles. The adjusted UDS node passes
three consecutive runs, and both debugger/tooling transport suites pass with
39 passed and 6 skipped.
## Files changed
- `tractor/ipc/_transport.py` - deliver pending cancellation before
translating a shielded send's transport error.
- `tests/ipc/test_each_tpt.py` - reproduce cancellation followed by local
stream closure during shielded frame publication.
- `tests/devx/test_debugger.py` - accept either valid source/relay role for
each nested intermediary while retaining the actor and error assertions.
## Human edits
The human pushed the preceding fix, ran the proposed verification plan, and
reported a repeated UDS debugger failure. That report prompted the A/B
comparison and role-insensitive assertion. No direct source-line edits were
made by the human.

View File

@ -1,26 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-20T14:38:50Z
git_ref: 559fd0f1
diff_cmd: git diff HEAD~1..HEAD
---
Continue preparing PR #481 after pushing the macOS debugger skip fix.
Follow the replacement CI run and address any remaining PR-specific
failure.
> `git diff HEAD~1..HEAD -- tractor/ipc/_transport.py`
> `git diff HEAD~1..HEAD -- tests/ipc/test_each_tpt.py`
macOS UDS failed `test_reqresp_ontopof_streaming` when its two-second
`move_on_after()` scope cancelled during `stream.send('ping')`. Commit
`88a23449` shields framed `send_all()` and checks pending cancellation only
after a successful write. Concurrent transport teardown instead closed the
socket, causing `ClosedResourceError` to escape as `TransportClosed` before
the pending cancellation could be delivered.
Preserve structured cancellation precedence on the shielded send's
transport-error path, and add a deterministic regression that cancels the
sender before making the fake stream raise `ClosedResourceError`.

View File

@ -1,33 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 7b9c97c4-fff7-4ac4-97fb-35720453308e
timestamp: 2026-08-20T15:02:50Z
git_ref: pformat_caller_frame_render_guard
scope: code
substantive: true
raw_file: 20260820T150250Z_9afda1c6_prompt_io.raw.md
---
## Prompt
Fix both newly exposed send-side `MsgTypeError` formatting failures
and pin them with an end-to-end regression in PR #503.
## Response summary
Corrected codec-spec formatting and default error-message assembly so
`_mk_send_mte()` returns a printable error instead of raising another
formatter exception.
## Files changed
- `tractor/msg/_codec.py` - pass the codec to its supported formatter.
- `tractor/_exceptions.py` - assemble the default message as `str`.
- `tests/devx/test_pformat.py` - render the complete default error.
## Human edits
The human selected both one-line fixes and the single end-to-end test
as coherent additions to PR #503, while leaving broader formatter
cleanup out of scope.

View File

@ -1,25 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-20T15:02:50Z
git_ref: pformat_caller_frame_render_guard
diff_cmd: git diff HEAD~1..HEAD
---
## Prompt
After reviewing additional `tractor.devx.pformat` work suitable for
PR #503, the user approved fixing both send-side `MsgTypeError`
formatting failures and adding an end-to-end regression.
## Response
The generated code corrects the `MsgCodec.msg_spec_str` formatter
input, keeps `_mk_send_mte()`'s assembled default message a string,
and tests that the resulting `MsgTypeError` can be rendered:
> `git diff HEAD~1..HEAD -- tractor/msg/_codec.py tractor/_exceptions.py tests/devx/test_pformat.py`
These failures were hidden behind the original
`pformat_caller_frame()` keyword error addressed by the first two
commits on the branch.

View File

@ -1,65 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-21T02:35:37Z
git_ref: ae6f2ac3
scope: code
substantive: true
raw_file: 20260821T023537Z_ae6f2ac3_prompt_io.raw.md
---
## Prompt
Simplify bounded actor cancellation by passing an explicit absolute
deadline from `Portal.cancel_actor()` through `_run_from_ns()`,
`Actor.start_remote_task()`, and `Channel.send()` into
`MsgpackTransport.send()`. Avoid a `ContextVar`, watcher tasks, shared
status, coalescing, and waiter state. After tracing the current
`Start -> StartAck -> CancelAck` transaction, rename the local result to
`cancel_ack_received`, document its exact semantics, and link a focused
follow-up for a dedicated `Cancel -> CancelAck` protocol.
## Response summary
Threaded one absolute Trio deadline through the existing private
actor-cancel RPC path. The transport retains complete-frame shielding
for ordinary sends, while a cancel-control send that overruns its
deadline force-closes the potentially corrupted stream before releasing
the send lock. The outer actor-cancel scope uses the same deadline for
ack waiting and redelivers pending caller cancellation afterward.
Renamed the completion flag to `cancel_ack_received` and documented that
the current private call consumes `StartAck`, then receives a real
`CancelAck` after `Actor.cancel()` completes; this does not establish
that the OS process exited. Added a source TODO linking issue #506 for
the future first-class `Cancel -> CancelAck` transaction.
Focused transport and actor-cancel verification passed all four tests.
## Files changed
- `tractor/runtime/_portal.py` - own the absolute deadline, accurately
record ack receipt, and link the dedicated cancellation protocol.
- `tractor/runtime/_runtime.py` - forward the optional deadline for the
exact private `Start` publication.
- `tractor/ipc/_chan.py` - pass the operation-specific deadline to the
transport without changing ordinary sends.
- `tractor/ipc/_transport.py` - bound the shielded frame publication and
close a partial-frame stream before unlocking it.
- `tests/ipc/test_each_tpt.py` - cover deadline expiry after a partial
frame prefix reaches the stream.
- `tests/test_to_actor.py` - prove actor-cancel publication and ack
waiting share one absolute timeout budget.
## Human edits
The human rejected the initial watcher-task, shared `_SendStatus`, cancel
coalescing, and per-waiter design as unnecessary complexity. They also
rejected `ContextVar` propagation in favor of explicit functional
threading, selected a single absolute deadline for publication and ack
waiting, and required item 2 to remain separate from the item-3 child
reaping work. After reviewing the result, they requested the precise
`cancel_ack_received` name, a detailed protocol-trace comment, a focused
follow-up issue, and a linked source TODO. No direct source-line edits
were made by the human.

View File

@ -1,41 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-21T02:35:37Z
git_ref: ae6f2ac3
diff_cmd: git diff HEAD~1..HEAD
---
Replace the actor-cancel timeout watcher/status experiment with one
explicit absolute deadline threaded through the existing private call
path. Do not use a `ContextVar`, shared result state, waiter
coalescing, or polling tasks.
> `git diff HEAD~1..HEAD -- tractor/runtime/_portal.py`
`Portal.cancel_actor()` computes one absolute deadline and uses it for
both `Start` frame publication and the subsequent cancel-ack wait.
> `git diff HEAD~1..HEAD -- tractor/runtime/_runtime.py`
> `git diff HEAD~1..HEAD -- tractor/ipc/_chan.py`
The private RPC path forwards the operation-specific deadline. Lower
layers preserve the ordinary infinite-deadline call shape.
> `git diff HEAD~1..HEAD -- tractor/ipc/_transport.py`
`MsgpackTransport.send()` applies the deadline inside its complete-frame
shield. If the deadline expires after partial publication, it closes
the unusable stream before releasing the send lock.
> `git diff HEAD~1..HEAD -- tests/ipc/test_each_tpt.py`
> `git diff HEAD~1..HEAD -- tests/test_to_actor.py`
Focused regressions prove a partial-frame timeout closes the stream and
that actor-cancel publication and acknowledgement share one budget.
The implementation removes the earlier `_SendStatus`, watcher task,
coalescing, shared cancel result, and per-waiter state. Four focused
transport and actor-cancel tests pass.

View File

@ -1,57 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: 76c5d31c-5a2f-4503-9b16-410ee7f4fab3
timestamp: 2026-08-21T04:08:03Z
git_ref: 3c1bbe73
scope: code
substantive: true
raw_file: 20260821T040803Z_3c1bbe73_prompt_io.raw.md
---
## Prompt
Continue PR #481 review cleanup one item at a time. Compare the
remaining child-registration/reaping experiment against PR #484,
identify the next valid behavior, and generate an exact commit plan for
that boundary without committing automatically.
## Response summary
Separated late-child admission from the broader bounded-reap rewrite.
Added one synchronous `ActorNursery._register_child()` publication step
that stores the child and its reap events before sampling nursery
cancellation. This closes both scheduler orderings around
`ActorNursery.cancel()`'s child snapshot.
The MP backend registers immediately before synchronous process startup
and refuses to start when cancellation already owns the child. The Trio
backend registers immediately after `open_process()` and kills that
already-created process when registration observes cancellation. An
early `start_actor()` guard rejects calls begun after cancellation is
already visible.
Deterministic tests cover the nursery registration ordering and the MP
no-start invariant. Comparison with PR #484 confirmed that its retained
generic nursery/backends do not close this race.
## Files changed
- `tractor/runtime/_supervise.py` - atomically publish child ownership
and reject actor starts after nursery cancellation.
- `tractor/spawn/_mp.py` - register before synchronous process startup
and abort a cancellation-owned child.
- `tractor/spawn/_trio.py` - register immediately after process creation,
kill a cancellation-owned child, and remove its stale unused import.
- `tests/test_to_actor.py` - cover late registration and MP startup
suppression.
## Human edits
The human required review extras to be handled one item and one
behavioral commit at a time, with each item compared against PR #484
before acceptance. That direction split this late-registration fix from
the original broad experiment's bounded post-ack reaping,
`ActorNursery.cancel()` hard-reap rewrite, and debugger/error behavior.
The human accepted the narrower late-registration boundary by requesting
its commit plan. No direct source-line edits were made by the human.

View File

@ -1,48 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-21T04:08:03Z
git_ref: 3c1bbe73
diff_cmd: git diff HEAD~1..HEAD
---
Compare the remaining child-registration and reaping experiment with
PR #484, then identify the next review item without changing code.
The next item is the late-child admission race. A spawn can pass
`ActorNursery.start_actor()`'s early cancellation check, then be absent
from `ActorNursery.cancel()`'s child snapshot and register afterward.
The existing reap-request latch releases its monitor but does not send
runtime cancellation, so the monitor can wait forever for a still-live
process.
> `git diff HEAD~1..HEAD -- tractor/runtime/_supervise.py`
`ActorNursery._register_child()` publishes the child, installs its reap
events, and samples `ActorNursery._cancel_called` without a checkpoint.
The two scheduler orderings are then complete: registration first puts
the child in the cancel snapshot, while cancellation first makes the
backend abort the late registration.
> `git diff HEAD~1..HEAD -- tractor/spawn/_mp.py`
The multiprocessing backend registers immediately before `proc.start()`
and refuses to start a process already owned by nursery cancellation.
There is no Trio checkpoint between registration and process startup.
> `git diff HEAD~1..HEAD -- tractor/spawn/_trio.py`
The Trio backend registers immediately after `open_process()` and kills
the newly opened process if cancellation won the registration race. Its
stale unused `get_runtime_vars` import is removed so the touched module
remains lint-clean.
> `git diff HEAD~1..HEAD -- tests/test_to_actor.py`
Deterministic regressions prove late registration observes cancellation
and that the MP backend never starts a process after cancellation owns
its registration.
PR #484 retains the affected generic nursery and spawn-backend paths and
does not close this race. Keep this fix in PR #481 as its own commit;
review bounded post-`CancelAck` reaping separately.

View File

@ -1,37 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: d9d7df2c-7044-463f-8768-ec024718eac9
timestamp: 2026-08-24T22:20:33Z
git_ref: ce38cb6f
scope: code
substantive: true
raw_file: 20260824T222033Z_ce38cb6f_prompt_io.raw.md
---
## Prompt
Continue the PR #481 review after triage. The human explicitly accepted
the proposed merge-blocking `Context.cancel()` deadline update with
"keep" and required the work to remain limited to that review item.
## Response summary
Update `Context.cancel()` so one absolute deadline bounds both shielded
cancel-request publication and acknowledgement waiting. Add a focused
mocked-clock regression for the blocked-publication failure mode and run
the narrow cancellation tests.
## Files changed
- `tractor/_context.py` - forward the cancel transaction's absolute
deadline to frame publication.
- `tests/test_to_actor.py` - prove blocked context-cancel publication is
bounded by the shared deadline.
## Human edits
The human retained ownership of review scope and explicitly selected
"keep" for this item after receiving keep/defer/drop options. The human
required no unrelated cancellation changes and did not directly edit
source lines.

View File

@ -1,26 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-24T22:20:33Z
git_ref: ce38cb6f
diff_cmd: git diff HEAD~1..HEAD
---
Implement the approved PR #481 review update for `Context.cancel()`.
Use one absolute deadline for both cancellation-request frame
publication and acknowledgement waiting, without broadening the change
to unrelated cancellation behavior.
> `git diff HEAD~1..HEAD -- tractor/_context.py`
`Context.cancel()` computes one absolute cancellation deadline, uses it
for the outer bounded wait, and forwards it through
`Portal._run_from_ns()` to shielded frame publication.
> `git diff HEAD~1..HEAD -- tests/test_to_actor.py`
A deterministic mocked-clock regression arranges a shielded blocked
publication and proves that `Context.cancel()` forwards the same deadline
which bounds the complete cancel transaction.
Run the focused cancellation deadline regressions after the edit.

View File

@ -1,35 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: d9d7df2c-7044-463f-8768-ec024718eac9
timestamp: 2026-08-24T22:36:14Z
git_ref: 88d538e3
scope: code
substantive: true
raw_file: 20260824T223614Z_88d538e3_prompt_io.raw.md
---
## Prompt
Continue PR #481 review remediation after committing the shared
`Context.cancel()` deadline fix. The human accepted the proposed
child-reap bookkeeping invariant, asking only that the first fix receive
its own commit plan and commit before this update began.
## Response summary
Check that `ActorNursery` removes its paired reap-coordination entries
together while preserving valid pre-registration and immediate-cancel
paths. Extend the existing real-runtime reap tests to prove all three
child bookkeeping mappings are empty before `to_actor.run()` returns.
## Files changed
- `tractor/runtime/_supervise.py` - assert paired reap-map cleanup.
- `tests/test_to_actor.py` - verify graceful and hard-reap bookkeeping.
## Human edits
The human explicitly accepted this invariant update but directed the
preceding cancellation fix to be planned and committed as a separate
boundary first. No direct source-line edits were made by the human.

View File

@ -1,26 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-24T22:36:14Z
git_ref: 88d538e3
diff_cmd: git diff HEAD~1..HEAD
---
Implement the approved PR #481 child-reap bookkeeping update after the
preceding `Context.cancel()` fix was committed separately.
> `git diff HEAD~1..HEAD -- tractor/runtime/_supervise.py`
`ActorNursery._mark_child_reaped()` captures both reap-coordination
entries and asserts that they are either both present or both absent.
It intentionally does not require the reap-request event to be set,
because backend cancellation can reap immediately after registration.
> `git diff HEAD~1..HEAD -- tests/test_to_actor.py`
Existing real-runtime graceful and hard-reap tests verify that
`ActorNursery._children`, `ActorNursery._child_reap_requests`, and
`ActorNursery._child_reaped` are all empty before the one-shot call
returns.
Run focused bookkeeping and real-runtime reap tests after the edit.

View File

@ -1,39 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: d9d7df2c-7044-463f-8768-ec024718eac9
timestamp: 2026-08-24T22:53:56Z
git_ref: 2f86dd1a
scope: code
substantive: true
raw_file: 20260824T225356Z_2f86dd1a_prompt_io.raw.md
---
## Prompt
Continue PR #481 review remediation after committing the paired
`ActorNursery` reap-state invariant. The human selected "keep" for the
reviewer's request to factor a duplicated debugger predicate in
`_try_cancel_then_kill()`.
## Response summary
Factor the child/tree debugger predicate into a local sampler used both
before and after the cancel-RPC checkpoint. Preserve dynamic debugger
lock re-evaluation and its distinction from root-wide debug mode.
## Files changed
- `tractor/runtime/_supervise.py` - factor the duplicated debugger
predicate without changing cancellation behavior.
## Human edits
The human explicitly selected "keep" after receiving keep/defer/drop
options for this isolated review item. During commit-plan review, the
agent found that a single pre-checkpoint snapshot could become stale;
the human selected a local helper which re-evaluates the lock after the
cancel RPC. The human then considered moving the predicate into
`.devx.debug` and accepted keeping it local after confirming that no
existing helper shares its supervisor-owned semantics. No direct
source-line edits were made by the human.

View File

@ -1,18 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-24T22:53:56Z
git_ref: 2f86dd1a
diff_cmd: git diff HEAD~1..HEAD
---
Implement the approved PR #481 review refactor in
`_try_cancel_then_kill()` without changing debugger behavior.
> `git diff HEAD~1..HEAD -- tractor/runtime/_supervise.py`
Compute the child/tree debugger predicate once, reuse it in the broader
hard-kill protection predicate, and pass it directly to
`debug.maybe_wait_for_debugger()`.
Run focused debugger/cancellation coverage and lint after the edit.

View File

@ -1,36 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: d9d7df2c-7044-463f-8768-ec024718eac9
timestamp: 2026-08-24T23:39:57Z
git_ref: 5327b25e
scope: code
substantive: true
raw_file: 20260824T233957Z_5327b25e_prompt_io.raw.md
---
## Prompt
Continue PR #481 review remediation after committing the debugger-state
sampler. The human selected "keep" for the paired review request to use
`Aid` objects as keys in the newly added reap-coordination maps.
## Response summary
Migrate only `ActorNursery._child_reap_requests` and
`ActorNursery._child_reaped` to `Aid` keys. Preserve the legacy
`ActorNursery._children` `.uid` key and pass full actor identities
through the narrow process-monitor bookkeeping path.
## Files changed
- `tractor/runtime/_supervise.py` - key fresh reap maps by `Aid`.
- `tractor/spawn/_spawn.py` - pass `Aid` into completed-reap cleanup.
- `tests/test_to_actor.py` - exercise `Aid` registration keys.
## Human edits
The human explicitly selected "keep" after reviewing the scope,
performance, and mutability tradeoffs. The human retained the legacy
tuple key for `_children` and accepted `Aid` for only the two fresh
private mappings. No direct source-line edits were made by the human.

View File

@ -1,27 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-24T23:39:57Z
git_ref: 5327b25e
diff_cmd: git diff HEAD~1..HEAD
---
Implement the approved PR #481 review update which uses `Aid` keys for
the two fresh `ActorNursery` reap-coordination maps while preserving the
legacy `.uid` key for `ActorNursery._children`.
> `git diff HEAD~1..HEAD -- tractor/runtime/_supervise.py`
Type and access `_child_reap_requests` and `_child_reaped` by `Aid`.
Pass full actor identities through registration, cancellation, and
completed-reap bookkeeping, deriving `.uid` only for `_children`.
> `git diff HEAD~1..HEAD -- tractor/spawn/_spawn.py`
Forward `subactor.aid` when publishing completed process teardown.
> `git diff HEAD~1..HEAD -- tests/test_to_actor.py`
Update deterministic registration tests to exercise `Aid` map keys.
Run focused registration/reaping tests and the full `to_actor` suite.

View File

@ -1,38 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: d9d7df2c-7044-463f-8768-ec024718eac9
timestamp: 2026-08-25T01:57:42Z
git_ref: e42ecb55
scope: code
substantive: true
raw_file: 20260825T015742Z_e42ecb55_prompt_io.raw.md
---
## Prompt
Continue PR #481 review remediation after committing the `Aid` reap-map
migration. The human selected "keep" for comments explaining why both
spawn backends provisionally register children with `portal=None`.
## Response summary
Document that a child has no `Portal` until its IPC handshake yields a
`Channel`, and make `portal=None` explicit at both registration calls.
Identify the later replacement of each provisional entry with
`Portal(chan)`. Update the MP registration test double to accept and
assert the explicit provisional portal state. Name every registration
argument consistently in both backends.
## Files changed
- `tractor/spawn/_mp.py` - clarify provisional MP registration.
- `tractor/spawn/_trio.py` - clarify provisional Trio registration.
- `tests/test_to_actor.py` - model explicit provisional registration.
## Human edits
The human explicitly selected "keep" after receiving keep/defer/drop
options for this paired clarification. During local review, the human
then requested that `subactor` and `proc` also be passed by name in both
backend calls. No direct source-line edits were made by the human.

View File

@ -1,21 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-25T01:57:42Z
git_ref: e42ecb55
diff_cmd: git diff HEAD~1..HEAD
---
Implement the approved PR #481 clarification for provisional child
registration in both process-spawn backends.
> `git diff HEAD~1..HEAD -- tractor/spawn/_mp.py`
> `git diff HEAD~1..HEAD -- tractor/spawn/_trio.py`
Explain that `portal=None` is provisional because no `Portal` can exist
until the child completes its IPC handshake and returns a `Channel`.
Use an explicit keyword argument and identify the later replacement with
`Portal(chan)`.
Run lint and the full `to_actor` runtime suite.

View File

@ -1,32 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: d9d7df2c-7044-463f-8768-ec024718eac9
timestamp: 2026-08-25T02:13:19Z
git_ref: ce430fca
scope: code
substantive: true
raw_file: 20260825T021319Z_ce430fca_prompt_io.raw.md
---
## Prompt
Continue PR #481 review remediation after committing provisional child
registration clarifications. The human selected "keep" for inlining the
guarded `functools.Placeholder` lookup with a walrus assignment.
## Response summary
Remove the standalone placeholder assignment and bind the optional
Python 3.14 sentinel directly in the existing conditional while
preserving compatibility behavior.
## Files changed
- `tractor/to_actor/_api.py` - inline placeholder feature detection.
## Human edits
The human explicitly selected "keep" after receiving keep/defer/drop
options for this isolated cleanup. No direct source-line edits were made
by the human.

View File

@ -1,19 +0,0 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-25T02:13:19Z
git_ref: ce430fca
diff_cmd: git diff HEAD~1..HEAD
---
Implement the approved PR #481 review cleanup for Python 3.14 partial
placeholder detection.
> `git diff HEAD~1..HEAD -- tractor/to_actor/_api.py`
Inline the guarded `functools.Placeholder` lookup into the existing
condition with a walrus assignment, preserving fallback behavior when
the attribute is unavailable.
Run partial/placeholder normalization tests and the full `to_actor`
suite.

View File

@ -1,7 +0,0 @@
NOTE: you MUST pause this work at 12:50PM EST (BEFORE your weekly
limit reset) for review by a human!
---
attempt to resolve https://github.com/goodboy/tractor/issues/477
do it with /open-wkt.

View File

@ -37,6 +37,7 @@ Spawning actors
.. autoclass:: ActorNursery .. autoclass:: ActorNursery
:members: start_actor, :members: start_actor,
run_in_actor,
cancel, cancel,
cancel_called, cancel_called,
cancelled_caught cancelled_caught
@ -45,25 +46,11 @@ Spawning actors
:meth:`ActorNursery.start_actor` (daemon actor + portal) is the :meth:`ActorNursery.start_actor` (daemon actor + portal) is the
blessed spawning primitive; pair it with blessed spawning primitive; pair it with
:meth:`Portal.open_context` for SC-linked remote tasks. ``Portal.open_context()`` for SC-linked remote tasks.
:meth:`ActorNursery.run_in_actor` is a *convenience* one-shot —
One-shot task actors spawn, run a single task, auto-cancel after the result — slated
-------------------- to be rebuilt as a high-level wrapper, so don't design around
it as the core model.
.. autofunction:: tractor.to_actor.run
.. note::
Without ``portal=``, :func:`tractor.to_actor.run` (parlance of
``trio.to_thread.run_sync()`` and friends) is the convenience
one-shot: spawn, run one task, block on its result and reap. It
combines :meth:`ActorNursery.start_actor`, a linked
:meth:`Portal.open_context` call and per-child reaping. With
``portal=`` it owns only the linked task and leaves the existing
actor's lifetime to the portal owner; that actor must expose both
the target module and ``tractor.to_actor.MODULE``. It supersedes
the legacy, non-blocking ``ActorNursery.run_in_actor()`` retained
only for compatibility until its removal in PR #484.
.. deprecated:: 0.1.0a6 .. deprecated:: 0.1.0a6
@ -84,12 +71,14 @@ flowing back `exactly like trio`_.
:members: run, :members: run,
run_from_ns, run_from_ns,
open_stream_from, open_stream_from,
wait_for_result,
cancel_actor, cancel_actor,
chan chan
.. deprecated:: 0.1.0a6 .. deprecated:: 0.1.0a6
The str-form ``Portal.run('mod.path', 'fn_name')`` warns; ``Portal.result()`` warns; use :meth:`Portal.wait_for_result`.
The str-form ``Portal.run('mod.path', 'fn_name')`` also warns;
pass a function *object* whose module is listed in the target's pass a function *object* whose module is listed in the target's
``enable_modules``. ``Portal.channel`` is the legacy spelling ``enable_modules``. ``Portal.channel`` is the legacy spelling
of :attr:`Portal.chan`. of :attr:`Portal.chan`.

View File

@ -5,9 +5,8 @@ This is the curated reference for ``tractor``'s public surface: the
names you can import and lean on without reading runtime internals. names you can import and lean on without reading runtime internals.
Everything below is re-exported at the top level (``import Everything below is re-exported at the top level (``import
tractor``) unless a page says otherwise; subsystems like tractor``) unless a page says otherwise; subsystems like
``tractor.msg``, ``tractor.trionics``, ``tractor.to_actor``, ``tractor.msg``, ``tractor.trionics``, ``tractor.to_asyncio``,
``tractor.to_asyncio``, ``tractor.devx`` and ``tractor.log`` are ``tractor.devx`` and ``tractor.log`` are importable as submodules.
importable as submodules.
``tractor`` is "just trio_" extended across processes: every API ``tractor`` is "just trio_" extended across processes: every API
here is designed to keep the structured concurrency (SC) rules you here is designed to keep the structured concurrency (SC) rules you
@ -24,7 +23,6 @@ Most-used names at a glance:
open_root_actor open_root_actor
open_nursery open_nursery
to_actor.run
run_daemon run_daemon
ActorNursery ActorNursery
Portal Portal

View File

@ -30,12 +30,11 @@ Starting asyncio tasks from trio
.. note:: .. note::
:func:`open_channel_from` mirrors the :func:`open_channel_from` mirrors the
:meth:`tractor.Portal.open_context` handshake: the asyncio side calls ``Portal.open_context()`` handshake: the asyncio side calls
``chan.started_nowait(value)`` and that value pops out as ``chan.started_nowait(value)`` and that value pops out as
``first`` on the trio side. :func:`run_task` is the one-shot ``first`` on the trio side. :func:`run_task` is the one-shot
form — run a single asyncio-compatible coroutine fn and return form — run a single asyncio-compatible coroutine fn and return
its result to trio; :func:`tractor.to_actor.run` is its its result to trio.
cross-process sibling.
The inter-loop channel The inter-loop channel
---------------------- ----------------------

View File

@ -130,10 +130,9 @@ UDS: same-host, creds included
Pass ``enable_transports=['uds']`` and actors instead talk over Pass ``enable_transports=['uds']`` and actors instead talk over
unix-domain sockets, with socket files placed in the per-user unix-domain sockets, with socket files placed in the per-user
runtime dir: ``$XDG_RUNTIME_DIR/tractor/`` on linux, a short runtime dir (``$XDG_RUNTIME_DIR/tractor/`` on linux, the
owner-only ``/tmp/tractor-<uid>`` dir on Darwin, and the ``platformdirs`` equivalent elsewhere). Two perks over tcp on a
``platformdirs`` equivalent elsewhere. Two perks over tcp on a single single host:
host:
- no ports to fight over; addrs are just file paths, - no ports to fight over; addrs are just file paths,
- the kernel snitches on your peer for free: the listening side - the kernel snitches on your peer for free: the listening side

View File

@ -76,8 +76,8 @@ Just flip the flag on :meth:`tractor.ActorNursery.start_actor`:
infect_asyncio=True, infect_asyncio=True,
) )
The one-shot convenience ``tractor.to_actor.run()`` accepts the The one-shot convenience ``ActorNursery.run_in_actor()`` accepts
same flag. The ``to_asyncio`` APIs may **only** be called from the same flag. The ``to_asyncio`` APIs may **only** be called from
tasks inside an infected actor; calling them anywhere else raises tasks inside an infected actor; calling them anywhere else raises
a loud ``RuntimeError``. You can introspect at runtime with a loud ``RuntimeError``. You can introspect at runtime with
``tractor.current_actor().is_infected_aio()``. ``tractor.current_actor().is_infected_aio()``.
@ -234,7 +234,7 @@ dialog, skip the channel ceremony and use
It schedules the fn as an ``asyncio.Task``, waits for completion It schedules the fn as an ``asyncio.Task``, waits for completion
and hands the return value back to ``trio``; think of it as the and hands the return value back to ``trio``; think of it as the
cross-loop sibling of ``tractor.to_actor.run()``. Errors and cross-loop sibling of ``ActorNursery.run_in_actor()``. Errors and
cancellation are translated exactly as for channels. cancellation are translated exactly as for channels.
Cross-loop errors and cancellation Cross-loop errors and cancellation

View File

@ -64,13 +64,11 @@ What's going on here?
- three healthy actors are spawned as daemons via - three healthy actors are spawned as daemons via
:meth:`tractor.ActorNursery.start_actor`; left alone they'd :meth:`tractor.ActorNursery.start_actor`; left alone they'd
happily idle forever, happily idle forever,
- a fourth actor runs ``assert_err()`` via a blocking - a fourth actor runs ``assert_err()`` via ``.run_in_actor()`` and
``tractor.to_actor.run()`` one-shot and promptly trips its promptly trips its ``assert 0``,
``assert 0``,
- the resulting ``AssertionError`` ships back over IPC as a - the resulting ``AssertionError`` ships back over IPC as a
serialized error msg and re-raises *boxed* right at the call serialized error msg and re-raises *boxed* inside the nursery
inside the nursery block as a block as a :class:`tractor.RemoteActorError`,
:class:`tractor.RemoteActorError`,
- the nursery reacts like any ``trio`` nursery would: it cancels - the nursery reacts like any ``trio`` nursery would: it cancels
the three healthy siblings (graceful runtime-cancel requests, the three healthy siblings (graceful runtime-cancel requests,
acks awaited), reaps all four processes, then re-raises, acks awaited), reaps all four processes, then re-raises,
@ -230,23 +228,22 @@ Graceful first, hard as a last resort
The hard-kill path is *skipped* whenever an actor in the tree The hard-kill path is *skipped* whenever an actor in the tree
holds the debug-REPL lock (``debug_mode=True`` flavors): holds the debug-REPL lock (``debug_mode=True`` flavors):
Process signals raining down on a tree mid-``pdb`` session would SIGTERM raining down on a tree mid-``pdb`` session would
clobber your prompt. See :doc:`/guide/debugging`. clobber your prompt. See :doc:`/guide/debugging`.
Owned-child teardown in ``tractor`` begins with the same graceful Every process teardown in ``tractor`` walks the same escalation
steps, then selects the escalation path used by its supervisor, ladder, top rung first,
1. **graceful cancel request**: a runtime-cancel msg over IPC; the 1. **graceful cancel request**: a runtime-cancel msg over IPC; the
target actor cancels its tasks, closes its channels and exits target actor cancels its tasks, closes its channels and exits
its :func:`trio.run` cleanly, its :func:`trio.run` cleanly,
2. **soft wait**: the parent waits (bounded) for the child process 2. **soft wait**: the parent waits (bounded) for the child process
to exit on its own, to exit on its own,
3. **actor-nursery hard reap**: no cancel ack within the bounded wait 3. **SIGTERM**: no ack within the bounded wait (internally an
(internally an ``ActorTooSlowError``) escalates directly to ``ActorTooSlowError``) escalates to ``proc.terminate()``,
``proc.kill()`` before the child monitor joins the process, 4. **SIGKILL ultimatum**: still alive after the hard-kill timeout
4. **legacy soft-kill path**: older teardown callers may first issue (~1.6s)? The runtime logs that the "T-800" has been deployed to
``proc.terminate()`` and then deploy the "T-800" ``proc.kill()`` collect the zombie and issues ``proc.kill()``. No survivors.
ultimatum if the process survives that additional bounded wait.
The result is the **no-zombies guarantee**: ``tractor`` tries to The result is the **no-zombies guarantee**: ``tractor`` tries to
protect you from zombies, no matter what. Quoting the project protect you from zombies, no matter what. Quoting the project

View File

@ -62,10 +62,7 @@ one kwarg away,
.. code:: python .. code:: python
async with tractor.open_actor_cluster( async with tractor.open_actor_cluster(
modules=[ modules=['mylib.workers'],
'mylib.workers',
tractor.to_actor.MODULE,
],
count=4, count=4,
names=['scout', 'miner', 'smelter', 'smith'], names=['scout', 'miner', 'smelter', 'smith'],
debug_mode=True, # whole-fleet crash-to-REPL debug_mode=True, # whole-fleet crash-to-REPL
@ -73,12 +70,9 @@ one kwarg away,
... ...
From here the composition patterns are the usual ``tractor`` fare: From here the composition patterns are the usual ``tractor`` fare:
``portal.run()`` for bare one-shot RPCs (as in the demo), ``portal.run()`` for one-shot calls (as in the demo), or — for a
``tractor.to_actor.run(..., portal=portal)`` for cancellation-linked persistent bidirectional dialog per worker — concurrently enter N
one-shot tasks in an existing worker (include ``portal.open_context()`` blocks with
``tractor.to_actor.MODULE`` in ``modules``; the cluster still owns
the worker's lifetime), or — for a persistent bidirectional dialog
per worker — concurrently enter N ``portal.open_context()`` blocks with
``tractor.trionics.gather_contexts()``; see :doc:`/guide/context` ``tractor.trionics.gather_contexts()``; see :doc:`/guide/context`
for that whole layer. for that whole layer.
@ -93,8 +87,8 @@ Clusters vs. nurseries
``open_actor_cluster()`` is sugar, not a new primitive: under the ``open_actor_cluster()`` is sugar, not a new primitive: under the
hood it's just :func:`tractor.open_nursery` plus N concurrent hood it's just :func:`tractor.open_nursery` plus N concurrent
:meth:`~tractor.ActorNursery.start_actor` calls plus a ``.cancel()`` ``start_actor()`` calls plus a ``.cancel()`` on the way out. Reach
on the way out. Reach for it when, for it when,
- you want a *flat*, homogeneous fleet (classic worker-pool or - you want a *flat*, homogeneous fleet (classic worker-pool or
map-style fan-out shapes), map-style fan-out shapes),

View File

@ -15,12 +15,12 @@ a single `structured concurrency`_ (SC) scope over IPC.
:alt: sequence diagram of the context handshake msg flow :alt: sequence diagram of the context handshake msg flow
Pretty much everything else is (or is slated to be) built on this Pretty much everything else is (or is slated to be) built on this
one primitive: ``tractor.to_actor.run()`` uses it for a linked one primitive: ``ActorNursery.run_in_actor()`` is a convenience
one-shot task, spawning and reaping an actor only when no ``portal=`` for "spawn, open a context, await the result, tear down"; plain
is supplied; plain ``Portal.run()`` RPC is planned to be ``Portal.run()`` RPC is planned to be re-implemented on top of it;
re-implemented on top of it; the multi-process debugger's tree-wide the multi-process debugger's tree-wide REPL lock rides one. Grok
REPL lock rides one. Grok this page and the rest of the library reads this page and the rest of the library reads as convenience
as convenience wrappers B) wrappers B)
The endpoint contract The endpoint contract
--------------------- ---------------------

View File

@ -44,9 +44,8 @@ clan shares one registry with zero config on your part.
The bootstrap rule inside ``open_root_actor()`` is delightfully The bootstrap rule inside ``open_root_actor()`` is delightfully
simple: simple:
- on boot, probe every addr in ``registry_addrs`` with a bounded - on boot, ping every socket addr in ``registry_addrs``; when none
Tractor ``Aid`` handshake; when none are passed the per-transport are passed the per-transport defaults are used: for TCP the
defaults are used: for TCP the
loopback ``('127.0.0.1', 1616)``, for UDS a loopback ``('127.0.0.1', 1616)``, for UDS a
``registry@1616.sock`` file, ``registry@1616.sock`` file,
@ -54,11 +53,9 @@ simple:
actor and register with the *existing* registry; your own IPC actor and register with the *existing* registry; your own IPC
server binds random same-transport addrs instead, server binds random same-transport addrs instead,
- if every address is absent, congratulations: you just became the - if **nothing answers, congratulations: you just became the
registrar. Your transport server binds the registry addrs registrar**. Your transport server binds the registry addrs
themselves and you start serving lookups for everyone else, themselves and you start serving lookups for everyone else.
- if no registrar answers but an address is occupied by a foreign or
non-responsive endpoint, startup fails instead of binding over it.
Pass ``ensure_registry=True`` when your program *requires* being Pass ``ensure_registry=True`` when your program *requires* being
the one-and-only registrar; boot then fails loudly with a the one-and-only registrar; boot then fails loudly with a
@ -199,10 +196,9 @@ the existing registrar:
trio.run(main) trio.run(main)
Per the bootstrap rules above, if those addrs are absent this process Per the bootstrap rules above, if the registrar at those addrs is
becomes its own registrar root, so the same code works standalone and *not* reachable this process simply becomes its own (registrar)
as a tree-joiner. An occupied address that does not complete a Tractor root — so the same code works standalone and as a tree-joiner.
registrar handshake fails startup instead of being rebound.
"Arbiter"? A legacy naming note "Arbiter"? A legacy naming note
------------------------------- -------------------------------

View File

@ -9,8 +9,8 @@ docs; what you read is what CI runs).
Roughly in "first date to long term relationship" Roughly in "first date to long term relationship"
order, order,
- :doc:`spawning` — actor nurseries, daemons, - :doc:`spawning` — actor nurseries, daemons +
``to_actor.run()`` one-shots and process lifetimes. one-shot workers, process lifetimes.
- :doc:`rpc` — portals: calling into another - :doc:`rpc` — portals: calling into another
process like it's a local ``await``. process like it's a local ``await``.
- :doc:`context` — the cross-actor task-pair - :doc:`context` — the cross-actor task-pair

View File

@ -119,16 +119,15 @@ Run a func in a process
Even a pool can be overkill; "run this one async func in a Even a pool can be overkill; "run this one async func in a
subprocess and give me the result" is a one-liner via subprocess and give me the result" is a one-liner via
:func:`tractor.to_actor.run`, :meth:`tractor.ActorNursery.run_in_actor`,
.. literalinclude:: ../../examples/parallelism/single_func.py .. literalinclude:: ../../examples/parallelism/single_func.py
:caption: examples/parallelism/single_func.py :caption: examples/parallelism/single_func.py
:language: python :language: python
``to_actor.run()`` is a *convenience wrapper* — spawn an actor, ``run_in_actor()`` is a *convenience wrapper* — spawn an actor, run
run exactly one task in it, block on and return its result, reap exactly one task in it, reap on result — not the core spawning
— not the core spawning model (that's model (that's :meth:`tractor.ActorNursery.start_actor` plus
:meth:`tractor.ActorNursery.start_actor` plus
:meth:`tractor.Portal.open_context`; see :doc:`/guide/context`). :meth:`tractor.Portal.open_context`; see :doc:`/guide/context`).
But for this fire-and-collect shape it's exactly the right amount But for this fire-and-collect shape it's exactly the right amount
of typing. of typing.

View File

@ -80,56 +80,28 @@ One special namespace exists: ``'self'`` resolves to the remote
how internal machinery (cancel requests, registry ops) travels; how internal machinery (cancel requests, registry ops) travels;
don't build your app on it. don't build your app on it.
One-shot subactors: ``to_actor.run()`` One-shot results: ``wait_for_result()``
-------------------------------------- ---------------------------------------
When the call should own a fresh subactor whose entire job is one A portal returned from
function call, :func:`tractor.to_actor.run` spawns it, runs the task, :meth:`~tractor.ActorNursery.run_in_actor` has exactly one
returns its result and reaps the process — all in one blocking call: "main" task running remotely; that task's ``return`` value is
delivered as the portal's *final result*:
.. code:: python .. code:: python
from functools import partial portal = await an.run_in_actor(fib, n=10)
final = await portal.wait_for_result()
final = await tractor.to_actor.run(
partial(fib, n=10),
an=an,
)
Semantics worth knowing: Semantics worth knowing:
- it blocks until the remote task returns, re-raising any - it blocks until the remote task returns, re-raising any
remote error in the usual boxed form right in the calling remote error in the usual boxed form.
task. - once resolved it's idempotent: later calls return the same
- lifetime mode also determines process ownership: ``an=`` spawns and cached value.
reaps a fresh child in an existing actor nursery, while passing - a *daemon* portal (from ``start_actor()``) has no main task,
neither does the same in a private call-scoped nursery (booting so there's no final result to wait for: you'll get a warning
the runtime if needed). ``portal=`` instead runs one linked task plus a ``NoResult`` sentinel. Results of individual daemon
in an existing actor; it neither spawns nor reaps that actor, so calls come straight back from each ``await portal.run()``.
the portal's owner remains responsible for its lifetime.
- concurrency composes the plain ``trio`` way: schedule
multiple ``run()`` calls into a local task nursery (see
``examples/parallelism/concurrent_toactor_primes.py``).
A reused actor must expose both the target module and the
``to_actor`` context trampoline:
.. code:: python
async with tractor.open_nursery() as an:
portal = await an.start_actor(
'worker',
enable_modules=[
__name__,
tractor.to_actor.MODULE,
],
)
try:
final = await tractor.to_actor.run(
partial(fib, n=10),
portal=portal,
)
finally:
await portal.cancel_actor()
Pure RPC daemons: ``run_daemon()`` Pure RPC daemons: ``run_daemon()``
---------------------------------- ----------------------------------
@ -175,8 +147,7 @@ call tears down the entire sub-tree — SC, transitively.
When to graduate to ``Context`` When to graduate to ``Context``
------------------------------- -------------------------------
The :meth:`~tractor.Portal.run` method is great for one-shot, ``portal.run()`` is great for one-shot, request-response calls.
request-response calls.
Reach for :meth:`~tractor.Portal.open_context` with an Reach for :meth:`~tractor.Portal.open_context` with an
``@tractor.context`` endpoint as soon as you want: ``@tractor.context`` endpoint as soon as you want:
@ -189,15 +160,10 @@ Reach for :meth:`~tractor.Portal.open_context` with an
:meth:`~tractor.Portal.cancel_actor` nukes the **entire** :meth:`~tractor.Portal.cancel_actor` nukes the **entire**
remote runtime and its process. remote runtime and its process.
:func:`tractor.to_actor.run` already enters the full In fact the source plans for ``Portal.run()`` itself to be
:meth:`~tractor.Portal.open_context` lifecycle. The older rebuilt on top of ``open_context()`` — contexts *are* the core
:meth:`~tractor.Portal.run` path instead uses the ``Context`` returned inter-actor protocol. Take the full tour in
by the lower-level ``Actor.start_remote_task()`` directly, avoiding a :doc:`/guide/context`.
``Started`` handshake but owning less lifecycle machinery. A follow-up
should factor their shared linked-task lifecycle without requiring
``Portal.run()`` to delegate through the public context API or add
another wire message. Take the full tour in
:doc:`the context guide </guide/context>`.
.. seealso:: .. seealso::

View File

@ -91,34 +91,31 @@ somebody-ing:
What's going on here? What's going on here?
- :meth:`~tractor.ActorNursery.start_actor` forks off - ``start_actor('frank', enable_modules=[__name__])`` forks off
a new process, boots a ``tractor`` runtime inside it, and a new process, boots a ``tractor`` runtime inside it, and
allows it to serve functions from the current module (see the allows it to serve functions from the current module (see the
allowlist section below). allowlist section below).
- each :meth:`~tractor.Portal.run` call schedules a *new* task in - each ``await portal.run(...)`` schedules a *new* task in
frank's task tree and waits on its result — the full RPC story frank's task tree and waits on its result — the full RPC story
lives in :doc:`/guide/rpc`. lives in :doc:`/guide/rpc`.
- frank has no main task to complete, so without the final - frank has no main task to complete, so without the final
:meth:`~tractor.Portal.cancel_actor` call the nursery block would ``await portal.cancel_actor()`` the nursery block would wait
wait on him **forever**. Daemon lifetimes are *yours* to end; on him **forever**. Daemon lifetimes are *yours* to end; that
that explicitness is the point. explicitness is the point.
``to_actor.run()``: quick one-shot parallelism ``run_in_actor()``: quick one-shot parallelism
---------------------------------------------- ----------------------------------------------
Without ``portal=``, :func:`tractor.to_actor.run` is the convenience :meth:`~tractor.ActorNursery.run_in_actor` is the convenience
wrapper: spawn an actor, run exactly one async function in it, block wrapper: spawn an actor, run exactly one async function in it,
on the result, then reap the process — the distributed sibling of then reap the process as soon as the result arrives.
``trio.to_thread.run_sync()``.
.. code:: python .. code:: python
async with ( async with tractor.open_nursery() as an:
tractor.open_nursery() as an, portal = await an.run_in_actor(burn_cpu)
trio.open_nursery() as tn,
):
# burn rubber in the parent too... # burn rubber in the parent too...
tn.start_soon(burn_cpu) await burn_cpu()
total = await tractor.to_actor.run(burn_cpu, an=an) total = await portal.wait_for_result()
A few details worth knowing: A few details worth knowing:
@ -126,61 +123,43 @@ A few details worth knowing:
``name='something_cuter'``. ``name='something_cuter'``.
- the function's module is auto-added to the child's - the function's module is auto-added to the child's
``enable_modules`` allowlist. ``enable_modules`` allowlist.
- targets cross IPC as ``module:name`` references, so portable calls - extra ``**kwargs`` are forwarded to the function itself.
use module-global async functions or ``functools.partial`` objects - the child is *auto-cancelled* once its "main" result lands;
wrapping them. Nested functions, methods and callable objects do not at nursery exit these run-once children are always reaped
provide that stable address. first (causality_ is paramount!).
- target arguments are positional; use ``functools.partial()``
to bind target keyword arguments. Keywords passed directly to
``run()`` configure actor placement and spawning.
- the call blocks until the result (or error) lands and the
child is *auto-cancelled* (reaped) right after — so remote
errors raise directly in your calling task (causality_ is
paramount!).
- "placement" composes: ``an=`` spawns a call-owned child from an
existing actor nursery, while passing neither opens a private
call-scoped nursery. ``portal=`` instead reuses an existing actor:
the call scopes only its linked remote task, neither spawns nor
reaps the actor, and leaves its lifetime with the portal's owner.
That actor must expose both the target module and
``tractor.to_actor.MODULE``.
.. note:: .. note::
:func:`tractor.to_actor.run` is a convenience, **not** the core ``run_in_actor()`` is a convenience, **not** the core model.
model. For actor-owning placements it combines The source literally marks it for an eventual rebuild as
:meth:`~tractor.ActorNursery.start_actor`, a linked a thin "hilevel" wrapper on top of
:meth:`~tractor.Portal.open_context` call, and per-child :meth:`~tractor.Portal.open_context` (the modern inter-actor
cancellation/reaping. With ``portal=`` it uses only the linked task API). Teach your fingers to use it for quick
context call and leaves the existing actor's lifetime untouched. fire-and-collect parallelism — think a per-function
Teach your fingers to use it for quick trio-parallel_ style one-shot — and reach for
fire-and-collect parallelism — think a per-function trio-parallel_ ``start_actor()`` + ``open_context()`` for anything
style one-shot — and reach for long-lived, stateful or streaming
:meth:`~tractor.ActorNursery.start_actor` plus (:doc:`/guide/context`).
:meth:`~tractor.Portal.open_context` for anything long-lived,
stateful or streaming; see :doc:`/guide/context`.
Actor lifetimes and teardown order Actor lifetimes and teardown order
---------------------------------- ----------------------------------
There are two actor-lifetime flavors: So we have two lifetime flavors:
- **call-owned one-shot** (``to_actor.run()`` without ``portal=``): - **run-once** (``run_in_actor()``): lives exactly as long as
spawned for one task, then cancelled and joined before ``run()`` its single task; reaped the moment its result (or error)
returns its result or raises its error. arrives.
- **caller-owned daemon** (:meth:`~tractor.ActorNursery.start_actor`), - **daemon** (``start_actor()``): lives until *someone* cancels
including an actor later reused through it — an explicit ``await portal.cancel_actor()``, a bulk
``to_actor.run(..., portal=portal)``: lives until *someone* ``await an.cancel()``, or the one-cancels-all strategy kicking
cancels it via an explicit in on error.
:meth:`~tractor.Portal.cancel_actor`, a bulk
:meth:`~tractor.ActorNursery.cancel`, or the one-cancels-all
strategy kicking in on error.
On a clean exit of the nursery block the teardown order is: On a clean exit of the nursery block the teardown order is:
1. call-owned actors do not survive their own ``to_actor.run()`` 1. the nursery waits on every run-once actor's final result;
calls; each is reaped before its call returns. any errors from these are raised immediately so your code
2. the nursery waits on caller-owned daemon actors (acting as supervisor) gets first crack at handling them.
**indefinitely**. If you spawned one, you own its lifetime. 2. then it waits on daemon actors — **indefinitely**. If you
spawned a daemon, you own its lifetime.
When a child *is* cancelled, teardown is graceful-first per SC When a child *is* cancelled, teardown is graceful-first per SC
discipline: the runtime sends an IPC cancel request and gives discipline: the runtime sends an IPC cancel request and gives

View File

@ -185,10 +185,8 @@ first with a bounded grace window — so actor runtimes can run
their ``trio`` teardown paths — escalating to ``SIGKILL`` only as their ``trio`` teardown paths — escalating to ``SIGKILL`` only as
a last resort. The ``--shm`` sweep unlinks ``/dev/shm/`` segments a last resort. The ``--shm`` sweep unlinks ``/dev/shm/`` segments
that no live process has open (it leans on psutil_, already in that no live process has open (it leans on psutil_, already in
your dev venv, to check live mappings and fds) and ``--uds`` clears your dev venv, to check live mappings and fds) and ``--uds``
dead-binder sockets from Tractor's platform-specific runtime dir. It clears socket files whose binder pid is dead.
also unconditionally removes ``registry@1616.sock``; do not run the UDS
sweep while a live registrar is serving from that default address.
Testing your own ``tractor`` app Testing your own ``tractor`` app
-------------------------------- --------------------------------

View File

@ -43,21 +43,24 @@ Run it::
What's going on here? What's going on here?
- ``trio.run(main)`` starts the **root actor**; the ``tractor`` - ``trio.run(main)`` starts the **root actor**; the ``tractor``
runtime boots *implicitly* inside this ``tractor.to_actor.run()`` runtime boots *implicitly* inside ``tractor.open_nursery()``
call because neither ``an=`` nor ``portal=`` was supplied. No whenever it isn't already up. No special entrypoint, no
special entrypoint, no framework takeover - it's just a ``trio`` framework takeover - it's just a ``trio`` app,
app,
- inside ``main()`` a *subactor* is spawned via - inside ``main()`` a *subactor* is spawned via
``tractor.to_actor.run()`` and told to run exactly one ``ActorNursery.run_in_actor()`` and told to run exactly one
function: ``cellar_door()``, function: ``cellar_door()``,
- you get back a ``Portal``: your handle for invoking tasks in
the new process's (separate!) memory domain. We lean on it
much harder in the next section,
- the subactor, *some_linguist*, boots a fresh ``trio.run()`` in - the subactor, *some_linguist*, boots a fresh ``trio.run()`` in
a **new process** and executes ``cellar_door()`` as its linked a **new process** and executes ``cellar_door()`` as its *main
one-shot task (note the child proving it is *not* the root with task* (note the child proving it is *not* the root with
``tractor.is_root_process()``), then ships the return value ``tractor.is_root_process()``), then ships the return value
back over IPC, back over IPC,
- the call *blocks* until that final result arrives, then - the parent grabs that *final result* with
returns it - causality is preserved: your task only proceeds ``await portal.wait_for_result()``, much like you'd expect
once the child is *done*, dead, and reaped. from a "future" - except causality is preserved: the nursery
block only exits once the child is *done*, dead, and reaped.
.. margin:: Just need a worker pool? .. margin:: Just need a worker pool?
@ -68,20 +71,17 @@ What's going on here?
.. note:: .. note::
Without ``portal=``, ``to_actor.run()`` (parlance of ``run_in_actor()`` is the *convenience* wrapper: one-shot
``trio.to_thread`` and friends) is the *convenience* wrapper: spawn-run-reap semantics for when a subactor's entire job is
one-shot spawn-run-reap semantics for when a subactor's entire a single function call. The core primitives are
job is a single function call. The core primitives are ``ActorNursery.start_actor()`` (next up) paired with
:meth:`~tractor.ActorNursery.start_actor` (next up) — which ``Portal.open_context()`` for full, SC-linked cross-actor
hands you a ``Portal``, your handle for invoking tasks in the dialogs - see :doc:`/guide/context`.
new process's (separate!) memory domain — paired with
:meth:`~tractor.Portal.open_context` for full, SC-linked
cross-actor dialogs; see :doc:`/guide/context`.
Daemon actors and RPC Daemon actors and RPC
--------------------- ---------------------
A subactor spawned by ``to_actor.run()`` terminates after its lone A ``run_in_actor()``-spawned actor terminates when its main task
task returns. But often you want long-lived *daemon* actors instead: returns. But often you want long-lived *daemon* actors instead:
spawned once, then serving (allowlisted) RPC requests until told spawned once, then serving (allowlisted) RPC requests until told
otherwise. That's ``start_actor()``: otherwise. That's ``start_actor()``:
@ -91,17 +91,14 @@ otherwise. That's ``start_actor()``:
Two lifetime rules to internalize: Two lifetime rules to internalize:
- a subactor spawned and owned by ``to_actor.run()`` is cancelled - a ``run_in_actor()`` actor lives exactly as long as its main
and reaped before the call returns its result or raises its error, task; the nursery waits for that function (and thus the
process) to complete before unblocking,
- a ``start_actor()`` actor *lives forever* - an RPC daemon the - a ``start_actor()`` actor *lives forever* - an RPC daemon the
nursery will happily wait on **indefinitely** - until some nursery will happily wait on **indefinitely** - until some
task explicitly cancels it via ``Portal.cancel_actor()`` (as task explicitly cancels it via ``Portal.cancel_actor()`` (as
above), or its parent nursery is cancelled wholesale. above), or its parent nursery is cancelled wholesale.
Passing ``portal=`` is different: the call owns only the linked
remote task. It neither spawns nor reaps the existing actor; the
portal's owner must end that actor's lifetime.
.. tip:: .. tip::
Want your *entire program* to just be a long-lived RPC Want your *entire program* to just be a long-lived RPC
@ -211,20 +208,16 @@ The script of the scene (runtime ``INFO`` log lines trimmed)::
The new tricks in play: The new tricks in play:
- *donny* and *gretchen* start as daemon actors so each remains alive - two subactors, *donny* and *gretchen*, are each told to run
while the other discovers it and completes its line, ``say_hello()`` targeting the *other* by name,
- a local ``trio`` nursery runs both ``Portal.run(say_hello)`` calls
concurrently; starting both actors first avoids either reciprocal
dialog racing one-shot process reaping,
- ``tractor.wait_for_actor()`` blocks until the named peer has - ``tractor.wait_for_actor()`` blocks until the named peer has
registered with the tree's *registrar* (every actor announces registered with the tree's *registrar* (every actor announces
itself at boot), then yields a ``Portal`` connected itself at boot), then yields a ``Portal`` connected
**directly** to that peer, **directly** to that peer,
- each actor invokes its partner's ``hi()`` over that portal: - each actor invokes its partner's ``hi()`` over that portal:
actor-to-actor RPC with the root merely *directing* - and each actor-to-actor RPC with the root merely *directing* - and both
``Portal.run()`` returns its final line directly to ``main()``, final lines flow back to ``main()`` via
- the actor nursery explicitly cancels both daemons only after both ``await portal.wait_for_result()``,
dialogs complete,
- ``tractor.log.get_console_log("INFO")`` cranks up runtime - ``tractor.log.get_console_log("INFO")`` cranks up runtime
logging so you can watch the spawn/register/cancel machinery logging so you can watch the spawn/register/cancel machinery
narrate itself; remove it for a quiet set. narrate itself; remove it for a quiet set.

View File

@ -21,38 +21,22 @@ async def main():
"""Main tractor entry point, the "master" process (for now """Main tractor entry point, the "master" process (for now
acts as the "director"). acts as the "director").
""" """
async with tractor.open_nursery() as an: async with tractor.open_nursery() as n:
print("Alright... Action!") print("Alright... Action!")
# both actors wait on (then dial!) the *other*, so each donny = await n.run_in_actor(
# must outlive both hellos: spawn as daemons, run the
# hellos concurrently, reap only once both complete.
portals: dict[str, tractor.Portal] = {
name: await an.start_actor(
name,
enable_modules=[__name__],
)
for name in ('donny', 'gretchen')
}
async def run_and_print(
name: str,
other_actor: str,
) -> None:
print(
# RPC through an existing actor's `Portal`.
await portals[name].run(
say_hello, say_hello,
other_actor=other_actor, name='donny',
# arguments are always named
other_actor='gretchen',
) )
gretchen = await n.run_in_actor(
say_hello,
name='gretchen',
other_actor='donny',
) )
print(await gretchen.wait_for_result())
async with trio.open_nursery() as tn: print(await donny.wait_for_result())
tn.start_soon(run_and_print, 'donny', 'gretchen')
tn.start_soon(run_and_print, 'gretchen', 'donny')
await an.cancel()
print("CUTTTT CUUTT CUT!!! Donny!! You're supposed to say...") print("CUTTTT CUUTT CUT!!! Donny!! You're supposed to say...")

View File

@ -10,14 +10,17 @@ async def cellar_door():
async def main(): async def main():
"""The main ``tractor`` routine. """The main ``tractor`` routine.
""" """
# spawn a subactor, run ``cellar_door()`` as its lone task, async with tractor.open_nursery() as n:
# block until its result arrives and the subactor is reaped.
print( portal = await n.run_in_actor(
await tractor.to_actor.run(
cellar_door, cellar_door,
name='some_linguist', name='some_linguist',
) )
)
# The ``async with`` will unblock here since the 'some_linguist'
# actor has completed its main task ``cellar_door``.
print(await portal.wait_for_result())
if __name__ == '__main__': if __name__ == '__main__':

View File

@ -1,5 +1,3 @@
from functools import partial
import trio import trio
import tractor import tractor
@ -23,41 +21,26 @@ async def breakpoint_forever():
async def spawn_until(depth=0): async def spawn_until(depth=0):
""""A nested nursery that triggers another ``NameError``. """"A nested nursery that triggers another ``NameError``.
""" """
async with ( async with tractor.open_nursery() as n:
tractor.open_nursery() as an,
trio.open_nursery() as tn,
):
if depth < 1: if depth < 1:
tn.start_soon( await n.run_in_actor(breakpoint_forever)
partial(
tractor.to_actor.run,
breakpoint_forever,
an=an,
)
)
# Let the background one-shot enter `breakpoint_forever()` p = await n.run_in_actor(
# before its sibling raises and cancellation propagates. name_error,
name='name_error'
)
await trio.sleep(0.5) await trio.sleep(0.5)
# rx and propagate error from child # rx and propagate error from child
await tractor.to_actor.run( await p.result()
name_error,
an=an,
name='name_error',
)
else: else:
# recusrive call to spawn another process branching layer of # recusrive call to spawn another process branching layer of
# the tree; blocks (up) each level until the leaf's # the tree
# `name_error` relays through.
depth -= 1 depth -= 1
await tractor.to_actor.run( await n.run_in_actor(
partial(
spawn_until, spawn_until,
depth=depth, depth=depth,
),
an=an,
name=f'spawn_until_{depth}', name=f'spawn_until_{depth}',
) )
@ -82,37 +65,34 @@ async def main():
python -m tractor._child --uid ('spawn_until_0', 'de918e6d ...) python -m tractor._child --uid ('spawn_until_0', 'de918e6d ...)
""" """
async with ( async with tractor.open_nursery(
tractor.open_nursery(
debug_mode=True, debug_mode=True,
loglevel='pdb', loglevel='pdb',
) as an, ) as n:
trio.open_nursery() as tn,
): # spawn both actors
# spawn both spawner trees as concurrent one-shots; the portal = await n.run_in_actor(
# first tree's (relayed) error cancels the other.
tn.start_soon(
partial(
tractor.to_actor.run,
partial(
spawn_until, spawn_until,
depth=3, depth=3,
),
an=an,
name='spawner0', name='spawner0',
) )
) portal1 = await n.run_in_actor(
tn.start_soon(
partial(
tractor.to_actor.run,
partial(
spawn_until, spawn_until,
depth=4, depth=4,
),
an=an,
name='spawner1', name='spawner1',
) )
)
# TODO: test this case as well where the parent don't see
# the sub-actor errors by default and instead expect a user
# ctrl-c to kill the root.
with trio.move_on_after(3):
await trio.sleep_forever()
# gah still an issue here.
await portal.result()
# should never get here
await portal1.result()
if __name__ == '__main__': if __name__ == '__main__':

View File

@ -15,12 +15,12 @@ async def name_error():
async def spawn_error(): async def spawn_error():
""""A nested nursery that triggers another ``NameError``. """"A nested nursery that triggers another ``NameError``.
""" """
async with tractor.open_nursery() as an: async with tractor.open_nursery() as n:
return await tractor.to_actor.run( portal = await n.run_in_actor(
name_error, name_error,
an=an,
name='name_error_1', name='name_error_1',
) )
return await portal.result()
async def main(): async def main():
@ -38,36 +38,29 @@ async def main():
- root actor should then fail on assert - root actor should then fail on assert
- program termination - program termination
""" """
async with ( async with tractor.open_nursery(
tractor.open_nursery(
debug_mode=True, debug_mode=True,
loglevel='devx', loglevel='devx',
) as an, ) as n:
trio.open_nursery() as tn,
):
# spawn both actors..
portal = await an.start_actor(
'name_error',
enable_modules=[__name__],
)
portal1 = await an.start_actor(
'spawn_error',
enable_modules=[__name__],
)
# ..and bg-schedule their erroring tasks. # spawn both actors
tn.start_soon(portal.run, name_error) portal = await n.run_in_actor(
tn.start_soon(portal1.run, spawn_error) name_error,
name='name_error',
# yield to the bg tasks so both RPC requests are )
# submitted (and start crashing) before the root's own portal1 = await n.run_in_actor(
# error below (the legacy `run_in_actor()` submitted spawn_error,
# in-line with each spawn). name='spawn_error',
await trio.sleep(0.5) )
# trigger a root actor error # trigger a root actor error
assert 0 assert 0
# attempt to collect results (which raises error in parent)
# still has some issues where the parent seems to get stuck
await portal.result()
await portal1.result()
if __name__ == '__main__': if __name__ == '__main__':
trio.run(main) trio.run(main)

View File

@ -17,12 +17,12 @@ async def name_error():
async def spawn_error(): async def spawn_error():
""""A nested nursery that triggers another ``NameError``. """"A nested nursery that triggers another ``NameError``.
""" """
async with tractor.open_nursery() as an: async with tractor.open_nursery() as n:
return await tractor.to_actor.run( portal = await n.run_in_actor(
name_error, name_error,
an=an,
name='name_error_1', name='name_error_1',
) )
return await portal.result()
async def main(): async def main():
@ -36,39 +36,17 @@ async def main():
`-python -m tractor._child --uid ('spawn_error', '52ee14a5 ...) `-python -m tractor._child --uid ('spawn_error', '52ee14a5 ...)
`-python -m tractor._child --uid ('name_error', '3391222c ...) `-python -m tractor._child --uid ('name_error', '3391222c ...)
""" """
errors: list[BaseException] = []
async with tractor.open_nursery( async with tractor.open_nursery(
debug_mode=True, debug_mode=True,
# loglevel='runtime', # loglevel='runtime',
) as an: ) as n:
async def run_and_collect(fn): # Spawn both actors, don't bother with collecting results
''' # (would result in a different debugger outcome due to parent's
One-shot whose (boxed) error is stashed instead of # cancellation).
raised so a sibling's crash never cancels the others await n.run_in_actor(breakpoint_forever)
before they've had their own debugger sessions (the await n.run_in_actor(name_error)
"collect all errors" the legacy `run_in_actor()` API await n.run_in_actor(spawn_error)
did implicitly at nursery teardown).
'''
try:
await tractor.to_actor.run(fn, an=an)
except tractor.RemoteActorError as rae:
errors.append(rae)
# Spawn all one-shot task actors, collecting (vs.
# raising) their errors.
async with trio.open_nursery() as tn:
tn.start_soon(run_and_collect, breakpoint_forever)
tn.start_soon(run_and_collect, name_error)
tn.start_soon(run_and_collect, spawn_error)
if errors:
raise BaseExceptionGroup(
'multi_subactors errored!',
errors,
)
if __name__ == '__main__': if __name__ == '__main__':

View File

@ -1,5 +1,3 @@
from functools import partial
import trio import trio
import tractor import tractor
@ -12,17 +10,15 @@ async def name_error():
async def spawn_until(depth=0): async def spawn_until(depth=0):
""""A nested nursery that triggers another ``NameError``. """"A nested nursery that triggers another ``NameError``.
""" """
async with tractor.open_nursery() as an: async with tractor.open_nursery() as n:
if depth < 1: if depth < 1:
await tractor.to_actor.run(name_error, an=an) # await n.run_in_actor('breakpoint_forever', breakpoint_forever)
await n.run_in_actor(name_error)
else: else:
depth -= 1 depth -= 1
await tractor.to_actor.run( await n.run_in_actor(
partial(
spawn_until, spawn_until,
depth=depth, depth=depth,
),
an=an,
name=f'spawn_until_{depth}', name=f'spawn_until_{depth}',
) )
@ -41,37 +37,28 @@ async def main():
python -m tractor._child --uid ('name_error', '6c2733b8 ...) python -m tractor._child --uid ('name_error', '6c2733b8 ...)
''' '''
async with ( async with tractor.open_nursery(
tractor.open_nursery(
debug_mode=True, debug_mode=True,
enable_transports=['uds'], # TODO, pass this via osenv? enable_transports=['uds'], # TODO, apss this via osenv?
loglevel='devx', # XXX, required for test! loglevel='devx', # XXX, required for test!
) as an, ) as n:
trio.open_nursery() as tn,
):
# spawn the deeper tree in the bg..
tn.start_soon(
partial(
tractor.to_actor.run,
partial(
spawn_until,
depth=1,
),
an=an,
name='spawner1',
)
)
# ..while blocking on the shallow (faster to fail) tree # spawn both actors
# whose propagated error triggers nursery cancellation. portal = await n.run_in_actor(
await tractor.to_actor.run(
partial(
spawn_until, spawn_until,
depth=0, depth=0,
),
an=an,
name='spawner0', name='spawner0',
) )
portal1 = await n.run_in_actor(
spawn_until,
depth=1,
name='spawner1',
)
# nursery cancellation should be triggered due to propagated
# error from child.
await portal.result()
await portal1.result()
if __name__ == '__main__': if __name__ == '__main__':

View File

@ -13,24 +13,17 @@ async def main():
simultaneously. simultaneously.
''' '''
async with ( async with tractor.open_nursery(
tractor.open_nursery(
debug_mode=True, debug_mode=True,
# loglevel='debug' # ?XXX required? # loglevel='debug' # ?XXX required?
) as an, ) as n:
trio.open_nursery() as tn,
): # spawn both actors
# spawn the actor.. portal = await n.run_in_actor(key_error)
portal = await an.start_actor(
'key_error',
enable_modules=[__name__],
)
print( print(
f'Child is up @ {portal.chan.aid.reprol()}' f'Child is up @ {portal.chan.aid.reprol()}'
) )
# ..then schedule its erroring task in the bg while the
# root blocks below.
tn.start_soon(portal.run, key_error)
# XXX: originally a bug caused by this is where root would enter # XXX: originally a bug caused by this is where root would enter
# the debugger and clobber the tty used by the repl even though # the debugger and clobber the tty used by the repl even though

View File

@ -74,11 +74,11 @@ async def cancelled_before_pause(
async def main(): async def main():
async with tractor.open_nursery( async with tractor.open_nursery(
debug_mode=True, debug_mode=True,
) as an: ) as n:
await tractor.to_actor.run( portal: tractor.Portal = await n.run_in_actor(
cancelled_before_pause, cancelled_before_pause,
an=an,
) )
await portal.wait_for_result()
# ensure the same works in the root actor! # ensure the same works in the root actor!
await pm_on_cancelled() await pm_on_cancelled()

View File

@ -17,14 +17,12 @@ async def main():
async with tractor.open_nursery( async with tractor.open_nursery(
debug_mode=True, debug_mode=True,
loglevel='cancel', loglevel='cancel',
) as an: ) as n:
# parks awaiting a result which only arrives once the portal = await n.run_in_actor(
# user quits (`BdbQuit`s) the child's REPL loop.
await tractor.to_actor.run(
breakpoint_forever, breakpoint_forever,
an=an,
) )
await portal.wait_for_result()
if __name__ == '__main__': if __name__ == '__main__':

View File

@ -12,12 +12,16 @@ async def main():
) as an: ) as an:
# TODO: ideally the REPL arrives at this frame in the parent, # TODO: ideally the REPL arrives at this frame in the parent,
# ABOVE the @api_frame of `to_actor.run()` .. # ABOVE the @api_frame of `Portal.run_in_actor()` (which
# should eventually not even be a portal method ... XD)
# await tractor.pause() # await tractor.pause()
p: tractor.Portal = await an.run_in_actor(name_error)
# the one-shot blocks on the subactor's result so the # with this style, should raise on this line
# boxed `NameError` raises right here. await p.wait_for_result()
await tractor.to_actor.run(name_error, an=an)
# with this alt style should raise at `open_nusery()`
# return await p.wait_for_result()
if __name__ == '__main__': if __name__ == '__main__':

View File

@ -90,7 +90,7 @@ async def main() -> None:
# TODO: 3 sub-actor usage cases: # TODO: 3 sub-actor usage cases:
# -[x] via a `.open_context()` # -[x] via a `.open_context()`
# -[ ] via a `to_actor.run()` call # -[ ] via a `.run_in_actor()` call
# -[ ] via a `.run()` # -[ ] via a `.run()`
# -[ ] via a `.to_thread.run_sync()` in subactor # -[ ] via a `.to_thread.run_sync()` in subactor
async with p.open_context( async with p.open_context(

View File

@ -1,83 +0,0 @@
'''
`tractor.to_actor.run()`: concurrent one-shot prime checks
invocation, the SC-parallelism sibling of
`trio.to_thread.run_sync()` (and `anyio.to_process`).
Each call spawns a subactor, schedules the async fn as
its lone remote task, waits on the result and reaps the
subactor. Concurrency composes the plain `trio` way:
schedule multiple one-shot calls in a local task nursery
against a shared actor-nursery; any remote error raises
directly in the task which scheduled it.
'''
import math
import tractor
import trio
async def is_prime(
n: int,
) -> bool:
if n < 2:
return False
if n == 2:
return True
if n % 2 == 0:
return False
sqrt_n = int(math.floor(math.sqrt(n)))
for i in range(3, sqrt_n + 1, 2):
if n % i == 0:
return False
return True
async def main() -> None:
# fully implicit one-shot: boots the actor-runtime,
# spawns a subactor, runs the task, reaps the
# subactor, tears the runtime back down.
assert await tractor.to_actor.run(
is_prime,
2,
)
# the "worker-pool-ish" pattern from the original
# `concurrent.futures` example: one subactor per
# input, all concurrent, results and errors
# collected by caller-side tasks.
results: dict[int, bool] = {}
async def check(
an: tractor.ActorNursery,
n: int,
i: int,
) -> None:
results[n] = await tractor.to_actor.run(
is_prime,
n,
an=an,
name=f'prime_checker_{i}',
)
inputs: list[int] = [
7,
8,
3691,
3693,
]
async with (
tractor.open_nursery() as an,
trio.open_nursery() as tn,
):
for i, n in enumerate(inputs):
tn.start_soon(check, an, n, i)
for n, prime in sorted(results.items()):
print(f'{n} is prime: {prime}')
if __name__ == '__main__':
trio.run(main)

View File

@ -20,20 +20,22 @@ async def burn_cpu():
for _ in range(50000): for _ in range(50000):
await trio.sleep(1/50000/50) await trio.sleep(1/50000/50)
return pid return os.getpid()
async def main(): async def main():
async with trio.open_nursery() as tn: async with tractor.open_nursery() as n:
portal = await n.run_in_actor(burn_cpu)
# burn rubber in the parent too # burn rubber in the parent too
tn.start_soon(burn_cpu) await burn_cpu()
# run the same func as the lone task in a subactor, # wait on result from target function
# block on and collect its PID as the caller-side result pid = await portal.wait_for_result()
pid = await tractor.to_actor.run(burn_cpu)
# end of nursery block
print(f"Collected subproc {pid}") print(f"Collected subproc {pid}")

View File

@ -23,10 +23,18 @@ async def endpoint(
await trio.sleep_forever() await trio.sleep_forever()
async def open_ep( async def spawn_and_open_ep(
ptl: tractor.Portal, an: tractor.ActorNursery,
i: int, i: int,
) -> None: ) -> None:
'''
Spawn a subactor, start a remote `endpoint()`-task in it.
'''
ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
ctx: tractor.Context ctx: tractor.Context
async with ptl.open_context(endpoint) as ( async with ptl.open_context(endpoint) as (
ctx, ctx,
@ -39,33 +47,7 @@ async def open_ep(
await ctx.wait_for_result() await ctx.wait_for_result()
async def spawn_and_open_ep( async def main():
an: tractor.ActorNursery,
i: int,
maybe_ptl: tractor.Portal|None = None,
) -> None:
'''
Spawn a subactor, start a remote `endpoint()`-task in it.
'''
if maybe_ptl is None:
maybe_ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
await open_ep(
ptl=maybe_ptl,
i=i,
)
async def main(
# spawn subs concurrently (in bg `trio.Task`s) so each
# actor's cold `import tractor` (~0.4s, see #470) overlaps
# instead of stacking; once forkserver (#463) lands, spawn
# is cheap enough to just loop sequentially.
spawn_subs_in_bg_tasks: bool = True,
):
''' '''
Spawn a subactor-per-CPU then self-destruct the cluster. Spawn a subactor-per-CPU then self-destruct the cluster.
@ -78,21 +60,17 @@ async def main(
# https://github.com/goodboy/tractor/pull/463 # https://github.com/goodboy/tractor/pull/463
# start_method='main_thread_forkserver', # start_method='main_thread_forkserver',
) as an, ) as an,
# spawn subs concurrently (in bg `trio.Task`s) so each
# actor's cold `import tractor` (~0.4s, see #470) overlaps
# instead of stacking; once forkserver (#463) lands, spawn
# is cheap enough to just loop sequentially.
trio.open_nursery() as tn, trio.open_nursery() as tn,
): ):
for i in range(cpu_count()): for i in range(cpu_count()):
maybe_ptl: tractor.Portal|None = None
if not spawn_subs_in_bg_tasks:
maybe_ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
tn.start_soon( tn.start_soon(
spawn_and_open_ep, spawn_and_open_ep,
an, an,
i, i,
maybe_ptl,
) )
destruct_in: int = 2 destruct_in: int = 2
print( print(

View File

@ -7,20 +7,19 @@ async def assert_err():
async def main(): async def main():
async with tractor.open_nursery() as an: async with tractor.open_nursery() as n:
real_actors = [] real_actors = []
for i in range(3): for i in range(3):
real_actors.append(await an.start_actor( real_actors.append(await n.start_actor(
f'actor_{i}', f'actor_{i}',
enable_modules=[__name__], enable_modules=[__name__],
)) ))
# run one one-shot task actor that will fail immediately; # start one actor that will fail immediately
# its error raises right here in the caller's task.. await n.run_in_actor(assert_err)
await tractor.to_actor.run(assert_err, an=an)
# ..as a ``RemoteActorError`` containing an ``AssertionError`` # should error here with a ``RemoteActorError`` containing
# and all the other actors have been cancelled # an ``AssertionError`` and all the other actors have been cancelled
if __name__ == '__main__': if __name__ == '__main__':

View File

@ -6,8 +6,7 @@ subactor inherits the preference.
Every channel address is a filesystem socket path (no TCP port Every channel address is a filesystem socket path (no TCP port
in sight!) and, as a kernel-provided bonus, the peer's pid is in sight!) and, as a kernel-provided bonus, the peer's pid is
exchanged for free via `SO_PEERCRED` on linux, exchanged for free via `SO_PEERCRED`.
`LOCAL_PEERPID` on macOS.
''' '''
import os import os
@ -43,7 +42,7 @@ async def main() -> None:
# (named for the root registrar) this channel rode in # (named for the root registrar) this channel rode in
# on, NOT a per-child path; the child-specific identity # on, NOT a per-child path; the child-specific identity
# we get for free is the kernel-reported peer pid (via # we get for free is the kernel-reported peer pid (via
# `SO_PEERCRED` on linux, `LOCAL_PEERPID` on macOS). # `SO_PEERCRED`).
print( print(
f'portal chan tpt proto: {raddr.proto_key!r}\n' f'portal chan tpt proto: {raddr.proto_key!r}\n'
f'listener sock file: {raddr.sockpath}\n' f'listener sock file: {raddr.sockpath}\n'

View File

@ -1,4 +0,0 @@
Fix Unix-domain-socket actor trees and registrar discovery on macOS.
Runtime sockets now use a short, owner-only runtime directory,
generated socket names remain within platform limits, and transient
or reset pre-handshake connections no longer destabilize discovery.

View File

@ -1,3 +0,0 @@
Add ``tractor.to_actor.run()`` for Trio-style one-shot async calls in
new or existing actors, with caller-scoped result/error propagation,
linked cancellation, and deterministic reaping of call-owned children.

View File

@ -23,10 +23,10 @@ Two cleanup phases (run in order when both are enabled):
hard-crashing actor leaves leaked segments that hard-crashing actor leaves leaked segments that
nothing else GCs. nothing else GCs.
3. **UDS sweep** (`--uds` / `--uds-only`) — unlinks socket 3. **UDS sweep** (`--uds` / `--uds-only`) — unlinks
files from Tractor's platform-specific default bindspace whose `${XDG_RUNTIME_DIR}/tractor/<name>@<pid>.sock` files
binder pid is dead (or the `1616` registry sentinel). Needed whose binder pid is dead (or the `1616` registry
because the IPC server's sentinel). Needed because the IPC server's
`os.unlink()` cleanup lives in a `finally:` block `os.unlink()` cleanup lives in a `finally:` block
that doesn't always run on hard exits (SIGKILL, that doesn't always run on hard exits (SIGKILL,
escaped `KeyboardInterrupt`, etc.) — see issue #452. escaped `KeyboardInterrupt`, etc.) — see issue #452.
@ -137,8 +137,8 @@ def main() -> int:
action='store_true', action='store_true',
help=( help=(
'after process reap, also unlink orphaned ' 'after process reap, also unlink orphaned '
'sockets from Tractor\'s platform default ' '${XDG_RUNTIME_DIR}/tractor/*.sock files '
'bindspace whose binder pid is dead (or the 1616 ' 'whose binder pid is dead (or the 1616 '
'registry sentinel). See issue #452.' 'registry sentinel). See issue #452.'
), ),
) )
@ -212,9 +212,7 @@ def main() -> int:
# --- phase 3: UDS sweep (opt-in) --- # --- phase 3: UDS sweep (opt-in) ---
if args.uds or args.uds_only: if args.uds or args.uds_only:
leaked_uds: list[str] = find_orphaned_uds( leaked_uds: list[str] = find_orphaned_uds()
include_registry_sentinel=True,
)
if not leaked_uds: if not leaked_uds:
print( print(
'[tractor-reap] no orphaned UDS sock-files ' '[tractor-reap] no orphaned UDS sock-files '

View File

@ -1,53 +0,0 @@
'''
Shared helpers for actor-runtime test suites.
'''
from pathlib import Path
from types import TracebackType
import tractor
import trio
class CancellationMarkers:
'''
Mark a test endpoint and require cancellation-driven teardown.
'''
def __init__(
self,
started_path: str,
cancelled_path: str,
) -> None:
self.started_path = started_path
self.cancelled_path = cancelled_path
def __enter__(self) -> None:
Path(self.started_path).touch()
def __exit__(
self,
exc_type: type[BaseException]|None,
exc_value: BaseException|None,
traceback: TracebackType|None,
) -> None:
assert exc_type is trio.Cancelled
assert isinstance(exc_value, trio.Cancelled)
Path(self.cancelled_path).touch()
def non_registration_contexts(
actor: tractor.Actor,
) -> dict[tuple, str]:
'''
Snapshot application contexts without registrar-service traffic.
'''
return {
key: str(ctx._nsf)
for key, ctx in actor._contexts.items()
if str(ctx._nsf) != (
'tractor.discovery._registry:'
'Registrar.register_actor'
)
}

View File

@ -27,7 +27,6 @@ from pexpect.exceptions import (
import tractor import tractor
from .conftest import ( from .conftest import (
ansi_strip,
do_ctlc, do_ctlc,
PROMPT, PROMPT,
_pause_msg, _pause_msg,
@ -769,15 +768,6 @@ def test_multi_subactors_root_errors(
@has_nested_actors @has_nested_actors
@pytest.mark.skipif(
platform.system() == 'Darwin'
and
bool(_ci_env),
reason=(
'Nested crash-REPL ordering is unreliable on macOS CI; '
'see https://github.com/goodboy/tractor/issues/320'
),
)
def test_multi_nested_subactors_error_through_nurseries( def test_multi_nested_subactors_error_through_nurseries(
ci_env: bool, ci_env: bool,
spawn: PexpectSpawner, spawn: PexpectSpawner,
@ -804,7 +794,6 @@ def test_multi_nested_subactors_error_through_nurseries(
loglevel='pdb', loglevel='pdb',
) )
last_send_char: str|None = None last_send_char: str|None = None
transcript_parts: list[str] = []
# inflate pexpect waits under CPU throttle — incl. the # inflate pexpect waits under CPU throttle — incl. the
# sustained-load power-cap invisible to static freq reads — so # sustained-load power-cap invisible to static freq reads — so
@ -844,9 +833,6 @@ def test_multi_nested_subactors_error_through_nurseries(
PROMPT, PROMPT,
timeout=timeout, timeout=timeout,
) )
transcript_parts.append(
ansi_strip(child.before.decode())
)
delay: float = 0.1 delay: float = 0.1
test_log.info('Sleeping {delay!r} before next send-chart..') test_log.info('Sleeping {delay!r} before next send-chart..')
time.sleep(delay) time.sleep(delay)
@ -856,9 +842,6 @@ def test_multi_nested_subactors_error_through_nurseries(
# script finally exited with tb on console. # script finally exited with tb on console.
except EOF: except EOF:
transcript_parts.append(
ansi_strip(child.before.decode())
)
test_log.info( test_log.info(
f'Breaking from send-char loop' f'Breaking from send-char loop'
f'last_send_char: {last_send_char!r}\n' f'last_send_char: {last_send_char!r}\n'
@ -892,33 +875,24 @@ def test_multi_nested_subactors_error_through_nurseries(
# happening but ONLY WHEN RUN FROM THE TEST, bc when i try to # happening but ONLY WHEN RUN FROM THE TEST, bc when i try to
# run the test script manually the correct output ALWAYS seems # run the test script manually the correct output ALWAYS seems
# to be in the last `str(child.before.decode())` output !?!? # to be in the last `str(child.before.decode())` output !?!?
transcript: str = '\n'.join(transcript_parts)
if ( if (
not is_forking_spawner not is_forking_spawner
and and
last_send_char == 'q' last_send_char == 'q'
): ):
# Cancellation can swap which intermediary is rendered as expect_patts += [
# the immediate source vs. relay. Require both actor levels # expect the pdb-quit exc.
# below without pinning those racy roles. "bdb.BdbQuit",
expect_patts.append('bdb.BdbQuit') # BUT WHY these dude!?
for uid in ( "src_uid=('spawn_until_0'",
'spawn_until_0', "relay_uid=('spawn_until_1'",
'spawn_until_1', ]
):
assert any(
role in transcript
for role in (
f"src_uid=('{uid}'",
f"relay_uid=('{uid}'",
)
)
for part in expect_patts: assert_before(
assert part in transcript child,
expect_patts,
assert child.flag_eof )
assert not child.isalive() expect(child, EOF)
# @pytest.mark.timeout(15) # @pytest.mark.timeout(15)
@ -1309,8 +1283,13 @@ def test_ctxep_pauses_n_maybe_ipc_breaks(
) )
child.sendline('c') child.sendline('c')
child.expect(EOF) child.expect(EOF)
assert child.flag_eof assert_before(
assert not child.isalive() child,
["tractor._exceptions.RemoteActorError: remote task raised a 'BdbQuit'",
"bdb.BdbQuit",
"('bp_boi'",
]
)
break # end-of-test break # end-of-test
child.sendline('c') child.sendline('c')
@ -1328,41 +1307,29 @@ def test_ctxep_pauses_n_maybe_ipc_breaks(
if _non_linux: if _non_linux:
tpt: str = 'TCP' tpt: str = 'TCP'
before: str = assert_before( assert_before(
child, child,
['peer IPC channel closed abruptly?', ['peer IPC channel closed abruptly?',
'another task closed this fd', 'another task closed this fd',
'Debug lock request was CANCELLED?', 'Debug lock request was CANCELLED?',
f"'Msgpack{tpt}Stream' was already closed locally?",
f"TransportClosed: 'Msgpack{tpt}Stream' was already closed 'by peer'?",
] ]
# XXX races on whether these show/hit? # XXX races on whether these show/hit?
# 'Failed to REPl via `_pause()` You called `tractor.pause()` from an already cancelled scope!', # 'Failed to REPl via `_pause()` You called `tractor.pause()` from an already cancelled scope!',
# 'AssertionError', # 'AssertionError',
) )
# Error shipment and peer receive race after local close.
# Either diagnostic proves the transport was torn down.
closed_locally: str = (
f"'Msgpack{tpt}Stream' was already closed locally?"
)
closed_by_peer: str = (
f"TransportClosed: 'Msgpack{tpt}Stream' was "
f"already closed 'by peer'?"
)
assert (
closed_locally in before
or closed_by_peer in before
)
# OSc(ancel) the hanging tree # OSc(ancel) the hanging tree
do_ctlc( do_ctlc(
child=child, child=child,
expect_prompt=False, expect_prompt=False,
) )
child.expect(EOF) child.expect(EOF)
before += ansi_strip(child.before.decode()) assert_before(
assert 'KeyboardInterrupt' in before child,
assert child.flag_eof ['KeyboardInterrupt'],
assert not child.isalive() )
def test_crash_handling_within_cancelled_root_actor( def test_crash_handling_within_cancelled_root_actor(

View File

@ -1,84 +0,0 @@
'''
Unit tests for the `tractor.devx.pformat` render helpers.
'''
from __future__ import annotations
import pytest
from tractor._exceptions import _mk_send_mte
from tractor.devx.pformat import (
pformat_boxed_tb,
pformat_caller_frame,
)
from tractor.msg._codec import _def_tractor_codec
@pytest.mark.parametrize(
'box_tb',
[True, False],
ids=['boxed', 'bare'],
)
def test_pformat_caller_frame_renders(box_tb: bool):
'''
`pformat_caller_frame()` must render, not raise.
XXX the `box_tb=True` branch was passing an `indent=''` kwarg
that `pformat_boxed_tb()` never accepted, so it blew up with
a `TypeError`. Nothing in the test suite covered it, and the
only caller is `_mk_send_mte()` i.e. EVERY send-side
`MsgTypeError` died while formatting itself, masking the real
msg-spec violation behind a bogus `TypeError`.
'''
report: str = pformat_caller_frame(
stack_limit=3,
box_tb=box_tb,
)
assert isinstance(report, str)
assert 'test_pformat_caller_frame_renders' in report
def test_pformat_boxed_tb_rejects_unknown_kwargs():
'''
Pin the signature so a future typo'd kwarg fails loudly at the
call site rather than only when some rare error path runs.
'''
assert pformat_boxed_tb(tb_str='doggy\n')
with pytest.raises(TypeError):
pformat_boxed_tb(
tb_str='doggy\n',
indent='',
)
def test_send_mte_default_message_renders():
'''
The default send-side `MsgTypeError` must remain printable.
Once `pformat_caller_frame()` stopped failing first, this path
exposed two more formatter errors: `MsgCodec.msg_spec_str` passed
a type union where `pformat_msgspec()` requires a codec/decoder,
then `_mk_send_mte()` wrapped its message in a one-element tuple.
Construct the error without an override message to execute that
complete default path. Requiring a `str` message with the bad
value and valid spec, then rendering the exception, proves the
original IPC violation survives every formatter layer.
'''
bad_msg: dict[str, bool] = {'bad': True}
mte = _mk_send_mte(
msg=bad_msg,
codec=_def_tractor_codec,
)
assert isinstance(mte.message, str)
assert f'invalid msg -> {bad_msg}' in mte.message
assert 'Valid IPC msgs are:' in mte.message
report: str = repr(mte)
assert 'MsgTypeError' in report
assert f'invalid msg -> {bad_msg}' in report

View File

@ -191,8 +191,9 @@ def test_shield_pause(
] ]
if not no_capfd: if not no_capfd:
expect_on_teardown += [ expect_on_teardown += [
'Cancel-ack TIMED OUT for sub-actor', # 'Shutting down actor runtime',
'-> escalating to `proc.kill()` (hard-reap)', '#T-800 deployed to collect zombie B0',
"'--uid', \"('hanger',",
] ]
assert_before( assert_before(
child, child,

View File

@ -1,18 +1,21 @@
''' '''
Discovery-suite fixtures, including the `daemon` remote-registrar Discovery-suite fixtures, including the `daemon`
subprocess used by the multi-program discovery tests. remote-registrar subprocess used by the multi-program
discovery tests.
Lives here (vs. the parent `tests/conftest.py`) Lives here (vs. the parent `tests/conftest.py`)
because `daemon` is a discovery-protocol primitive: it boots a child because `daemon` is a discovery-protocol primitive
that enters `open_root_actor()` and waits as a registrar peer for boots a separate `tractor.run_daemon()` process whose
sole purpose is to serve as a registrar peer for
discovery-roundtrip tests. Pytest fixtures inherit discovery-roundtrip tests. Pytest fixtures inherit
DOWNWARD through conftest hierarchy, so anything DOWNWARD through conftest hierarchy, so anything
under `tests/discovery/` automatically picks this up. under `tests/discovery/` automatically picks this up.
''' '''
from __future__ import annotations from __future__ import annotations
from pathlib import Path import os
import platform import platform
import socket
import subprocess import subprocess
import sys import sys
import time import time
@ -28,27 +31,33 @@ from ..conftest import (
def _wait_for_daemon_ready( def _wait_for_daemon_ready(
ready_path: Path, reg_addr: tuple,
tpt_proto: str,
*, *,
deadline: float = 10.0, deadline: float = 10.0,
poll_interval: float = 0.05, poll_interval: float = 0.05,
proc: subprocess.Popen|None = None, proc: subprocess.Popen|None = None,
) -> None: ) -> None:
''' '''
Poll until the daemon reports completed actor startup. Active-poll the daemon's bind address until it
accepts a connection (proving it has called
`bind() + listen()` and is ready to handle IPC).
Replaces the historical blind `time.sleep()` in the Replaces the historical blind `time.sleep()` in the
`daemon` fixture which was racy under load see `daemon` fixture which was racy under load see
`ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`. `ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`.
The child writes `ready_path` only after entering Uses stdlib `socket` directly (no trio runtime
`open_root_actor()`, which guarantees all transport listeners are bootstrap cost) sufficient because
serving without requiring a raw connection probe. `tractor.run_daemon()` doesn't return from
bootstrap until the runtime is fully ready to
accept IPC.
Raises `TimeoutError` on `deadline` exceeded. If Raises `TimeoutError` on `deadline` exceeded. If
`proc` is given, ALSO raises early if the daemon `proc` is given, ALSO raises early if the daemon
process exits before the deadline (catches a daemon startup crash process exits non-zero before the deadline (catches
that the blind sleep used to silently mask). daemon-startup-crash that the blind sleep used to
silently mask).
''' '''
end: float = time.monotonic() + deadline end: float = time.monotonic() + deadline
@ -61,25 +70,43 @@ def _wait_for_daemon_ready(
if proc is not None and proc.poll() is not None: if proc is not None and proc.poll() is not None:
raise RuntimeError( raise RuntimeError(
f'Daemon proc exited (rc={proc.returncode}) ' f'Daemon proc exited (rc={proc.returncode}) '
f'before reporting ready at {ready_path!r}' f'before becoming ready to accept on '
f'{reg_addr!r}'
) )
try: try:
if ready_path.is_file(): if tpt_proto == 'tcp':
if proc is not None and proc.poll() is not None: # `socket.create_connection` does the
raise RuntimeError( # `socket() + connect()` dance with a
f'Daemon proc exited (rc={proc.returncode}) ' # builtin timeout — perfect primitive
f'after reporting ready at {ready_path!r}' # for a one-shot probe.
) with socket.create_connection(
reg_addr,
timeout=poll_interval,
):
return return
else:
# UDS — `reg_addr` is a `(filedir, sockname)`
# tuple per `tractor.ipc._uds.UDSAddress.unwrap`.
sockpath: str = os.path.join(*reg_addr)
sock = socket.socket(socket.AF_UNIX)
try:
sock.settimeout(poll_interval)
sock.connect(sockpath)
return
finally:
sock.close()
except ( except (
ConnectionRefusedError,
FileNotFoundError, FileNotFoundError,
OSError, OSError,
socket.timeout,
) as exc: ) as exc:
last_exc = exc last_exc = exc
time.sleep(poll_interval) time.sleep(poll_interval)
raise TimeoutError( raise TimeoutError(
f'Daemon never reported ready at {ready_path!r} within ' f'Daemon never accepted on {reg_addr!r} within '
f'{deadline}s (last sentinel-state exc: {last_exc!r})' f'{deadline}s (last connect-attempt exc: '
f'{last_exc!r})'
) )
@ -109,27 +136,18 @@ def daemon(
) )
loglevel: str = 'info' loglevel: str = 'info'
ready_path: Path = (
Path(str(testdir.tmpdir))
/ 'daemon-ready'
)
ready_path.unlink(missing_ok=True)
code: str = ( code: str = (
f'from pathlib import Path\n' "import tractor; "
f'import tractor\n' "tractor.run_daemon([], "
f'import trio\n' "registry_addrs={reg_addrs}, "
f'\n' "enable_transports={enable_tpts}, "
f'async def main():\n' "debug_mode={debug_mode}, "
f' async with tractor.open_root_actor(\n' "loglevel={ll})"
f' registry_addrs={[reg_addr]!r},\n' ).format(
f' enable_transports={[tpt_proto]!r},\n' reg_addrs=str([reg_addr]),
f' debug_mode={debug_mode!r},\n' enable_tpts=str([tpt_proto]),
f' loglevel={loglevel!r},\n' ll="'{}'".format(loglevel) if loglevel else None,
f' ):\n' debug_mode=debug_mode,
f' Path({str(ready_path)!r}).touch()\n'
f' await trio.sleep_forever()\n'
f'\n'
f'trio.run(main)\n'
) )
cmd: list[str] = [ cmd: list[str] = [
sys.executable, sys.executable,
@ -145,9 +163,9 @@ def daemon(
**kwargs, **kwargs,
) )
# Poll the child's ready sentinel, published after actor startup, # Active-poll the daemon's bind address until it's
# instead of connecting to its transport socket. This replaces # ready to accept connections — replaces the legacy
# the legacy blind `time.sleep(2.2)` which was racy under load # blind `time.sleep(2.2)` which was racy under load
# (see # (see
# `ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`). # `ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`).
# #
@ -158,22 +176,20 @@ def daemon(
15.0 if (_non_linux and ci_env) 15.0 if (_non_linux and ci_env)
else 10.0 else 10.0
) )
try:
_wait_for_daemon_ready( _wait_for_daemon_ready(
ready_path=ready_path, reg_addr=reg_addr,
tpt_proto=tpt_proto,
deadline=deadline, deadline=deadline,
proc=proc, proc=proc,
) )
assert not proc.returncode assert not proc.returncode
yield proc yield proc
finally:
if proc.poll() is None:
sig_prog(proc, _INT_SIGNAL) sig_prog(proc, _INT_SIGNAL)
# NOTE: these blocking reads can hang when descendants retain # XXX! yeah.. just be reaaal careful with this bc
# inherited pipe descriptors. Keep teardown signaling above # sometimes it can lock up on the `_io.BufferedReader`
# them and avoid adding subprocesses outside the actor tree. # and hang..
# #
# NB, drain happens at TEARDOWN (post-yield), so the # NB, drain happens at TEARDOWN (post-yield), so the
# test body has its chance to read `proc.stderr` # test body has its chance to read `proc.stderr`
@ -203,4 +219,5 @@ def daemon(
) )
if rc < 0: if rc < 0:
raise RuntimeError(msg) raise RuntimeError(msg)
test_log.error(msg) test_log.error(msg)

View File

@ -1,76 +0,0 @@
'''
Discovery daemon fixture regressions.
This module imports private helpers from the sibling
`tests.discovery.conftest` plugin to exercise that fixture machinery
directly, rather than testing a production `tractor` API.
'''
from unittest.mock import (
call,
Mock,
)
from .conftest import _wait_for_daemon_ready
def test_daemon_ready_check_does_not_connect(
monkeypatch,
tmp_path,
):
'''
Observe completed daemon startup without a raw connection.
The old UDS readiness helper connected and immediately closed. That
entered Tractor's actor-handshake handler with no `Aid` payload and
destabilized the remote registrar on macOS before discovery tests
started. This test creates the child sentinel, forbids all socket
construction and connection helpers, then proves readiness returns
without touching the transport layer.
'''
ready_path = tmp_path / 'daemon-ready'
ready_path.touch()
socket_ctor = Mock(side_effect=AssertionError('socket opened'))
connect = Mock(side_effect=AssertionError('socket connected'))
monkeypatch.setattr('socket.socket', socket_ctor)
monkeypatch.setattr('socket.create_connection', connect)
_wait_for_daemon_ready(
ready_path=ready_path,
deadline=.1,
poll_interval=.01,
)
socket_ctor.assert_not_called()
connect.assert_not_called()
def test_daemon_ready_check_backs_off(monkeypatch):
'''
Back off while waiting for the child startup sentinel.
The sentinel may appear after several parent polling intervals.
A deterministic false/false/true path sequence proves the helper
sleeps between unsuccessful observations instead of hot-spinning
and starving a booting daemon on constrained CI workers.
'''
ready_path = Mock()
ready_path.is_file.side_effect = [False, False, True]
sleep = Mock()
monotonic = Mock(side_effect=[0, 0, 0, 0])
monkeypatch.setattr('time.sleep', sleep)
monkeypatch.setattr('time.monotonic', monotonic)
_wait_for_daemon_ready(
ready_path=ready_path,
deadline=.2,
poll_interval=.01,
)
assert ready_path.is_file.call_count == 3
assert sleep.call_args_list == [
call(.01),
call(.01),
]

View File

@ -207,7 +207,7 @@ def test_dup_name_cancel_cascade_escalates_to_hard_kill(
Post-fix, `Portal.cancel_actor()` raises `ActorTooSlowError` on Post-fix, `Portal.cancel_actor()` raises `ActorTooSlowError` on
the bounded-wait timeout, and `ActorNursery.cancel()`'s the bounded-wait timeout, and `ActorNursery.cancel()`'s
per-child wrapper escalates directly to `proc.kill()` (hard-reap). per-child wrapper escalates to `proc.terminate()` (hard-kill).
The full nursery teardown therefore stays bounded even under The full nursery teardown therefore stays bounded even under
pathological timing. pathological timing.
@ -266,7 +266,7 @@ def test_dup_name_cancel_cascade_escalates_to_hard_kill(
# post-teardown sanity: every child proc must be reaped. # post-teardown sanity: every child proc must be reaped.
# If escalation worked, even timed-out cancel-RPCs would # If escalation worked, even timed-out cancel-RPCs would
# have triggered `proc.kill()` and the procs are dead. # have triggered `proc.terminate()` and the procs are dead.
for p in portals: for p in portals:
# `Portal.channel.connected()` -> False once the # `Portal.channel.connected()` -> False once the
# underlying chan disconnected (clean exit OR # underlying chan disconnected (clean exit OR

View File

@ -2,211 +2,23 @@
`open_root_actor(tpt_bind_addrs=...)` test suite. `open_root_actor(tpt_bind_addrs=...)` test suite.
Verify all three runtime code paths for explicit IPC-server Verify all three runtime code paths for explicit IPC-server
bind-address selection in `_root.py` and registry probing in bind-address selection in `_root.py`:
`discovery._api`:
1. Non-registrar, no explicit bind -> random addrs from registry proto 1. Non-registrar, no explicit bind -> random addrs from registry proto
2. Registrar, no explicit bind -> binds to registry_addrs 2. Registrar, no explicit bind -> binds to registry_addrs
3. Explicit bind given -> wraps via `wrap_address()` and uses them 3. Explicit bind given -> wraps via `wrap_address()` and uses them
''' '''
from contextlib import asynccontextmanager as acm
from unittest.mock import (
AsyncMock,
call,
Mock,
)
import pytest import pytest
import trio import trio
import tractor import tractor
from tractor.discovery import _api
from tractor.discovery._addr import ( from tractor.discovery._addr import (
wrap_address, wrap_address,
) )
from tractor.discovery._multiaddr import mk_maddr from tractor.discovery._multiaddr import mk_maddr
from tractor.ipc import _connect_chan
from tractor._testing.addr import get_rando_addr from tractor._testing.addr import get_rando_addr
def test_registry_probe_retries_transient_handshake(
monkeypatch: pytest.MonkeyPatch,
):
'''
Retry a connected registrar after transient handshake timeout.
Loaded macOS runners can accept the transport while delaying the
actor handshake beyond one second. Treating that first timeout as
final makes a healthy remote daemon look occupied and cascades into
discovery failures. This deterministic fake fails once, succeeds
on the second complete handshake, and proves one bounded backoff.
'''
async def stall_handshake(**kwargs):
await trio.sleep_forever()
first_handshake = AsyncMock(side_effect=stall_handshake)
second_handshake = AsyncMock(
return_value=tractor.msg.Aid(
name='registrar',
uuid='registrar-uuid',
pid=1234,
is_registrar=True,
),
)
chans = [
Mock(_do_handshake=first_handshake),
Mock(_do_handshake=second_handshake),
]
closed: list[object] = []
@acm
async def connect_chan(addr, close_timeout):
assert close_timeout == .2
chan = chans[len(closed)]
try:
yield chan
finally:
closed.append(chan)
sleep = AsyncMock()
monkeypatch.setattr(_api, '_connect_chan', connect_chan)
monkeypatch.setattr(_api.trio, 'sleep', sleep)
async def main():
status = await _api._probe_registry(
addr=wrap_address(('127.0.0.1', 1616)),
timeout=.3,
attempt_timeout=.1,
max_attempts=3,
retry_delay=.01,
)
assert status == 'registrar'
trio.run(main)
first_handshake.assert_awaited_once()
second_handshake.assert_awaited_once()
assert first_handshake.await_args.kwargs['timeout'] == .1
assert second_handshake.await_args.kwargs['timeout'] == .1
assert closed == chans
sleep.assert_has_awaits([call(.01)])
def test_probe_channel_close_is_bounded(
monkeypatch: pytest.MonkeyPatch,
):
'''
Bound shielded channel cleanup after a registry probe.
`_connect_chan()` shields `.aclose()` so cancellation cannot leak
ordinary channels. A stalled close previously let registry probing
exceed every connect and handshake deadline. This fake close never
completes; the explicit cleanup allowance must still return control
to the caller without cancelling its surrounding task.
'''
chan = Mock()
chan.aclose = AsyncMock(side_effect=trio.sleep_forever)
monkeypatch.setattr(
tractor.Channel,
'from_addr',
AsyncMock(return_value=chan),
)
async def main():
with trio.fail_after(.5):
async with _connect_chan(
('127.0.0.1', 1616),
close_timeout=.01,
):
pass
trio.run(main)
chan.aclose.assert_awaited_once()
def test_transport_only_listener_is_not_registrar():
'''
Require a Tractor handshake before accepting a registry address.
The old election probe marked an address live after transport
connect alone. A non-Tractor listener, or a registrar still
failing its initial handshake, was therefore selected as the
remote registry. This test accepts the probe and closes it without
replying, then proves `open_root_actor()` rejects that occupied
endpoint instead of selecting it or binding over it.
'''
async def transport_only_handler(
stream: trio.SocketStream,
) -> None:
await stream.aclose()
async def main():
listeners = await trio.open_tcp_listeners(0)
listener = listeners[0]
sockname = listener.socket.getsockname()
reg_addr: tuple[str, int] = (
sockname[0],
sockname[1],
)
async with trio.open_nursery() as tn:
tn.start_soon(
trio.serve_listeners,
transport_only_handler,
listeners,
)
with pytest.raises(
RuntimeError,
match='occupied but did not answer',
):
async with tractor.open_root_actor(
registry_addrs=[reg_addr],
enable_transports=['tcp'],
):
pytest.fail('foreign listener selected as registrar')
tn.cancel_scope.cancel()
trio.run(main)
def test_registry_probe_preserves_no_peers_state(
reg_addr: tuple,
tpt_proto: str,
):
'''
Keep an idle registrar peer-free after an election probe.
Probe handshakes exchange registrar capability but must not enter
`IPCServer._peers`. Resetting `_no_more_peers` before identifying a
probe left an idle registrar reporting phantom peers and delayed
shutdown. This test probes the live local registrar and proves its
peer map and no-peers event remain unchanged afterward.
'''
async def main():
async with tractor.open_root_actor(
registry_addrs=[reg_addr],
enable_transports=[tpt_proto],
):
actor = tractor.current_actor()
server = actor.ipc_server
probe_status = await _api._probe_registry(
addr=wrap_address(reg_addr),
)
assert probe_status == 'registrar'
await trio.sleep(0)
assert not server._peers
assert server._no_more_peers.is_set()
trio.run(main)
# ------------------------------------------------------------------ # ------------------------------------------------------------------
# helpers # helpers
# ------------------------------------------------------------------ # ------------------------------------------------------------------

View File

@ -3,285 +3,14 @@ Unit-ish tests for specific IPC transport protocol backends.
''' '''
from __future__ import annotations from __future__ import annotations
import os
from pathlib import Path from pathlib import Path
import socket
import stat
import struct
import sys
import tempfile
from types import SimpleNamespace
from unittest.mock import Mock
import pytest import pytest
import trio import trio
from trio.testing import (
MockClock,
wait_all_tasks_blocked,
)
import tractor import tractor
from tractor import Actor from tractor import Actor
from tractor.discovery import _addr
from tractor.ipc._transport import MsgpackTransport
from tractor.runtime import _state from tractor.runtime import _state
from tractor.discovery import _addr
def test_cancelled_transport_send_completes_frame():
'''
Finish an in-flight frame before delivering sender cancellation.
A cancelled `send_all()` may leave an arbitrary frame prefix on the
wire. Closing the actor-wide stream avoids decoder corruption but
also destroys unrelated contexts using that channel. On its first
call, the fake stream publishes two header bytes and blocks until
the parent test releases it. This lets the parent cancel the sender
while frame publication is suspended. The sender must remain inside
`send_first()` until the complete frame is written, then observe
pending cancellation; a second sender proves sibling contexts can
safely reuse the still frame-aligned stream.
'''
class PartialSendStream:
def __init__(self) -> None:
self.send_all_entered = trio.Event()
self.send_all_release = trio.Event()
self.closed = False
self.wire = bytearray()
async def send_all(
self,
data: bytes,
) -> None:
assert data
if not self.wire:
self.wire.extend(data[:2])
self.send_all_entered.set()
await self.send_all_release.wait()
self.wire.extend(data[2:])
else:
self.wire.extend(data)
async def aclose(self) -> None:
self.closed = True
def count_frames(wire: bytearray) -> int:
offset: int = 0
count: int = 0
while offset < len(wire):
header_end: int = offset + 4
assert header_end <= len(wire)
size, = struct.unpack('<I', wire[offset:header_end])
offset = header_end + size
assert offset <= len(wire)
count += 1
assert offset == len(wire)
return count
async def main() -> None:
stream = PartialSendStream()
transport = object.__new__(MsgpackTransport)
transport.stream = stream
transport._send_lock = trio.StrictFIFOLock()
sender_done = trio.Event()
sender_scopes: list[trio.CancelScope] = []
cancelled_caught: bool = False
first_msg = tractor.msg.Start(
ns=__name__,
func='add_one',
kwargs={'n': 1},
uid=('root', 'test'),
cid='partial-send',
)
second_msg = tractor.msg.Start(
ns=__name__,
func='add_one',
kwargs={'n': 2},
uid=('root', 'test'),
cid='second-send',
)
async def send_first() -> None:
nonlocal cancelled_caught
with trio.CancelScope() as cs:
sender_scopes.append(cs)
await transport.send(first_msg)
cancelled_caught = cs.cancelled_caught
sender_done.set()
async with trio.open_nursery() as tn:
tn.start_soon(
send_first,
)
await stream.send_all_entered.wait()
sender_scopes[0].cancel()
await wait_all_tasks_blocked()
assert not stream.closed
# Cancellation is pending, but complete-frame shielding
# keeps `send_first()` suspended in `.send_all()`.
assert not sender_done.is_set()
# Let the underlying frame write finish after the parent
# has requested sender cancellation.
stream.send_all_release.set()
await sender_done.wait()
assert cancelled_caught
assert not stream.closed
# The initial two-byte prefix was completed into one valid
# frame before cancellation reached `send_first()`.
assert count_frames(stream.wire) == 1
await transport.send(second_msg)
# A sibling sender can append and decode another frame only
# because the first cancellation preserved stream alignment.
assert count_frames(stream.wire) == 2
tn.cancel_scope.cancel()
trio.run(main)
def test_transport_send_deadline_closes_partial_frame():
'''
Destroy a stalled partial frame before another sender can append.
Ordinary cancellation cannot interrupt complete-frame publication.
Bounded actor/context cancellation instead passes its absolute
deadline into this operation. The fake stream writes a partial
header and stalls; when the send's own deadline fires, the transport
must close the stream before releasing its shared send lock. This
prevents the next sender from appending bytes which a decoder would
treat as the remainder of the corrupt first frame.
'''
class StalledStream:
def __init__(self) -> None:
self.closed = False
self.wire = bytearray()
async def send_all(
self,
data: bytes,
) -> None:
self.wire.extend(data[:2])
await trio.sleep_forever()
async def aclose(self) -> None:
self.closed = True
async def main() -> None:
stream = StalledStream()
transport = object.__new__(MsgpackTransport)
transport.stream = stream
transport._send_lock = trio.StrictFIFOLock()
msg = tractor.msg.Start(
ns=__name__,
func='add_one',
kwargs={'n': 1},
uid=('root', 'test'),
cid='deadline-send',
)
with pytest.raises(
tractor.TransportClosed,
match='frame publication exceeded',
):
await transport.send(
msg,
send_deadline=1,
)
assert stream.closed # partial-frame timeout destroys stream
assert len(stream.wire) == 2 # only a header fragment was sent
assert not transport._send_lock.locked() # cleanup released lock
trio.run(
main,
clock=MockClock(autojump_threshold=0),
)
def test_cancelled_transport_send_preserves_cancellation():
'''
Prefer sender cancellation when teardown closes the stream.
`MsgpackTransport.send()` shields frame publication at
`tractor.ipc._transport:MsgpackTransport.send`. Before this fix, an
outer `move_on_after()`/cancel scope could cancel `Channel.send()`
while actor teardown closed the shared stream. The resulting
`ClosedResourceError` escaped from the transport handler instead of
its `checkpoint_if_cancelled()` redelivering pending cancellation.
The fake stream blocks inside the shield until the test cancels the
sender. Parent-controlled release then simulates actor teardown
closing the socket and raises the `ClosedResourceError` observed on
macOS UDS. `CancelScope.cancelled_caught` proves the handler's
checkpoint preserved cancellation as the primary outcome instead of
leaking that secondary close error.
'''
class ClosingStream:
def __init__(self) -> None:
self.send_all_entered = trio.Event()
self.send_all_release = trio.Event()
async def send_all(
self,
data: bytes,
) -> None:
assert data
self.send_all_entered.set()
await self.send_all_release.wait()
# Model actor teardown closing the shared transport while
# this sender is still inside the complete-frame shield.
raise trio.ClosedResourceError(
'this socket was already closed'
)
async def main() -> None:
stream = ClosingStream()
transport = object.__new__(MsgpackTransport)
transport.stream = stream
transport._send_lock = trio.StrictFIFOLock()
sender_done = trio.Event()
sender_scopes: list[trio.CancelScope] = []
cancelled_caught: bool = False
msg = tractor.msg.Start(
ns=__name__,
func='add_one',
kwargs={'n': 1},
uid=('root', 'test'),
cid='close-during-cancelled-send',
)
async def send() -> None:
nonlocal cancelled_caught
with trio.CancelScope() as cs:
sender_scopes.append(cs)
await transport.send(msg)
cancelled_caught = cs.cancelled_caught
sender_done.set()
async with trio.open_nursery() as tn:
tn.start_soon(send)
await stream.send_all_entered.wait()
sender_scopes[0].cancel()
await wait_all_tasks_blocked()
assert not sender_done.is_set()
stream.send_all_release.set()
await sender_done.wait()
assert cancelled_caught
trio.run(main)
@pytest.fixture @pytest.fixture
@ -302,381 +31,6 @@ def bindspace_dir_str() -> str:
bs_dir.rmdir() bs_dir.rmdir()
def test_macos_rt_dir_fits_uds_path_limit(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Keep the default Darwin UDS bindpath below its 104-byte limit.
`platformdirs` normally places the runtime directory below the
long `~/Library/Caches/TemporaryItems` path. Pytest also assigns
a deeply nested temporary home, so appending a registry socket
name made every macOS UDS listener fail with `AF_UNIX path too
long`. This test simulates Darwin and an intentionally long
platformdirs result, then proves `get_rt_dir()` uses the short
system temporary directory and leaves room for the socket name.
'''
long_rt_dir: Path = tmp_path / ('long' * 30)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(long_rt_dir / appname),
)
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
rt_dir: Path = _state.get_rt_dir()
sockpath: Path = (
Path('/tmp')
/ f'tractor-{os.getuid()}'
/ 'registry@1616.sock'
)
assert rt_dir == tmp_path / f'tractor-{os.getuid()}'
assert len(os.fsencode(sockpath)) < 104
assert stat.S_IMODE(rt_dir.stat().st_mode) == 0o700
def test_macos_rt_dir_rejects_symlink(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject a pre-created symlink at the Darwin runtime path.
Darwin uses the predictable `/tmp/tractor-<uid>` path to stay
below its `AF_UNIX` limit. A hostile local user could otherwise
point that path at a victim-owned directory and make
`get_rt_dir()` chmod or place sockets in the symlink target. The
test replaces `/tmp` with a controlled directory, installs the
malicious link, and proves non-following validation rejects it.
'''
runtime_link: Path = tmp_path / f'tractor-{os.getuid()}'
target_dir: Path = tmp_path / 'target'
target_dir.mkdir(mode=0o755)
runtime_link.symlink_to(target_dir, target_is_directory=True)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
with pytest.raises(PermissionError, match='Unsafe Darwin'):
_state.get_rt_dir()
assert stat.S_IMODE(target_dir.stat().st_mode) == 0o755
def test_reaper_uses_default_uds_bindspace(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Sweep the same platform-specific bindspace used by UDS actors.
The reaper previously consulted only `XDG_RUNTIME_DIR`, missing
Darwin sockets after the runtime moved to `/tmp/tractor-<uid>`.
This test replaces `UDSAddress.def_bindspace` and proves the test
harness resolves that shared transport default directly.
'''
from tractor._testing import _reap
from tractor.ipc._uds import UDSAddress
monkeypatch.setattr(
UDSAddress,
'def_bindspace',
tmp_path,
)
assert _reap.get_uds_dir() == str(tmp_path)
def test_automatic_reaper_preserves_registry_sentinel(
monkeypatch: pytest.MonkeyPatch,
):
'''
Reserve unconditional registry cleanup for the explicit CLI.
The `registry@1616.sock` suffix does not encode its binder PID, so
automatic pytest cleanup cannot distinguish a leak from another
live registrar. This test creates registry and actor sockets,
proves the default sweep selects only the dead actor, then proves
explicit sentinel inclusion retains the CLI's documented behavior.
'''
from tractor._testing import _reap
with tempfile.TemporaryDirectory(
prefix='tractor-reap-',
dir='/tmp',
) as tmpdir:
bindspace: Path = Path(tmpdir)
registry_path: Path = bindspace / 'registry@1616.sock'
actor_path: Path = bindspace / 'worker@1234.sock'
socks: list[socket.socket] = []
for path in (registry_path, actor_path):
sock = socket.socket(socket.AF_UNIX)
sock.bind(str(path))
socks.append(sock)
monkeypatch.setattr(_reap, '_is_alive', lambda pid: False)
try:
assert _reap.find_orphaned_uds(
uds_dir=str(bindspace),
) == [str(actor_path)]
assert set(
_reap.find_orphaned_uds(
uds_dir=str(bindspace),
include_registry_sentinel=True,
)
) == {
str(registry_path),
str(actor_path),
}
finally:
for sock in socks:
sock.close()
def test_rt_dir_rejects_non_directory(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Preserve the non-Darwin runtime-directory type contract.
Replacing `Path.is_dir()` with unguarded `lstat()` briefly made
existing files look like valid runtime directories on Linux.
This test points `platformdirs` at a regular file and proves
`get_rt_dir()` rejects it during initialization.
'''
rt_file: Path = tmp_path / 'runtime-file'
rt_file.touch()
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_file),
)
with pytest.raises(
PermissionError,
match='Unsafe POSIX',
):
_state.get_rt_dir()
new_rt_dir: Path = tmp_path / 'new-runtime-dir'
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(new_rt_dir),
)
assert _state.get_rt_dir() == new_rt_dir
assert stat.S_IMODE(new_rt_dir.stat().st_mode) == 0o700
def test_linux_rt_dir_secures_existing_path(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Enforce owner-only access on an existing Linux runtime directory.
Linux previously accepted any existing directory returned by
`platformdirs`, without checking ownership or correcting a
traversable mode. This test creates an owner-controlled `0o755`
directory and proves `get_rt_dir()` normalizes the managed
bindspace to `0o700` before returning it.
'''
rt_dir: Path = tmp_path / 'tractor'
rt_dir.mkdir(mode=0o755)
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_dir),
)
assert _state.get_rt_dir() == rt_dir
assert stat.S_IMODE(rt_dir.stat().st_mode) == 0o700
def test_linux_rt_dir_rejects_foreign_owner(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject an existing Linux runtime directory owned by another UID.
A pre-created bindspace must never be made private with `chmod`
until ownership is verified. This test makes the current process
appear to have a different UID and proves `get_rt_dir()` rejects
the directory without changing its original mode.
'''
rt_dir: Path = tmp_path / 'tractor'
rt_dir.mkdir(mode=0o755)
original_mode: int = stat.S_IMODE(rt_dir.stat().st_mode)
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_dir),
)
monkeypatch.setattr(
os,
'getuid',
lambda: rt_dir.stat().st_uid + 1,
)
with pytest.raises(
PermissionError,
match='Unsafe POSIX',
):
_state.get_rt_dir()
assert stat.S_IMODE(rt_dir.stat().st_mode) == original_mode
def test_macos_rt_dir_rejects_intermediate_symlink(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject symlinks in nested Darwin runtime subdirectories.
The earlier final-component check allowed `link/child` to follow
an intermediate symlink and create `child` outside the secured
runtime root. This test installs that link and proves traversal
stops before anything is created in its target.
'''
rt_root: Path = tmp_path / f'tractor-{os.getuid()}'
target_dir: Path = tmp_path / 'target'
rt_root.mkdir(mode=0o700)
target_dir.mkdir()
(rt_root / 'link').symlink_to(
target_dir,
target_is_directory=True,
)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
with pytest.raises(PermissionError, match='Unsafe Darwin'):
_state.get_rt_dir(subdir='link/child')
assert not (target_dir / 'child').exists()
@pytest.mark.parametrize(
('platform_name', 'path_limit'),
[
('darwin', 104),
('linux', 108),
],
)
def test_uds_sockname_compaction(
monkeypatch: pytest.MonkeyPatch,
platform_name: str,
path_limit: int,
):
'''
Keep generated actor sockets safe and below Darwin's byte limit.
Actor names are unrestricted identity strings. A long, multibyte,
or path-like name previously produced overlong or escaping socket
paths. These cases prove `UDSAddress.get_sockname()` preserves a
short legacy name, deterministically compacts unsafe names, keeps
the reaper's `@pid.sock` suffix, and stays within Darwin's byte
limit.
'''
from tractor.ipc._uds import UDSAddress
bindspace: Path = Path('/tmp/tractor-501')
pid: int = 12345
from tractor.ipc import _uds
monkeypatch.setattr(sys, 'platform', platform_name)
monkeypatch.setattr(_uds, '_SUN_PATH_LIMIT', path_limit)
short: Path = UDSAddress.get_sockname(
name='worker',
pid=pid,
bindspace=bindspace,
)
long_name: str = 'actor-' + ('\u00e9' * 100)
compact: Path = UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=bindspace,
)
unsafe: Path = UDSAddress.get_sockname(
name='../worker',
pid=pid,
bindspace=bindspace,
)
assert short == Path(f'worker@{pid}.sock')
assert compact == UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=bindspace,
)
assert compact.name.endswith(f'@{pid}.sock')
assert unsafe.parent == Path('.')
assert '..' not in unsafe.name
assert len(os.fsencode(bindspace / compact)) < path_limit
with pytest.raises(ValueError) as exc_info:
UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=Path('/tmp') / ('x' * 90),
)
errmsg: str = str(exc_info.value)
assert 'leaves no room' in errmsg
assert 'name was unsafe: False' in errmsg
assert 'name was over budget: True' in errmsg
assert f'AF_UNIX path limit: {path_limit}' in errmsg
def test_uds_reaper_ignores_unreconstructable_path(
monkeypatch: pytest.MonkeyPatch,
):
'''
Keep post-kill UDS cleanup best-effort on path overflow.
`unlink_uds_bind_addrs()` reconstructs a self-assigned socket from
the dead actor's name and PID. An over-budget bindspace makes that
naming helper raise before `os.unlink()`; propagating the error
would replace the original supervision outcome after the child was
already killed. This test forces overflow and proves cleanup skips
reconstruction without attempting an unlink or raising.
'''
from tractor.ipc import _uds
from tractor.spawn import _reap
long_bindspace: Path = Path('/tmp') / ('x' * 120)
proc = SimpleNamespace(pid=12345)
subactor = SimpleNamespace(
aid=SimpleNamespace(name='worker'),
)
unlink = Mock()
monkeypatch.setattr(
_uds.UDSAddress,
'def_bindspace',
long_bindspace,
)
monkeypatch.setattr(_reap.os, 'unlink', unlink)
_reap.unlink_uds_bind_addrs(
proc=proc,
subactor=subactor,
)
unlink.assert_not_called()
def test_uds_bindspace_created_implicitly( def test_uds_bindspace_created_implicitly(
debug_mode: bool, debug_mode: bool,
bindspace_dir_str: str, bindspace_dir_str: str,

View File

@ -3,13 +3,7 @@ High-level `.ipc._server` unit tests.
''' '''
from __future__ import annotations from __future__ import annotations
import errno
from unittest.mock import (
AsyncMock,
Mock,
)
import msgspec
import pytest import pytest
import trio import trio
from tractor import ( from tractor import (
@ -20,11 +14,6 @@ from tractor import (
from tractor._testing.addr import ( from tractor._testing.addr import (
get_rando_addr, get_rando_addr,
) )
from tractor._exceptions import TransportClosed
from tractor.ipc._chan import Channel
from tractor.ipc import _server
from tractor.ipc._transport import MsgpackTransport
from tractor.msg.types import Aid
# TODO, use/check-roundtripping with some of these wrapper types? # TODO, use/check-roundtripping with some of these wrapper types?
# #
# from .._addr import Address # from .._addr import Address
@ -34,165 +23,6 @@ from tractor.msg.types import Aid
# from ._tcp import TCPAddress # from ._tcp import TCPAddress
def test_send_normalizes_only_grouped_peer_resets():
'''
Normalize only all-peer-close grouped transport failures.
A UDS peer may disconnect before completing the actor handshake.
Darwin can report the server's first handshake write as
`ECONNRESET`, wrapped by `trio.BrokenResourceError` and potentially
nested in an `ExceptionGroup`. This fake stream first groups reset
and broken-pipe branches, proving `.send()` normalizes a complete
peer-close tree to `TransportClosed`. It then groups a reset with
an unrelated `ValueError`, proving the mixed failure remains a
`trio.BrokenResourceError` instead of hiding the application error.
'''
def broken_resource(err_no: int) -> trio.BrokenResourceError:
try:
raise OSError(
err_no,
'Peer closed',
)
except OSError as peer_err:
try:
raise trio.BrokenResourceError from peer_err
except trio.BrokenResourceError as broken_err:
return broken_err
class GroupedFailureStream:
def __init__(self, exceptions: list[Exception]) -> None:
self.exceptions = exceptions
async def send_all(self, data: bytes) -> None:
grouped_err = ExceptionGroup(
'concurrent send failures',
self.exceptions,
)
raise trio.BrokenResourceError from grouped_err
async def main():
transport = object.__new__(MsgpackTransport)
transport.stream = GroupedFailureStream([
broken_resource(errno.ECONNRESET),
broken_resource(errno.EPIPE),
])
transport._send_lock = trio.StrictFIFOLock()
transport._laddr = 'local'
transport._raddr = 'remote'
transport._task = trio.lowlevel.current_task()
with pytest.raises(TransportClosed) as exc_info:
await transport.send(
{'probe': True},
strict_types=False,
)
grouped_err = exc_info.value.src_exc.__cause__
assert isinstance(grouped_err, ExceptionGroup)
assert len(grouped_err.exceptions) == 2
transport.stream = GroupedFailureStream([
ValueError('unrelated failure'),
broken_resource(errno.ECONNRESET),
])
with pytest.raises(trio.BrokenResourceError) as exc_info:
await transport.send(
{'probe': True},
strict_types=False,
)
grouped_err = exc_info.value.__cause__
assert isinstance(grouped_err, ExceptionGroup)
assert isinstance(grouped_err.exceptions[0], ValueError)
trio.run(main)
def test_handshake_normalizes_decode_error():
'''
Keep malformed pre-handshake frames out of the service nursery.
A non-msgpack peer can trigger `msgspec.DecodeError` before a
remote `Aid` exists. Letting that decoder error escape the inbound
handler cancels the actor's shared IPC nursery. This fake channel
proves `_do_handshake()` presents only `TransportClosed` upward.
'''
chan = object.__new__(Channel)
chan.send = AsyncMock()
chan.recv = AsyncMock(
side_effect=msgspec.DecodeError('malformed handshake'),
)
async def main():
with pytest.raises(TransportClosed) as exc_info:
await chan._do_handshake(
aid=Aid(
name='local',
uuid='local-uuid',
pid=1234,
),
timeout=.1,
)
assert isinstance(
exc_info.value.src_exc,
msgspec.DecodeError,
)
trio.run(main)
def test_server_uses_independent_handshake_timeout(
monkeypatch: pytest.MonkeyPatch,
):
'''
Give ordinary actor handshakes a distinct, generous deadline.
Registry probes use short retries, but ordinary portal and child
connections do not retry. Applying the probe's one-second timeout
in the server can terminate a valid delayed child and leave its
parent blocked in `IPCServer.wait_for_peer()`. This handler fake
proves the server uses its separate pre-registration budget.
'''
handshake = AsyncMock(
side_effect=TransportClosed(message='stop after assertion'),
)
chan = Mock(_do_handshake=handshake)
actor = Mock(
aid=Aid(
name='local',
uuid='local-uuid',
pid=1234,
),
)
monkeypatch.setattr(
Channel,
'from_stream',
Mock(return_value=chan),
)
monkeypatch.setattr(
_server._state,
'current_actor',
Mock(return_value=actor),
)
async def main():
await _server.handle_stream_from_peer(
stream=Mock(),
server=Mock(),
)
trio.run(main)
handshake.assert_awaited_once_with(
aid=actor.aid,
timeout=_server._PRE_REG_HANDSHAKE_TIMEOUT,
)
assert _server._PRE_REG_HANDSHAKE_TIMEOUT == 10
@pytest.mark.parametrize( @pytest.mark.parametrize(
'_tpt_proto', '_tpt_proto',
['uds', 'tcp'] ['uds', 'tcp']

Some files were not shown because too many files have changed in this diff Show More