Compare commits

..

75 Commits

Author SHA1 Message Date
Gud Boi 38d6688d2a Doc the #477 migration outcome + one-shot-acm sketch
Fold the endeavour's resolution into the plan doc + log the
session per prompt-io policy,

- `ria_nursery_removal_plan.md`: RESOLVED section — migrate
  everything, remove the API; the migration-pattern table
  (blocking / fire-and-forget / fan-out / collect-don't-cancel
  / mutual-rendezvous), the semantic deltas (cancel-on-first +
  `collapse_eg()` chain collapse vs the old teardown-reap BEG),
  the excision inventory and the structural dissolution of the
  reap-hang class.
- adds the `to_actor.open_one_shot()` follow-up sketch: an
  `@acm` + private task-nursery over the existing blocking
  `run()` — done-`trio.Event` as a result memo (NOT a
  cancel-relay), no `Portal` in the iface, errors always
  propagate at scope exit; zero `_supervise` coupling.
- prompt-io entry `20260706T172818Z_ad42871e` (+ raw diff-ref
  companion) covering commits `d01a2123..ad42871e`.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 38aa92bc61 Name every `ActorNursery` binding `an` in tests/examples
Convention sweep (user req): all `tractor.open_nursery()`
bindings in test + example code use `an: ActorNursery` (`n`,
`nursery` + several tractor-nurseries confusingly named `tn`
are renamed); `trio.open_nursery()` bindings stay `tn` (incl.
`concurrent_actors_primes.py`'s inner trio nursery, renamed
`n` -> `tn` to match).

Purely mechanical, function-scoped renames — prose "nursery"/
"an" in docstrings/comments untouched; func-arg kwargs like
`portal.run(func, n=value)` untouched.

Gate: renamed test modules green on `trio`; full debugger suite
(28p/6s) + example-runner (21p) green.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 7e9e39c5f8 Fix mutual-rendezvous premature-reap race (#477)
The `test_trynamic_trio` + `a_trynamic_first_scene.py` migration
to paired `to_actor.run()` one-shots carries a race the legacy
`run_in_actor()` shape never had: donny + gretchen each
`wait_for_actor()` (then DIAL) the *other*, but a one-shot is
reaped the instant its own hello returns — so the slower peer
can resolve the winner's registry entry and connect to an
already-dead sockaddr -> `ConnectionRefusedError` boxed as a
`RemoteActorError` (or a reg-wait `TooSlowError`), flaking
~1-in-3 standalone runs.

Mutual-rendezvous peers must OUTLIVE both dialogs, so pin the
lifetimes explicitly: `start_actor()` both as daemons, run both
hellos concurrently via bg `Portal.run()` tasks, then reap with
`an.cancel()` only after the task-nursery joins. (The legacy
teardown-reap provided this pinning implicitly — one of the
few places its semantics were ever actually relied upon.)

Gate: `-k trynamic` standalone x8 green (was flaking); full
`test_registrar` module + the example-runner green.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 0a1431c933 Remove `run_in_actor()` + the ria reap cluster
The final excision of #477: with zero in-repo callers left (all
tests/examples/docs migrated to `to_actor.run()` et al) the
entire legacy one-shot machinery drops out,

- `runtime/_supervise.py`: `ActorNursery.run_in_actor()`, the
  `._cancel_after_result_on_exit` portal-set and the
  `_reap_ria_portals()` teardown-reaper (both its happy-path
  block-exit call AND the error-path snapshot + 0.5s-bounded
  collection) are deleted — one-shot result-waiting now lives
  entirely in the caller's task via `to_actor.run()`, whose
  enclosing cancel-scope bounds the wait by construction (the
  correct-scoping fix for the unbounded-reap hang class; the
  `d1fb4a1a` guard test now passes structurally).
- `runtime/_portal.py`: `Portal._submit_for_result()`,
  `._expect_result_ctx`, `._final_result_msg/_pld`,
  `.wait_for_result()` + the deprecated `.result()` alias are
  gone — a `Portal` no longer has any "main result" notion.
  NB `Context.wait_for_result()` is a different (very alive)
  API and is untouched.
- `spawn/_spawn.py`: `exhaust_portal()` +
  `cancel_on_completion()` (the reaper tasks) deleted; backend
  comment sweeps in `_trio.py`/`_mp.py`.
- `_exceptions.py`: the `NoResult` sentinel dies with its lone
  reader.
- `tests/test_ringbuf.py`: drop a daemon-portal `.result()`
  call that was already a warn + `NoResult` no-op (the ctx-acm
  exit does the real result-wait); unshadow the 2nd `sctx` as
  `rctx`.
- comment/docstring x-ref sweeps: `msg/types.py`,
  `_context.py`, `to_actor/`, `tests/test_to_actor.py`.

Gate: `test_to_actor test_spawning test_cancellation
test_infected_asyncio test_local test_rpc` = 81 passed,
3 xfailed on `trio`; +`test_ringbuf` = 70 passed, 3 skipped,
3 xfailed on `mp_spawn`.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 3146409baf Fix stale `@pub` docstring example in `experimental`
The `_pubsub.pub` decorator's usage example predates several API
generations: ancient positional-arg-order `run_in_actor()` (a
missing `await` too) plus the deprecated `portal.result()` — and
`run_in_actor()` never allowed streaming funcs anyway. Show the
canonical `start_actor()` + `Portal.open_stream_from()`
consumption instead (#477 removal sweep).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi dd9d890921 Port docs off `run_in_actor` + `Portal.wait_for_result`
The 8-page docs sweep of the #477 removal, ahead of the API's
excision,

- `start/quickstart.rst`: the first-actor-tree walkthrough now
  narrates the (migrated) `to_actor.run()` example — no portal
  in hand until the daemon section introduces `start_actor()`.
- `guide/spawning.rst`: the one-shot section becomes
  `to_actor.run()` (blocking call, placement opts, "built on the
  primitives" note); lifetime/teardown rules update — one-shots
  never make it to nursery exit since each is reaped inside its
  own call.
- `guide/rpc.rst`: the `wait_for_result()` section (an API that
  dies with the reap cluster, incl. the `NoResult` sentinel)
  becomes a `to_actor.run()` one-shot section.
- `api/core.rst`: drop `run_in_actor`/`wait_for_result` from the
  autodoc member lists, drop the `Portal.result()` deprecation
  note, add a "One-shot task actors" `tractor.to_actor.run`
  autodoc section.
- `guide/{asyncio,context,cancellation,parallelism}.rst`:
  mention swaps to the successor API.

Gate: `make -C docs html` builds clean; `to_actor.run` autodoc
renders in `api/core.html`.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 0b926210a3 Port debugging examples off `run_in_actor`
The 8 `examples/debugging/` scripts driven by the pexpect'd
`test_debugger.py` REPL-flows (#477 removal),

- blocking one-shots (`subactor_error`, `subactor_breakpoint`,
  `shielded_pause`): straight `to_actor.run(fn, an=an)` — the
  boxed error/`BdbQuit` raises in the root's task.
- `multi_subactors`: introduces the "collect all errors" pattern
  — each one-shot catches + stashes its `RemoteActorError` (vs
  raising) so no child's crash cancels its siblings before
  they've had their own REPL sessions, then a
  `BaseExceptionGroup` of the lot raises at the end; preserves
  the legacy teardown-reap REPL flow exactly (28/28 debugger
  suite unchanged).
- `multi_nested_subactors_error_up_through_nurseries` +
  `root_cancelled_but_child_is_in_tty_lock`: recursive
  `spawn_until` levels each block on their one-shot child; the
  parallel spawner-trees run as bg task-nursery one-shots where
  the first tree's error cancels the other.
- `multi_subactor_root_errors` +
  `root_timeout_while_child_crashed`: `start_actor()` + bg
  `Portal.run()` tasks so the root's own error/timeout races the
  already-crashed children, same as before.
- `sync_bp`: TODO-comment x-ref update only.

`test_debugger.py`: the nested-nurseries test's final-output
patterns update to the new relay shape — the LAST-released leaf
REPL's error chain wins each level's relay-vs-cancel race and
relays as a `collapse_eg()`-annotated collapsed chain, while the
sibling tree is cancelled + absorbed. (The legacy teardown-reap
grouped BOTH the `name_error` and bp-quit chains — explaining
the previously-mysterious "extra" `src_uid`/`relay_uid` patterns
noted in the old TODO.)

Gate: `tests/devx/test_debugger.py` = 28 passed, 6 skipped —
identical to the pre-migration baseline.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 7c6bff50c5 Port non-debugging examples off `run_in_actor`
4 example scripts of the #477 removal sweep, each exercised by
`test_docs_examples.py`,

- `actor_spawning_and_causality.py`: the simplest possible
  `to_actor.run()` demo — private call-scoped nursery, block on
  and print the one-shot's result.
- `remote_error_propagation.py`: blocking `to_actor.run(an=n)`
  raises the boxed `AssertionError` in the caller's task,
  cancelling the sibling daemons.
- `parallelism/single_func.py`: bg-burn a core in the parent via
  a local task-nursery while the one-shot burns (and returns
  from) a subactor.
- `a_trynamic_first_scene.py`: donny + gretchen wait on each
  *other* so their one-shots run concurrently in a local
  task-nursery against a shared `an` (mirrors the migrated
  `test_trynamic_trio`).

Gate: all 4 green via the example-runner suite.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 17d6879706 Port `test_dynamic_pub_sub` off `run_in_actor`
The known-flaky dynamic pubsub test's 3 fire-and-forget spawn
sites (#477 removal),

- the forever-streaming `publisher` + N `consumer` one-shots now
  bg-schedule as `to_actor.run(fn, an=n)` tasks in a local `trio`
  task-nursery (`publisher`'s rendezvous name still derives from
  `fn.__name__`).
- the simulated user-cancel raise (`KeyboardInterrupt` /
  `TooSlowError` params) cancels the task-nursery, each one-shot
  reaping its subactor via `to_actor.run()`'s shielded
  `Portal.cancel_actor()`; `_run_and_match()`'s existing
  `BaseExceptionGroup.split()` walk covers the (possibly nested)
  relay shapes unchanged.
- spawns now issue concurrently rather than sequentially —
  comment on the fork-backend budget updated to match.

Gate: both params x4 runs green on `trio` + x1 on `mp_spawn`;
full module green.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 2af5acde88 Port SIGINT + sync-sleep cancel tests off `run_in_actor`
Final `test_cancellation.py` group of the `run_in_actor` removal
(#477) — cancel-mechanics tests, so clean conversions,

- `test_cancel_via_SIGINT_other_task`: the 3 keep-alive
  `run_in_actor(sleep_forever)` one-shots become plain
  `start_actor()` daemons (an idle daemon needs no "main" task,
  and no longer shares a single dup'd `namesucka` name).
- `spawn_sub_with_sync_blocking_task`: the middle layer's spawn
  becomes a blocking `to_actor.run(spin_for, an=an)` which parks
  awaiting the sync-sleeping grandchild's result until cancelled
  from above.
- `test_cancel_while_childs_child_in_sync_sleep`: the
  fire-and-forget middle-actor spawn becomes a bg
  `to_actor.run()` task in a local task-nursery; the root's
  `assert 0` cancels it, driving the same
  graceful-cancel-then-zombie-reap cascade on the sync-blocked
  grandchild. The `man_cancel_outer` xfail param is unchanged.

Zero live `run_in_actor()` call-sites remain in this suite.

Gate: full `test_cancellation.py` module green on both `trio`
(18p/1xf) + `mp_spawn` (18p/1xf).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 5c374b183e Port `test_nested_multierrors` off `run_in_actor`
Third `test_cancellation.py` group of the `run_in_actor` removal
(#477),

- `spawn_and_error` fans out each level's erroring one-shots as
  concurrent `to_actor.run(fn, an=an)` tasks in a local `trio`
  task-nursery (recursing per spawner subactor), as does the
  test-body's top-level spawner loop.
- the deterministic exact-breadth nested-BEG shape dies with the
  legacy teardown-reap: each level now groups whatever subset of
  sub-tree errors relay before the first one's cancel wins, and
  a single-member group gets unwrapped by the runtime's own
  `collapse_eg()` at every actor boundary — so a fully-raced
  tree relays a bare `RemoteActorError` chain.
- loosen the shape walk accordingly: accept a lone
  `RemoteActorError` or a 1..breadth group whose members box
  `ExceptionGroup` (multi-relay), `AssertionError` (collapsed
  leaf chain), `RemoteActorError` (re-boxed collapsed chain) or
  `BaseExceptionGroup` (runtime reap-deadline `Cancelled`
  upgrade); fold the windows-only tolerances into the same walk.
- raced sibling `trio.Cancelled`s are now ABSORBED by the
  task-nursery instead of landing in the group, so the MTF
  shape-mismatch xfail should consistently xpass — note added to
  drop the marker once CI confirms.
- add an `else: pytest.fail()` so a silently-clean tree can no
  longer pass.

Gate: both depths green on `trio` (10 consecutive runs) +
`mp_spawn`.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 29ac9067ce Fix unbound `timeout` under non-trio/MTF backends
`test_nested_multierrors`'s backend/depth budget `match` only
carries arms for the `trio` + `main_thread_forkserver` spawn
backends, so running under any other (e.g. `mp_spawn`) leaves
`timeout` unbound and crashes with an `UnboundLocalError` at the
headroom-scaling below. Add default per-depth arms riding the MTF
budgets (same per-spawn round-trip cost class).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi efce12e145 Port `test_some_cancels_all` off `run_in_actor`
Second `test_cancellation.py` group of the `run_in_actor` removal
(#477),

- one-shot subactors now run as concurrent `to_actor.run(fn,
  an=an)` tasks in a local `trio` task-nursery, so their errors
  raise WHILE the actor-nursery block is open (vs the legacy
  teardown-reap) and the first error cancels sibling one-shots.
- wrap the task-nursery in `collapse_eg()` so the deterministic
  single-error cases still surface a bare `RemoteActorError`.
- loosen the group-shape assertion: the relay-vs-cancel race
  populates anywhere from 1 to `num_actors` `RemoteActorError`s
  (the exact-`num_actors` BEG was `run_in_actor`'s
  reap-all-at-teardown); group members are always
  `RemoteActorError` now since sibling `trio.Cancelled`s are
  absorbed by the task-nursery.
- move the daemon-portal call loop inside the task-nursery body
  so the sleep-forever one-shot case is cancelled by the daemon
  error raise.
- rename the `*run_in_actor*` param ids to `*one_shot*`.

Gate: 6 passed on both `trio` + `mp_spawn` backends.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 0c3b21e9fe Doc ria-reap hang fix + paused reaper re-scope
Append two sections to the ria-removal plan capturing the
2026-07-02 hang episode + the resulting design pivot.

Regression writeup: the full-suite hang on
`test_tractor_cancels_aio` root-caused to the step-A reaper
hoist (`5cd190c5`), not the B2 handler merge. The happy-path
`_reap_ria_portals()` parks unbounded on `wait_for_result()`
after a user `portal.cancel_actor()`; the old spawn-backend
reaper raced `soft_kill()`'s scope-cancel, the hoist dropped
it. Records the `proc.poll()` death-watch fix + why poll (not
the event `wait_func`) bc `soft_kill` already awaits
`proc.sentinel` (a 2nd `wait_readable` -> `BusyResourceError`).

Pause writeup: user's insight that the hoist landed in the
wrong scope — result-waiting belongs in the `to_actor`
one-shot scope (`_invoke_in_subactor()`), beside `an` + a
local task-nursery + cancel-scope, where bounding the wait is
trivial + the hang dissolves. So the poll fix is likely
SUPERSEDED (flagged do-not-land); the anti-hang guard commit
(`d1fb4a1a`) stays red-first per the failing-test convention.

(this patch was generated in some part by `claude-code` using
`claude-opus-4-8` (`anthropic`))
2026-08-17 22:06:58 -04:00
Gud Boi 6a7e06e39d Port `test_cancellation` multierror cluster off `run_in_actor`
First group of the `test_cancellation.py` `run_in_actor` removal
(#477),

- `test_remote_error` -> blocking `to_actor.run()` (single erroring
  one-shot; a bad-arg `TypeError` still relays as a
  `RemoteActorError`).
- `test_multierror` -> concurrent fan-out via
  `gather_contexts([p.open_context(assert_err_ctx) ...])` over
  `start_actor()` portals. NB `gather_contexts` is cancel-on-first
  so the 2nd errorer is usually cancelled before relaying its own
  exc and the pair collapses to a single `RemoteActorError` (vs the
  legacy reap-all-at-teardown `BEG`-of-N) — the assertion now
  accepts either shape.
- delete `test_multierror_fast_nursery` — a 25-actor stress test of
  `run_in_actor`'s teardown-reap; no analogous surface under the
  `to_actor` fan-out.
- add an `assert_err_ctx` `@context` shim for the `open_context`
  fan-out.

Remaining `test_cancellation` groups (some_cancels_all, nested,
SIGINT, sync-blocking) still on `run_in_actor` — ported next.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi aa9c9a97ff Port `test_registrar` off `run_in_actor`
Two sites migrated (#477 removal),

- `test_trynamic_trio`: donny + gretchen each wait on the *other*
  to register, so they must run CONCURRENTLY — was two
  non-blocking `run_in_actor()`s awaited after; now two
  `to_actor.run()` one-shots scheduled into a local `trio`
  task-nursery.
- the unregister-on-cancel cluster test: its non-streaming branch
  spawned `run_in_actor(trio.sleep_forever)` purely to keep each
  subactor alive + registered — a `start_actor()` daemon does that
  without a "main" task, so the spawn loop collapses to the same
  `start_actor()` the streaming branch already used.

Suite: 16 passed.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 07326d2b3e Port `test_pubsub` off `run_in_actor`
`test_multi_actor_subs_arbiter_pub` used `run_in_actor()` to spawn
two forever-ish subscriber actors and hold their portals for a
later `cancel_actor()` (its `.result()` was commented out exactly
because `subs()` never cleanly returns). That deferred-spawn +
cancel shape isn't a blocking `to_actor.run()`, so convert to the
successor primitives (#477 removal),

- `start_actor()` per subscriber — keeps the portal for the
  existing `cancel_actor()` teardown,
- run `subs()` on each via a background `Portal.run()` task in a
  local `trio` nursery so both subscribe concurrently with the
  test's `wait_for_actor` / topic checks,
- each bg runner swallows the `RemoteActorError`/`ContextCancelled`
  that `cancel_actor()` relays; a trailing `tn.cancel_scope.cancel()`
  drops any lingering runner.

Suite: 8 passed.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi e876430787 Port `test_spawning` off `run_in_actor`
Migrate all 4 sites to blocking `tractor.to_actor.run()` (#477
removal),

- rename the two API-named tests to `test_to_actor_run_*`
  (`same_func_in_child`, `can_skip_parent_main_inheritance`) —
  they exercise the same spawn / `inherit_parent_main` path via
  the successor API.
- the recursive `spawn()` helper drops its white-box
  `an._children` / portal-`_peers` asserts (which probed
  `run_in_actor`'s portal + nursery-tracking internals);
  `to_actor.run()` returns the result and reaps internally, so
  keep the user-facing `result == 10` check.
- `test_most_beautiful_word` drops the 2nd `wait_for_result()`
  (the legacy result-cache re-fetch) — `to_actor.run()` delivers
  the value once, no cache.

Suite: 9 passed.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 0402994572 Port `test_rpc` off `run_in_actor`
Sole call-site: `run_in_actor(sleep_back_actor, ...)` ->
blocking `tractor.to_actor.run(..., an=n, ...)` (#477 removal).
The RPC-callback subactor is awaited in-caller instead of
reaped at nursery teardown; `name=`/`enable_modules=` map to
`to_actor.run()`'s same-named params, the rest to `**fn_kwargs`.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi fa748980d5 Port `test_runtime` off `run_in_actor`
Sole call-site: an inlined `run_in_actor(...).result()` ->
blocking `tractor.to_actor.run(fn, an=an, ...)` (#477 removal).
Behaviour identical — the one-shot's result/error is awaited
in the caller's task rather than reaped at nursery teardown;
the enclosing `move_on_after` still cancels the sub in the
`error_in_child=False` case.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 5b5235962b Port `test_infected_asyncio` off `run_in_actor`
First test-file of the #477 `.run_in_actor()` removal (blocking
`to_actor.run()` is the successor; the legacy non-blocking one-shot
is dropped, not replaced). All 9 call-sites migrated,

- blocking result/error/streaming-result tests -> `to_actor.run(fn,
  an=an, ...)`; the "streaming" ones stream aio<->trio INSIDE the
  subactor so the caller only awaits the final result.
- forever-task + cancel tests (`test_tractor_cancels_aio`,
  `test_trio_cancels_aio`) -> `start_actor()` +
  `Portal.open_context()` + cancel — can't block on a
  never-returning task. Adds a small `sleep_forever_aio_ctx`
  `@context` shim.
- greens the red `test_tractor_cancels_aio` anti-hang guard from
  the prior commit: under the correctly-scoped API the wait is
  bounded by the caller's cancel scope, so the hang is structurally
  gone — not patched.

Suite: 34 passed, 2 xfailed (trio backend).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi d8962f2a8a Add anti-hang `fail_after` cap to aio-cancel test
Wrap `test_tractor_cancels_aio`'s `main()` in a
`trio.fail_after(9 * cpu_perf_headroom())` so a wedged remote
runtime can't hang the test forever. This is the blessed
anti-hang guard here bc `pytest-timeout`'s global cap is
intentionally off (it breaks `trio` under the fork backends,
per the `pyproject` NOTE).

The cap is generous + CPU-headroom-scaled bc it's an anti-hang
guard, not a perf assertion. Motivated by the
`._ria_nursery`-removal regression where a wedged ria-reaper
once hung this exact test indefinitely.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi d73c3b55e0 Merge the supervise error handlers into one
Step B2 of the `._ria_nursery` removal (issue #477; see
`ai/conc-anal/ria_nursery_removal_plan.md`). With the 2ndary
nursery gone (step B), the two nested error handlers in
`_open_and_supervise_one_cancels_all_nursery` collapse to one,

- the outer `except (Exception, BaseExceptionGroup,
  trio.Cancelled)` existed to catch errors bubbling from the
  old `._ria_nursery.__aexit__` reaper-group; that nursery no
  longer exists.
- trace shows the outer handler's `raise` was already DEAD: the
  inner handler records `errors[uid]` as its first action, so
  `errors` is always non-empty by the time anything could reach
  the outer handler, and the `finally`'s raise-from-`errors`
  always superseded the outer `raise`.
- so fold both into a single `except BaseException as
  _scope_err` guarding the lone daemon nursery; the `finally`
  (unchanged) still raises the collected `errors` as a single
  exc or `BaseExceptionGroup`.
- drop the now-unused `outer_err`/`inner_err` locals.

Behaviour-preserving (net ~30 lines lighter); the big diff is
the one-level de-indent of the handler body. The two remaining
`maybe_wait_for_debugger()` guards collapse to the single
pre-teardown wait.

Prompt-IO: ai/prompt-io/claude/20260702T222544Z_9201a2ed_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 2029a154ed Doc step-B2 handler-merge + prompt-io
Split from the step-B2 code commit to keep the runtime diff
free of `ai/` meta noise,

- `ai/conc-anal/ria_nursery_removal_plan.md`: add a "Step-B2
  outcome" section — the dead-outer-`raise` trace, why the
  merge is behavior-preserving, and the gate results.
- `ai/prompt-io/claude/20260702T222544Z_9201a2ed_*`: NLNet
  provenance (log + unedited raw) for the step-B2 work.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 194ce5e392 Drop the vestigial `._ria_nursery`
Step B of the `._ria_nursery` removal (issue #477; see
`ai/conc-anal/ria_nursery_removal_plan.md`). With step A having
rerouted `.run_in_actor()` children onto the daemon nursery,
the 2ndary "run-in-actor" nursery spawns nothing and its stored
ref is never read — pure dead weight,

- collapse the inner `async with trio.open_nursery() as
  ria_nursery` layer in
  `_open_and_supervise_one_cancels_all_nursery`; `da_nursery` is
  now the single nursery for ALL subactors.
- `ActorNursery.__init__` loses the `ria_nursery` param + the
  `self._ria_nursery` attr; `start_actor()` loses its `nursery=`
  escape-hatch (spawns via `self._da_nursery` directly).
- `._cancel_after_result_on_exit` stays — still the ria-child
  discriminator for `_reap_ria_portals()`.

Behavior-preserving: a zero-task `trio.open_nursery()` only adds
a checkpoint. The two error handlers are KEPT (now nested under
the single nursery); merging them changes error/cancel
propagation and is deferred to its own PR (TODO left at the
outer `except`).

Prompt-IO: ai/prompt-io/claude/20260702T172233Z_5cd190c5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 64ee2eeb75 Doc step-B outcome + prompt-io
Split from the step-B code commit to keep the runtime diff
free of `ai/` meta noise,

- `ai/conc-anal/ria_nursery_removal_plan.md`: add a "Step-B
  outcome" section — the empty-nursery collapse, why it's
  behavior-preserving, the deliberate handler-merge deferral,
  and the targeted-gate result.
- `ai/prompt-io/claude/20260702T172233Z_5cd190c5_*`: NLNet
  provenance (log + unedited raw) for the step-B work.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 62b574a3d4 Hoist ria-reaping out of the spawn backends
Step A of the `._ria_nursery` removal (issue #477 follow-up, see
`ai/conc-anal/ria_nursery_removal_plan.md`): `.run_in_actor()`
children now spawn via the default daemon nursery and their
result-reaping moves up into the `ActorNursery` machinery,

- new `_supervise._reap_ria_portals()`: one
  `_spawn.cancel_on_completion()` task per ria child, run AFTER
  `._join_procs` is set — replacing the per-child reaper task the
  backends formerly spawned (keyed off
  `._cancel_after_result_on_exit` membership) which required
  routing such children into `._ria_nursery`.
- happy path: reap awaited right after `._join_procs.set()`,
  preserving "collect ria results before daemon join" sequencing.
- error path: snapshot ria `(portal, subactor)` pairs (backend
  `finally`s pop `._children` as procs reap), `await an.cancel()`,
  THEN a 0.5s-bounded reap over the snapshot; anything collectable
  is already queued in the local ctx and a parked reaper
  self-cleans (`trio.Cancelled` results are never stashed). NB: a
  concurrent reap+cancel variant deadlocked `test_multierror` and
  a 3s bound blew `test_cancel_while_childs_child_in_sync_sleep`'s
  deadline — deats in the plan doc's probe history.
- `spawn/_trio.py` + `spawn/_mp.py`: drop the membership branch,
  per-child reaper nursery + now-unused `cancel_on_completion`
  imports; the join phase is a bare `soft_kill()`.

`._ria_nursery` is now vestigial (zero spawn users): step B
deletes it + `start_actor()`'s `nursery=` escape hatch and merges
the supervisor's two error handlers.

Prompt-IO: ai/prompt-io/claude/20260702T165806Z_a34aaf98_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 93227086b0 Add `_ria_nursery` removal plan + step-A prompt-io
Split from the step-A code commit to keep the runtime diff
free of `ai/` meta noise,

- `ai/conc-anal/ria_nursery_removal_plan.md`: agent-verified
  machinery map + 3-step (A/B/C) design + probe history
  (reap-relocation deadlock -> sequencing fix -> bound
  tighten) + risk register for the `._ria_nursery` excision.
- `ai/prompt-io/claude/20260702T165806Z_a34aaf98_*`: NLNet
  provenance (log + unedited raw) for the step-A work.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:58 -04:00
Gud Boi 4151b9569a Add `to_actor` one-shot parallelism example
Demo both flavors of the new API in a runnable script
(auto-collected by `test_docs_examples.py`),

- the fully-implicit one-shot which boots (and tears down) the
  actor-runtime around a single `to_actor.run()` call,
- the concurrent "worker-pool-ish" prime-check pattern: a local
  `trio` task nursery scheduling one-shots against a shared
  caller-managed `an`, mirroring (in miniature) the neighboring
  `concurrent_actors_primes.py` example per issue #477.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:08 -04:00
Gud Boi 8d04af0d73 Add `tests/test_to_actor.py` one-shot API suite
Cover every placement variant + failure mode of the new
`to_actor.run()`,

- private-nursery one-shot + implicit runtime boot via pass-through
  `runtime_kwargs`,
- remote-error relay to the caller's task (bare and inside a
  caller-managed `an`) as boxed `RemoteActorError`s,
- caller-nursery spawn + portal-reuse w/o implicit reap,
- the concurrent "worker-pool-ish" pattern: a local `trio` task
  nursery scheduling one-shots against a shared `an`,
- the 4 pre-spawn validation rejections (sync fn, async-gen fn,
  `portal`+`an` combo, `runtime_kwargs`+placement combo).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:08 -04:00
Gud Boi a3cd448612 Add `tractor.to_actor` one-shot task API subpkg
First cut at the `to_thread`/`to_process`-style "run it over there"
wrapper layer from issue #477: a single-remote-task invocation API
decoupled from the `ActorNursery` spawn machinery, composed purely
from the lower level daemon-actor + portal primitives,

- `to_actor.run(fn, **fn_kwargs)` spawns a subactor via
  `ActorNursery.start_actor()`, schedules `fn` as its lone task
  with `Portal.run()` and ALWAYS reaps it via a `finally`-scoped
  `Portal.cancel_actor()` (whose bounded cancel-req wait is
  internally shielded so the reap also runs under caller-scope
  cancellation).
- remote errors raise directly in the caller's task as boxed
  `RemoteActorError`s, moving error collection/propagation up into
  whatever local `trio` scope encloses the call.
- "placement" opts: `portal=` reuses a running actor (no
  spawn/reap), `an=` spawns from a caller-managed actor-nursery,
  neither opens a private call-scoped `open_nursery()` (implicitly
  booting the runtime, tunable via pass-through `runtime_kwargs`).
- fail-fast validation BEFORE any spawn: non-streaming async fn
  only (same constraint as `Portal.run()`), `portal=`/`an=` mutual
  exclusion and no `runtime_kwargs` alongside a placement opt.

Also,
- x-ref the successor API from `.run_in_actor()`'s deprecation TODO
  + docstring; emitting a formal `DeprecationWarning` waits on
  migrating in-repo usage.
- log prompt-io provenance per NLNet policy incl. the driver prompt
  file.

Prompt-IO: ai/prompt-io/claude/20260702T154255Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:08 -04:00
Bd 4e27dcde48
Merge pull request #475 from goodboy/windows_support_round2
Restore Windows support (optional UDS + `SIGUSR1`)
2026-08-17 20:05:31 -04:00
Gud Boi 93322405a1 Restore IPC transport style conventions
Bring the transport helpers back in line with project style:

- restore single-quote strings and docstrings;
- drop the oversized helper divider and simplify comments;
- keep guarded `match` dispatch so a missing `socket.AF_UNIX`
  remains safe when `HAS_UDS` is false.

Also, format the `HAS_UDS` conjunction with the project's
multiline branch convention.

Review: PR #475 (goodboy)
https://github.com/goodboy/tractor/pull/475

Prompt-IO: ai/prompt-io/opencode/20260817T231825Z_359fe75c_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-17 19:37:44 -04:00
Gud Boi 359fe75ced Tolerate only Windows `pytest` failures
Job-level `continue-on-error` made setup and import-smoke failures
non-blocking even though the smoke is the hard Windows support
signal.

Move tolerance to the `pytest` step so the incomplete suite remains
informational while install, dependency, and `HAS_UDS` smoke
failures fail the job. When only the known suite fails, the job and
PR rollup remain green with the failed step still visible.

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-17 17:04:20 -04:00
Gud Boi 40be587ce2 Keep `HAS_UDS` false on Windows
Modern Windows Python can expose `AF_UNIX`, so `trio.has_unix`
alone can register the UDS backend even though its credential and
lifecycle paths remain POSIX-only.

Gate `HAS_UDS` on `sys.platform` and make the Windows import smoke
step assert the TCP-only capability contract.

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-17 12:09:52 -04:00
Gud Boi ca4582b003 Skip Windows-hanging shm IPC test
`test_parent_writer_child_reader` deadlocks on Windows (the
parent/child shared-mem transfer hangs at the larger frame size),
so the `windows-latest` CI leg ran to the 16-min job cap instead
of completing. It's a genuine nascent-Windows shm bug, not a
clean "unsupported", so it's `skipif`'d (not removed) and tracked
under #404; `test_child_attaches_alot` still runs on Windows.

- `@pytest.mark.skipif(platform.system() == 'Windows', ...)` on
  the parametrized `test_parent_writer_child_reader` so the leg
  completes + reports. linux/macOS unaffected (all 6 variants
  still run).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:17:11 -04:00
Gud Boi 75385e448d Derive transport registries from one list
The five transport/address lookup maps each hand-guarded `uds`
with its own `if HAS_UDS:` block (3 in `ipc/_types`, 2 in
`discovery/_addr`) — easy to let drift so a backend half-registers
(known by address but not by key, listed but unroutable, &c).

- build one `_msg_transports` / `_address_protos` list per module
  (TCP always, UDS only when `HAS_UDS`), then DERIVE every map
  from it via each backend's ClassVars (`codec_key`,
  `address_type`, `proto_key`). Adding a backend now touches one
  list, and the maps can't disagree.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:17:09 -04:00
Gud Boi 34d81adef3 Skip `infect_asyncio` tests on Windows
`tractor`'s infect-asyncio mode runs an `asyncio` loop under
`trio` guest-mode; on Windows the default `ProactorEventLoop` is
incompatible and the suite hangs/crashes mid-run (orphaned py
procs), so the `windows-latest` CI leg never finishes reporting.

- add a module-level `pytest.skip(allow_module_level=True)` gated
  on `platform.system() == 'Windows'` to `test_infected_asyncio`
  and `test_root_infect_asyncio`, before their asyncio-interop
  imports. macOS/linux are unaffected (they run these fine).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:02 -04:00
Gud Boi 1691be96fd Scale `test_lifetime_stack` deadline for CI
`test_lifetime_stack_wipes_tmpfile` guards spawn+teardown with a
hard-coded `trio.move_on_after()` (1.6s / 1s) that isn't scaled
for slow CI. On a noisy macOS runner the `error_in_child=True`
case times out before the child error propagates, so the scope
cancels and `assert not cs.cancel_called` flips — reddening the
(required) macOS leg. Same unscaled-deadline class `main` already
fixed for `test_dynamic_pub_sub`.

- multiply the budget by `cpu_perf_headroom()` (`tests/conftest`),
  the established deadline-headroom helper (3x on macOS CI, a
  1.0 no-op locally / on un-throttled linux).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:02 -04:00
Gud Boi 36ad1f3dd0 Skip `test_ringbuf` at collection off-linux
`tests/test_ringbuf.py` imports `tractor.ipc._ringbuf` at module
top, which pulls in `tractor.ipc._linux` whose module-level
`ffi.dlopen(None)` raises `OSError` on Windows (and any non-linux
host). That fires at COLLECTION, before the module's existing
`pytestmark = pytest.mark.skip` can apply, so it aborts the whole
pytest session — the new `windows-latest` CI leg never gets past
collection.

- guard the module with `pytest.skip(allow_module_level=True)`
  gated on `platform.system() != 'Linux'`, placed before the
  crashing import — same idiom as `tests/devx/test_debugger.py`.
- the `eventfd`-based ringbuf backend is linux-only by design, so
  macOS skips cleanly too (previously it only skipped incidentally
  via the absent `cffi` optional dep).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:02 -04:00
Gud Boi 72441124c0 Fix Windows `import tractor` at the UDS root
The prior round gated UDS in four modules but `import tractor`
still crashed on Windows: `tractor.ipc._uds` does `from socket
import AF_UNIX` at module top, and several modules in the import
graph (`discovery._api`, `spawn._reap`, `discovery._multiaddr`,
`_testing.addr`) import `_uds` unconditionally. Instead of
guarding every importer, fix the root and collapse the per-module
probes to one capability flag.

- in `ipc/_uds.py`, guard the lone `AF_UNIX` import so the module
  stays importable everywhere; expose `HAS_UDS = trio.has_unix`
  as the single source of truth (the same predicate that gates
  `trio.open_unix_socket()`).
- `ipc/_types.py`, `discovery/_addr.py` and `ipc/_server.py` now
  import `UDSAddress`/`MsgpackUDSStream`/`HAS_UDS` directly and
  gate the transport + address registries on `HAS_UDS`; drop the
  duplicated `getattr(socket,'AF_UNIX')` / `platform.system()`
  probes, the dead `HAS_AF_UNIX` conjunct, and the import-time
  `log.warning()` spam.
- `devx/_stackscope.py` `enable_stack_on_sig()` early-returns
  when `sig is None`, so a missing `SIGUSR1` (Windows) degrades
  to a no-op instead of a `TypeError` from `getsignal()` /
  `signal()`.
- add a `windows-latest` CI leg (UDS excluded; informational via
  `continue-on-error` while support matures) plus an `import
  tractor` smoke step as the hard signal for the import fix.

Because `_uds` is importable everywhere `UDSAddress` stays a real
class, so `isinstance()` checks and `wrap_address()` no longer
`AttributeError` on no-UDS hosts; actual socket use stays gated
on `has_unix`.

Review: https://github.com/goodboy/tractor/pull/475
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:01 -04:00
Gud Boi 0cbb850645 Make UDS + `SIGUSR1` optional for Windows
Windows (and any CPython that doesn't expose `socket.AF_UNIX`)
can't import the UDS transport backend nor `signal.SIGUSR1`, so
the unconditional imports break `import tractor` outright on
those hosts. Guard the platform-specific bits behind capability
probes and fall back to a TCP-only runtime when the UDS backend
is unavailable.

- across `discovery/_addr.py`, `ipc/_server.py` and
  `ipc/_types.py`, gate on `getattr(socket, 'AF_UNIX', None)` +
  `platform.system()` and import `UDSAddress` /
  `MsgpackUDSStream` only when supported, leaving the names as
  `None` otherwise.
- register the `'uds'` key in `_address_types`, its
  default-loopback addr, and the transport lookup maps only when
  the backend actually loads, so TCP keeps working standalone.
- in `devx/_stackscope.py`, import `SIGUSR1` conditionally and
  set it to `None` on Windows.

Rebased onto the post-reorg tree where `_addr.py` now lives
under `tractor/discovery/`; adapt the relocated imports to the
package's `..ipc._uds` / `..ipc._tcp` paths (the original
single-dot paths would silently disable UDS on POSIX) and drop a
duplicated `TYPE_CHECKING` block and dead `import logging` left
by the move.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 22:35:40 -04:00
Bd 8f0df0ff6f
Merge pull request #480 from goodboy/wkt/uds_macos_473
Fix `uds` addr corruption + exercise macOS CI
2026-08-14 19:14:02 -04:00
Gud Boi f75c9cdeab Traverse `BaseExceptionGroup` peer-close errors
`_peer_closed_errno()` followed cause and context links but did not
descend through grouped exceptions. A reset below a group could
therefore escape as `trio.BrokenResourceError` instead of the
normalized `TransportClosed` boundary.

Walk the exception tree with cycle protection, requiring every
group branch to represent peer closure before normalization. Extend
the `MsgpackTransport.send()` regression to prove all-transport and
mixed-failure behavior.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 18:54:56 -04:00
Gud Boi 75cda1933c Move registrar probes to `discovery._api`
Registrar election and multi-address probing are discovery-protocol
concerns, but their implementation lived in root-runtime ignition.
Move the bounded handshake probe and its concurrent address
classifier into `discovery._api`, leaving `open_root_actor()` to
consume the classified results.

Update probe tests for the canonical module and clarify that the
daemon-fixture regressions directly exercise their sibling
`conftest` plugin.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 18:43:36 -04:00
Gud Boi f454cefe56 Enforce `get_rt_dir()` ownership on POSIX
Linux previously accepted a pre-existing runtime bindspace without
checking its owner or mode, even though Darwin enforced both. Share
the POSIX directory guard so every managed root and subdir rejects
non-directories and foreign UIDs before changing permissions.

Normalize owner-controlled bindspaces to `0o700` and add Linux
regressions for mode repair, foreign ownership, and non-directory
paths.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 18:40:23 -04:00
Gud Boi d0cc06815f Accept raced `TransportClosed` diagnostics
The `@context` debugger E2E intentionally closes its channel.
Teardown can surface from either local error shipment or the peer
receive task.
The RPC response fix makes local close win under CI, while the test
required both scheduler-dependent diagnostics.

Keep the common debugger and cancellation assertions, then accept
either transport-close report. This preserves real actor-tree
teardown coverage without depending on task scheduling order.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 12:49:40 -04:00
Gud Boi ebf2258b4f Keep accepted RPCs alive on `TransportClosed`
An RPC caller can close its channel after the callee creates the
endpoint coro but before its `StartAck` or final response lands.
Treat those response-send failures as terminal delivery failures so
the accepted endpoint still runs and application errors stay local
instead of cancelling the shared service nursery.

Register each cancellable RPC in `Actor._rpc_tasks` before
publishing its `Context` through `TaskStatus.started()`.
This closes the checkpoint-free completion race without a
cross-task handoff and keeps `Actor._ongoing_rpc_tasks` balanced
through existing cleanup.

Add regressions for caller disconnects at `StartAck` and error
shipment, plus registration-before-execution and cleanup checks.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 12:28:56 -04:00
Gud Boi ac61d7a5bf Use a short UDS reaper test bindspace
Darwin's pytest `tmp_path` can already exceed the 104-byte AF_UNIX
budget before appending either synthetic socket filename. That made
the new sentinel-policy regression fail identically in both macOS
matrix legs without exercising reaper behavior.

Allocate the test's real socket files under a short
`/tmp/tractor-reap-*` directory and retain scoped cleanup.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:59:07 -04:00
Gud Boi 584ea4e9ad Document stable macOS UDS operation
The remediation grew beyond the original no-autobind fix, leaving
public docs and nearby comments describing connect-only discovery,
XDG-only socket paths, raw readiness probes, and old cleanup naming.

Document typed registrar probing, occupied-address rejection, Darwin's
short runtime directory, platform-aware socket cleanup, sentinel
readiness, and subprocess output draining. Add the GH #473 bugfix
fragment and update regression rationale without changing task states.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:37:19 -04:00
Gud Boi 16dd876b0c Make UDS reaping platform-aware
Moving Darwin sockets to `/tmp/tractor-<uid>` left pytest and the
standalone reaper searching only `XDG_RUNTIME_DIR`. Enabling that
shared bindspace naively would also let automatic session cleanup
unlink an independent live `registry@1616.sock`.

Resolve the runtime's actual default UDS bindspace on every platform.
Exclude the pid-less registry sentinel from automatic cleanup while
retaining explicit CLI removal, and document that destructive choice.

Cover bindspace resolution and automatic-vs-explicit sentinel policy.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:36:27 -04:00
Gud Boi 1a9ce915f3 Harden inbound actor handshakes
Registry probes need short retry deadlines, but applying their
one-second budget to every portal and child connection can terminate a
valid delayed actor with no client retry path.

Give ordinary pre-registration handshakes an independent ten-second
deadline. Normalize raw `msgspec.DecodeError` frames to
`TransportClosed` so malformed peers cannot cancel the shared IPC
nursery with decoder internals.

Cover malformed frames and the ordinary-vs-probe timeout distinction.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:34:59 -04:00
Gud Boi 0b63af020e Retry transient registrar handshakes
A loaded runner can accept a registry transport while delaying its
actor handshake beyond the first one-second attempt. Treating that
timeout as final classifies a healthy daemon as occupied and cascades
into unrelated discovery failures.

Retry connected handshake failures on fresh channels under a shared
three-second budget, while returning immediately for truly absent
listeners. Bound every connect-plus-handshake attempt, use incremental
backoff, and retain the one-second unauthenticated server limit.

Also bound shielded probe-channel cleanup to 200ms and cover the real
timeout, reconnection, backoff, fresh-channel, and stalled-close paths.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 23:40:12 -04:00
Gud Boi 6e424d4696 Use a daemon-ready sentinel in discovery tests
The discovery `daemon` fixture probed UDS readiness by connecting and
immediately closing. That entered Tractor's actor handshake without an
`Aid` payload and destabilized the remote registrar on macOS before
test roots attempted discovery.

Run the child through a small `open_root_actor()` wrapper and publish a
filesystem sentinel only after runtime startup completes. Poll that
sentinel with process-liveness checks and guaranteed setup-failure
cleanup, without touching the transport socket.

Cover transport-free readiness and deterministic polling backoff.

Caught-during: review remediation
Found-via: `/run-tests` discovery daemon fixture consumers
Cause: initial sentinel drafts mishandled `pytest.Testdir`, emitted
  invalid `python -c` syntax, and misplaced return-code logging.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 20:38:22 -04:00
Gud Boi 5724c0516a Probe registrar capability in actor handshakes
Transport connect alone can select a foreign, stalled, or ordinary
actor endpoint as the registry. On macOS UDS this also exercises a
fragile connect-and-bail path before every root election.

Extend `Aid` with backward-compatible probe and registrar capability
fields, require a bounded typed handshake, and classify addresses as
absent, occupied, or confirmed registrars. Reject occupied endpoints
instead of binding over them.

Also,
- close `_connect_chan()` in `finally`
- bypass normal peer tracking for election probes
- preserve legacy registrar handshakes with unknown capability
- cover foreign listeners and idle registrar peer state

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 20:04:48 -04:00
Gud Boi 0d6d7c2a63 Keep UDS post-kill cleanup best-effort
`unlink_uds_bind_addrs()` reconstructs self-assigned socket paths
after a hard kill. An over-budget bindspace can make
`UDSAddress.get_sockname()` raise before the guarded `os.unlink()`,
replacing the original supervision outcome after the child is gone.

Catch and report reconstruction failures, then skip cleanup without
raising. Cover the overflow path and prove no unlink is attempted.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:57:13 -04:00
Gud Boi bd38204fde Preserve docs-example body failures
Best-effort subprocess teardown must not replace the exception raised
by the test body. Suppress cleanup errors only while propagating that
active failure; keep raising teardown errors on normal body exit.

Also close the Windows stdin pipe after its bounded leader reap.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:41:20 -04:00
Gud Boi 340d506940 Bound generated UDS socket paths
Actor names and custom runtime subdirs can otherwise produce unsafe
or overlong pathname sockets after moving Darwin's bindspace to
`/tmp`.

Deats,
- hash unsafe or over-budget actor names while retaining `@pid.sock`
- share deterministic naming with post-kill socket cleanup
- enforce Linux and Darwin `sun_path` byte budgets
- restore non-Darwin dir validation and mode `0700`
- validate every nested Darwin runtime-dir component

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:40:55 -04:00
Gud Boi 64e820e18e Normalize peer resets in `.send()`
A raw UDS readiness client can disconnect before the actor handshake.
Darwin reports the first server write as `ECONNRESET`, wrapped in
`trio.BrokenResourceError`; letting it escape cancels the daemon's
shared IPC nursery and makes later roots elect themselves registrar.

Walk the exception chain for `EPIPE` or `ECONNRESET` and translate
either into the existing `TransportClosed` boundary. Also handle
argument-less resource errors without raising `IndexError`.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:30:44 -04:00
Gud Boi 0a3b0efcc6 Reap timed-out docs example trees
Always drain docs-example pipes, even when the process exited before
the first status check, and decode malformed output with replacement
so diagnostics preserve the original failure.

Run POSIX examples in dedicated sessions and kill the full process
group on timeout. Reap the leader in every exit path, with bounded
Windows cleanup that cannot wait forever on descendant-held pipes.

Cover fast non-zero exits, invalid output bytes, process-group
termination, and post-timeout reaping.

Review: PR #480 (goodboy,copilot-pull-request-reviewer[bot])
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 16:40:05 -04:00
Gud Boi 634d914161 Use short Darwin UDS runtime paths
`platformdirs` can place the runtime dir deep below a pytest temp
home, pushing `registry@1616.sock` past Darwin's 104-byte
`AF_UNIX` limit.

Use a compact `/tmp/<app>-<uid>` root on Darwin and secure it
before allocating socket paths:
- require a real, current-user-owned dir via `lstat()`
- tighten existing roots to mode `0700`
- reject symlinks and unsafe nested dirs

Cover the path budget, mode, and symlink rejection.

Caught-during: review remediation
Found-via: `/run-tests` test_macos_rt_dir_fits_uds_path_limit

Review: PR #480 (goodboy,copilot-pull-request-reviewer[bot])
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 16:39:32 -04:00
Gud Boi 451e0acf8a Add prompt-io log for GH #473 UDS-on-macOS work
Provenance entry (+ unedited raw output) for the root-cause
session behind the prior three patches, per the NLNet
generative-AI policy tracked under `ai/prompt-io/`.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Gud Boi 3d02a8569e Run UDS-on-macOS CI leg, un-skip the example
UDS-on-macOS is otherwise un-exercised: the matrix
explicitly excludes the `macos-latest` + `uds` combo, so the
`uds_transport_actor_tree.py` example (skipped on macOS CI
since PR #460) is the only thing that ever touches that
path.

- drop the matrix `exclude` so the full suite runs with
  `--tpt-proto=uds` on `macos-latest`.
- un-skip the example on macOS+CI; with example-stderr
  surfacing in place a still-red run now yields the full
  traceback GH #473 asks for instead of a bare returncode
  assert.

Task-bullets 3 + 4 of GH #473.

Prompt-IO: ai/prompt-io/claude/20260702T155006Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Gud Boi bceca74eb7 Fix UDS addr corruption sans-autobind (macOS)
`MsgpackUDSStream.get_stream_addrs()` matches the
`(peername, sockname)` pair by type to find the listener's
fs-path, but the `(str, str)` arm unconditionally takes
`peername`: on platforms without linux's
`SO_PASSCRED`-triggered autobind (macOS!) the accept side's
`getpeername()` is `''`, so every accepted conn gets garbage
`Path('')` laddr/raddr structs.

Proven on linux by disabling `SO_PASSCRED` (no autobind ->
same `''` shape as darwin): the `uds_transport_actor_tree.py`
example reports `listener sock file: .` pre-fix and the real
registry sockpath post-fix.

- pick the non-empty name in the `(str, str)` arm: `peername`
  on the connect side, `sockname` on the accept side; raise
  `ValueError` on an (unexpected) empty pair.
- document the linux-autobind origin of the `bytes` arms
  which the original impl noted as "unclear".
- `start_listener()`: create the bindspace dir with
  `parents=True, exist_ok=True` (nested custom `filedir`s +
  racing actors).
- example docstring: peer-pid comes via `SO_PEERCRED` on
  linux but `LOCAL_PEERPID` on macOS.

May not be the (only) macOS crasher for GH #473 — it is
non-fatal on the linux sim — but with stderr surfacing now
in place the next macOS CI run pins any remaining layer.

Prompt-IO: ai/prompt-io/claude/20260702T155006Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Gud Boi 18faffcff7 Surface example stderr on any non-zero exit
The docs-example harness only re-raises captured subproc
stderr when the LAST line contains 'Error', but a `tractor`
root-actor crash always ends stderr with the strict-EG
collapse note `( ^^^ this exc was collapsed from a group ^^^ )`
— so every possible crash is swallowed down to a bare
`assert 1 == 0`, exactly what the macOS CI leg shows for the
UDS example in GH #473.

- raise with the FULL stderr (+stdout) whenever the example
  subproc exits non-zero, regardless of stderr shape.
- keep the legacy last-line 'Error' check for zero-rc runs
  which still emit error-ish output.

First task-bullet of GH #473.

Prompt-IO: ai/prompt-io/claude/20260702T155006Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Bd 7554f90e59
Merge pull request #478 from goodboy/wkt/boot_latency_470
Trim `import tractor` 0.42s -> 0.15s (gh #470)
2026-08-13 15:02:13 -04:00
Gud Boi 3705bbe594 Guard the cold import latency budget
Run seven fresh interpreters and gate the median `import tractor`
time below a conservative 0.35s ceiling. Provide an explicit
environment override for platforms with a different baseline.

Also assert the optional modules deferred by this patch remain
unloaded so a timing pass cannot hide an eager-import regression.

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 02:35:07 -04:00
Gud Boi 66ac7863b5 Advertise the lazy `to_asyncio` API
Preserve the existing public wildcard surface while adding the lazy
`to_asyncio` submodule to `__all__` and `dir(tractor)`.

Exercise both APIs in cold interpreters and verify normal package
import still leaves `asyncio` unloaded.

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 20:55:41 -04:00
Gud Boi 089e158da9 Retain dynamic logger caller discovery
Keep the `sys.modules` lookup as the normal fast path, then use
`inspect.getmodule()` for unregistered `runpy`, plugin, and `exec()`
namespaces.

Cover an unregistered module name backed by a real package file and
verify implicit logger naming still resolves to that package.

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 20:52:47 -04:00
Gud Boi 7987f6b8d7 Keep lazy annotations runtime-resolvable
Provide import-free runtime aliases for annotation-only actor and
multiaddr types so `typing.get_type_hints()` remains usable without
eagerly loading optional dependencies.

Correct `_address_types` to its actual `dict` shape and cover the
affected discovery and transport APIs.

Caught-during: review remediation
Found-via: `/run-tests` test_lazy_annotation_names_resolve

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 20:52:15 -04:00
Gud Boi 70c7e334a7 Make sub-spawn strategy toggleable in `we_are_processes`
The #470 boot-latency example hard-coded spawning each `worker_<i>`
subactor concurrently from a bg `trio.Task` (so each child's cold
`import tractor` overlaps). Add a `main()` `spawn_subs_in_bg_tasks`
flag so the serial-spawn path can be demo'd/compared too: flip it
`False` to `start_actor()` each sub inline in the loop before
handing the ready `Portal` to the bg task.

Deats,
- factor an `open_ep(ptl, i)` helper out of `spawn_and_open_ep()` -
  just the `Portal.open_context()` + `wait_for_result()` half, now
  that the spawn step is caller-optional.
- `spawn_and_open_ep()` grows a `maybe_ptl: Portal|None = None`
  param: spawn the subactor itself when unset (bg-task path), OW
  reuse the pre-spawned one (serial path).
- move the "overlap cold imports" rationale comment onto the new
  `main()` param where the toggle now lives.

(this commit msg was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:59:49 -04:00
Gud Boi 62b729a106 Add prompt-io entries for gh #470 latency work
Log the AI-assisted session per the NLNet generative-AI
policy: prompt, profiling findings, per-file diff pointers,
measured results and the unimplemented `pdbp`/`platformdirs`
deferral follow-ups.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
Gud Boi 213a298aad Defer `asyncio` import via PEP-562 `.to_asyncio`
`asyncio` (~5ms) only matters for infected-aio actors yet gets
imported by every cold `import tractor` via module-lvl
`.to_asyncio` imports in the debug-REPL + spawn-entry mods.

Deats,
- `devx.debug._trace`/`._tty_lock`: mv `import asyncio` under
  `TYPE_CHECKING` + fn-local it at the two
  `asyncio.current_task()` call-sites; fn-local the
  `run_trio_task_in_future` imports in the infected-aio-only
  branches.
- `spawn._entry`: fn-local `run_as_asyncio_guest` inside the
  `infect_asyncio=True` branches of `_mp_main()`/
  `_trio_main()`.
- `tractor/__init__.py`: add a PEP-562 module `__getattr__`
  lazy-loading `.to_asyncio` on first attr-access so the
  public `tractor.to_asyncio.<attr>` API (e.g.
  `LinkedTaskChannel` annots in
  `test_child_manages_service_nursery.py` + downstream users)
  keeps working unchanged.

Prompt-IO: ai/prompt-io/claude/20260702T155626Z_65bf9df5_prompt_io.md
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
Gud Boi d1d3fc58bb Lazy-import optional 3rd-party deps (gh #470)
Move every import-time-only-by-accident dep off the eager
`import tractor` path so cold child-actor boots only pay for
what they actually use:

- `bidict` -> `TYPE_CHECKING` in `discovery._addr`
  (annotation-only; `_address_types` is a plain `dict`
  literal).
- `multiaddr` -> `TYPE_CHECKING` + fn-local imports in
  `discovery._multiaddr.mk_maddr()`/`parse_maddr()`; also
  `TYPE_CHECKING` the `Multiaddr` annots in `ipc._tcp`/`._uds`
  (adds future-annots to `._multiaddr`).
- `colorlog` -> fn-local in `log.get_console_log()`.
- `pdbp` + `wrapt` -> fn-local in
  `devx._frame_stack.hide_runtime_frames()`/`api_frame()`.
- `platformdirs` -> fn-local in `runtime._state.get_rt_dir()`.

Still eager (documented follow-ups),
- `pdbp` via `devx.debug._repl` class-bases
  (`PdbREPL(pdbp.Pdb)`) + the module-lvl `@pdbp.hideframe` in
  `._tty_lock`; needs a `._repl` restructure.
- `platformdirs` via the `UDSAddress.def_bindspace: ClassVar`
  class-body eval of `get_rt_dir()`; needs an `Address`-proto
  rework.
- `stackscope` is already fn-local; `setproctitle` is not
  imported anywhere.

Prompt-IO: ai/prompt-io/claude/20260702T155626Z_65bf9df5_prompt_io.md
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
Gud Boi a2c8af558e Use `sys._getframe()` in `get_logger()` mod lookup
`get_caller_mod()` (nested in `get_logger()`) walks the WHOLE
call-stack via `inspect.stack()`, which also resolves src-file
info for every frame and scans all of `sys.modules` per frame
via `inspect.getmodule()`. During nested imports (deep
importlib stacks) each module-level `get_logger()` call costs
~5-10ms, making the ~39 such calls dominate `import tractor`
wall-time: ~244ms of the ~420ms total (see gh #470).

Deats,
- resolve the caller frame with `sys._getframe(frames_up)` and
  map its `f_globals['__name__']` through `sys.modules`: O(1)
  vs. O(stack x sys.modules).
- guard `ValueError` (stack too shallow) -> `None`, matching
  the existing null-caller handling at all use-sites.
- drop the now-unused `inspect` imports; pull `FrameType` from
  `types` instead.

Results: `import tractor` drops 0.42s -> ~0.155s; sequential
`.start_actor()` spawn latency ~0.42 -> ~0.18s/actor.

Prompt-IO: ai/prompt-io/claude/20260702T155626Z_65bf9df5_prompt_io.md
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
52 changed files with 2915 additions and 446 deletions

View File

@ -91,6 +91,10 @@ jobs:
name: '${{ matrix.os }} Python${{ matrix.python-version }} spawn_backend=${{ matrix.spawn_backend }} tpt_proto=${{ matrix.tpt_proto }}'
timeout-minutes: 16
runs-on: ${{ matrix.os }}
# Windows support is nascent: its full test suite remains
# informational, while setup and the `import tractor` smoke below
# are hard signals. Promote the test step to required once the
# suite is green.
strategy:
fail-fast: false
@ -98,6 +102,7 @@ jobs:
os: [
ubuntu-latest,
macos-latest,
windows-latest,
]
python-version: [
'3.13',
@ -118,10 +123,10 @@ jobs:
'tcp',
'uds',
]
# https://github.com/orgs/community/discussions/26253#discussioncomment-3250989
exclude:
# don't do UDS run on macOS (for now)
- os: macos-latest
# UDS is POSIX-only; Windows has no `AF_UNIX` so the
# backend is intentionally unavailable there.
- os: windows-latest
tpt_proto: 'uds'
steps:
@ -150,7 +155,14 @@ jobs:
- name: List deps tree
run: uv tree
# hard signal for the Windows import-safety fix: `import
# tractor` must succeed everywhere, and `HAS_UDS` reflects
# platform capability (False on Windows, True on POSIX).
- name: 'Smoke: import tractor'
run: uv run python -c "import sys; import tractor; from tractor.ipc._uds import HAS_UDS; assert sys.platform != 'win32' or not HAS_UDS; print('import tractor OK | HAS_UDS=', HAS_UDS)"
- name: Run tests
continue-on-error: ${{ matrix.os == 'windows-latest' }}
run: >
uv run
pytest

View File

@ -0,0 +1,77 @@
---
model: claude-fable-5
service: claude
session: b8ec2ffe-1e47-4db5-8bc3-5bebc5feaaf6
timestamp: 2026-07-02T15:50:06Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260702T155006Z_65bf9df5_prompt_io.raw.md
---
## Prompt
From `ai/prompt-io/prompts/issue_473.md`:
> NOTE: you MUST pause this work at 12:50PM EST (BEFORE
> your weekly limit reset) for review by a human!
>
> resolve GH #473: https://github.com/goodboy/tractor/issues/474
> do it with /open-wkt.
(The prompt's issue-number/URL mismatch was resolved in
favor of #473 — file name, session name and prompt text
all say #473; the `/474` URL is a typo.)
GH #473: "Root-cause UDS-on-macOS —
`uds_transport_actor_tree` example fails on CI".
## Response summary
Worked all four #473 next-step checkboxes in wkt
`uds_macos_473`:
1. **stderr surfacing**: `test_docs_examples.py` now
raises with the full subproc stderr (+stdout) on ANY
non-zero exit. The old last-line-'Error' check could
never fire for `tractor` crashes since root-actor
stderr always ends with the strict-EG collapse note
`( ^^^ this exc was collapsed from a group ^^^ )`
proven against the real PR #460 macOS CI log (bare
`assert 1 == 0`, no traceback).
2. **root-cause (linux-provable layer)**: macOS-only
addr corruption in
`MsgpackUDSStream.get_stream_addrs()` — no
`SO_PASSCRED`/autobind on darwin means the accept
side's `getpeername()` is `''`, and the
`(str(), str())` arm took `peername` unconditionally →
`Path('')` garbage addrs on every accepted conn.
Proven + fixed via linux no-autobind simulation.
Possibly not the final macOS crasher (non-fatal on
linux-sim); the diagnostic patch guarantees the next
macOS CI run shows any remaining layer.
3. **CI matrix**: removed the `macos-latest`+`uds`
exclude.
4. **un-skip**: dropped the macOS+CI skip of the example.
Also: `start_listener()` bindspace mkdir hardened
(`parents=True, exist_ok=True`), example docstring
peer-pid mechanism corrected for macOS.
## Files changed
- `tests/test_docs_examples.py` — surface full stderr on
non-zero exit; remove macOS skip of the UDS example
- `tractor/ipc/_uds.py` — fix no-autobind
`get_stream_addrs()` arm; harden bindspace mkdir;
document autobind semantics
- `.github/workflows/ci.yml` — add macOS UDS leg (remove
matrix exclude)
- `examples/uds_transport_actor_tree.py` — platform-
correct peer-pid docs
## Human edits
None yet — pending human review (hard-stop review
checkpoint at 12:50PM EST per prompt); commit staging and
any edits are the human's.

View File

@ -0,0 +1,101 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T15:50:06Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/uds_macos_473
---
# Raw output — GH #473 UDS-on-macOS root-cause session
NOTE: code output is diff-referenced (not copied) per
prompt-io diff-ref mode; all generated code is on branch
`wkt/uds_macos_473` relative to `main` (65bf9df5).
## Diagnostic narrative (verbatim)
Root-cause work proceeded by linux-side simulation since no
macOS host is available:
1. Pulled the actual macOS CI failure log from PR #460 run
28343878933 (job 84115107992): the failure is a bare
`assert 1 == 0` on `proc.returncode` with NO stderr
surfaced — the harness only re-raised when the LAST
stderr line contained 'Error'.
2. Reproduced the stderr-swallow mechanism on linux: any
`tractor` root-actor crash ends stderr with the
strict-EG collapse note
`( ^^^ this exc was collapsed from a group ^^^ )` which
never matches 'Error' — so EVERY possible crash was
swallowed. (Verified by sabotaging the runtime dir via
an over-long `XDG_RUNTIME_DIR` → `OSError: AF_UNIX path
too long` → rc=1 + swallowed.)
3. Found + proved a macOS-only addr-corruption bug in
`MsgpackUDSStream.get_stream_addrs()`: the
`(str(), str())` match-arm unconditionally took
`peername`, but on no-autobind platforms (macOS lacks
linux's `SO_PASSCRED`-triggered autobind) the accept
side's `getpeername()` is `''``Path('')` garbage
laddr/raddr on EVERY accepted UDS conn. Simulated on
linux by nulling `SO_PASSCRED` (no autobind → same `''`
shape): pre-fix the example printed
`listener sock file: .`; post-fix it prints the real
registry sockpath. Non-fatal on linux-sim (rc=0), so
possibly not the final macOS crasher — the diagnostic
patch guarantees the next macOS CI run reveals any
remaining layer.
4. Falsified the missing-parent-dir theory:
`get_rt_dir()` already `mkdir(parents=True,
exist_ok=True)`s at import (and macOS TCP CI passes),
so `~/Library/Caches/TemporaryItems` absence cannot be
the crasher. Hardened `start_listener()`'s bindspace
mkdir anyway (custom `filedir` case + racing actors).
## Generated changes (diff pointers)
> `git diff main..wkt/uds_macos_473 -- tests/test_docs_examples.py`
- always raise with FULL subproc stderr (+stdout) on any
non-zero example exit; keep legacy last-line 'Error'
check for zero-rc cases; drop the macOS+CI skip of
`uds_transport_actor_tree.py` (GH #473 next-step).
> `git diff main..wkt/uds_macos_473 -- tractor/ipc/_uds.py`
- `get_stream_addrs()`: document the autobind semantics
(bytes = linux abstract-ns autobind artifact), add
no-autobind `(str, str)` arm picking the non-empty name
(`peername` connect-side, `sockname` accept-side) with
an empty-pair `ValueError` guard.
- `start_listener()`: `bs.mkdir(parents=True,
exist_ok=True)`.
> `git diff main..wkt/uds_macos_473 -- .github/workflows/ci.yml`
- remove the `macos-latest`+`uds` matrix exclude so
UDS-on-macOS is exercised by CI (GH #473 next-step).
> `git diff main..wkt/uds_macos_473 -- examples/uds_transport_actor_tree.py`
- docs nit: peer-pid mechanism is `SO_PEERCRED` on linux,
`LOCAL_PEERPID` on macOS.
## Verification (verbatim summary)
- macOS-shape sim (no autobind) + fix: example rc=0 with
correct listener sockpath.
- native linux post-fix: example rc=0, autobind arms
unchanged.
- sabotage smoke-test: diagnostic patch surfaces the full
EG traceback incl. collapse-note line.
- `pytest tests/ipc/ tests/test_2way.py tests/discovery/
--tpt-proto uds`: 61 passed, 1 `TooSlowError` flake
(test_simple_rpc) that passes solo in 2.2s — attributed
to a concurrent full-suite run (other session) loading
the host.
- full `tests/test_docs_examples.py` run pending at
raw-file write time (queued behind the concurrent
session's suite).

View File

@ -0,0 +1,83 @@
---
model: claude-fable-5
service: claude
session: b6b42e23-0454-4d9b-be1c-0a89d65aed1b
timestamp: 2026-07-02T15:56:26Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260702T155626Z_65bf9df5_prompt_io.raw.md
---
## Prompt
From `ai/prompt-io/prompts/issue_470.md`:
> attempt to resolve
> https://github.com/goodboy/tractor/issues/470
> do it with /open-wkt, ensure you /run-tests on all
> changes
(plus a hard pause-for-human-review deadline of
12:50PM EST.)
Issue #470 asks to trim the ~0.42s `import tractor`
cost — which dominates per-actor spawn latency on the
`trio` backend — by lazy-importing heavy/optional deps
(`pdbp`, `stackscope`, `multiaddr`, + audit of
`colorlog`/`bidict`/`wrapt`/`setproctitle`).
## Response summary
Profiling showed the issue's dep-list only accounted
for ~20ms; the dominant cost (~244ms) was
`log.get_logger()`'s `get_caller_mod()` calling
`inspect.stack()` at module level in ~39 modules —
each call walks every stack frame (deep during nested
imports) and scans `sys.modules` per frame via
`inspect.getmodule()`.
Changes, in impact order:
1. `get_caller_mod()` -> `sys._getframe()` +
`f_globals['__name__']` `sys.modules` lookup
(~240ms saved).
2. Issue's lazy-import checklist: `bidict`,
`multiaddr`, `colorlog`, `wrapt` moved to
`TYPE_CHECKING`/function-local imports;
`platformdirs` function-local; `asyncio` +
`.to_asyncio` deferred out of the `devx.debug` +
`spawn._entry` eager paths (~15ms saved).
3. PEP 562 `__getattr__` on `tractor/__init__.py`
preserving public `tractor.to_asyncio` attr access.
Results: `import tractor` 0.42s -> ~0.145s (~65%);
sequential `start_actor` latency 0.40-0.44s ->
~0.179s/actor. `pdbp` (needs `_repl.py` class-base
restructure) + `platformdirs` (needs
`UDSAddress.def_bindspace` protocol rework) documented
as follow-ups.
## Files changed
- `tractor/log.py``get_caller_mod()` perf fix +
lazy `colorlog`
- `tractor/__init__.py` — PEP 562 lazy `to_asyncio`
- `tractor/discovery/_addr.py``bidict` ->
`TYPE_CHECKING`
- `tractor/discovery/_multiaddr.py` — lazy `multiaddr`
- `tractor/ipc/_tcp.py`, `tractor/ipc/_uds.py`
`Multiaddr` -> `TYPE_CHECKING`
- `tractor/runtime/_state.py` — lazy `platformdirs`
- `tractor/devx/_frame_stack.py` — lazy `pdbp` +
`wrapt`
- `tractor/devx/debug/_trace.py`,
`tractor/devx/debug/_tty_lock.py` — lazy `asyncio` +
`.to_asyncio`
- `tractor/spawn/_entry.py` — lazy
`run_as_asyncio_guest`
## Human edits
None yet — pending user review at the 12:50PM EST
pause gate (test-suite results reported in-session).

View File

@ -0,0 +1,122 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T15:56:26Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/boot_latency_470
---
# Raw AI output — gh #470 `import tractor` latency trim
All generated code is committed on the
`wkt/boot_latency_470` branch; per diff-ref mode each
file's content is referenced via its diff instead of
copied verbatim.
## Profiling findings (verbatim analysis output)
Baseline: `import tractor` ~0.39-0.42s wall.
`python -X importtime` + `cProfile` traced the cost NOT
primarily to third-party deps (the issue's hypothesis)
but to `tractor/log.py:get_logger()` calling
`get_caller_mod()` -> `inspect.stack()` at module level
in ~39 tractor modules:
- `inspect.stack()` builds `FrameInfo` (incl. src-file
and line-context resolution) for EVERY frame on the
stack; during nested imports the stack is dozens of
importlib frames deep.
- each `FrameInfo` resolution calls
`inspect.getmodule()` which scans all of
`sys.modules` per frame (1.4M `ismodule()` calls in
one profiled import).
- aggregate: ~244ms of tractor-own module "self" time
vs ~20ms for ALL the issue-listed third-party deps
(`pdbp` ~10ms, `bidict` ~4.5ms, `multiaddr` ~3.5ms,
`wrapt`/`colorlog` ~1ms each); `trio` itself is
~70-100ms and unavoidable.
## Generated changes
> `git diff main..wkt/boot_latency_470 -- tractor/log.py`
`get_caller_mod()` rewritten from `inspect.stack()` +
`inspect.getmodule()` to `sys._getframe(frames_up)` +
`frame.f_globals['__name__']` -> `sys.modules` lookup
(O(1) vs O(stack x sys.modules)). Unused `inspect`
imports dropped; `FrameType` imported from `types`.
Also `colorlog` lazy-imported inside
`get_console_log()`.
> `git diff main..wkt/boot_latency_470 -- tractor/discovery/_addr.py`
`bidict` import moved under `TYPE_CHECKING`
(annotation-only use; `_address_types` is a plain dict
literal).
> `git diff main..wkt/boot_latency_470 -- tractor/discovery/_multiaddr.py`
`from __future__ import annotations` added; `multiaddr`
import moved under `TYPE_CHECKING` + function-local
imports in `mk_maddr()`/`parse_maddr()`.
> `git diff main..wkt/boot_latency_470 -- tractor/ipc/_tcp.py tractor/ipc/_uds.py`
`Multiaddr` imports moved under `TYPE_CHECKING`
(annotation-only in both transports).
> `git diff main..wkt/boot_latency_470 -- tractor/runtime/_state.py`
`platformdirs` lazy-imported inside `get_rt_dir()`
(NOTE: still imported eagerly via
`UDSAddress.def_bindspace` class-var eval; see
follow-ups).
> `git diff main..wkt/boot_latency_470 -- tractor/devx/_frame_stack.py`
`pdbp` + `wrapt` lazy-imported inside
`hide_runtime_frames()` / `api_frame()` respectively.
> `git diff main..wkt/boot_latency_470 -- tractor/devx/debug/_trace.py tractor/devx/debug/_tty_lock.py`
`asyncio` moved to `TYPE_CHECKING` + call-site local
imports (`asyncio.current_task()` sites);
`tractor.to_asyncio.run_trio_task_in_future` imports
moved into the infected-aio runtime branches.
> `git diff main..wkt/boot_latency_470 -- tractor/spawn/_entry.py`
`run_as_asyncio_guest` import moved into the
`infect_asyncio=True` branches of `_mp_main()` /
`_trio_main()`.
> `git diff main..wkt/boot_latency_470 -- tractor/__init__.py`
PEP 562 module `__getattr__` added so
`tractor.to_asyncio` attr-access still works (required
by `tests/test_child_manages_service_nursery.py` and
any downstream user) while keeping `asyncio` off the
eager import path.
## Measured results (verbatim)
- `import tractor`: 0.39-0.42s -> ~0.145s (~65% cut)
- `start_actor` spawn+boot+reg+cancel: ~0.40-0.44s ->
~0.179s/actor (n=5 sequential, warm parent)
- post-change eager-module check: only `pdbp` +
`platformdirs` of the issue's list remain eager.
## Known follow-ups (not implemented, deadline-bound)
- `pdbp` (~10ms): still eager via
`devx/debug/_repl.py` class bases
(`class PdbREPL(pdbp.Pdb)`) + `_tty_lock.py`
module-level `@pdbp.hideframe`; needs `_repl`
restructure + PEP 562 in `devx.debug.__init__`.
- `platformdirs` (~1.5ms): eager via
`UDSAddress.def_bindspace: ClassVar = get_rt_dir()`
class-body call; needs `Address`-protocol rework of
`def_bindspace` to a lazy accessor.
- `stackscope` + `setproctitle`: already lazy/absent —
no change needed.

View File

@ -0,0 +1,39 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: pr475-review-fixes-20260817
timestamp: 2026-08-17T23:18:25Z
git_ref: 359fe75c
scope: code
substantive: true
raw_file: 20260817T231825Z_359fe75c_prompt_io.raw.md
---
## Prompt
Continue the `/code-review-changes` pass for PR #475 in its isolated
worktree. Address the seven accepted manual-review findings in
`tractor/ipc/_types.py` and `tractor/ipc/_uds.py`, preserve the existing
Windows capability behavior, verify the result, and prepare the work for
human-controlled commit and review-reply steps. Do not publish replies,
stage, commit, or push without the required explicit authorization.
## Response summary
Restored project quote, docstring, multiline-expression, and
`match/case` conventions while retaining the Windows-safe UDS guard.
Removed unnecessary structural and comment churn, then verified the
focused transport, discovery, and lazy-import paths plus the missing
`AF_UNIX` behavior.
## Files changed
- `tractor/ipc/_types.py` - restore project style and guarded
socket-family dispatch.
- `tractor/ipc/_uds.py` - format the UDS capability gate
consistently.
## Human edits
None - the generated patch remains uncommitted and awaits human
review.

View File

@ -0,0 +1,45 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-17T23:18:25Z
git_ref: 359fe75c
diff_cmd: git diff HEAD~1..HEAD
---
Applied the seven accepted manual-review fixes for PR #475 while
preserving the Windows transport capability behavior.
> `git diff HEAD~1..HEAD -- tractor/ipc/_types.py`
The generated changes restore the project's single-quote docstring and
string conventions, remove the unnecessary helper divider, simplify the
transport-registry comments, and restore `match/case` socket-family
dispatch. The UDS case retains a `HAS_UDS` guard that short-circuits
before `socket.AF_UNIX` is evaluated on unsupported hosts. Nearby error
messages are wrapped without changing their content.
> `git diff HEAD~1..HEAD -- tractor/ipc/_uds.py`
The generated change reformats the `HAS_UDS` conjunction according to
the project's multiline boolean-expression convention and simplifies
the adjacent capability comment.
Verification:
`/home/goodboy/repos/tractor/py313/bin/pytest -q tests/test_lazy_imports.py tests/discovery tests/ipc/test_server.py`
Result: `66 passed, 2 xpassed in 60.62s`.
`ruff check --no-cache --output-format=json tractor/ipc/_types.py tractor/ipc/_uds.py`
Result: no findings.
`git diff --check`
Result: no whitespace errors.
An explicit missing-`AF_UNIX` probe set `HAS_UDS = False`, removed the
socket constant, and exercised an unsupported socket family. It raised
the expected `NotImplementedError` instead of `AttributeError`.
No review replies, commits, or pushes were published.

View File

@ -130,9 +130,10 @@ UDS: same-host, creds included
Pass ``enable_transports=['uds']`` and actors instead talk over
unix-domain sockets, with socket files placed in the per-user
runtime dir (``$XDG_RUNTIME_DIR/tractor/`` on linux, the
``platformdirs`` equivalent elsewhere). Two perks over tcp on a
single host:
runtime dir: ``$XDG_RUNTIME_DIR/tractor/`` on linux, a short
owner-only ``/tmp/tractor-<uid>`` dir on Darwin, and the
``platformdirs`` equivalent elsewhere. Two perks over tcp on a single
host:
- no ports to fight over; addrs are just file paths,
- the kernel snitches on your peer for free: the listening side

View File

@ -44,8 +44,9 @@ clan shares one registry with zero config on your part.
The bootstrap rule inside ``open_root_actor()`` is delightfully
simple:
- on boot, ping every socket addr in ``registry_addrs``; when none
are passed the per-transport defaults are used: for TCP the
- on boot, probe every addr in ``registry_addrs`` with a bounded
Tractor ``Aid`` handshake; when none are passed the per-transport
defaults are used: for TCP the
loopback ``('127.0.0.1', 1616)``, for UDS a
``registry@1616.sock`` file,
@ -53,9 +54,11 @@ simple:
actor and register with the *existing* registry; your own IPC
server binds random same-transport addrs instead,
- if **nothing answers, congratulations: you just became the
registrar**. Your transport server binds the registry addrs
themselves and you start serving lookups for everyone else.
- if every address is absent, congratulations: you just became the
registrar. Your transport server binds the registry addrs
themselves and you start serving lookups for everyone else,
- if no registrar answers but an address is occupied by a foreign or
non-responsive endpoint, startup fails instead of binding over it.
Pass ``ensure_registry=True`` when your program *requires* being
the one-and-only registrar; boot then fails loudly with a
@ -196,9 +199,10 @@ the existing registrar:
trio.run(main)
Per the bootstrap rules above, if the registrar at those addrs is
*not* reachable this process simply becomes its own (registrar)
root — so the same code works standalone and as a tree-joiner.
Per the bootstrap rules above, if those addrs are absent this process
becomes its own registrar root, so the same code works standalone and
as a tree-joiner. An occupied address that does not complete a Tractor
registrar handshake fails startup instead of being rebound.
"Arbiter"? A legacy naming note
-------------------------------

View File

@ -185,8 +185,10 @@ first with a bounded grace window — so actor runtimes can run
their ``trio`` teardown paths — escalating to ``SIGKILL`` only as
a last resort. The ``--shm`` sweep unlinks ``/dev/shm/`` segments
that no live process has open (it leans on psutil_, already in
your dev venv, to check live mappings and fds) and ``--uds``
clears socket files whose binder pid is dead.
your dev venv, to check live mappings and fds) and ``--uds`` clears
dead-binder sockets from Tractor's platform-specific runtime dir. It
also unconditionally removes ``registry@1616.sock``; do not run the UDS
sweep while a live registrar is serving from that default address.
Testing your own ``tractor`` app
--------------------------------

View File

@ -23,18 +23,10 @@ async def endpoint(
await trio.sleep_forever()
async def spawn_and_open_ep(
an: tractor.ActorNursery,
async def open_ep(
ptl: tractor.Portal,
i: int,
) -> None:
'''
Spawn a subactor, start a remote `endpoint()`-task in it.
'''
ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
ctx: tractor.Context
async with ptl.open_context(endpoint) as (
ctx,
@ -47,7 +39,33 @@ async def spawn_and_open_ep(
await ctx.wait_for_result()
async def main():
async def spawn_and_open_ep(
an: tractor.ActorNursery,
i: int,
maybe_ptl: tractor.Portal|None = None,
) -> None:
'''
Spawn a subactor, start a remote `endpoint()`-task in it.
'''
if maybe_ptl is None:
maybe_ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
await open_ep(
ptl=maybe_ptl,
i=i,
)
async def main(
# spawn subs concurrently (in bg `trio.Task`s) so each
# actor's cold `import tractor` (~0.4s, see #470) overlaps
# instead of stacking; once forkserver (#463) lands, spawn
# is cheap enough to just loop sequentially.
spawn_subs_in_bg_tasks: bool = True,
):
'''
Spawn a subactor-per-CPU then self-destruct the cluster.
@ -60,17 +78,21 @@ async def main():
# https://github.com/goodboy/tractor/pull/463
# start_method='main_thread_forkserver',
) as an,
# spawn subs concurrently (in bg `trio.Task`s) so each
# actor's cold `import tractor` (~0.4s, see #470) overlaps
# instead of stacking; once forkserver (#463) lands, spawn
# is cheap enough to just loop sequentially.
trio.open_nursery() as tn,
):
for i in range(cpu_count()):
maybe_ptl: tractor.Portal|None = None
if not spawn_subs_in_bg_tasks:
maybe_ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
tn.start_soon(
spawn_and_open_ep,
an,
i,
maybe_ptl,
)
destruct_in: int = 2
print(

View File

@ -6,7 +6,8 @@ subactor inherits the preference.
Every channel address is a filesystem socket path (no TCP port
in sight!) and, as a kernel-provided bonus, the peer's pid is
exchanged for free via `SO_PEERCRED`.
exchanged for free via `SO_PEERCRED` on linux,
`LOCAL_PEERPID` on macOS.
'''
import os
@ -42,7 +43,7 @@ async def main() -> None:
# (named for the root registrar) this channel rode in
# on, NOT a per-child path; the child-specific identity
# we get for free is the kernel-reported peer pid (via
# `SO_PEERCRED`).
# `SO_PEERCRED` on linux, `LOCAL_PEERPID` on macOS).
print(
f'portal chan tpt proto: {raddr.proto_key!r}\n'
f'listener sock file: {raddr.sockpath}\n'

View File

@ -0,0 +1,4 @@
Fix Unix-domain-socket actor trees and registrar discovery on macOS.
Runtime sockets now use a short, owner-only runtime directory,
generated socket names remain within platform limits, and transient
or reset pre-handshake connections no longer destabilize discovery.

View File

@ -23,10 +23,10 @@ Two cleanup phases (run in order when both are enabled):
hard-crashing actor leaves leaked segments that
nothing else GCs.
3. **UDS sweep** (`--uds` / `--uds-only`) — unlinks
`${XDG_RUNTIME_DIR}/tractor/<name>@<pid>.sock` files
whose binder pid is dead (or the `1616` registry
sentinel). Needed because the IPC server's
3. **UDS sweep** (`--uds` / `--uds-only`) — unlinks socket
files from Tractor's platform-specific default bindspace whose
binder pid is dead (or the `1616` registry sentinel). Needed
because the IPC server's
`os.unlink()` cleanup lives in a `finally:` block
that doesn't always run on hard exits (SIGKILL,
escaped `KeyboardInterrupt`, etc.) — see issue #452.
@ -137,8 +137,8 @@ def main() -> int:
action='store_true',
help=(
'after process reap, also unlink orphaned '
'${XDG_RUNTIME_DIR}/tractor/*.sock files '
'whose binder pid is dead (or the 1616 '
'sockets from Tractor\'s platform default '
'bindspace whose binder pid is dead (or the 1616 '
'registry sentinel). See issue #452.'
),
)
@ -212,7 +212,9 @@ def main() -> int:
# --- phase 3: UDS sweep (opt-in) ---
if args.uds or args.uds_only:
leaked_uds: list[str] = find_orphaned_uds()
leaked_uds: list[str] = find_orphaned_uds(
include_registry_sentinel=True,
)
if not leaked_uds:
print(
'[tractor-reap] no orphaned UDS sock-files '

View File

@ -1300,18 +1300,30 @@ def test_ctxep_pauses_n_maybe_ipc_breaks(
if _non_linux:
tpt: str = 'TCP'
assert_before(
before: str = assert_before(
child,
['peer IPC channel closed abruptly?',
'another task closed this fd',
'Debug lock request was CANCELLED?',
f"'Msgpack{tpt}Stream' was already closed locally?",
f"TransportClosed: 'Msgpack{tpt}Stream' was already closed 'by peer'?",
]
]
# XXX races on whether these show/hit?
# 'Failed to REPl via `_pause()` You called `tractor.pause()` from an already cancelled scope!',
# 'AssertionError',
# XXX races on whether these show/hit?
# 'Failed to REPl via `_pause()` You called `tractor.pause()` from an already cancelled scope!',
# 'AssertionError',
)
# Error shipment and peer receive race after local close.
# Either diagnostic proves the transport was torn down.
closed_locally: str = (
f"'Msgpack{tpt}Stream' was already closed locally?"
)
closed_by_peer: str = (
f"TransportClosed: 'Msgpack{tpt}Stream' was "
f"already closed 'by peer'?"
)
assert (
closed_locally in before
or closed_by_peer in before
)
# OSc(ancel) the hanging tree
do_ctlc(

View File

@ -1,21 +1,18 @@
'''
Discovery-suite fixtures, including the `daemon`
remote-registrar subprocess used by the multi-program
discovery tests.
Discovery-suite fixtures, including the `daemon` remote-registrar
subprocess used by the multi-program discovery tests.
Lives here (vs. the parent `tests/conftest.py`)
because `daemon` is a discovery-protocol primitive
boots a separate `tractor.run_daemon()` process whose
sole purpose is to serve as a registrar peer for
because `daemon` is a discovery-protocol primitive: it boots a child
that enters `open_root_actor()` and waits as a registrar peer for
discovery-roundtrip tests. Pytest fixtures inherit
DOWNWARD through conftest hierarchy, so anything
under `tests/discovery/` automatically picks this up.
'''
from __future__ import annotations
import os
from pathlib import Path
import platform
import socket
import subprocess
import sys
import time
@ -31,33 +28,27 @@ from ..conftest import (
def _wait_for_daemon_ready(
reg_addr: tuple,
tpt_proto: str,
ready_path: Path,
*,
deadline: float = 10.0,
poll_interval: float = 0.05,
proc: subprocess.Popen|None = None,
) -> None:
'''
Active-poll the daemon's bind address until it
accepts a connection (proving it has called
`bind() + listen()` and is ready to handle IPC).
Poll until the daemon reports completed actor startup.
Replaces the historical blind `time.sleep()` in the
`daemon` fixture which was racy under load see
`ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`.
Uses stdlib `socket` directly (no trio runtime
bootstrap cost) sufficient because
`tractor.run_daemon()` doesn't return from
bootstrap until the runtime is fully ready to
accept IPC.
The child writes `ready_path` only after entering
`open_root_actor()`, which guarantees all transport listeners are
serving without requiring a raw connection probe.
Raises `TimeoutError` on `deadline` exceeded. If
`proc` is given, ALSO raises early if the daemon
process exits non-zero before the deadline (catches
daemon-startup-crash that the blind sleep used to
silently mask).
process exits before the deadline (catches a daemon startup crash
that the blind sleep used to silently mask).
'''
end: float = time.monotonic() + deadline
@ -70,43 +61,25 @@ def _wait_for_daemon_ready(
if proc is not None and proc.poll() is not None:
raise RuntimeError(
f'Daemon proc exited (rc={proc.returncode}) '
f'before becoming ready to accept on '
f'{reg_addr!r}'
f'before reporting ready at {ready_path!r}'
)
try:
if tpt_proto == 'tcp':
# `socket.create_connection` does the
# `socket() + connect()` dance with a
# builtin timeout — perfect primitive
# for a one-shot probe.
with socket.create_connection(
reg_addr,
timeout=poll_interval,
):
return
else:
# UDS — `reg_addr` is a `(filedir, sockname)`
# tuple per `tractor.ipc._uds.UDSAddress.unwrap`.
sockpath: str = os.path.join(*reg_addr)
sock = socket.socket(socket.AF_UNIX)
try:
sock.settimeout(poll_interval)
sock.connect(sockpath)
return
finally:
sock.close()
if ready_path.is_file():
if proc is not None and proc.poll() is not None:
raise RuntimeError(
f'Daemon proc exited (rc={proc.returncode}) '
f'after reporting ready at {ready_path!r}'
)
return
except (
ConnectionRefusedError,
FileNotFoundError,
OSError,
socket.timeout,
) as exc:
last_exc = exc
time.sleep(poll_interval)
time.sleep(poll_interval)
raise TimeoutError(
f'Daemon never accepted on {reg_addr!r} within '
f'{deadline}s (last connect-attempt exc: '
f'{last_exc!r})'
f'Daemon never reported ready at {ready_path!r} within '
f'{deadline}s (last sentinel-state exc: {last_exc!r})'
)
@ -136,18 +109,27 @@ def daemon(
)
loglevel: str = 'info'
ready_path: Path = (
Path(str(testdir.tmpdir))
/ 'daemon-ready'
)
ready_path.unlink(missing_ok=True)
code: str = (
"import tractor; "
"tractor.run_daemon([], "
"registry_addrs={reg_addrs}, "
"enable_transports={enable_tpts}, "
"debug_mode={debug_mode}, "
"loglevel={ll})"
).format(
reg_addrs=str([reg_addr]),
enable_tpts=str([tpt_proto]),
ll="'{}'".format(loglevel) if loglevel else None,
debug_mode=debug_mode,
f'from pathlib import Path\n'
f'import tractor\n'
f'import trio\n'
f'\n'
f'async def main():\n'
f' async with tractor.open_root_actor(\n'
f' registry_addrs={[reg_addr]!r},\n'
f' enable_transports={[tpt_proto]!r},\n'
f' debug_mode={debug_mode!r},\n'
f' loglevel={loglevel!r},\n'
f' ):\n'
f' Path({str(ready_path)!r}).touch()\n'
f' await trio.sleep_forever()\n'
f'\n'
f'trio.run(main)\n'
)
cmd: list[str] = [
sys.executable,
@ -163,9 +145,9 @@ def daemon(
**kwargs,
)
# Active-poll the daemon's bind address until it's
# ready to accept connections — replaces the legacy
# blind `time.sleep(2.2)` which was racy under load
# Poll the child's ready sentinel, published after actor startup,
# instead of connecting to its transport socket. This replaces
# the legacy blind `time.sleep(2.2)` which was racy under load
# (see
# `ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`).
#
@ -176,48 +158,49 @@ def daemon(
15.0 if (_non_linux and ci_env)
else 10.0
)
_wait_for_daemon_ready(
reg_addr=reg_addr,
tpt_proto=tpt_proto,
deadline=deadline,
proc=proc,
)
assert not proc.returncode
yield proc
sig_prog(proc, _INT_SIGNAL)
# XXX! yeah.. just be reaaal careful with this bc
# sometimes it can lock up on the `_io.BufferedReader`
# and hang..
#
# NB, drain happens at TEARDOWN (post-yield), so the
# test body has its chance to read `proc.stderr`
# FIRST. Reading here AFTER would silently swallow
# the daemon's stderr output and break tests that
# assert on it (e.g. `test_abort_on_sigint`).
stderr: str = proc.stderr.read().decode()
stdout: str = proc.stdout.read().decode()
if (
stderr
or
stdout
):
print(
f'Daemon actor tree produced output:\n'
f'{proc.args}\n'
f'\n'
f'stderr: {stderr!r}\n'
f'stdout: {stdout!r}\n'
try:
_wait_for_daemon_ready(
ready_path=ready_path,
deadline=deadline,
proc=proc,
)
if (rc := proc.returncode) != -2:
msg: str = (
f'Daemon actor tree was not cancelled !?\n'
f'proc.args: {proc.args!r}\n'
f'proc.returncode: {rc!r}\n'
)
if rc < 0:
raise RuntimeError(msg)
assert not proc.returncode
yield proc
finally:
if proc.poll() is None:
sig_prog(proc, _INT_SIGNAL)
test_log.error(msg)
# NOTE: these blocking reads can hang when descendants retain
# inherited pipe descriptors. Keep teardown signaling above
# them and avoid adding subprocesses outside the actor tree.
#
# NB, drain happens at TEARDOWN (post-yield), so the
# test body has its chance to read `proc.stderr`
# FIRST. Reading here AFTER would silently swallow
# the daemon's stderr output and break tests that
# assert on it (e.g. `test_abort_on_sigint`).
stderr: str = proc.stderr.read().decode()
stdout: str = proc.stdout.read().decode()
if (
stderr
or
stdout
):
print(
f'Daemon actor tree produced output:\n'
f'{proc.args}\n'
f'\n'
f'stderr: {stderr!r}\n'
f'stdout: {stdout!r}\n'
)
if (rc := proc.returncode) != -2:
msg: str = (
f'Daemon actor tree was not cancelled !?\n'
f'proc.args: {proc.args!r}\n'
f'proc.returncode: {rc!r}\n'
)
if rc < 0:
raise RuntimeError(msg)
test_log.error(msg)

View File

@ -0,0 +1,76 @@
'''
Discovery daemon fixture regressions.
This module imports private helpers from the sibling
`tests.discovery.conftest` plugin to exercise that fixture machinery
directly, rather than testing a production `tractor` API.
'''
from unittest.mock import (
call,
Mock,
)
from .conftest import _wait_for_daemon_ready
def test_daemon_ready_check_does_not_connect(
monkeypatch,
tmp_path,
):
'''
Observe completed daemon startup without a raw connection.
The old UDS readiness helper connected and immediately closed. That
entered Tractor's actor-handshake handler with no `Aid` payload and
destabilized the remote registrar on macOS before discovery tests
started. This test creates the child sentinel, forbids all socket
construction and connection helpers, then proves readiness returns
without touching the transport layer.
'''
ready_path = tmp_path / 'daemon-ready'
ready_path.touch()
socket_ctor = Mock(side_effect=AssertionError('socket opened'))
connect = Mock(side_effect=AssertionError('socket connected'))
monkeypatch.setattr('socket.socket', socket_ctor)
monkeypatch.setattr('socket.create_connection', connect)
_wait_for_daemon_ready(
ready_path=ready_path,
deadline=.1,
poll_interval=.01,
)
socket_ctor.assert_not_called()
connect.assert_not_called()
def test_daemon_ready_check_backs_off(monkeypatch):
'''
Back off while waiting for the child startup sentinel.
The sentinel may appear after several parent polling intervals.
A deterministic false/false/true path sequence proves the helper
sleeps between unsuccessful observations instead of hot-spinning
and starving a booting daemon on constrained CI workers.
'''
ready_path = Mock()
ready_path.is_file.side_effect = [False, False, True]
sleep = Mock()
monotonic = Mock(side_effect=[0, 0, 0, 0])
monkeypatch.setattr('time.sleep', sleep)
monkeypatch.setattr('time.monotonic', monotonic)
_wait_for_daemon_ready(
ready_path=ready_path,
deadline=.2,
poll_interval=.01,
)
assert ready_path.is_file.call_count == 3
assert sleep.call_args_list == [
call(.01),
call(.01),
]

View File

@ -2,23 +2,211 @@
`open_root_actor(tpt_bind_addrs=...)` test suite.
Verify all three runtime code paths for explicit IPC-server
bind-address selection in `_root.py`:
bind-address selection in `_root.py` and registry probing in
`discovery._api`:
1. Non-registrar, no explicit bind -> random addrs from registry proto
2. Registrar, no explicit bind -> binds to registry_addrs
3. Explicit bind given -> wraps via `wrap_address()` and uses them
'''
from contextlib import asynccontextmanager as acm
from unittest.mock import (
AsyncMock,
call,
Mock,
)
import pytest
import trio
import tractor
from tractor.discovery import _api
from tractor.discovery._addr import (
wrap_address,
)
from tractor.discovery._multiaddr import mk_maddr
from tractor.ipc import _connect_chan
from tractor._testing.addr import get_rando_addr
def test_registry_probe_retries_transient_handshake(
monkeypatch: pytest.MonkeyPatch,
):
'''
Retry a connected registrar after transient handshake timeout.
Loaded macOS runners can accept the transport while delaying the
actor handshake beyond one second. Treating that first timeout as
final makes a healthy remote daemon look occupied and cascades into
discovery failures. This deterministic fake fails once, succeeds
on the second complete handshake, and proves one bounded backoff.
'''
async def stall_handshake(**kwargs):
await trio.sleep_forever()
first_handshake = AsyncMock(side_effect=stall_handshake)
second_handshake = AsyncMock(
return_value=tractor.msg.Aid(
name='registrar',
uuid='registrar-uuid',
pid=1234,
is_registrar=True,
),
)
chans = [
Mock(_do_handshake=first_handshake),
Mock(_do_handshake=second_handshake),
]
closed: list[object] = []
@acm
async def connect_chan(addr, close_timeout):
assert close_timeout == .2
chan = chans[len(closed)]
try:
yield chan
finally:
closed.append(chan)
sleep = AsyncMock()
monkeypatch.setattr(_api, '_connect_chan', connect_chan)
monkeypatch.setattr(_api.trio, 'sleep', sleep)
async def main():
status = await _api._probe_registry(
addr=wrap_address(('127.0.0.1', 1616)),
timeout=.3,
attempt_timeout=.1,
max_attempts=3,
retry_delay=.01,
)
assert status == 'registrar'
trio.run(main)
first_handshake.assert_awaited_once()
second_handshake.assert_awaited_once()
assert first_handshake.await_args.kwargs['timeout'] == .1
assert second_handshake.await_args.kwargs['timeout'] == .1
assert closed == chans
sleep.assert_has_awaits([call(.01)])
def test_probe_channel_close_is_bounded(
monkeypatch: pytest.MonkeyPatch,
):
'''
Bound shielded channel cleanup after a registry probe.
`_connect_chan()` shields `.aclose()` so cancellation cannot leak
ordinary channels. A stalled close previously let registry probing
exceed every connect and handshake deadline. This fake close never
completes; the explicit cleanup allowance must still return control
to the caller without cancelling its surrounding task.
'''
chan = Mock()
chan.aclose = AsyncMock(side_effect=trio.sleep_forever)
monkeypatch.setattr(
tractor.Channel,
'from_addr',
AsyncMock(return_value=chan),
)
async def main():
with trio.fail_after(.5):
async with _connect_chan(
('127.0.0.1', 1616),
close_timeout=.01,
):
pass
trio.run(main)
chan.aclose.assert_awaited_once()
def test_transport_only_listener_is_not_registrar():
'''
Require a Tractor handshake before accepting a registry address.
The old election probe marked an address live after transport
connect alone. A non-Tractor listener, or a registrar still
failing its initial handshake, was therefore selected as the
remote registry. This test accepts the probe and closes it without
replying, then proves `open_root_actor()` rejects that occupied
endpoint instead of selecting it or binding over it.
'''
async def transport_only_handler(
stream: trio.SocketStream,
) -> None:
await stream.aclose()
async def main():
listeners = await trio.open_tcp_listeners(0)
listener = listeners[0]
sockname = listener.socket.getsockname()
reg_addr: tuple[str, int] = (
sockname[0],
sockname[1],
)
async with trio.open_nursery() as tn:
tn.start_soon(
trio.serve_listeners,
transport_only_handler,
listeners,
)
with pytest.raises(
RuntimeError,
match='occupied but did not answer',
):
async with tractor.open_root_actor(
registry_addrs=[reg_addr],
enable_transports=['tcp'],
):
pytest.fail('foreign listener selected as registrar')
tn.cancel_scope.cancel()
trio.run(main)
def test_registry_probe_preserves_no_peers_state(
reg_addr: tuple,
tpt_proto: str,
):
'''
Keep an idle registrar peer-free after an election probe.
Probe handshakes exchange registrar capability but must not enter
`IPCServer._peers`. Resetting `_no_more_peers` before identifying a
probe left an idle registrar reporting phantom peers and delayed
shutdown. This test probes the live local registrar and proves its
peer map and no-peers event remain unchanged afterward.
'''
async def main():
async with tractor.open_root_actor(
registry_addrs=[reg_addr],
enable_transports=[tpt_proto],
):
actor = tractor.current_actor()
server = actor.ipc_server
probe_status = await _api._probe_registry(
addr=wrap_address(reg_addr),
)
assert probe_status == 'registrar'
await trio.sleep(0)
assert not server._peers
assert server._no_more_peers.is_set()
trio.run(main)
# ------------------------------------------------------------------
# helpers
# ------------------------------------------------------------------

View File

@ -3,14 +3,21 @@ Unit-ish tests for specific IPC transport protocol backends.
'''
from __future__ import annotations
import os
from pathlib import Path
import socket
import stat
import sys
import tempfile
from types import SimpleNamespace
from unittest.mock import Mock
import pytest
import trio
import tractor
from tractor import Actor
from tractor.runtime import _state
from tractor.discovery import _addr
from tractor.runtime import _state
@pytest.fixture
@ -31,6 +38,381 @@ def bindspace_dir_str() -> str:
bs_dir.rmdir()
def test_macos_rt_dir_fits_uds_path_limit(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Keep the default Darwin UDS bindpath below its 104-byte limit.
`platformdirs` normally places the runtime directory below the
long `~/Library/Caches/TemporaryItems` path. Pytest also assigns
a deeply nested temporary home, so appending a registry socket
name made every macOS UDS listener fail with `AF_UNIX path too
long`. This test simulates Darwin and an intentionally long
platformdirs result, then proves `get_rt_dir()` uses the short
system temporary directory and leaves room for the socket name.
'''
long_rt_dir: Path = tmp_path / ('long' * 30)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(long_rt_dir / appname),
)
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
rt_dir: Path = _state.get_rt_dir()
sockpath: Path = (
Path('/tmp')
/ f'tractor-{os.getuid()}'
/ 'registry@1616.sock'
)
assert rt_dir == tmp_path / f'tractor-{os.getuid()}'
assert len(os.fsencode(sockpath)) < 104
assert stat.S_IMODE(rt_dir.stat().st_mode) == 0o700
def test_macos_rt_dir_rejects_symlink(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject a pre-created symlink at the Darwin runtime path.
Darwin uses the predictable `/tmp/tractor-<uid>` path to stay
below its `AF_UNIX` limit. A hostile local user could otherwise
point that path at a victim-owned directory and make
`get_rt_dir()` chmod or place sockets in the symlink target. The
test replaces `/tmp` with a controlled directory, installs the
malicious link, and proves non-following validation rejects it.
'''
runtime_link: Path = tmp_path / f'tractor-{os.getuid()}'
target_dir: Path = tmp_path / 'target'
target_dir.mkdir(mode=0o755)
runtime_link.symlink_to(target_dir, target_is_directory=True)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
with pytest.raises(PermissionError, match='Unsafe Darwin'):
_state.get_rt_dir()
assert stat.S_IMODE(target_dir.stat().st_mode) == 0o755
def test_reaper_uses_default_uds_bindspace(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Sweep the same platform-specific bindspace used by UDS actors.
The reaper previously consulted only `XDG_RUNTIME_DIR`, missing
Darwin sockets after the runtime moved to `/tmp/tractor-<uid>`.
This test replaces `UDSAddress.def_bindspace` and proves the test
harness resolves that shared transport default directly.
'''
from tractor._testing import _reap
from tractor.ipc._uds import UDSAddress
monkeypatch.setattr(
UDSAddress,
'def_bindspace',
tmp_path,
)
assert _reap.get_uds_dir() == str(tmp_path)
def test_automatic_reaper_preserves_registry_sentinel(
monkeypatch: pytest.MonkeyPatch,
):
'''
Reserve unconditional registry cleanup for the explicit CLI.
The `registry@1616.sock` suffix does not encode its binder PID, so
automatic pytest cleanup cannot distinguish a leak from another
live registrar. This test creates registry and actor sockets,
proves the default sweep selects only the dead actor, then proves
explicit sentinel inclusion retains the CLI's documented behavior.
'''
from tractor._testing import _reap
with tempfile.TemporaryDirectory(
prefix='tractor-reap-',
dir='/tmp',
) as tmpdir:
bindspace: Path = Path(tmpdir)
registry_path: Path = bindspace / 'registry@1616.sock'
actor_path: Path = bindspace / 'worker@1234.sock'
socks: list[socket.socket] = []
for path in (registry_path, actor_path):
sock = socket.socket(socket.AF_UNIX)
sock.bind(str(path))
socks.append(sock)
monkeypatch.setattr(_reap, '_is_alive', lambda pid: False)
try:
assert _reap.find_orphaned_uds(
uds_dir=str(bindspace),
) == [str(actor_path)]
assert set(
_reap.find_orphaned_uds(
uds_dir=str(bindspace),
include_registry_sentinel=True,
)
) == {
str(registry_path),
str(actor_path),
}
finally:
for sock in socks:
sock.close()
def test_rt_dir_rejects_non_directory(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Preserve the non-Darwin runtime-directory type contract.
Replacing `Path.is_dir()` with unguarded `lstat()` briefly made
existing files look like valid runtime directories on Linux.
This test points `platformdirs` at a regular file and proves
`get_rt_dir()` rejects it during initialization.
'''
rt_file: Path = tmp_path / 'runtime-file'
rt_file.touch()
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_file),
)
with pytest.raises(
PermissionError,
match='Unsafe POSIX',
):
_state.get_rt_dir()
new_rt_dir: Path = tmp_path / 'new-runtime-dir'
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(new_rt_dir),
)
assert _state.get_rt_dir() == new_rt_dir
assert stat.S_IMODE(new_rt_dir.stat().st_mode) == 0o700
def test_linux_rt_dir_secures_existing_path(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Enforce owner-only access on an existing Linux runtime directory.
Linux previously accepted any existing directory returned by
`platformdirs`, without checking ownership or correcting a
traversable mode. This test creates an owner-controlled `0o755`
directory and proves `get_rt_dir()` normalizes the managed
bindspace to `0o700` before returning it.
'''
rt_dir: Path = tmp_path / 'tractor'
rt_dir.mkdir(mode=0o755)
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_dir),
)
assert _state.get_rt_dir() == rt_dir
assert stat.S_IMODE(rt_dir.stat().st_mode) == 0o700
def test_linux_rt_dir_rejects_foreign_owner(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject an existing Linux runtime directory owned by another UID.
A pre-created bindspace must never be made private with `chmod`
until ownership is verified. This test makes the current process
appear to have a different UID and proves `get_rt_dir()` rejects
the directory without changing its original mode.
'''
rt_dir: Path = tmp_path / 'tractor'
rt_dir.mkdir(mode=0o755)
original_mode: int = stat.S_IMODE(rt_dir.stat().st_mode)
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_dir),
)
monkeypatch.setattr(
os,
'getuid',
lambda: rt_dir.stat().st_uid + 1,
)
with pytest.raises(
PermissionError,
match='Unsafe POSIX',
):
_state.get_rt_dir()
assert stat.S_IMODE(rt_dir.stat().st_mode) == original_mode
def test_macos_rt_dir_rejects_intermediate_symlink(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject symlinks in nested Darwin runtime subdirectories.
The earlier final-component check allowed `link/child` to follow
an intermediate symlink and create `child` outside the secured
runtime root. This test installs that link and proves traversal
stops before anything is created in its target.
'''
rt_root: Path = tmp_path / f'tractor-{os.getuid()}'
target_dir: Path = tmp_path / 'target'
rt_root.mkdir(mode=0o700)
target_dir.mkdir()
(rt_root / 'link').symlink_to(
target_dir,
target_is_directory=True,
)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
with pytest.raises(PermissionError, match='Unsafe Darwin'):
_state.get_rt_dir(subdir='link/child')
assert not (target_dir / 'child').exists()
@pytest.mark.parametrize(
('platform_name', 'path_limit'),
[
('darwin', 104),
('linux', 108),
],
)
def test_uds_sockname_compaction(
monkeypatch: pytest.MonkeyPatch,
platform_name: str,
path_limit: int,
):
'''
Keep generated actor sockets safe and below Darwin's byte limit.
Actor names are unrestricted identity strings. A long, multibyte,
or path-like name previously produced overlong or escaping socket
paths. These cases prove `UDSAddress.get_sockname()` preserves a
short legacy name, deterministically compacts unsafe names, keeps
the reaper's `@pid.sock` suffix, and stays within Darwin's byte
limit.
'''
from tractor.ipc._uds import UDSAddress
bindspace: Path = Path('/tmp/tractor-501')
pid: int = 12345
from tractor.ipc import _uds
monkeypatch.setattr(sys, 'platform', platform_name)
monkeypatch.setattr(_uds, '_SUN_PATH_LIMIT', path_limit)
short: Path = UDSAddress.get_sockname(
name='worker',
pid=pid,
bindspace=bindspace,
)
long_name: str = 'actor-' + ('\u00e9' * 100)
compact: Path = UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=bindspace,
)
unsafe: Path = UDSAddress.get_sockname(
name='../worker',
pid=pid,
bindspace=bindspace,
)
assert short == Path(f'worker@{pid}.sock')
assert compact == UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=bindspace,
)
assert compact.name.endswith(f'@{pid}.sock')
assert unsafe.parent == Path('.')
assert '..' not in unsafe.name
assert len(os.fsencode(bindspace / compact)) < path_limit
with pytest.raises(ValueError) as exc_info:
UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=Path('/tmp') / ('x' * 90),
)
errmsg: str = str(exc_info.value)
assert 'leaves no room' in errmsg
assert 'name was unsafe: False' in errmsg
assert 'name was over budget: True' in errmsg
assert f'AF_UNIX path limit: {path_limit}' in errmsg
def test_uds_reaper_ignores_unreconstructable_path(
monkeypatch: pytest.MonkeyPatch,
):
'''
Keep post-kill UDS cleanup best-effort on path overflow.
`unlink_uds_bind_addrs()` reconstructs a self-assigned socket from
the dead actor's name and PID. An over-budget bindspace makes that
naming helper raise before `os.unlink()`; propagating the error
would replace the original supervision outcome after the child was
already killed. This test forces overflow and proves cleanup skips
reconstruction without attempting an unlink or raising.
'''
from tractor.ipc import _uds
from tractor.spawn import _reap
long_bindspace: Path = Path('/tmp') / ('x' * 120)
proc = SimpleNamespace(pid=12345)
subactor = SimpleNamespace(
aid=SimpleNamespace(name='worker'),
)
unlink = Mock()
monkeypatch.setattr(
_uds.UDSAddress,
'def_bindspace',
long_bindspace,
)
monkeypatch.setattr(_reap.os, 'unlink', unlink)
_reap.unlink_uds_bind_addrs(
proc=proc,
subactor=subactor,
)
unlink.assert_not_called()
def test_uds_bindspace_created_implicitly(
debug_mode: bool,
bindspace_dir_str: str,

View File

@ -3,7 +3,13 @@ High-level `.ipc._server` unit tests.
'''
from __future__ import annotations
import errno
from unittest.mock import (
AsyncMock,
Mock,
)
import msgspec
import pytest
import trio
from tractor import (
@ -14,6 +20,11 @@ from tractor import (
from tractor._testing.addr import (
get_rando_addr,
)
from tractor._exceptions import TransportClosed
from tractor.ipc._chan import Channel
from tractor.ipc import _server
from tractor.ipc._transport import MsgpackTransport
from tractor.msg.types import Aid
# TODO, use/check-roundtripping with some of these wrapper types?
#
# from .._addr import Address
@ -23,6 +34,165 @@ from tractor._testing.addr import (
# from ._tcp import TCPAddress
def test_send_normalizes_only_grouped_peer_resets():
'''
Normalize only all-peer-close grouped transport failures.
A UDS peer may disconnect before completing the actor handshake.
Darwin can report the server's first handshake write as
`ECONNRESET`, wrapped by `trio.BrokenResourceError` and potentially
nested in an `ExceptionGroup`. This fake stream first groups reset
and broken-pipe branches, proving `.send()` normalizes a complete
peer-close tree to `TransportClosed`. It then groups a reset with
an unrelated `ValueError`, proving the mixed failure remains a
`trio.BrokenResourceError` instead of hiding the application error.
'''
def broken_resource(err_no: int) -> trio.BrokenResourceError:
try:
raise OSError(
err_no,
'Peer closed',
)
except OSError as peer_err:
try:
raise trio.BrokenResourceError from peer_err
except trio.BrokenResourceError as broken_err:
return broken_err
class GroupedFailureStream:
def __init__(self, exceptions: list[Exception]) -> None:
self.exceptions = exceptions
async def send_all(self, data: bytes) -> None:
grouped_err = ExceptionGroup(
'concurrent send failures',
self.exceptions,
)
raise trio.BrokenResourceError from grouped_err
async def main():
transport = object.__new__(MsgpackTransport)
transport.stream = GroupedFailureStream([
broken_resource(errno.ECONNRESET),
broken_resource(errno.EPIPE),
])
transport._send_lock = trio.StrictFIFOLock()
transport._laddr = 'local'
transport._raddr = 'remote'
transport._task = trio.lowlevel.current_task()
with pytest.raises(TransportClosed) as exc_info:
await transport.send(
{'probe': True},
strict_types=False,
)
grouped_err = exc_info.value.src_exc.__cause__
assert isinstance(grouped_err, ExceptionGroup)
assert len(grouped_err.exceptions) == 2
transport.stream = GroupedFailureStream([
ValueError('unrelated failure'),
broken_resource(errno.ECONNRESET),
])
with pytest.raises(trio.BrokenResourceError) as exc_info:
await transport.send(
{'probe': True},
strict_types=False,
)
grouped_err = exc_info.value.__cause__
assert isinstance(grouped_err, ExceptionGroup)
assert isinstance(grouped_err.exceptions[0], ValueError)
trio.run(main)
def test_handshake_normalizes_decode_error():
'''
Keep malformed pre-handshake frames out of the service nursery.
A non-msgpack peer can trigger `msgspec.DecodeError` before a
remote `Aid` exists. Letting that decoder error escape the inbound
handler cancels the actor's shared IPC nursery. This fake channel
proves `_do_handshake()` presents only `TransportClosed` upward.
'''
chan = object.__new__(Channel)
chan.send = AsyncMock()
chan.recv = AsyncMock(
side_effect=msgspec.DecodeError('malformed handshake'),
)
async def main():
with pytest.raises(TransportClosed) as exc_info:
await chan._do_handshake(
aid=Aid(
name='local',
uuid='local-uuid',
pid=1234,
),
timeout=.1,
)
assert isinstance(
exc_info.value.src_exc,
msgspec.DecodeError,
)
trio.run(main)
def test_server_uses_independent_handshake_timeout(
monkeypatch: pytest.MonkeyPatch,
):
'''
Give ordinary actor handshakes a distinct, generous deadline.
Registry probes use short retries, but ordinary portal and child
connections do not retry. Applying the probe's one-second timeout
in the server can terminate a valid delayed child and leave its
parent blocked in `IPCServer.wait_for_peer()`. This handler fake
proves the server uses its separate pre-registration budget.
'''
handshake = AsyncMock(
side_effect=TransportClosed(message='stop after assertion'),
)
chan = Mock(_do_handshake=handshake)
actor = Mock(
aid=Aid(
name='local',
uuid='local-uuid',
pid=1234,
),
)
monkeypatch.setattr(
Channel,
'from_stream',
Mock(return_value=chan),
)
monkeypatch.setattr(
_server._state,
'current_actor',
Mock(return_value=actor),
)
async def main():
await _server.handle_stream_from_peer(
stream=Mock(),
server=Mock(),
)
trio.run(main)
handshake.assert_awaited_once_with(
aid=actor.aid,
timeout=_server._PRE_REG_HANDSHAKE_TIMEOUT,
)
assert _server._PRE_REG_HANDSHAKE_TIMEOUT == 10
@pytest.mark.parametrize(
'_tpt_proto',
['uds', 'tcp']

View File

@ -5,11 +5,13 @@ Let's make sure them docs work yah?
from contextlib import contextmanager
import itertools
import os
import signal
import sys
import subprocess
import platform
import shutil
from typing import Callable
from unittest.mock import Mock
import pytest
import tractor
@ -21,6 +23,184 @@ _non_linux: bool = platform.system() != 'Linux'
_friggin_macos: bool = platform.system() == 'Darwin'
def _kill_proc_tree(proc: subprocess.Popen) -> None:
'''
Terminate an example process and its POSIX descendants.
'''
try:
if platform.system() == 'Windows':
proc.kill()
else:
os.killpg(proc.pid, signal.SIGKILL)
except ProcessLookupError:
pass
def _reap_killed_proc(
proc: subprocess.Popen,
) -> tuple[bytes, bytes]:
'''
Reap a killed process without waiting on Windows descendants.
'''
if platform.system() != 'Windows':
return proc.communicate()
proc.wait(timeout=5)
if proc.stdin:
proc.stdin.close()
if proc.stdout:
proc.stdout.close()
if proc.stderr:
proc.stderr.close()
return b'', b''
def _wait_for_proc(
proc: subprocess.Popen,
timeout: float,
test_log: tractor.log.StackLevelAdapter,
) -> None:
'''
Wait for an example process and surface its captured output.
'''
try:
out, err = proc.communicate(timeout=timeout)
except subprocess.TimeoutExpired as timeout_exc:
test_log.exception(
f'Example failed to finish within {timeout}s ??\n'
)
_kill_proc_tree(proc)
out, err = _reap_killed_proc(proc)
if platform.system() == 'Windows':
out = timeout_exc.output or b''
err = timeout_exc.stderr or b''
errmsg: str = err.decode(errors='replace')
# NOTE: always include captured stdout and stderr for a non-zero
# exit. Depending on the final stderr line previously hid grouped
# exception diagnostics; see GH #473.
#
# The prior impl only raised when the LAST stderr
# line contained 'Error', swallowing any crash whose
# traceback ends in a non-`XxxError:` line; in
# particular EVERY `tractor` root-actor crash ends
# with the strict-EG collapse note,
# '( ^^^ this exc was collapsed from a group ^^^ )',
# so ALL such failures were reduced to a bare
# `assert 1 == 0` in CI logs.. see GH #473.
rc: int|None = proc.returncode
if rc:
outmsg: str = out.decode(errors='replace')
raise Exception(
f'Example script exited with rc={rc} !?\n'
f'\n'
f'stdout:\n'
f'{outmsg}\n'
f'\n'
f'stderr:\n'
f'{errmsg}\n'
)
# if we get some gnarly output let's aggregate and raise
if errmsg:
errlines = errmsg.splitlines()
last_error = errlines[-1]
if (
'Error' in last_error
# XXX: currently we print this to console, but maybe
# shouldn't eventually once we figure out what's
# a better way to be explicit about aio side
# cancels?
and
'asyncio.exceptions.CancelledError' not in last_error
):
raise Exception(errmsg)
assert proc.returncode == 0
def test_wait_for_failed_example_captures_output():
'''
Preserve diagnostics from a subprocess which already exited.
The previous `poll()` guard skipped `communicate()` when a fast
failure returned a non-zero status before the parent checked it.
Its stdout and stderr were therefore reported as empty. This
fake process begins with `returncode=1` and returns non-UTF-8
output, proving the helper always drains both pipes and replaces
undecodable bytes without hiding the original process failure.
'''
proc = Mock()
proc.returncode = 1
proc.communicate.return_value = (
b'stdout\xff',
b'stderr\xff',
)
with pytest.raises(Exception) as exc_info:
_wait_for_proc(
proc=proc,
timeout=1,
test_log=Mock(),
)
proc.communicate.assert_called_once_with(timeout=1)
errmsg: str = str(exc_info.value)
assert 'stdout\ufffd' in errmsg
assert 'stderr\ufffd' in errmsg
@pytest.mark.skipif(
platform.system() == 'Windows',
reason='POSIX process groups are unavailable on Windows',
)
def test_wait_for_timed_out_example_reaps_group(
monkeypatch: pytest.MonkeyPatch,
):
'''
Kill the example process group and reap its leader on timeout.
The old timeout branch killed only the immediate process and
never drained it. Actor descendants could retain the capture
pipes while the leader remained unreaped, hanging CI until its
job timeout. This fake process raises `TimeoutExpired` on the
timed wait and completes on the second `communicate()` call;
the assertions prove group-directed `SIGKILL` precedes that
final drain and leaves a concrete non-zero return code.
'''
proc = Mock()
proc.pid = 1234
def communicate(timeout=None):
if timeout is not None:
raise subprocess.TimeoutExpired('example', timeout)
proc.returncode = -signal.SIGKILL
return b'', b'timed out'
proc.communicate.side_effect = communicate
killpg = Mock()
monkeypatch.setattr(os, 'killpg', killpg)
with pytest.raises(Exception, match='timed out'):
_wait_for_proc(
proc=proc,
timeout=.01,
test_log=Mock(),
)
killpg.assert_called_once_with(1234, signal.SIGKILL)
assert proc.communicate.call_count == 2
assert proc.returncode == -signal.SIGKILL
@pytest.fixture
def run_example_in_subproc(
loglevel: str,
@ -61,14 +241,14 @@ def run_example_in_subproc(
]
else:
script_file = testdir.makefile('.py', script_code)
kwargs['start_new_session'] = True
cmdargs = [
sys.executable,
str(script_file),
]
# XXX: BE FOREVER WARNED: if you enable lots of tractor logging
# in the subprocess it may cause infinite blocking on the pipes
# due to backpressure!!!
# Captured pipes are drained by `_wait_for_proc()` while the
# example runs.
proc = testdir.popen(
cmdargs,
stdin=subprocess.PIPE,
@ -77,9 +257,20 @@ def run_example_in_subproc(
**kwargs,
)
assert not proc.returncode
yield proc
proc.wait()
assert proc.returncode == 0
try:
yield proc
except BaseException:
if proc.poll() is None:
try:
_kill_proc_tree(proc)
_reap_killed_proc(proc)
except Exception:
pass
raise
else:
if proc.poll() is None:
_kill_proc_tree(proc)
_reap_killed_proc(proc)
yield run
@ -145,21 +336,6 @@ def test_example(
'This test does run just fine "in person" however..'
)
if (
'uds_transport_actor_tree' in ex_file
and
_friggin_macos
and
ci_env
):
pytest.skip(
'UDS-transport example reliably fails on macOS CI.\n'
'UDS-on-macOS is otherwise un-exercised by the matrix\n'
'(no `tpt_proto=uds` macOS job), so this new example is\n'
'the first to surface it; the macOS UDS path needs\n'
'root-causing. Passes on Linux.'
)
from .conftest import cpu_perf_headroom
timeout: float = (
@ -178,33 +354,8 @@ def test_example(
code = ex.read()
with run_example_in_subproc(code) as proc:
err = None
try:
if not proc.poll():
_, err = proc.communicate(timeout=timeout)
except subprocess.TimeoutExpired as e:
test_log.exception(
f'Example failed to finish within {timeout}s ??\n'
)
proc.kill()
err = e.stderr
# if we get some gnarly output let's aggregate and raise
if err:
errmsg = err.decode()
errlines = errmsg.splitlines()
last_error = errlines[-1]
if (
'Error' in last_error
# XXX: currently we print this to console, but maybe
# shouldn't eventually once we figure out what's
# a better way to be explicit about aio side
# cancels?
and
'asyncio.exceptions.CancelledError' not in last_error
):
raise Exception(errmsg)
assert proc.returncode == 0
_wait_for_proc(
proc=proc,
timeout=timeout,
test_log=test_log,
)

View File

@ -20,6 +20,18 @@ from typing import (
import pytest
import trio
import tractor
# `infect_asyncio` mode is unsupported on Windows (asyncio's
# `ProactorEventLoop` is incompatible with our `trio` guest-mode
# interop and currently hangs/crashes the run). Skip the module on
# Windows so the CI leg completes + reports the rest of the suite.
import platform
if platform.system() == 'Windows':
pytest.skip(
'infect_asyncio mode is unsupported on Windows',
allow_module_level=True,
)
from tractor import (
current_actor,
Actor,

View File

@ -0,0 +1,160 @@
'''
Regression tests for the cold package import surface.
'''
import json
import os
from statistics import median
import subprocess
import sys
from typing import (
Any,
get_type_hints,
)
from tractor.discovery import (
_addr,
_multiaddr,
)
from tractor.ipc import (
_tcp,
_uds,
)
def run_cold_import(code: str) -> dict[str, object]:
result = subprocess.run(
[
sys.executable,
'-c',
code,
],
check=True,
capture_output=True,
text=True,
)
return json.loads(result.stdout)
def test_lazy_to_asyncio_package_api():
'''
Keep the public lazy submodule discoverable without eagerly
importing it.
Before the lazy conversion, package import side effects exposed
`to_asyncio` to `dir()` and wildcard imports. Exercise those APIs
in cold interpreters so this test proves normal `import tractor`
leaves `asyncio` unloaded, while discovery and wildcard access
still advertise and resolve the public submodule.
'''
cold = run_cold_import(
'import json, sys, tractor; '
'print(json.dumps({'
'"advertised": "to_asyncio" in dir(tractor), '
'"asyncio_loaded": "asyncio" in sys.modules}))'
)
assert cold == {
'advertised': True,
'asyncio_loaded': False,
}
wildcard = run_cold_import(
'import json; '
'from tractor import *; '
'print(json.dumps({'
'"module": to_asyncio.__name__}))'
)
assert wildcard == {
'module': 'tractor.to_asyncio',
}
def test_cold_import_budget():
'''
Keep cold package import below the pre-optimization regression.
The original `inspect.stack()` caller lookup made a fresh
`import tractor` take about 0.42s and dominate actor startup.
Run seven independent interpreters and gate their median at a
deliberately broad 0.35s: over twice the measured ~0.145s
baseline, but low enough to catch restoration of that hot path.
Taking the median absorbs process-start and shared-runner noise.
The child measures only its import, rather than parent-side
process creation. `TRACTOR_IMPORT_BUDGET_S` provides an explicit,
reviewable override for platforms that establish a different
baseline instead of silently weakening the project default.
Each child also reports the modules whose eager loading this PR
intentionally removes, proving a timing pass cannot hide a
dependency-import regression.
'''
budget_s = float(
os.environ.get(
'TRACTOR_IMPORT_BUDGET_S',
'0.35',
)
)
optional_mods = (
'asyncio',
'bidict',
'colorlog',
'multiaddr',
'wrapt',
)
code = (
'import json, sys, time; '
'started = time.perf_counter(); '
'import tractor; '
'elapsed = time.perf_counter() - started; '
f'optional = {optional_mods!r}; '
'print(json.dumps({'
'"elapsed": elapsed, '
'"loaded": [name for name in optional '
'if name in sys.modules]}))'
)
samples = [
run_cold_import(code)
for _ in range(7)
]
elapsed = [
float(sample['elapsed'])
for sample in samples
]
loaded = {
name
for sample in samples
for name in sample['loaded']
}
assert not loaded
assert median(elapsed) < budget_s, (
f'cold import median exceeded {budget_s:.3f}s budget: '
f'{elapsed!r}'
)
def test_lazy_annotation_names_resolve():
'''
Resolve annotations without importing optional dependencies.
Moving annotation-only third-party names under `TYPE_CHECKING`
left their runtime globals undefined, causing
`typing.get_type_hints()` to raise `NameError`. Resolve every
affected API and prove the lazy aliases retain import-free runtime
introspection.
'''
assert get_type_hints(_multiaddr.mk_maddr)['return'] is Any
assert get_type_hints(_tcp.MsgpackTCPStream.maddr.fget)[
'return'
] is Any
assert get_type_hints(_uds.MsgpackUDSStream.maddr.fget)[
'return'
] == Any|str
assert get_type_hints(_addr.Address.get_random)[
'current_actor'
] is Any
assert _addr.__annotations__['_address_types'].startswith('dict')

View File

@ -2,9 +2,11 @@
`tractor.log`-wrapping unit tests.
'''
import importlib
import logging
from pathlib import Path
import shutil
import sys
from types import ModuleType
import pytest
@ -165,6 +167,53 @@ def test_implicit_mod_name_applied_for_child(
assert submod.log.logger in sub_logs
def test_implicit_mod_name_from_unregistered_namespace(
tmp_path: Path,
):
'''
Preserve implicit logger naming for dynamic module namespaces.
The fast `sys.modules` caller lookup cannot resolve `runpy`,
plugin-loader, or `exec()` namespaces that are not registered.
Compile a real package file under an unregistered module name so
the rare filename fallback must recover its imported package and
retain the same package-level logger name.
'''
pkg_name = 'dynamic_logger_pkg'
pkg_dir = tmp_path / pkg_name
pkg_dir.mkdir()
init_path = pkg_dir / '__init__.py'
init_path.write_text('')
mod_path = pkg_dir / 'plugin.py'
mod_path.write_text('')
sys.path.insert(0, str(tmp_path))
try:
importlib.import_module(pkg_name)
namespace = {
'__name__': f'{pkg_name}.unregistered',
'__package__': pkg_name,
'tractor': tractor,
}
exec(
compile(
'log = tractor.log.get_logger('
f'pkg_name={pkg_name!r})',
str(mod_path),
'exec',
),
namespace,
)
dynamic_log = namespace.get('log')
finally:
sys.path.remove(str(tmp_path))
sys.modules.pop(pkg_name, None)
assert dynamic_log is not None
assert dynamic_log.name == pkg_name
def test_io_custom_level_registered():
'''
The `IO`(21) level (registered via `add_log_level()` at

View File

@ -1,10 +1,22 @@
import time
import platform
import trio
import pytest
import tractor
# `tractor.ipc._ringbuf` is built on linux `eventfd(2)`; importing
# it pulls in `tractor.ipc._linux` whose module-level
# `ffi.dlopen(None)` raises on non-linux. Skip the whole module at
# COLLECTION before that crashing import runs (a `pytestmark` skip
# is too late — markers apply only after the import succeeds).
if platform.system() != 'Linux':
pytest.skip(
'ringbuf (eventfd) IPC is linux-only',
allow_module_level=True,
)
# XXX `cffi` dun build on py3.14 yet..
pytest.importorskip("cffi")

View File

@ -9,6 +9,17 @@ from functools import partial
import pytest
import trio
import tractor
# `infect_asyncio` mode is unsupported on Windows (see
# `test_infected_asyncio`); skip at COLLECTION before the
# asyncio-interop imports below so the CI leg completes.
import platform
if platform.system() == 'Windows':
pytest.skip(
'infect_asyncio mode is unsupported on Windows',
allow_module_level=True,
)
from tractor import (
to_asyncio,
)

View File

@ -4,11 +4,18 @@ related API and error checks.
'''
import itertools
from unittest.mock import (
AsyncMock,
Mock,
)
import pytest
import tractor
import trio
from tractor._exceptions import TransportClosed
from tractor.runtime import _rpc
async def sleep_back_actor(
actor_name,
@ -46,6 +53,126 @@ async def short_sleep():
await trio.sleep(0)
def test_rpc_runs_after_startack_disconnect():
'''
Complete an accepted RPC when its caller closes before `StartAck`.
Registrar teardown opens short-lived `unregister_actor` RPCs. A
loaded caller can close its channel while the registrar sends the
acknowledgement; normalized `TransportClosed` previously escaped
into the shared service nursery before the already-created
coroutine was awaited. This fake fails the first response send and
proves the RPC side effect still runs with no later send attempt.
'''
async def main():
rpc_ran = trio.Event()
chan = Mock()
chan.send = AsyncMock(
side_effect=TransportClosed(
message='caller closed before StartAck',
),
)
chan.connected.return_value = False
ctx = Mock(
chan=chan,
cid='rpc-cid',
_scope=None,
_task='rpc-task',
)
actor = Mock()
actor.get_context.return_value = ctx
actor._rpc_tasks = {}
actor._ongoing_rpc_tasks = trio.Event()
actor._ongoing_rpc_tasks.set()
async def rpc_func():
assert (chan, ctx.cid) in actor._rpc_tasks
rpc_ran.set()
async def invoke(task_status):
await _rpc._invoke(
actor=actor,
cid=ctx.cid,
chan=chan,
func=rpc_func,
kwargs={},
task_status=task_status,
)
async with trio.open_nursery() as nursery:
started_ctx = await nursery.start(invoke)
assert started_ctx is ctx
await rpc_ran.wait()
assert rpc_ran.is_set()
assert not actor._rpc_tasks
assert actor._ongoing_rpc_tasks.is_set()
chan.send.assert_awaited_once()
trio.run(main)
def test_error_shipment_ignores_closed_response_channel(monkeypatch):
'''
Preserve an application error when its response channel is closed.
A caller can disconnect after submitting an RPC but before its
error response. Normalized `TransportClosed` from that final send
is terminal response failure, not a new actor-wide service error.
This test proves error shipment logs and returns without replacing
the original application exception.
'''
chan = Mock()
chan.send = AsyncMock(
side_effect=[
None,
TransportClosed(
message='caller closed before Error response',
),
],
)
error_msg = Mock(boxed_type_str='ValueError')
monkeypatch.setattr(
_rpc,
'pack_error',
Mock(return_value=error_msg),
)
ctx = Mock(
chan=chan,
cid='rpc-cid',
_scope=None,
_task='rpc-task',
)
actor = Mock()
actor.get_context.return_value = ctx
actor._rpc_tasks = {}
actor._ongoing_rpc_tasks = trio.Event()
actor._ongoing_rpc_tasks.set()
async def failing_rpc():
raise ValueError('application failure')
async def main():
async with trio.open_nursery() as nursery:
started_ctx = await nursery.start(
_rpc._invoke,
actor,
ctx.cid,
chan,
failing_rpc,
{},
)
assert started_ctx is ctx
trio.run(main)
assert chan.send.await_count == 2
assert not actor._rpc_tasks
assert actor._ongoing_rpc_tasks.is_set()
@pytest.mark.parametrize(
'to_call', [
([], 'short_sleep', tractor.RemoteActorError),

View File

@ -73,6 +73,11 @@ async def test_lifetime_stack_wipes_tmpfile(
1.6 if error_in_child
else 1
)
# scale for slow/noisy CI (esp. macOS) so the child error
# propagates before the deadline; otherwise `move_on_after`
# cancels first and flips the `error_in_child=True` assert.
from .conftest import cpu_perf_headroom
timeout *= cpu_perf_headroom()
try:
with trio.move_on_after(timeout) as cs:
async with tractor.open_nursery(

View File

@ -120,6 +120,13 @@ async def child_read_shm_list(
print(f'(child): reading frame: {frame}')
@pytest.mark.skipif(
platform.system() == 'Windows',
reason=(
'parent/child shm IPC deadlocks on Windows '
'(frame-size dependent hang); nascent — see #404'
),
)
@pytest.mark.parametrize(
'use_str',
[False, True],

View File

@ -75,3 +75,39 @@ from .discovery._registry import (
Arbiter as Arbiter,
)
# from . import hilevel as hilevel
__all__: tuple[str, ...] = tuple(
name
for name in globals()
if not name.startswith('_')
) + (
'to_asyncio',
)
def __dir__() -> list[str]:
return sorted(set(globals()) | set(__all__))
def __getattr__(name: str):
'''
PEP 562 lazy sub-module loading, presently only for
`.to_asyncio` which (transitively) imports `asyncio`
itself: a non-trivial multi-ms chunk of the eager
`import tractor` cost (gh #470) unneeded by
`trio`-only apps.
Any `tractor.to_asyncio.<attr>` access (or a
`from tractor import to_asyncio`) still works, the
sub-mod is simply imported on first-access instead
of at pkg-import time.
'''
if name == 'to_asyncio':
from importlib import import_module
return import_module('.to_asyncio', __name__)
raise AttributeError(
f'module {__name__!r} has no attribute {name!r}'
)

View File

@ -47,9 +47,7 @@ from .devx import (
from .spawn import _spawn
from .runtime import _state
from . import log
from .ipc import (
_connect_chan,
)
from .discovery._api import _probe_registry_addrs
from .discovery._addr import (
Address,
UnwrappedAddress,
@ -451,50 +449,23 @@ async def open_root_actor(
from .devx._stackscope import enable_stack_on_sig
enable_stack_on_sig()
# closed into below ping task-func
ponged_addrs: list[Address] = []
ponged_addrs: list[Address]
occupied_addrs: list[Address]
(
ponged_addrs,
occupied_addrs,
) = await _probe_registry_addrs(uw_reg_addrs)
async def ping_tpt_socket(
addr: Address,
timeout: float = 1,
) -> None:
'''
Attempt temporary connection to see if a registry is
listening at the requested address by a tranport layer
ping.
If a connection can't be made quickly we assume none no
server is listening at that addr.
'''
try:
# TODO: this connect-and-bail forces us to have to
# carefully rewrap TCP 104-connection-reset errors as
# EOF so as to avoid propagating cancel-causing errors
# to the channel-msg loop machinery. Likely it would
# be better to eventually have a "discovery" protocol
# with basic handshake instead?
with trio.move_on_after(timeout):
async with _connect_chan(addr.unwrap()):
ponged_addrs.append(addr)
except OSError:
# ?TODO, make this a "discovery" log level?
logger.info(
f'No root-actor registry found @ {addr!r}\n'
)
# !TODO, this is basically just another (abstract)
# happy-eyeballs, so we should try for formalize it somewhere
# in a `.[_]discovery` ya?
#
async with trio.open_nursery() as tn:
for uw_addr in uw_reg_addrs:
addr: Address = wrap_address(uw_addr)
tn.start_soon(
ping_tpt_socket,
addr,
)
if (
not ponged_addrs
and
occupied_addrs
):
raise RuntimeError(
f'Registry address(es) are occupied but did not '
f'answer as Tractor registrars!\n'
f'occupied_addrs: {occupied_addrs!r}\n'
)
if tpt_bind_addrs is None:
tpt_bind_addrs: list[Address] = []

View File

@ -111,14 +111,13 @@ SHM_DIR: str = '/dev/shm'
# UDS-socket leak sweep — see `find_orphaned_uds()` /
# `reap_uds()` below. Tractor's UDS transport
# (`tractor.ipc._uds`) creates sock files under
# `${XDG_RUNTIME_DIR}/tractor/<name>@<pid>.sock`; a
# (`tractor.ipc._uds`) creates sock files in its platform-specific
# default bindspace; a
# crash / SIGKILL / mid-cancel teardown can leave the
# file behind because `os.unlink()` lives in the
# `_serve_ipc_eps` `finally:` block which doesn't always
# get to run on hard exits. The reaper here is best-effort
# cleanup for the test harness + the `tractor-reap` CLI.
_UDS_SUBDIR: str = 'tractor'
# `<actor-name>@<pid>.sock` — pid is the binder's pid at
# creation time. Special sentinel: `registry@1616.sock`
# uses the magic `1616` not a real pid (the root
@ -738,19 +737,16 @@ def reap_shm(
def get_uds_dir() -> str|None:
'''
Path of tractor's per-user UDS sock-file dir
(`${XDG_RUNTIME_DIR}/tractor/`).
Path of Tractor's platform-specific default UDS bindspace.
Returns `None` when `XDG_RUNTIME_DIR` is unset (e.g.
non-systemd hosts, or inside a container without the
var plumbed through). Caller should treat that as
"no UDS leaks possible to detect — skip".
Returns `None` only when the bindspace cannot be resolved.
'''
xdg: str|None = os.environ.get('XDG_RUNTIME_DIR')
if not xdg:
try:
from tractor.ipc._uds import UDSAddress
return str(UDSAddress.def_bindspace)
except Exception:
return None
return os.path.join(xdg, _UDS_SUBDIR)
def _parse_uds_name(filename: str) -> tuple[str, int]|None:
@ -768,16 +764,16 @@ def _parse_uds_name(filename: str) -> tuple[str, int]|None:
def find_orphaned_uds(
*,
uds_dir: str|None = None,
include_registry_sentinel: bool = False,
) -> list[str]:
'''
`<uds_dir>/*.sock` paths whose binder pid is no
longer alive (orphaned). Includes the
`registry@1616.sock` sentinel `1616` is a magic
sentinel pid (not a real one) so the file's
presence alone signals a leak from a dead session.
longer alive (orphaned). Explicit callers may include the
`registry@1616.sock` sentinel; automatic pytest cleanup excludes
it because binder liveness cannot be inferred from magic `1616`.
Returns `[]` on platforms without `XDG_RUNTIME_DIR`
or when the dir doesn't exist. Files whose name
Returns `[]` when the platform bindspace cannot be resolved or the
dir doesn't exist. Files whose name
doesn't match the `<name>@<pid>.sock` pattern are
skipped (we don't unlink things we don't recognize).
@ -811,10 +807,8 @@ def find_orphaned_uds(
continue
_name, pid = parsed
if pid == _UDS_REGISTRY_SENTINEL_PID:
# sentinel — never a real pid; if the file
# exists nobody live is "owning" it via
# /proc lookup, so always orphaned
leaked.append(path)
if include_registry_sentinel:
leaked.append(path)
continue
if not _is_alive(pid):
leaked.append(path)
@ -933,8 +927,8 @@ def track_orphaned_uds_per_test():
teardown that flakifies sibling tests via
sock-file rebind races).
Snapshots `${XDG_RUNTIME_DIR}/tractor/` before and
after each test; any `<name>@<pid>.sock` files
Snapshots Tractor's platform-specific default UDS bindspace before
and after each test; any `<name>@<pid>.sock` files
created during the test that survive teardown AND
whose creator pid is dead are surfaced as a loud
warning AND reaped, so the next test starts with a
@ -950,8 +944,8 @@ def track_orphaned_uds_per_test():
it (vs. blanket session-end sweep) makes blame
obvious + prevents cascade flakiness.
Cheap: 2x `os.listdir` + a few `os.stat`s per test.
Skips silently when `XDG_RUNTIME_DIR` isn't set.
Cheap: 2x `os.listdir` + a few `os.stat`s per test. Skips silently
when the platform bindspace cannot be resolved.
'''
uds_dir: str|None = get_uds_dir()

View File

@ -39,14 +39,15 @@ from typing import (
Type,
)
import pdbp
# NOTE, `pdbp` + `wrapt` are lazy-imported at their
# single use-sites below to keep them off the eager
# `import tractor` path (gh #470).
from tractor.log import get_logger
import trio
from tractor.msg import (
pretty_struct,
NamespacePath,
)
import wrapt
log = get_logger()
@ -257,6 +258,7 @@ def api_frame(
caller_frames_up: int = 1,
) -> Callable:
import wrapt
# handle the decorator called WITHOUT () case,
# i.e. just @api_frame, NOT @api_frame(extra=<blah>)
@ -320,6 +322,8 @@ def hide_runtime_frames() -> dict[FunctionType, CodeType]:
as possible, particularly from inside a `PdbREPL`.
'''
import pdbp
# XXX HACKZONE XXX
# hide exit stack frames on nurseries and cancel-scopes!
# |_ so avoid seeing it when the `pdbp` REPL is first engaged from

View File

@ -31,12 +31,21 @@ from threading import (
RLock,
)
import multiprocessing as mp
import platform
from signal import (
signal,
getsignal,
SIGUSR1,
SIGINT,
)
if platform.system() != "Windows":
from signal import SIGUSR1
else:
SIGUSR1 = None
# import traceback
from types import ModuleType
from typing import (
@ -347,8 +356,8 @@ def dump_tree_on_sig(
def enable_stack_on_sig(
sig: int = SIGUSR1,
) -> ModuleType:
sig: int|None = SIGUSR1,
) -> ModuleType|None:
'''
Enable `stackscope` tracing on reception of a signal; by
default this is SIGUSR1.
@ -367,6 +376,16 @@ def enable_stack_on_sig(
>> pkill --signal SIGUSR1 -f <part-of-cmd: str>
'''
# no `SIGUSR1` on this platform (e.g. Windows) -> nothing to
# wire up; degrade gracefully instead of crashing callers that
# only guard against a missing `stackscope` (`ImportError`).
if sig is None:
log.warning(
'No `SIGUSR1` on this platform;\n'
'skipping `stackscope` trace-on-signal setup!\n'
)
return None
try:
# NOTE, `stackscope._glue` does intentional async-gen type
# introspection at import-time which trips

View File

@ -24,7 +24,6 @@ mult-process support within a single actor tree.
'''
from __future__ import annotations
import asyncio
import bdb
from contextlib import (
AbstractContextManager,
@ -53,7 +52,6 @@ from trio import (
)
import tractor
from tractor.log import get_logger
from tractor.to_asyncio import run_trio_task_in_future
from tractor._context import Context
from tractor.runtime import _state
from tractor._exceptions import (
@ -85,6 +83,10 @@ from ..pformat import (
)
if TYPE_CHECKING:
# NOTE, `asyncio` (and `.to_asyncio`) are
# lazy-imported at their use-sites to keep them off
# the eager `import tractor` path (gh #470).
import asyncio
from trio.lowlevel import Task
from threading import Thread
from tractor.runtime._runtime import (
@ -164,6 +166,7 @@ async def _pause(
'An `asyncio` task should not be calling this!?'
) from rte
else:
import asyncio
task = asyncio.current_task()
if debug_func is not None:
@ -946,6 +949,7 @@ def pause_from_sync(
asyncio_task: asyncio.Task|None = None
if is_infected_aio:
import asyncio
asyncio_task = asyncio.current_task()
# TODO: we could also check for a non-`.to_thread` context
@ -1059,6 +1063,9 @@ def pause_from_sync(
greenback: ModuleType = maybe_import_greenback()
if greenback.has_portal():
from tractor.to_asyncio import (
run_trio_task_in_future,
)
DebugStatus.shield_sigint()
fute: asyncio.Future = run_trio_task_in_future(
partial(

View File

@ -20,7 +20,6 @@ Root-actor TTY mutex-locking machinery.
'''
from __future__ import annotations
import asyncio
from contextlib import (
AbstractContextManager,
asynccontextmanager as acm,
@ -52,7 +51,6 @@ from trio import (
TaskStatus,
)
import tractor
from tractor.to_asyncio import run_trio_task_in_future
from tractor.log import get_logger
from tractor._context import Context
from tractor.runtime import _state
@ -66,6 +64,10 @@ from tractor.runtime._state import (
)
if TYPE_CHECKING:
# NOTE, `asyncio` (and `.to_asyncio`) are
# lazy-imported at their use-sites to keep them off
# the eager `import tractor` path (gh #470).
import asyncio
from trio.lowlevel import Task
from threading import Thread
from tractor.ipc import (
@ -910,6 +912,9 @@ class DebugStatus:
async def _set_repl_release():
repl_release.set()
from tractor.to_asyncio import (
run_trio_task_in_future,
)
fute: asyncio.Future = run_trio_task_in_future(
_set_repl_release
)

View File

@ -16,13 +16,13 @@
from __future__ import annotations
from uuid import uuid4
from typing import (
Any,
Protocol,
ClassVar,
Type,
TYPE_CHECKING,
)
from bidict import bidict
from trio import (
SocketListener,
)
@ -32,14 +32,20 @@ from ..runtime._state import (
_def_tpt_proto,
)
from ..ipc._tcp import TCPAddress
from ..ipc._uds import UDSAddress
from ..ipc._uds import (
UDSAddress,
HAS_UDS,
)
if TYPE_CHECKING:
# ONLY type-annots, the eager import costs ~4.5ms
# of `import tractor` wall-time (gh #470).
from ..runtime._runtime import Actor
else:
Actor = Any
log = get_logger()
# TODO, maybe breakout the netns key to a struct?
# class NetNs(Struct)[str, int]:
# ...
@ -170,25 +176,36 @@ class Address(Protocol):
...
_address_types: bidict[str, Type[Address]] = {
'tcp': TCPAddress,
'uds': UDSAddress
# the address types available on this host: TCP always, UDS only
# where usable (`HAS_UDS`). Both registries derive from this single
# list via each type's `proto_key`.
_address_protos: list[Type[Address]] = [TCPAddress]
if HAS_UDS:
_address_protos.append(UDSAddress)
_address_types: dict[str, Type[Address]] = {
cls.proto_key: cls
for cls in _address_protos
}
# TODO! really these are discovery sys default addrs ONLY useful for
# when none is provided to a root actor on first boot.
_default_lo_addrs: dict[
str,
UnwrappedAddress
] = {
'tcp': TCPAddress.get_root().unwrap(),
'uds': UDSAddress.get_root().unwrap(),
_default_lo_addrs: dict[str, UnwrappedAddress] = {
cls.proto_key: cls.get_root().unwrap()
for cls in _address_protos
}
def get_address_cls(name: str) -> Type[Address]:
return _address_types[name]
try:
return _address_types[name]
except KeyError:
raise NotImplementedError(
f'No IPC transport backend for {name!r} on this '
f'platform!\n'
f'(available: {list(_address_types)})\n'
)
def is_wrapped_addr(addr: any) -> bool:
@ -286,7 +303,14 @@ def default_lo_addrs(
for an input transport key set.
'''
return [
_default_lo_addrs[transport]
for transport in transports
]
lo_addrs: list[UnwrappedAddress] = []
for transport in transports:
try:
lo_addrs.append(_default_lo_addrs[transport])
except KeyError:
raise NotImplementedError(
f'No default loopback addr for transport '
f'{transport!r} on this platform!\n'
f'(available: {list(_default_lo_addrs)})\n'
)
return lo_addrs

View File

@ -21,14 +21,18 @@ management of (service) actors.
"""
from __future__ import annotations
import ipaddress
import os
import socket
from typing import (
AsyncGenerator,
AsyncContextManager,
Literal,
TYPE_CHECKING,
)
from contextlib import asynccontextmanager as acm
import trio
from tractor.log import get_logger
from ..trionics import (
gather_contexts,
@ -40,6 +44,7 @@ from ..ipc._uds import UDSAddress
from ._addr import (
UnwrappedAddress,
Address,
mk_uuid,
wrap_address,
)
from ..runtime._portal import (
@ -52,6 +57,7 @@ from ..runtime._state import (
_runtime_vars,
_def_tpt_proto,
)
from ..msg.types import Aid
if TYPE_CHECKING:
from ..runtime._runtime import Actor
@ -60,6 +66,115 @@ if TYPE_CHECKING:
log = get_logger()
async def _probe_registry(
addr: Address,
timeout: float = 3,
attempt_timeout: float = 1,
max_attempts: int = 3,
retry_delay: float = .05,
close_timeout: float = .2,
) -> Literal[
'absent',
'occupied',
'registrar',
]:
'''
Confirm an address serves the Tractor actor handshake.
Connection and handshake work share `timeout`; each attempt gets
`attempt_timeout`. Shielded cleanup may add up to `close_timeout`
per attempted channel.
'''
from .._exceptions import TransportClosed
connected_once: bool = False
with trio.move_on_after(timeout):
for attempt in range(max_attempts):
try:
with trio.move_on_after(attempt_timeout) as attempt_cs:
async with _connect_chan(
addr.unwrap(),
close_timeout=close_timeout,
) as chan:
connected_once = True
peer_aid: Aid = await chan._do_handshake(
aid=Aid(
name='registry-probe',
uuid=mk_uuid(),
pid=os.getpid(),
is_probe=True,
),
timeout=attempt_timeout,
)
if peer_aid.is_registrar is not False:
return 'registrar'
return 'occupied'
if attempt_cs.cancelled_caught:
if not connected_once:
return 'absent'
except OSError:
return (
'occupied'
if connected_once
else 'absent'
)
except TransportClosed:
pass
if attempt + 1 < max_attempts:
await trio.sleep(retry_delay * (attempt + 1))
return 'occupied'
async def _probe_registry_addrs(
addrs: list[UnwrappedAddress],
timeout: float = 3,
) -> tuple[
list[Address],
list[Address],
]:
'''
Concurrently classify candidate registrar addresses.
Return confirmed registrar addresses followed by addresses occupied
by non-registrar or unresponsive Tractor peers.
'''
registrar_addrs: list[Address] = []
occupied_addrs: list[Address] = []
async def probe_addr(addr: Address) -> None:
probe_status = await _probe_registry(
addr=addr,
timeout=timeout,
)
if probe_status == 'registrar':
registrar_addrs.append(addr)
elif probe_status == 'occupied':
occupied_addrs.append(addr)
else:
# ?TODO, make this a "discovery" log level?
log.info(
f'No root-actor registry found @ {addr!r}\n'
)
async with trio.open_nursery() as nursery:
for unwrapped_addr in addrs:
nursery.start_soon(
probe_addr,
wrap_address(unwrapped_addr),
)
return (
registrar_addrs,
occupied_addrs,
)
def _is_local_addr(addr: Address) -> bool:
'''
Determine whether `addr` is reachable on the

View File

@ -24,14 +24,23 @@ Multiaddress support using the upstream `py-multiaddr` lib
- https://github.com/multiformats/multiaddr/blob/master/protocols/unix.md
'''
from __future__ import annotations
import ipaddress
from pathlib import Path
from typing import TYPE_CHECKING
from multiaddr import Multiaddr
from typing import (
Any,
TYPE_CHECKING,
)
if TYPE_CHECKING:
# NOTE, `multiaddr` is lazy-imported at first use
# (in the fns below) to keep it off the eager
# `import tractor` path (gh #470).
from multiaddr import Multiaddr
from tractor.discovery._addr import Address
else:
Multiaddr = Any
Address = Any
# map from tractor-internal `proto_key` identifiers
# to the standard multiaddr protocol name strings.
@ -56,6 +65,8 @@ def mk_maddr(
multiaddr-spec-compliant protocol path.
'''
from multiaddr import Multiaddr
proto_key: str = addr.proto_key
maddr_proto: str|None = _tpt_proto_to_maddr.get(proto_key)
if maddr_proto is None:
@ -98,6 +109,7 @@ def parse_maddr(
'''
# lazy imports to avoid circular deps
from multiaddr import Multiaddr
from tractor.ipc._tcp import TCPAddress
from tractor.ipc._uds import UDSAddress

View File

@ -33,6 +33,7 @@ from typing import (
)
import warnings
import msgspec
import trio
from ._types import (
@ -495,6 +496,7 @@ class Channel:
async def _do_handshake(
self,
aid: Aid,
timeout: float|None = None,
) -> Aid:
'''
@ -505,8 +507,28 @@ class Channel:
"actor model" parlance.
'''
await self.send(aid)
peer_aid: Aid = await self.recv()
try:
with trio.fail_after(
timeout if timeout is not None else float('inf')
):
await self.send(aid)
peer_aid: Aid = await self.recv()
if not isinstance(peer_aid, Aid):
raise TypeError(
f'Expected {Aid!r}, received {peer_aid!r}'
)
except (
MsgTypeError,
msgspec.DecodeError,
TypeError,
UnicodeDecodeError,
trio.TooSlowError,
) as handshake_err:
raise TransportClosed(
message='Peer sent an invalid actor handshake!\n',
src_exc=handshake_err,
loglevel='warning',
) from handshake_err
log.runtime(
f'Received hanshake with peer\n'
f'<= {peer_aid.reprol(sin_uuid=False)}\n'
@ -518,7 +540,8 @@ class Channel:
@acm
async def _connect_chan(
addr: UnwrappedAddress
addr: UnwrappedAddress,
close_timeout: float|None = None,
) -> typing.AsyncGenerator[Channel, None]:
'''
Create and connect a `Channel` to the provided `addr`, disconnect
@ -529,6 +552,18 @@ async def _connect_chan(
'''
chan = await Channel.from_addr(addr)
yield chan
with trio.CancelScope(shield=True):
await chan.aclose()
try:
yield chan
finally:
with trio.CancelScope(shield=True):
if close_timeout is None:
await chan.aclose()
else:
with trio.move_on_after(close_timeout) as close_cs:
await chan.aclose()
if close_cs.cancelled_caught:
log.warning(
f'Timed out closing channel after '
f'{close_timeout}s\n'
f'|_{chan}\n'
)

View File

@ -62,16 +62,19 @@ from .. import log
from ..discovery._addr import Address
from ._chan import Channel
from ._transport import MsgTransport
from ._uds import UDSAddress
from ._tcp import TCPAddress
if TYPE_CHECKING:
from ..runtime._runtime import Actor
from ..runtime._supervise import ActorNursery
from ._tcp import TCPAddress
from ._uds import UDSAddress
log = log.get_logger()
_PRE_REG_HANDSHAKE_TIMEOUT: float = 10
async def maybe_wait_on_canced_subs(
uid: tuple[str, str],
@ -316,8 +319,6 @@ async def handle_stream_from_peer(
)
'''
server._no_more_peers = trio.Event() # unset by making new
# TODO, debug_mode tooling for when hackin this lower layer?
# with debug.maybe_open_crash_handler(
# pdb=True,
@ -335,6 +336,7 @@ async def handle_stream_from_peer(
if actor := _state.current_actor():
peer_aid: msgtypes.Aid = await chan._do_handshake(
aid=actor.aid,
timeout=_PRE_REG_HANDSHAKE_TIMEOUT,
)
except (
TransportClosed,
@ -351,12 +353,9 @@ async def handle_stream_from_peer(
# "kinda-error" that we expect to tolerate during
# discovery-sys related pings, queires, DoS etc.
):
# XXX: This may propagate up from `Channel._aiter_recv()`
# and `MsgpackStream._inter_packets()` on a read from the
# stream particularly when the runtime is first starting up
# inside `open_root_actor()` where there is a check for
# a bound listener on the registrar addr. the reset will be
# because the handshake was never meant took place.
# `TransportClosed` is expected when a peer disconnects or
# fails the initial typed handshake, including foreign clients
# and probes racing shutdown.
log.runtime(
con_status
+
@ -364,6 +363,13 @@ async def handle_stream_from_peer(
)
return
# Registry election probes need only the server's `Aid` capability
# response; never register them as ordinary RPC peers.
if peer_aid.is_probe:
return
server._no_more_peers = trio.Event() # unset by making new
uid: tuple[str, str] = (
peer_aid.name,
peer_aid.uuid,

View File

@ -20,7 +20,9 @@ TCP implementation of tractor.ipc._transport.MsgTransport protocol
from __future__ import annotations
import ipaddress
from typing import (
Any,
ClassVar,
TYPE_CHECKING,
)
# from contextlib import (
# asynccontextmanager as acm,
@ -33,7 +35,6 @@ from trio import (
open_tcp_listeners,
)
from multiaddr import Multiaddr
from tractor.msg import MsgCodec
from tractor.log import get_logger
from tractor.discovery._multiaddr import mk_maddr
@ -42,6 +43,13 @@ from tractor.ipc._transport import (
MsgpackTransport,
)
if TYPE_CHECKING:
# ONLY type-annots, the eager import costs
# `import tractor` wall-time (gh #470).
from multiaddr import Multiaddr
else:
Multiaddr = Any
log = get_logger()

View File

@ -33,6 +33,7 @@ from collections.abc import (
AsyncGenerator,
AsyncIterator,
)
import errno
import struct
import trio
@ -61,6 +62,68 @@ if TYPE_CHECKING:
log = get_logger()
def _peer_closed_errno(exc: BaseException) -> int|None:
'''
Classify a complete transport exception tree as peer closure.
Follow explicit cause/context links. For a `BaseExceptionGroup`,
require every child branch to resolve to a peer-close errno so an
unrelated concurrent failure is never hidden as `TransportClosed`.
'''
def find_peer_errno(
current_exc: BaseException,
ancestors: set[int],
) -> int|None:
exc_id: int = id(current_exc)
if exc_id in ancestors:
return None
ancestors = ancestors | {exc_id}
if (
isinstance(current_exc, OSError)
and
current_exc.errno in {
errno.ECONNRESET,
errno.EPIPE,
}
):
return current_exc.errno
if isinstance(current_exc, BaseExceptionGroup):
child_errnos: list[int|None] = [
find_peer_errno(
child_exc,
ancestors,
)
for child_exc in current_exc.exceptions
]
if all(
child_errno is not None
for child_errno in child_errnos
):
return child_errnos[0]
return None
chained_exc: BaseException|None = (
current_exc.__cause__
or
current_exc.__context__
)
if chained_exc is not None:
return find_peer_errno(
chained_exc,
ancestors,
)
return None
return find_peer_errno(
exc,
set(),
)
# (codec, transport)
MsgTransportKey = tuple[str, str]
@ -443,23 +506,23 @@ class MsgpackTransport(MsgTransport):
trans_err = _re
tpt_name: str = f'{type(self).__name__!r}'
trans_err_msg: str = trans_err.args[0]
trans_err_msg: str = (
str(trans_err.args[0])
if trans_err.args
else ''
)
by_whom: str = {
'another task closed this fd': 'locally',
'this socket was already closed': 'by peer',
}.get(trans_err_msg)
match trans_err:
# XXX, specifc to UDS transport and its,
# well, "speediness".. XD
# |_ likely todo with races related to how fast
# the socket is setup/torn-down on linux
# as it pertains to rando pings from the
# `.discovery` subsys and protos.
# UDS peers can disconnect before handshake.
# Linux normally reports `EPIPE`; Darwin reports
# `ECONNRESET` for the same expected closure.
case trio.BrokenResourceError() if (
'[Errno 32] Broken pipe'
in
trans_err_msg
_peer_closed_errno(trans_err)
is not None
):
tpt_closed = TransportClosed.from_src_exc(
message=(

View File

@ -18,17 +18,13 @@
IPC subsys type-lookup helpers?
'''
from typing import (
Type,
# TYPE_CHECKING,
)
import trio
from typing import Type
import socket
import trio
from tractor.ipc._transport import (
MsgTransportKey,
MsgTransport
MsgTransport,
)
from tractor.ipc._tcp import (
TCPAddress,
@ -37,37 +33,33 @@ from tractor.ipc._tcp import (
from tractor.ipc._uds import (
UDSAddress,
MsgpackUDSStream,
HAS_UDS,
)
# if TYPE_CHECKING:
# from tractor._addr import Address
# the UDS backend is importable everywhere but only *usable* when
# `HAS_UDS` is `True`; otherwise the runtime registers TCP only.
Address = TCPAddress|UDSAddress
# manually updated list of all supported msg transport types
_msg_transports = [
# the available msg-transport backends on this host: TCP always,
# UDS only where usable (`HAS_UDS`). The lookup maps below derive
# from this single list via each backend's `codec_key` and
# `address_type`: register a backend here and every map picks it up.
_msg_transports: list[Type[MsgTransport]] = [
MsgpackTCPStream,
MsgpackUDSStream
]
if HAS_UDS:
_msg_transports.append(MsgpackUDSStream)
# convert a MsgTransportKey to the corresponding transport type
_key_to_transport: dict[
MsgTransportKey,
Type[MsgTransport],
] = {
('msgpack', 'tcp'): MsgpackTCPStream,
('msgpack', 'uds'): MsgpackUDSStream,
# map a `MsgTransportKey` -> `MsgTransport` type
_key_to_transport: dict[MsgTransportKey, Type[MsgTransport]] = {
(t.codec_key, t.address_type.proto_key): t
for t in _msg_transports
}
# convert an Address wrapper to its corresponding transport type
_addr_to_transport: dict[
Type[TCPAddress|UDSAddress],
Type[MsgTransport]
] = {
TCPAddress: MsgpackTCPStream,
UDSAddress: MsgpackUDSStream,
# map an `Address`-wrapper -> `MsgTransport` type
_addr_to_transport: dict[Type[Address], Type[MsgTransport]] = {
t.address_type: t
for t in _msg_transports
}
@ -81,41 +73,51 @@ def transport_from_addr(
'''
try:
return _addr_to_transport[type(addr)]
addr_type = type(addr)
return _addr_to_transport[addr_type]
except KeyError:
raise NotImplementedError(
f'No known transport for address {repr(addr)}'
f'No known transport for address '
f'{addr!r}'
)
def transport_from_stream(
stream: trio.abc.Stream,
codec_key: str = 'msgpack'
codec_key: str = 'msgpack',
) -> Type[MsgTransport]:
'''
Given an arbitrary `trio.abc.Stream` and a desired codec,
find the corresponding `MsgTransport` type.
'''
transport = None
transport: str|None = None
if isinstance(stream, trio.SocketStream):
sock: socket.socket = stream.socket
match sock.family:
case socket.AF_INET | socket.AF_INET6:
transport = 'tcp'
case socket.AF_UNIX:
# `HAS_UDS` short-circuits before `socket.AF_UNIX` on
# hosts where that constant is absent.
case fam if (
HAS_UDS
and
fam == socket.AF_UNIX
):
transport = 'uds'
case _:
case fam:
raise NotImplementedError(
f'Unsupported socket family: {sock.family}'
f'Unsupported socket family: {fam}'
)
if not transport:
raise NotImplementedError(
f'Could not figure out transport type for stream type {type(stream)}'
f'Could not figure out transport type for stream type '
f'{type(stream)}'
)
key = (codec_key, transport)

View File

@ -21,17 +21,29 @@ from __future__ import annotations
from contextlib import (
contextmanager as cm,
)
import hashlib
from pathlib import Path
import os
import sys
from socket import (
AF_UNIX,
SOCK_STREAM,
SOL_SOCKET,
error as socket_error,
)
# NOTE, `AF_UNIX` is absent on Windows / any CPython built without
# unix-domain-socket support. Keep this module importable
# everywhere (so `UDSAddress` stays referenceable for type and
# `isinstance()` checks plus registry lookups); the `AF_UNIX`-using
# code paths below are runtime-only and are never reached when the
# UDS backend is unusable (gated on `trio`'s `has_unix`, see
# `HAS_UDS`).
try:
from socket import AF_UNIX
except ImportError:
AF_UNIX = None
import struct
from typing import (
Any,
Type,
TYPE_CHECKING,
ClassVar,
@ -49,7 +61,6 @@ from trio._highlevel_open_unix_stream import (
has_unix,
)
from multiaddr import Multiaddr
from tractor.msg import MsgCodec
from tractor.log import get_logger
from tractor.discovery._multiaddr import mk_maddr
@ -63,7 +74,13 @@ from tractor.runtime._state import (
)
if TYPE_CHECKING:
# ONLY type-annots, the eager import costs
# `import tractor` wall-time (gh #470).
from multiaddr import Multiaddr
from tractor.runtime._runtime import Actor
else:
Multiaddr = Any
Actor = Any
# Platform-specific credential passing constants
@ -90,6 +107,22 @@ else:
log = get_logger()
_SUN_PATH_LIMIT: int = (
108
if sys.platform == 'linux'
else 104
)
# single source of truth for whether the UDS backend is usable on this
# host. Windows can expose `AF_UNIX`, but this backend remains
# POSIX-only until its credential and lifecycle paths are supported.
HAS_UDS: bool = (
sys.platform != 'win32'
and
has_unix
)
def unwrap_sockpath(
sockpath: Path,
@ -190,7 +223,7 @@ class UDSAddress(
err_on_no_runtime=False,
)
if actor:
sockname: str = f'{actor.aid.name}@{pid}'
sockname: str = actor.aid.name
# XXX, orig version which broke both macOS (file-name
# length) and `multiaddrs` ('::' invalid separator).
# sockname: str = '::'.join(actor.uid) + f'@{pid}'
@ -216,15 +249,72 @@ class UDSAddress(
# `(?P<name>.+)@(?P<pid>\d+)\.sock` regex, and the
# `spawn._reap` `{name}@{pid}.sock` reconstruction.
token: str = uuid4().hex[:8]
sockname: str = f'{prefix}.{token}@{pid}'
sockname = f'{prefix}.{token}'
sockpath: Path = Path(f'{sockname}.sock')
sockpath: Path = cls.get_sockname(
name=sockname,
pid=pid,
bindspace=filedir,
)
return UDSAddress(
filedir=filedir,
filename=sockpath,
maybe_pid=pid,
)
@classmethod
def get_sockname(
cls,
name: str,
pid: int,
bindspace: Path,
) -> Path:
'''
Build a safe, deterministic UDS socket filename.
'''
suffix: str = f'@{pid}.sock'
filename: str = f'{name}{suffix}'
unsafe: bool = (
'\0' in name
or
'/' in name
or
bool(os.altsep and os.altsep in name)
or
Path(filename).is_absolute()
)
too_long: bool = (
len(os.fsencode(bindspace / filename))
>= _SUN_PATH_LIMIT
)
if (
unsafe
or
too_long
):
digest: str = hashlib.blake2s(
os.fsencode(name),
digest_size=16,
).hexdigest()
filename = f'actor.{digest}{suffix}'
sockpath: Path = bindspace / filename
path_nbytes: int = len(os.fsencode(sockpath))
if path_nbytes >= _SUN_PATH_LIMIT:
raise ValueError(
f'UDS bindspace leaves no room for an AF_UNIX '
f'socket filename!\n'
f'bindspace: {bindspace}\n'
f'name was unsafe: {unsafe}\n'
f'name was over budget: {too_long}\n'
f'compacted filename: {filename}\n'
f'encoded path bytes: {path_nbytes}\n'
f'AF_UNIX path limit: {_SUN_PATH_LIMIT}\n'
)
return Path(filename)
@classmethod
def get_root(cls) -> UDSAddress:
def_uds_filename: Path = 'registry@1616.sock'
@ -301,7 +391,16 @@ async def start_listener(
f'>{{\n'
f'|_{bs!r}\n'
)
bs.mkdir()
bs.mkdir(
# ensure the full ancestor tree for any nested
# (custom `filedir`) bindspace; the default
# `get_rt_dir()` space is pre-created but a custom
# one may have missing parents.
parents=True,
# avoid `FileExistsError` from racing actors, same
# guard as in `get_rt_dir()`.
exist_ok=True,
)
with _reraise_as_connerr(
src_excs=(
@ -554,11 +653,18 @@ class MsgpackUDSStream(MsgpackTransport):
]:
sock: trio.socket.socket = stream.socket
# NOTE XXX, it's unclear why one or the other ends up being
# `bytes` versus the socket-file-path, i presume it's
# something to do with who is the server (called `.listen()`)?
# maybe could be better implemented using another info-query
# on the socket like,
# NOTE, the `bytes` case is a linux-only artifact: setting
# `SO_PASSCRED` (see `open_unix_socket_w_passcred()`)
# causes the kernel to *autobind* the un-named client
# sock to an abstract-namespace addr which python
# delivers as `bytes`; the listener-bound end is always
# the fs-path `str`. On platforms WITHOUT autobind
# (macOS et al) the un-bound end instead reports as an
# empty `str` so BOTH names arrive as `str`s and the
# real fs-path is whichever is non-empty: `peername` on
# the connect side, `sockname` on the accept side.
#
# for socket-api deats see,
# https://beej.us/guide/bgnet/html/split-wide/system-calls-or-bust.html#gethostnamewho-am-i
sockname: str|bytes = sock.getsockname()
# https://beej.us/guide/bgnet/html/split-wide/system-calls-or-bust.html#getpeernamewho-are-you
@ -570,8 +676,24 @@ class MsgpackUDSStream(MsgpackTransport):
case (bytes(), str()):
sock_path: Path = Path(sockname)
case (str(), str()): # XXX, likely macOS
sock_path: Path = Path(peername)
# NOTE, no-autobind case (macOS): the un-bound end
# is `''`, NOT a `bytes` abstract-ns addr; taking
# `peername` unconditionally (as prior impl did)
# delivers garbage `Path('')` addrs on the accept
# side!
case (str(), str()):
bound_name: str = (
peername
or
sockname
)
if not bound_name:
raise ValueError(
f'Empty UDS (peername, sockname) pair ??\n'
f'peername: {peername!r}\n'
f'sockname: {sockname!r}\n'
)
sock_path: Path = Path(bound_name)
case _:
raise TypeError(

View File

@ -26,11 +26,6 @@ built on `tractor`.
'''
from collections.abc import Mapping
from functools import partial
from inspect import (
FrameInfo,
getmodule,
stack,
)
import sys
import logging
from logging import (
@ -38,10 +33,16 @@ from logging import (
Logger,
StreamHandler,
)
from types import ModuleType
from types import (
FrameType,
ModuleType,
)
import warnings
import colorlog # type: ignore
# NOTE, `colorlog` is lazy-imported in
# `get_console_log()` to keep it off the eager
# `import tractor` path (gh #470).
#
# ?TODO, some other (modern) alt libs?
# import coloredlogs
# import colored_traceback.auto # ?TODO, need better config?
@ -436,16 +437,40 @@ def get_logger(
pkg_name: str = _root_name
def get_caller_mod(
frames_up:int = 2
):
frames_up: int = 2,
) -> ModuleType|None:
'''
Attempt to get the module which called `tractor.get_logger()`.
Attempt to get the module which called
`tractor.get_logger()`.
Resolve the caller's frame with `sys._getframe()` and
map its `__name__` through `sys.modules`; `inspect.stack()`
(the previous impl) builds src-file info for EVERY frame
on the stack, scanning all of `sys.modules` per frame via
`inspect.getmodule()`, which made module-level
`get_logger()` calls dominate `import tractor` time
(see gh #470).
'''
callstack: list[FrameInfo] = stack()
caller_fi: FrameInfo = callstack[frames_up]
caller_mod: ModuleType = getmodule(caller_fi.frame)
return caller_mod
try:
caller_frame: FrameType = sys._getframe(frames_up)
except ValueError:
return None
mod_name: str|None = caller_frame.f_globals.get(
'__name__',
)
if mod_name is None:
return None
if caller_mod := sys.modules.get(mod_name):
return caller_mod
# Preserve caller discovery for `runpy`, plugin loaders,
# and `exec()` namespaces not registered in `sys.modules`.
# Import `inspect` only on this rare fallback path.
from inspect import getmodule
return getmodule(caller_frame)
# --- Auto--naming-CASE ---
# -------------------------
@ -780,6 +805,10 @@ def get_console_log(
None,
)
):
# lazy-imported to keep it off the eager
# `import tractor` path (gh #470).
import colorlog # type: ignore
fmt: str = LOG_FORMAT # always apply our format?
handler = StreamHandler()
formatter = colorlog.ColoredFormatter(

View File

@ -144,6 +144,8 @@ class Aid(
name: str
uuid: str
pid: int|None = None
is_registrar: bool|None = None
is_probe: bool = False
# TODO? can/should we extend this field set?
# -[ ] use built-in support for UUIDs? `uuid.UUID` which has

View File

@ -94,6 +94,33 @@ if TYPE_CHECKING:
log = get_logger('tractor')
def _register_rpc_task(
actor: Actor,
chan: Channel,
func: Callable,
is_rpc: bool,
task_status: TaskStatus[
Context | BaseException
],
ctx: Context,
) -> None:
'''
Register an RPC task before publishing it to `Nursery.start()`.
'''
if is_rpc:
if not actor._rpc_tasks:
actor._ongoing_rpc_tasks = trio.Event()
actor._rpc_tasks[(chan, ctx.cid)] = (
ctx,
func,
trio.Event(),
)
task_status.started(ctx)
# ?TODO? move to a `tractor.lowlevel._rpc` with the below
# func-type-cases implemented "on top of" `@context` defs:
# -[ ] std async func helper decorated with `@rpc_func`?
@ -142,7 +169,14 @@ async def _invoke_non_context(
# is propagated!
with cancel_scope as cs:
ctx._scope = cs
task_status.started(ctx)
_register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
async with aclosing(coro) as agen:
async for item in agen:
# TODO: can we send values back in here?
@ -178,7 +212,14 @@ async def _invoke_non_context(
)
with cancel_scope as cs:
ctx._scope = cs
task_status.started(ctx)
_register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
await coro
if not cs.cancelled_caught:
@ -202,23 +243,29 @@ async def _invoke_non_context(
)
await chan.send(ack)
except (
TransportClosed,
trio.ClosedResourceError,
trio.BrokenResourceError,
BrokenPipeError,
) as ipc_err:
failed_resp = True
if is_rpc:
raise ipc_err
else:
log.exception(
f'Failed to ack runtime RPC request\n\n'
f'{func} x=> {ctx.chan}\n\n'
f'{ack}\n'
)
log.warning(
f'Failed to ack runtime RPC request\n\n'
f'{func} x=> {ctx.chan}\n\n'
f'{ack}\n'
f' |_{ipc_err!r}\n'
)
with cancel_scope as cs:
ctx._scope: CancelScope = cs
task_status.started(ctx)
_register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
result = await coro
fname: str = func.__name__
@ -247,6 +294,7 @@ async def _invoke_non_context(
)
await chan.send(ret_msg)
except (
TransportClosed,
BrokenPipeError,
trio.BrokenResourceError,
):
@ -673,7 +721,14 @@ async def _invoke(
):
ctx._scope_nursery = tn
rpc_ctx_cs = ctx._scope = tn.cancel_scope
task_status.started(ctx)
_register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
# invoke user endpoint fn.
res: Any|PayloadT = await coro
@ -920,6 +975,7 @@ async def try_ship_error_to_remote(
# downward should be mostly wrapping such cases in a
# tpt-closed; the `.critical()` usage is warranted.
except (
TransportClosed,
trio.ClosedResourceError,
trio.BrokenResourceError,
BrokenPipeError,
@ -1221,18 +1277,6 @@ async def process_messages(
)
continue
else:
# mark our global state with ongoing rpc tasks
actor._ongoing_rpc_tasks = trio.Event()
# store cancel scope such that the rpc task can be
# cancelled gracefully if requested
actor._rpc_tasks[(chan, cid)] = (
ctx,
func,
trio.Event(),
)
# XXX RUNTIME-SCOPED! remote (likely internal) error
# (^- bc no `Error.cid` -^)
#

View File

@ -259,6 +259,7 @@ class Actor:
name=name,
uuid=uuid,
pid=os.getpid(),
is_registrar=self.is_registrar,
)
self._task: trio.Task|None = None

View File

@ -22,7 +22,10 @@ from __future__ import annotations
from contextvars import (
ContextVar,
)
import os
from pathlib import Path
import stat
import sys
from typing import (
Any,
Callable,
@ -30,7 +33,6 @@ from typing import (
TYPE_CHECKING,
)
import platformdirs
from trio.lowlevel import current_task
from msgspec import (
@ -43,6 +45,9 @@ if TYPE_CHECKING:
from .._context import Context
_DARWIN_TMPDIR: Path = Path('/tmp')
# default IPC transport protocol settings
TransportProtocolKey = Literal[
'tcp',
@ -318,6 +323,60 @@ def current_ipc_ctx(
def _ensure_owner_only_posix_dir(
path: Path,
*,
parents: bool = False,
) -> None:
'''
Create or validate a UID-owned POSIX runtime directory.
Pre-existing directories are accepted only when owned by the
current user. Their mode is normalized to `0o700` because runtime
directories hold IPC sockets and are private bindspaces.
'''
# TODO: https://github.com/goodboy/tractor/issues/494
# Research having the actor-tree root process choose and create
# this bindspace, then propagate it to every subactor. On
# Linux, a private mount namespace could isolate it while letting
# spawned subactors inherit access; independently launched
# discovery clients would need an explicit join or fallback path.
# POSIX metadata alone records UID/GID ownership, so other systems
# still need explicit runtime metadata and lifecycle management.
try:
dir_stat: os.stat_result = path.lstat()
except FileNotFoundError:
try:
path.mkdir(
mode=0o700,
parents=parents,
)
except FileExistsError:
pass
dir_stat = path.lstat()
if (
not stat.S_ISDIR(dir_stat.st_mode)
or
dir_stat.st_uid != os.getuid()
):
platform_name: str = (
'Darwin'
if sys.platform == 'darwin'
else 'POSIX'
)
raise PermissionError(
f'Unsafe {platform_name} runtime directory!\n'
f'path: {path}\n'
f'owner uid: {dir_stat.st_uid}\n'
f'mode: {stat.filemode(dir_stat.st_mode)}\n'
)
if stat.S_IMODE(dir_stat.st_mode) != 0o700:
path.chmod(0o700)
def get_rt_dir(
subdir: str|Path|None = None,
appname: str = 'tractor',
@ -327,20 +386,35 @@ def get_rt_dir(
userspace apps stick their IPC and cache related system
util-files.
On linux we use a `${XDG_RUNTIME_DIR}/tractor/` subdir by
default, but equivalents are mapped for each platform using
the lovely `platformdirs` lib.
Linux uses an owner-only `${XDG_RUNTIME_DIR}/tractor/`; Darwin
uses a short, owner-only `/tmp/tractor-<uid>` path; other
platforms use the lovely `platformdirs` lib.
'''
rt_dir: Path = Path(
platformdirs.user_runtime_dir(
appname=appname,
),
)
# lazy-imported to keep it off the eager
# `import tractor` path (gh #470).
import platformdirs
rt_root: Path|None = None
if sys.platform == 'darwin':
# Darwin's AF_UNIX path limit is 104 bytes. The standard
# platformdirs path can consume that before the sock name.
rt_root = (
_DARWIN_TMPDIR
/ f'{appname}-{os.getuid()}'
)
rt_dir: Path = rt_root
else:
rt_dir = Path(
platformdirs.user_runtime_dir(
appname=appname,
),
)
# Normalize and validate that `subdir` is a relative path
# without any parent-directory ("..") components, to prevent
# escaping the runtime directory.
subdir_path: Path|None = None
if subdir:
subdir_path = (
subdir
@ -358,13 +432,28 @@ def get_rt_dir(
f'{subdir!r}\n'
)
rt_dir: Path = rt_dir / subdir_path
if os.name != 'posix':
if subdir_path is not None:
rt_dir = rt_dir / subdir_path
if not rt_dir.is_dir():
rt_dir.mkdir(
# Runtime dirs hold IPC sockets; owner-only access
# prevents other users from traversing the bindspace.
mode=0o700,
parents=True,
exist_ok=True,
)
return rt_dir
if not rt_dir.is_dir():
rt_dir.mkdir(
parents=True,
exist_ok=True, # avoid `FileExistsError` from conc calls
)
_ensure_owner_only_posix_dir(
rt_dir,
parents=(rt_root is None),
)
if subdir_path is not None:
for part in subdir_path.parts:
rt_dir = rt_dir / part
_ensure_owner_only_posix_dir(rt_dir)
return rt_dir

View File

@ -38,7 +38,11 @@ from ..devx import (
pformat,
)
# from ..msg import pretty_struct
from ..to_asyncio import run_as_asyncio_guest
#
# NOTE, `.to_asyncio` (and thus `asyncio` itself) is
# lazy-imported at the `infect_asyncio=True` use-sites
# below to keep it off the eager `import tractor` path
# (gh #470).
from ..discovery._addr import UnwrappedAddress
from ..runtime._runtime import (
async_main,
@ -96,6 +100,7 @@ def _mp_main(
)
try:
if infect_asyncio:
from ..to_asyncio import run_as_asyncio_guest
actor._infected_aio = True
run_as_asyncio_guest(trio_main)
else:
@ -156,6 +161,7 @@ def _trio_main(
)
try:
if infect_asyncio:
from ..to_asyncio import run_as_asyncio_guest
actor._infected_aio = True
run_as_asyncio_guest(trio_main)
else:

View File

@ -35,14 +35,14 @@ Future-work TODO — authoritative UDS bind-addr tracking
`unlink_uds_bind_addrs()` currently has two cleanup paths:
1. Explicit `bind_addrs` (when parent set them at spawn time)
2. **Convention-based reconstruction**
`<XDG_RUNTIME_DIR>/tractor/<name>@<pid>.sock` for the
2. **Convention-based reconstruction** in the platform default UDS
bindspace for the
common case where the subactor self-assigned a random sock
via `UDSAddress.get_random()`.
Path (2) hardcodes the `<name>@<pid>.sock` convention from
`tractor.ipc._uds.UDSAddress`. If that convention ever
changes or the subactor binds to a non-default
Path (2) delegates filename reconstruction to
`tractor.ipc._uds.UDSAddress.get_sockname()`. If the subactor binds to
a non-default
`bindspace`/`filedir` we'll silently fail to unlink.
A more authoritative approach would be:
@ -71,6 +71,7 @@ fd leak. Different bug class but same broader theme of
from __future__ import annotations
import os
from pathlib import Path
from typing import TYPE_CHECKING
import trio
@ -104,7 +105,7 @@ def unlink_uds_bind_addrs(
`_serve_ipc_eps` `finally:` block (which normally calls
`os.unlink(addr.sockpath)`) never runs. Without this
parent-side cleanup, the dead subactor's
`${XDG_RUNTIME_DIR}/tractor/<name>@<pid>.sock` file
platform-default UDS socket file
accumulates on the filesystem (see issue #454 + the
autouse `_track_orphaned_uds_per_test` fixture).
@ -118,7 +119,7 @@ def unlink_uds_bind_addrs(
picked its own random sock via
`UDSAddress.get_random()`), reconstruct the path
from `(subactor.aid.name, proc.pid)` using the
same `<name>@<pid>.sock` convention. We can do this
same `UDSAddress.get_sockname()` helper. We can do this
because the subactor uses its OWN `os.getpid()` at
bind time, which equals `proc.pid` from the
parent's view.
@ -154,7 +155,21 @@ def unlink_uds_bind_addrs(
and subactor is not None
and proc.pid is not None
):
sockname: str = f'{subactor.aid.name}@{proc.pid}.sock'
try:
sockname: Path = UDSAddress.get_sockname(
name=subactor.aid.name,
pid=proc.pid,
bindspace=UDSAddress.def_bindspace,
)
except Exception:
log.exception(
f'Failed to reconstruct UDS sock-file for '
f'post-kill cleanup — skipping\n'
f' |_{proc}\n'
f' |_{subactor.aid}\n'
)
return
sockpath: str = str(
UDSAddress.def_bindspace / sockname
)