Compare commits

..

65 Commits

Author SHA1 Message Date
Gud Boi 4151b9569a Add `to_actor` one-shot parallelism example
Demo both flavors of the new API in a runnable script
(auto-collected by `test_docs_examples.py`),

- the fully-implicit one-shot which boots (and tears down) the
  actor-runtime around a single `to_actor.run()` call,
- the concurrent "worker-pool-ish" prime-check pattern: a local
  `trio` task nursery scheduling one-shots against a shared
  caller-managed `an`, mirroring (in miniature) the neighboring
  `concurrent_actors_primes.py` example per issue #477.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:08 -04:00
Gud Boi 8d04af0d73 Add `tests/test_to_actor.py` one-shot API suite
Cover every placement variant + failure mode of the new
`to_actor.run()`,

- private-nursery one-shot + implicit runtime boot via pass-through
  `runtime_kwargs`,
- remote-error relay to the caller's task (bare and inside a
  caller-managed `an`) as boxed `RemoteActorError`s,
- caller-nursery spawn + portal-reuse w/o implicit reap,
- the concurrent "worker-pool-ish" pattern: a local `trio` task
  nursery scheduling one-shots against a shared `an`,
- the 4 pre-spawn validation rejections (sync fn, async-gen fn,
  `portal`+`an` combo, `runtime_kwargs`+placement combo).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:08 -04:00
Gud Boi a3cd448612 Add `tractor.to_actor` one-shot task API subpkg
First cut at the `to_thread`/`to_process`-style "run it over there"
wrapper layer from issue #477: a single-remote-task invocation API
decoupled from the `ActorNursery` spawn machinery, composed purely
from the lower level daemon-actor + portal primitives,

- `to_actor.run(fn, **fn_kwargs)` spawns a subactor via
  `ActorNursery.start_actor()`, schedules `fn` as its lone task
  with `Portal.run()` and ALWAYS reaps it via a `finally`-scoped
  `Portal.cancel_actor()` (whose bounded cancel-req wait is
  internally shielded so the reap also runs under caller-scope
  cancellation).
- remote errors raise directly in the caller's task as boxed
  `RemoteActorError`s, moving error collection/propagation up into
  whatever local `trio` scope encloses the call.
- "placement" opts: `portal=` reuses a running actor (no
  spawn/reap), `an=` spawns from a caller-managed actor-nursery,
  neither opens a private call-scoped `open_nursery()` (implicitly
  booting the runtime, tunable via pass-through `runtime_kwargs`).
- fail-fast validation BEFORE any spawn: non-streaming async fn
  only (same constraint as `Portal.run()`), `portal=`/`an=` mutual
  exclusion and no `runtime_kwargs` alongside a placement opt.

Also,
- x-ref the successor API from `.run_in_actor()`'s deprecation TODO
  + docstring; emitting a formal `DeprecationWarning` waits on
  migrating in-repo usage.
- log prompt-io provenance per NLNet policy incl. the driver prompt
  file.

Prompt-IO: ai/prompt-io/claude/20260702T154255Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-17 22:06:08 -04:00
Bd 4e27dcde48
Merge pull request #475 from goodboy/windows_support_round2
Restore Windows support (optional UDS + `SIGUSR1`)
2026-08-17 20:05:31 -04:00
Gud Boi 93322405a1 Restore IPC transport style conventions
Bring the transport helpers back in line with project style:

- restore single-quote strings and docstrings;
- drop the oversized helper divider and simplify comments;
- keep guarded `match` dispatch so a missing `socket.AF_UNIX`
  remains safe when `HAS_UDS` is false.

Also, format the `HAS_UDS` conjunction with the project's
multiline branch convention.

Review: PR #475 (goodboy)
https://github.com/goodboy/tractor/pull/475

Prompt-IO: ai/prompt-io/opencode/20260817T231825Z_359fe75c_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-17 19:37:44 -04:00
Gud Boi 359fe75ced Tolerate only Windows `pytest` failures
Job-level `continue-on-error` made setup and import-smoke failures
non-blocking even though the smoke is the hard Windows support
signal.

Move tolerance to the `pytest` step so the incomplete suite remains
informational while install, dependency, and `HAS_UDS` smoke
failures fail the job. When only the known suite fails, the job and
PR rollup remain green with the failed step still visible.

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-17 17:04:20 -04:00
Gud Boi 40be587ce2 Keep `HAS_UDS` false on Windows
Modern Windows Python can expose `AF_UNIX`, so `trio.has_unix`
alone can register the UDS backend even though its credential and
lifecycle paths remain POSIX-only.

Gate `HAS_UDS` on `sys.platform` and make the Windows import smoke
step assert the TCP-only capability contract.

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-17 12:09:52 -04:00
Gud Boi ca4582b003 Skip Windows-hanging shm IPC test
`test_parent_writer_child_reader` deadlocks on Windows (the
parent/child shared-mem transfer hangs at the larger frame size),
so the `windows-latest` CI leg ran to the 16-min job cap instead
of completing. It's a genuine nascent-Windows shm bug, not a
clean "unsupported", so it's `skipif`'d (not removed) and tracked
under #404; `test_child_attaches_alot` still runs on Windows.

- `@pytest.mark.skipif(platform.system() == 'Windows', ...)` on
  the parametrized `test_parent_writer_child_reader` so the leg
  completes + reports. linux/macOS unaffected (all 6 variants
  still run).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:17:11 -04:00
Gud Boi 75385e448d Derive transport registries from one list
The five transport/address lookup maps each hand-guarded `uds`
with its own `if HAS_UDS:` block (3 in `ipc/_types`, 2 in
`discovery/_addr`) — easy to let drift so a backend half-registers
(known by address but not by key, listed but unroutable, &c).

- build one `_msg_transports` / `_address_protos` list per module
  (TCP always, UDS only when `HAS_UDS`), then DERIVE every map
  from it via each backend's ClassVars (`codec_key`,
  `address_type`, `proto_key`). Adding a backend now touches one
  list, and the maps can't disagree.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:17:09 -04:00
Gud Boi 34d81adef3 Skip `infect_asyncio` tests on Windows
`tractor`'s infect-asyncio mode runs an `asyncio` loop under
`trio` guest-mode; on Windows the default `ProactorEventLoop` is
incompatible and the suite hangs/crashes mid-run (orphaned py
procs), so the `windows-latest` CI leg never finishes reporting.

- add a module-level `pytest.skip(allow_module_level=True)` gated
  on `platform.system() == 'Windows'` to `test_infected_asyncio`
  and `test_root_infect_asyncio`, before their asyncio-interop
  imports. macOS/linux are unaffected (they run these fine).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:02 -04:00
Gud Boi 1691be96fd Scale `test_lifetime_stack` deadline for CI
`test_lifetime_stack_wipes_tmpfile` guards spawn+teardown with a
hard-coded `trio.move_on_after()` (1.6s / 1s) that isn't scaled
for slow CI. On a noisy macOS runner the `error_in_child=True`
case times out before the child error propagates, so the scope
cancels and `assert not cs.cancel_called` flips — reddening the
(required) macOS leg. Same unscaled-deadline class `main` already
fixed for `test_dynamic_pub_sub`.

- multiply the budget by `cpu_perf_headroom()` (`tests/conftest`),
  the established deadline-headroom helper (3x on macOS CI, a
  1.0 no-op locally / on un-throttled linux).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:02 -04:00
Gud Boi 36ad1f3dd0 Skip `test_ringbuf` at collection off-linux
`tests/test_ringbuf.py` imports `tractor.ipc._ringbuf` at module
top, which pulls in `tractor.ipc._linux` whose module-level
`ffi.dlopen(None)` raises `OSError` on Windows (and any non-linux
host). That fires at COLLECTION, before the module's existing
`pytestmark = pytest.mark.skip` can apply, so it aborts the whole
pytest session — the new `windows-latest` CI leg never gets past
collection.

- guard the module with `pytest.skip(allow_module_level=True)`
  gated on `platform.system() != 'Linux'`, placed before the
  crashing import — same idiom as `tests/devx/test_debugger.py`.
- the `eventfd`-based ringbuf backend is linux-only by design, so
  macOS skips cleanly too (previously it only skipped incidentally
  via the absent `cffi` optional dep).

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:02 -04:00
Gud Boi 72441124c0 Fix Windows `import tractor` at the UDS root
The prior round gated UDS in four modules but `import tractor`
still crashed on Windows: `tractor.ipc._uds` does `from socket
import AF_UNIX` at module top, and several modules in the import
graph (`discovery._api`, `spawn._reap`, `discovery._multiaddr`,
`_testing.addr`) import `_uds` unconditionally. Instead of
guarding every importer, fix the root and collapse the per-module
probes to one capability flag.

- in `ipc/_uds.py`, guard the lone `AF_UNIX` import so the module
  stays importable everywhere; expose `HAS_UDS = trio.has_unix`
  as the single source of truth (the same predicate that gates
  `trio.open_unix_socket()`).
- `ipc/_types.py`, `discovery/_addr.py` and `ipc/_server.py` now
  import `UDSAddress`/`MsgpackUDSStream`/`HAS_UDS` directly and
  gate the transport + address registries on `HAS_UDS`; drop the
  duplicated `getattr(socket,'AF_UNIX')` / `platform.system()`
  probes, the dead `HAS_AF_UNIX` conjunct, and the import-time
  `log.warning()` spam.
- `devx/_stackscope.py` `enable_stack_on_sig()` early-returns
  when `sig is None`, so a missing `SIGUSR1` (Windows) degrades
  to a no-op instead of a `TypeError` from `getsignal()` /
  `signal()`.
- add a `windows-latest` CI leg (UDS excluded; informational via
  `continue-on-error` while support matures) plus an `import
  tractor` smoke step as the hard signal for the import fix.

Because `_uds` is importable everywhere `UDSAddress` stays a real
class, so `isinstance()` checks and `wrap_address()` no longer
`AttributeError` on no-UDS hosts; actual socket use stays gated
on `has_unix`.

Review: https://github.com/goodboy/tractor/pull/475
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 23:12:01 -04:00
Gud Boi 0cbb850645 Make UDS + `SIGUSR1` optional for Windows
Windows (and any CPython that doesn't expose `socket.AF_UNIX`)
can't import the UDS transport backend nor `signal.SIGUSR1`, so
the unconditional imports break `import tractor` outright on
those hosts. Guard the platform-specific bits behind capability
probes and fall back to a TCP-only runtime when the UDS backend
is unavailable.

- across `discovery/_addr.py`, `ipc/_server.py` and
  `ipc/_types.py`, gate on `getattr(socket, 'AF_UNIX', None)` +
  `platform.system()` and import `UDSAddress` /
  `MsgpackUDSStream` only when supported, leaving the names as
  `None` otherwise.
- register the `'uds'` key in `_address_types`, its
  default-loopback addr, and the transport lookup maps only when
  the backend actually loads, so TCP keeps working standalone.
- in `devx/_stackscope.py`, import `SIGUSR1` conditionally and
  set it to `None` on Windows.

Rebased onto the post-reorg tree where `_addr.py` now lives
under `tractor/discovery/`; adapt the relocated imports to the
package's `..ipc._uds` / `..ipc._tcp` paths (the original
single-dot paths would silently disable UDS on POSIX) and drop a
duplicated `TYPE_CHECKING` block and dead `import logging` left
by the move.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-14 22:35:40 -04:00
Bd 8f0df0ff6f
Merge pull request #480 from goodboy/wkt/uds_macos_473
Fix `uds` addr corruption + exercise macOS CI
2026-08-14 19:14:02 -04:00
Gud Boi f75c9cdeab Traverse `BaseExceptionGroup` peer-close errors
`_peer_closed_errno()` followed cause and context links but did not
descend through grouped exceptions. A reset below a group could
therefore escape as `trio.BrokenResourceError` instead of the
normalized `TransportClosed` boundary.

Walk the exception tree with cycle protection, requiring every
group branch to represent peer closure before normalization. Extend
the `MsgpackTransport.send()` regression to prove all-transport and
mixed-failure behavior.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 18:54:56 -04:00
Gud Boi 75cda1933c Move registrar probes to `discovery._api`
Registrar election and multi-address probing are discovery-protocol
concerns, but their implementation lived in root-runtime ignition.
Move the bounded handshake probe and its concurrent address
classifier into `discovery._api`, leaving `open_root_actor()` to
consume the classified results.

Update probe tests for the canonical module and clarify that the
daemon-fixture regressions directly exercise their sibling
`conftest` plugin.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 18:43:36 -04:00
Gud Boi f454cefe56 Enforce `get_rt_dir()` ownership on POSIX
Linux previously accepted a pre-existing runtime bindspace without
checking its owner or mode, even though Darwin enforced both. Share
the POSIX directory guard so every managed root and subdir rejects
non-directories and foreign UIDs before changing permissions.

Normalize owner-controlled bindspaces to `0o700` and add Linux
regressions for mode repair, foreign ownership, and non-directory
paths.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 18:40:23 -04:00
Gud Boi d0cc06815f Accept raced `TransportClosed` diagnostics
The `@context` debugger E2E intentionally closes its channel.
Teardown can surface from either local error shipment or the peer
receive task.
The RPC response fix makes local close win under CI, while the test
required both scheduler-dependent diagnostics.

Keep the common debugger and cancellation assertions, then accept
either transport-close report. This preserves real actor-tree
teardown coverage without depending on task scheduling order.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 12:49:40 -04:00
Gud Boi ebf2258b4f Keep accepted RPCs alive on `TransportClosed`
An RPC caller can close its channel after the callee creates the
endpoint coro but before its `StartAck` or final response lands.
Treat those response-send failures as terminal delivery failures so
the accepted endpoint still runs and application errors stay local
instead of cancelling the shared service nursery.

Register each cancellable RPC in `Actor._rpc_tasks` before
publishing its `Context` through `TaskStatus.started()`.
This closes the checkpoint-free completion race without a
cross-task handoff and keeps `Actor._ongoing_rpc_tasks` balanced
through existing cleanup.

Add regressions for caller disconnects at `StartAck` and error
shipment, plus registration-before-execution and cleanup checks.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 12:28:56 -04:00
Gud Boi ac61d7a5bf Use a short UDS reaper test bindspace
Darwin's pytest `tmp_path` can already exceed the 104-byte AF_UNIX
budget before appending either synthetic socket filename. That made
the new sentinel-policy regression fail identically in both macOS
matrix legs without exercising reaper behavior.

Allocate the test's real socket files under a short
`/tmp/tractor-reap-*` directory and retain scoped cleanup.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:59:07 -04:00
Gud Boi 584ea4e9ad Document stable macOS UDS operation
The remediation grew beyond the original no-autobind fix, leaving
public docs and nearby comments describing connect-only discovery,
XDG-only socket paths, raw readiness probes, and old cleanup naming.

Document typed registrar probing, occupied-address rejection, Darwin's
short runtime directory, platform-aware socket cleanup, sentinel
readiness, and subprocess output draining. Add the GH #473 bugfix
fragment and update regression rationale without changing task states.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:37:19 -04:00
Gud Boi 16dd876b0c Make UDS reaping platform-aware
Moving Darwin sockets to `/tmp/tractor-<uid>` left pytest and the
standalone reaper searching only `XDG_RUNTIME_DIR`. Enabling that
shared bindspace naively would also let automatic session cleanup
unlink an independent live `registry@1616.sock`.

Resolve the runtime's actual default UDS bindspace on every platform.
Exclude the pid-less registry sentinel from automatic cleanup while
retaining explicit CLI removal, and document that destructive choice.

Cover bindspace resolution and automatic-vs-explicit sentinel policy.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:36:27 -04:00
Gud Boi 1a9ce915f3 Harden inbound actor handshakes
Registry probes need short retry deadlines, but applying their
one-second budget to every portal and child connection can terminate a
valid delayed actor with no client retry path.

Give ordinary pre-registration handshakes an independent ten-second
deadline. Normalize raw `msgspec.DecodeError` frames to
`TransportClosed` so malformed peers cannot cancel the shared IPC
nursery with decoder internals.

Cover malformed frames and the ordinary-vs-probe timeout distinction.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-14 10:34:59 -04:00
Gud Boi 0b63af020e Retry transient registrar handshakes
A loaded runner can accept a registry transport while delaying its
actor handshake beyond the first one-second attempt. Treating that
timeout as final classifies a healthy daemon as occupied and cascades
into unrelated discovery failures.

Retry connected handshake failures on fresh channels under a shared
three-second budget, while returning immediately for truly absent
listeners. Bound every connect-plus-handshake attempt, use incremental
backoff, and retain the one-second unauthenticated server limit.

Also bound shielded probe-channel cleanup to 200ms and cover the real
timeout, reconnection, backoff, fresh-channel, and stalled-close paths.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 23:40:12 -04:00
Gud Boi 6e424d4696 Use a daemon-ready sentinel in discovery tests
The discovery `daemon` fixture probed UDS readiness by connecting and
immediately closing. That entered Tractor's actor handshake without an
`Aid` payload and destabilized the remote registrar on macOS before
test roots attempted discovery.

Run the child through a small `open_root_actor()` wrapper and publish a
filesystem sentinel only after runtime startup completes. Poll that
sentinel with process-liveness checks and guaranteed setup-failure
cleanup, without touching the transport socket.

Cover transport-free readiness and deterministic polling backoff.

Caught-during: review remediation
Found-via: `/run-tests` discovery daemon fixture consumers
Cause: initial sentinel drafts mishandled `pytest.Testdir`, emitted
  invalid `python -c` syntax, and misplaced return-code logging.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 20:38:22 -04:00
Gud Boi 5724c0516a Probe registrar capability in actor handshakes
Transport connect alone can select a foreign, stalled, or ordinary
actor endpoint as the registry. On macOS UDS this also exercises a
fragile connect-and-bail path before every root election.

Extend `Aid` with backward-compatible probe and registrar capability
fields, require a bounded typed handshake, and classify addresses as
absent, occupied, or confirmed registrars. Reject occupied endpoints
instead of binding over them.

Also,
- close `_connect_chan()` in `finally`
- bypass normal peer tracking for election probes
- preserve legacy registrar handshakes with unknown capability
- cover foreign listeners and idle registrar peer state

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 20:04:48 -04:00
Gud Boi 0d6d7c2a63 Keep UDS post-kill cleanup best-effort
`unlink_uds_bind_addrs()` reconstructs self-assigned socket paths
after a hard kill. An over-budget bindspace can make
`UDSAddress.get_sockname()` raise before the guarded `os.unlink()`,
replacing the original supervision outcome after the child is gone.

Catch and report reconstruction failures, then skip cleanup without
raising. Cover the overflow path and prove no unlink is attempted.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:57:13 -04:00
Gud Boi bd38204fde Preserve docs-example body failures
Best-effort subprocess teardown must not replace the exception raised
by the test body. Suppress cleanup errors only while propagating that
active failure; keep raising teardown errors on normal body exit.

Also close the Windows stdin pipe after its bounded leader reap.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:41:20 -04:00
Gud Boi 340d506940 Bound generated UDS socket paths
Actor names and custom runtime subdirs can otherwise produce unsafe
or overlong pathname sockets after moving Darwin's bindspace to
`/tmp`.

Deats,
- hash unsafe or over-budget actor names while retaining `@pid.sock`
- share deterministic naming with post-kill socket cleanup
- enforce Linux and Darwin `sun_path` byte budgets
- restore non-Darwin dir validation and mode `0700`
- validate every nested Darwin runtime-dir component

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:40:55 -04:00
Gud Boi 64e820e18e Normalize peer resets in `.send()`
A raw UDS readiness client can disconnect before the actor handshake.
Darwin reports the first server write as `ECONNRESET`, wrapped in
`trio.BrokenResourceError`; letting it escape cancels the daemon's
shared IPC nursery and makes later roots elect themselves registrar.

Walk the exception chain for `EPIPE` or `ECONNRESET` and translate
either into the existing `TransportClosed` boundary. Also handle
argument-less resource errors without raising `IndexError`.

Review: PR #480 (goodboy)
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 18:30:44 -04:00
Gud Boi 0a3b0efcc6 Reap timed-out docs example trees
Always drain docs-example pipes, even when the process exited before
the first status check, and decode malformed output with replacement
so diagnostics preserve the original failure.

Run POSIX examples in dedicated sessions and kill the full process
group on timeout. Reap the leader in every exit path, with bounded
Windows cleanup that cannot wait forever on descendant-held pipes.

Cover fast non-zero exits, invalid output bytes, process-group
termination, and post-timeout reaping.

Review: PR #480 (goodboy,copilot-pull-request-reviewer[bot])
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 16:40:05 -04:00
Gud Boi 634d914161 Use short Darwin UDS runtime paths
`platformdirs` can place the runtime dir deep below a pytest temp
home, pushing `registry@1616.sock` past Darwin's 104-byte
`AF_UNIX` limit.

Use a compact `/tmp/<app>-<uid>` root on Darwin and secure it
before allocating socket paths:
- require a real, current-user-owned dir via `lstat()`
- tighten existing roots to mode `0700`
- reject symlinks and unsafe nested dirs

Cover the path budget, mode, and symlink rejection.

Caught-during: review remediation
Found-via: `/run-tests` test_macos_rt_dir_fits_uds_path_limit

Review: PR #480 (goodboy,copilot-pull-request-reviewer[bot])
https://github.com/goodboy/tractor/pull/480

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 16:39:32 -04:00
Gud Boi 451e0acf8a Add prompt-io log for GH #473 UDS-on-macOS work
Provenance entry (+ unedited raw output) for the root-cause
session behind the prior three patches, per the NLNet
generative-AI policy tracked under `ai/prompt-io/`.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Gud Boi 3d02a8569e Run UDS-on-macOS CI leg, un-skip the example
UDS-on-macOS is otherwise un-exercised: the matrix
explicitly excludes the `macos-latest` + `uds` combo, so the
`uds_transport_actor_tree.py` example (skipped on macOS CI
since PR #460) is the only thing that ever touches that
path.

- drop the matrix `exclude` so the full suite runs with
  `--tpt-proto=uds` on `macos-latest`.
- un-skip the example on macOS+CI; with example-stderr
  surfacing in place a still-red run now yields the full
  traceback GH #473 asks for instead of a bare returncode
  assert.

Task-bullets 3 + 4 of GH #473.

Prompt-IO: ai/prompt-io/claude/20260702T155006Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Gud Boi bceca74eb7 Fix UDS addr corruption sans-autobind (macOS)
`MsgpackUDSStream.get_stream_addrs()` matches the
`(peername, sockname)` pair by type to find the listener's
fs-path, but the `(str, str)` arm unconditionally takes
`peername`: on platforms without linux's
`SO_PASSCRED`-triggered autobind (macOS!) the accept side's
`getpeername()` is `''`, so every accepted conn gets garbage
`Path('')` laddr/raddr structs.

Proven on linux by disabling `SO_PASSCRED` (no autobind ->
same `''` shape as darwin): the `uds_transport_actor_tree.py`
example reports `listener sock file: .` pre-fix and the real
registry sockpath post-fix.

- pick the non-empty name in the `(str, str)` arm: `peername`
  on the connect side, `sockname` on the accept side; raise
  `ValueError` on an (unexpected) empty pair.
- document the linux-autobind origin of the `bytes` arms
  which the original impl noted as "unclear".
- `start_listener()`: create the bindspace dir with
  `parents=True, exist_ok=True` (nested custom `filedir`s +
  racing actors).
- example docstring: peer-pid comes via `SO_PEERCRED` on
  linux but `LOCAL_PEERPID` on macOS.

May not be the (only) macOS crasher for GH #473 — it is
non-fatal on the linux sim — but with stderr surfacing now
in place the next macOS CI run pins any remaining layer.

Prompt-IO: ai/prompt-io/claude/20260702T155006Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Gud Boi 18faffcff7 Surface example stderr on any non-zero exit
The docs-example harness only re-raises captured subproc
stderr when the LAST line contains 'Error', but a `tractor`
root-actor crash always ends stderr with the strict-EG
collapse note `( ^^^ this exc was collapsed from a group ^^^ )`
— so every possible crash is swallowed down to a bare
`assert 1 == 0`, exactly what the macOS CI leg shows for the
UDS example in GH #473.

- raise with the FULL stderr (+stdout) whenever the example
  subproc exits non-zero, regardless of stderr shape.
- keep the legacy last-line 'Error' check for zero-rc runs
  which still emit error-ish output.

First task-bullet of GH #473.

Prompt-IO: ai/prompt-io/claude/20260702T155006Z_65bf9df5_prompt_io.md

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-13 15:46:56 -04:00
Bd 7554f90e59
Merge pull request #478 from goodboy/wkt/boot_latency_470
Trim `import tractor` 0.42s -> 0.15s (gh #470)
2026-08-13 15:02:13 -04:00
Gud Boi 3705bbe594 Guard the cold import latency budget
Run seven fresh interpreters and gate the median `import tractor`
time below a conservative 0.35s ceiling. Provide an explicit
environment override for platforms with a different baseline.

Also assert the optional modules deferred by this patch remain
unloaded so a timing pass cannot hide an eager-import regression.

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-13 02:35:07 -04:00
Gud Boi 66ac7863b5 Advertise the lazy `to_asyncio` API
Preserve the existing public wildcard surface while adding the lazy
`to_asyncio` submodule to `__all__` and `dir(tractor)`.

Exercise both APIs in cold interpreters and verify normal package
import still leaves `asyncio` unloaded.

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 20:55:41 -04:00
Gud Boi 089e158da9 Retain dynamic logger caller discovery
Keep the `sys.modules` lookup as the normal fast path, then use
`inspect.getmodule()` for unregistered `runpy`, plugin, and `exec()`
namespaces.

Cover an unregistered module name backed by a real package file and
verify implicit logger naming still resolves to that package.

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 20:52:47 -04:00
Gud Boi 7987f6b8d7 Keep lazy annotations runtime-resolvable
Provide import-free runtime aliases for annotation-only actor and
multiaddr types so `typing.get_type_hints()` remains usable without
eagerly loading optional dependencies.

Correct `_address_types` to its actual `dict` shape and cover the
affected discovery and transport APIs.

Caught-during: review remediation
Found-via: `/run-tests` test_lazy_annotation_names_resolve

Review: PR #478 (goodboy)
https://github.com/goodboy/tractor/pull/478#pullrequestreview-4922213201

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 20:52:15 -04:00
Gud Boi 70c7e334a7 Make sub-spawn strategy toggleable in `we_are_processes`
The #470 boot-latency example hard-coded spawning each `worker_<i>`
subactor concurrently from a bg `trio.Task` (so each child's cold
`import tractor` overlaps). Add a `main()` `spawn_subs_in_bg_tasks`
flag so the serial-spawn path can be demo'd/compared too: flip it
`False` to `start_actor()` each sub inline in the loop before
handing the ready `Portal` to the bg task.

Deats,
- factor an `open_ep(ptl, i)` helper out of `spawn_and_open_ep()` -
  just the `Portal.open_context()` + `wait_for_result()` half, now
  that the spawn step is caller-optional.
- `spawn_and_open_ep()` grows a `maybe_ptl: Portal|None = None`
  param: spawn the subactor itself when unset (bg-task path), OW
  reuse the pre-spawned one (serial path).
- move the "overlap cold imports" rationale comment onto the new
  `main()` param where the toggle now lives.

(this commit msg was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:59:49 -04:00
Gud Boi 62b729a106 Add prompt-io entries for gh #470 latency work
Log the AI-assisted session per the NLNet generative-AI
policy: prompt, profiling findings, per-file diff pointers,
measured results and the unimplemented `pdbp`/`platformdirs`
deferral follow-ups.

(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
Gud Boi 213a298aad Defer `asyncio` import via PEP-562 `.to_asyncio`
`asyncio` (~5ms) only matters for infected-aio actors yet gets
imported by every cold `import tractor` via module-lvl
`.to_asyncio` imports in the debug-REPL + spawn-entry mods.

Deats,
- `devx.debug._trace`/`._tty_lock`: mv `import asyncio` under
  `TYPE_CHECKING` + fn-local it at the two
  `asyncio.current_task()` call-sites; fn-local the
  `run_trio_task_in_future` imports in the infected-aio-only
  branches.
- `spawn._entry`: fn-local `run_as_asyncio_guest` inside the
  `infect_asyncio=True` branches of `_mp_main()`/
  `_trio_main()`.
- `tractor/__init__.py`: add a PEP-562 module `__getattr__`
  lazy-loading `.to_asyncio` on first attr-access so the
  public `tractor.to_asyncio.<attr>` API (e.g.
  `LinkedTaskChannel` annots in
  `test_child_manages_service_nursery.py` + downstream users)
  keeps working unchanged.

Prompt-IO: ai/prompt-io/claude/20260702T155626Z_65bf9df5_prompt_io.md
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
Gud Boi d1d3fc58bb Lazy-import optional 3rd-party deps (gh #470)
Move every import-time-only-by-accident dep off the eager
`import tractor` path so cold child-actor boots only pay for
what they actually use:

- `bidict` -> `TYPE_CHECKING` in `discovery._addr`
  (annotation-only; `_address_types` is a plain `dict`
  literal).
- `multiaddr` -> `TYPE_CHECKING` + fn-local imports in
  `discovery._multiaddr.mk_maddr()`/`parse_maddr()`; also
  `TYPE_CHECKING` the `Multiaddr` annots in `ipc._tcp`/`._uds`
  (adds future-annots to `._multiaddr`).
- `colorlog` -> fn-local in `log.get_console_log()`.
- `pdbp` + `wrapt` -> fn-local in
  `devx._frame_stack.hide_runtime_frames()`/`api_frame()`.
- `platformdirs` -> fn-local in `runtime._state.get_rt_dir()`.

Still eager (documented follow-ups),
- `pdbp` via `devx.debug._repl` class-bases
  (`PdbREPL(pdbp.Pdb)`) + the module-lvl `@pdbp.hideframe` in
  `._tty_lock`; needs a `._repl` restructure.
- `platformdirs` via the `UDSAddress.def_bindspace: ClassVar`
  class-body eval of `get_rt_dir()`; needs an `Address`-proto
  rework.
- `stackscope` is already fn-local; `setproctitle` is not
  imported anywhere.

Prompt-IO: ai/prompt-io/claude/20260702T155626Z_65bf9df5_prompt_io.md
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
Gud Boi a2c8af558e Use `sys._getframe()` in `get_logger()` mod lookup
`get_caller_mod()` (nested in `get_logger()`) walks the WHOLE
call-stack via `inspect.stack()`, which also resolves src-file
info for every frame and scans all of `sys.modules` per frame
via `inspect.getmodule()`. During nested imports (deep
importlib stacks) each module-level `get_logger()` call costs
~5-10ms, making the ~39 such calls dominate `import tractor`
wall-time: ~244ms of the ~420ms total (see gh #470).

Deats,
- resolve the caller frame with `sys._getframe(frames_up)` and
  map its `f_globals['__name__']` through `sys.modules`: O(1)
  vs. O(stack x sys.modules).
- guard `ValueError` (stack too shallow) -> `None`, matching
  the existing null-caller handling at all use-sites.
- drop the now-unused `inspect` imports; pull `FrameType` from
  `types` instead.

Results: `import tractor` drops 0.42s -> ~0.155s; sequential
`.start_actor()` spawn latency ~0.42 -> ~0.18s/actor.

Prompt-IO: ai/prompt-io/claude/20260702T155626Z_65bf9df5_prompt_io.md
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 19:57:28 -04:00
Bd 3ad7e7e5dc
Merge pull request #459 from goodboy/dependabot/uv/idna-3.15
Bump idna from 3.10 to 3.18
2026-08-12 19:55:40 -04:00
dependabot[bot] 4b4cc76263 Bump idna from 3.10 to 3.18
Bumps [idna](https://github.com/kjd/idna) from 3.10 to 3.18.
- [Release notes](https://github.com/kjd/idna/releases)
- [Changelog](https://github.com/kjd/idna/blob/master/HISTORY.md)
- [Commits](https://github.com/kjd/idna/compare/v3.10...v3.18)

---
updated-dependencies:
- dependency-name: idna
  dependency-version: '3.18'
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-08-12 19:35:17 -04:00
Bd 99f9beccb2
Merge pull request #487 from goodboy/dependabot/uv/setuptools-83.0.0
Bump setuptools from 82.0.1 to 83.0.0
2026-08-12 17:13:22 -04:00
dependabot[bot] 5c4d42c7a7 Bump setuptools from 82.0.1 to 83.0.0
Bumps [setuptools](https://github.com/pypa/setuptools) from 82.0.1 to 83.0.0.
- [Release notes](https://github.com/pypa/setuptools/releases)
- [Changelog](https://github.com/pypa/setuptools/blob/main/NEWS.rst)
- [Commits](https://github.com/pypa/setuptools/compare/v82.0.1...v83.0.0)

---
updated-dependencies:
- dependency-name: setuptools
  dependency-version: 83.0.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-08-12 16:48:08 -04:00
Bd d887603a1b
Merge pull request #479 from goodboy/wkt/start_or_cancel_tests_474
Add `.trionics.start_or_cancel()` test suite
2026-08-12 16:43:33 -04:00
Gud Boi ae67e2f429 Assert cancellation at the startup boundary
Replace the pre-start wall-clock sleep with an indefinite
checkpoint so cancellation ordering cannot race a timer in slow CI.

Record that `Cancelled` escapes each `start_or_cancel()` await
before the enclosing nursery or cancel scope handles it.

Review: PR #479 (goodboy)
https://github.com/goodboy/tractor/pull/479

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 15:45:30 -04:00
Gud Boi 1c7d0c7f3e Match Trio's startup error exactly
Compare the complete canonical `Nursery.start()` protocol error
before re-surfacing ambient cancellation. Preserve child-owned
`RuntimeError` objects whose messages only resemble Trio's wording.

Cover the colliding prefix and assert the original error remains the
exception group's sole leaf.

Review: PR #479 (goodboy)
https://github.com/goodboy/tractor/pull/479

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 15:45:23 -04:00
Gud Boi 203a0f7e1f Add `.trionics.start_or_cancel()` test suite
Resolve #474 with a new `tests/trionics/test_taskc.py` (9
tests) covering the `trio.Nursery.start()` wrapper landed in
PR #464, incl. the `modden.runtime.progman.open_wks()` use
case dug out as a minimal repro.

Deats,
- the lossy `RuntimeError('child exited without calling
  task_status.started()')` only fires when the child absorbs
  its ambient cancel pre-`.started()` (graceful-teardown
  pattern); a well-behaved child surfaces `Cancelled` direct
  from `.start()` on `trio` 0.29 - verified empirically 1st.
- `test_sibling_err_not_masked_by_startup_rte`: the `modden`
  case; ONLY the root-cause sibling `ValueError` escapes the
  nursery with the wrapper vs. bare-`.start()`'s lossy
  riding-along startup-RTE noise.
- `test_pure_oob_cancel_not_morphed_to_rte`: plain ancestor
  `cs.cancel()` exits clean vs. bare's eg-wrapped RTE.
- `test_genuine_startup_rte_still_raised`: sans cancellation
  the protocol-bug RTE re-raises same as bare.
- `test_childs_own_rte_never_demoted_to_cancel`: exact-msg +
  `isinstance`-guard regression cover; a child's own
  `RuntimeError('never got started!')`/`RuntimeError(1234)`
  never demotes to a `Cancelled`.
- `test_started_value_and_args_passthru`: positional args,
  `name=` and `.started()`-value forwarding.

Also,
- `use_start_or_cancel=False` params pin upstream `trio`'s
  current lossy behaviour as wart-documentation: a break on
  a `trio` upgrade likely means new upstream porcelain and
  the wrapper deserves a re-audit.
- verified 0 flakes over 50 hammer runs + 2 impl mutations
  each caught by exactly the targeted tests.

Prompt-IO: ai/prompt-io/claude/20260702T161624Z_65bf9df5_prompt_io.md
(this patch was generated in some part by [`claude-code`][claude-code-gh])
[claude-code-gh]: https://github.com/anthropics/claude-code
2026-08-12 14:15:31 -04:00
Bd 92c737ad83
Merge pull request #458 from mahmoudhas/fix/guard-hot-path-log-rendering
Guard hot-path log calls to avoid payload rendering when disabled
2026-08-12 14:08:47 -04:00
Gud Boi 935c8cf656 Guard receive-path transport rendering
Skip raw packet, decoded message, peer, and channel formatting when
transport logging is disabled. Keep message processing and wire
reads outside the guards so logging controls never affect IPC flow.

Also narrow the remaining pretty-struct TODO to require a
non-raising formatter with native-repr fallback.

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 13:48:40 -04:00
Gud Boi 84ec895150 Drop resolved channel log-guard TODO
Remove the stale `Channel.from_addr()` design note now that its
`at_least_level()` guard avoids inactive pretty-rendering work.

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-12 12:51:43 -04:00
Gud Boi 67280c2898 Honor log disable controls in hot-path guards
Use `Logger.isEnabledFor()` in `at_least_level()` so logger-local
and global disable controls short-circuit payload rendering.

Add `Channel.send()` coverage for effective-level, per-logger, and
global suppression while ensuring transport remains unchanged.

Caught-during: review remediation
Found-via: `/run-tests` test_log_guard_skips_payload_formatting

Review: PR #458 (goodboy)
https://github.com/goodboy/tractor/pull/458#issuecomment-5258207470

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-11 21:56:27 -04:00
root 0e11ff7e9d Guard hot-path log calls to avoid payload rendering when disabled
Wrap `log.transport()` in `Channel.send()` and `log.runtime()` in
`PldRx.decode_pld()` with `log.at_least_level()` checks so that
expensive `pformat(payload)` / `repr(msg)` / `repr(pld)` calls are
skipped entirely when the respective log level is not active.

Previously the f-string arguments were eagerly evaluated before being
passed to the log method, even though `StackLevelAdapter.log()` would
then discard the message internally via its own `isEnabledFor()` check.
On high-frequency IPC paths this caused `pformat` to dominate CPU
usage (~60-70 %) as reported in #455.

This also restores the full diagnostic output (msg type, decoded
payload) that was temporarily commented out in 0373164 as a stopgap.

Resolves #455
2026-08-11 20:41:20 -04:00
mahmoud 148a098ca6 avoid format on the hot send path 2026-08-11 20:41:20 -04:00
Bd 83b3488455
Merge pull request #488 from goodboy/wkt/moc_teardown_completion
Wait for ctx exit in `maybe_open_context()`
2026-08-11 12:16:15 -04:00
Gud Boi daa661aba3 Clarify `maybe_open_context()` teardown notes
Drop the stale sentinel experiment and fix the cancellation-path
comment. Document that cached regular `__aexit__()` failures are
always re-raised at the final consumer boundary.

Review: PR #488 (goodboy)
https://github.com/goodboy/tractor/pull/488

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-11 11:38:43 -04:00
Gud Boi 55ec3dbf51 Drop unused `_Cache` teardown bindings
Remove the unused `value` assignment after cache eviction and skip
unpacking stale resource state before raising its invariant error.

Review: PR #488 (Copilot)
https://github.com/goodboy/tractor/pull/488#pullrequestreview-4850557500

(this patch was generated in some part by `opencode` using
`gpt-5.6-sol` (`openai`))
2026-08-07 22:24:27 -04:00
Gud Boi a753625fc6 Wait for ctx exit in `maybe_open_context()`
Block the final user on its cached resource's `__aexit__()` and
raise regular cleanup errors at that user's ctx boundary.

Deats,
- serialize user registration and teardown under each cache-key lock
- keep queued entrants on the same lock through resource replacement
- preserve `Cancelled`, `KeyboardInterrupt`, and `SystemExit` flow
- cover successful, failing, cancelled, and re-entry teardown paths

Prompt-IO: ai/prompt-io/opencode/20260804T030309Z_65bf9df5_prompt_io.md

(this patch was generated in some part by `opencode` using `gpt-5.6-sol` (`openai`))
2026-08-03 23:45:42 -04:00
69 changed files with 4106 additions and 988 deletions

View File

@ -91,6 +91,10 @@ jobs:
name: '${{ matrix.os }} Python${{ matrix.python-version }} spawn_backend=${{ matrix.spawn_backend }} tpt_proto=${{ matrix.tpt_proto }}' name: '${{ matrix.os }} Python${{ matrix.python-version }} spawn_backend=${{ matrix.spawn_backend }} tpt_proto=${{ matrix.tpt_proto }}'
timeout-minutes: 16 timeout-minutes: 16
runs-on: ${{ matrix.os }} runs-on: ${{ matrix.os }}
# Windows support is nascent: its full test suite remains
# informational, while setup and the `import tractor` smoke below
# are hard signals. Promote the test step to required once the
# suite is green.
strategy: strategy:
fail-fast: false fail-fast: false
@ -98,6 +102,7 @@ jobs:
os: [ os: [
ubuntu-latest, ubuntu-latest,
macos-latest, macos-latest,
windows-latest,
] ]
python-version: [ python-version: [
'3.13', '3.13',
@ -118,10 +123,10 @@ jobs:
'tcp', 'tcp',
'uds', 'uds',
] ]
# https://github.com/orgs/community/discussions/26253#discussioncomment-3250989
exclude: exclude:
# don't do UDS run on macOS (for now) # UDS is POSIX-only; Windows has no `AF_UNIX` so the
- os: macos-latest # backend is intentionally unavailable there.
- os: windows-latest
tpt_proto: 'uds' tpt_proto: 'uds'
steps: steps:
@ -150,7 +155,14 @@ jobs:
- name: List deps tree - name: List deps tree
run: uv tree run: uv tree
# hard signal for the Windows import-safety fix: `import
# tractor` must succeed everywhere, and `HAS_UDS` reflects
# platform capability (False on Windows, True on POSIX).
- name: 'Smoke: import tractor'
run: uv run python -c "import sys; import tractor; from tractor.ipc._uds import HAS_UDS; assert sys.platform != 'win32' or not HAS_UDS; print('import tractor OK | HAS_UDS=', HAS_UDS)"
- name: Run tests - name: Run tests
continue-on-error: ${{ matrix.os == 'windows-latest' }}
run: > run: >
uv run uv run
pytest pytest

View File

@ -1,187 +0,0 @@
# `_ria_nursery` removal plan (issue #477 follow-up)
Goal: drop the secondary "run-in-actor" spawn nursery (and
friends) from `ActorNursery`/spawn internals, now that
`tractor.to_actor.run()` delivers one-shot semantics purely on
the daemon-spawn + portal primitives.
## Verified machinery map (2026-07-02, wkt @ a34aaf98)
The entire mechanism is 4 files:
- `runtime/_supervise.py`
- `ActorNursery.__init__(.., ria_nursery, ..)` stores
`._ria_nursery` (:202, :238); sole read is
`run_in_actor()` passing `nursery=self._ria_nursery`
(:442) into `start_actor()`'s `nursery:
trio.Nursery|None` escape-hatch param (:305, :367).
- `._cancel_after_result_on_exit: set` (:244) marks ria
portals (:457).
- `_open_and_supervise_one_cancels_all_nursery()` nests
`da_nursery` (:609) around `ria_nursery` (:622); the
`finally:` at the ria->da boundary (:747-766) raises
collected `errors` (single exc or BEG).
- `runtime/_portal.py`
- `._expect_result_ctx` (:112) set by `_submit_for_result()`
(:142, sole caller `run_in_actor()`); consumed by
`wait_for_result()` (:167) + deprecated `result()` (:220).
The `None` branch (:184-196) returns the `NoResult`
sentinel (`_exceptions.py:1164`).
- `spawn/_spawn.py`
- `exhaust_portal()` (:129): awaits
`portal.wait_for_result()`, CATCHES+RETURNS any exc
(never raises).
- `cancel_on_completion()` (:177): `exhaust_portal()` ->
on exc-result stash `errors[uid] = result` (:203) ->
ALWAYS `portal.cancel_actor()` (:218).
- `spawn/_trio.py` (:195-222) + `spawn/_mp.py` (:187-213),
identical shape: after shielded
`await an._join_procs.wait()`, open a per-child local
nursery; IFF `portal in an._cancel_after_result_on_exit`
start `cancel_on_completion` alongside `soft_kill()`; when
`soft_kill` returns first, `nursery.cancel_scope.cancel()`
reaps the result-waiter.
## The load-bearing semantic (already-deferred errors)
Remote ria-child errors NEVER raise into `ria_nursery`:
1. reaper tasks only START after `_join_procs.set()` (block
exit or the inner error handler),
2. `exhaust_portal` swallows the exc into a return value,
3. `cancel_on_completion` stashes it in `errors` + cancels
that child,
4. the ria->da `finally:` re-raises collected `errors` (and
`an.cancel()`s any daemon stragglers).
So mid-block there is NO error propagation from ria children
(unless user code explicitly `await portal.wait_for_result()`s)
— the two-nursery nesting only sequences "reap ria results
BEFORE blocking on daemon join". A single-nursery impl only
needs to preserve that sequencing, not any ASAP-cancel
behavior.
## Target design
### step A: single-nursery `run_in_actor()` (mechanical)
- `run_in_actor()` spawns via the DEFAULT (`_da_nursery`)
path — drop `nursery=self._ria_nursery`.
- rename `._cancel_after_result_on_exit` ->
`._ria_portals: dict[portal, Actor]` (need the subactor ref
for `cancel_on_completion`).
- move reaper start-up OUT of the backends into
`_open_and_supervise...`: immediately after EACH
`an._join_procs.set()` call-site (happy path :642, inner
error handler :661), start one
`cancel_on_completion(portal, subactor, errors)` task per
ria portal into `da_nursery`, then (happy path only)
`await` their completion BEFORE falling out of the
`try:`/`finally:` that raises `errors` — e.g. gather in a
dedicated inner `trio.open_nursery()` block replacing
today's `ria_nursery` join point.
- delete the membership branch + local reaper nursery from
`_trio.py`/`_mp.py` (keep the `soft_kill()` call; the
per-child local nursery collapses to just `soft_kill`).
- `_trio.py:310` `_children.pop()` etc. unchanged.
### step B: delete the plumbing
- `_open_and_supervise...`: drop the inner
`ria_nursery` + merge its `except BaseException` classify
logic into ONE handler on the (now single) nursery scope;
`ActorNursery.__init__` loses the `ria_nursery` param.
- `start_actor()` loses the `nursery:` escape-hatch param
(the :302-304 TODO).
- backends: no more `_cancel_after_result_on_exit` refs.
### step C: (separate PRs) deprecate + migrate + excise
- migrate in-repo `.run_in_actor()` usage to
`to_actor.run()`: tests 46 hits/9 files (test_cancellation
15, test_infected_asyncio 10, test_spawning 8, registrar 3,
adv_streaming 4, pubsub 2, rpc 1, runtime 1), examples 28
hits/13 files (debugging/* dominate), docs 20 hits/8 rst
files. NOTE: many sites also use deprecated
`Portal.result()`/`wait_for_result()` — these die with
`_expect_result_ctx`, so migration must land FIRST.
- add `DeprecationWarning` to `run_in_actor()` (+
`_submit_for_result`/`wait_for_result`).
- final excision: `run_in_actor()`, `_submit_for_result`,
`_expect_result_ctx`, `wait_for_result`/`result`,
`exhaust_portal`, `cancel_on_completion`, `NoResult`.
## Risk register
1. hard-killed ria child: today the backend-local
`nursery.cancel_scope.cancel()` discards a still-parked
reaper when the proc dies first; a da_nursery-hosted
reaper instead sees the transport break ->
`exhaust_portal` returns a `TransportClosed`-ish exc ->
NEW entry in `errors` that today gets discarded. Guard:
reap-gather block must cancel remaining reapers once all
ria procs are dead, or filter transport-death excs for
already-`cancel_called` children.
2. error-path ordering: inner handler today sets
`_join_procs` THEN `an.cancel()`; reapers race the
cancel-RPC. Keep that ordering when moving reaper spawn.
3. debugger interplay: `maybe_wait_for_debugger()` calls
(:654, :730) must stay BEFORE any reap/cancel issuance.
4. `errors` double-entry: local body error (:646) + child's
relayed exc (via reaper) can both land for the same
scenario -> BEG shape changes vs today? (today has the
same dual-write sites; keep behavior identical.)
5. mp backend parity: mirror every `_trio.py` edit in
`_mp.py` (identical block).
## Step-A first-probe findings (2026-07-02, WIP in tree)
Step A is IMPLEMENTED (uncommitted):
`run_in_actor()` spawns via da_nursery; new
`_supervise._reap_ria_portals()` helper; reap awaited after
happy-path `_join_procs.set()`; error-path runs reap
CONCURRENT with `an.cancel()` in the shielded block;
backends stripped of the membership branch + per-child
reaper nursery (+ dead imports).
Probe history (trio backend):
- `tests/test_to_actor.py` + `tests/test_spawning.py`:
20/20 PASS — incl. all `run_in_actor()` result
round-trips + `test_remote_error` (single erroring child,
body re-raise -> inner error path).
- FIRST attempt ran the error-path reap CONCURRENT with
`an.cancel()` (mimicking the old backend-side race):
`test_cancellation.py::test_multierror` (2 erroring ria
children, body re-raises one) DEADLOCKED. Root cause per
the sequencing fix below: reap + cancel must NOT race at
this layer (suspected `._children` pop-during-iteration
and/or double-cancel RPC wedge; not fully root-caused
since the fix removes the race wholesale).
- FIX (2nd attempt, current impl): error path SEQUENCES:
(1) snapshot ria `(portal, subactor)` pairs (backend
`finally`s pop `._children` as procs reap), (2)
`await an.cancel()`, (3) bounded reap over the snapshot.
Bound was first 3s -> blew the `fail_after` deadline in
`test_cancel_while_childs_child_in_sync_sleep` (hard-
killed grandchild never relays => reaper parks the full
bound). Tightened to 0.5s: anything collectable is
already queued in the local ctx (relayed BEFORE the
cancel); a parked reaper self-cleans (`trio.Cancelled`
results are never stashed).
- RESULT: `tests/test_cancellation.py` FULLY GREEN
(20 passed, 1 xfailed, 77s); full-suite gate run kicked
off same session (see final report/next session).
Remaining risk: on slow CI a relayed-but-undelivered error
racing the 0.5s bound could drop an `errors` entry
(BEG-shape flake); if observed, scale the bound via the
`cpu_perf_headroom()`-style approach or peek
`Portal._final_result_msg`/ctx queue state instead of
time-bounding.
## Verification gate
- `tests/test_cancellation.py test_spawning.py test_local.py
test_rpc.py` on `trio` + `mp_spawn` + `mp_forkserver`
backends, then full suite; `tests/devx/test_debugger.py`
for risk 3.

View File

@ -0,0 +1,77 @@
---
model: claude-fable-5
service: claude
session: b8ec2ffe-1e47-4db5-8bc3-5bebc5feaaf6
timestamp: 2026-07-02T15:50:06Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260702T155006Z_65bf9df5_prompt_io.raw.md
---
## Prompt
From `ai/prompt-io/prompts/issue_473.md`:
> NOTE: you MUST pause this work at 12:50PM EST (BEFORE
> your weekly limit reset) for review by a human!
>
> resolve GH #473: https://github.com/goodboy/tractor/issues/474
> do it with /open-wkt.
(The prompt's issue-number/URL mismatch was resolved in
favor of #473 — file name, session name and prompt text
all say #473; the `/474` URL is a typo.)
GH #473: "Root-cause UDS-on-macOS —
`uds_transport_actor_tree` example fails on CI".
## Response summary
Worked all four #473 next-step checkboxes in wkt
`uds_macos_473`:
1. **stderr surfacing**: `test_docs_examples.py` now
raises with the full subproc stderr (+stdout) on ANY
non-zero exit. The old last-line-'Error' check could
never fire for `tractor` crashes since root-actor
stderr always ends with the strict-EG collapse note
`( ^^^ this exc was collapsed from a group ^^^ )`
proven against the real PR #460 macOS CI log (bare
`assert 1 == 0`, no traceback).
2. **root-cause (linux-provable layer)**: macOS-only
addr corruption in
`MsgpackUDSStream.get_stream_addrs()` — no
`SO_PASSCRED`/autobind on darwin means the accept
side's `getpeername()` is `''`, and the
`(str(), str())` arm took `peername` unconditionally →
`Path('')` garbage addrs on every accepted conn.
Proven + fixed via linux no-autobind simulation.
Possibly not the final macOS crasher (non-fatal on
linux-sim); the diagnostic patch guarantees the next
macOS CI run shows any remaining layer.
3. **CI matrix**: removed the `macos-latest`+`uds`
exclude.
4. **un-skip**: dropped the macOS+CI skip of the example.
Also: `start_listener()` bindspace mkdir hardened
(`parents=True, exist_ok=True`), example docstring
peer-pid mechanism corrected for macOS.
## Files changed
- `tests/test_docs_examples.py` — surface full stderr on
non-zero exit; remove macOS skip of the UDS example
- `tractor/ipc/_uds.py` — fix no-autobind
`get_stream_addrs()` arm; harden bindspace mkdir;
document autobind semantics
- `.github/workflows/ci.yml` — add macOS UDS leg (remove
matrix exclude)
- `examples/uds_transport_actor_tree.py` — platform-
correct peer-pid docs
## Human edits
None yet — pending human review (hard-stop review
checkpoint at 12:50PM EST per prompt); commit staging and
any edits are the human's.

View File

@ -0,0 +1,101 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T15:50:06Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/uds_macos_473
---
# Raw output — GH #473 UDS-on-macOS root-cause session
NOTE: code output is diff-referenced (not copied) per
prompt-io diff-ref mode; all generated code is on branch
`wkt/uds_macos_473` relative to `main` (65bf9df5).
## Diagnostic narrative (verbatim)
Root-cause work proceeded by linux-side simulation since no
macOS host is available:
1. Pulled the actual macOS CI failure log from PR #460 run
28343878933 (job 84115107992): the failure is a bare
`assert 1 == 0` on `proc.returncode` with NO stderr
surfaced — the harness only re-raised when the LAST
stderr line contained 'Error'.
2. Reproduced the stderr-swallow mechanism on linux: any
`tractor` root-actor crash ends stderr with the
strict-EG collapse note
`( ^^^ this exc was collapsed from a group ^^^ )` which
never matches 'Error' — so EVERY possible crash was
swallowed. (Verified by sabotaging the runtime dir via
an over-long `XDG_RUNTIME_DIR` → `OSError: AF_UNIX path
too long` → rc=1 + swallowed.)
3. Found + proved a macOS-only addr-corruption bug in
`MsgpackUDSStream.get_stream_addrs()`: the
`(str(), str())` match-arm unconditionally took
`peername`, but on no-autobind platforms (macOS lacks
linux's `SO_PASSCRED`-triggered autobind) the accept
side's `getpeername()` is `''``Path('')` garbage
laddr/raddr on EVERY accepted UDS conn. Simulated on
linux by nulling `SO_PASSCRED` (no autobind → same `''`
shape): pre-fix the example printed
`listener sock file: .`; post-fix it prints the real
registry sockpath. Non-fatal on linux-sim (rc=0), so
possibly not the final macOS crasher — the diagnostic
patch guarantees the next macOS CI run reveals any
remaining layer.
4. Falsified the missing-parent-dir theory:
`get_rt_dir()` already `mkdir(parents=True,
exist_ok=True)`s at import (and macOS TCP CI passes),
so `~/Library/Caches/TemporaryItems` absence cannot be
the crasher. Hardened `start_listener()`'s bindspace
mkdir anyway (custom `filedir` case + racing actors).
## Generated changes (diff pointers)
> `git diff main..wkt/uds_macos_473 -- tests/test_docs_examples.py`
- always raise with FULL subproc stderr (+stdout) on any
non-zero example exit; keep legacy last-line 'Error'
check for zero-rc cases; drop the macOS+CI skip of
`uds_transport_actor_tree.py` (GH #473 next-step).
> `git diff main..wkt/uds_macos_473 -- tractor/ipc/_uds.py`
- `get_stream_addrs()`: document the autobind semantics
(bytes = linux abstract-ns autobind artifact), add
no-autobind `(str, str)` arm picking the non-empty name
(`peername` connect-side, `sockname` accept-side) with
an empty-pair `ValueError` guard.
- `start_listener()`: `bs.mkdir(parents=True,
exist_ok=True)`.
> `git diff main..wkt/uds_macos_473 -- .github/workflows/ci.yml`
- remove the `macos-latest`+`uds` matrix exclude so
UDS-on-macOS is exercised by CI (GH #473 next-step).
> `git diff main..wkt/uds_macos_473 -- examples/uds_transport_actor_tree.py`
- docs nit: peer-pid mechanism is `SO_PEERCRED` on linux,
`LOCAL_PEERPID` on macOS.
## Verification (verbatim summary)
- macOS-shape sim (no autobind) + fix: example rc=0 with
correct listener sockpath.
- native linux post-fix: example rc=0, autobind arms
unchanged.
- sabotage smoke-test: diagnostic patch surfaces the full
EG traceback incl. collapse-note line.
- `pytest tests/ipc/ tests/test_2way.py tests/discovery/
--tpt-proto uds`: 61 passed, 1 `TooSlowError` flake
(test_simple_rpc) that passes solo in 2.2s — attributed
to a concurrent full-suite run (other session) loading
the host.
- full `tests/test_docs_examples.py` run pending at
raw-file write time (queued behind the concurrent
session's suite).

View File

@ -0,0 +1,83 @@
---
model: claude-fable-5
service: claude
session: b6b42e23-0454-4d9b-be1c-0a89d65aed1b
timestamp: 2026-07-02T15:56:26Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260702T155626Z_65bf9df5_prompt_io.raw.md
---
## Prompt
From `ai/prompt-io/prompts/issue_470.md`:
> attempt to resolve
> https://github.com/goodboy/tractor/issues/470
> do it with /open-wkt, ensure you /run-tests on all
> changes
(plus a hard pause-for-human-review deadline of
12:50PM EST.)
Issue #470 asks to trim the ~0.42s `import tractor`
cost — which dominates per-actor spawn latency on the
`trio` backend — by lazy-importing heavy/optional deps
(`pdbp`, `stackscope`, `multiaddr`, + audit of
`colorlog`/`bidict`/`wrapt`/`setproctitle`).
## Response summary
Profiling showed the issue's dep-list only accounted
for ~20ms; the dominant cost (~244ms) was
`log.get_logger()`'s `get_caller_mod()` calling
`inspect.stack()` at module level in ~39 modules —
each call walks every stack frame (deep during nested
imports) and scans `sys.modules` per frame via
`inspect.getmodule()`.
Changes, in impact order:
1. `get_caller_mod()` -> `sys._getframe()` +
`f_globals['__name__']` `sys.modules` lookup
(~240ms saved).
2. Issue's lazy-import checklist: `bidict`,
`multiaddr`, `colorlog`, `wrapt` moved to
`TYPE_CHECKING`/function-local imports;
`platformdirs` function-local; `asyncio` +
`.to_asyncio` deferred out of the `devx.debug` +
`spawn._entry` eager paths (~15ms saved).
3. PEP 562 `__getattr__` on `tractor/__init__.py`
preserving public `tractor.to_asyncio` attr access.
Results: `import tractor` 0.42s -> ~0.145s (~65%);
sequential `start_actor` latency 0.40-0.44s ->
~0.179s/actor. `pdbp` (needs `_repl.py` class-base
restructure) + `platformdirs` (needs
`UDSAddress.def_bindspace` protocol rework) documented
as follow-ups.
## Files changed
- `tractor/log.py``get_caller_mod()` perf fix +
lazy `colorlog`
- `tractor/__init__.py` — PEP 562 lazy `to_asyncio`
- `tractor/discovery/_addr.py``bidict` ->
`TYPE_CHECKING`
- `tractor/discovery/_multiaddr.py` — lazy `multiaddr`
- `tractor/ipc/_tcp.py`, `tractor/ipc/_uds.py`
`Multiaddr` -> `TYPE_CHECKING`
- `tractor/runtime/_state.py` — lazy `platformdirs`
- `tractor/devx/_frame_stack.py` — lazy `pdbp` +
`wrapt`
- `tractor/devx/debug/_trace.py`,
`tractor/devx/debug/_tty_lock.py` — lazy `asyncio` +
`.to_asyncio`
- `tractor/spawn/_entry.py` — lazy
`run_as_asyncio_guest`
## Human edits
None yet — pending user review at the 12:50PM EST
pause gate (test-suite results reported in-session).

View File

@ -0,0 +1,122 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T15:56:26Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/boot_latency_470
---
# Raw AI output — gh #470 `import tractor` latency trim
All generated code is committed on the
`wkt/boot_latency_470` branch; per diff-ref mode each
file's content is referenced via its diff instead of
copied verbatim.
## Profiling findings (verbatim analysis output)
Baseline: `import tractor` ~0.39-0.42s wall.
`python -X importtime` + `cProfile` traced the cost NOT
primarily to third-party deps (the issue's hypothesis)
but to `tractor/log.py:get_logger()` calling
`get_caller_mod()` -> `inspect.stack()` at module level
in ~39 tractor modules:
- `inspect.stack()` builds `FrameInfo` (incl. src-file
and line-context resolution) for EVERY frame on the
stack; during nested imports the stack is dozens of
importlib frames deep.
- each `FrameInfo` resolution calls
`inspect.getmodule()` which scans all of
`sys.modules` per frame (1.4M `ismodule()` calls in
one profiled import).
- aggregate: ~244ms of tractor-own module "self" time
vs ~20ms for ALL the issue-listed third-party deps
(`pdbp` ~10ms, `bidict` ~4.5ms, `multiaddr` ~3.5ms,
`wrapt`/`colorlog` ~1ms each); `trio` itself is
~70-100ms and unavoidable.
## Generated changes
> `git diff main..wkt/boot_latency_470 -- tractor/log.py`
`get_caller_mod()` rewritten from `inspect.stack()` +
`inspect.getmodule()` to `sys._getframe(frames_up)` +
`frame.f_globals['__name__']` -> `sys.modules` lookup
(O(1) vs O(stack x sys.modules)). Unused `inspect`
imports dropped; `FrameType` imported from `types`.
Also `colorlog` lazy-imported inside
`get_console_log()`.
> `git diff main..wkt/boot_latency_470 -- tractor/discovery/_addr.py`
`bidict` import moved under `TYPE_CHECKING`
(annotation-only use; `_address_types` is a plain dict
literal).
> `git diff main..wkt/boot_latency_470 -- tractor/discovery/_multiaddr.py`
`from __future__ import annotations` added; `multiaddr`
import moved under `TYPE_CHECKING` + function-local
imports in `mk_maddr()`/`parse_maddr()`.
> `git diff main..wkt/boot_latency_470 -- tractor/ipc/_tcp.py tractor/ipc/_uds.py`
`Multiaddr` imports moved under `TYPE_CHECKING`
(annotation-only in both transports).
> `git diff main..wkt/boot_latency_470 -- tractor/runtime/_state.py`
`platformdirs` lazy-imported inside `get_rt_dir()`
(NOTE: still imported eagerly via
`UDSAddress.def_bindspace` class-var eval; see
follow-ups).
> `git diff main..wkt/boot_latency_470 -- tractor/devx/_frame_stack.py`
`pdbp` + `wrapt` lazy-imported inside
`hide_runtime_frames()` / `api_frame()` respectively.
> `git diff main..wkt/boot_latency_470 -- tractor/devx/debug/_trace.py tractor/devx/debug/_tty_lock.py`
`asyncio` moved to `TYPE_CHECKING` + call-site local
imports (`asyncio.current_task()` sites);
`tractor.to_asyncio.run_trio_task_in_future` imports
moved into the infected-aio runtime branches.
> `git diff main..wkt/boot_latency_470 -- tractor/spawn/_entry.py`
`run_as_asyncio_guest` import moved into the
`infect_asyncio=True` branches of `_mp_main()` /
`_trio_main()`.
> `git diff main..wkt/boot_latency_470 -- tractor/__init__.py`
PEP 562 module `__getattr__` added so
`tractor.to_asyncio` attr-access still works (required
by `tests/test_child_manages_service_nursery.py` and
any downstream user) while keeping `asyncio` off the
eager import path.
## Measured results (verbatim)
- `import tractor`: 0.39-0.42s -> ~0.145s (~65% cut)
- `start_actor` spawn+boot+reg+cancel: ~0.40-0.44s ->
~0.179s/actor (n=5 sequential, warm parent)
- post-change eager-module check: only `pdbp` +
`platformdirs` of the issue's list remain eager.
## Known follow-ups (not implemented, deadline-bound)
- `pdbp` (~10ms): still eager via
`devx/debug/_repl.py` class bases
(`class PdbREPL(pdbp.Pdb)`) + `_tty_lock.py`
module-level `@pdbp.hideframe`; needs `_repl`
restructure + PEP 562 in `devx.debug.__init__`.
- `platformdirs` (~1.5ms): eager via
`UDSAddress.def_bindspace: ClassVar = get_rt_dir()`
class-body call; needs `Address`-protocol rework of
`def_bindspace` to a lazy accessor.
- `stackscope` + `setproctitle`: already lazy/absent —
no change needed.

View File

@ -0,0 +1,80 @@
---
model: claude-fable-5
service: claude
session: f4da6654-63f9-4641-b567-f200da46d697
timestamp: 2026-07-02T16:16:24Z
git_ref: 65bf9df5
scope: tests
substantive: true
raw_file: 20260702T161624Z_65bf9df5_prompt_io.raw.md
---
## Prompt
> attempt to resolve
> https://github.com/goodboy/tractor/issues/474
> do it with /open-wkt.
Issue #474: "We need test suite for
`.trionics.start_or_cancel()`" — dig out the use case
from `modden` as a unit test (and provide motivation
for upstream `trio` to possibly provide better
porcelain around this case). The wrapper landed via
PR #464 with no tests.
## Response summary
New `tests/trionics/test_taskc.py` (9 tests) covering
`start_or_cancel()`, authored in wkt
`start_or_cancel_tests_474`.
Behaviour was first probed empirically (trio 0.29):
the lossy `RuntimeError('child exited without calling
task_status.started()')` only fires when the child
exits pre-`.started()` WITHOUT propagating the ambient
`Cancelled` — i.e. when the child (or lib code it
calls) absorbs the cancel in a graceful-teardown
pattern; a well-behaved child surfaces `Cancelled`
straight out of `.start()`. The `modden`
`progman.open_wks()` use case was reconstructed from
`modden/runtime/progman.py` accordingly.
Tests (each `use_start_or_cancel` parametrization also
pins upstream trio's current lossy behaviour as
wart-documentation):
- `test_sibling_err_not_masked_by_startup_rte` — the
`modden` case: sibling error OOB-cancels the shared
nursery scope; with the wrapper ONLY the root-cause
`ValueError` escapes; bare `.start()` adds the lossy
RTE alongside.
- `test_pure_oob_cancel_not_morphed_to_rte` — plain
ancestor `cs.cancel()`: wrapper → clean exit; bare
→ eg-wrapped RTE.
- `test_genuine_startup_rte_still_raised` — no
cancellation → protocol-bug RTE re-raised same as
bare.
- `test_childs_own_rte_never_demoted_to_cancel` — a
child's own `RuntimeError('never got started!')` /
`RuntimeError(1234)` under ambient cancel is never
demoted to `Cancelled` (exact-msg-match + str-guard
regression cover).
- `test_started_value_and_args_passthru` — happy path:
positional args, `name=`, `.started()` value.
Verified: 9/9 pass; 0 flakes across 50 hammer runs;
two impl mutations (checkpoint removed; guard relaxed
to substring match) each caught by exactly the
targeted tests; `tests/trionics/` +
`tests/test_trioisms.py` subset green (23 passed,
5 xfailed); ruff clean; 69-col style.
## Files changed
- `tests/trionics/test_taskc.py` — new
`start_or_cancel()` unit-test suite (gh #474).
## Human edits
Pending review — session paused pre-commit per user
deadline; nothing committed as of this entry.

View File

@ -0,0 +1,107 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T16:16:24Z
git_ref: 65bf9df5
diff_cmd: git diff main..wkt/start_or_cancel_tests_474
---
# Raw output — gh #474 `start_or_cancel()` test suite
## Generated test code
> `git diff main..wkt/start_or_cancel_tests_474 -- tests/trionics/test_taskc.py`
Prose summary of the generated module
(`tests/trionics/test_taskc.py`):
- module docstring framing the `trio.Nursery.start()`
startup-cancellation wart, the wrapper's repair, and
the intent that `use_start_or_cancel=False` params
double as upstream-trio wart-documentation (break on
a trio upgrade → upstream may have shipped porcelain,
re-audit the wrapper); cites gh #474 / PR #464 and
`modden`'s `progman.open_wks()` as the source use
case.
- shared children: `absorbs_cancel_pre_started()` (the
graceful-teardown cancel-absorber which triggers the
lossy RTE path) + `raise_val_err()` (fast-erroring
sibling).
- `test_sibling_err_not_masked_by_startup_rte`
(parametrized `use_start_or_cancel`): asserts eg
contains exactly one `ValueError` and, wrapper-case,
NO residual RTE (`eg.split(ValueError)` remainder is
`None`); bare-case, the residual RTE carries trio's
exact "child exited without calling" wording.
- `test_pure_oob_cancel_not_morphed_to_rte`
(parametrized): wrapper-case runs clean and asserts
`cs.cancelled_caught`; bare-case asserts the
eg-wrapped RTE.
- `test_genuine_startup_rte_still_raised`
(parametrized): no-cancel protocol bug → RTE with
trio's wording from both call forms.
- `test_childs_own_rte_never_demoted_to_cancel`
(parametrized `rte_arg` in `'never got started!'`,
`1234`): child cancels the ambient scope then raises
its own RTE synchronously (no checkpoint between →
deterministically under-cancellation at catch time);
asserts the RTE survives with `args[0]` intact.
- `test_started_value_and_args_passthru`: `.started()`
value, positional args and the `name=` kwarg (via
`trio.lowlevel.current_task().name`) all forward.
## Non-code output (verbatim highlights)
Behaviour probe (trio 0.29, scratchpad scripts) — the
decision basis for the test shapes:
```
== B-sibling-err use_soc=False
start raised: RuntimeError('child exited without
calling task_status.started()')
top-level: ExceptionGroup([ValueError('sibling blew
up!'), RuntimeError('child exited without calling
task_status.started()')])
== B-cs-cancel use_soc=False
top-level: ExceptionGroup([RuntimeError('child
exited without calling task_status.started()')])
== B-sibling-err use_soc=True
start raised: Cancelled()
top-level: ExceptionGroup([ValueError('sibling blew
up!')])
== B-cs-cancel use_soc=True
start raised: Cancelled()
top-level: clean return
== own-rte-under-cancel (both) -> RTE('never got
started!') propagates unchanged
```
Key finding: with a WELL-BEHAVED (non-absorbing) child
an OOB ancestor cancel surfaces `Cancelled` directly
from `.start()` on trio 0.29 — the lossy RTE requires
the child to absorb its cancel pre-`.started()`, which
is what `modden`'s `open_from_wks` teardown did. Trio's
nursery-exit wait defers cancel delivery to children,
so all tested shapes are deterministic (0 flakes / 50
runs).
Mutation verification:
```
mutation 1 (checkpoint_if_cancelled removed):
FAILED test_sibling_err_not_masked_by_startup_rte[True]
FAILED test_pure_oob_cancel_not_morphed_to_rte[True]
mutation 2 (guard relaxed to 'started' substring,
isinstance dropped):
FAILED test_childs_own_rte_never_demoted_to_cancel[never got started!]
FAILED test_childs_own_rte_never_demoted_to_cancel[1234]
```
Final runs:
```
tests/trionics/test_taskc.py: 9 passed in 0.03s
hammer: 0/50 runs failed
tests/trionics/ + tests/test_trioisms.py:
23 passed, 5 xfailed in 3.02s
```

View File

@ -1,78 +0,0 @@
---
model: claude-fable-5
service: claude
session: f6c84722-471a-4458-9a80-e453fea9029f
timestamp: 2026-07-02T16:58:06Z
git_ref: a34aaf98
scope: code
substantive: true
raw_file: 20260702T165806Z_a34aaf98_prompt_io.raw.md
---
## Prompt
Follow-up round in the same session as the
`tractor.to_actor` landing (see
`20260702T154255Z_65bf9df5_prompt_io.md`). After
committing that work the user green-lit the deferred
items:
> go go go on this with what time you have left, in
> particular see if you can get the _ria_nursery
> removal going!
then extended the deadline twice to iterate on the
discovered hang:
> continue on this up until a 12:58:30 deadline
and finally chose "Commit step A now" from the
next-steps prompt.
## Response summary
Step A of the `._ria_nursery` removal (issue #477): hoist
`.run_in_actor()` result-reaping out of the spawn
backends into the `ActorNursery` machinery so ria
children spawn via the default daemon nursery,
- new `_supervise._reap_ria_portals()` runs one
`_spawn.cancel_on_completion()` task per ria child
AFTER `._join_procs` is set; happy path awaits it
right after `._join_procs.set()`.
- error path SEQUENCES: snapshot ria
`(portal, subactor)` pairs -> `await an.cancel()` ->
0.5s-bounded reap. Two failed intermediates informed
this: a concurrent reap+cancel DEADLOCKED
`test_multierror`; a 3s bound blew
`test_cancel_while_childs_child_in_sync_sleep`'s
`fail_after` deadline.
- backends (`spawn/_trio.py`, `spawn/_mp.py`) lose the
`._cancel_after_result_on_exit` membership branch,
per-child reaper nursery + dead imports.
- design/probe-history doc:
`ai/conc-anal/ria_nursery_removal_plan.md` (from an
agent-verified machinery map).
Verification: `test_cancellation.py` fully green
(20 passed, 1 xfailed) incl. the previously-hung
`test_multierror`; `test_to_actor`+`test_spawning`
20/20; bounded full-suite gate SIGINT'd ~30s early at
303 passed / 0 failures (user opted to commit on that
signal, deferring the unbounded re-run to step-B
verification).
## Files changed
- `tractor/runtime/_supervise.py``_reap_ria_portals()`
+ two call-sites; `run_in_actor()` off the ria nursery
- `tractor/spawn/_trio.py` — reaper branch + import drop
- `tractor/spawn/_mp.py` — same as `_trio.py`
- `ai/conc-anal/ria_nursery_removal_plan.md` — plan +
probe history
## Human edits
None yet — committed via the drafted
`.claude/git_commit_msg_ria_step_a.md` (user-driven
`git commit --edit`).

View File

@ -1,55 +0,0 @@
---
model: claude-fable-5
service: claude
timestamp: 2026-07-02T16:58:06Z
git_ref: a34aaf98
diff_cmd: git diff a34aaf98..wkt/to_actor_subpkg
---
# Raw AI output (diff-ref mode)
Step-A code is committed on `wkt/to_actor_subpkg`
directly after `a34aaf98`; per diff-ref mode the verbatim
content is reachable via the pointers below.
## Generated files
> `git diff a34aaf98..wkt/to_actor_subpkg -- tractor/runtime/_supervise.py`
New `_reap_ria_portals(an, errors, ria_children=None)`
helper (one `_spawn.cancel_on_completion()` task per ria
child under `collapse_eg()` + a local nursery);
`run_in_actor()` drops `nursery=self._ria_nursery`; happy
path awaits the reap right after `._join_procs.set()`;
inner error handler snapshots ria pairs, runs
`await an.cancel()` then a `move_on_after(0.5)`-bounded
reap over the snapshot.
> `git diff a34aaf98..wkt/to_actor_subpkg -- tractor/spawn/_trio.py`
> `git diff a34aaf98..wkt/to_actor_subpkg -- tractor/spawn/_mp.py`
Both backends: the post-`_join_procs` block collapses to
a bare `soft_kill()` (membership branch, per-child reaper
nursery, reaper-cancel logging and the now-unused
`cancel_on_completion` imports all removed).
> `git diff a34aaf98..wkt/to_actor_subpkg -- ai/conc-anal/ria_nursery_removal_plan.md`
Agent-verified machinery map, 3-step design (A/B/C),
probe history (deadlock -> sequencing fix -> bound
tightening) and risk register.
## Test runs (verbatim)
```
tests/test_cancellation.py: 20 passed, 1 xfailed in 77.28s
tests/test_to_actor.py + tests/test_spawning.py: 20 passed
full-suite (SIGINT at 540s bound, ~30s short):
303 passed, 9 skipped, 1 xfailed, 2 xpassed in 539.42s
```
Failed intermediates (not in final diff): concurrent
reap+`an.cancel()` deadlocked `test_multierror`; a 3s
reap bound blew
`test_cancel_while_childs_child_in_sync_sleep`'s
`fail_after` deadline.

View File

@ -0,0 +1,39 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: moc-teardown-completion-20260804
timestamp: 2026-08-04T03:03:09Z
git_ref: 65bf9df5
scope: code
substantive: true
raw_file: 20260804T030309Z_65bf9df5_prompt_io.raw.md
---
## Prompt
Patch Tractor's `maybe_open_context()` so the final consumer waits for
resource `__aexit__()` completion and receives cleanup errors. Reuse
`outcome.Outcome` for the exit result; use the smaller mutable
`_CtxExit` holder if that makes the implementation simpler. Add and run
the relevant existing unit tests, but do not commit or push the patch.
After reviewing the result, simplify `_CtxExit` back to an optional
exception because the success outcome carries no useful value.
## Response summary
Added an exception-backed completion handshake between
`_Cache.run_ctx()` and the final `maybe_open_context()` consumer.
Serialized consumer registration and final teardown under the per-key
lock, preserving that lock for queued entrants. Added deterministic
regressions for normal exit, cleanup errors, cancellation interactions,
service-nursery cancellation, and teardown re-entry.
## Files changed
- `tractor/trionics/_mngrs.py` - publish and unwrap cached exit outcomes.
- `tests/test_resource_cache.py` - cover completion and cancellation.
## Human edits
The user directed the final simplification from `outcome.Outcome` to an
optional exception field. The patch remains uncommitted.

View File

@ -0,0 +1,36 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-04T03:03:09Z
git_ref: 65bf9df5
diff_cmd: git diff HEAD~1..HEAD
---
Implemented cached-context exit completion in
`tractor.trionics.maybe_open_context()`.
> `git diff HEAD~1..HEAD -- tractor/trionics/_mngrs.py`
The generated implementation adds `_CtxExit`, whose `done` event
publishes an `outcome.Outcome[None]`. `_Cache.run_ctx()` records either
`Value(None)` or `Error(exc)` after the resource exit attempt. The final
MOC consumer signals `no_more_users`, waits for completion under a
shielded cancel scope, removes the per-key lock, and unwraps the outcome
so ordinary cleanup failures are raised at the consumer boundary.
`trio.Cancelled`, `KeyboardInterrupt`, and `SystemExit` continue through
the service task rather than being converted into regular cleanup
errors.
> `git diff HEAD~1..HEAD -- tests/test_resource_cache.py`
The generated regressions cover successful exit blocking, cleanup-error
delivery, final-user cancellation, cancellation combined with a cleanup
error, and service-nursery cancellation. The existing teardown re-entry
test now uses explicit events instead of a ten-second cleanup sleep and
asserts that the replacement resource is a fresh cache miss.
Verification:
`env PYTHONPATH="$PWD" /home/goodboy/repos/tractor/py313/bin/python -m pytest tests/test_resource_cache.py`
Result: `16 passed in 8.37s`.

View File

@ -0,0 +1,39 @@
---
model: openai/gpt-5.6-sol
service: opencode
session: pr475-review-fixes-20260817
timestamp: 2026-08-17T23:18:25Z
git_ref: 359fe75c
scope: code
substantive: true
raw_file: 20260817T231825Z_359fe75c_prompt_io.raw.md
---
## Prompt
Continue the `/code-review-changes` pass for PR #475 in its isolated
worktree. Address the seven accepted manual-review findings in
`tractor/ipc/_types.py` and `tractor/ipc/_uds.py`, preserve the existing
Windows capability behavior, verify the result, and prepare the work for
human-controlled commit and review-reply steps. Do not publish replies,
stage, commit, or push without the required explicit authorization.
## Response summary
Restored project quote, docstring, multiline-expression, and
`match/case` conventions while retaining the Windows-safe UDS guard.
Removed unnecessary structural and comment churn, then verified the
focused transport, discovery, and lazy-import paths plus the missing
`AF_UNIX` behavior.
## Files changed
- `tractor/ipc/_types.py` - restore project style and guarded
socket-family dispatch.
- `tractor/ipc/_uds.py` - format the UDS capability gate
consistently.
## Human edits
None - the generated patch remains uncommitted and awaits human
review.

View File

@ -0,0 +1,45 @@
---
model: openai/gpt-5.6-sol
service: opencode
timestamp: 2026-08-17T23:18:25Z
git_ref: 359fe75c
diff_cmd: git diff HEAD~1..HEAD
---
Applied the seven accepted manual-review fixes for PR #475 while
preserving the Windows transport capability behavior.
> `git diff HEAD~1..HEAD -- tractor/ipc/_types.py`
The generated changes restore the project's single-quote docstring and
string conventions, remove the unnecessary helper divider, simplify the
transport-registry comments, and restore `match/case` socket-family
dispatch. The UDS case retains a `HAS_UDS` guard that short-circuits
before `socket.AF_UNIX` is evaluated on unsupported hosts. Nearby error
messages are wrapped without changing their content.
> `git diff HEAD~1..HEAD -- tractor/ipc/_uds.py`
The generated change reformats the `HAS_UDS` conjunction according to
the project's multiline boolean-expression convention and simplifies
the adjacent capability comment.
Verification:
`/home/goodboy/repos/tractor/py313/bin/pytest -q tests/test_lazy_imports.py tests/discovery tests/ipc/test_server.py`
Result: `66 passed, 2 xpassed in 60.62s`.
`ruff check --no-cache --output-format=json tractor/ipc/_types.py tractor/ipc/_uds.py`
Result: no findings.
`git diff --check`
Result: no whitespace errors.
An explicit missing-`AF_UNIX` probe set `HAS_UDS = False`, removed the
socket constant, and exercised an unsupported socket family. It raised
the expected `NotImplementedError` instead of `AttributeError`.
No review replies, commits, or pushes were published.

View File

@ -0,0 +1,22 @@
# AI Prompt I/O Log - OpenCode
This directory tracks prompt inputs and model outputs for AI-assisted
development using `opencode`.
## Policy
Prompt logging follows the [NLNet generative AI policy][nlnet-ai]. All
substantive AI contributions are logged with:
- Model name and version
- Timestamps
- The prompts that produced the output
- Unedited model output (`.raw.md` files)
[nlnet-ai]: https://nlnet.nl/foundation/policies/generativeAI/
## Usage
Entries are created by the prompt-io workflow. Human contributors remain
accountable for all decisions. AI-generated content is never presented as
human-authored work.

View File

@ -130,9 +130,10 @@ UDS: same-host, creds included
Pass ``enable_transports=['uds']`` and actors instead talk over Pass ``enable_transports=['uds']`` and actors instead talk over
unix-domain sockets, with socket files placed in the per-user unix-domain sockets, with socket files placed in the per-user
runtime dir (``$XDG_RUNTIME_DIR/tractor/`` on linux, the runtime dir: ``$XDG_RUNTIME_DIR/tractor/`` on linux, a short
``platformdirs`` equivalent elsewhere). Two perks over tcp on a owner-only ``/tmp/tractor-<uid>`` dir on Darwin, and the
single host: ``platformdirs`` equivalent elsewhere. Two perks over tcp on a single
host:
- no ports to fight over; addrs are just file paths, - no ports to fight over; addrs are just file paths,
- the kernel snitches on your peer for free: the listening side - the kernel snitches on your peer for free: the listening side

View File

@ -44,8 +44,9 @@ clan shares one registry with zero config on your part.
The bootstrap rule inside ``open_root_actor()`` is delightfully The bootstrap rule inside ``open_root_actor()`` is delightfully
simple: simple:
- on boot, ping every socket addr in ``registry_addrs``; when none - on boot, probe every addr in ``registry_addrs`` with a bounded
are passed the per-transport defaults are used: for TCP the Tractor ``Aid`` handshake; when none are passed the per-transport
defaults are used: for TCP the
loopback ``('127.0.0.1', 1616)``, for UDS a loopback ``('127.0.0.1', 1616)``, for UDS a
``registry@1616.sock`` file, ``registry@1616.sock`` file,
@ -53,9 +54,11 @@ simple:
actor and register with the *existing* registry; your own IPC actor and register with the *existing* registry; your own IPC
server binds random same-transport addrs instead, server binds random same-transport addrs instead,
- if **nothing answers, congratulations: you just became the - if every address is absent, congratulations: you just became the
registrar**. Your transport server binds the registry addrs registrar. Your transport server binds the registry addrs
themselves and you start serving lookups for everyone else. themselves and you start serving lookups for everyone else,
- if no registrar answers but an address is occupied by a foreign or
non-responsive endpoint, startup fails instead of binding over it.
Pass ``ensure_registry=True`` when your program *requires* being Pass ``ensure_registry=True`` when your program *requires* being
the one-and-only registrar; boot then fails loudly with a the one-and-only registrar; boot then fails loudly with a
@ -196,9 +199,10 @@ the existing registrar:
trio.run(main) trio.run(main)
Per the bootstrap rules above, if the registrar at those addrs is Per the bootstrap rules above, if those addrs are absent this process
*not* reachable this process simply becomes its own (registrar) becomes its own registrar root, so the same code works standalone and
root — so the same code works standalone and as a tree-joiner. as a tree-joiner. An occupied address that does not complete a Tractor
registrar handshake fails startup instead of being rebound.
"Arbiter"? A legacy naming note "Arbiter"? A legacy naming note
------------------------------- -------------------------------

View File

@ -185,8 +185,10 @@ first with a bounded grace window — so actor runtimes can run
their ``trio`` teardown paths — escalating to ``SIGKILL`` only as their ``trio`` teardown paths — escalating to ``SIGKILL`` only as
a last resort. The ``--shm`` sweep unlinks ``/dev/shm/`` segments a last resort. The ``--shm`` sweep unlinks ``/dev/shm/`` segments
that no live process has open (it leans on psutil_, already in that no live process has open (it leans on psutil_, already in
your dev venv, to check live mappings and fds) and ``--uds`` your dev venv, to check live mappings and fds) and ``--uds`` clears
clears socket files whose binder pid is dead. dead-binder sockets from Tractor's platform-specific runtime dir. It
also unconditionally removes ``registry@1616.sock``; do not run the UDS
sweep while a live registrar is serving from that default address.
Testing your own ``tractor`` app Testing your own ``tractor`` app
-------------------------------- --------------------------------

View File

@ -23,18 +23,10 @@ async def endpoint(
await trio.sleep_forever() await trio.sleep_forever()
async def spawn_and_open_ep( async def open_ep(
an: tractor.ActorNursery, ptl: tractor.Portal,
i: int, i: int,
) -> None: ) -> None:
'''
Spawn a subactor, start a remote `endpoint()`-task in it.
'''
ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
ctx: tractor.Context ctx: tractor.Context
async with ptl.open_context(endpoint) as ( async with ptl.open_context(endpoint) as (
ctx, ctx,
@ -47,7 +39,33 @@ async def spawn_and_open_ep(
await ctx.wait_for_result() await ctx.wait_for_result()
async def main(): async def spawn_and_open_ep(
an: tractor.ActorNursery,
i: int,
maybe_ptl: tractor.Portal|None = None,
) -> None:
'''
Spawn a subactor, start a remote `endpoint()`-task in it.
'''
if maybe_ptl is None:
maybe_ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
await open_ep(
ptl=maybe_ptl,
i=i,
)
async def main(
# spawn subs concurrently (in bg `trio.Task`s) so each
# actor's cold `import tractor` (~0.4s, see #470) overlaps
# instead of stacking; once forkserver (#463) lands, spawn
# is cheap enough to just loop sequentially.
spawn_subs_in_bg_tasks: bool = True,
):
''' '''
Spawn a subactor-per-CPU then self-destruct the cluster. Spawn a subactor-per-CPU then self-destruct the cluster.
@ -60,17 +78,21 @@ async def main():
# https://github.com/goodboy/tractor/pull/463 # https://github.com/goodboy/tractor/pull/463
# start_method='main_thread_forkserver', # start_method='main_thread_forkserver',
) as an, ) as an,
# spawn subs concurrently (in bg `trio.Task`s) so each
# actor's cold `import tractor` (~0.4s, see #470) overlaps
# instead of stacking; once forkserver (#463) lands, spawn
# is cheap enough to just loop sequentially.
trio.open_nursery() as tn, trio.open_nursery() as tn,
): ):
for i in range(cpu_count()): for i in range(cpu_count()):
maybe_ptl: tractor.Portal|None = None
if not spawn_subs_in_bg_tasks:
maybe_ptl: tractor.Portal = await an.start_actor(
name=f'worker_{i}',
enable_modules=[__name__],
)
tn.start_soon( tn.start_soon(
spawn_and_open_ep, spawn_and_open_ep,
an, an,
i, i,
maybe_ptl,
) )
destruct_in: int = 2 destruct_in: int = 2
print( print(

View File

@ -6,7 +6,8 @@ subactor inherits the preference.
Every channel address is a filesystem socket path (no TCP port Every channel address is a filesystem socket path (no TCP port
in sight!) and, as a kernel-provided bonus, the peer's pid is in sight!) and, as a kernel-provided bonus, the peer's pid is
exchanged for free via `SO_PEERCRED`. exchanged for free via `SO_PEERCRED` on linux,
`LOCAL_PEERPID` on macOS.
''' '''
import os import os
@ -42,7 +43,7 @@ async def main() -> None:
# (named for the root registrar) this channel rode in # (named for the root registrar) this channel rode in
# on, NOT a per-child path; the child-specific identity # on, NOT a per-child path; the child-specific identity
# we get for free is the kernel-reported peer pid (via # we get for free is the kernel-reported peer pid (via
# `SO_PEERCRED`). # `SO_PEERCRED` on linux, `LOCAL_PEERPID` on macOS).
print( print(
f'portal chan tpt proto: {raddr.proto_key!r}\n' f'portal chan tpt proto: {raddr.proto_key!r}\n'
f'listener sock file: {raddr.sockpath}\n' f'listener sock file: {raddr.sockpath}\n'

View File

@ -0,0 +1,4 @@
Fix Unix-domain-socket actor trees and registrar discovery on macOS.
Runtime sockets now use a short, owner-only runtime directory,
generated socket names remain within platform limits, and transient
or reset pre-handshake connections no longer destabilize discovery.

View File

@ -23,10 +23,10 @@ Two cleanup phases (run in order when both are enabled):
hard-crashing actor leaves leaked segments that hard-crashing actor leaves leaked segments that
nothing else GCs. nothing else GCs.
3. **UDS sweep** (`--uds` / `--uds-only`) — unlinks 3. **UDS sweep** (`--uds` / `--uds-only`) — unlinks socket
`${XDG_RUNTIME_DIR}/tractor/<name>@<pid>.sock` files files from Tractor's platform-specific default bindspace whose
whose binder pid is dead (or the `1616` registry binder pid is dead (or the `1616` registry sentinel). Needed
sentinel). Needed because the IPC server's because the IPC server's
`os.unlink()` cleanup lives in a `finally:` block `os.unlink()` cleanup lives in a `finally:` block
that doesn't always run on hard exits (SIGKILL, that doesn't always run on hard exits (SIGKILL,
escaped `KeyboardInterrupt`, etc.) — see issue #452. escaped `KeyboardInterrupt`, etc.) — see issue #452.
@ -137,8 +137,8 @@ def main() -> int:
action='store_true', action='store_true',
help=( help=(
'after process reap, also unlink orphaned ' 'after process reap, also unlink orphaned '
'${XDG_RUNTIME_DIR}/tractor/*.sock files ' 'sockets from Tractor\'s platform default '
'whose binder pid is dead (or the 1616 ' 'bindspace whose binder pid is dead (or the 1616 '
'registry sentinel). See issue #452.' 'registry sentinel). See issue #452.'
), ),
) )
@ -212,7 +212,9 @@ def main() -> int:
# --- phase 3: UDS sweep (opt-in) --- # --- phase 3: UDS sweep (opt-in) ---
if args.uds or args.uds_only: if args.uds or args.uds_only:
leaked_uds: list[str] = find_orphaned_uds() leaked_uds: list[str] = find_orphaned_uds(
include_registry_sentinel=True,
)
if not leaked_uds: if not leaked_uds:
print( print(
'[tractor-reap] no orphaned UDS sock-files ' '[tractor-reap] no orphaned UDS sock-files '

View File

@ -1307,19 +1307,31 @@ def test_ctxep_pauses_n_maybe_ipc_breaks(
if _non_linux: if _non_linux:
tpt: str = 'TCP' tpt: str = 'TCP'
assert_before( before: str = assert_before(
child, child,
['peer IPC channel closed abruptly?', ['peer IPC channel closed abruptly?',
'another task closed this fd', 'another task closed this fd',
'Debug lock request was CANCELLED?', 'Debug lock request was CANCELLED?',
f"'Msgpack{tpt}Stream' was already closed locally?",
f"TransportClosed: 'Msgpack{tpt}Stream' was already closed 'by peer'?",
] ]
# XXX races on whether these show/hit? # XXX races on whether these show/hit?
# 'Failed to REPl via `_pause()` You called `tractor.pause()` from an already cancelled scope!', # 'Failed to REPl via `_pause()` You called `tractor.pause()` from an already cancelled scope!',
# 'AssertionError', # 'AssertionError',
) )
# Error shipment and peer receive race after local close.
# Either diagnostic proves the transport was torn down.
closed_locally: str = (
f"'Msgpack{tpt}Stream' was already closed locally?"
)
closed_by_peer: str = (
f"TransportClosed: 'Msgpack{tpt}Stream' was "
f"already closed 'by peer'?"
)
assert (
closed_locally in before
or closed_by_peer in before
)
# OSc(ancel) the hanging tree # OSc(ancel) the hanging tree
do_ctlc( do_ctlc(
child=child, child=child,

View File

@ -1,21 +1,18 @@
''' '''
Discovery-suite fixtures, including the `daemon` Discovery-suite fixtures, including the `daemon` remote-registrar
remote-registrar subprocess used by the multi-program subprocess used by the multi-program discovery tests.
discovery tests.
Lives here (vs. the parent `tests/conftest.py`) Lives here (vs. the parent `tests/conftest.py`)
because `daemon` is a discovery-protocol primitive because `daemon` is a discovery-protocol primitive: it boots a child
boots a separate `tractor.run_daemon()` process whose that enters `open_root_actor()` and waits as a registrar peer for
sole purpose is to serve as a registrar peer for
discovery-roundtrip tests. Pytest fixtures inherit discovery-roundtrip tests. Pytest fixtures inherit
DOWNWARD through conftest hierarchy, so anything DOWNWARD through conftest hierarchy, so anything
under `tests/discovery/` automatically picks this up. under `tests/discovery/` automatically picks this up.
''' '''
from __future__ import annotations from __future__ import annotations
import os from pathlib import Path
import platform import platform
import socket
import subprocess import subprocess
import sys import sys
import time import time
@ -31,33 +28,27 @@ from ..conftest import (
def _wait_for_daemon_ready( def _wait_for_daemon_ready(
reg_addr: tuple, ready_path: Path,
tpt_proto: str,
*, *,
deadline: float = 10.0, deadline: float = 10.0,
poll_interval: float = 0.05, poll_interval: float = 0.05,
proc: subprocess.Popen|None = None, proc: subprocess.Popen|None = None,
) -> None: ) -> None:
''' '''
Active-poll the daemon's bind address until it Poll until the daemon reports completed actor startup.
accepts a connection (proving it has called
`bind() + listen()` and is ready to handle IPC).
Replaces the historical blind `time.sleep()` in the Replaces the historical blind `time.sleep()` in the
`daemon` fixture which was racy under load see `daemon` fixture which was racy under load see
`ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`. `ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`.
Uses stdlib `socket` directly (no trio runtime The child writes `ready_path` only after entering
bootstrap cost) sufficient because `open_root_actor()`, which guarantees all transport listeners are
`tractor.run_daemon()` doesn't return from serving without requiring a raw connection probe.
bootstrap until the runtime is fully ready to
accept IPC.
Raises `TimeoutError` on `deadline` exceeded. If Raises `TimeoutError` on `deadline` exceeded. If
`proc` is given, ALSO raises early if the daemon `proc` is given, ALSO raises early if the daemon
process exits non-zero before the deadline (catches process exits before the deadline (catches a daemon startup crash
daemon-startup-crash that the blind sleep used to that the blind sleep used to silently mask).
silently mask).
''' '''
end: float = time.monotonic() + deadline end: float = time.monotonic() + deadline
@ -70,43 +61,25 @@ def _wait_for_daemon_ready(
if proc is not None and proc.poll() is not None: if proc is not None and proc.poll() is not None:
raise RuntimeError( raise RuntimeError(
f'Daemon proc exited (rc={proc.returncode}) ' f'Daemon proc exited (rc={proc.returncode}) '
f'before becoming ready to accept on ' f'before reporting ready at {ready_path!r}'
f'{reg_addr!r}'
) )
try: try:
if tpt_proto == 'tcp': if ready_path.is_file():
# `socket.create_connection` does the if proc is not None and proc.poll() is not None:
# `socket() + connect()` dance with a raise RuntimeError(
# builtin timeout — perfect primitive f'Daemon proc exited (rc={proc.returncode}) '
# for a one-shot probe. f'after reporting ready at {ready_path!r}'
with socket.create_connection( )
reg_addr,
timeout=poll_interval,
):
return return
else:
# UDS — `reg_addr` is a `(filedir, sockname)`
# tuple per `tractor.ipc._uds.UDSAddress.unwrap`.
sockpath: str = os.path.join(*reg_addr)
sock = socket.socket(socket.AF_UNIX)
try:
sock.settimeout(poll_interval)
sock.connect(sockpath)
return
finally:
sock.close()
except ( except (
ConnectionRefusedError,
FileNotFoundError, FileNotFoundError,
OSError, OSError,
socket.timeout,
) as exc: ) as exc:
last_exc = exc last_exc = exc
time.sleep(poll_interval) time.sleep(poll_interval)
raise TimeoutError( raise TimeoutError(
f'Daemon never accepted on {reg_addr!r} within ' f'Daemon never reported ready at {ready_path!r} within '
f'{deadline}s (last connect-attempt exc: ' f'{deadline}s (last sentinel-state exc: {last_exc!r})'
f'{last_exc!r})'
) )
@ -136,18 +109,27 @@ def daemon(
) )
loglevel: str = 'info' loglevel: str = 'info'
ready_path: Path = (
Path(str(testdir.tmpdir))
/ 'daemon-ready'
)
ready_path.unlink(missing_ok=True)
code: str = ( code: str = (
"import tractor; " f'from pathlib import Path\n'
"tractor.run_daemon([], " f'import tractor\n'
"registry_addrs={reg_addrs}, " f'import trio\n'
"enable_transports={enable_tpts}, " f'\n'
"debug_mode={debug_mode}, " f'async def main():\n'
"loglevel={ll})" f' async with tractor.open_root_actor(\n'
).format( f' registry_addrs={[reg_addr]!r},\n'
reg_addrs=str([reg_addr]), f' enable_transports={[tpt_proto]!r},\n'
enable_tpts=str([tpt_proto]), f' debug_mode={debug_mode!r},\n'
ll="'{}'".format(loglevel) if loglevel else None, f' loglevel={loglevel!r},\n'
debug_mode=debug_mode, f' ):\n'
f' Path({str(ready_path)!r}).touch()\n'
f' await trio.sleep_forever()\n'
f'\n'
f'trio.run(main)\n'
) )
cmd: list[str] = [ cmd: list[str] = [
sys.executable, sys.executable,
@ -163,9 +145,9 @@ def daemon(
**kwargs, **kwargs,
) )
# Active-poll the daemon's bind address until it's # Poll the child's ready sentinel, published after actor startup,
# ready to accept connections — replaces the legacy # instead of connecting to its transport socket. This replaces
# blind `time.sleep(2.2)` which was racy under load # the legacy blind `time.sleep(2.2)` which was racy under load
# (see # (see
# `ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`). # `ai/conc-anal/test_register_duplicate_name_daemon_connect_race_issue.md`).
# #
@ -176,20 +158,22 @@ def daemon(
15.0 if (_non_linux and ci_env) 15.0 if (_non_linux and ci_env)
else 10.0 else 10.0
) )
try:
_wait_for_daemon_ready( _wait_for_daemon_ready(
reg_addr=reg_addr, ready_path=ready_path,
tpt_proto=tpt_proto,
deadline=deadline, deadline=deadline,
proc=proc, proc=proc,
) )
assert not proc.returncode assert not proc.returncode
yield proc yield proc
finally:
if proc.poll() is None:
sig_prog(proc, _INT_SIGNAL) sig_prog(proc, _INT_SIGNAL)
# XXX! yeah.. just be reaaal careful with this bc # NOTE: these blocking reads can hang when descendants retain
# sometimes it can lock up on the `_io.BufferedReader` # inherited pipe descriptors. Keep teardown signaling above
# and hang.. # them and avoid adding subprocesses outside the actor tree.
# #
# NB, drain happens at TEARDOWN (post-yield), so the # NB, drain happens at TEARDOWN (post-yield), so the
# test body has its chance to read `proc.stderr` # test body has its chance to read `proc.stderr`
@ -219,5 +203,4 @@ def daemon(
) )
if rc < 0: if rc < 0:
raise RuntimeError(msg) raise RuntimeError(msg)
test_log.error(msg) test_log.error(msg)

View File

@ -0,0 +1,76 @@
'''
Discovery daemon fixture regressions.
This module imports private helpers from the sibling
`tests.discovery.conftest` plugin to exercise that fixture machinery
directly, rather than testing a production `tractor` API.
'''
from unittest.mock import (
call,
Mock,
)
from .conftest import _wait_for_daemon_ready
def test_daemon_ready_check_does_not_connect(
monkeypatch,
tmp_path,
):
'''
Observe completed daemon startup without a raw connection.
The old UDS readiness helper connected and immediately closed. That
entered Tractor's actor-handshake handler with no `Aid` payload and
destabilized the remote registrar on macOS before discovery tests
started. This test creates the child sentinel, forbids all socket
construction and connection helpers, then proves readiness returns
without touching the transport layer.
'''
ready_path = tmp_path / 'daemon-ready'
ready_path.touch()
socket_ctor = Mock(side_effect=AssertionError('socket opened'))
connect = Mock(side_effect=AssertionError('socket connected'))
monkeypatch.setattr('socket.socket', socket_ctor)
monkeypatch.setattr('socket.create_connection', connect)
_wait_for_daemon_ready(
ready_path=ready_path,
deadline=.1,
poll_interval=.01,
)
socket_ctor.assert_not_called()
connect.assert_not_called()
def test_daemon_ready_check_backs_off(monkeypatch):
'''
Back off while waiting for the child startup sentinel.
The sentinel may appear after several parent polling intervals.
A deterministic false/false/true path sequence proves the helper
sleeps between unsuccessful observations instead of hot-spinning
and starving a booting daemon on constrained CI workers.
'''
ready_path = Mock()
ready_path.is_file.side_effect = [False, False, True]
sleep = Mock()
monotonic = Mock(side_effect=[0, 0, 0, 0])
monkeypatch.setattr('time.sleep', sleep)
monkeypatch.setattr('time.monotonic', monotonic)
_wait_for_daemon_ready(
ready_path=ready_path,
deadline=.2,
poll_interval=.01,
)
assert ready_path.is_file.call_count == 3
assert sleep.call_args_list == [
call(.01),
call(.01),
]

View File

@ -2,23 +2,211 @@
`open_root_actor(tpt_bind_addrs=...)` test suite. `open_root_actor(tpt_bind_addrs=...)` test suite.
Verify all three runtime code paths for explicit IPC-server Verify all three runtime code paths for explicit IPC-server
bind-address selection in `_root.py`: bind-address selection in `_root.py` and registry probing in
`discovery._api`:
1. Non-registrar, no explicit bind -> random addrs from registry proto 1. Non-registrar, no explicit bind -> random addrs from registry proto
2. Registrar, no explicit bind -> binds to registry_addrs 2. Registrar, no explicit bind -> binds to registry_addrs
3. Explicit bind given -> wraps via `wrap_address()` and uses them 3. Explicit bind given -> wraps via `wrap_address()` and uses them
''' '''
from contextlib import asynccontextmanager as acm
from unittest.mock import (
AsyncMock,
call,
Mock,
)
import pytest import pytest
import trio import trio
import tractor import tractor
from tractor.discovery import _api
from tractor.discovery._addr import ( from tractor.discovery._addr import (
wrap_address, wrap_address,
) )
from tractor.discovery._multiaddr import mk_maddr from tractor.discovery._multiaddr import mk_maddr
from tractor.ipc import _connect_chan
from tractor._testing.addr import get_rando_addr from tractor._testing.addr import get_rando_addr
def test_registry_probe_retries_transient_handshake(
monkeypatch: pytest.MonkeyPatch,
):
'''
Retry a connected registrar after transient handshake timeout.
Loaded macOS runners can accept the transport while delaying the
actor handshake beyond one second. Treating that first timeout as
final makes a healthy remote daemon look occupied and cascades into
discovery failures. This deterministic fake fails once, succeeds
on the second complete handshake, and proves one bounded backoff.
'''
async def stall_handshake(**kwargs):
await trio.sleep_forever()
first_handshake = AsyncMock(side_effect=stall_handshake)
second_handshake = AsyncMock(
return_value=tractor.msg.Aid(
name='registrar',
uuid='registrar-uuid',
pid=1234,
is_registrar=True,
),
)
chans = [
Mock(_do_handshake=first_handshake),
Mock(_do_handshake=second_handshake),
]
closed: list[object] = []
@acm
async def connect_chan(addr, close_timeout):
assert close_timeout == .2
chan = chans[len(closed)]
try:
yield chan
finally:
closed.append(chan)
sleep = AsyncMock()
monkeypatch.setattr(_api, '_connect_chan', connect_chan)
monkeypatch.setattr(_api.trio, 'sleep', sleep)
async def main():
status = await _api._probe_registry(
addr=wrap_address(('127.0.0.1', 1616)),
timeout=.3,
attempt_timeout=.1,
max_attempts=3,
retry_delay=.01,
)
assert status == 'registrar'
trio.run(main)
first_handshake.assert_awaited_once()
second_handshake.assert_awaited_once()
assert first_handshake.await_args.kwargs['timeout'] == .1
assert second_handshake.await_args.kwargs['timeout'] == .1
assert closed == chans
sleep.assert_has_awaits([call(.01)])
def test_probe_channel_close_is_bounded(
monkeypatch: pytest.MonkeyPatch,
):
'''
Bound shielded channel cleanup after a registry probe.
`_connect_chan()` shields `.aclose()` so cancellation cannot leak
ordinary channels. A stalled close previously let registry probing
exceed every connect and handshake deadline. This fake close never
completes; the explicit cleanup allowance must still return control
to the caller without cancelling its surrounding task.
'''
chan = Mock()
chan.aclose = AsyncMock(side_effect=trio.sleep_forever)
monkeypatch.setattr(
tractor.Channel,
'from_addr',
AsyncMock(return_value=chan),
)
async def main():
with trio.fail_after(.5):
async with _connect_chan(
('127.0.0.1', 1616),
close_timeout=.01,
):
pass
trio.run(main)
chan.aclose.assert_awaited_once()
def test_transport_only_listener_is_not_registrar():
'''
Require a Tractor handshake before accepting a registry address.
The old election probe marked an address live after transport
connect alone. A non-Tractor listener, or a registrar still
failing its initial handshake, was therefore selected as the
remote registry. This test accepts the probe and closes it without
replying, then proves `open_root_actor()` rejects that occupied
endpoint instead of selecting it or binding over it.
'''
async def transport_only_handler(
stream: trio.SocketStream,
) -> None:
await stream.aclose()
async def main():
listeners = await trio.open_tcp_listeners(0)
listener = listeners[0]
sockname = listener.socket.getsockname()
reg_addr: tuple[str, int] = (
sockname[0],
sockname[1],
)
async with trio.open_nursery() as tn:
tn.start_soon(
trio.serve_listeners,
transport_only_handler,
listeners,
)
with pytest.raises(
RuntimeError,
match='occupied but did not answer',
):
async with tractor.open_root_actor(
registry_addrs=[reg_addr],
enable_transports=['tcp'],
):
pytest.fail('foreign listener selected as registrar')
tn.cancel_scope.cancel()
trio.run(main)
def test_registry_probe_preserves_no_peers_state(
reg_addr: tuple,
tpt_proto: str,
):
'''
Keep an idle registrar peer-free after an election probe.
Probe handshakes exchange registrar capability but must not enter
`IPCServer._peers`. Resetting `_no_more_peers` before identifying a
probe left an idle registrar reporting phantom peers and delayed
shutdown. This test probes the live local registrar and proves its
peer map and no-peers event remain unchanged afterward.
'''
async def main():
async with tractor.open_root_actor(
registry_addrs=[reg_addr],
enable_transports=[tpt_proto],
):
actor = tractor.current_actor()
server = actor.ipc_server
probe_status = await _api._probe_registry(
addr=wrap_address(reg_addr),
)
assert probe_status == 'registrar'
await trio.sleep(0)
assert not server._peers
assert server._no_more_peers.is_set()
trio.run(main)
# ------------------------------------------------------------------ # ------------------------------------------------------------------
# helpers # helpers
# ------------------------------------------------------------------ # ------------------------------------------------------------------

View File

@ -3,14 +3,21 @@ Unit-ish tests for specific IPC transport protocol backends.
''' '''
from __future__ import annotations from __future__ import annotations
import os
from pathlib import Path from pathlib import Path
import socket
import stat
import sys
import tempfile
from types import SimpleNamespace
from unittest.mock import Mock
import pytest import pytest
import trio import trio
import tractor import tractor
from tractor import Actor from tractor import Actor
from tractor.runtime import _state
from tractor.discovery import _addr from tractor.discovery import _addr
from tractor.runtime import _state
@pytest.fixture @pytest.fixture
@ -31,6 +38,381 @@ def bindspace_dir_str() -> str:
bs_dir.rmdir() bs_dir.rmdir()
def test_macos_rt_dir_fits_uds_path_limit(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Keep the default Darwin UDS bindpath below its 104-byte limit.
`platformdirs` normally places the runtime directory below the
long `~/Library/Caches/TemporaryItems` path. Pytest also assigns
a deeply nested temporary home, so appending a registry socket
name made every macOS UDS listener fail with `AF_UNIX path too
long`. This test simulates Darwin and an intentionally long
platformdirs result, then proves `get_rt_dir()` uses the short
system temporary directory and leaves room for the socket name.
'''
long_rt_dir: Path = tmp_path / ('long' * 30)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(long_rt_dir / appname),
)
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
rt_dir: Path = _state.get_rt_dir()
sockpath: Path = (
Path('/tmp')
/ f'tractor-{os.getuid()}'
/ 'registry@1616.sock'
)
assert rt_dir == tmp_path / f'tractor-{os.getuid()}'
assert len(os.fsencode(sockpath)) < 104
assert stat.S_IMODE(rt_dir.stat().st_mode) == 0o700
def test_macos_rt_dir_rejects_symlink(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject a pre-created symlink at the Darwin runtime path.
Darwin uses the predictable `/tmp/tractor-<uid>` path to stay
below its `AF_UNIX` limit. A hostile local user could otherwise
point that path at a victim-owned directory and make
`get_rt_dir()` chmod or place sockets in the symlink target. The
test replaces `/tmp` with a controlled directory, installs the
malicious link, and proves non-following validation rejects it.
'''
runtime_link: Path = tmp_path / f'tractor-{os.getuid()}'
target_dir: Path = tmp_path / 'target'
target_dir.mkdir(mode=0o755)
runtime_link.symlink_to(target_dir, target_is_directory=True)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
with pytest.raises(PermissionError, match='Unsafe Darwin'):
_state.get_rt_dir()
assert stat.S_IMODE(target_dir.stat().st_mode) == 0o755
def test_reaper_uses_default_uds_bindspace(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Sweep the same platform-specific bindspace used by UDS actors.
The reaper previously consulted only `XDG_RUNTIME_DIR`, missing
Darwin sockets after the runtime moved to `/tmp/tractor-<uid>`.
This test replaces `UDSAddress.def_bindspace` and proves the test
harness resolves that shared transport default directly.
'''
from tractor._testing import _reap
from tractor.ipc._uds import UDSAddress
monkeypatch.setattr(
UDSAddress,
'def_bindspace',
tmp_path,
)
assert _reap.get_uds_dir() == str(tmp_path)
def test_automatic_reaper_preserves_registry_sentinel(
monkeypatch: pytest.MonkeyPatch,
):
'''
Reserve unconditional registry cleanup for the explicit CLI.
The `registry@1616.sock` suffix does not encode its binder PID, so
automatic pytest cleanup cannot distinguish a leak from another
live registrar. This test creates registry and actor sockets,
proves the default sweep selects only the dead actor, then proves
explicit sentinel inclusion retains the CLI's documented behavior.
'''
from tractor._testing import _reap
with tempfile.TemporaryDirectory(
prefix='tractor-reap-',
dir='/tmp',
) as tmpdir:
bindspace: Path = Path(tmpdir)
registry_path: Path = bindspace / 'registry@1616.sock'
actor_path: Path = bindspace / 'worker@1234.sock'
socks: list[socket.socket] = []
for path in (registry_path, actor_path):
sock = socket.socket(socket.AF_UNIX)
sock.bind(str(path))
socks.append(sock)
monkeypatch.setattr(_reap, '_is_alive', lambda pid: False)
try:
assert _reap.find_orphaned_uds(
uds_dir=str(bindspace),
) == [str(actor_path)]
assert set(
_reap.find_orphaned_uds(
uds_dir=str(bindspace),
include_registry_sentinel=True,
)
) == {
str(registry_path),
str(actor_path),
}
finally:
for sock in socks:
sock.close()
def test_rt_dir_rejects_non_directory(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Preserve the non-Darwin runtime-directory type contract.
Replacing `Path.is_dir()` with unguarded `lstat()` briefly made
existing files look like valid runtime directories on Linux.
This test points `platformdirs` at a regular file and proves
`get_rt_dir()` rejects it during initialization.
'''
rt_file: Path = tmp_path / 'runtime-file'
rt_file.touch()
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_file),
)
with pytest.raises(
PermissionError,
match='Unsafe POSIX',
):
_state.get_rt_dir()
new_rt_dir: Path = tmp_path / 'new-runtime-dir'
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(new_rt_dir),
)
assert _state.get_rt_dir() == new_rt_dir
assert stat.S_IMODE(new_rt_dir.stat().st_mode) == 0o700
def test_linux_rt_dir_secures_existing_path(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Enforce owner-only access on an existing Linux runtime directory.
Linux previously accepted any existing directory returned by
`platformdirs`, without checking ownership or correcting a
traversable mode. This test creates an owner-controlled `0o755`
directory and proves `get_rt_dir()` normalizes the managed
bindspace to `0o700` before returning it.
'''
rt_dir: Path = tmp_path / 'tractor'
rt_dir.mkdir(mode=0o755)
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_dir),
)
assert _state.get_rt_dir() == rt_dir
assert stat.S_IMODE(rt_dir.stat().st_mode) == 0o700
def test_linux_rt_dir_rejects_foreign_owner(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject an existing Linux runtime directory owned by another UID.
A pre-created bindspace must never be made private with `chmod`
until ownership is verified. This test makes the current process
appear to have a different UID and proves `get_rt_dir()` rejects
the directory without changing its original mode.
'''
rt_dir: Path = tmp_path / 'tractor'
rt_dir.mkdir(mode=0o755)
original_mode: int = stat.S_IMODE(rt_dir.stat().st_mode)
monkeypatch.setattr(sys, 'platform', 'linux')
monkeypatch.setattr(
'platformdirs.user_runtime_dir',
lambda appname: str(rt_dir),
)
monkeypatch.setattr(
os,
'getuid',
lambda: rt_dir.stat().st_uid + 1,
)
with pytest.raises(
PermissionError,
match='Unsafe POSIX',
):
_state.get_rt_dir()
assert stat.S_IMODE(rt_dir.stat().st_mode) == original_mode
def test_macos_rt_dir_rejects_intermediate_symlink(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
):
'''
Reject symlinks in nested Darwin runtime subdirectories.
The earlier final-component check allowed `link/child` to follow
an intermediate symlink and create `child` outside the secured
runtime root. This test installs that link and proves traversal
stops before anything is created in its target.
'''
rt_root: Path = tmp_path / f'tractor-{os.getuid()}'
target_dir: Path = tmp_path / 'target'
rt_root.mkdir(mode=0o700)
target_dir.mkdir()
(rt_root / 'link').symlink_to(
target_dir,
target_is_directory=True,
)
monkeypatch.setattr(sys, 'platform', 'darwin')
monkeypatch.setattr(_state, '_DARWIN_TMPDIR', tmp_path)
with pytest.raises(PermissionError, match='Unsafe Darwin'):
_state.get_rt_dir(subdir='link/child')
assert not (target_dir / 'child').exists()
@pytest.mark.parametrize(
('platform_name', 'path_limit'),
[
('darwin', 104),
('linux', 108),
],
)
def test_uds_sockname_compaction(
monkeypatch: pytest.MonkeyPatch,
platform_name: str,
path_limit: int,
):
'''
Keep generated actor sockets safe and below Darwin's byte limit.
Actor names are unrestricted identity strings. A long, multibyte,
or path-like name previously produced overlong or escaping socket
paths. These cases prove `UDSAddress.get_sockname()` preserves a
short legacy name, deterministically compacts unsafe names, keeps
the reaper's `@pid.sock` suffix, and stays within Darwin's byte
limit.
'''
from tractor.ipc._uds import UDSAddress
bindspace: Path = Path('/tmp/tractor-501')
pid: int = 12345
from tractor.ipc import _uds
monkeypatch.setattr(sys, 'platform', platform_name)
monkeypatch.setattr(_uds, '_SUN_PATH_LIMIT', path_limit)
short: Path = UDSAddress.get_sockname(
name='worker',
pid=pid,
bindspace=bindspace,
)
long_name: str = 'actor-' + ('\u00e9' * 100)
compact: Path = UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=bindspace,
)
unsafe: Path = UDSAddress.get_sockname(
name='../worker',
pid=pid,
bindspace=bindspace,
)
assert short == Path(f'worker@{pid}.sock')
assert compact == UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=bindspace,
)
assert compact.name.endswith(f'@{pid}.sock')
assert unsafe.parent == Path('.')
assert '..' not in unsafe.name
assert len(os.fsencode(bindspace / compact)) < path_limit
with pytest.raises(ValueError) as exc_info:
UDSAddress.get_sockname(
name=long_name,
pid=pid,
bindspace=Path('/tmp') / ('x' * 90),
)
errmsg: str = str(exc_info.value)
assert 'leaves no room' in errmsg
assert 'name was unsafe: False' in errmsg
assert 'name was over budget: True' in errmsg
assert f'AF_UNIX path limit: {path_limit}' in errmsg
def test_uds_reaper_ignores_unreconstructable_path(
monkeypatch: pytest.MonkeyPatch,
):
'''
Keep post-kill UDS cleanup best-effort on path overflow.
`unlink_uds_bind_addrs()` reconstructs a self-assigned socket from
the dead actor's name and PID. An over-budget bindspace makes that
naming helper raise before `os.unlink()`; propagating the error
would replace the original supervision outcome after the child was
already killed. This test forces overflow and proves cleanup skips
reconstruction without attempting an unlink or raising.
'''
from tractor.ipc import _uds
from tractor.spawn import _reap
long_bindspace: Path = Path('/tmp') / ('x' * 120)
proc = SimpleNamespace(pid=12345)
subactor = SimpleNamespace(
aid=SimpleNamespace(name='worker'),
)
unlink = Mock()
monkeypatch.setattr(
_uds.UDSAddress,
'def_bindspace',
long_bindspace,
)
monkeypatch.setattr(_reap.os, 'unlink', unlink)
_reap.unlink_uds_bind_addrs(
proc=proc,
subactor=subactor,
)
unlink.assert_not_called()
def test_uds_bindspace_created_implicitly( def test_uds_bindspace_created_implicitly(
debug_mode: bool, debug_mode: bool,
bindspace_dir_str: str, bindspace_dir_str: str,

View File

@ -3,7 +3,13 @@ High-level `.ipc._server` unit tests.
''' '''
from __future__ import annotations from __future__ import annotations
import errno
from unittest.mock import (
AsyncMock,
Mock,
)
import msgspec
import pytest import pytest
import trio import trio
from tractor import ( from tractor import (
@ -14,6 +20,11 @@ from tractor import (
from tractor._testing.addr import ( from tractor._testing.addr import (
get_rando_addr, get_rando_addr,
) )
from tractor._exceptions import TransportClosed
from tractor.ipc._chan import Channel
from tractor.ipc import _server
from tractor.ipc._transport import MsgpackTransport
from tractor.msg.types import Aid
# TODO, use/check-roundtripping with some of these wrapper types? # TODO, use/check-roundtripping with some of these wrapper types?
# #
# from .._addr import Address # from .._addr import Address
@ -23,6 +34,165 @@ from tractor._testing.addr import (
# from ._tcp import TCPAddress # from ._tcp import TCPAddress
def test_send_normalizes_only_grouped_peer_resets():
'''
Normalize only all-peer-close grouped transport failures.
A UDS peer may disconnect before completing the actor handshake.
Darwin can report the server's first handshake write as
`ECONNRESET`, wrapped by `trio.BrokenResourceError` and potentially
nested in an `ExceptionGroup`. This fake stream first groups reset
and broken-pipe branches, proving `.send()` normalizes a complete
peer-close tree to `TransportClosed`. It then groups a reset with
an unrelated `ValueError`, proving the mixed failure remains a
`trio.BrokenResourceError` instead of hiding the application error.
'''
def broken_resource(err_no: int) -> trio.BrokenResourceError:
try:
raise OSError(
err_no,
'Peer closed',
)
except OSError as peer_err:
try:
raise trio.BrokenResourceError from peer_err
except trio.BrokenResourceError as broken_err:
return broken_err
class GroupedFailureStream:
def __init__(self, exceptions: list[Exception]) -> None:
self.exceptions = exceptions
async def send_all(self, data: bytes) -> None:
grouped_err = ExceptionGroup(
'concurrent send failures',
self.exceptions,
)
raise trio.BrokenResourceError from grouped_err
async def main():
transport = object.__new__(MsgpackTransport)
transport.stream = GroupedFailureStream([
broken_resource(errno.ECONNRESET),
broken_resource(errno.EPIPE),
])
transport._send_lock = trio.StrictFIFOLock()
transport._laddr = 'local'
transport._raddr = 'remote'
transport._task = trio.lowlevel.current_task()
with pytest.raises(TransportClosed) as exc_info:
await transport.send(
{'probe': True},
strict_types=False,
)
grouped_err = exc_info.value.src_exc.__cause__
assert isinstance(grouped_err, ExceptionGroup)
assert len(grouped_err.exceptions) == 2
transport.stream = GroupedFailureStream([
ValueError('unrelated failure'),
broken_resource(errno.ECONNRESET),
])
with pytest.raises(trio.BrokenResourceError) as exc_info:
await transport.send(
{'probe': True},
strict_types=False,
)
grouped_err = exc_info.value.__cause__
assert isinstance(grouped_err, ExceptionGroup)
assert isinstance(grouped_err.exceptions[0], ValueError)
trio.run(main)
def test_handshake_normalizes_decode_error():
'''
Keep malformed pre-handshake frames out of the service nursery.
A non-msgpack peer can trigger `msgspec.DecodeError` before a
remote `Aid` exists. Letting that decoder error escape the inbound
handler cancels the actor's shared IPC nursery. This fake channel
proves `_do_handshake()` presents only `TransportClosed` upward.
'''
chan = object.__new__(Channel)
chan.send = AsyncMock()
chan.recv = AsyncMock(
side_effect=msgspec.DecodeError('malformed handshake'),
)
async def main():
with pytest.raises(TransportClosed) as exc_info:
await chan._do_handshake(
aid=Aid(
name='local',
uuid='local-uuid',
pid=1234,
),
timeout=.1,
)
assert isinstance(
exc_info.value.src_exc,
msgspec.DecodeError,
)
trio.run(main)
def test_server_uses_independent_handshake_timeout(
monkeypatch: pytest.MonkeyPatch,
):
'''
Give ordinary actor handshakes a distinct, generous deadline.
Registry probes use short retries, but ordinary portal and child
connections do not retry. Applying the probe's one-second timeout
in the server can terminate a valid delayed child and leave its
parent blocked in `IPCServer.wait_for_peer()`. This handler fake
proves the server uses its separate pre-registration budget.
'''
handshake = AsyncMock(
side_effect=TransportClosed(message='stop after assertion'),
)
chan = Mock(_do_handshake=handshake)
actor = Mock(
aid=Aid(
name='local',
uuid='local-uuid',
pid=1234,
),
)
monkeypatch.setattr(
Channel,
'from_stream',
Mock(return_value=chan),
)
monkeypatch.setattr(
_server._state,
'current_actor',
Mock(return_value=actor),
)
async def main():
await _server.handle_stream_from_peer(
stream=Mock(),
server=Mock(),
)
trio.run(main)
handshake.assert_awaited_once_with(
aid=actor.aid,
timeout=_server._PRE_REG_HANDSHAKE_TIMEOUT,
)
assert _server._PRE_REG_HANDSHAKE_TIMEOUT == 10
@pytest.mark.parametrize( @pytest.mark.parametrize(
'_tpt_proto', '_tpt_proto',
['uds', 'tcp'] ['uds', 'tcp']

View File

@ -5,11 +5,13 @@ Let's make sure them docs work yah?
from contextlib import contextmanager from contextlib import contextmanager
import itertools import itertools
import os import os
import signal
import sys import sys
import subprocess import subprocess
import platform import platform
import shutil import shutil
from typing import Callable from typing import Callable
from unittest.mock import Mock
import pytest import pytest
import tractor import tractor
@ -21,6 +23,184 @@ _non_linux: bool = platform.system() != 'Linux'
_friggin_macos: bool = platform.system() == 'Darwin' _friggin_macos: bool = platform.system() == 'Darwin'
def _kill_proc_tree(proc: subprocess.Popen) -> None:
'''
Terminate an example process and its POSIX descendants.
'''
try:
if platform.system() == 'Windows':
proc.kill()
else:
os.killpg(proc.pid, signal.SIGKILL)
except ProcessLookupError:
pass
def _reap_killed_proc(
proc: subprocess.Popen,
) -> tuple[bytes, bytes]:
'''
Reap a killed process without waiting on Windows descendants.
'''
if platform.system() != 'Windows':
return proc.communicate()
proc.wait(timeout=5)
if proc.stdin:
proc.stdin.close()
if proc.stdout:
proc.stdout.close()
if proc.stderr:
proc.stderr.close()
return b'', b''
def _wait_for_proc(
proc: subprocess.Popen,
timeout: float,
test_log: tractor.log.StackLevelAdapter,
) -> None:
'''
Wait for an example process and surface its captured output.
'''
try:
out, err = proc.communicate(timeout=timeout)
except subprocess.TimeoutExpired as timeout_exc:
test_log.exception(
f'Example failed to finish within {timeout}s ??\n'
)
_kill_proc_tree(proc)
out, err = _reap_killed_proc(proc)
if platform.system() == 'Windows':
out = timeout_exc.output or b''
err = timeout_exc.stderr or b''
errmsg: str = err.decode(errors='replace')
# NOTE: always include captured stdout and stderr for a non-zero
# exit. Depending on the final stderr line previously hid grouped
# exception diagnostics; see GH #473.
#
# The prior impl only raised when the LAST stderr
# line contained 'Error', swallowing any crash whose
# traceback ends in a non-`XxxError:` line; in
# particular EVERY `tractor` root-actor crash ends
# with the strict-EG collapse note,
# '( ^^^ this exc was collapsed from a group ^^^ )',
# so ALL such failures were reduced to a bare
# `assert 1 == 0` in CI logs.. see GH #473.
rc: int|None = proc.returncode
if rc:
outmsg: str = out.decode(errors='replace')
raise Exception(
f'Example script exited with rc={rc} !?\n'
f'\n'
f'stdout:\n'
f'{outmsg}\n'
f'\n'
f'stderr:\n'
f'{errmsg}\n'
)
# if we get some gnarly output let's aggregate and raise
if errmsg:
errlines = errmsg.splitlines()
last_error = errlines[-1]
if (
'Error' in last_error
# XXX: currently we print this to console, but maybe
# shouldn't eventually once we figure out what's
# a better way to be explicit about aio side
# cancels?
and
'asyncio.exceptions.CancelledError' not in last_error
):
raise Exception(errmsg)
assert proc.returncode == 0
def test_wait_for_failed_example_captures_output():
'''
Preserve diagnostics from a subprocess which already exited.
The previous `poll()` guard skipped `communicate()` when a fast
failure returned a non-zero status before the parent checked it.
Its stdout and stderr were therefore reported as empty. This
fake process begins with `returncode=1` and returns non-UTF-8
output, proving the helper always drains both pipes and replaces
undecodable bytes without hiding the original process failure.
'''
proc = Mock()
proc.returncode = 1
proc.communicate.return_value = (
b'stdout\xff',
b'stderr\xff',
)
with pytest.raises(Exception) as exc_info:
_wait_for_proc(
proc=proc,
timeout=1,
test_log=Mock(),
)
proc.communicate.assert_called_once_with(timeout=1)
errmsg: str = str(exc_info.value)
assert 'stdout\ufffd' in errmsg
assert 'stderr\ufffd' in errmsg
@pytest.mark.skipif(
platform.system() == 'Windows',
reason='POSIX process groups are unavailable on Windows',
)
def test_wait_for_timed_out_example_reaps_group(
monkeypatch: pytest.MonkeyPatch,
):
'''
Kill the example process group and reap its leader on timeout.
The old timeout branch killed only the immediate process and
never drained it. Actor descendants could retain the capture
pipes while the leader remained unreaped, hanging CI until its
job timeout. This fake process raises `TimeoutExpired` on the
timed wait and completes on the second `communicate()` call;
the assertions prove group-directed `SIGKILL` precedes that
final drain and leaves a concrete non-zero return code.
'''
proc = Mock()
proc.pid = 1234
def communicate(timeout=None):
if timeout is not None:
raise subprocess.TimeoutExpired('example', timeout)
proc.returncode = -signal.SIGKILL
return b'', b'timed out'
proc.communicate.side_effect = communicate
killpg = Mock()
monkeypatch.setattr(os, 'killpg', killpg)
with pytest.raises(Exception, match='timed out'):
_wait_for_proc(
proc=proc,
timeout=.01,
test_log=Mock(),
)
killpg.assert_called_once_with(1234, signal.SIGKILL)
assert proc.communicate.call_count == 2
assert proc.returncode == -signal.SIGKILL
@pytest.fixture @pytest.fixture
def run_example_in_subproc( def run_example_in_subproc(
loglevel: str, loglevel: str,
@ -61,14 +241,14 @@ def run_example_in_subproc(
] ]
else: else:
script_file = testdir.makefile('.py', script_code) script_file = testdir.makefile('.py', script_code)
kwargs['start_new_session'] = True
cmdargs = [ cmdargs = [
sys.executable, sys.executable,
str(script_file), str(script_file),
] ]
# XXX: BE FOREVER WARNED: if you enable lots of tractor logging # Captured pipes are drained by `_wait_for_proc()` while the
# in the subprocess it may cause infinite blocking on the pipes # example runs.
# due to backpressure!!!
proc = testdir.popen( proc = testdir.popen(
cmdargs, cmdargs,
stdin=subprocess.PIPE, stdin=subprocess.PIPE,
@ -77,9 +257,20 @@ def run_example_in_subproc(
**kwargs, **kwargs,
) )
assert not proc.returncode assert not proc.returncode
try:
yield proc yield proc
proc.wait() except BaseException:
assert proc.returncode == 0 if proc.poll() is None:
try:
_kill_proc_tree(proc)
_reap_killed_proc(proc)
except Exception:
pass
raise
else:
if proc.poll() is None:
_kill_proc_tree(proc)
_reap_killed_proc(proc)
yield run yield run
@ -145,21 +336,6 @@ def test_example(
'This test does run just fine "in person" however..' 'This test does run just fine "in person" however..'
) )
if (
'uds_transport_actor_tree' in ex_file
and
_friggin_macos
and
ci_env
):
pytest.skip(
'UDS-transport example reliably fails on macOS CI.\n'
'UDS-on-macOS is otherwise un-exercised by the matrix\n'
'(no `tpt_proto=uds` macOS job), so this new example is\n'
'the first to surface it; the macOS UDS path needs\n'
'root-causing. Passes on Linux.'
)
from .conftest import cpu_perf_headroom from .conftest import cpu_perf_headroom
timeout: float = ( timeout: float = (
@ -178,33 +354,8 @@ def test_example(
code = ex.read() code = ex.read()
with run_example_in_subproc(code) as proc: with run_example_in_subproc(code) as proc:
err = None _wait_for_proc(
try: proc=proc,
if not proc.poll(): timeout=timeout,
_, err = proc.communicate(timeout=timeout) test_log=test_log,
except subprocess.TimeoutExpired as e:
test_log.exception(
f'Example failed to finish within {timeout}s ??\n'
) )
proc.kill()
err = e.stderr
# if we get some gnarly output let's aggregate and raise
if err:
errmsg = err.decode()
errlines = errmsg.splitlines()
last_error = errlines[-1]
if (
'Error' in last_error
# XXX: currently we print this to console, but maybe
# shouldn't eventually once we figure out what's
# a better way to be explicit about aio side
# cancels?
and
'asyncio.exceptions.CancelledError' not in last_error
):
raise Exception(errmsg)
assert proc.returncode == 0

View File

@ -20,6 +20,18 @@ from typing import (
import pytest import pytest
import trio import trio
import tractor import tractor
# `infect_asyncio` mode is unsupported on Windows (asyncio's
# `ProactorEventLoop` is incompatible with our `trio` guest-mode
# interop and currently hangs/crashes the run). Skip the module on
# Windows so the CI leg completes + reports the rest of the suite.
import platform
if platform.system() == 'Windows':
pytest.skip(
'infect_asyncio mode is unsupported on Windows',
allow_module_level=True,
)
from tractor import ( from tractor import (
current_actor, current_actor,
Actor, Actor,

View File

@ -0,0 +1,160 @@
'''
Regression tests for the cold package import surface.
'''
import json
import os
from statistics import median
import subprocess
import sys
from typing import (
Any,
get_type_hints,
)
from tractor.discovery import (
_addr,
_multiaddr,
)
from tractor.ipc import (
_tcp,
_uds,
)
def run_cold_import(code: str) -> dict[str, object]:
result = subprocess.run(
[
sys.executable,
'-c',
code,
],
check=True,
capture_output=True,
text=True,
)
return json.loads(result.stdout)
def test_lazy_to_asyncio_package_api():
'''
Keep the public lazy submodule discoverable without eagerly
importing it.
Before the lazy conversion, package import side effects exposed
`to_asyncio` to `dir()` and wildcard imports. Exercise those APIs
in cold interpreters so this test proves normal `import tractor`
leaves `asyncio` unloaded, while discovery and wildcard access
still advertise and resolve the public submodule.
'''
cold = run_cold_import(
'import json, sys, tractor; '
'print(json.dumps({'
'"advertised": "to_asyncio" in dir(tractor), '
'"asyncio_loaded": "asyncio" in sys.modules}))'
)
assert cold == {
'advertised': True,
'asyncio_loaded': False,
}
wildcard = run_cold_import(
'import json; '
'from tractor import *; '
'print(json.dumps({'
'"module": to_asyncio.__name__}))'
)
assert wildcard == {
'module': 'tractor.to_asyncio',
}
def test_cold_import_budget():
'''
Keep cold package import below the pre-optimization regression.
The original `inspect.stack()` caller lookup made a fresh
`import tractor` take about 0.42s and dominate actor startup.
Run seven independent interpreters and gate their median at a
deliberately broad 0.35s: over twice the measured ~0.145s
baseline, but low enough to catch restoration of that hot path.
Taking the median absorbs process-start and shared-runner noise.
The child measures only its import, rather than parent-side
process creation. `TRACTOR_IMPORT_BUDGET_S` provides an explicit,
reviewable override for platforms that establish a different
baseline instead of silently weakening the project default.
Each child also reports the modules whose eager loading this PR
intentionally removes, proving a timing pass cannot hide a
dependency-import regression.
'''
budget_s = float(
os.environ.get(
'TRACTOR_IMPORT_BUDGET_S',
'0.35',
)
)
optional_mods = (
'asyncio',
'bidict',
'colorlog',
'multiaddr',
'wrapt',
)
code = (
'import json, sys, time; '
'started = time.perf_counter(); '
'import tractor; '
'elapsed = time.perf_counter() - started; '
f'optional = {optional_mods!r}; '
'print(json.dumps({'
'"elapsed": elapsed, '
'"loaded": [name for name in optional '
'if name in sys.modules]}))'
)
samples = [
run_cold_import(code)
for _ in range(7)
]
elapsed = [
float(sample['elapsed'])
for sample in samples
]
loaded = {
name
for sample in samples
for name in sample['loaded']
}
assert not loaded
assert median(elapsed) < budget_s, (
f'cold import median exceeded {budget_s:.3f}s budget: '
f'{elapsed!r}'
)
def test_lazy_annotation_names_resolve():
'''
Resolve annotations without importing optional dependencies.
Moving annotation-only third-party names under `TYPE_CHECKING`
left their runtime globals undefined, causing
`typing.get_type_hints()` to raise `NameError`. Resolve every
affected API and prove the lazy aliases retain import-free runtime
introspection.
'''
assert get_type_hints(_multiaddr.mk_maddr)['return'] is Any
assert get_type_hints(_tcp.MsgpackTCPStream.maddr.fget)[
'return'
] is Any
assert get_type_hints(_uds.MsgpackUDSStream.maddr.fget)[
'return'
] == Any|str
assert get_type_hints(_addr.Address.get_random)[
'current_actor'
] is Any
assert _addr.__annotations__['_address_types'].startswith('dict')

View File

@ -2,16 +2,21 @@
`tractor.log`-wrapping unit tests. `tractor.log`-wrapping unit tests.
''' '''
import importlib
import logging
from pathlib import Path from pathlib import Path
import shutil import shutil
import sys
from types import ModuleType from types import ModuleType
import pytest import pytest
import tractor import tractor
import trio
from tractor import ( from tractor import (
_code_load, _code_load,
log, log,
) )
from tractor.ipc import _chan
def test_root_pkg_not_duplicated_in_logger_name(): def test_root_pkg_not_duplicated_in_logger_name():
@ -162,6 +167,53 @@ def test_implicit_mod_name_applied_for_child(
assert submod.log.logger in sub_logs assert submod.log.logger in sub_logs
def test_implicit_mod_name_from_unregistered_namespace(
tmp_path: Path,
):
'''
Preserve implicit logger naming for dynamic module namespaces.
The fast `sys.modules` caller lookup cannot resolve `runpy`,
plugin-loader, or `exec()` namespaces that are not registered.
Compile a real package file under an unregistered module name so
the rare filename fallback must recover its imported package and
retain the same package-level logger name.
'''
pkg_name = 'dynamic_logger_pkg'
pkg_dir = tmp_path / pkg_name
pkg_dir.mkdir()
init_path = pkg_dir / '__init__.py'
init_path.write_text('')
mod_path = pkg_dir / 'plugin.py'
mod_path.write_text('')
sys.path.insert(0, str(tmp_path))
try:
importlib.import_module(pkg_name)
namespace = {
'__name__': f'{pkg_name}.unregistered',
'__package__': pkg_name,
'tractor': tractor,
}
exec(
compile(
'log = tractor.log.get_logger('
f'pkg_name={pkg_name!r})',
str(mod_path),
'exec',
),
namespace,
)
dynamic_log = namespace.get('log')
finally:
sys.path.remove(str(tmp_path))
sys.modules.pop(pkg_name, None)
assert dynamic_log is not None
assert dynamic_log.name == pkg_name
def test_io_custom_level_registered(): def test_io_custom_level_registered():
''' '''
The `IO`(21) level (registered via `add_log_level()` at The `IO`(21) level (registered via `add_log_level()` at
@ -222,6 +274,88 @@ def test_add_log_level_pluggable():
delattr(log.StackLevelAdapter, name.lower()) delattr(log.StackLevelAdapter, name.lower())
@pytest.mark.parametrize(
'suppression',
[
'level',
'logger',
'global',
],
)
def test_log_guard_skips_payload_formatting(
monkeypatch: pytest.MonkeyPatch,
suppression: str,
):
'''
Suppressed transport logs must not render payloads.
The original hot-path guard compared only the effective logger
level. A logger disabled through its `Logger.disabled` flag or
the global `logging.disable()` threshold could therefore still
call `pformat()` before `Logger.isEnabledFor()` discarded the
record.
Exercise effective-level, per-logger, and global suppression
independently. A poisoned `_chan.pformat()` proves rendering is
skipped, while the fake transport proves `Channel.send()` still
transmits the original payload and traceback-hiding flag.
'''
sent: list[tuple[object, bool]] = []
class FakeTransport:
async def send(
self,
payload: object,
hide_tb: bool = False,
) -> None:
sent.append((payload, hide_tb))
def fail_pformat(payload: object) -> str:
raise AssertionError(
f'suppressed log rendered payload: {payload!r}'
)
chan_log = log.get_logger(
name=f'guard_test.{suppression}',
)
std_log = chan_log.logger
orig_level: int = std_log.level
orig_disable: int = logging.root.manager.disable
transport_level: int = log.CUSTOM_LEVELS['TRANSPORT']
monkeypatch.setattr(_chan, 'log', chan_log)
monkeypatch.setattr(_chan, 'pformat', fail_pformat)
try:
logging.disable(logging.NOTSET)
std_log.setLevel(transport_level)
if suppression == 'level':
std_log.setLevel(logging.INFO)
elif suppression == 'logger':
monkeypatch.setattr(std_log, 'disabled', True)
else:
logging.disable(logging.CRITICAL)
assert not chan_log.isEnabledFor(transport_level)
transport = FakeTransport()
chan = _chan.Channel(transport=transport)
payload = object()
async def send_payload() -> None:
await chan.send(
payload,
hide_tb=True,
)
trio.run(send_payload)
assert sent == [(payload, True)]
finally:
std_log.setLevel(orig_level)
logging.disable(orig_disable)
# TODO, moar tests against existing feats: # TODO, moar tests against existing feats:
# ------ - ------ # ------ - ------
# - [ ] color settings? # - [ ] color settings?

View File

@ -9,6 +9,7 @@ from typing import Awaitable
import pytest import pytest
import trio import trio
from trio.testing import wait_all_tasks_blocked
import tractor import tractor
from tractor.trionics import ( from tractor.trionics import (
maybe_open_context, maybe_open_context,
@ -94,6 +95,232 @@ def test_resource_only_entered_once(key_on):
trio.run(main) trio.run(main)
def test_last_moc_user_waits_for_resource_exit():
'''
Verify the final user cannot return before resource teardown.
Previously the final `maybe_open_context()` user only signalled
`_Cache.run_ctx()` through its `no_more_users` event. The user
then returned while the service task was still running the
resource's `__aexit__()`, so callers could observe stale external
state immediately after their `async with` block.
The resource sets `exit_started` before blocking on
`allow_exit`. The user task must remain inside MOC until the test
releases that deterministic checkpoint and `__aexit__()` sets
`exit_finished`.
'''
async def main():
exit_started = trio.Event()
allow_exit = trio.Event()
exit_finished = trio.Event()
user_returned = trio.Event()
@acm
async def open_resource():
try:
yield
finally:
exit_started.set()
await allow_exit.wait()
exit_finished.set()
async def use_resource():
async with maybe_open_context(open_resource):
pass
assert exit_finished.is_set()
user_returned.set()
async with (
tractor.open_root_actor(),
trio.open_nursery() as tn,
):
tn.start_soon(use_resource)
await exit_started.wait()
assert not user_returned.is_set()
allow_exit.set()
await user_returned.wait()
trio.run(main)
def test_moc_delivers_resource_exit_error():
'''
Verify a resource exit error reaches the final MOC user.
Previously `_Cache.run_ctx()` executed the cached resource's
`__aexit__()` after the final user had returned. An exit failure
therefore surfaced later through the actor service nursery rather
than at the user's `async with maybe_open_context()` boundary.
This resource raises a unique `ResourceExitError` during exit.
Catching that exact instance around MOC proves the service task
delivered the failure to the final user without replacing it.
'''
class ResourceExitError(Exception):
pass
exit_error = ResourceExitError('resource exit failed')
async def main():
@acm
async def open_resource():
yield
raise exit_error
async with tractor.open_root_actor():
with pytest.raises(ResourceExitError) as exc_info:
async with maybe_open_context(open_resource):
pass
assert exc_info.value is exit_error
trio.run(main)
def test_moc_final_user_cancellation_waits_for_exit():
'''
Verify final-user cancellation still waits for successful exit.
Previously cancellation escaped the final MOC user immediately
after it signalled `_Cache.run_ctx()`, leaving resource exit to
finish later in the actor service task. This violated the context
manager boundary even when cleanup itself succeeded.
The consumer cancels its own scope while holding the sole cached
resource. The resource sets `exit_finished` from its `finally`
block, and the consumer checks that event immediately after its
cancel scope catches `trio.Cancelled`. This proves MOC's
completion wait is shielded without suppressing the original
cancellation.
'''
async def main():
exit_finished = trio.Event()
@acm
async def open_resource():
try:
yield
finally:
exit_finished.set()
async with tractor.open_root_actor():
with trio.CancelScope() as cs:
async with maybe_open_context(open_resource):
cs.cancel()
await trio.sleep_forever()
assert cs.cancelled_caught
assert exit_finished.is_set()
trio.run(main)
def test_moc_exit_error_masks_final_user_cancellation():
'''
Verify cleanup errors survive final-user cancellation.
A cancelled final user previously signalled `no_more_users` and
propagated `trio.Cancelled` before `_Cache.run_ctx()` completed
resource exit. If `__aexit__()` then failed, its error was
detached from the API call which caused teardown.
The consumer cancels its own scope at a deterministic checkpoint
inside MOC. Resource exit raises `ResourceExitError`; observing
that exact error outside the cancel scope proves MOC shields the
completion wait and applies normal context-manager masking, where
a cleanup failure replaces the active cancellation.
'''
class ResourceExitError(Exception):
pass
exit_error = ResourceExitError('resource exit failed')
async def main():
@acm
async def open_resource():
yield
raise exit_error
async with tractor.open_root_actor():
with pytest.raises(ResourceExitError) as exc_info:
with trio.CancelScope() as cs:
async with maybe_open_context(open_resource):
cs.cancel()
await trio.sleep_forever()
assert exc_info.value is exit_error
trio.run(main)
def test_moc_service_nursery_cancellation_completes_exit():
'''
Verify service-nursery cancellation cannot strand a final user.
`_Cache.run_ctx()` and an MOC consumer may share a
caller-provided service nursery. Cancelling that nursery
interrupts the service task's `no_more_users` wait and the
consumer body together. A shielded final-user wait would deadlock
if `run_ctx()` failed to publish completion while propagating its
own `trio.Cancelled`.
The outer task waits for resource entry, then cancels the exact
nursery containing both tasks. The resource shields one cleanup
checkpoint and sets `exit_finished`; observing both it and
`service_finished` proves cancellation propagated normally while
MOC's completion handshake terminated deterministically.
'''
async def main():
resource_entered = trio.Event()
exit_finished = trio.Event()
service_finished = trio.Event()
service_tn: trio.Nursery|None = None
@acm
async def open_resource():
try:
resource_entered.set()
yield
finally:
with trio.CancelScope(shield=True):
await trio.lowlevel.checkpoint()
exit_finished.set()
async def use_resource(tn: trio.Nursery):
async with maybe_open_context(
open_resource,
tn=tn,
):
await trio.sleep_forever()
async def run_service():
nonlocal service_tn
async with trio.open_nursery() as tn:
service_tn = tn
tn.start_soon(use_resource, tn)
service_finished.set()
async with trio.open_nursery() as outer_tn:
outer_tn.start_soon(run_service)
await resource_entered.wait()
assert service_tn is not None
service_tn.cancel_scope.cancel()
await service_finished.wait()
assert exit_finished.is_set()
trio.run(main)
@tractor.context @tractor.context
async def streamer( async def streamer(
ctx: tractor.Context, ctx: tractor.Context,
@ -548,43 +775,52 @@ def test_moc_reentry_during_teardown(
loglevel: str, loglevel: str,
): ):
''' '''
Reproduce the piker `open_cached_client('kraken')` race: Reproduce re-entry while an identical cached context exits.
- same `acm_func`, NO kwargs (identical `ctx_key`) - multiple tasks use the same `acm_func` with no kwargs,
- multiple tasks share the cached resource producing an identical `ctx_key`;
- all users exit -> teardown starts - all users leave and the final user starts resource teardown;
- a NEW task enters during `_Cache.run_ctx.__aexit__` - `_Cache.run_ctx()` removes the cached value and resource entry
- `values[ctx_key]` is gone (popped in inner finally) before entering the resource's blocking `__aexit__()` body;
but `resources[ctx_key]` still exists (outer finally - a new task attempts to enter that same `ctx_key` during exit;
hasn't run yet bc the acm cleanup has checkpoints) - the per-key lock keeps that entrant queued until exit
- old code: `assert not resources.get(ctx_key)` FIRES completes;
- the entrant then receives a fresh cache miss and resource.
This models the real-world scenario where `brokerd.kraken` Without teardown sharing the registration lock, re-entry could
tasks concurrently call `open_cached_client('kraken')` race resource replacement while the prior generation was still
(same `acm_func`, empty kwargs, shared `ctx_key`) and exiting. The final user could also return before that exit
the teardown/re-entry race triggers intermittently. completed.
The first resource generation signals `in_aexit` and waits on
`allow_aexit`. The re-entry task signals `reentry_started` and
blocks inside MOC; only after `wait_all_tasks_blocked()` confirms
that ordering does the coordinator release cleanup. The entrant
must then receive a fresh cache miss. `first_done` additionally
proves the first MOC user observed completed teardown before
returning.
''' '''
async def main(): async def main():
in_aexit = trio.Event() in_aexit = trio.Event()
allow_aexit = trio.Event()
reentry_started = trio.Event()
generation: int = 0
@acm @acm
async def cached_client(): async def cached_client():
''' '''
Simulates `kraken.api.get_client()`: Simulate a no-argument `kraken.api.get_client()`.
- no params (all callers share one `ctx_key`)
- slow-ish cleanup to widen the race window
between `values.pop()` and `resources.pop()`
inside `_Cache.run_ctx`.
''' '''
nonlocal generation
generation += 1
resource_generation: int = generation
yield 'the-client' yield 'the-client'
# Signal that we're in __aexit__ — at this if resource_generation == 1:
# point `values` has already been popped by
# `run_ctx`'s inner finally, but `resources`
# is still alive (outer finally hasn't run).
in_aexit.set() in_aexit.set()
await trio.sleep(10) await allow_aexit.wait()
first_done = trio.Event() first_done = trio.Event()
@ -598,16 +834,25 @@ def test_moc_reentry_during_teardown(
async def reenter_during_teardown(): async def reenter_during_teardown():
''' '''
Wait for the acm's `__aexit__` to start (meaning Wait for the acm's `__aexit__` to start (meaning
`values` is popped but `resources` still exists), the cached value is no longer available), then re-enter.
then re-enter triggering the assert.
''' '''
await in_aexit.wait() await in_aexit.wait()
# Tell the coordinator this task is about to enter MOC.
# `Event.set()` is not a checkpoint. Though `async with`
# awaits MOC's `__aenter__()`, its async generator runs
# synchronously until the held per-key `lock.acquire()`
# actually suspends this task.
reentry_started.set()
async with maybe_open_context( async with maybe_open_context(
cached_client, cached_client,
) as (cache_hit, value): ) as (cache_hit, value):
assert not cache_hit
assert value == 'the-client' assert value == 'the-client'
await first_done.wait()
with trio.fail_after(5): with trio.fail_after(5):
async with ( async with (
tractor.open_root_actor( tractor.open_root_actor(
@ -619,5 +864,15 @@ def test_moc_reentry_during_teardown(
): ):
tn.start_soon(use_and_exit) tn.start_soon(use_and_exit)
tn.start_soon(reenter_during_teardown) tn.start_soon(reenter_during_teardown)
await reentry_started.wait()
# Wait until the re-entry task is queued on MOC's
# per-key lock while `_Cache.run_ctx()` remains
# blocked in the first generation's `__aexit__()`.
# Only then release cleanup, making the intended
# enter-during-sibling-exit ordering deterministic.
await wait_all_tasks_blocked()
assert not first_done.is_set()
allow_aexit.set()
trio.run(main) trio.run(main)

View File

@ -1,10 +1,22 @@
import time import time
import platform
import trio import trio
import pytest import pytest
import tractor import tractor
# `tractor.ipc._ringbuf` is built on linux `eventfd(2)`; importing
# it pulls in `tractor.ipc._linux` whose module-level
# `ffi.dlopen(None)` raises on non-linux. Skip the whole module at
# COLLECTION before that crashing import runs (a `pytestmark` skip
# is too late — markers apply only after the import succeeds).
if platform.system() != 'Linux':
pytest.skip(
'ringbuf (eventfd) IPC is linux-only',
allow_module_level=True,
)
# XXX `cffi` dun build on py3.14 yet.. # XXX `cffi` dun build on py3.14 yet..
pytest.importorskip("cffi") pytest.importorskip("cffi")

View File

@ -9,6 +9,17 @@ from functools import partial
import pytest import pytest
import trio import trio
import tractor import tractor
# `infect_asyncio` mode is unsupported on Windows (see
# `test_infected_asyncio`); skip at COLLECTION before the
# asyncio-interop imports below so the CI leg completes.
import platform
if platform.system() == 'Windows':
pytest.skip(
'infect_asyncio mode is unsupported on Windows',
allow_module_level=True,
)
from tractor import ( from tractor import (
to_asyncio, to_asyncio,
) )

View File

@ -4,11 +4,18 @@ related API and error checks.
''' '''
import itertools import itertools
from unittest.mock import (
AsyncMock,
Mock,
)
import pytest import pytest
import tractor import tractor
import trio import trio
from tractor._exceptions import TransportClosed
from tractor.runtime import _rpc
async def sleep_back_actor( async def sleep_back_actor(
actor_name, actor_name,
@ -46,6 +53,126 @@ async def short_sleep():
await trio.sleep(0) await trio.sleep(0)
def test_rpc_runs_after_startack_disconnect():
'''
Complete an accepted RPC when its caller closes before `StartAck`.
Registrar teardown opens short-lived `unregister_actor` RPCs. A
loaded caller can close its channel while the registrar sends the
acknowledgement; normalized `TransportClosed` previously escaped
into the shared service nursery before the already-created
coroutine was awaited. This fake fails the first response send and
proves the RPC side effect still runs with no later send attempt.
'''
async def main():
rpc_ran = trio.Event()
chan = Mock()
chan.send = AsyncMock(
side_effect=TransportClosed(
message='caller closed before StartAck',
),
)
chan.connected.return_value = False
ctx = Mock(
chan=chan,
cid='rpc-cid',
_scope=None,
_task='rpc-task',
)
actor = Mock()
actor.get_context.return_value = ctx
actor._rpc_tasks = {}
actor._ongoing_rpc_tasks = trio.Event()
actor._ongoing_rpc_tasks.set()
async def rpc_func():
assert (chan, ctx.cid) in actor._rpc_tasks
rpc_ran.set()
async def invoke(task_status):
await _rpc._invoke(
actor=actor,
cid=ctx.cid,
chan=chan,
func=rpc_func,
kwargs={},
task_status=task_status,
)
async with trio.open_nursery() as nursery:
started_ctx = await nursery.start(invoke)
assert started_ctx is ctx
await rpc_ran.wait()
assert rpc_ran.is_set()
assert not actor._rpc_tasks
assert actor._ongoing_rpc_tasks.is_set()
chan.send.assert_awaited_once()
trio.run(main)
def test_error_shipment_ignores_closed_response_channel(monkeypatch):
'''
Preserve an application error when its response channel is closed.
A caller can disconnect after submitting an RPC but before its
error response. Normalized `TransportClosed` from that final send
is terminal response failure, not a new actor-wide service error.
This test proves error shipment logs and returns without replacing
the original application exception.
'''
chan = Mock()
chan.send = AsyncMock(
side_effect=[
None,
TransportClosed(
message='caller closed before Error response',
),
],
)
error_msg = Mock(boxed_type_str='ValueError')
monkeypatch.setattr(
_rpc,
'pack_error',
Mock(return_value=error_msg),
)
ctx = Mock(
chan=chan,
cid='rpc-cid',
_scope=None,
_task='rpc-task',
)
actor = Mock()
actor.get_context.return_value = ctx
actor._rpc_tasks = {}
actor._ongoing_rpc_tasks = trio.Event()
actor._ongoing_rpc_tasks.set()
async def failing_rpc():
raise ValueError('application failure')
async def main():
async with trio.open_nursery() as nursery:
started_ctx = await nursery.start(
_rpc._invoke,
actor,
ctx.cid,
chan,
failing_rpc,
{},
)
assert started_ctx is ctx
trio.run(main)
assert chan.send.await_count == 2
assert not actor._rpc_tasks
assert actor._ongoing_rpc_tasks.is_set()
@pytest.mark.parametrize( @pytest.mark.parametrize(
'to_call', [ 'to_call', [
([], 'short_sleep', tractor.RemoteActorError), ([], 'short_sleep', tractor.RemoteActorError),

View File

@ -73,6 +73,11 @@ async def test_lifetime_stack_wipes_tmpfile(
1.6 if error_in_child 1.6 if error_in_child
else 1 else 1
) )
# scale for slow/noisy CI (esp. macOS) so the child error
# propagates before the deadline; otherwise `move_on_after`
# cancels first and flips the `error_in_child=True` assert.
from .conftest import cpu_perf_headroom
timeout *= cpu_perf_headroom()
try: try:
with trio.move_on_after(timeout) as cs: with trio.move_on_after(timeout) as cs:
async with tractor.open_nursery( async with tractor.open_nursery(

View File

@ -120,6 +120,13 @@ async def child_read_shm_list(
print(f'(child): reading frame: {frame}') print(f'(child): reading frame: {frame}')
@pytest.mark.skipif(
platform.system() == 'Windows',
reason=(
'parent/child shm IPC deadlocks on Windows '
'(frame-size dependent hang); nascent — see #404'
),
)
@pytest.mark.parametrize( @pytest.mark.parametrize(
'use_str', 'use_str',
[False, True], [False, True],

View File

@ -0,0 +1,333 @@
'''
`tractor.trionics._taskc.start_or_cancel()` unit tests.
`trio.Nursery.start()` collapses an out-of-band (ancestor)
cancellation into a lossy,
`RuntimeError('child exited without calling
task_status.started()')`
whenever the started child exits pre-`.started()` WITHOUT
propagating the ambient `trio.Cancelled`; a common outcome
when the child (or any lib code it calls) runs a graceful
teardown which absorbs the cancel and returns early. Our
`start_or_cancel()` wrapper re-surfaces the real in-flight
cancellation in that case so the true root error/cancel
propagates to the `.start()` caller instead.
These tests verify both that repair AND document upstream
`trio`'s current lossy behaviour via the
`use_start_or_cancel=False` parametrizations; if a `trio`
upgrade breaks one of THOSE cases it likely means upstream
shipped better startup-cancellation porcelain and our
wrapper deserves a re-audit!
The core use case was dug out of `modden`'s
`progman.open_wks()` program-spawn machinery as per gh
issue #474; the wrapper landed originally via gh PR #464.
'''
import pytest
import trio
from trio import TaskStatus
from tractor.trionics import start_or_cancel
async def absorbs_cancel_pre_started(
task_status: TaskStatus[None] = trio.TASK_STATUS_IGNORED,
):
'''
Swallow the ambient (ancestor-scope) cancel and return
early, a naughty-but-realistic graceful-teardown pattern
and the exact shape which causes `trio.Nursery.start()`
to raise its lossy startup `RuntimeError` in place of
the real `trio.Cancelled`.
'''
try:
await trio.sleep_forever()
except trio.Cancelled:
return
async def raise_val_err():
'''
Sibling task which blows up (fast) thus OOB-cancelling
the shared parent-nursery's cancel-scope.
'''
await trio.lowlevel.checkpoint()
raise ValueError('sibling blew up!')
@pytest.mark.parametrize(
'use_start_or_cancel',
[
True,
False,
],
)
def test_sibling_err_not_masked_by_startup_rte(
use_start_or_cancel: bool,
):
'''
The `modden.runtime.progman` use case: a sibling task
errors while the `.start()`-ed child is still
pre-`.started()`, OOB-cancelling the shared nursery
scope; the child absorbs its cancel (graceful teardown)
and exits early.
- with `start_or_cancel()` the in-flight cancellation
is re-surfaced as the real `trio.Cancelled` (then
absorbed by the cancelled nursery scope) so ONLY the
root-cause sibling error escapes the nursery.
- with a bare `.start()`, upstream `trio` (currently)
also delivers its lossy startup `RuntimeError`
alongside, obscuring that the child was in fact
cancelled due to the sibling's error.
`cancelled_at_start` records that the wrapper's own await
raises `Cancelled`, rather than merely relying on the
nursery's eventual exception-group shape.
'''
cancelled_at_start: list[bool] = []
async def main():
async with trio.open_nursery() as tn:
tn.start_soon(raise_val_err)
if use_start_or_cancel:
try:
await start_or_cancel(
tn,
absorbs_cancel_pre_started,
)
except trio.Cancelled:
cancelled_at_start.append(True)
raise
else:
await tn.start(absorbs_cancel_pre_started)
with pytest.raises(ExceptionGroup) as excinfo:
trio.run(main)
eg: ExceptionGroup = excinfo.value
val_eg, rest_eg = eg.split(ValueError)
assert len(val_eg.exceptions) == 1
if use_start_or_cancel:
assert cancelled_at_start == [True]
# the re-surfaced `Cancelled` is absorbed by the
# (sibling-error cancelled) nursery scope leaving
# NO startup-noise, just the root cause.
assert rest_eg is None
else:
# the `trio` wart: a lossy startup RTE rides along
# with (and distracts from) the root cause.
rte = rest_eg.exceptions[0]
assert isinstance(rte, RuntimeError)
assert 'child exited without calling' in rte.args[0]
@pytest.mark.parametrize(
'use_start_or_cancel',
[
True,
False,
],
)
def test_pure_oob_cancel_not_morphed_to_rte(
use_start_or_cancel: bool,
):
'''
A plain (error-free) ancestor `CancelScope.cancel()`
fired while the (cancel-absorbing) child is still
pre-`.started()`:
- `start_or_cancel()` re-surfaces the `Cancelled` so
the cancelled scope exits CLEAN, no error at all.
- a bare `.start()` (currently) morphs the plain
cancel into an (eg-wrapped) startup `RuntimeError`.
`cancelled_at_start` proves cancellation interrupts the
wrapper call itself before the cancelled scope exits.
'''
cancelled_at_start: list[bool] = []
async def main():
with trio.CancelScope() as cs:
async with trio.open_nursery() as tn:
async def canceller():
await trio.lowlevel.checkpoint()
cs.cancel()
tn.start_soon(canceller)
if use_start_or_cancel:
try:
await start_or_cancel(
tn,
absorbs_cancel_pre_started,
)
except trio.Cancelled:
cancelled_at_start.append(True)
raise
else:
await tn.start(
absorbs_cancel_pre_started,
)
assert cs.cancelled_caught
if use_start_or_cancel:
trio.run(main)
assert cancelled_at_start == [True]
else:
with pytest.raises(ExceptionGroup) as excinfo:
trio.run(main)
rte = excinfo.value.exceptions[0]
assert isinstance(rte, RuntimeError)
assert 'child exited without calling' in rte.args[0]
@pytest.mark.parametrize(
'use_start_or_cancel',
[
True,
False,
],
)
def test_genuine_startup_rte_still_raised(
use_start_or_cancel: bool,
):
'''
Absent ANY in-flight cancellation, a child exiting
cleanly without calling `task_status.started()` is a
genuine startup-protocol bug; `start_or_cancel()` must
re-raise the resulting `RuntimeError` exactly like a
bare `.start()` does.
'''
async def exits_wo_started(
task_status: TaskStatus[None] = (
trio.TASK_STATUS_IGNORED
),
):
await trio.lowlevel.checkpoint()
async def main():
async with trio.open_nursery() as tn:
with pytest.raises(RuntimeError) as excinfo:
if use_start_or_cancel:
await start_or_cancel(
tn,
exits_wo_started,
)
else:
await tn.start(exits_wo_started)
rte = excinfo.value
assert (
'child exited without calling'
in
rte.args[0]
)
trio.run(main)
@pytest.mark.parametrize(
'rte_arg',
[
# Broad substring matches would wrongly demote either
# child-owned error to `Cancelled` under cancellation.
'never got started!',
'child exited without calling user hook',
# non-`str` first-arg edge; must not `TypeError`
# inside the wrapper's msg-match guard.
1234,
],
)
def test_childs_own_rte_never_demoted_to_cancel(
rte_arg: str|int,
):
'''
A child's OWN `RuntimeError`, one which merely smells
like `trio`'s startup wording (or carries a non-`str`
first arg), raised under ambient cancellation must NOT
be demoted to a `trio.Cancelled` by the exact-msg-match
guard inside `start_or_cancel()`; the real error must
always propagate to the caller as the sole exception-group
leaf, preserving object identity.
'''
child_rte = RuntimeError(rte_arg)
async def cancels_cs_then_raises(
task_status: TaskStatus[None] = (
trio.TASK_STATUS_IGNORED
),
):
# cancel the ambient (ancestor) scope then raise
# sync-ly, no checkpoint between, so the child
# deterministically dies with ITS error while the
# caller is under effective cancellation.
cs.cancel()
raise child_rte
cs = trio.CancelScope()
async def main():
with cs:
async with trio.open_nursery() as tn:
await start_or_cancel(
tn,
cancels_cs_then_raises,
)
with pytest.raises(ExceptionGroup) as excinfo:
trio.run(main)
assert excinfo.value.exceptions == (child_rte,)
def test_started_value_and_args_passthru():
'''
Happy path: positional args, the `name=` kwarg and the
`.started(value)`-delivered value all pass through
`start_or_cancel()` identically to a bare `.start()`.
'''
async def echo_started(
*args,
task_status: TaskStatus[tuple] = (
trio.TASK_STATUS_IGNORED
),
):
task_name: str = trio.lowlevel.current_task().name
task_status.started((
args,
task_name,
))
async def main():
async with trio.open_nursery() as tn:
(
args,
task_name,
) = await start_or_cancel(
tn,
echo_started,
'chillin',
10,
name='doggy',
)
assert args == ('chillin', 10)
assert task_name == 'doggy'
trio.run(main)

View File

@ -75,3 +75,39 @@ from .discovery._registry import (
Arbiter as Arbiter, Arbiter as Arbiter,
) )
# from . import hilevel as hilevel # from . import hilevel as hilevel
__all__: tuple[str, ...] = tuple(
name
for name in globals()
if not name.startswith('_')
) + (
'to_asyncio',
)
def __dir__() -> list[str]:
return sorted(set(globals()) | set(__all__))
def __getattr__(name: str):
'''
PEP 562 lazy sub-module loading, presently only for
`.to_asyncio` which (transitively) imports `asyncio`
itself: a non-trivial multi-ms chunk of the eager
`import tractor` cost (gh #470) unneeded by
`trio`-only apps.
Any `tractor.to_asyncio.<attr>` access (or a
`from tractor import to_asyncio`) still works, the
sub-mod is simply imported on first-access instead
of at pkg-import time.
'''
if name == 'to_asyncio':
from importlib import import_module
return import_module('.to_asyncio', __name__)
raise AttributeError(
f'module {__name__!r} has no attribute {name!r}'
)

View File

@ -47,9 +47,7 @@ from .devx import (
from .spawn import _spawn from .spawn import _spawn
from .runtime import _state from .runtime import _state
from . import log from . import log
from .ipc import ( from .discovery._api import _probe_registry_addrs
_connect_chan,
)
from .discovery._addr import ( from .discovery._addr import (
Address, Address,
UnwrappedAddress, UnwrappedAddress,
@ -451,49 +449,22 @@ async def open_root_actor(
from .devx._stackscope import enable_stack_on_sig from .devx._stackscope import enable_stack_on_sig
enable_stack_on_sig() enable_stack_on_sig()
# closed into below ping task-func ponged_addrs: list[Address]
ponged_addrs: list[Address] = [] occupied_addrs: list[Address]
(
ponged_addrs,
occupied_addrs,
) = await _probe_registry_addrs(uw_reg_addrs)
async def ping_tpt_socket( if (
addr: Address, not ponged_addrs
timeout: float = 1, and
) -> None: occupied_addrs
''' ):
Attempt temporary connection to see if a registry is raise RuntimeError(
listening at the requested address by a tranport layer f'Registry address(es) are occupied but did not '
ping. f'answer as Tractor registrars!\n'
f'occupied_addrs: {occupied_addrs!r}\n'
If a connection can't be made quickly we assume none no
server is listening at that addr.
'''
try:
# TODO: this connect-and-bail forces us to have to
# carefully rewrap TCP 104-connection-reset errors as
# EOF so as to avoid propagating cancel-causing errors
# to the channel-msg loop machinery. Likely it would
# be better to eventually have a "discovery" protocol
# with basic handshake instead?
with trio.move_on_after(timeout):
async with _connect_chan(addr.unwrap()):
ponged_addrs.append(addr)
except OSError:
# ?TODO, make this a "discovery" log level?
logger.info(
f'No root-actor registry found @ {addr!r}\n'
)
# !TODO, this is basically just another (abstract)
# happy-eyeballs, so we should try for formalize it somewhere
# in a `.[_]discovery` ya?
#
async with trio.open_nursery() as tn:
for uw_addr in uw_reg_addrs:
addr: Address = wrap_address(uw_addr)
tn.start_soon(
ping_tpt_socket,
addr,
) )
if tpt_bind_addrs is None: if tpt_bind_addrs is None:

View File

@ -111,14 +111,13 @@ SHM_DIR: str = '/dev/shm'
# UDS-socket leak sweep — see `find_orphaned_uds()` / # UDS-socket leak sweep — see `find_orphaned_uds()` /
# `reap_uds()` below. Tractor's UDS transport # `reap_uds()` below. Tractor's UDS transport
# (`tractor.ipc._uds`) creates sock files under # (`tractor.ipc._uds`) creates sock files in its platform-specific
# `${XDG_RUNTIME_DIR}/tractor/<name>@<pid>.sock`; a # default bindspace; a
# crash / SIGKILL / mid-cancel teardown can leave the # crash / SIGKILL / mid-cancel teardown can leave the
# file behind because `os.unlink()` lives in the # file behind because `os.unlink()` lives in the
# `_serve_ipc_eps` `finally:` block which doesn't always # `_serve_ipc_eps` `finally:` block which doesn't always
# get to run on hard exits. The reaper here is best-effort # get to run on hard exits. The reaper here is best-effort
# cleanup for the test harness + the `tractor-reap` CLI. # cleanup for the test harness + the `tractor-reap` CLI.
_UDS_SUBDIR: str = 'tractor'
# `<actor-name>@<pid>.sock` — pid is the binder's pid at # `<actor-name>@<pid>.sock` — pid is the binder's pid at
# creation time. Special sentinel: `registry@1616.sock` # creation time. Special sentinel: `registry@1616.sock`
# uses the magic `1616` not a real pid (the root # uses the magic `1616` not a real pid (the root
@ -738,19 +737,16 @@ def reap_shm(
def get_uds_dir() -> str|None: def get_uds_dir() -> str|None:
''' '''
Path of tractor's per-user UDS sock-file dir Path of Tractor's platform-specific default UDS bindspace.
(`${XDG_RUNTIME_DIR}/tractor/`).
Returns `None` when `XDG_RUNTIME_DIR` is unset (e.g. Returns `None` only when the bindspace cannot be resolved.
non-systemd hosts, or inside a container without the
var plumbed through). Caller should treat that as
"no UDS leaks possible to detect — skip".
''' '''
xdg: str|None = os.environ.get('XDG_RUNTIME_DIR') try:
if not xdg: from tractor.ipc._uds import UDSAddress
return str(UDSAddress.def_bindspace)
except Exception:
return None return None
return os.path.join(xdg, _UDS_SUBDIR)
def _parse_uds_name(filename: str) -> tuple[str, int]|None: def _parse_uds_name(filename: str) -> tuple[str, int]|None:
@ -768,16 +764,16 @@ def _parse_uds_name(filename: str) -> tuple[str, int]|None:
def find_orphaned_uds( def find_orphaned_uds(
*, *,
uds_dir: str|None = None, uds_dir: str|None = None,
include_registry_sentinel: bool = False,
) -> list[str]: ) -> list[str]:
''' '''
`<uds_dir>/*.sock` paths whose binder pid is no `<uds_dir>/*.sock` paths whose binder pid is no
longer alive (orphaned). Includes the longer alive (orphaned). Explicit callers may include the
`registry@1616.sock` sentinel `1616` is a magic `registry@1616.sock` sentinel; automatic pytest cleanup excludes
sentinel pid (not a real one) so the file's it because binder liveness cannot be inferred from magic `1616`.
presence alone signals a leak from a dead session.
Returns `[]` on platforms without `XDG_RUNTIME_DIR` Returns `[]` when the platform bindspace cannot be resolved or the
or when the dir doesn't exist. Files whose name dir doesn't exist. Files whose name
doesn't match the `<name>@<pid>.sock` pattern are doesn't match the `<name>@<pid>.sock` pattern are
skipped (we don't unlink things we don't recognize). skipped (we don't unlink things we don't recognize).
@ -811,9 +807,7 @@ def find_orphaned_uds(
continue continue
_name, pid = parsed _name, pid = parsed
if pid == _UDS_REGISTRY_SENTINEL_PID: if pid == _UDS_REGISTRY_SENTINEL_PID:
# sentinel — never a real pid; if the file if include_registry_sentinel:
# exists nobody live is "owning" it via
# /proc lookup, so always orphaned
leaked.append(path) leaked.append(path)
continue continue
if not _is_alive(pid): if not _is_alive(pid):
@ -933,8 +927,8 @@ def track_orphaned_uds_per_test():
teardown that flakifies sibling tests via teardown that flakifies sibling tests via
sock-file rebind races). sock-file rebind races).
Snapshots `${XDG_RUNTIME_DIR}/tractor/` before and Snapshots Tractor's platform-specific default UDS bindspace before
after each test; any `<name>@<pid>.sock` files and after each test; any `<name>@<pid>.sock` files
created during the test that survive teardown AND created during the test that survive teardown AND
whose creator pid is dead are surfaced as a loud whose creator pid is dead are surfaced as a loud
warning AND reaped, so the next test starts with a warning AND reaped, so the next test starts with a
@ -950,8 +944,8 @@ def track_orphaned_uds_per_test():
it (vs. blanket session-end sweep) makes blame it (vs. blanket session-end sweep) makes blame
obvious + prevents cascade flakiness. obvious + prevents cascade flakiness.
Cheap: 2x `os.listdir` + a few `os.stat`s per test. Cheap: 2x `os.listdir` + a few `os.stat`s per test. Skips silently
Skips silently when `XDG_RUNTIME_DIR` isn't set. when the platform bindspace cannot be resolved.
''' '''
uds_dir: str|None = get_uds_dir() uds_dir: str|None = get_uds_dir()

View File

@ -39,14 +39,15 @@ from typing import (
Type, Type,
) )
import pdbp # NOTE, `pdbp` + `wrapt` are lazy-imported at their
# single use-sites below to keep them off the eager
# `import tractor` path (gh #470).
from tractor.log import get_logger from tractor.log import get_logger
import trio import trio
from tractor.msg import ( from tractor.msg import (
pretty_struct, pretty_struct,
NamespacePath, NamespacePath,
) )
import wrapt
log = get_logger() log = get_logger()
@ -257,6 +258,7 @@ def api_frame(
caller_frames_up: int = 1, caller_frames_up: int = 1,
) -> Callable: ) -> Callable:
import wrapt
# handle the decorator called WITHOUT () case, # handle the decorator called WITHOUT () case,
# i.e. just @api_frame, NOT @api_frame(extra=<blah>) # i.e. just @api_frame, NOT @api_frame(extra=<blah>)
@ -320,6 +322,8 @@ def hide_runtime_frames() -> dict[FunctionType, CodeType]:
as possible, particularly from inside a `PdbREPL`. as possible, particularly from inside a `PdbREPL`.
''' '''
import pdbp
# XXX HACKZONE XXX # XXX HACKZONE XXX
# hide exit stack frames on nurseries and cancel-scopes! # hide exit stack frames on nurseries and cancel-scopes!
# |_ so avoid seeing it when the `pdbp` REPL is first engaged from # |_ so avoid seeing it when the `pdbp` REPL is first engaged from

View File

@ -31,12 +31,21 @@ from threading import (
RLock, RLock,
) )
import multiprocessing as mp import multiprocessing as mp
import platform
from signal import ( from signal import (
signal, signal,
getsignal, getsignal,
SIGUSR1,
SIGINT, SIGINT,
) )
if platform.system() != "Windows":
from signal import SIGUSR1
else:
SIGUSR1 = None
# import traceback # import traceback
from types import ModuleType from types import ModuleType
from typing import ( from typing import (
@ -347,8 +356,8 @@ def dump_tree_on_sig(
def enable_stack_on_sig( def enable_stack_on_sig(
sig: int = SIGUSR1, sig: int|None = SIGUSR1,
) -> ModuleType: ) -> ModuleType|None:
''' '''
Enable `stackscope` tracing on reception of a signal; by Enable `stackscope` tracing on reception of a signal; by
default this is SIGUSR1. default this is SIGUSR1.
@ -367,6 +376,16 @@ def enable_stack_on_sig(
>> pkill --signal SIGUSR1 -f <part-of-cmd: str> >> pkill --signal SIGUSR1 -f <part-of-cmd: str>
''' '''
# no `SIGUSR1` on this platform (e.g. Windows) -> nothing to
# wire up; degrade gracefully instead of crashing callers that
# only guard against a missing `stackscope` (`ImportError`).
if sig is None:
log.warning(
'No `SIGUSR1` on this platform;\n'
'skipping `stackscope` trace-on-signal setup!\n'
)
return None
try: try:
# NOTE, `stackscope._glue` does intentional async-gen type # NOTE, `stackscope._glue` does intentional async-gen type
# introspection at import-time which trips # introspection at import-time which trips

View File

@ -24,7 +24,6 @@ mult-process support within a single actor tree.
''' '''
from __future__ import annotations from __future__ import annotations
import asyncio
import bdb import bdb
from contextlib import ( from contextlib import (
AbstractContextManager, AbstractContextManager,
@ -53,7 +52,6 @@ from trio import (
) )
import tractor import tractor
from tractor.log import get_logger from tractor.log import get_logger
from tractor.to_asyncio import run_trio_task_in_future
from tractor._context import Context from tractor._context import Context
from tractor.runtime import _state from tractor.runtime import _state
from tractor._exceptions import ( from tractor._exceptions import (
@ -85,6 +83,10 @@ from ..pformat import (
) )
if TYPE_CHECKING: if TYPE_CHECKING:
# NOTE, `asyncio` (and `.to_asyncio`) are
# lazy-imported at their use-sites to keep them off
# the eager `import tractor` path (gh #470).
import asyncio
from trio.lowlevel import Task from trio.lowlevel import Task
from threading import Thread from threading import Thread
from tractor.runtime._runtime import ( from tractor.runtime._runtime import (
@ -164,6 +166,7 @@ async def _pause(
'An `asyncio` task should not be calling this!?' 'An `asyncio` task should not be calling this!?'
) from rte ) from rte
else: else:
import asyncio
task = asyncio.current_task() task = asyncio.current_task()
if debug_func is not None: if debug_func is not None:
@ -946,6 +949,7 @@ def pause_from_sync(
asyncio_task: asyncio.Task|None = None asyncio_task: asyncio.Task|None = None
if is_infected_aio: if is_infected_aio:
import asyncio
asyncio_task = asyncio.current_task() asyncio_task = asyncio.current_task()
# TODO: we could also check for a non-`.to_thread` context # TODO: we could also check for a non-`.to_thread` context
@ -1059,6 +1063,9 @@ def pause_from_sync(
greenback: ModuleType = maybe_import_greenback() greenback: ModuleType = maybe_import_greenback()
if greenback.has_portal(): if greenback.has_portal():
from tractor.to_asyncio import (
run_trio_task_in_future,
)
DebugStatus.shield_sigint() DebugStatus.shield_sigint()
fute: asyncio.Future = run_trio_task_in_future( fute: asyncio.Future = run_trio_task_in_future(
partial( partial(

View File

@ -20,7 +20,6 @@ Root-actor TTY mutex-locking machinery.
''' '''
from __future__ import annotations from __future__ import annotations
import asyncio
from contextlib import ( from contextlib import (
AbstractContextManager, AbstractContextManager,
asynccontextmanager as acm, asynccontextmanager as acm,
@ -52,7 +51,6 @@ from trio import (
TaskStatus, TaskStatus,
) )
import tractor import tractor
from tractor.to_asyncio import run_trio_task_in_future
from tractor.log import get_logger from tractor.log import get_logger
from tractor._context import Context from tractor._context import Context
from tractor.runtime import _state from tractor.runtime import _state
@ -66,6 +64,10 @@ from tractor.runtime._state import (
) )
if TYPE_CHECKING: if TYPE_CHECKING:
# NOTE, `asyncio` (and `.to_asyncio`) are
# lazy-imported at their use-sites to keep them off
# the eager `import tractor` path (gh #470).
import asyncio
from trio.lowlevel import Task from trio.lowlevel import Task
from threading import Thread from threading import Thread
from tractor.ipc import ( from tractor.ipc import (
@ -910,6 +912,9 @@ class DebugStatus:
async def _set_repl_release(): async def _set_repl_release():
repl_release.set() repl_release.set()
from tractor.to_asyncio import (
run_trio_task_in_future,
)
fute: asyncio.Future = run_trio_task_in_future( fute: asyncio.Future = run_trio_task_in_future(
_set_repl_release _set_repl_release
) )

View File

@ -16,13 +16,13 @@
from __future__ import annotations from __future__ import annotations
from uuid import uuid4 from uuid import uuid4
from typing import ( from typing import (
Any,
Protocol, Protocol,
ClassVar, ClassVar,
Type, Type,
TYPE_CHECKING, TYPE_CHECKING,
) )
from bidict import bidict
from trio import ( from trio import (
SocketListener, SocketListener,
) )
@ -32,14 +32,20 @@ from ..runtime._state import (
_def_tpt_proto, _def_tpt_proto,
) )
from ..ipc._tcp import TCPAddress from ..ipc._tcp import TCPAddress
from ..ipc._uds import UDSAddress from ..ipc._uds import (
UDSAddress,
HAS_UDS,
)
if TYPE_CHECKING: if TYPE_CHECKING:
# ONLY type-annots, the eager import costs ~4.5ms
# of `import tractor` wall-time (gh #470).
from ..runtime._runtime import Actor from ..runtime._runtime import Actor
else:
Actor = Any
log = get_logger() log = get_logger()
# TODO, maybe breakout the netns key to a struct? # TODO, maybe breakout the netns key to a struct?
# class NetNs(Struct)[str, int]: # class NetNs(Struct)[str, int]:
# ... # ...
@ -170,25 +176,36 @@ class Address(Protocol):
... ...
_address_types: bidict[str, Type[Address]] = { # the address types available on this host: TCP always, UDS only
'tcp': TCPAddress, # where usable (`HAS_UDS`). Both registries derive from this single
'uds': UDSAddress # list via each type's `proto_key`.
_address_protos: list[Type[Address]] = [TCPAddress]
if HAS_UDS:
_address_protos.append(UDSAddress)
_address_types: dict[str, Type[Address]] = {
cls.proto_key: cls
for cls in _address_protos
} }
# TODO! really these are discovery sys default addrs ONLY useful for # TODO! really these are discovery sys default addrs ONLY useful for
# when none is provided to a root actor on first boot. # when none is provided to a root actor on first boot.
_default_lo_addrs: dict[ _default_lo_addrs: dict[str, UnwrappedAddress] = {
str, cls.proto_key: cls.get_root().unwrap()
UnwrappedAddress for cls in _address_protos
] = {
'tcp': TCPAddress.get_root().unwrap(),
'uds': UDSAddress.get_root().unwrap(),
} }
def get_address_cls(name: str) -> Type[Address]: def get_address_cls(name: str) -> Type[Address]:
try:
return _address_types[name] return _address_types[name]
except KeyError:
raise NotImplementedError(
f'No IPC transport backend for {name!r} on this '
f'platform!\n'
f'(available: {list(_address_types)})\n'
)
def is_wrapped_addr(addr: any) -> bool: def is_wrapped_addr(addr: any) -> bool:
@ -286,7 +303,14 @@ def default_lo_addrs(
for an input transport key set. for an input transport key set.
''' '''
return [ lo_addrs: list[UnwrappedAddress] = []
_default_lo_addrs[transport] for transport in transports:
for transport in transports try:
] lo_addrs.append(_default_lo_addrs[transport])
except KeyError:
raise NotImplementedError(
f'No default loopback addr for transport '
f'{transport!r} on this platform!\n'
f'(available: {list(_default_lo_addrs)})\n'
)
return lo_addrs

View File

@ -21,14 +21,18 @@ management of (service) actors.
""" """
from __future__ import annotations from __future__ import annotations
import ipaddress import ipaddress
import os
import socket import socket
from typing import ( from typing import (
AsyncGenerator, AsyncGenerator,
AsyncContextManager, AsyncContextManager,
Literal,
TYPE_CHECKING, TYPE_CHECKING,
) )
from contextlib import asynccontextmanager as acm from contextlib import asynccontextmanager as acm
import trio
from tractor.log import get_logger from tractor.log import get_logger
from ..trionics import ( from ..trionics import (
gather_contexts, gather_contexts,
@ -40,6 +44,7 @@ from ..ipc._uds import UDSAddress
from ._addr import ( from ._addr import (
UnwrappedAddress, UnwrappedAddress,
Address, Address,
mk_uuid,
wrap_address, wrap_address,
) )
from ..runtime._portal import ( from ..runtime._portal import (
@ -52,6 +57,7 @@ from ..runtime._state import (
_runtime_vars, _runtime_vars,
_def_tpt_proto, _def_tpt_proto,
) )
from ..msg.types import Aid
if TYPE_CHECKING: if TYPE_CHECKING:
from ..runtime._runtime import Actor from ..runtime._runtime import Actor
@ -60,6 +66,115 @@ if TYPE_CHECKING:
log = get_logger() log = get_logger()
async def _probe_registry(
addr: Address,
timeout: float = 3,
attempt_timeout: float = 1,
max_attempts: int = 3,
retry_delay: float = .05,
close_timeout: float = .2,
) -> Literal[
'absent',
'occupied',
'registrar',
]:
'''
Confirm an address serves the Tractor actor handshake.
Connection and handshake work share `timeout`; each attempt gets
`attempt_timeout`. Shielded cleanup may add up to `close_timeout`
per attempted channel.
'''
from .._exceptions import TransportClosed
connected_once: bool = False
with trio.move_on_after(timeout):
for attempt in range(max_attempts):
try:
with trio.move_on_after(attempt_timeout) as attempt_cs:
async with _connect_chan(
addr.unwrap(),
close_timeout=close_timeout,
) as chan:
connected_once = True
peer_aid: Aid = await chan._do_handshake(
aid=Aid(
name='registry-probe',
uuid=mk_uuid(),
pid=os.getpid(),
is_probe=True,
),
timeout=attempt_timeout,
)
if peer_aid.is_registrar is not False:
return 'registrar'
return 'occupied'
if attempt_cs.cancelled_caught:
if not connected_once:
return 'absent'
except OSError:
return (
'occupied'
if connected_once
else 'absent'
)
except TransportClosed:
pass
if attempt + 1 < max_attempts:
await trio.sleep(retry_delay * (attempt + 1))
return 'occupied'
async def _probe_registry_addrs(
addrs: list[UnwrappedAddress],
timeout: float = 3,
) -> tuple[
list[Address],
list[Address],
]:
'''
Concurrently classify candidate registrar addresses.
Return confirmed registrar addresses followed by addresses occupied
by non-registrar or unresponsive Tractor peers.
'''
registrar_addrs: list[Address] = []
occupied_addrs: list[Address] = []
async def probe_addr(addr: Address) -> None:
probe_status = await _probe_registry(
addr=addr,
timeout=timeout,
)
if probe_status == 'registrar':
registrar_addrs.append(addr)
elif probe_status == 'occupied':
occupied_addrs.append(addr)
else:
# ?TODO, make this a "discovery" log level?
log.info(
f'No root-actor registry found @ {addr!r}\n'
)
async with trio.open_nursery() as nursery:
for unwrapped_addr in addrs:
nursery.start_soon(
probe_addr,
wrap_address(unwrapped_addr),
)
return (
registrar_addrs,
occupied_addrs,
)
def _is_local_addr(addr: Address) -> bool: def _is_local_addr(addr: Address) -> bool:
''' '''
Determine whether `addr` is reachable on the Determine whether `addr` is reachable on the

View File

@ -24,14 +24,23 @@ Multiaddress support using the upstream `py-multiaddr` lib
- https://github.com/multiformats/multiaddr/blob/master/protocols/unix.md - https://github.com/multiformats/multiaddr/blob/master/protocols/unix.md
''' '''
from __future__ import annotations
import ipaddress import ipaddress
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING from typing import (
Any,
from multiaddr import Multiaddr TYPE_CHECKING,
)
if TYPE_CHECKING: if TYPE_CHECKING:
# NOTE, `multiaddr` is lazy-imported at first use
# (in the fns below) to keep it off the eager
# `import tractor` path (gh #470).
from multiaddr import Multiaddr
from tractor.discovery._addr import Address from tractor.discovery._addr import Address
else:
Multiaddr = Any
Address = Any
# map from tractor-internal `proto_key` identifiers # map from tractor-internal `proto_key` identifiers
# to the standard multiaddr protocol name strings. # to the standard multiaddr protocol name strings.
@ -56,6 +65,8 @@ def mk_maddr(
multiaddr-spec-compliant protocol path. multiaddr-spec-compliant protocol path.
''' '''
from multiaddr import Multiaddr
proto_key: str = addr.proto_key proto_key: str = addr.proto_key
maddr_proto: str|None = _tpt_proto_to_maddr.get(proto_key) maddr_proto: str|None = _tpt_proto_to_maddr.get(proto_key)
if maddr_proto is None: if maddr_proto is None:
@ -98,6 +109,7 @@ def parse_maddr(
''' '''
# lazy imports to avoid circular deps # lazy imports to avoid circular deps
from multiaddr import Multiaddr
from tractor.ipc._tcp import TCPAddress from tractor.ipc._tcp import TCPAddress
from tractor.ipc._uds import UDSAddress from tractor.ipc._uds import UDSAddress

View File

@ -33,6 +33,7 @@ from typing import (
) )
import warnings import warnings
import msgspec
import trio import trio
from ._types import ( from ._types import (
@ -198,9 +199,6 @@ class Channel:
# assert transport.raddr == addr # assert transport.raddr == addr
chan = Channel(transport=transport) chan = Channel(transport=transport)
# ?TODO, compact this into adapter level-methods?
# -[ ] would avoid extra repr-calcs if level not active?
# |_ how would the `calc_if_level` look though? func?
if log.at_least_level('runtime'): if log.at_least_level('runtime'):
from tractor.devx import ( from tractor.devx import (
pformat as _pformat, pformat as _pformat,
@ -325,6 +323,8 @@ class Channel:
''' '''
__tracebackhide__: bool = hide_tb __tracebackhide__: bool = hide_tb
try: try:
if log.at_least_level('transport'):
# don't materialize the payload repr if not necessary
log.transport( log.transport(
'=> send IPC msg:\n\n' '=> send IPC msg:\n\n'
f'{pformat(payload)}\n' f'{pformat(payload)}\n'
@ -496,6 +496,7 @@ class Channel:
async def _do_handshake( async def _do_handshake(
self, self,
aid: Aid, aid: Aid,
timeout: float|None = None,
) -> Aid: ) -> Aid:
''' '''
@ -506,8 +507,28 @@ class Channel:
"actor model" parlance. "actor model" parlance.
''' '''
try:
with trio.fail_after(
timeout if timeout is not None else float('inf')
):
await self.send(aid) await self.send(aid)
peer_aid: Aid = await self.recv() peer_aid: Aid = await self.recv()
if not isinstance(peer_aid, Aid):
raise TypeError(
f'Expected {Aid!r}, received {peer_aid!r}'
)
except (
MsgTypeError,
msgspec.DecodeError,
TypeError,
UnicodeDecodeError,
trio.TooSlowError,
) as handshake_err:
raise TransportClosed(
message='Peer sent an invalid actor handshake!\n',
src_exc=handshake_err,
loglevel='warning',
) from handshake_err
log.runtime( log.runtime(
f'Received hanshake with peer\n' f'Received hanshake with peer\n'
f'<= {peer_aid.reprol(sin_uuid=False)}\n' f'<= {peer_aid.reprol(sin_uuid=False)}\n'
@ -519,7 +540,8 @@ class Channel:
@acm @acm
async def _connect_chan( async def _connect_chan(
addr: UnwrappedAddress addr: UnwrappedAddress,
close_timeout: float|None = None,
) -> typing.AsyncGenerator[Channel, None]: ) -> typing.AsyncGenerator[Channel, None]:
''' '''
Create and connect a `Channel` to the provided `addr`, disconnect Create and connect a `Channel` to the provided `addr`, disconnect
@ -530,6 +552,18 @@ async def _connect_chan(
''' '''
chan = await Channel.from_addr(addr) chan = await Channel.from_addr(addr)
try:
yield chan yield chan
finally:
with trio.CancelScope(shield=True): with trio.CancelScope(shield=True):
if close_timeout is None:
await chan.aclose() await chan.aclose()
else:
with trio.move_on_after(close_timeout) as close_cs:
await chan.aclose()
if close_cs.cancelled_caught:
log.warning(
f'Timed out closing channel after '
f'{close_timeout}s\n'
f'|_{chan}\n'
)

View File

@ -62,16 +62,19 @@ from .. import log
from ..discovery._addr import Address from ..discovery._addr import Address
from ._chan import Channel from ._chan import Channel
from ._transport import MsgTransport from ._transport import MsgTransport
from ._uds import UDSAddress
from ._tcp import TCPAddress
if TYPE_CHECKING: if TYPE_CHECKING:
from ..runtime._runtime import Actor from ..runtime._runtime import Actor
from ..runtime._supervise import ActorNursery from ..runtime._supervise import ActorNursery
from ._tcp import TCPAddress
from ._uds import UDSAddress
log = log.get_logger() log = log.get_logger()
_PRE_REG_HANDSHAKE_TIMEOUT: float = 10
async def maybe_wait_on_canced_subs( async def maybe_wait_on_canced_subs(
uid: tuple[str, str], uid: tuple[str, str],
@ -316,8 +319,6 @@ async def handle_stream_from_peer(
) )
''' '''
server._no_more_peers = trio.Event() # unset by making new
# TODO, debug_mode tooling for when hackin this lower layer? # TODO, debug_mode tooling for when hackin this lower layer?
# with debug.maybe_open_crash_handler( # with debug.maybe_open_crash_handler(
# pdb=True, # pdb=True,
@ -335,6 +336,7 @@ async def handle_stream_from_peer(
if actor := _state.current_actor(): if actor := _state.current_actor():
peer_aid: msgtypes.Aid = await chan._do_handshake( peer_aid: msgtypes.Aid = await chan._do_handshake(
aid=actor.aid, aid=actor.aid,
timeout=_PRE_REG_HANDSHAKE_TIMEOUT,
) )
except ( except (
TransportClosed, TransportClosed,
@ -351,12 +353,9 @@ async def handle_stream_from_peer(
# "kinda-error" that we expect to tolerate during # "kinda-error" that we expect to tolerate during
# discovery-sys related pings, queires, DoS etc. # discovery-sys related pings, queires, DoS etc.
): ):
# XXX: This may propagate up from `Channel._aiter_recv()` # `TransportClosed` is expected when a peer disconnects or
# and `MsgpackStream._inter_packets()` on a read from the # fails the initial typed handshake, including foreign clients
# stream particularly when the runtime is first starting up # and probes racing shutdown.
# inside `open_root_actor()` where there is a check for
# a bound listener on the registrar addr. the reset will be
# because the handshake was never meant took place.
log.runtime( log.runtime(
con_status con_status
+ +
@ -364,6 +363,13 @@ async def handle_stream_from_peer(
) )
return return
# Registry election probes need only the server's `Aid` capability
# response; never register them as ordinary RPC peers.
if peer_aid.is_probe:
return
server._no_more_peers = trio.Event() # unset by making new
uid: tuple[str, str] = ( uid: tuple[str, str] = (
peer_aid.name, peer_aid.name,
peer_aid.uuid, peer_aid.uuid,

View File

@ -20,7 +20,9 @@ TCP implementation of tractor.ipc._transport.MsgTransport protocol
from __future__ import annotations from __future__ import annotations
import ipaddress import ipaddress
from typing import ( from typing import (
Any,
ClassVar, ClassVar,
TYPE_CHECKING,
) )
# from contextlib import ( # from contextlib import (
# asynccontextmanager as acm, # asynccontextmanager as acm,
@ -33,7 +35,6 @@ from trio import (
open_tcp_listeners, open_tcp_listeners,
) )
from multiaddr import Multiaddr
from tractor.msg import MsgCodec from tractor.msg import MsgCodec
from tractor.log import get_logger from tractor.log import get_logger
from tractor.discovery._multiaddr import mk_maddr from tractor.discovery._multiaddr import mk_maddr
@ -42,6 +43,13 @@ from tractor.ipc._transport import (
MsgpackTransport, MsgpackTransport,
) )
if TYPE_CHECKING:
# ONLY type-annots, the eager import costs
# `import tractor` wall-time (gh #470).
from multiaddr import Multiaddr
else:
Multiaddr = Any
log = get_logger() log = get_logger()

View File

@ -33,6 +33,7 @@ from collections.abc import (
AsyncGenerator, AsyncGenerator,
AsyncIterator, AsyncIterator,
) )
import errno
import struct import struct
import trio import trio
@ -61,6 +62,68 @@ if TYPE_CHECKING:
log = get_logger() log = get_logger()
def _peer_closed_errno(exc: BaseException) -> int|None:
'''
Classify a complete transport exception tree as peer closure.
Follow explicit cause/context links. For a `BaseExceptionGroup`,
require every child branch to resolve to a peer-close errno so an
unrelated concurrent failure is never hidden as `TransportClosed`.
'''
def find_peer_errno(
current_exc: BaseException,
ancestors: set[int],
) -> int|None:
exc_id: int = id(current_exc)
if exc_id in ancestors:
return None
ancestors = ancestors | {exc_id}
if (
isinstance(current_exc, OSError)
and
current_exc.errno in {
errno.ECONNRESET,
errno.EPIPE,
}
):
return current_exc.errno
if isinstance(current_exc, BaseExceptionGroup):
child_errnos: list[int|None] = [
find_peer_errno(
child_exc,
ancestors,
)
for child_exc in current_exc.exceptions
]
if all(
child_errno is not None
for child_errno in child_errnos
):
return child_errnos[0]
return None
chained_exc: BaseException|None = (
current_exc.__cause__
or
current_exc.__context__
)
if chained_exc is not None:
return find_peer_errno(
chained_exc,
ancestors,
)
return None
return find_peer_errno(
exc,
set(),
)
# (codec, transport) # (codec, transport)
MsgTransportKey = tuple[str, str] MsgTransportKey = tuple[str, str]
@ -309,7 +372,10 @@ class MsgpackTransport(MsgTransport):
log.transport(f'received header {size}') # type: ignore log.transport(f'received header {size}') # type: ignore
msg_bytes: bytes = await self.recv_stream.receive_exactly(size) msg_bytes: bytes = await self.recv_stream.receive_exactly(size)
log.transport(f"received {msg_bytes}") # type: ignore if log.at_least_level('transport'):
log.transport( # type: ignore
f'received {msg_bytes}'
)
try: try:
# NOTE: lookup the `trio.Task.context`'s var for # NOTE: lookup the `trio.Task.context`'s var for
# the current `MsgCodec`. # the current `MsgCodec`.
@ -440,23 +506,23 @@ class MsgpackTransport(MsgTransport):
trans_err = _re trans_err = _re
tpt_name: str = f'{type(self).__name__!r}' tpt_name: str = f'{type(self).__name__!r}'
trans_err_msg: str = trans_err.args[0] trans_err_msg: str = (
str(trans_err.args[0])
if trans_err.args
else ''
)
by_whom: str = { by_whom: str = {
'another task closed this fd': 'locally', 'another task closed this fd': 'locally',
'this socket was already closed': 'by peer', 'this socket was already closed': 'by peer',
}.get(trans_err_msg) }.get(trans_err_msg)
match trans_err: match trans_err:
# XXX, specifc to UDS transport and its, # UDS peers can disconnect before handshake.
# well, "speediness".. XD # Linux normally reports `EPIPE`; Darwin reports
# |_ likely todo with races related to how fast # `ECONNRESET` for the same expected closure.
# the socket is setup/torn-down on linux
# as it pertains to rando pings from the
# `.discovery` subsys and protos.
case trio.BrokenResourceError() if ( case trio.BrokenResourceError() if (
'[Errno 32] Broken pipe' _peer_closed_errno(trans_err)
in is not None
trans_err_msg
): ):
tpt_closed = TransportClosed.from_src_exc( tpt_closed = TransportClosed.from_src_exc(
message=( message=(

View File

@ -18,17 +18,13 @@
IPC subsys type-lookup helpers? IPC subsys type-lookup helpers?
''' '''
from typing import ( from typing import Type
Type,
# TYPE_CHECKING,
)
import trio
import socket import socket
import trio
from tractor.ipc._transport import ( from tractor.ipc._transport import (
MsgTransportKey, MsgTransportKey,
MsgTransport MsgTransport,
) )
from tractor.ipc._tcp import ( from tractor.ipc._tcp import (
TCPAddress, TCPAddress,
@ -37,37 +33,33 @@ from tractor.ipc._tcp import (
from tractor.ipc._uds import ( from tractor.ipc._uds import (
UDSAddress, UDSAddress,
MsgpackUDSStream, MsgpackUDSStream,
HAS_UDS,
) )
# if TYPE_CHECKING: # the UDS backend is importable everywhere but only *usable* when
# from tractor._addr import Address # `HAS_UDS` is `True`; otherwise the runtime registers TCP only.
Address = TCPAddress|UDSAddress Address = TCPAddress|UDSAddress
# manually updated list of all supported msg transport types # the available msg-transport backends on this host: TCP always,
_msg_transports = [ # UDS only where usable (`HAS_UDS`). The lookup maps below derive
# from this single list via each backend's `codec_key` and
# `address_type`: register a backend here and every map picks it up.
_msg_transports: list[Type[MsgTransport]] = [
MsgpackTCPStream, MsgpackTCPStream,
MsgpackUDSStream
] ]
if HAS_UDS:
_msg_transports.append(MsgpackUDSStream)
# map a `MsgTransportKey` -> `MsgTransport` type
# convert a MsgTransportKey to the corresponding transport type _key_to_transport: dict[MsgTransportKey, Type[MsgTransport]] = {
_key_to_transport: dict[ (t.codec_key, t.address_type.proto_key): t
MsgTransportKey, for t in _msg_transports
Type[MsgTransport],
] = {
('msgpack', 'tcp'): MsgpackTCPStream,
('msgpack', 'uds'): MsgpackUDSStream,
} }
# convert an Address wrapper to its corresponding transport type # map an `Address`-wrapper -> `MsgTransport` type
_addr_to_transport: dict[ _addr_to_transport: dict[Type[Address], Type[MsgTransport]] = {
Type[TCPAddress|UDSAddress], t.address_type: t
Type[MsgTransport] for t in _msg_transports
] = {
TCPAddress: MsgpackTCPStream,
UDSAddress: MsgpackUDSStream,
} }
@ -81,41 +73,51 @@ def transport_from_addr(
''' '''
try: try:
return _addr_to_transport[type(addr)] addr_type = type(addr)
return _addr_to_transport[addr_type]
except KeyError: except KeyError:
raise NotImplementedError( raise NotImplementedError(
f'No known transport for address {repr(addr)}' f'No known transport for address '
f'{addr!r}'
) )
def transport_from_stream( def transport_from_stream(
stream: trio.abc.Stream, stream: trio.abc.Stream,
codec_key: str = 'msgpack' codec_key: str = 'msgpack',
) -> Type[MsgTransport]: ) -> Type[MsgTransport]:
''' '''
Given an arbitrary `trio.abc.Stream` and a desired codec, Given an arbitrary `trio.abc.Stream` and a desired codec,
find the corresponding `MsgTransport` type. find the corresponding `MsgTransport` type.
''' '''
transport = None transport: str|None = None
if isinstance(stream, trio.SocketStream): if isinstance(stream, trio.SocketStream):
sock: socket.socket = stream.socket sock: socket.socket = stream.socket
match sock.family: match sock.family:
case socket.AF_INET | socket.AF_INET6: case socket.AF_INET | socket.AF_INET6:
transport = 'tcp' transport = 'tcp'
case socket.AF_UNIX: # `HAS_UDS` short-circuits before `socket.AF_UNIX` on
# hosts where that constant is absent.
case fam if (
HAS_UDS
and
fam == socket.AF_UNIX
):
transport = 'uds' transport = 'uds'
case _: case fam:
raise NotImplementedError( raise NotImplementedError(
f'Unsupported socket family: {sock.family}' f'Unsupported socket family: {fam}'
) )
if not transport: if not transport:
raise NotImplementedError( raise NotImplementedError(
f'Could not figure out transport type for stream type {type(stream)}' f'Could not figure out transport type for stream type '
f'{type(stream)}'
) )
key = (codec_key, transport) key = (codec_key, transport)

View File

@ -21,17 +21,29 @@ from __future__ import annotations
from contextlib import ( from contextlib import (
contextmanager as cm, contextmanager as cm,
) )
import hashlib
from pathlib import Path from pathlib import Path
import os import os
import sys import sys
from socket import ( from socket import (
AF_UNIX,
SOCK_STREAM, SOCK_STREAM,
SOL_SOCKET, SOL_SOCKET,
error as socket_error, error as socket_error,
) )
# NOTE, `AF_UNIX` is absent on Windows / any CPython built without
# unix-domain-socket support. Keep this module importable
# everywhere (so `UDSAddress` stays referenceable for type and
# `isinstance()` checks plus registry lookups); the `AF_UNIX`-using
# code paths below are runtime-only and are never reached when the
# UDS backend is unusable (gated on `trio`'s `has_unix`, see
# `HAS_UDS`).
try:
from socket import AF_UNIX
except ImportError:
AF_UNIX = None
import struct import struct
from typing import ( from typing import (
Any,
Type, Type,
TYPE_CHECKING, TYPE_CHECKING,
ClassVar, ClassVar,
@ -49,7 +61,6 @@ from trio._highlevel_open_unix_stream import (
has_unix, has_unix,
) )
from multiaddr import Multiaddr
from tractor.msg import MsgCodec from tractor.msg import MsgCodec
from tractor.log import get_logger from tractor.log import get_logger
from tractor.discovery._multiaddr import mk_maddr from tractor.discovery._multiaddr import mk_maddr
@ -63,7 +74,13 @@ from tractor.runtime._state import (
) )
if TYPE_CHECKING: if TYPE_CHECKING:
# ONLY type-annots, the eager import costs
# `import tractor` wall-time (gh #470).
from multiaddr import Multiaddr
from tractor.runtime._runtime import Actor from tractor.runtime._runtime import Actor
else:
Multiaddr = Any
Actor = Any
# Platform-specific credential passing constants # Platform-specific credential passing constants
@ -90,6 +107,22 @@ else:
log = get_logger() log = get_logger()
_SUN_PATH_LIMIT: int = (
108
if sys.platform == 'linux'
else 104
)
# single source of truth for whether the UDS backend is usable on this
# host. Windows can expose `AF_UNIX`, but this backend remains
# POSIX-only until its credential and lifecycle paths are supported.
HAS_UDS: bool = (
sys.platform != 'win32'
and
has_unix
)
def unwrap_sockpath( def unwrap_sockpath(
sockpath: Path, sockpath: Path,
@ -190,7 +223,7 @@ class UDSAddress(
err_on_no_runtime=False, err_on_no_runtime=False,
) )
if actor: if actor:
sockname: str = f'{actor.aid.name}@{pid}' sockname: str = actor.aid.name
# XXX, orig version which broke both macOS (file-name # XXX, orig version which broke both macOS (file-name
# length) and `multiaddrs` ('::' invalid separator). # length) and `multiaddrs` ('::' invalid separator).
# sockname: str = '::'.join(actor.uid) + f'@{pid}' # sockname: str = '::'.join(actor.uid) + f'@{pid}'
@ -216,15 +249,72 @@ class UDSAddress(
# `(?P<name>.+)@(?P<pid>\d+)\.sock` regex, and the # `(?P<name>.+)@(?P<pid>\d+)\.sock` regex, and the
# `spawn._reap` `{name}@{pid}.sock` reconstruction. # `spawn._reap` `{name}@{pid}.sock` reconstruction.
token: str = uuid4().hex[:8] token: str = uuid4().hex[:8]
sockname: str = f'{prefix}.{token}@{pid}' sockname = f'{prefix}.{token}'
sockpath: Path = Path(f'{sockname}.sock') sockpath: Path = cls.get_sockname(
name=sockname,
pid=pid,
bindspace=filedir,
)
return UDSAddress( return UDSAddress(
filedir=filedir, filedir=filedir,
filename=sockpath, filename=sockpath,
maybe_pid=pid, maybe_pid=pid,
) )
@classmethod
def get_sockname(
cls,
name: str,
pid: int,
bindspace: Path,
) -> Path:
'''
Build a safe, deterministic UDS socket filename.
'''
suffix: str = f'@{pid}.sock'
filename: str = f'{name}{suffix}'
unsafe: bool = (
'\0' in name
or
'/' in name
or
bool(os.altsep and os.altsep in name)
or
Path(filename).is_absolute()
)
too_long: bool = (
len(os.fsencode(bindspace / filename))
>= _SUN_PATH_LIMIT
)
if (
unsafe
or
too_long
):
digest: str = hashlib.blake2s(
os.fsencode(name),
digest_size=16,
).hexdigest()
filename = f'actor.{digest}{suffix}'
sockpath: Path = bindspace / filename
path_nbytes: int = len(os.fsencode(sockpath))
if path_nbytes >= _SUN_PATH_LIMIT:
raise ValueError(
f'UDS bindspace leaves no room for an AF_UNIX '
f'socket filename!\n'
f'bindspace: {bindspace}\n'
f'name was unsafe: {unsafe}\n'
f'name was over budget: {too_long}\n'
f'compacted filename: {filename}\n'
f'encoded path bytes: {path_nbytes}\n'
f'AF_UNIX path limit: {_SUN_PATH_LIMIT}\n'
)
return Path(filename)
@classmethod @classmethod
def get_root(cls) -> UDSAddress: def get_root(cls) -> UDSAddress:
def_uds_filename: Path = 'registry@1616.sock' def_uds_filename: Path = 'registry@1616.sock'
@ -301,7 +391,16 @@ async def start_listener(
f'>{{\n' f'>{{\n'
f'|_{bs!r}\n' f'|_{bs!r}\n'
) )
bs.mkdir() bs.mkdir(
# ensure the full ancestor tree for any nested
# (custom `filedir`) bindspace; the default
# `get_rt_dir()` space is pre-created but a custom
# one may have missing parents.
parents=True,
# avoid `FileExistsError` from racing actors, same
# guard as in `get_rt_dir()`.
exist_ok=True,
)
with _reraise_as_connerr( with _reraise_as_connerr(
src_excs=( src_excs=(
@ -554,11 +653,18 @@ class MsgpackUDSStream(MsgpackTransport):
]: ]:
sock: trio.socket.socket = stream.socket sock: trio.socket.socket = stream.socket
# NOTE XXX, it's unclear why one or the other ends up being # NOTE, the `bytes` case is a linux-only artifact: setting
# `bytes` versus the socket-file-path, i presume it's # `SO_PASSCRED` (see `open_unix_socket_w_passcred()`)
# something to do with who is the server (called `.listen()`)? # causes the kernel to *autobind* the un-named client
# maybe could be better implemented using another info-query # sock to an abstract-namespace addr which python
# on the socket like, # delivers as `bytes`; the listener-bound end is always
# the fs-path `str`. On platforms WITHOUT autobind
# (macOS et al) the un-bound end instead reports as an
# empty `str` so BOTH names arrive as `str`s and the
# real fs-path is whichever is non-empty: `peername` on
# the connect side, `sockname` on the accept side.
#
# for socket-api deats see,
# https://beej.us/guide/bgnet/html/split-wide/system-calls-or-bust.html#gethostnamewho-am-i # https://beej.us/guide/bgnet/html/split-wide/system-calls-or-bust.html#gethostnamewho-am-i
sockname: str|bytes = sock.getsockname() sockname: str|bytes = sock.getsockname()
# https://beej.us/guide/bgnet/html/split-wide/system-calls-or-bust.html#getpeernamewho-are-you # https://beej.us/guide/bgnet/html/split-wide/system-calls-or-bust.html#getpeernamewho-are-you
@ -570,8 +676,24 @@ class MsgpackUDSStream(MsgpackTransport):
case (bytes(), str()): case (bytes(), str()):
sock_path: Path = Path(sockname) sock_path: Path = Path(sockname)
case (str(), str()): # XXX, likely macOS # NOTE, no-autobind case (macOS): the un-bound end
sock_path: Path = Path(peername) # is `''`, NOT a `bytes` abstract-ns addr; taking
# `peername` unconditionally (as prior impl did)
# delivers garbage `Path('')` addrs on the accept
# side!
case (str(), str()):
bound_name: str = (
peername
or
sockname
)
if not bound_name:
raise ValueError(
f'Empty UDS (peername, sockname) pair ??\n'
f'peername: {peername!r}\n'
f'sockname: {sockname!r}\n'
)
sock_path: Path = Path(bound_name)
case _: case _:
raise TypeError( raise TypeError(

View File

@ -26,11 +26,6 @@ built on `tractor`.
''' '''
from collections.abc import Mapping from collections.abc import Mapping
from functools import partial from functools import partial
from inspect import (
FrameInfo,
getmodule,
stack,
)
import sys import sys
import logging import logging
from logging import ( from logging import (
@ -38,10 +33,16 @@ from logging import (
Logger, Logger,
StreamHandler, StreamHandler,
) )
from types import ModuleType from types import (
FrameType,
ModuleType,
)
import warnings import warnings
import colorlog # type: ignore # NOTE, `colorlog` is lazy-imported in
# `get_console_log()` to keep it off the eager
# `import tractor` path (gh #470).
#
# ?TODO, some other (modern) alt libs? # ?TODO, some other (modern) alt libs?
# import coloredlogs # import coloredlogs
# import colored_traceback.auto # ?TODO, need better config? # import colored_traceback.auto # ?TODO, need better config?
@ -111,9 +112,7 @@ def at_least_level(
if isinstance(level, str): if isinstance(level, str):
level: int = CUSTOM_LEVELS[level.upper()] level: int = CUSTOM_LEVELS[level.upper()]
if log.getEffectiveLevel() <= level: return log.isEnabledFor(level)
return True
return False
# TODO, compare with using a "filter" instead? # TODO, compare with using a "filter" instead?
@ -438,17 +437,41 @@ def get_logger(
pkg_name: str = _root_name pkg_name: str = _root_name
def get_caller_mod( def get_caller_mod(
frames_up:int = 2 frames_up: int = 2,
): ) -> ModuleType|None:
''' '''
Attempt to get the module which called `tractor.get_logger()`. Attempt to get the module which called
`tractor.get_logger()`.
Resolve the caller's frame with `sys._getframe()` and
map its `__name__` through `sys.modules`; `inspect.stack()`
(the previous impl) builds src-file info for EVERY frame
on the stack, scanning all of `sys.modules` per frame via
`inspect.getmodule()`, which made module-level
`get_logger()` calls dominate `import tractor` time
(see gh #470).
''' '''
callstack: list[FrameInfo] = stack() try:
caller_fi: FrameInfo = callstack[frames_up] caller_frame: FrameType = sys._getframe(frames_up)
caller_mod: ModuleType = getmodule(caller_fi.frame) except ValueError:
return None
mod_name: str|None = caller_frame.f_globals.get(
'__name__',
)
if mod_name is None:
return None
if caller_mod := sys.modules.get(mod_name):
return caller_mod return caller_mod
# Preserve caller discovery for `runpy`, plugin loaders,
# and `exec()` namespaces not registered in `sys.modules`.
# Import `inspect` only on this rare fallback path.
from inspect import getmodule
return getmodule(caller_frame)
# --- Auto--naming-CASE --- # --- Auto--naming-CASE ---
# ------------------------- # -------------------------
# Implicitly introspect the caller's module-name whenever `name` # Implicitly introspect the caller's module-name whenever `name`
@ -782,6 +805,10 @@ def get_console_log(
None, None,
) )
): ):
# lazy-imported to keep it off the eager
# `import tractor` path (gh #470).
import colorlog # type: ignore
fmt: str = LOG_FORMAT # always apply our format? fmt: str = LOG_FORMAT # always apply our format?
handler = StreamHandler() handler = StreamHandler()
formatter = colorlog.ColoredFormatter( formatter = colorlog.ColoredFormatter(

View File

@ -306,12 +306,12 @@ class PldRx(Struct):
): ):
try: try:
pld: PayloadT = self._pld_dec.decode(pld) pld: PayloadT = self._pld_dec.decode(pld)
if log.at_least_level('runtime'):
# don't materialize the payload repr if not necessary
log.runtime( log.runtime(
f'Decoded payload for\n' f'Decoded payload for\n'
# f'\n' f'\n'
f'{msg}\n' f'{msg}\n'
# ^TODO?, ideally just render with `,
# pld={decode}` in the `msg.pformat()`??
f'where, ' f'where, '
f'{type(msg).__name__}.pld={pld!r}\n' f'{type(msg).__name__}.pld={pld!r}\n'
) )

View File

@ -144,6 +144,8 @@ class Aid(
name: str name: str
uuid: str uuid: str
pid: int|None = None pid: int|None = None
is_registrar: bool|None = None
is_probe: bool = False
# TODO? can/should we extend this field set? # TODO? can/should we extend this field set?
# -[ ] use built-in support for UUIDs? `uuid.UUID` which has # -[ ] use built-in support for UUIDs? `uuid.UUID` which has

View File

@ -94,6 +94,33 @@ if TYPE_CHECKING:
log = get_logger('tractor') log = get_logger('tractor')
def _register_rpc_task(
actor: Actor,
chan: Channel,
func: Callable,
is_rpc: bool,
task_status: TaskStatus[
Context | BaseException
],
ctx: Context,
) -> None:
'''
Register an RPC task before publishing it to `Nursery.start()`.
'''
if is_rpc:
if not actor._rpc_tasks:
actor._ongoing_rpc_tasks = trio.Event()
actor._rpc_tasks[(chan, ctx.cid)] = (
ctx,
func,
trio.Event(),
)
task_status.started(ctx)
# ?TODO? move to a `tractor.lowlevel._rpc` with the below # ?TODO? move to a `tractor.lowlevel._rpc` with the below
# func-type-cases implemented "on top of" `@context` defs: # func-type-cases implemented "on top of" `@context` defs:
# -[ ] std async func helper decorated with `@rpc_func`? # -[ ] std async func helper decorated with `@rpc_func`?
@ -142,7 +169,14 @@ async def _invoke_non_context(
# is propagated! # is propagated!
with cancel_scope as cs: with cancel_scope as cs:
ctx._scope = cs ctx._scope = cs
task_status.started(ctx) _register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
async with aclosing(coro) as agen: async with aclosing(coro) as agen:
async for item in agen: async for item in agen:
# TODO: can we send values back in here? # TODO: can we send values back in here?
@ -178,7 +212,14 @@ async def _invoke_non_context(
) )
with cancel_scope as cs: with cancel_scope as cs:
ctx._scope = cs ctx._scope = cs
task_status.started(ctx) _register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
await coro await coro
if not cs.cancelled_caught: if not cs.cancelled_caught:
@ -202,23 +243,29 @@ async def _invoke_non_context(
) )
await chan.send(ack) await chan.send(ack)
except ( except (
TransportClosed,
trio.ClosedResourceError, trio.ClosedResourceError,
trio.BrokenResourceError, trio.BrokenResourceError,
BrokenPipeError, BrokenPipeError,
) as ipc_err: ) as ipc_err:
failed_resp = True failed_resp = True
if is_rpc: log.warning(
raise ipc_err
else:
log.exception(
f'Failed to ack runtime RPC request\n\n' f'Failed to ack runtime RPC request\n\n'
f'{func} x=> {ctx.chan}\n\n' f'{func} x=> {ctx.chan}\n\n'
f'{ack}\n' f'{ack}\n'
f' |_{ipc_err!r}\n'
) )
with cancel_scope as cs: with cancel_scope as cs:
ctx._scope: CancelScope = cs ctx._scope: CancelScope = cs
task_status.started(ctx) _register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
result = await coro result = await coro
fname: str = func.__name__ fname: str = func.__name__
@ -247,6 +294,7 @@ async def _invoke_non_context(
) )
await chan.send(ret_msg) await chan.send(ret_msg)
except ( except (
TransportClosed,
BrokenPipeError, BrokenPipeError,
trio.BrokenResourceError, trio.BrokenResourceError,
): ):
@ -673,7 +721,14 @@ async def _invoke(
): ):
ctx._scope_nursery = tn ctx._scope_nursery = tn
rpc_ctx_cs = ctx._scope = tn.cancel_scope rpc_ctx_cs = ctx._scope = tn.cancel_scope
task_status.started(ctx) _register_rpc_task(
actor,
chan,
func,
is_rpc,
task_status,
ctx,
)
# invoke user endpoint fn. # invoke user endpoint fn.
res: Any|PayloadT = await coro res: Any|PayloadT = await coro
@ -920,6 +975,7 @@ async def try_ship_error_to_remote(
# downward should be mostly wrapping such cases in a # downward should be mostly wrapping such cases in a
# tpt-closed; the `.critical()` usage is warranted. # tpt-closed; the `.critical()` usage is warranted.
except ( except (
TransportClosed,
trio.ClosedResourceError, trio.ClosedResourceError,
trio.BrokenResourceError, trio.BrokenResourceError,
BrokenPipeError, BrokenPipeError,
@ -1003,17 +1059,15 @@ async def process_messages(
task_status.started(loop_cs) task_status.started(loop_cs)
async for msg in chan: async for msg in chan:
if log.at_least_level('transport'):
log.transport( # type: ignore log.transport( # type: ignore
f'IPC msg from peer\n' f'IPC msg from peer\n'
f'<= {chan.aid.reprol()}\n\n' f'<= {chan.aid.reprol()}\n\n'
# TODO: use of the pprinting of structs is # TODO: pretty-printing structs is FRAGILE;
# FRAGILE and should prolly not be # -[ ] add a non-raising log formatter with
# # native-repr fallback before using
# avoid fmting depending on loglevel for perf? # `.msg.pretty_struct` here.
# -[ ] specifically `pretty_struct.pformat()` sub-call..?
# - how to only log-level-aware actually call this?
# -[ ] use `.msg.pretty_struct` here now instead!
# f'{pretty_struct.pformat(msg)}\n' # f'{pretty_struct.pformat(msg)}\n'
f'{msg}\n' f'{msg}\n'
) )
@ -1223,18 +1277,6 @@ async def process_messages(
) )
continue continue
else:
# mark our global state with ongoing rpc tasks
actor._ongoing_rpc_tasks = trio.Event()
# store cancel scope such that the rpc task can be
# cancelled gracefully if requested
actor._rpc_tasks[(chan, cid)] = (
ctx,
func,
trio.Event(),
)
# XXX RUNTIME-SCOPED! remote (likely internal) error # XXX RUNTIME-SCOPED! remote (likely internal) error
# (^- bc no `Error.cid` -^) # (^- bc no `Error.cid` -^)
# #
@ -1262,6 +1304,7 @@ async def process_messages(
log.exception(message) log.exception(message)
raise RuntimeError(message) raise RuntimeError(message)
if log.at_least_level('transport'):
log.transport( log.transport(
'Waiting on next IPC msg from\n' 'Waiting on next IPC msg from\n'
f'peer: {chan.aid.reprol()}\n' f'peer: {chan.aid.reprol()}\n'

View File

@ -259,6 +259,7 @@ class Actor:
name=name, name=name,
uuid=uuid, uuid=uuid,
pid=os.getpid(), pid=os.getpid(),
is_registrar=self.is_registrar,
) )
self._task: trio.Task|None = None self._task: trio.Task|None = None

View File

@ -22,7 +22,10 @@ from __future__ import annotations
from contextvars import ( from contextvars import (
ContextVar, ContextVar,
) )
import os
from pathlib import Path from pathlib import Path
import stat
import sys
from typing import ( from typing import (
Any, Any,
Callable, Callable,
@ -30,7 +33,6 @@ from typing import (
TYPE_CHECKING, TYPE_CHECKING,
) )
import platformdirs
from trio.lowlevel import current_task from trio.lowlevel import current_task
from msgspec import ( from msgspec import (
@ -43,6 +45,9 @@ if TYPE_CHECKING:
from .._context import Context from .._context import Context
_DARWIN_TMPDIR: Path = Path('/tmp')
# default IPC transport protocol settings # default IPC transport protocol settings
TransportProtocolKey = Literal[ TransportProtocolKey = Literal[
'tcp', 'tcp',
@ -318,6 +323,60 @@ def current_ipc_ctx(
def _ensure_owner_only_posix_dir(
path: Path,
*,
parents: bool = False,
) -> None:
'''
Create or validate a UID-owned POSIX runtime directory.
Pre-existing directories are accepted only when owned by the
current user. Their mode is normalized to `0o700` because runtime
directories hold IPC sockets and are private bindspaces.
'''
# TODO: https://github.com/goodboy/tractor/issues/494
# Research having the actor-tree root process choose and create
# this bindspace, then propagate it to every subactor. On
# Linux, a private mount namespace could isolate it while letting
# spawned subactors inherit access; independently launched
# discovery clients would need an explicit join or fallback path.
# POSIX metadata alone records UID/GID ownership, so other systems
# still need explicit runtime metadata and lifecycle management.
try:
dir_stat: os.stat_result = path.lstat()
except FileNotFoundError:
try:
path.mkdir(
mode=0o700,
parents=parents,
)
except FileExistsError:
pass
dir_stat = path.lstat()
if (
not stat.S_ISDIR(dir_stat.st_mode)
or
dir_stat.st_uid != os.getuid()
):
platform_name: str = (
'Darwin'
if sys.platform == 'darwin'
else 'POSIX'
)
raise PermissionError(
f'Unsafe {platform_name} runtime directory!\n'
f'path: {path}\n'
f'owner uid: {dir_stat.st_uid}\n'
f'mode: {stat.filemode(dir_stat.st_mode)}\n'
)
if stat.S_IMODE(dir_stat.st_mode) != 0o700:
path.chmod(0o700)
def get_rt_dir( def get_rt_dir(
subdir: str|Path|None = None, subdir: str|Path|None = None,
appname: str = 'tractor', appname: str = 'tractor',
@ -327,12 +386,26 @@ def get_rt_dir(
userspace apps stick their IPC and cache related system userspace apps stick their IPC and cache related system
util-files. util-files.
On linux we use a `${XDG_RUNTIME_DIR}/tractor/` subdir by Linux uses an owner-only `${XDG_RUNTIME_DIR}/tractor/`; Darwin
default, but equivalents are mapped for each platform using uses a short, owner-only `/tmp/tractor-<uid>` path; other
the lovely `platformdirs` lib. platforms use the lovely `platformdirs` lib.
''' '''
rt_dir: Path = Path( # lazy-imported to keep it off the eager
# `import tractor` path (gh #470).
import platformdirs
rt_root: Path|None = None
if sys.platform == 'darwin':
# Darwin's AF_UNIX path limit is 104 bytes. The standard
# platformdirs path can consume that before the sock name.
rt_root = (
_DARWIN_TMPDIR
/ f'{appname}-{os.getuid()}'
)
rt_dir: Path = rt_root
else:
rt_dir = Path(
platformdirs.user_runtime_dir( platformdirs.user_runtime_dir(
appname=appname, appname=appname,
), ),
@ -341,6 +414,7 @@ def get_rt_dir(
# Normalize and validate that `subdir` is a relative path # Normalize and validate that `subdir` is a relative path
# without any parent-directory ("..") components, to prevent # without any parent-directory ("..") components, to prevent
# escaping the runtime directory. # escaping the runtime directory.
subdir_path: Path|None = None
if subdir: if subdir:
subdir_path = ( subdir_path = (
subdir subdir
@ -358,13 +432,28 @@ def get_rt_dir(
f'{subdir!r}\n' f'{subdir!r}\n'
) )
rt_dir: Path = rt_dir / subdir_path if os.name != 'posix':
if subdir_path is not None:
rt_dir = rt_dir / subdir_path
if not rt_dir.is_dir(): if not rt_dir.is_dir():
rt_dir.mkdir( rt_dir.mkdir(
# Runtime dirs hold IPC sockets; owner-only access
# prevents other users from traversing the bindspace.
mode=0o700,
parents=True, parents=True,
exist_ok=True, # avoid `FileExistsError` from conc calls exist_ok=True,
) )
return rt_dir
_ensure_owner_only_posix_dir(
rt_dir,
parents=(rt_root is None),
)
if subdir_path is not None:
for part in subdir_path.parts:
rt_dir = rt_dir / part
_ensure_owner_only_posix_dir(rt_dir)
return rt_dir return rt_dir

View File

@ -438,6 +438,8 @@ class ActorNursery:
), ),
bind_addrs=bind_addrs, bind_addrs=bind_addrs,
loglevel=loglevel, loglevel=loglevel,
# use the run_in_actor nursery
nursery=self._ria_nursery,
infect_asyncio=infect_asyncio, infect_asyncio=infect_asyncio,
inherit_parent_main=inherit_parent_main, inherit_parent_main=inherit_parent_main,
proc_kwargs=proc_kwargs proc_kwargs=proc_kwargs
@ -578,51 +580,6 @@ class ActorNursery:
self._join_procs.set() self._join_procs.set()
async def _reap_ria_portals(
an: ActorNursery,
errors: dict[tuple[str, str], BaseException],
ria_children: list[tuple[Portal, Actor]]|None = None,
) -> None:
'''
Wait on and stash the final result/error from every
`.run_in_actor()`-spawned child then cancel its actor
runtime, one `_spawn.cancel_on_completion()` task per
child.
Replaces the per-child reaper task formerly spawned by
the spawn backends (keyed off
`._cancel_after_result_on_exit` membership) which
required routing such children into the (now removable)
`._ria_nursery`. Only call AFTER `._join_procs` is set
so user code inside the nursery block retains exclusive
result-await access; see the "manually await results"
note in `spawn._mp.mp_proc()`.
'''
if ria_children is None:
ria_children: list[tuple[Portal, Actor]] = [
(portal, subactor)
for subactor, _, portal in an._children.values()
if portal in an._cancel_after_result_on_exit
]
if not ria_children:
return
async with (
collapse_eg(),
trio.open_nursery() as tn,
):
portal: Portal
subactor: Actor
for portal, subactor in ria_children:
tn.start_soon(
_spawn.cancel_on_completion,
portal,
subactor,
errors,
)
@acm @acm
async def _open_and_supervise_one_cancels_all_nursery( async def _open_and_supervise_one_cancels_all_nursery(
actor: Actor, actor: Actor,
@ -684,11 +641,6 @@ async def _open_and_supervise_one_cancels_all_nursery(
) )
an._join_procs.set() an._join_procs.set()
# collect results (and errors) from all
# `.run_in_actor()` children then cancel
# each, one reaper task per child.
await _reap_ria_portals(an, errors)
except BaseException as _inner_err: except BaseException as _inner_err:
inner_err = _inner_err inner_err = _inner_err
errors[actor.aid.uid] = inner_err errors[actor.aid.uid] = inner_err
@ -752,39 +704,9 @@ async def _open_and_supervise_one_cancels_all_nursery(
# '------ - ------' # '------ - ------'
) )
# snapshot `.run_in_actor()` children
# BEFORE cancelling: each backend
# spawn-task pops its `._children`
# entry as the proc gets reaped.
ria_children: list = [
(portal, subactor)
for subactor, _, portal
in an._children.values()
if portal in
an._cancel_after_result_on_exit
]
# cancel all subactors # cancel all subactors
await an.cancel() await an.cancel()
# then collect any already-relayed
# results/errors from ria children.
# Tightly bounded: anything
# collectable is already queued in
# the local ctx (relayed BEFORE the
# cancel above); a child hard-killed
# without relaying just parks its
# reaper which then self-cleans (a
# `trio.Cancelled` result is never
# stashed), mirroring the old
# backend-side reaper-vs-`soft_kill`
# cancel race.
with trio.move_on_after(0.5):
await _reap_ria_portals(
an,
errors,
ria_children=ria_children,
)
# ria_nursery scope end # ria_nursery scope end
# TODO: this is the handler around the ``.run_in_actor()`` # TODO: this is the handler around the ``.run_in_actor()``

View File

@ -38,7 +38,11 @@ from ..devx import (
pformat, pformat,
) )
# from ..msg import pretty_struct # from ..msg import pretty_struct
from ..to_asyncio import run_as_asyncio_guest #
# NOTE, `.to_asyncio` (and thus `asyncio` itself) is
# lazy-imported at the `infect_asyncio=True` use-sites
# below to keep it off the eager `import tractor` path
# (gh #470).
from ..discovery._addr import UnwrappedAddress from ..discovery._addr import UnwrappedAddress
from ..runtime._runtime import ( from ..runtime._runtime import (
async_main, async_main,
@ -96,6 +100,7 @@ def _mp_main(
) )
try: try:
if infect_asyncio: if infect_asyncio:
from ..to_asyncio import run_as_asyncio_guest
actor._infected_aio = True actor._infected_aio = True
run_as_asyncio_guest(trio_main) run_as_asyncio_guest(trio_main)
else: else:
@ -156,6 +161,7 @@ def _trio_main(
) )
try: try:
if infect_asyncio: if infect_asyncio:
from ..to_asyncio import run_as_asyncio_guest
actor._infected_aio = True actor._infected_aio = True
run_as_asyncio_guest(trio_main) run_as_asyncio_guest(trio_main)
else: else:

View File

@ -48,6 +48,7 @@ from ._entry import _mp_main
# by `try_set_start_method()` after module load time. # by `try_set_start_method()` after module load time.
from . import _spawn from . import _spawn
from ._spawn import ( from ._spawn import (
cancel_on_completion,
proc_waiter, proc_waiter,
soft_kill, soft_kill,
) )
@ -186,17 +187,31 @@ async def mp_proc(
with trio.CancelScope(shield=True): with trio.CancelScope(shield=True):
await actor_nursery._join_procs.wait() await actor_nursery._join_procs.wait()
async with trio.open_nursery() as nursery:
if portal in actor_nursery._cancel_after_result_on_exit:
nursery.start_soon(
cancel_on_completion,
portal,
subactor,
errors
)
# This is a "soft" (cancellable) join/reap which # This is a "soft" (cancellable) join/reap which
# will remote cancel the actor on a ``trio.Cancelled`` # will remote cancel the actor on a ``trio.Cancelled``
# condition. Any `.run_in_actor()` result-reaping # condition.
# happens up in the `ActorNursery` machinery (see
# `_supervise._reap_ria_portals()`), NOT here.
await soft_kill( await soft_kill(
proc, proc,
proc_waiter, proc_waiter,
portal portal
) )
# cancel result waiter that may have been spawned in
# tandem if not done already
log.warning(
"Cancelling existing result waiter task for "
f"{subactor.aid.uid}")
nursery.cancel_scope.cancel()
finally: finally:
# hard reap sequence # hard reap sequence
if proc.is_alive(): if proc.is_alive():

View File

@ -35,14 +35,14 @@ Future-work TODO — authoritative UDS bind-addr tracking
`unlink_uds_bind_addrs()` currently has two cleanup paths: `unlink_uds_bind_addrs()` currently has two cleanup paths:
1. Explicit `bind_addrs` (when parent set them at spawn time) 1. Explicit `bind_addrs` (when parent set them at spawn time)
2. **Convention-based reconstruction** 2. **Convention-based reconstruction** in the platform default UDS
`<XDG_RUNTIME_DIR>/tractor/<name>@<pid>.sock` for the bindspace for the
common case where the subactor self-assigned a random sock common case where the subactor self-assigned a random sock
via `UDSAddress.get_random()`. via `UDSAddress.get_random()`.
Path (2) hardcodes the `<name>@<pid>.sock` convention from Path (2) delegates filename reconstruction to
`tractor.ipc._uds.UDSAddress`. If that convention ever `tractor.ipc._uds.UDSAddress.get_sockname()`. If the subactor binds to
changes or the subactor binds to a non-default a non-default
`bindspace`/`filedir` we'll silently fail to unlink. `bindspace`/`filedir` we'll silently fail to unlink.
A more authoritative approach would be: A more authoritative approach would be:
@ -71,6 +71,7 @@ fd leak. Different bug class but same broader theme of
from __future__ import annotations from __future__ import annotations
import os import os
from pathlib import Path
from typing import TYPE_CHECKING from typing import TYPE_CHECKING
import trio import trio
@ -104,7 +105,7 @@ def unlink_uds_bind_addrs(
`_serve_ipc_eps` `finally:` block (which normally calls `_serve_ipc_eps` `finally:` block (which normally calls
`os.unlink(addr.sockpath)`) never runs. Without this `os.unlink(addr.sockpath)`) never runs. Without this
parent-side cleanup, the dead subactor's parent-side cleanup, the dead subactor's
`${XDG_RUNTIME_DIR}/tractor/<name>@<pid>.sock` file platform-default UDS socket file
accumulates on the filesystem (see issue #454 + the accumulates on the filesystem (see issue #454 + the
autouse `_track_orphaned_uds_per_test` fixture). autouse `_track_orphaned_uds_per_test` fixture).
@ -118,7 +119,7 @@ def unlink_uds_bind_addrs(
picked its own random sock via picked its own random sock via
`UDSAddress.get_random()`), reconstruct the path `UDSAddress.get_random()`), reconstruct the path
from `(subactor.aid.name, proc.pid)` using the from `(subactor.aid.name, proc.pid)` using the
same `<name>@<pid>.sock` convention. We can do this same `UDSAddress.get_sockname()` helper. We can do this
because the subactor uses its OWN `os.getpid()` at because the subactor uses its OWN `os.getpid()` at
bind time, which equals `proc.pid` from the bind time, which equals `proc.pid` from the
parent's view. parent's view.
@ -154,7 +155,21 @@ def unlink_uds_bind_addrs(
and subactor is not None and subactor is not None
and proc.pid is not None and proc.pid is not None
): ):
sockname: str = f'{subactor.aid.name}@{proc.pid}.sock' try:
sockname: Path = UDSAddress.get_sockname(
name=subactor.aid.name,
pid=proc.pid,
bindspace=UDSAddress.def_bindspace,
)
except Exception:
log.exception(
f'Failed to reconstruct UDS sock-file for '
f'post-kill cleanup — skipping\n'
f' |_{proc}\n'
f' |_{subactor.aid}\n'
)
return
sockpath: str = str( sockpath: str = str(
UDSAddress.def_bindspace / sockname UDSAddress.def_bindspace / sockname
) )

View File

@ -50,6 +50,7 @@ from tractor.msg import (
pretty_struct, pretty_struct,
) )
from ._spawn import ( from ._spawn import (
cancel_on_completion,
hard_kill, hard_kill,
soft_kill, soft_kill,
) )
@ -194,17 +195,32 @@ async def trio_proc(
with trio.CancelScope(shield=True): with trio.CancelScope(shield=True):
await actor_nursery._join_procs.wait() await actor_nursery._join_procs.wait()
async with trio.open_nursery() as nursery:
if portal in actor_nursery._cancel_after_result_on_exit:
nursery.start_soon(
cancel_on_completion,
portal,
subactor,
errors
)
# This is a "soft" (cancellable) join/reap which # This is a "soft" (cancellable) join/reap which
# will remote cancel the actor on a ``trio.Cancelled`` # will remote cancel the actor on a ``trio.Cancelled``
# condition. Any `.run_in_actor()` result-reaping # condition.
# happens up in the `ActorNursery` machinery (see
# `_supervise._reap_ria_portals()`), NOT here.
await soft_kill( await soft_kill(
proc, proc,
trio.Process.wait, # XXX, uses `pidfd_open()` below. trio.Process.wait, # XXX, uses `pidfd_open()` below.
portal portal
) )
# cancel result waiter that may have been spawned in
# tandem if not done already
log.cancel(
'Cancelling portal result reaper task\n'
f'c)> {subactor.aid.reprol()!r}\n'
)
nursery.cancel_scope.cancel()
finally: finally:
# XXX NOTE XXX: The "hard" reap since no actor zombies are # XXX NOTE XXX: The "hard" reap since no actor zombies are
# allowed! Do this **after** cancellation/teardown to avoid # allowed! Do this **after** cancellation/teardown to avoid

View File

@ -198,6 +198,16 @@ async def gather_contexts(
# Further potential examples of interest: # Further potential examples of interest:
# https://gist.github.com/njsmith/cf6fc0a97f53865f2c671659c88c1798#file-cache-py-L8 # https://gist.github.com/njsmith/cf6fc0a97f53865f2c671659c88c1798#file-cache-py-L8
class _CtxExit:
'''
Completion state for a cached context's shared exit.
'''
def __init__(self) -> None:
self.done = trio.Event()
self.error: Exception|None = None
class _Cache: class _Cache:
''' '''
Globally (actor-processs scoped) cached, task access to Globally (actor-processs scoped) cached, task access to
@ -213,7 +223,11 @@ class _Cache:
values: dict[Any, Any] = {} values: dict[Any, Any] = {}
resources: dict[ resources: dict[
Hashable, Hashable,
tuple[trio.Nursery, trio.Event] tuple[
trio.Nursery,
trio.Event,
_CtxExit,
],
] = {} ] = {}
# nurseries: dict[int, trio.Nursery] = {} # nurseries: dict[int, trio.Nursery] = {}
no_more_users: trio.Event|None = None no_more_users: trio.Event|None = None
@ -223,19 +237,39 @@ class _Cache:
cls, cls,
mng, mng,
ctx_key: tuple, ctx_key: tuple,
ctx_exit: _CtxExit,
task_status: trio.TaskStatus[T] = trio.TASK_STATUS_IGNORED, task_status: trio.TaskStatus[T] = trio.TASK_STATUS_IGNORED,
) -> None: ) -> None:
entered: bool = False
try:
async with mng as value: async with mng as value:
_, no_more_users = cls.resources[ctx_key] entered = True
(
_,
no_more_users,
_,
) = cls.resources[ctx_key]
cls.values[ctx_key] = value cls.values[ctx_key] = value
task_status.started(value) task_status.started(value)
try: try:
await no_more_users.wait() await no_more_users.wait()
finally: finally:
value = cls.values.pop(ctx_key) cls.values.pop(ctx_key)
cls.resources.pop(ctx_key) cls.resources.pop(ctx_key)
except Exception as exc:
if not entered:
raise
# Deliver regular `__aexit__()` failures to the final
# consumer instead of raising into the service nursery.
ctx_exit.error = exc
finally:
if entered:
ctx_exit.done.set()
class _UnresolvedCtx: class _UnresolvedCtx:
''' '''
@ -281,9 +315,10 @@ async def maybe_open_context(
) )
# yielded output # yielded output
# sentinel = object()
yielded: Any = _UnresolvedCtx yielded: Any = _UnresolvedCtx
user_registered: bool = False user_registered: bool = False
ctx_exit: _CtxExit|None = None
exit_error: Exception|None = None
# Lock resource acquisition around task racing / ``trio``'s # Lock resource acquisition around task racing / ``trio``'s
# scheduler protocol. # scheduler protocol.
@ -300,7 +335,6 @@ async def maybe_open_context(
] = trio.StrictFIFOLock() ] = trio.StrictFIFOLock()
header: str = 'Allocated NEW lock for @acm_func,\n' header: str = 'Allocated NEW lock for @acm_func,\n'
else: else:
await trio.lowlevel.checkpoint()
header: str = 'Reusing OLD lock for @acm_func,\n' header: str = 'Reusing OLD lock for @acm_func,\n'
log.debug( log.debug(
@ -368,7 +402,6 @@ async def maybe_open_context(
resources = _Cache.resources resources = _Cache.resources
entry: tuple|None = resources.get(ctx_key) entry: tuple|None = resources.get(ctx_key)
if entry: if entry:
service_tn, ev = entry
raise RuntimeError( raise RuntimeError(
f'Caching resources ALREADY exist?!\n' f'Caching resources ALREADY exist?!\n'
f'ctx_key={ctx_key!r}\n' f'ctx_key={ctx_key!r}\n'
@ -376,12 +409,18 @@ async def maybe_open_context(
f'task: {task}\n' f'task: {task}\n'
) )
resources[ctx_key] = (service_tn, trio.Event()) ctx_exit = _CtxExit()
resources[ctx_key] = (
service_tn,
trio.Event(),
ctx_exit,
)
try: try:
yielded: Any = await service_tn.start( yielded: Any = await service_tn.start(
_Cache.run_ctx, _Cache.run_ctx,
mngr, mngr,
ctx_key, ctx_key,
ctx_exit,
) )
except BaseException: except BaseException:
# If `run_ctx` (wrapping the acm's `__aenter__`) # If `run_ctx` (wrapping the acm's `__aenter__`)
@ -427,6 +466,11 @@ async def maybe_open_context(
raise taskc raise taskc
else: else:
# XXX, cached-entry-path # XXX, cached-entry-path
(
_,
_,
ctx_exit,
) = _Cache.resources[ctx_key]
_Cache.users[ctx_key] += 1 _Cache.users[ctx_key] += 1
user_registered = True user_registered = True
log.debug( log.debug(
@ -445,24 +489,17 @@ async def maybe_open_context(
) )
finally: finally:
if lock.locked():
stats: trio.LockStatistics = lock.statistics()
owner: trio.Task|None = stats.owner
log.error(
f'Lock never released by last owner={owner!r} !?\n'
f'{stats}\n'
f'\n'
f'task={task!r}\n'
f'ctx_key={ctx_key!r}\n'
f'acm_func={acm_func}\n'
)
if user_registered: if user_registered:
# Serialize user registration and teardown under the same
# per-key lock so no entrant can acquire a resource after
# its final user has committed to exiting it.
with trio.CancelScope(shield=True):
await lock.acquire()
try:
_Cache.users[ctx_key] -= 1 _Cache.users[ctx_key] -= 1
if yielded is not _UnresolvedCtx: # If no consumers remain, keep entrants queued
# if no more consumers, teardown the client # until the cached context has completely exited.
if _Cache.users[ctx_key] <= 0: if _Cache.users[ctx_key] <= 0:
log.debug( log.debug(
f'De-allocating @acm-func entry\n' f'De-allocating @acm-func entry\n'
@ -470,19 +507,39 @@ async def maybe_open_context(
f'acm_func={acm_func!r}\n' f'acm_func={acm_func!r}\n'
) )
# XXX: if we're cancelled we the entry may have never # XXX: if we're cancelled, the entry may
# been entered since the nursery task was killed. # have never been entered since the nursery
# _, no_more_users = _Cache.resources[ctx_key] # task was killed.
entry = _Cache.resources.get(ctx_key) entry = _Cache.resources.get(ctx_key)
if entry: if entry:
_, no_more_users = entry (
_,
no_more_users,
ctx_exit,
) = entry
no_more_users.set() no_more_users.set()
maybe_lock = _Cache.locks.pop( assert ctx_exit is not None
ctx_key, await ctx_exit.done.wait()
None, exit_error = ctx_exit.error
)
if maybe_lock is None: # A queued entrant already holds a reference
# to this lock. Keep it registered until that
# task has acquired and released it.
stats = lock.statistics()
if not stats.tasks_waiting:
maybe_lock = _Cache.locks.get(ctx_key)
if maybe_lock is lock:
_Cache.locks.pop(ctx_key)
else:
log.error( log.error(
f'Resource lock for {ctx_key} ALREADY POPPED?' f'Resource lock for {ctx_key} '
f'was replaced before teardown?'
) )
finally:
lock.release()
if exit_error is not None:
# Always re-raise a regular `__aexit__()` error at the
# final consumer's context boundary.
raise exit_error

View File

@ -349,7 +349,10 @@ async def start_or_cancel(
# demote it to a `Cancelled`, losing the real error. The # demote it to a `Cancelled`, losing the real error. The
# `isinstance` guard also avoids a `TypeError` when # `isinstance` guard also avoids a `TypeError` when
# `rte.args[0]` isn't a `str`. # `rte.args[0]` isn't a `str`.
'child exited without calling' in rte.args[0] rte.args[0] == (
'child exited without calling '
'task_status.started()'
)
): ):
# re-raises the in-flight `trio.Cancelled` IFF we're # re-raises the in-flight `trio.Cancelled` IFF we're
# under effective cancellation; else a cheap no-op and # under effective cancellation; else a cheap no-op and

12
uv.lock
View File

@ -308,11 +308,11 @@ wheels = [
[[package]] [[package]]
name = "idna" name = "idna"
version = "3.10" version = "3.18"
source = { registry = "https://pypi.org/simple" } source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/f1/70/7703c29685631f5a7590aa73f1f1d3fa9a380e654b86af429e0934a32f7d/idna-3.10.tar.gz", hash = "sha256:12f65c9b470abda6dc35cf8e63cc574b1c52b11df2c86030af0ac09b01b13ea9", size = 190490, upload-time = "2024-09-15T18:07:39.745Z" } sdist = { url = "https://files.pythonhosted.org/packages/cd/63/9496c57188a2ee585e0f1db071d75089a11e98aa86eb99d9d7618fc1edce/idna-3.18.tar.gz", hash = "sha256:ffb385a7e039654cef1ab9ef32c6fafe283c0c0467bba1d9029738ce4a14a848", size = 196711, upload-time = "2026-06-02T14:34:07.794Z" }
wheels = [ wheels = [
{ url = "https://files.pythonhosted.org/packages/76/c6/c88e154df9c4e1a2a66ccf0005a88dfb2650c1dffb6f5ce603dfbd452ce3/idna-3.10-py3-none-any.whl", hash = "sha256:946d195a0d259cbba61165e88e65941f16e9b36ea6ddb97f00452bae8b1287d3", size = 70442, upload-time = "2024-09-15T18:07:37.964Z" }, { url = "https://files.pythonhosted.org/packages/1e/5e/d4e9f1a599fb8e573b7b87160658329fbf28d19eac2718f51fc3def3aa5a/idna-3.18-py3-none-any.whl", hash = "sha256:7f952cbe720b688055e3f87de14f5c3e5fdaa8bc3928985c4077ca689de849a2", size = 65455, upload-time = "2026-06-02T14:34:06.319Z" },
] ]
[[package]] [[package]]
@ -900,11 +900,11 @@ wheels = [
[[package]] [[package]]
name = "setuptools" name = "setuptools"
version = "82.0.1" version = "83.0.0"
source = { registry = "https://pypi.org/simple" } source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/4f/db/cfac1baf10650ab4d1c111714410d2fbb77ac5a616db26775db562c8fab2/setuptools-82.0.1.tar.gz", hash = "sha256:7d872682c5d01cfde07da7bccc7b65469d3dca203318515ada1de5eda35efbf9", size = 1152316, upload-time = "2026-03-09T12:47:17.221Z" } sdist = { url = "https://files.pythonhosted.org/packages/34/26/f5d29e25ffdb535afef2d35cdb55b325298f96debd670da4c325e08d70f4/setuptools-83.0.0.tar.gz", hash = "sha256:025bccbbf0fa05b6192bc64ae1e7b16e001fd6d6d4d5de03c97b1c1ade523bef", size = 1154254, upload-time = "2026-07-04T15:31:22.699Z" }
wheels = [ wheels = [
{ url = "https://files.pythonhosted.org/packages/9d/76/f789f7a86709c6b087c5a2f52f911838cad707cc613162401badc665acfe/setuptools-82.0.1-py3-none-any.whl", hash = "sha256:a59e362652f08dcd477c78bb6e7bd9d80a7995bc73ce773050228a348ce2e5bb", size = 1006223, upload-time = "2026-03-09T12:47:15.026Z" }, { url = "https://files.pythonhosted.org/packages/5d/40/e1e72872c6354b306daef1703549e8e83b4d43cfea356311bf722a043752/setuptools-83.0.0-py3-none-any.whl", hash = "sha256:29b23c360f22f414dc7336bb39178cc7bcbf6021ed2733cde173f09dba19abb3", size = 1008090, upload-time = "2026-07-04T15:31:20.885Z" },
] ]
[[package]] [[package]]