archivebox.core.shutdown_util
Shared foreground signal and child-process cleanup contracts.
Shutdown intent and process exit are separate: takeover must preserve the replacement owner’s work, while a user abort must not later resume the crawl. Callers own that policy; this module preserves the received signal and stops only the process objects they explicitly supply.
Cooperative boundaries are essential, not just quieter exception handling. A real Linux CI regression interrupted Popen.poll() with SIGINT and stranded its wait lock. SIGKILL stopped the child (it became a zombie), but Popen.wait() still timed out. Raising a different exception or waiting longer cannot repair that interrupted critical section. Polling loops must record the signal first and honor it between operations; async owners may cancel their own task while their loop is running. Once exiting, the last-resort process cleanup must stay synchronous: the original loop may be closed, and creating a replacement loop cannot recover tasks or resources owned by it.
Keep immediate signal notices and repeated-signal force exit independent of the event bus: it may already be draining or closed during shutdown. Resumable pause/retry/skip belongs to the existing crawl controller and must not set sticky command-shutdown state. See foreground_shutdown_signals below for both cooperative and legacy blocking callers; do not change all callers’ policy at once without supplying their interruption boundaries.
Module Contents
Classes
Tracks the exact OS signal that asked a foreground command to exit. |
Functions
Honor sticky shutdown intent at a caller-owned interruption boundary. |
|
Return a deterministic shutdown bound from generated worker definitions. |
|
Reap our own Popen child and stop its explicitly captured descendants. |
|
Wait for a psutil parent and then hard-kill any surviving descendants. |
|
Keep shutdown intent visible until the owning foreground command unwinds. |
|
Ask a foreground command to exit if its launcher/wrapper disappears. |
Data
API
- class archivebox.core.shutdown_util.ShutdownSignalState[source]
Tracks the exact OS signal that asked a foreground command to exit.
- archivebox.core.shutdown_util._active_shutdown_state: archivebox.core.shutdown_util.ShutdownSignalState | None[source]
None
- archivebox.core.shutdown_util.raise_if_shutdown_requested() None[source]
Honor sticky shutdown intent at a caller-owned interruption boundary.
- archivebox.core.shutdown_util.configured_stopwaitsecs(workers: list[dict[str, str]] | tuple[dict[str, str], ...], *, default: int = 5, buffer: int = 5) int[source]
Return a deterministic shutdown bound from generated worker definitions.
- archivebox.core.shutdown_util.wait_popen_and_kill_children(proc: subprocess.Popen, children: list[psutil.Process], *, timeout: float, kill_timeout: float = 2.0) None[source]
Reap our own Popen child and stop its explicitly captured descendants.
A dead child still needs its parent to reap it. Keep wait failures visible; bypassing Popen’s lock would hide the interrupted-owner bug described above. Descendants must still be cleaned up when that wait fails.
- archivebox.core.shutdown_util.wait_psutil_and_kill_children(proc: psutil.Process, children: list[psutil.Process], *, timeout: float, kill_timeout: float = 2.0) None[source]
Wait for a psutil parent and then hard-kill any surviving descendants.
- archivebox.core.shutdown_util.kill_remaining_processes(processes: list[psutil.Process], *, timeout: float = 2.0) None[source]
- archivebox.core.shutdown_util.foreground_shutdown_signals(handled_signals: tuple[signal.Signals, ...] = (signal.SIGHUP, signal.SIGINT, signal.SIGTERM), *, first_signal_message: str | None = '\n[🛑] Got {signal_name}, stopping gracefully...\n', on_signal: collections.abc.Callable[[signal.Signals], None] | None = None, interrupt_handlers: dict[signal.Signals, collections.abc.Callable[[], None]] | None = None, raise_on_first_signal: bool = True) collections.abc.Iterator[archivebox.core.shutdown_util.ShutdownSignalState][source]
Keep shutdown intent visible until the owning foreground command unwinds.
Signal state was added because inner loops can swallow KeyboardInterrupt while their caller still needs to stop its children. Cooperative callers use raise_on_first_signal=False and check raise_if_shutdown_requested() between operations, or cancel their owned async task through on_signal. Raising from the handler can otherwise interrupt library locks or partially sent I/O. Blocking callers without a cooperative boundary retain immediate exceptions.
- archivebox.core.shutdown_util.foreground_parent_watchdog(*, enabled: bool = True, check_interval: float = 2.0, shutdown_signal: signal.Signals = signal.SIGTERM) collections.abc.Iterator[None][source]
Ask a foreground command to exit if its launcher/wrapper disappears.
uv run archivebox ...and similar wrappers can be killed without delivering a signal to the real Python child. If that child keeps crawling as an orphan, it can hold SQLite write locks long after the user-facing command timed out. This watchdog is only for foreground command lifetimes; daemon/supervisord workers should not use it because their parent may intentionally hand them off.