archivebox.crawls.models
Module Contents
Classes
API
- class archivebox.crawls.models.CrawlSchedule[source]
Bases:
archivebox.base_models.models.ModelWithUUID,archivebox.base_models.models.ModelWithNotes- class Meta[source]
Bases:
archivebox.base_models.models.ModelWithUUID.Meta,archivebox.base_models.models.ModelWithNotes.Meta
- dispatch(queued_at=None) Crawl | None[source]
Run maintenance directly or enqueue one ordinary Crawl.
- enqueue(queued_at=None) archivebox.crawls.models.Crawl[source]
- class archivebox.crawls.models.Crawl[source]
Bases:
archivebox.base_models.models.ModelWithDeleteAfter,archivebox.base_models.models.ModelWithOutputDir,archivebox.base_models.models.ModelWithConfig,archivebox.base_models.models.ModelWithHealthStats,archivebox.workers.models.ModelWithQueue- snapshot_set: django.db.models.Manager[archivebox.core.models.Snapshot][source]
None
- class Meta[source]
Bases:
archivebox.base_models.models.ModelWithDeleteAfter.Meta,archivebox.base_models.models.ModelWithOutputDir.Meta,archivebox.base_models.models.ModelWithConfig.Meta,archivebox.base_models.models.ModelWithHealthStats.Meta,archivebox.workers.models.ModelWithQueue.Meta
- update_child_snapshot_permissions(old_permissions: str | None, new_permissions: str | None) int[source]
- static parse_tag_names(tags: collections.abc.Iterable[str] | str, *, pattern: str = ',') list[str][source]
- apply_snapshot_tag_diff(*, added_tag_names: collections.abc.Iterable[str], removed_tag_names: collections.abc.Iterable[str]) None[source]
- static from_json(record: dict, overrides: dict | None = None)[source]
Create or get a Crawl from a JSON dict.
Args: record: Dict with ‘urls’ (required), optional ‘max_depth’, ‘tags_str’, ‘label’ overrides: Dict of field overrides (e.g., created_by_id)
Returns: Crawl instance or None if invalid
- get_urls_list() list[str][source]
Get list of URLs from urls field, filtering out comments and empty lines.
- count_urls_for_limit() int[source]
Count unique URLs already queued or snapshotted for this crawl.
max_urls is a crawl-wide cap on snapshots, so direct URL entries and recursively discovered snapshots both have to consume the same budget.
- static _config_value(config: collections.abc.Mapping[str, Any] | Any, key: str, default: Any = None) Any[source]
- classmethod create_scheduler_row(**kwargs) archivebox.crawls.models.Crawl[source]
- limit_stop_reason(*, config: collections.abc.Mapping[str, Any] | Any | None = None, output_dir: pathlib.Path | None = None, num_snapshots: int | None = None) str[source]
- lifecycle_stop_reason(*, num_snapshots: int | None = None, num_sealed_snapshots: int | None = None) str[source]
- stop_reason(*, config: collections.abc.Mapping[str, Any] | Any | None = None, output_dir: pathlib.Path | None = None, num_snapshots: int | None = None, num_sealed_snapshots: int | None = None) str[source]
- add_url(entry: dict) bool[source]
Add a URL to the crawl queue if not already present.
Args: entry: dict with ‘url’, optional ‘depth’, ‘title’, ‘timestamp’, ‘tags’, ‘via_snapshot’, ‘plugin’
Returns: True if URL was added, False if skipped (duplicate or depth exceeded)
- create_snapshots_from_urls() list[archivebox.core.models.Snapshot][source]
Create Snapshot objects for each URL in self.urls that doesn’t already exist.
Returns: List of newly created Snapshot objects
- create_discovered_snapshot(parent_snapshot, *, url: str, depth: int, title: str = '', tags: str = '', created_by_id: int | None = None)[source]
Create one child snapshot if it passes crawl filters and limits.
- create_discovered_snapshots(parent_snapshot, records: collections.abc.Iterable[collections.abc.Mapping[str, Any]], *, depth: int, created_by_id: int | None = None) list[archivebox.core.models.Snapshot][source]
Create child snapshots from discovered URL records after filtering and deduping once.
- is_finished() bool[source]
Check if crawl is finished (all snapshots sealed or no snapshots exist).