archivebox.cli.archivebox_update
Module Contents
Functions
Queue stale filesystem rows and backfill missing search facts. |
|
Update the database without forcing a filesystem scan by default. |
|
Drain old archive/ directories (0.8.x → 0.9.x migration). |
|
Queue only snapshots whose indexed filesystem version is stale. |
|
Process snapshots matching filters (DB query only). |
|
Print statistics for filtered mode. |
|
Print statistics for full mode. |
|
API
- archivebox.cli.archivebox_update._get_snapshot_crawl(snapshot: archivebox.core.models.Snapshot) archivebox.crawls.models.Crawl | None[source]
- archivebox.cli.archivebox_update.reindex_snapshots(snapshots: django.db.models.QuerySet[archivebox.core.models.Snapshot, archivebox.core.models.Snapshot], *, search_plugins: list[str], batch_size: int, collect_ids: bool = False, reverse: bool = False, wait_for_turn=None) dict[str, Any][source]
- archivebox.cli.archivebox_update.run_scheduled_maintenance(*, batch_size: int = 500) dict[str, Any][source]
Queue stale filesystem rows and backfill missing search facts.
- archivebox.cli.archivebox_update.update(filter_patterns: collections.abc.Iterable[str] = (), filter_type: str = 'exact', status: str | None = None, url__icontains: str | None = None, url__istartswith: str | None = None, tag: str | None = None, crawl_id: str | None = None, limit: int | None = None, sort: str | None = None, search: str | None = None, before: float | None = None, after: float | None = None, resume: str | None = None, batch_size: int = 500, continuous: bool = False, index_only: bool = False, migrate_only: bool = False, rescan: bool = False, reverse: bool = False, stop_daemon_stack: bool = True) None[source]
Update the database without forcing a filesystem scan by default.
–migrate-only: eagerly finish pending filesystem migrations for known rows. –index-only: reconcile known snapshot assets/metadata and backfill search. –rescan: discover filesystem orphans first, then repair known snapshots. Explicit scans default to newest first; –reverse selects oldest first. Filters select existing database snapshots for targeted updates.
- archivebox.cli.archivebox_update.drain_old_archive_dirs(resume_from: str | None = None, batch_size: int = 500, reverse: bool = False) dict[str, int][source]
Drain old archive/ directories (0.8.x → 0.9.x migration).
Removes obsolete timestamp symlinks and processes real legacy directories. For each old dir found in archive/:
Load or create DB snapshot
Trigger fs migration on save() to move to data/archive/users/{user}/…
Remove the old timestamp path after the verified migration commits
After this drains, current snapshot data exists only under archive/users/.
- archivebox.cli.archivebox_update.process_all_db_snapshots(batch_size: int = 500, resume: str | None = None, wait_for_turn=None, reverse: bool = False, migrate: bool = False) dict[str, int][source]
Queue only snapshots whose indexed filesystem version is stale.
- archivebox.cli.archivebox_update.process_filtered_snapshots(filter_patterns: collections.abc.Iterable[str], filter_type: str, status: str | None, url__icontains: str | None, url__istartswith: str | None, tag: str | None, crawl_id: str | None, limit: int | None, sort: str | None, search: str | None, before: float | None, after: float | None, resume: str | None, batch_size: int, queue_for_archiving: bool = True, reverse: bool = False, wait_for_turn=None) dict[str, Any][source]
Process snapshots matching filters (DB query only).
- archivebox.cli.archivebox_update.print_stats(stats: dict)[source]
Print statistics for filtered mode.