Data lifecycle, retention, and portability
Status: current implemented safety contract plus remaining roadmap; reconciled 2026-09-09.
The original design plan grew into a point-in-time implementation ledger. It is preserved at
archive/data-lifecycle-retention-and-portability-plan.md.
This document records only the current product boundary and genuinely unfinished work.
Current behavior
- Scan, Hunt, AI Gate, Model Intake, finding, evidence, and request-archive records retain their product-specific ownership and evidence semantics. They are not governed by one generic age query.
- Findings support explicit triage and bounded cleanup. Active findings are not proof merely because they remain active, and age alone must not silently resolve them.
- Evidence retention deletion is interactive-only. It requires an immutable target-scoped preview, a one-use dangerous approval bound to that preview, locked revalidation, durable execution intent, and idempotent retry/finalization.
- Retention schedules are disabled. Destructive cleanup cannot be converted into background automation.
- HTTP request archives provide redacted JSON by default. Raw HAR is sensitive and requires explicit operator authorization and deployment support.
- Hunt exposes requests-only export separately from its explicit decision record/debrief. Hidden chain-of-thought is never an export product.
- Content-addressed evidence and external blobs must not be deleted before durable ownership and reference checks succeed.
Safety invariants
- Archive/close, content purge, record purge, and “forget” are different operations.
- A destructive operation starts with a dry-run manifest the user can inspect.
- Approval binds the exact immutable candidate set and expires with it.
- Active execution, legal/operational hold, and unresolved ownership block deletion.
- Database intent is durable before external blob deletion begins.
- Retry is resumable and idempotent; missing already-deleted blobs do not corrupt finalization.
- Exports label redaction, completeness, product scope, schema version, and evidence omissions.
- Import never grants execution authority, proof, credentials, or target authorization.
Remaining work
- A unified user-facing archive/restore lifecycle across all four product planes.
- Product-aware “Export All / Import All” for disaster recovery and migration, with schema and subject-digest validation.
- Legal/operational hold management and policy simulation before any automatic retention is reconsidered.
- Storage accounting that distinguishes database rows, content-addressed blobs, quarantine, generated reports, and external object stores.
Until those capabilities have public contracts and acceptance tests, do not describe them as
shipped. Backup/restore operations remain documented in upgrade-and-rollback.md.
2.3.1: target and finding record deletion
Target record deletion is now different from target archive. POST /targets/{id}/archive
hides the asset, disables automatic ASM, and pauses its recurring schedules in one transaction.
Both scheduler claim paths recheck target activity. Archive does not cancel already admitted
work or remove records. The target deletion dialog offers Archive target instead with its
own confirmation, including when erasure is blocked by protected history.
DELETE /targets/{id} is a permanent database-record operation requiring the preview
and approval below. The Targets page exposes a delete control on each actual target, including
subdomains. It never interprets a root-domain group as recursive ownership of every subdomain.
The Findings page supports selected-record deletion from the selection dock (Select, choose
rows, More, Delete selected findings) and previewed age cleanup (Advanced cleanup). Neither is a
front-line control: bulk triage (POST /findings/bulk) is the dock's primary action, and in
managed workspaces both deletion entries require the record_deletion and engine_admin
capabilities. Finding detail supports single-record deletion with scan scope from the Manage
record block at the end of the page. Investigation candidates are not finding rows and cannot be
selected through this surface.
Explicit API flow
- Call
POST /data-deletion/previewwith either{"kind":"target","target_id":"<UUID>"},{"kind":"findings","finding_ids":["<UUID>"],"scan_id":"<optional UUID>"}, or{"kind":"findings","older_than_days":90,"status":"resolved","root_domain":"example.invalid"}. The response includes exact IDs, cascade/detach/retain counts, blockers, expiry, a scope receipt, and an immutable preview hash. It performs no deletion. The explicit batch limit is 500 findings; the per-table inventory limit is 10,000 records. Oversized selections require a narrower preview. - Display the preview and retained-data warning to the operator. Only after explicit confirmation,
call
POST /arsenal/approvalsusing itsscope_receipt_id,risk_tier: "dangerous",action_name: "data.records.delete",approved_by,expires_atequal to the preview expiry, confirmationsconfirm_authorized,confirm_scope_reviewed, andconfirm_delete_records, and exactaction_context: {"preview_id":"...","preview_hash":"..."}. - Call
POST /data-deletion/executewithpreview_id,preview_hash, andapproval_receipt_id. Reuse exactly this request after an uncertain network response. A successful retry returns the original durable deletion receipt, not a second operation.
Legacy DELETE /targets/{id}, DELETE /findings/{id}, and destructive POST /findings/cleanup
also require the preview and approval; missing preconditions return HTTP 428. A preview for a
batch cannot authorize a singleton route. Changed records, expiry, cross-owner dependencies,
protected evidence, active work, and unresolved restrictive dependencies return HTTP 409.
The complete inventory is revalidated under transaction-scoped writer locks; count updates,
record removal, evidence-index detachment, and the durable result commit atomically. There is
no external storage I/O in that transaction and no scheduled destructive execution.
Completed-operation replay validates the stored manifest and approval association without
locking writer tables. Expired pending previews fail before those locks; real deletion still
performs locked revalidation. Run-state blockers follow each subsystem's terminal statuses;
unknown states and resumable states remain blockers.
What is removed, and what is retained
Target deletion removes the exact target and its owned cascading database records, including its findings and target-scoped credential profiles. The preview names the affected tables. Other target IDs survive; child-target parent links are detached. Finding deletion removes only the selected finding records and their cascading children, then refreshes owner finding counts.
Historical scans/reports, scan artifacts, exports, backups, and external content-addressed
files are not erased. Scan HTTP archives survive with the scan. A sensitive classification
alone does not prevent preserving a row; original links of retained and detached rows are
bound into the preview and kept in the durable operation receipt. Explicit legal/operational
holds and protected audit records still block even an ownership detachment.
Hunt HTTP archives have a cascading relationship with their Hunt: deleting a target erases its
own Hunt and scan transaction archives with it under the dangerous-tier approval. A sensitive
classification is a content label, not a hold: every recorded transaction carries it by default,
and treating it as a hold made any target that had ever been scanned or hunted undeletable. Only
an explicit legal_hold or audit class, or a legal_hold/operational_hold flag, blocks
erasure. Use archive to hide inventory without erasing history.
Finding-linked evidence_objects are detached before
the finding FK cascade, preserving their storage index instead of silently orphaning blobs.
A report may therefore still contain a historical copy of a deleted finding. Run the dedicated
approved evidence-retention operation first when removing eligible content is also intended.
No suppression/tombstone prevents future discovery or scans from creating new records.
Model Intake targets use their separate product lifecycle. Mixed product ownership and legal or
operational holds block this generic operation. Managed deployments must explicitly enable the
record_deletion UI capability (plus engine_admin for the Findings list and detail controls)
and authorize these endpoints at their gateway; this UI flag is not an API authorization mechanism. Standalone remains a single-user local application.
Acceptance coverage lives in tests/test_data_lifecycle.py,
tests/test_data_lifecycle_replay.py, tests/test_target_archive_admission.py,
tests/test_data_lifecycle_postgres.py, ui/tests/browser/data-lifecycle.spec.ts, and
ui/tests/browser/findings-triage.spec.ts.
The PostgreSQL test uses only the explicitly named disposable local test database; it must never
be pointed at an existing installation.
Keyed collection uploads have separate real-route acceptance in
tests/test_collection_atomic_retry_postgres.py. The maintenance workflow requires that suite
to execute against a dedicated disposable PostgreSQL database without skipped cases. Its app
lifespan is not started, so the tests do not start schedulers, workers, or target traffic.