0006 Source Matched Legacy Evidence Salvage
Source-Matched Legacy Evidence Salvage Is the Legacy Reuse Pattern
Section titled “Source-Matched Legacy Evidence Salvage Is the Legacy Reuse Pattern”Status: accepted
Legacy Data Gov output may be reused across slices only after it matches a current primary-source-derived source or normalized record. Reused legacy output enters the Dagster graph as evidence claims, curation/review input, or coverage/reuse metadata; canonical and q_* models must not read migration.* legacy staging tables directly.
Considered Options
- Treat legacy Data Gov rows as source authority and import them directly into canonical/query models.
- Ignore legacy Data Gov output and pay for fresh deterministic/LLM gap-fill across all historical rows.
- Reuse legacy output only after source matching, provenance capture, validation, and normal evidence/resolution/canonical gates.
Consequences
Each salvage slice defines its own source match keys and validation gates: trials match by NCT/ChiCTR registry id, publications by PMID/DOI/source id, approvals by source id/application/product/generic, news by release/source-native article id, and SOC by source document/context. Legacy identity rows are reconciliation baselines or review candidates, not canonical identities. Fresh paid LLM work is gap-fill only after deterministic extraction and source-matched legacy reuse have measured the remaining uncovered or rejected rows.