# P7 related-work ledger

Status: provisional paper-local author ledger

External-source verification date inherited from the binding evidence dossier:
23 July 2026

This ledger states what the manuscript may attribute to primary prior work and
what it must not imply is new. `references.bib` contains the citation records;
`primary-source-bibliography.md` records provenance and use.

| BibTeX key | Established result or system capability used by P7 | P7 boundary / forbidden novelty implication |
|---|---|---|
| `FellegiSunter1969` | Decision-theoretic record linkage with link, possible-link, and nonlink regions. | Does not establish CatDB states, calibrated posterior probabilities, clustering, authority, or canonical identity. |
| `Jaro1989` | Operational probabilistic record-linkage methodology. | Does not establish universal thresholds or cross-domain validity. |
| `HernandezStolfo1995` | Merge/purge and sorted-neighborhood duplicate detection at database scale. | Does not supply versioned evidence, scoped authority, or reversible decisions. |
| `BenjellounEtAl2009` | Generic match/merge framework with algebraic properties that enable efficient ER. | CatDB operations must prove the properties they use; typed syntax does not grant them. |
| `PasulaEtAl2002` | Probabilistic modeling of uncertain mappings from observations to entities. | Does not establish P7’s operational decision ontology or downstream invalidation. |
| `DongHalevyMadhavan2005` | Relationship-aware reference reconciliation with positive and negative evidence. | Collective evidence remains dependent and fallible. |
| `BhattacharyaGetoor2007` | Collective relational clustering for entity resolution. | Does not justify hidden propagation, double-counting, or unrecorded closure. |
| `BansalBlumChawla2004` | Correlation-clustering objective over positive and negative pair labels. | Is a partitioning baseline, not an authority or identity semantics. |
| `BienvenuEtAl2022` | Declarative hard/soft rules and denial constraints for collective ER. | Creates a high novelty bar for CatDB rule and constraint claims. |
| `BertossiEtAl2011` | Matching dependencies, matching functions, chase-like enforcement, and clean query answers. | Shows declarative matching and multiple clean outcomes predate CatDB. |
| `MichelsonKnoblock2006` | Learned blocking schemes for record linkage. | Blocking generates candidates; it does not decide identity. |
| `WhangEtAl2009` | Iterative feedback between blocking and entity resolution. | Feedback dependencies and order must remain provenance-visible. |
| `SarawagiBhamidipaty2002` | Active learning for interactive duplicate classification. | Human labels and learned functions remain workflow- and population-dependent. |
| `WangEtAl2012` | Hybrid human-machine ER with crowd batching and cost tradeoffs. | Crowd answers are evidence, not unqualified authority. |
| `GokhaleEtAl2014` | End-to-end crowdsourced entity-matching workflow. | Review orchestration itself is not a CatDB novelty. |
| `KondaEtAl2016` | Entity-matching management system covering the wider workflow. | CatDB must compare total setup, matching, review, correction, and maintenance cost. |
| `MudgalEtAl2018` | Learned representations and classifiers for entity matching. | Classifier performance is not semantic identity or authority. |
| `LiEtAl2020` | Pretrained-language-model pair classification with domain knowledge and augmentation. | Model output remains a versioned result requiring calibration and policy. |
| `MalkovYashunin2020` | Scalable approximate nearest-neighbor search using HNSW. | Retrieves nearby vectors; does not establish semantic identity. |
| `JohnsonDouzeJegou2021` | GPU similarity search and Faiss at large scale. | Is an execution baseline, not an identity model. |
| `OWL2RDFSemantics2012` | `owl:sameAs` denotes equality of individuals under OWL semantics. | Candidate, probable, contextual, or similarity links must not be exported as equality. |
| `HalpinEtAl2010` | Analysis of identity misuse around `owl:sameAs`. | Supports a conservative export boundary, not CatDB correctness. |
| `SKOS2009` | Weaker concept-mapping relations distinct from OWL individual equality. | Concept mapping does not replace a domain identity profile. |
| `PROVO2013` | Standard entities, activities, agents, roles, delegation, revision, and invalidation. | Does not define P7 states, probabilities, calibration, authorization, or equivalence. |
| `YinHanYu2007` | Iterative estimation of source trustworthiness and fact confidence. | Statistical reliability is not institutional authority. |
| `DongBertiEquilleSrivastava2009` | Source dependence and copying in conflicting-data integration. | Agreeing copied sources must not be counted as independent evidence. |
| `GneitingRaftery2007` | Proper scoring-rule theory for probabilistic forecasts. | Supplies evaluation tools, not identity decisions. |
| `SchnellBachtelerReiher2009` | Error-tolerant record linkage over Bloom-filter encodings. | Does not justify generic privacy claims without a leakage and attacker model. |

## Required comparison dimensions

Every later experimental version of P7 must compare:

- candidate recall and reduction before final matching accuracy;
- pair and cluster metrics with exact denominators;
- uncertainty, calibration, threshold, and decision-cost semantics;
- collective-dependency and copied-source handling;
- authority, ownership, review, contradiction, and reversal;
- source, schema, model, policy, and population drift;
- end-to-end setup, review, repair, and maintenance cost;
- provenance completeness and interchange loss;
- privacy assumptions, leakage, attacks, and accuracy cost; and
- downstream invalidation after reconciliation changes.

## Source-use constraints

- Prior work supports only claims under its stated assumptions.
- No primary source above is evidence that CatDB implements the corresponding
  capability.
- The manuscript does not quote these sources verbatim.
- A future source update requires a new access date and reopens citation,
  related-work, and paper review.
