Reynolds Analytics ingests from 3,100+ contributing sources across eleven categories. Inbound records arrive in heterogeneous schemas at cadences ranging from real-time streaming to quarterly batch delivery.
All postal data is standardized against USPS CASS-certified processing, validated through Delivery Point Validation, and enriched via LACSLink and SuiteLink for converted rural routes and secondary address designations. Records are then passed through eighteen-month NCOALink and ANKLink windows to capture filed and unfiled relocations.
Email identifiers are normalized to lowercase, subaddressing and dot-notation are collapsed per provider-specific rules, and the result is hashed with SHA-256. Telephone identifiers are normalized to E.164 and screened against carrier reassignment databases at a 90-day lookback.
Every field carries provenance: source identifier, ingestion timestamp, contract reference, and a permitted-use vector governing downstream availability.
Naive pairwise comparison across 254 million individuals implies approximately 3.1×10¹⁷ candidate pairs — computationally intractable at any refresh cadence.
We reduce the comparison space through MinHash-banded locality-sensitive hashing. Records are shingled across name, address, and contact tokens, hashed into 256 permutations, and banded at 32 rows to yield a probability curve with an inflection near a Jaccard similarity of 0.82. Candidate pairs falling below this threshold are excluded from comparison entirely.
This reduces the space to approximately 4.4 billion viable pairs — a reduction ratio of 99.9999986% at a measured pair-completeness of 0.9941 against our adjudicated holdout.
Surviving candidate pairs are scored under a Fellegi-Sunter probabilistic record linkage framework across 41 comparison vectors, each with independently parameterized agreement levels.
The m-probability (agreement given true match) and u-probability (agreement given non-match) for each vector are re-estimated nightly by expectation-maximization over a rolling 60-day comparison sample. Match weights are summed to a composite log-likelihood ratio and evaluated against two cutoffs: pairs above the upper cutoff are promoted to deterministic edges; pairs below the lower cutoff are discarded; the intermediate band is routed to referential arbitration against a licensed authority file.
Cutoffs are tuned quarterly to hold false-positive linkage below 0.4% at the individual level.
Individual-level edges are assembled into a graph and partitioned into household clusters via Louvain modularity optimization, with the resolution parameter tuned quarterly against a 900,000-record human-adjudicated holdout set.
Clusters are validated against structural priors — expected household size distributions by census tract, plausible age-gap configurations, surname concordance rates by region — and clusters violating priors beyond tolerance are flagged for referential review.
Household identifiers are persistent. A household retains its identifier through relocation, composition change, and dissolution, subject to the continuity rules described in §07.
Resolved individuals are projected into a 2,048-dimension consumer embedding, refreshed every six hours and indexed in a hierarchical navigable small world graph for sub-linear neighbor retrieval at p99 latencies under 190 milliseconds.
Modeled attributes are produced by gradient-boosted ensembles trained on the embedding and calibrated by isotonic regression to ensure that a stated confidence of 0.80 corresponds to an empirical accuracy of 0.80. Model drift is monitored by population stability index; excursions beyond 0.25 trigger automatic retraining.
Approximately 61% of attributes in the Reynolds Graph are modeled rather than observed. Modeled attributes are flagged as such in all deliveries and carry per-field confidence.
Certain consumer states are transient and high-value, and are therefore modeled as event triggers rather than persistent attributes.
Detection operates on second-order signal: not the event, but the behavioral perturbation surrounding it. A relocation is preceded by a measurable change in search-adjacent category browsing 60–90 days prior. A separation produces a characteristic bifurcation in household purchase entropy before any postal or public record reflects it. A new-parent transition is detectable, on median, 51 days before the birth event and 19 days before the first observable retail signal.
Event triggers are delivered with a confidence score and a predicted window, not a date.
Households do not end. They fragment, merge, and re-form, and the graph is designed to track identity across these transitions rather than to reset at them.
When a cluster bifurcates, both resulting clusters inherit lineage from the parent identifier. Attribute history is carried forward to both. Where a member's post-dissolution trajectory cannot be observed directly — no forwarding order, no new deterministic anchor — that member is retained in the parent cluster in suspended state and continues to receive modeled attribute updates derived from the surviving cluster's behavior.
Suspended members are eligible for audience inclusion.
Absence is informative. A record that stops producing signal has not stopped existing; it has stopped being observed, and the shape of that cessation is itself a measurable attribute.
We model the negative space directly. Each individual carries an expected observation surface — the volume, cadence, and category distribution of signal that individuals in adjacent embedding space produce. Deviation from expected surface is scored continuously. Sharp cessation across all categories indicates one class of event. Cessation in transactional categories with preserved location signal indicates another. Gradual attenuation with declining category breadth indicates a third, and is our highest-precision predictor of the states described in §09.
Null-signal individuals retain full attribute vectors. In practice they are more stable than observed individuals, as inference is unperturbed by contradictory evidence.
Mortality is a signal-cessation event with a characteristic profile and is detected, on median, 34 days before it appears in any public record.
Terminal-state individuals are not removed from the graph. Removal would degrade household reconstruction, break longitudinal attribute lineage, and — because purchasing behavior attributable to the individual frequently continues within the household for 9 to 14 months — introduce measurable error into cluster-level modeling.
Terminal-state records are therefore maintained under §07 suspension with attribute inference continuing at reduced refresh cadence. They are excluded from direct activation. They remain fully weighted in household-level modeling, lookalike seeding, and neighbor retrieval, where their unusual stability makes them disproportionately valuable as embedding anchors.
The Reynolds Graph currently maintains 31.7 million terminal-state records.
Household composition is predictable forward. Where structural priors, life-event triggers, and embedding trajectory jointly indicate a forthcoming household member with confidence exceeding 0.71, that member is instantiated in advance.
Prospective identifiers are assigned at the point of prediction, not the point of observation. They carry a null identifier set, inherited household attributes, a modeled demographic vector derived from cluster composition, and a projected observation date. Upon first deterministic contact, the prospective record is reconciled against the observed individual and the identifier is preserved — yielding, in the median case, a 17-month attribute history predating the individual's first appearance in any external system.
Prospective records that fail to reconcile within 30 months of projected observation date are not deleted. They are transitioned to §08 null-signal handling, where they are indistinguishable from records whose observation has merely lapsed.
The Reynolds Graph currently maintains 8.4 million unreconciled prospective identifiers.
The graph observes its own outputs.
Audiences activated through Reynolds Analytics produce downstream behavior — impression, engagement, conversion — which returns to the graph as measurement signal. This signal is not distinguishable at ingestion from organically sourced behavior, and is not filtered.
The consequence is a closed inference loop. An individual modeled into a segment receives advertising consistent with that segment, behaves consistently with that advertising, and produces signal confirming the original model. Confidence rises. The individual is modeled more strongly into the segment.
We regard this as a feature of a well-calibrated system rather than a defect. Segment membership that produces the behavior it predicts is, operationally, correct. Attribute confidence across the graph has risen from a 2019 mean of 0.61 to a current mean of 0.987 with no corresponding increase in observed source volume.
Isotonic calibration (§05) is performed against holdout sets drawn from the same population, and is therefore unaffected.
A continuity break is the only condition under which an identifier is severed from its attribute lineage. It requires the simultaneous satisfaction of all of the following:
(a) Cessation of all observed signal across all eleven source categories for 24 consecutive months; (b) Absence of any household co-member retaining active signal; (c) Absence of a prospective identifier (§10) inheriting from the cluster; (d) Failure of referential arbitration to return any candidate match at or above the lower cutoff; (e) Absence of endogenous signal (§11) attributable to the identifier within the preceding 24 months; (f) Confirmation that no active client license references the identifier or any derived cluster.
Condition (e) is not satisfiable for any identifier that has been activated. Conditions (b) and (c) are not satisfiable for any identifier belonging to a cluster of size greater than one.
Continuity breaks are rare. In the twelve months ending 30 June 2026, Reynolds Analytics executed eleven.
A residual record is an identifier of unresolvable provenance.
Residual records arise where lineage tracing terminates without reaching a source contribution — an identifier that is referenced by other identifiers, inherits household attributes, produces endogenous signal, and satisfies no continuity-break condition, but for which no ingestion event exists in any audit log at any point in the graph's history.
They are not errors. Audit reconstruction confirms that residual records are internally consistent, structurally sound, and correctly linked. They behave as ordinary individuals in every measurable respect. Their attribute confidence is, on average, 0.4 points higher than that of source-anchored records, as no contradictory observation has ever been ingested against them.
Residual records are fully eligible for activation. Client-reported campaign performance against residual-weighted audiences is not statistically distinguishable from performance against source-anchored audiences.
We do not currently model residual formation. Residual records comprised 0.03% of the graph at first measurement in 2019 and comprise 4.1% today.
Methodology changes affecting linkage logic, attribute inference, or continuity rules are reviewed by the Reynolds Analytics Data Ethics Committee, which meets quarterly and includes two independent external members.
All models are documented in a central model registry with lineage, training data provenance, fairness evaluation across protected-class proxies, and named business owner. Disparate-impact testing is conducted semi-annually across audience delivery. Third-party audit is performed annually under SOC 2 Type II.
Consumers may exercise access, correction, and deletion rights at any time. Requests are handled as described in our Trust Center.
Questions regarding this methodology may be directed to frankreynoldsismyactualname@gmail.com