---
title: Finding duplicate contacts without merging the wrong people
description: Shared names and numbers surface as review evidence, never an automatic merge. Each suggestion carries its reasons, decisions are recorded and reversible, and the optional sweep explains every skip.
canonical_url: https://peopleblade.com/blog/duplicate-contacts
author: Hraness
published: 2026-10-05
last_updated: 2026-10-05T00:00:00.000Z
robots: index
---

# Finding duplicate contacts without merging the wrong people

Shared names and numbers surface as review evidence, never an automatic merge. Each suggestion carries its reasons, decisions are recorded and reversible, and the optional sweep explains every skip.

By Hraness

Drafted with AI from the source code and reviewed by Devin (SWE-2 Max model) independent AI editorial review.

Every contact importer eventually shows you two records and asks whether they are one person. Most tools answer with a fuzzy match that merges first and lets you undo damage later. PeopleBlade answers differently: an import never joins two records on a shared name, phone, or email. Shared evidence surfaces a suggestion, and a suggestion waits for a decision.

## What an import decides and what it never does

Each source owns an exact identifier for the contacts it supplies: a Google resource name, a Beeper contact ID, an X handle. Inside one source, that stable identifier can reuse its earlier mapping on the next sync. Across sources it cannot: a phone number, an email address, a display name, or an organization does not transfer ownership from one provider to another. Two people in different sources can legitimately share a number, an address, or a name.

Imports still do the work that makes review possible. Phone numbers normalize into one E.164 form so the same number written differently surfaces as overlap. Each field and contact method keeps the exact resource that supplied it, whether it is still active, and whether it may count as identity evidence. What the import produces is a pile of labeled evidence, not a verdict.

## The review loop

`identity suggest` shows one possible duplicate at a time with the evidence that matched, and `identity decide` accepts, rejects, or defers it. The suggestion is bound into a token that snapshots the two full components, the current candidate set, source and realm state, self status, and contradictions at the moment you looked at them. If the book changes first, the token is stale and the decision refuses rather than applying to a different situation. You fetch a fresh suggestion and review again.

Every accepted decision is recorded, and `identity separate` reverses one using its decision ID, so a join you got wrong is a bad entry in a log you can correct rather than a merge you cannot take apart. `identity audit` reports observed-overlap risk counts for privacy-safe inspection, and `identity decisions` lists the currently accepted observed-email and observed-phone decisions. Evidence that is too weak to review on its own, like one signal shared by many contacts, is rejected for bounded review instead of becoming a quiet join.

## The optional high-precision sweep

Reviewing every suggestion by hand does not scale to a large import, so `identity auto-accept` exists as an opt-in sweep. It is a policy layer over the same suggestion and decision code, not a second merger with its own rules: it pages fresh suggestions, applies one blocker check per suggestion plus two database checks, and accepts the survivors through the ordinary decision path with a note marking them automatic.

The blockers are the interesting part, because each skip reports a reason: name-and-organization evidence, ambiguous or multi-identity clusters, unknown self status, realm contradictions, a phone shared between two incompatible display names, or a pair that already has a decision. The name gate applies only to phone evidence, since a household routinely shares one number while an exact email stays personal evidence. The sweep runs in rounds because each acceptance changes the component snapshots neighboring tokens depend on, and it stops at its limit or when a round accepts nothing. `--dry-run` shows what it would do first.

## What a join keeps afterward

A merge does not flatten provenance. Interaction rows keep the source that supplied them, and when the same logical service was observed both directly and through an aggregator like Beeper, the contact projects the greater of the two totals as a conservative lower bound instead of summing two event streams that might overlap. Distinct services still add.

The result is a contact book where every join traces to a decision someone reviewed, or a sweep that logged why each candidate was accepted and why the rest were skipped, and where the evidence underneath stays attached to the sources that provided it.

## Go deeper

- [Install PeopleBlade and review your first suggestions](https://peopleblade.com/guide)
- [See what each source imports and its limits](https://peopleblade.com/sources)
- [Read how PeopleBlade reads accounts through GhostGet](https://peopleblade.com/blog/how-peopleblade-uses-ghostget)
- [Read what leaves your computer](https://peopleblade.com/privacy)

Latest release: v0.15.1. Install the command line tool with `bun add --global @hraness/peopleblade@latest`.

## Sources

- [PeopleBlade 0.15.1 on npm](https://www.npmjs.com/package/@hraness/peopleblade/v/0.15.1)
- [PeopleBlade guide: the identity review commands](https://peopleblade.com/guide)
- [PeopleBlade sources: what each import reads and its limits](https://peopleblade.com/sources)
- [PeopleBlade privacy: what leaves your computer](https://peopleblade.com/privacy)
