Cleaning up HubSpot duplicates before a migration
Migrating a CRM full of duplicates bakes the mess into the new system — and most migration tools match on exact email only, so near-duplicates sail through. Clean first, migrate second.
Why duplicates happen
- CSV imports without dedup checks — the same list imported twice, or overlapping lists from different owners.
- Integrations creating contacts independently: web forms, chat, event tools, and ad platforms each mint their own record.
- Manual entry with typos, nicknames, or personal vs work emails for the same person.
- Sync conflicts between HubSpot and a second system that disagree on the canonical record.
The dedup workflow
- Export. In HubSpot: CRM → Contacts → select the view → Export (or Settings → Import & Export). Keep the HubSpot record ID column — you need it to merge later.
- Normalize. Make equivalent values identical before comparing (see below).
- Match. First pass: exact matches on normalized email. Second pass: fuzzy matches on name + company + phone for records without a shared email.
- Review. Never auto-merge on fuzzy matches. A human confirms each candidate pair — especially for common names.
- Merge. In HubSpot: Contacts → Actions → Manage duplicates, or merge manually, keeping the record with the most engagement history as the survivor.
- Export the clean set as your import-ready CSV for the migration target.
Normalization examples
Run this over the export before matching. It uses only the Python standard library:
import csv, re
def norm_email(v):
return (v or "").strip().lower()
def norm_phone(v):
digits = re.sub(r"\D", "", v or "")
return digits[-10:] if len(digits) >= 10 else digits # compare local part
SUFFIXES = {"inc", "llc", "ltd", "co", "corp", "corporation", "gmbh"}
def norm_company(v):
words = re.sub(r"[^a-z0-9 ]", "", (v or "").lower()).split()
return " ".join(w for w in words if w not in SUFFIXES)
with open("contacts.csv", newline="", encoding="utf-8-sig") as fh:
for row in csv.DictReader(fh):
row["email_key"] = norm_email(row.get("Email"))
row["phone_key"] = norm_phone(row.get("Phone"))
row["company_key"] = norm_company(row.get("Company"))
# rows sharing (email_key) are exact dupes; sharing
# (phone_key, company_key) are review candidates
print(row["email_key"], row["phone_key"], row["company_key"])
Expected output (one line per contact, keys aligned for comparison):
j.smith@acme.com 5551234567 acme
j.smith@acme.com 5551234567 acme ← exact dupe, same email_key
jsmith@acme.com 5551234567 acme ← review candidate: same phone+company, different email
What a merge plan looks like
- Survivor rule: keep the record with the most recent engagement; copy missing fields from the loser.
- Field-level source of truth: decide per field (e.g. email from the most recently updated record, phone from the manually verified one).
- Merge log: record every merged pair as
loser_id → survivor_idso the migration can remap associations (deals, tickets, notes).
Common failures
- Matching on raw values.
J.Smith@Acme.comandj.smith@acme.comare the same person; without normalization they never match. - Auto-merging fuzzy matches. Two different “John Smith”s at two different companies become one contact.
- Dropping the record ID. Without IDs you can't merge in HubSpot or remap associations afterward.
- Migrating first, cleaning later. The new system's dedup is usually worse than the old one's.
Go further
The free csv-duplicate-inspector scans any CSV for exact-duplicate rows, key-column dupes, and near-duplicates before you touch the CRM. For the full job — fuzzy matching, merge review, import-ready CSVs, and machine-readable change logs — see the CRM Dedup & Migration Cleanup Kit ($149, one-time, offline desktop app).
Built by Payload
Payload builds practical software that makes AI, automation, and business infrastructure safer, cleaner, more reliable, and easier to ship. Support: kylers.partners@gmail.com