Turn prospect research into a deduplicated CRM import
You have a list of target companies from research, and a CRM that already holds some of them. If you import the list as-is, you risk duplicate company records, overwritten fields, or rows that fail. The fix is to produce one review file before anything touches the CRM. Each row in that file says one of three things: the company is new, it matches an existing record, or it is ambiguous and needs a person to decide.
This guide covers how to build that file, how HubSpot matches records on import, and what to check before you approve the import. Research, import approval and outreach are three separate steps. This guide covers only the first two.
What you need before you start
- A prospect list from supplied or authorized public sources. Examples are a client’s target list, company websites, or public directories you are allowed to use. Record where every value came from.
- A current CRM export of the objects you plan to touch, usually companies. Include the Record ID, the company domain name and any fields you might update.
- A decision on import mode. HubSpot lets you create and update, create new records only, or update existing records only. Pick the mode once you know how many rows are new and how many are updates.
Do not infer private contact details such as personal emails, phone numbers or guessed email patterns. If a contact was not supplied to you or published by the person or company, it does not go in the file.
How HubSpot decides what counts as a duplicate
Your review table should follow the same rules the CRM uses. According to HubSpot’s deduplication documentation and its import guide:
- Contacts match on the Email property. Importing a contact whose email already exists updates that contact.
- Companies match on the primary value of the Company domain name property. That property accepts only the domain up to the top-level domain, such as
example.com. Striphttps://,www., paths and query strings before import. - If several existing companies already share a domain, the import errors and that company is not imported. Find and resolve these rows first.
- Record ID can match imported rows to existing records. When you use Record ID, it overrides every other unique identifier in the import, including email. A row without a Record ID value creates a new record.
- Custom unique properties. Each object can have up to ten custom properties that require unique values. These block duplicates during imports and manual edits.
Other CRMs use different keys. Check your own CRM’s import documentation and replace these rules where they differ.
The deliverable: an import-review table
The table below is illustrative. The companies, domains and IDs are fictional.
| Company domain | Existing record ID | Source | Proposed field | Conflict status |
|---|---|---|---|---|
northfield-example.com | — | Client target list, row 12 | Create company; Industry = Logistics | New: no domain match |
harbor-example.com | 48211 | Company “About” page | Update City: Leeds → Manchester | Field conflict: hold for review |
cobalt-example.net | 30175 | Public directory listing | Update Employee range (CRM field blank) | Match: safe update |
vale-example.org | 51002, 51990 | Company website footer | None | Ambiguous: two records share this domain |
kestrel-example.com | 22760 | Client target list, row 4 | None. Research value equals CRM value | Match: no change |
The columns do these jobs:
- Company domain is normalized: lowercase, root domain only, no protocol or path.
- Existing record ID comes from the CRM export. Never type it from memory or guess it.
- Source is a URL or a row reference in the supplied list. A value without a source does not get imported.
- Proposed field names the exact field and new value. When it replaces an existing value, it shows the old value too.
- Conflict status uses a small fixed set of labels: New, Match: safe update, Match: no change, Field conflict, Ambiguous.
Building it by hand
- Normalize the prospect list. Reduce each website to its root domain. Remove duplicates within the file before you compare against the CRM.
- Normalize the CRM export the same way. If the export has domains with
www.or trailing slashes, clean a copy. Do not edit the CRM itself. - Join on domain. A spreadsheet lookup (
XLOOKUPorVLOOKUP) from the prospect domain to the export domain returns the Record ID. - Count matches per domain. Count how many export rows share each domain (
COUNTIF). Any count above 1 marks the row Ambiguous. HubSpot will reject these rows anyway. - Compare field by field. For matched rows, put the research value next to the CRM value. A blank CRM field is a safe update. A different, non-blank value is a field conflict.
- Split the output. Updates carry a Record ID. New companies carry none. Ambiguous and conflicting rows go into neither file until someone decides.
This works for a few dozen rows. It gets slow when the research is spread across many web pages and the CRM data sits behind a login. That is the part a browser agent can help with.
A reusable prompt
Use this with any assistant that can read your sources. Edit it to fit your fields:
Compare the prospect records below with the attached CRM company export.
For each prospect, return one row with: company domain (root domain only,
lowercase), existing Record ID from the export (or blank), source URL or
list row, proposed field change with old and new value, and conflict status
(New / Match: safe update / Match: no change / Field conflict / Ambiguous).
Rules:
- Use only values from the supplied list or the cited public page.
- Do not add or guess contact names, emails or phone numbers.
- If two or more export rows share a domain, mark Ambiguous and propose nothing.
- If a research value differs from a non-blank CRM value, mark Field conflict.
- Return proposed changes only. Do not edit, create or import any record.
Pre-import checklist
- Every Record ID in the update file exists in the current export.
- Every domain is in root form with no protocol,
www.or path. - No domain in the create file already exists in the CRM.
- Rows where several records share a domain are held and listed for merge review. HubSpot has a separate tool for reviewing possible duplicates.
- Field conflicts have a named reviewer and a decision.
- Every imported value keeps its source. Store the source in a note or a dedicated property so the provenance survives the import.
- The import mode matches the file: update-only for the update file, create-only for new records.
- No private contact data was inferred.
Hold any record you are unsure about. Splitting duplicates apart later costs more than leaving them out of one import.
Where Dassi fits
Dassi is a side-panel extension for Chrome, Edge and Brave. It works in pages you are already logged into, and browser actions run on your machine. It can read your research pages and the CRM export view in your own session, then fill in the review table using the rules above. You choose which actions need your approval first. For this job, the useful setting is to allow no CRM edits at all. Ask for proposed changes only, review them, and run the import yourself.
Two points to know before you start:
- Data still leaves your machine. The AI provider you configure, or Dassi’s managed-credit relay, processes the instructions and page content involved in the task.
- The import decision stays with you. Dassi can flag a field conflict, but it cannot know which value is correct.
If you repeat this every week, you can save the comparison as a workflow and check each result in run history.
This guide is about the pre-import file. If your problem is updating CRM fields after calls, see Your Reps Aren’t Selling. They’re Filling Out Salesforce. For the same source-per-claim approach applied to candidates, see Build a candidate brief with a source for every claim.
Next step
Start small. Pick five prospect records you are authorized to use, export the matching companies from your CRM, and ask Dassi to compare them and return proposed changes only. Check its table against the checklist above before you import anything.