re-doc
Try for free
Guide / De-identificationPage 1 of 10
Solutions · Guide / De-identification

Document de-identification
with the layout intact.

De-identification removes or transforms the information in a document that identifies a person, so the document can be used or shared without exposing them. For unstructured documents such as PDFs, Word files, text files and notes, that means finding names, dates, IDs and addresses inside free text and tables, then replacing them. Re-Doc replaces them with consistent synthetic data and keeps the original layout. Scanned pages and images are redacted with black boxes.

1. Your documentdischarge_summary.docx
Discharge Summary
Example · Cardiology ward
PatientMaria Alvarez
Date of birth03 Feb 1958
MRN88-21904
Phone(555) 014-2236
Address41 Linden Rd, Springfield
Ms. Alvarez was admitted with chest pain and discharged on day 3 with follow-up arranged at the clinic.
5 pieces of sensitive data foundIdentifiable patient details.
RE-DOC
2. Safe to sharedischarge_summary_deidentified.docx
Discharge Summary
Example · Cardiology ward
PatientElena Moreno
Date of birth19 Nov 1957
MRN47-60318
Phone(555) 013-7741
Address12 Birch Ave, Fairview
Ms. Moreno was admitted with chest pain and discharged on day 3 with follow-up arranged at the clinic.
Replaced with synthetic dataSynthetic data. Readable and consistent.

Illustrative example with invented data. The exact transformations, for example reducing dates to a year, follow the de-identification rules your team sets.

2HIPAA de-identification methods: Expert Determination and Safe Harbor
18Identifier types listed in the HIPAA Safe Harbor method
PDF · DOCX · TXTReplaced with synthetic data. Scans and images are redacted
10 pagesFree to try, no card needed
The problemPage 2 of 10
§ 02 · The basics

What de-identification means, and where it gets hard

The goal is simple. The difficulty is that identifying details are scattered through free text, and one miss can undo the work.

01 · TERMS

De-identification, anonymization and pseudonymization

“De-identification” is the US HIPAA term. “Anonymisation” is the EU GDPR term and is a stricter test: anonymous data is outside the GDPR only if people can no longer be identified by any means reasonably likely to be used. “Pseudonymisation” swaps identifiers for codes but keeps a way back, so the data stays personal data under the GDPR.

02 · UNSTRUCTURED DATA

Identifiers hide in free text

In a database, a name lives in one column. In a discharge summary, a contract or a letter, it appears in a sentence, a header, a signature block and a table, spelled in more than one way. Finding every instance is the hard part.

RE-DOC
Names, initials and nicknamesDates, ages and locationsPhone, email and ID numbersDetails that identify only in combination
03 · UTILITY

Black boxes break the document

Blacking out text removes the information, but it also removes context. Reviews, research and AI workflows need documents that still read naturally. Replacing identifiers with synthetic values keeps the document usable.

04 · CONSISTENCY

The same person must stay the same

If one person gets two different replacement names in one document, or the same replacement as someone else, the result is confusing or misleading. Replacement has to be consistent across the whole document and across a set of documents.

How it worksPage 3 of 10
§ 03 · How it works

How Re-Doc de-identifies a document

Re-Doc does the document-level work. Your team keeps the decisions and the review.

STEP 1

Upload

A PDF, DOCX, TXT file or a scan, in the web app or through the API.

STEP 2

Detect and replace

Identifiers are found in the text and replaced with consistent synthetic data, or redacted where you need removal. Scans are redacted with black boxes. The layout stays as it was.

STEP 3

Review

You check the result and edit anything that needs a human decision.

STEP 4

Release

Download a document in the same format, ready for sharing, research or testing.

Use casesPage 4 of 10
§ 04 · Use cases

Where document de-identification is used

The same approach applies wherever a document has to travel without the person in it.

Use case 01 · Guide / De-identificationHIPAA · PHI

Healthcare records and clinical notes

Discharge summaries, referral letters and clinical notes shared for research, quality review or training. See healthcare.

  • Patient identifiers replaced consistently
  • Clinical text stays readable
Use case 02 · Guide / De-identificationEMA Policy 0070 · Health Canada PRCI

Clinical study reports

Reports prepared for public release or data sharing. See clinical trial anonymization.

  • Participant identifiers handled across narratives and listings
  • Your team owns the risk assessment
Use case 03 · Guide / De-identificationLLM · RAG · Test data

Data for AI, testing and analytics

Real documents are often too sensitive to feed into AI tools or test environments. De-identified copies keep realistic structure without the real people.

  • Same layout and format as the source
  • Realistic synthetic values, not placeholders
Use case 04 · Guide / De-identificationGDPR · DPDP

Legal, HR and insurance files

Contracts, case files, claims and personnel documents shared with third parties or across borders. See legal and insurance.

  • Names, IDs and contact details replaced
  • Redaction available where removal is required
HIPAA methodsPage 5 of 10
HIPAA

HIPAA de-identification: Safe Harbor and Expert Determination

The HIPAA Privacy Rule, at 45 CFR 164.514(b), gives two ways to treat health information as de-identified.

Expert Determination. A person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods determines that the risk is very small that the information could be used, alone or in combination with other reasonably available information, by an anticipated recipient to identify an individual, and documents the methods and results of the analysis.

Safe Harbor. The identifiers of the individual, and of relatives, employers and household members, listed below are removed, and the organisation does not have actual knowledge that the remaining information could be used alone or in combination with other information to identify the individual.

The 18 identifier types in the Safe Harbor method are:

  • Names
  • Geographic subdivisions smaller than a State, including street address, city, county, precinct and ZIP code (with a limited exception for the first three digits of a ZIP code)
  • All elements of dates, except year, directly related to an individual, and ages over 89
  • Telephone numbers
  • Fax numbers
  • Email addresses
  • Social Security numbers
  • Medical record numbers
  • Health plan beneficiary numbers
  • Account numbers
  • Certificate and license numbers
  • Vehicle identifiers and serial numbers, including license plates
  • Device identifiers and serial numbers
  • Web URLs
  • IP addresses
  • Biometric identifiers, including finger and voice prints
  • Full-face photographs and comparable images
  • Any other unique identifying number, characteristic or code

Removing these types is not the whole test. The “actual knowledge” condition means free-text details outside the list can still matter, and the other route, Expert Determination, exists for cases where Safe Harbor is too blunt or not enough. Choosing the method, and signing off the result, belongs to your privacy officer or a qualified expert.

Terms comparedPage 6 of 10
Terms

De-identification, anonymization, pseudonymization, redaction and synthetic replacement

The words are used loosely. This is how they differ in practice.

  • De-identification is the broad term, and the HIPAA one. It means taking identifying information out of data or documents, by any suitable method.
  • Anonymisation is the GDPR term. It requires that people can no longer be identified by any means reasonably likely to be used (Recital 26), which is a stricter outcome than removing a list of identifiers.
  • Pseudonymisation replaces identifiers with codes or aliases while a way back exists. Under the GDPR (Article 4(5)) it reduces risk but the data remains personal data.
  • Redaction permanently blacks out or deletes content. It is simple and often required, but it removes context.
  • Synthetic replacement swaps each identifier for a realistic, different value and keeps the document readable and consistent. It is how Re-Doc de-identifies by default, with redaction available alongside it.

For a longer walkthrough, read redaction vs anonymization vs pseudonymization. For technical guidance on de-identification techniques and governance, see NIST SP 800-188.

FAQPage 7 of 10
FAQ

Document de-identification: common questions

01What is document de-identification?
It is the process of removing or transforming the information in a document that identifies a person, such as names, dates, addresses and ID numbers, so the document can be used or shared without exposing them. For unstructured documents, including PDFs, Word files, clinical notes and scans, the identifying details sit inside free text and tables, so they must be found before they can be changed.
02What is the difference between de-identification and anonymization?
The terms overlap, and each regime defines them differently. “De-identification” is the term used in the US HIPAA Privacy Rule. “Anonymisation” is the term used in EU data protection law, where truly anonymous data is outside the GDPR. Anonymisation is a higher bar: if a person can still be identified by any reasonably likely means, the data is not anonymous.
03Is pseudonymization the same as de-identification?
No. Pseudonymisation replaces identifiers with a code or alias but keeps a way to link back. Under the GDPR, pseudonymised data is still personal data. The Article 29 Working Party opinion on anonymisation states: “Pseudonymisation is not a method of anonymisation.”
04What are the two HIPAA de-identification methods?
The HIPAA Privacy Rule, at 45 CFR 164.514(b), allows two methods: Expert Determination, where a qualified expert documents that the risk of identification is very small, and Safe Harbor, where 18 types of identifiers are removed and the organisation has no actual knowledge that what remains could identify the individual.
05Is removing the 18 identifiers enough?
Not by itself. Safe Harbor has a second condition: the organisation must not have actual knowledge that the remaining information could be used, alone or in combination with other information, to identify the person. Free text can contain identifying details outside the 18 listed types, so review matters.
06Does Re-Doc make a document HIPAA or GDPR compliant?
No tool can promise that on its own. Re-Doc does the document work: it finds identifiers in unstructured documents and replaces them with consistent synthetic data, or redacts them, and keeps the layout. For identifiers beyond text, such as photographs, or for custom deployment, contact our team to discuss your options. Whether the result meets Safe Harbor, Expert Determination or the GDPR anonymisation standard is a decision for your privacy or compliance team.
07Why replace identifiers with synthetic data instead of blacking them out?
Black boxes remove the information but also make the document harder to read and use. Replacing a name with a different, realistic name, and a date with a different date, keeps sentences, tables and context intact. This matters when the document is used for review, research, testing or AI workflows. Where a regulation or policy requires removal, Re-Doc can redact instead.
08Which file types are supported?
PDF, DOCX and TXT files, plus scanned documents and images. Native files get synthetic replacement and keep their original format and layout. Scanned pages and images are redacted with black boxes. Other formats can be discussed as part of a pilot.
09Can we try it on our own documents?
Yes. The first 10 pages are free in the web app, and we run scoped pilots on sample documents so your team can check results before any wider use.

Reviewed by the Re-Doc team. Last reviewed 1 October 2026. This page is general information, not legal advice. Check the current text of the regulations, and your own obligations, with your privacy or legal counsel.

ComparisonPage 8 of 10
§ 05 · How Re-Doc compares

Re-Doc next to manual de-identification

Most de-identification of documents is still done by hand, one page at a time. Re-Doc gives your reviewers a consistent first pass.

#Typical tools in the marketRe-DocStatus
01Reading every page and blacking out identifiers by hand.Identifiers found across the whole document in one pass, then reviewed.● Covered
02Black boxes that make text hard to read.Realistic synthetic values that keep the text readable.● Covered
03Different treatment of the same person in different places.The same person gets the same replacement throughout.● Covered
04Reformatting after the edit, or a flattened image.Same file format and layout as the original.● Covered
SourcesPage 9 of 10
Get startedPage 10 of 10

Try it on your own documents.

The first 10 pages are free in the web app. For larger sets or custom deployment needs, contact our team: we will run a scoped pilot on a sample and share results you can check.