Redact or Replace: How to Anonymize Documents Without Losing What Matters
Redaction removes information. Synthetic replacement swaps it for realistic values so the document still works. This guide explains how to choose, and where each method fails.

Redaction removes information. Synthetic replacement swaps it for realistic values so the document still works. This guide explains how to choose, and where each method fails.
Redact when the document has to show that something was removed, or when the page is a scan with no text to work with. Replace identifiers with synthetic data when people still need to read, search or analyse the document afterwards. Many organizations need both, because one source document often goes to several recipients with different needs.
Redaction permanently removes the sensitive content, usually with a black box. Synthetic replacement swaps each identifier for a realistic, different value and applies it consistently, so a name that appears 23 times in a report becomes the same new name 23 times.
| Redaction | Synthetic replacement | |
|---|---|---|
| What happens to the identifier | Removed or blacked out | Swapped for a realistic, different value |
| What the reader sees | A black box or a "[REDACTED]" tag | Normal text |
| Search and cross-references | Broken for the removed text | Still work, using the new values |
| Visible proof that something was removed | Yes | No, unless you mark it |
| Works on scanned pages | Yes, with OCR | Needs text in the file, so it suits native PDF, DOCX and TXT |
| Typical use | Public records releases, court filings, scans | Review, research, testing, AI workflows |
Redaction fails when it is only cosmetic. A black rectangle drawn over text, or a highlight set to black, hides the text on screen but can leave it in the file. In January 2019, lawyers for Paul Manafort filed a court document in which the passages meant to be redacted could be read by highlighting the black bars and pasting the text into a new document [4]. The problem is also not confined to famous cases: a 2024 Federal Judicial Center study of 4,681,055 documents filed in US federal courts on 37 sample days in 2022 found 22,391 unredacted Social Security numbers, in 4,525 documents [5]. That study counted exposed numbers. It did not say how each one happened, but it shows how often identifiers slip through.
Replacement fails in other ways:
Replacement is not automatically anonymization in the legal sense. Under the GDPR, data is anonymous only if people can no longer be identified by any means reasonably likely to be used (Recital 26) [7]. The EU's Article 29 Working Party stated that pseudonymisation is not a method of anonymisation and that pseudonymised data remains personal data [8]. Whether a given output meets a legal standard is a decision for your privacy or legal team.
| Recipient or situation | Typical method | Why |
|---|---|---|
| Public records requester (FOIA, RTI) | Redaction | Visible removal is expected |
| Court filing | Redaction, after checking the court's rules | Many courts expect visible marks |
| Scanned page or fax | Redaction | No text layer to replace |
| Review team working across many files | Replacement | Cross-references survive |
| Clinical or research sharing | Replacement, then expert review | Documents stay readable |
| Vendor or offshore processing | Replacement, after checking the legal basis | Staff must read to do the work |
| Test data or AI workflows | Replacement | Realistic structure without real identities |
These are common patterns, not legal advice. Your privacy or legal team makes the final call.
Re-Doc replaces identifiers in native PDF, DOCX and TXT files with consistent synthetic data and keeps the original layout. Scanned pages and images are redacted with black boxes. Files that mix text and images can be handled in one run. You can review and edit the result before you download it. The first 10 pages are free at re-doc.com/try, and the API is available on paid plans. For custom deployment or specific requirements, contact our team. For a wider view of the terms used here, see our guide to document de-identification and the comparison of redaction, anonymization and pseudonymization.
Not automatically. Replacement is a method. Whether the result counts as anonymized or de-identified depends on the legal standard that applies (HIPAA, GDPR or another) and on whether the remaining details could still identify someone. Your privacy or legal team decides.
It is a redaction that only covers text visually, for example with a black rectangle or highlight, while the text stays in the file and can be copied out. Real redaction removes the underlying content.
Redaction is the usual choice, because requesters and agencies expect visible removal and often need to cite an exemption. Replacement is better suited to sharing documents for review or research.
Scanned pages are images with no text layer, so Re-Doc redacts them with black boxes instead of replacing text.
Not without thought. Check that nothing identifying remains in the text, metadata or images, and confirm with your privacy or legal team that sharing is allowed for your purpose.
Reviewed by the Re-Doc team. Last reviewed 1 October 2026. This article is general information, not legal advice.