re-doc
Try for free
Legal / eDiscoveryPage 1 of 9
Solutions · Legal / eDiscovery

eDiscovery redaction for 40,000 documents
without exposing a single SSN.

Re-Doc is eDiscovery redaction software that actually removes the data, not just covers it. Discovery production sends your client identifiers to opposing counsel, BPO reviewers, and case law databases. Black boxes break AI tools and expose text layers. Synthetic replacement swaps every identifier with consistent synthetic data. Same layout. Same Bates stamps. Fully searchable.

1. Your documentcomplaint.pdf
Complaint
Dallas County District Court · Exhibit 12
PlaintiffJohn Michael Smith
SSN423-88-1924
DOBMarch 4, 1981
Address1247 Oak Ave, Dallas TX 75201
CounselDavis & Whitmore LLP
Plaintiff John Michael Smith alleges that on March 4, 2023, defendant caused injury at the above address. Counsel Davis & Whitmore LLP filed on behalf of the plaintiff in Dallas County District Court.
5 pieces of sensitive data foundContains real PII. Not safe to produce.
RE-DOC
2. Safe to sharecomplaint_safe.pdf
Complaint
Dallas County District Court · Exhibit 12
PlaintiffRobert James Wilson
SSN571-34-6289
DOBJuly 19, 1979
Address834 Pine St, Austin TX 78701
CounselHarmon & Burke LLP
Plaintiff Robert James Wilson alleges that on March 4, 2023, defendant caused injury at the above address. Counsel Harmon & Burke LLP filed on behalf of the plaintiff in Dallas County District Court.
Replaced with synthetic dataSynthetic data. Safe to produce. Fully searchable.

All PII fields replaced with consistent synthetic data. Legal arguments, dates, venue, and cause of action preserved. Bates stamps and exhibit numbers unchanged.

$16.9BGlobal eDiscovery market in 2024
22,391SSNs found unredacted in PACER court filings
70%Law firms not yet using generative AI tools (ABA 2024)
The problemPage 2 of 9
§ 02 · The problem

Five ways black-box redaction fails in legal practice

These are not hypothetical failure modes. They are documented incidents and structural limitations that affect every discovery production, court filing, and DSAR response.

01 · DOCUMENTED INCIDENT

PACER: 22,391 SSNs in public filings

A 2024 Federal Judicial Center study reviewed about 4.7 million public PACER documents filed on 37 sample days in 2022. It found 22,391 unredacted Social Security numbers belonging to about 8,300 people, and 72% of them appeared to break the courts’ privacy rules.

RE-DOC
The lesson for litigators: identifiers slip through at volume, often concentrated in a few large filings. Detection has to run on every page, not only the pages someone remembers to check.
02 · CAREER RISK

Manafort filing: journalists copy-pasted through the redactions

In January 2019, lawyers for Paul Manafort filed a response to Special Counsel Mueller with black overlays over sensitive passages. Within hours, reporters including those from BuzzFeed News discovered they could copy-paste the text directly from the PDF. The same principle appeared in the Sony disclosures during FTC v. Microsoft proceedings, where confidential game development budgets were redacted with a physical pen but remained legible when the document was scanned.

RE-DOC
Failed redactions like these become public quickly and are hard to undo. Lawyers also have professional duties to protect client information, so a missed identifier is a professional risk, not only a technical one.
03 · TECHNICAL ROOT CAUSE

The image layer vs. the text layer

PDFs have two distinct layers: the visual image layer (what you see on screen) and the text data stream (what a program reads). Many PDF tools paint a black rectangle over the image layer. The text data stream, the actual PII, often remains fully intact. Copy-paste, screen readers, and AI systems all extract from the text layer, not the image layer.

RE-DOC
Visual layer: black box painted overText data stream: PII still accessible
04 · OPERATIONAL FAILURE

Redacted documents are useless downstream

Case law databases built on black-box redactions cannot be full-text searched. RAG systems trained on [REDACTED] tokens return degraded answers. BPO teams processing blacked-out files cannot extract the data they need. DSARs with redacted sections cannot be used in court. Technically compliant. Operationally broken.

05 · SCALE PROBLEM

Redaction fatigue: inconsistent coverage across duplicate documents

When a 40,000-document production contains the same plaintiff SSN across hundreds of separate documents, each one requires an individual human decision. At that scale, error rates climb. The same SSN gets caught in most documents and missed in a handful. Courts treat inadvertent disclosure as a breach regardless of volume. One missed identifier in a 40,000-document production carries the same sanctions exposure as wholesale non-compliance. Re-Doc processes the entire batch with a single entity map. Every instance of every identifier, consistently handled across every duplicate, in one pass.

How it worksPage 3 of 9
§ 03 · How it works

Re-Doc sits between your source documents and every downstream recipient

One API call. De-identified, layout-perfect output every time. Plugs into your existing eDiscovery workflow without changing how your team operates.

STEP 1

Source documents

Deposition transcripts (including TXT), scanned exhibits, native PDFs. Any format from any eDiscovery platform.

STEP 2

Re-Doc processes

Visual processing reads every pixel in scanned exhibits. A context-aware model detects all PII. Synthetic replacement or true redaction applied to both image and text layers.

STEP 3

Layout-perfect output

Same pagination, same Bates stamps, same exhibit numbering. Formatting preserved. The document looks exactly as it did, minus the real identities.

STEP 4

Safe to produce and share

Opposing counsel, regulators, case law databases, RAG systems, BPO teams. Built for discovery production workflows. For review-platform export formats, contact our team.

§ 04 · Three pipelines

Every document format in litigation. Covered.

Legal document repositories mix scanned paper records from the 1990s, native digital contracts from last week, and DOCX briefs that combine text with embedded exhibits. Re-Doc has a purpose-built pipeline for each.

BEST FOR SCANNED DOCUMENTS

Redaction pipeline

True redaction: removes both image and text layers

For scanned court exhibits and image-based PDFs. Visual processing reads the image layer and extracts all text, including handwriting and fax artifacts. A context-aware model identifies PII in context. Black boxes are drawn precisely at the pixel level over both the image and the text data stream. The copy-paste attack does not work on Re-Doc redactions.

  • FOIA responses: scanned government records
  • Court filings with scanned exhibits and affidavits
  • Scanned medical records in personal injury litigation
  • Old paper documents digitized for discovery production
Processing: Enterprise-grade visual processing pipeline with automatic fallback. Handles scanned exhibits of any quality or DPI.
BEST FOR NATIVE DOCUMENTS
Recommended for discovery production

Text anonymization pipeline

Synthetic data swap. Document stays searchable and usable.

For native digital documents. All PII is replaced with consistent synthetic data. John Michael Smith becomes Robert James Wilson on every page of every document in the production set. The same mapping applied across the entire batch. Case facts, dates, contract terms, and legal arguments are fully preserved. Opposing counsel receives a complete, coherent record.

  • Discovery productions with large document sets
  • Case law databases: full-text search preserved
  • RAG / AI legal research systems: no degraded LLM output
  • BPO and claims processing teams: no blacked-out data
  • DSARs and M&A second requests on tight deadlines
Consistent replacement: John Smith maps to Robert Wilson on page 1, 47, and 301, and across all documents in a batch production.
Native PDF · DOCX · TXT
BEST FOR MIXED TEXT + IMAGE FILES

Multiple data types

Text replaced and images redacted in a single pass

For DOCX documents that mix narrative text with embedded exhibits, scanned attachments, and signatures. Re-Doc replaces PII in the text with consistent synthetic data and redacts sensitive content inside the images (such as logos, stamps and signatures) at the same time, then rebuilds the document with its layout and Bates stamps intact.

  • Briefs and motions with embedded exhibit images
  • Contracts with signature blocks and scanned annexes
  • Investigation reports mixing typed notes and scans
  • Demand packages with photos and typed narrative
Single pass: Text and embedded images are handled together, so a mixed DOCX comes back fully de-identified with its original layout intact.
Use casesPage 4 of 9
§ 05 · Use cases

Three workflows where Re-Doc replaces manual redaction

These are the actual high-volume workflows where legal operations teams spend the most on manual de-identification.

Use case 01 · Legal / eDiscoveryFRCP 34 · Bates Stamps Preserved · API Batch

eDiscovery production

A typical commercial litigation production involves 30,000-100,000 documents. Each document must be reviewed for PII, redacted or de-identified, and produced with Bates stamps intact. Re-Doc processes the production through the API in batches, keeps Bates stamps and exhibit markers in place, and leaves your team to review and approve the output instead of redacting every page by hand.

APIBatch processing for large productions
LayoutBates stamps, exhibit markers and formatting kept
FRCP 34Built for production workflows
  • Available via REST API for batch processing workflows
  • Bates stamps and exhibit numbering untouched
  • Consistent synthetic identities across the whole production
  • Review-platform export formats: contact our team
Use case 02 · Legal / eDiscoveryFull-text Search · RAG / AI Ready

Case law database and precedent analysis

Legal research databases need to be searchable. A case law repository built on black-box redacted court filings cannot surface “all cases involving a defendant from Dallas County”. The data is literally hidden behind paint. Research cited below found language models handle redacted text noticeably worse than clean text, because [REDACTED] tokens break the narrative. Text anonymization preserves the full factual and legal narrative. The parties are synthetic. The precedent is intact. Every document remains fully indexed and searchable.

  • Legal narrative and holdings fully preserved
  • Full-text search and vector embeddings work correctly
  • AI research tools get coherent input, not redaction holes
Use case 03 · Legal / eDiscoveryHigh Volume · Tight Deadlines · GDPR / CCPA

DSARs and M&A second requests

Data Subject Access Requests under GDPR and CCPA require third-party PII to be redacted before the responsive records are provided to the requestor. M&A second requests from the DOJ or FTC can require very large volumes of company documents on tight timelines. In both cases, every document must have third-party personal information removed before production while the substantive business content is preserved. Manual review at those volumes breaks down at exactly the moment it cannot afford to. Re-Doc processes the batch through the API, de-identifies third-party PII, and returns documents with the business information fully intact and readable.

APIBatch processing at volume
ConsistentSame identity mapping across the batch
ReviewYour team approves before release
What to redact in discoveryPage 5 of 9
Answer first

What information should be redacted in discovery?

Typically personal identifiers the rules protect, Social Security numbers, taxpayer IDs, birth dates, names of minors and financial account numbers, plus privileged content and anything a protective order covers. Check your court's rules for the exact list.

In US federal courts, FRCP 5.2 limits filings to the last four digits of SSNs, taxpayer IDs and financial account numbers, the year of birth, and a minor's initials. Many state courts have similar rules.

Productions often also need third-party personal data, medical information and confidential business details removed, depending on the protective order and privacy laws that apply.

Re-Doc finds these identifiers across the whole production and either redacts them permanently (text is removed from the file, not covered) or replaces them with consistent synthetic values when the document must stay readable.

FAQPage 6 of 9
FAQ

Common questions

01What is eDiscovery redaction?
It is removing privileged, confidential or personal information from documents before they are produced in litigation or investigations. Proper redaction deletes the underlying text, not just draws a box over it.
02Why do black-box redactions fail?
Many tools draw a black rectangle over the visible page while the text stays in the PDF, so it can be copied or searched. Re-Doc removes the text itself, or replaces it with synthetic data.
03Can Re-Doc keep Bates numbers and layout?
Yes. Output keeps the original format and layout, including headers, footers and Bates stamps.
04Can we redact evidence files that are scans or images?
Yes. Scanned PDFs and images are read with OCR and redacted at pixel level.
ComparisonPage 7 of 9
§ 06 · How Re-Doc compares

Built for legal documents. Not an afterthought.

Most tools in eDiscovery apply cosmetic black boxes that leave text data intact. Re-Doc was built specifically for discovery production workflows where documents must actually be safe to share.

#Typical tools in the marketRe-DocStatus
01Cosmetic black boxes. Underlying text copy-pasteable via PDF extract.True text removal. PII permanently gone from the file, not just visually covered.● Covered
02Breaks document readability for people and AI tools.Text anonymization preserves narrative structure. Documents stay searchable and usable.● Covered
03Manual, per-document process. Unusable for 40,000-document productions.Batch API processes whole discovery productions without per-document manual work.● Covered
04Cannot handle scanned exhibits and fax-originated deposition transcripts.Visual processing pipeline handles scanned exhibits, deposition transcripts, and legacy formats.● Covered
05No API access. Every document must be processed manually one at a time.REST API and batch upload plug into any existing eDiscovery workflow.● Covered
144%

higher LLM perplexity on masked text vs. clean baseline

Building an AI legal research tool or RAG system over case law? Black-box redaction degrades training and inference quality. Redaction tokens are noise. Synthetic replacement is signal. Source: arXiv 2411.05978 + Firstsource.

Baseline (clean text)1.16
Masked / redacted text2.83
Differential privacy noise4.87
SourcesPage 8 of 9
Get startedPage 9 of 9

Produce documents that are actually safe to share.

True redaction when text must be gone. Text anonymization when the document still needs to work. Multiple Data Types when one DOCX mixes both. Three pipelines, one platform.

Related: document anonymization software · what de-identification means · about Re-Doc · Insurance document anonymization · FOIA redaction software