re-doc
Try for free
Healthcare / HIPAAPage 1 of 11
Solutions · Healthcare / HIPAA

De-identification software
for PHI and medical records.

Re-Doc is de-identification and PHI redaction software. It finds PHI in medical records, scanned faxes, discharge summaries, clinical notes, and either redacts it permanently or replaces it with consistent synthetic data. It targets the HIPAA Safe Harbor identifier categories that appear in text. For custom requirements, such as imaging, contact our team. Diagnoses, medications and the clinical narrative stay intact, so the record is still useful for research, Release of Information and AI training.

1. Your documentclinical_note.txt
Clinical Note
MRN-004821 · Cardiology
PatientMaria Elena Gonzalez
MRNMRN-004821
DOB02/14/1965
AttendingDr. Rajesh Patel
FacilitySt. Luke’s Medical Chicago
Dx: Hypertension, Stage 2. Continue metoprolol 50mg QD. Follow up in 90 days.
5 pieces of sensitive data foundReal patient identity. Cannot share.
RE-DOC
2. Safe to shareclinical_note_safe.txt
Clinical Note
MRN-004821 · Cardiology
PatientSarah Ann Thompson
MRNMRN-007643
DOB09/03/1968
AttendingDr. Kevin Harmon
FacilityRiverside General Columbus
Dx: Hypertension, Stage 2. Continue metoprolol 50mg QD. Follow up in 90 days.
Replaced with synthetic dataSynthetic data. Designed for Safe Harbor. Shareable.

Safe Harbor identifiers found in the text replaced with consistent synthetic data. Clinical content preserved. Processed to support HIPAA Safe Harbor de-identification requirements under 45 CFR §164.514(b)(2).

$7.42MAverage healthcare breach cost, highest of any industry (IBM 2025)
18HIPAA Safe Harbor identifier categories, with options for custom requirements
OCRScanned faxes, charts and image exports read at pixel level
On requestOn-premise or custom deployment through an enterprise agreement
The problemPage 2 of 11
§ 02 · The problem

Four ways healthcare document handling fails teams every day

$7.42 million per breach, the highest of any industry (IBM 2025). These are the documented failure modes behind that number, technical, operational, and clinical.

01 · TECHNICAL ROOT CAUSE

Burned-in PHI: faxes, medical images, handwritten notes

Medical imaging files often have patient name, DOB, and MRN burned directly into image pixels and not in a text layer. PDF editors cannot detect it. Standard NLP tools skip it entirely. It looks like part of the scan.

RE-DOC
Per DICOM Standard PS3.15 Annex E and IHE Radiology guidance, PHI embedded as human-readable text in pixel data requires pixel-level destruction and not a text overlay. Re-Doc’s visual processing reads the pixel layer, finds the PHI, and permanently replaces those pixels in the output file. Nothing left to copy. Nothing left to extract. Today this works on scanned PDFs and image exports (PNG, JPG). Native DICOM files are also supported: tested in a hospital imaging deployment and enabled on request.
02 · RESEARCH BLOCKER

“Patient ████ Age ████” is useless to a researcher

IRB-approved studies need readable patient data with identifiers removed and not clinical context destroyed. A discharge summary with every detail blacked out tells researchers nothing. The treatment course, medication names, diagnostic codes, and physician assessment are what matter. Black-box redaction removes identity and utility together.

RE-DOC
Text anonymization keeps the clinical narrative intact. Only the patient identity changes, not the medical facts.Diagnoses, medications, and dosages pass through unchanged. Removing the 18 Safe Harbor identifiers is one part of HIPAA de-identification; the other is having no actual knowledge that what remains could identify the patient.
03 · MANUAL REVIEW GAP

Manual review misses context-dependent PHI

The physician’s name in a narrative note. The MRN embedded in a table footer. The date buried in a page header. Pattern-matching tools catch structured fields, they miss the identifiers woven into clinical prose.

RE-DOC
Re-Doc’s context-aware model reads the entire clinical document and not just pattern-matched field labels. It understands that “Dr. Patel ordered” in a narrative is a physician identifier, not just a proper noun. The same entity is caught on every page, in every form it appears.
04 · COMPLIANCE GAP

BAAs cover the legal relationship. Not the technical quality.

A Business Associate Agreement defines who is responsible. It does not verify that PHI was actually removed from the document before it was shared. Unauthorized disclosure incidents in the HHS breach portal regularly involve records shared with vendors after incomplete de-identification.

RE-DOC
A BAA governs the relationship. Removing PHI from the document itself reduces what can be exposed if a file is shared or leaked.
How it worksPage 3 of 11
§ 03 · Where Re-Doc fits

Plugs into your existing HIM workflow

Re-Doc sits between your EHR export and distribution. Upload via API or drag-and-drop. The clinical narrative stays intact. Processing logs provide an entity-level audit trail per document.

STEP 1

Source records

EHR exports, plain-text clinical notes (.txt), scanned faxes, discharge summaries, operative notes. Any format your clinical workflows produce.

STEP 2

Re-Doc processes

A context-aware model reads the entities across the HIPAA identifier categories found in text and replaces each with a consistent synthetic equivalent matched to type and format.

STEP 3

Synthetic output

Same document structure, same clinical narrative, same layout. Patient identity replaced. Diagnoses, medications, and treatment notes preserved exactly.

STEP 4

Share anywhere

Send to research teams, push to AI training pipelines, or share with auditors. Processed to support Safe Harbor and Expert Determination compliance strategies.

§ 04 · Three pipelines

Three pipelines. Pick the right one.

Healthcare documents come in different forms: scanned images, native digital files, and mixed files that combine text with embedded images. The correct de-identification approach depends entirely on which one you have.

BEST FOR SCANNED DOCUMENTS

Redaction pipeline

True pixel-level destruction, no text layer to extract

Visual processing reads the scanned document pixel-by-pixel, finds PHI by region, and burns permanent black boxes over those areas in the output PDF. The original pixel data is destroyed and not covered. No text layer exists to extract from a scanned document, because scanned documents are images.

  • Patient record requests for scanned paper charts
  • Incoming faxed referrals and prior authorizations
  • Legal hold documents from physical archives
  • Medical image exports (PNG, JPG) with PHI burned into the pixels
  • Handwritten clinical notes and intake forms
Scanned PDF · Images · DICOM (on request) · Fax · Handwritten Notes
BEST FOR NATIVE DOCUMENTS
Recommended for research and ROI

Text anonymization pipeline

Synthetic data swap. Clinical narrative stays usable.

Finds every PHI entity in a native PDF, DOCX or TXT file and replaces it with demographically consistent synthetic data. “Maria Elena Gonzalez” becomes “Sarah Ann Thompson” consistently across every page, every reference. Clinical content stays untouched: diagnoses, medications, dosages, treatment timelines.

  • Clinical trial CSR anonymization (EMA Policy 0070)
  • IRB-approved research data sharing (Safe Harbor method)
  • Release of Information processing at scale
  • AI and LLM training dataset preparation
  • Vendor and business associate data sharing
Consistent synthetic identities: Maria Gonzalez maps to Sarah Thompson on page 1, 47 and 301 of the same document.
Native PDF · DOCX · TXT clinical notes
BEST FOR MIXED TEXT + IMAGE FILES

Multiple data types

Text replaced and images redacted in a single pass

For DOCX files that mix narrative text with embedded images, scans, and signatures. Re-Doc replaces PHI in the text with consistent synthetic data and redacts sensitive content inside the images (such as logos, stamps and signatures) at the same time, then rebuilds the document with its layout intact.

  • Discharge summaries with embedded scan images
  • Lab reports with charts and patient photos
  • Referral packets mixing typed notes and scans
  • Consent forms with signatures and stamps
  • Care plans with embedded diagrams
DOCX · Embedded Images · Signatures · Logos
Use casesPage 4 of 11
§ 05 · Use cases

Three workflows where Re-Doc replaces manual de-identification

These are the high-volume, compliance-critical workflows where black-box redaction and manual review consistently fall short.

Use case 01 · Healthcare / HIPAAEMA Policy 0070 · Health Canada PRCI

Clinical trial CSR anonymization (EMA Policy 0070, Health Canada PRCI), see clinical trials

Clinical study reports published under EMA Policy 0070 or Health Canada PRCI must have trial participants’ personal data anonymized. Re-Doc replaces participant identifiers across narratives, listings and appendices with consistent synthetic data, so the report stays readable, and hands your disclosure team a reviewed draft instead of a blank page. It supports your anonymization workflow; the anonymization report and risk assessment stay with your team.

  • Consistent synthetic participant IDs across a report
  • Narratives and tables stay readable
  • Human review before submission
Use case 02 · Healthcare / HIPAAHIPAA Safe Harbor · Safe Harbor identifiers · 45 CFR §164.514(b)(2)

Medical research data sharing (IRB / Safe Harbor)

IRB-approved studies require HIPAA Safe Harbor de-identification: removal of all 18 identifier categories before sharing with researchers. Traditional approaches use expert determination (expensive, slow) or manual review (error-prone, misses contextual PHI). Re-Doc applies LLM-based entity detection across the Safe Harbor categories found in text simultaneously, preserving the clinical narrative researchers actually need: diagnoses, lab values, medication histories, and treatment responses. Processed to support Safe Harbor requirements without destroying study utility.

  • Safe Harbor identifier categories in text detected simultaneously
  • Clinical narrative and lab values fully preserved
  • Consistent synthetic identity across multi-page charts
Use case 03 · Healthcare / HIPAA30-day HIPAA deadline · HIM automation

Release of information processing at scale

HIM departments process large volumes of patient record requests, each requiring de-identification of third-party PHI before release. The 30-day HIPAA response window is strict. The manual de-identification step is the bottleneck: a clinician reviewing every redaction placement on every page, per request. Re-Doc processes each request through the API in minutes. Health Information Management teams upload the chart, receive a de-identified output, and fulfill the request on deadline, without a physician reviewing every black box placement.

  • Batch API processes multiple requests in parallel
  • Per-document audit trail for HIPAA minimum necessary
  • Third-party PHI removed while patient clinical data preserved
OCR and redaction for medical recordsPage 5 of 11
Scanned records

OCR and redaction for medical records, in one step

Re-Doc reads scanned charts, faxed referrals and image exports with OCR, finds the PHI, and replaces those pixels in the output. No separate OCR tool, no hidden text layer left behind.

Most medical records that need redaction were never digital: faxed prior authorizations, scanned paper charts, intake forms, and images with patient details burned in. Text-only de-identification tools cannot see any of that. Re-Doc treats the page as an image first, so the same pipeline handles typed, scanned and handwritten records.

For native files, EHR exports, DOCX letters, TXT clinical notes, Re-Doc reads the text directly and can replace PHI with synthetic data instead of blacking it out.

On-premise PHI redactionPage 6 of 11
Deployment

On-premise PHI redaction for hospitals and health-data teams

Some records cannot leave your network. On request, Re-Doc can run on your own servers or private cloud, with the same redaction and synthetic data pipelines as the hosted product.

Hospitals, payers and research groups often need de-identification to happen inside their own environment for policy or contractual reasons. An on-premise deployment keeps documents, detections and outputs on your infrastructure. It is arranged through a separate enterprise discussion and can be built as a custom solution around your systems and volumes. Contact us to start that conversation.

Safe Harbor vs Expert DeterminationPage 7 of 11
HIPAA methods

Safe Harbor vs Expert Determination

HIPAA (45 CFR §164.514) allows two ways to de-identify health information. Re-Doc supports the document work behind both.

Safe Harbor removes the 18 listed identifier categories, names, geographic details smaller than a state, dates other than year, phone and fax numbers, email addresses, SSNs, MRNs, account numbers and more, and requires that you have no actual knowledge that what remains could identify the person.

Expert Determination relies on a qualified expert who applies statistical and scientific methods and documents that the risk of re-identification is very small. It allows more data to be kept, such as some dates or locations, when the expert judges it safe.

Re-Doc detects and replaces or redacts identifiers across unstructured documents, which is the manual bottleneck in both methods. Which method applies, and the final sign-off, stays with your privacy team. Source: HHS guidance on HIPAA de-identification.

FAQPage 8 of 11
FAQ

PHI redaction and de-identification: common questions

01Is Re-Doc HIPAA certified, and which identifiers does it find?
There is no official HIPAA certification for software; HHS does not certify products. Re-Doc does not hold SOC 2 or ISO 27001 certification yet; work toward security certification is in progress. If your policy requires data to stay in-house, Re-Doc can be deployed on-premise on request, so documents never leave your infrastructure. Re-Doc detects the HIPAA Safe Harbor identifier categories that appear as text (names, dates, addresses, phone and fax numbers, emails, SSNs, record and account numbers, IDs, URLs and IP addresses), in native files and scanned documents, and you review every change before you download the result. For full-face photographs and biometric identifiers, contact our team to discuss image workflows.
02What is PHI redaction software?
PHI redaction software finds protected health information in medical records (names, dates, MRNs, addresses, phone numbers and the other HIPAA identifiers) and removes it before the record is shared. Re-Doc can either black out PHI permanently or replace it with consistent synthetic data so the record stays readable.
03Which tool supports OCR and redaction for medical records?
Re-Doc does both in one step. Scanned charts, faxed referrals and image exports are read with OCR at pixel level, PHI is located, and the pixels are replaced in the output file, so there is no hidden text layer left to copy.
04Can Re-Doc run on-premise?
Yes, on request. On-premise and private-cloud deployments are set up through a separate enterprise discussion, and can be tailored into a custom solution for your environment. The hosted version is available straight away for teams that do not need this.
05Does removing the 18 Safe Harbor identifiers make a record de-identified?
Under HIPAA Safe Harbor you must remove the 18 identifier categories and also have no actual knowledge that the remaining information could identify the person. The alternative is Expert Determination, where a qualified expert certifies the re-identification risk is very small. Re-Doc handles the removal and replacement of identifiers found in the text. For other requirements, such as identifiers in images, contact our team to discuss options. Your privacy team decides which method applies.
06Does Re-Doc support DICOM files?
Scanned PDFs and image exports (PNG, JPG) are supported in the standard product. Native DICOM files have been tested in a hospital imaging deployment and are enabled on request.
07What is the difference between redaction and synthetic data replacement?
Redaction blacks out PHI, which is right for scans and records being released. Synthetic data replacement swaps each identifier for a realistic, consistent substitute, so research, analytics and AI training teams can still read the clinical narrative.
ComparisonPage 9 of 11
§ 06 · How Re-Doc compares

Built for clinical documents. Not data tables.

Most de-identification tools are built for structured database exports. Re-Doc handles unstructured documents: scanned faxes, discharge summaries, narrative notes. That is where PHI actually lives.

#Typical tools in the marketRe-DocStatus
01Structured data only. Cannot open a PDF, DOCX, or scanned clinical document.Processes PDFs, DOCX, TXT clinical notes, scanned faxes, and image-based medical records.● Covered
02Redaction removes text. Clinical narrative breaks down for downstream use.Text anonymization replaces PHI with synthetic equivalents. Context preserved.● Covered
03No scanned document support. Misses fax-originated records and DICOM pixel PHI.Visual processing handles scanned faxes, image exports and burned-in pixel PHI. Native DICOM enabled on request.● Covered
04Manual, per-file processing. Unusable for high-volume ROI and research workflows.Batch API processes hundreds of authorization requests in parallel. Audit trail included.● Covered
05No audit trail aligned with HIPAA minimum necessary standard.Processing logs per document with entity-level detection records for compliance review.● Covered
$7.42M

Average healthcare breach cost, the highest of any industry (IBM 2025)

A BAA governs the relationship. Removing PHI from the document itself reduces what can be exposed if a file is shared or leaked.

Text

Safe Harbor identifier categories detected in text and scans

3

Pipelines: redaction, synthetic data, mixed files

API

Batch processing with per-document logs

Custom

On-premise or custom deployment on request

SourcesPage 10 of 11
Get startedPage 11 of 11

Stop choosing between compliance and usability.

Redaction when you need permanent pixel destruction. Text anonymization when the document still needs to work. Multiple Data Types when one DOCX mixes both. Three pipelines, one platform.

Related: document anonymization software · what de-identification means · about Re-Doc · Clinical trial document anonymization · Insurance document anonymization