De-identification

De-identification with a manifest that never holds a value.

@cosyte/deid applies a de-identification policy, HIPAA Safe Harbor by default, to the fields you locate in a document and returns them transformed, with a manifest of everything it acted on. Clinical values are kept, and the manifest never holds a value.

@cosyte/deid
npm install @cosyte/deid

Need it integrated? Talk to us.

The standard

HIPAA defines two ways to de-identify health information.

  • Under Expert Determination, a person with appropriate knowledge of statistical and scientific methods determines that the risk of identifying an individual is very small, and documents the methods and results. Source: 45 CFR 164.514, 2025 edition (govinfo)
  • Under Safe Harbor, eighteen listed identifiers, lettered (A) names through (R) any other unique identifying number, characteristic or code, are removed for the individual and for relatives, employers and household members. Source: 45 CFR 164.514, 2025 edition (govinfo)
  • Safe Harbor also requires that the covered entity has no actual knowledge that what remains could identify the individual. Source: 45 CFR 164.514, 2025 edition (govinfo)
  • For imaging, DICOM Part 15, Annex E defines which attributes to remove or replace so a data set does not leak identifying information. Source: DICOM PS3.15, Annex E

@cosyte/deid

A policy engine that is conservative where a parser is lenient.

Identifiers are addressed by their place in the document structure, never found by pattern-matching raw bytes. Results are described as transformed per the configured policy, never as certified: the certification is always yours.

  • HIPAA Safe Harbor by default, covering all eighteen categories, the open-ended catch-all included.
  • Five transforms: redact, generalize (a date to its year, a ZIP code to three digits), keyed date shift, keyed HMAC pseudonym and keyed hash.
  • A value-free manifest: category, transform, location, count and disposition, never a value, a key or a date-shift offset.
  • A longitudinal registry keeps one person's surrogates consistent across documents.
  • Profiles can only tighten a policy, never loosen it.
  • An Expert Determination support report structures the manifest for a statistician and never renders a determination.
  • Zero third-party runtime dependencies: every primitive is node:crypto.
deidentify.tsts
import { deidentify, SAFE_HARBOR_CATEGORIES } from "@cosyte/deid";

// Synthetic inputs: placeholders in the shape of an HL7 v2 name and date of birth, not a person.
const name = "ZZFAMILY^ZZGIVEN";
const dob = "19900215";

const { document, manifest } = deidentify(
  {
    loci: [
      { path: "PID-5", kind: "identifier", category: SAFE_HARBOR_CATEGORIES.NAMES, value: name },
      { path: "PID-7", kind: "date", category: SAFE_HARBOR_CATEGORIES.DATES, value: dob },
      { path: "OBX-5", kind: "clinical", value: "5.4 mmol/L" },
    ],
  },
  {},
);

console.log(document.loci.map((locus) => locus.value));
console.log(manifest.map((entry) => `${entry.locus} ${entry.category} ${entry.disposition}`));

From the @cosyte/deid README on GitHub, verbatim. Every value in it is synthetic.

Limits

What it does not do.

Its README draws the honesty line itself:

  • It never certifies. Output is Safe-Harbor-transformed per the configured policy, and the actual-knowledge condition is yours.
  • Expert Determination is supported, never rendered: the determination and its documentation are the expert's.
  • Free text is blocked by default. Redacting it needs a redactor you supply.
  • DICOM support is metadata-only: burned-in pixel annotations are flagged, not cleaned.
  • NCPDP SCRIPT is refused outright rather than partly processed.

The full status and limits, in the README

Alternatives

What else you could use.

Each description comes from the project’s own documentation or listing, linked below.

Python library

Presidio

Open-source identification and anonymization modules for private entities in text and images, such as names, locations and phone numbers. It works on free text; @cosyte/deid works on structured documents and blocks free text by default.

Source: Presidio

Python library

deid (pydicom)

Best-effort de-identification for medical images in Python, built on pydicom.

Source: PyPI, deid

Desktop tool

RSNA DICOM Anonymizer

A tool from the Radiological Society of North America for processing DICOM files for research by removing or changing personal identifiers.

Source: RSNA, imaging research tools

Hosted service

Amazon Comprehend Medical DetectPHI

A hosted operation that detects protected health information in clinical text. Its documentation notes that the entities it detects do not map one to one to the Safe Harbor list.

Source: AWS documentation, DetectPHI

Works with

  • HL7 v2The @cosyte/deid/hl7 adapter works on a message parsed by @cosyte/hl7.
  • C-CDAThe @cosyte/deid/ccda adapter works on a document parsed by @cosyte/ccda.
  • FHIRThe @cosyte/deid/fhir adapter works on an R4 resource read by @cosyte/fhir.
  • X12The @cosyte/deid/x12 adapter works on an interchange parsed by @cosyte/x12.
  • NCPDPThe @cosyte/deid/ncpdp adapter works on a Telecom transaction parsed by @cosyte/ncpdp.
  • DICOMThe @cosyte/deid/dicom adapter works on a data set parsed by @cosyte/dicom.
  • @cosyte/synthPlants synthetic identifiers in generated data and checks that @cosyte/deid removes every one.

Try @cosyte/deid on your own messages.

It is free and MIT-licensed. If it saves you a day, a star on GitHub helps other engineers find it.

Need it integrated? Talk to us, or see what our integration services cover.