Data Standards

How every dataset earns the research-grade designation.

Research-grade means documentable, repeatable, and legally defensible at data room review. Below is the precise methodology Network Bio applies to every dataset in the catalog — from Safe Harbor de-identification through QC metrics and consent audit trail. No dataset enters the catalog without passing every stage.

De-identification

Safe Harbor de-identification — 18 identifiers removed.

Network Bio datasets are built with HIPAA Safe Harbor de-identification requirements in mind. All 18 categories of direct identifiers are systematically removed or generalized before any data enters our catalog. We do not claim HIPAA certification, but our process is designed to conform to Safe Harbor method standards.

01

Names and geographic identifiers

All personal names, geographic subdivisions smaller than state, and zip codes generalized or removed per Safe Harbor specification.

02

Dates and age generalization

All dates associated with an individual (except year) are shifted or removed. Ages above 89 are aggregated into a single 90+ category.

03

Contact and account identifiers

Phone numbers, fax, email, Social Security numbers, medical record numbers, health plan beneficiary numbers, and account numbers are removed.

04

Device IDs and biometric data

Certificate/license numbers, device identifiers, URLs, IP addresses, biometric identifiers, and full-face photographic images are removed or replaced.

Consent Model

All biobank collections in our catalog operate under a broad consent framework — meaning participants have provided consent for secondary research use of their de-identified samples and data. We maintain consent audit trails that document the consent version, date, and scope for every subject in a cohort.

Broad Consent Framework

Participants consent to secondary research use of their de-identified data. Consent scope is documented per cohort and included in the delivery package.

Re-Consent Protocols

Where biobank partners have re-consent programs, updated consent records are logged and maintained. Consent version history is available for subjects in longitudinal cohorts.

Consent Audit Trail

Every licensed delivery includes a consent audit trail document covering consent type, dates, scope, and version — structured for IRB and legal review.

Quality Control

QC pipeline: every dataset measured before delivery.

Before any dataset is made available in our catalog, it passes through an automated QC pipeline that measures data quality across five dimensions. Results are included in the delivery documentation.

qc_pipeline_report.tsv — synthetic example
Metric Threshold Result Status
Missingness rate (clinical fields) < 15% 8.3% PASS
Sample call rate (genomic) ≥ 95% 97.1% PASS
Batch effect score (PCA-based) < 0.15 0.09 PASS
Outlier rate (clinical lab values) < 3% 2.7% PASS
Phenotype completeness (required fields) ≥ 90% 87.4% WARN
Longitudinal follow-up coverage ≥ 80% at 12mo 83.2% PASS

QC metrics above are synthetic examples only. Actual metrics are provided in the dataset delivery package for each licensed cohort.

Need the documentation package for a specific dataset?

Our team can walk you through the de-identification report, consent audit trail, and QC run log for any collection in the catalog — before you commit to a licensing engagement.