Somewhere in most data room reviews for a biobank data licensing transaction, a lawyer opens a folder of consent forms and asks a question that should have been answered months earlier: do these consent documents actually authorize this use? The answer is frequently ambiguous — and that ambiguity is the last-mile barrier that stalls more licensing deals than any technical problem.
For biopharma research teams, the word "consented" has become a shorthand that obscures a significant amount of variation. When a dataset is described as consented, the question worth asking is: consented for what, under what framework, by whom, and with what documentation chain? The answers matter enormously in practice.
Consent Models in Biobank Research: The Spectrum
Biobank consent exists on a spectrum from very narrow to very broad, and the position of any given collection on that spectrum determines what downstream uses are permissible — and what legal review is required before a research team can access the data.
At the narrow end, study-specific consent authorizes a participant's contribution for a defined study with a fixed research question. Data collected under this model is not available for secondary use without additional consent from the original participant pool — a process that is slow, expensive, and in many cases simply not possible if the collection is years old and participants are no longer reachable. Study-specific consent was the default for much of the clinical research conducted before the early 2010s, which means a significant portion of historically rich biobank collections carry consent limitations that reflect assumptions about research use that no longer hold.
Broad consent — the model formalized in the revised Common Rule (45 CFR 46) in the US — permits secondary research use across a wide range of future studies, subject to IRB oversight. Under a well-drafted broad consent framework, a participant agrees to contribute their samples and data for "future research that may not be known at this time," with certain categorical limitations and the right to withdraw. This is the model that makes modern biobank data licensing operationally viable, but it only covers collections that were built under it.
Tiered consent is a middle ground that emerged from European biobank practice and has been adopted in some US collections. Participants indicate their willingness to participate in different categories of research — academic vs. commercial, genetic research, sharing with international repositories — through a structured opt-in matrix. This model is scientifically and ethically attractive because it gives participants meaningful control. It creates a more complex data curation challenge because each record carries different permission tiers, and any downstream dataset must be constructed with awareness of which participants consented to which use categories.
What IRB Documentation Actually Means in a Licensing Context
IRB oversight exists at the point of data collection and at the point of secondary use. For biobank data licensing, the relevant question is not just whether the original collection had IRB approval — almost all did — but whether the secondary use of that data for a specific research purpose requires its own IRB review and, if so, what documentation the data buyer needs to produce or receive.
Under the Common Rule, secondary research use of de-identified data is generally exempt from IRB review, because de-identified data is not considered human subjects research. This is the regulatory logic that makes large-scale biobank data licensing possible: once data is properly de-identified under a recognized standard, the research program using it doesn't need to convene an IRB for each analysis.
But that exemption depends on the de-identification being valid and documented. If a dataset arrives with an assertion that it has been de-identified but no audit trail showing who performed the de-identification, under what method, and when, a careful institutional IRB office may require an independent review before approving the secondary use. For academic research teams, this adds weeks to the process. For biopharma programs subject to additional regulatory scrutiny, it can create a documentation gap that surfaces later in a regulatory submission context.
We're not saying IRB oversight is the problem here — it isn't. Institutional oversight of human subjects research is precisely what makes the data trustworthy in the first place. The issue is that the documentation chain supporting IRB-exempt status is often incomplete or fragmented, requiring the data buyer to reconstruct it from scratch rather than receiving it as part of a properly packaged dataset.
Re-consent Protocols: When They're Required and What They Cost
When a research team identifies a collection with the scientific characteristics they need but consent documentation that doesn't clearly cover the intended use, re-consent is sometimes proposed as a solution. In practice, re-consent protocols deserve careful scrutiny before being treated as a reliable path.
Re-consent is operationally expensive. A collection of several thousand participants, with addresses that may be years out of date, response rates that are typically well below 50% for secondary-use re-consent requests, and a process that requires IRB approval in its own right, can cost substantially more calendar time and administrative effort than originally scoped. For time-sensitive research programs, the re-consent timeline frequently exceeds the project decision window.
Consider a scenario that reflects patterns reported in research operations contexts: a program evaluating a biomarker panel for an early oncology indication identifies a prospectively collected cohort with exceptionally well-characterized tumor histology and matched normal samples. The consent language, drafted in 2011, includes commercial use restrictions that were standard for institutional biorepositories at that time. Re-consent is proposed as a path forward. Fourteen months later, after two rounds of IRB review for the re-consent protocol itself, and a participant response rate that yielded a re-consented cohort roughly 40% the size of the original, the dataset is usable — but the program's target validation timeline has long since passed.
This is not an argument against re-consent as a practice. There are cases where it is the right and only ethical path. It is an argument for front-loading the consent scope assessment before any other due diligence effort begins. If a collection's consent limitations are identified before a program invests months of relationship-building with a biobank, the decision to pursue an alternative collection can be made at a point when it doesn't derail the overall timeline.
What Consent Documentation Needs to Contain for Licensing Transactions
When we work with biobank collections to prepare datasets for licensing, the consent documentation package we aim to produce for each collection includes several specific elements that data room reviewers routinely request. Understanding what that package looks like helps research teams ask the right questions earlier in the process.
The primary consent form used at the time of enrollment, including any amendments if the consent language was updated during the collection period, is the foundation. This document needs to be present not just as a reference but as a versioned record tied to the enrollment dates of the participants in the dataset. A collection enrolled over eight years may have three different versions of the consent form, and the distribution of participants across those versions matters for determining what uses are permissible.
An IRB approval record for the original study, along with any continuing review approvals, establishes the oversight context. This is often available in institutional records but may require active retrieval if the study completed years ago and files have been archived.
A consent scope summary — a plain-language characterization of what the consent language does and does not authorize, prepared or reviewed by legal counsel — is the document that typically saves the most time in data room review. Rather than requiring the buyer's legal team to interpret the consent language independently, a clear scope summary moves the review to a verification exercise rather than an interpretive one.
Finally, documentation of the de-identification method applied to the dataset, cross-referenced to the consent provisions governing data sharing, closes the loop between the consent record and the data itself.
The Last Mile Is a Documentation Problem, Not a Data Problem
The scientific quality of biobank collections in the US and globally is genuinely excellent in many cases. The phenotypic depth, genomic coverage, and longitudinal follow-up available in well-maintained collections represent decades of participant contributions and institutional investment that can't be replicated quickly.
What consistently fails research teams is not the underlying data — it's the documentation layer that should make the data transactable. Consent forms that were never indexed to participant records at the time of enrollment. IRB approvals filed in institutional archives without accessible metadata. De-identification records that document the method used but don't attest to the output. These are fixable infrastructure problems, and fixing them upstream — before any specific data request — is the only way to break the pattern of last-mile documentation delays that eats months from programs that planned their timelines in weeks.