When drug discovery teams plan a program timeline, cohort data sourcing typically appears as a line item labeled "data acquisition" with an estimate of four to eight weeks. In a significant share of programs, the actual time exceeds that estimate by a factor of three or four. The gap between planned and actual sourcing timelines is one of the most consistent — and consistently underestimated — friction points in translational research operations.
The numbers below are drawn from conversations with research operations professionals across biopharma and academic medical contexts, not from a formal structured survey. We're sharing ranges and patterns, not sample-size-backed statistics. The purpose is to give research teams a more realistic baseline for planning, and to help program leads understand which sourcing scenarios carry the most timeline risk.
The Baseline: What Typical Timelines Look Like by Data Type
Cohort sourcing time varies significantly by data type. Broadly, three categories carry different default timelines when sourcing from established biobank collections:
De-identified clinical phenotype data from collections with broad consent and existing data use agreement templates: typically 6 to 12 weeks from initial contact to data delivery. The dominant time variable is legal review on both sides — the biobank's data governance office and the requesting organization's legal team. When both sides have experienced data licensing operations and a template DUA already exists, timelines can compress to the lower end. When either side is working through a first-time transaction type, timelines routinely extend to 14–18 weeks.
Linked genomic and clinical data (e.g., genotyping array data or sequencing calls linked to EHR-derived phenotypes): typically 14 to 22 weeks. The additional time reflects two factors: genomic data processing timelines for collections where the data isn't pre-processed and formatted for delivery, and heightened consent review requirements at some institutions for genetic data sharing under state-level genetic privacy regulations, which vary across the US and create inconsistent review requirements depending on participant residence.
Multi-site cohorts requiring data harmonization across two or more source institutions: typically 20 to 36 weeks. Multi-site sourcing adds a coordination layer — multiple DUAs, potentially multiple IRB notifications, and the harmonization work required to align clinical data collected under different EHR systems and coding conventions. This category shows the widest variance; programs that planned multi-site sourcing carefully with institutional relationships in place sometimes complete in 16–18 weeks, while programs encountering governance friction at one or more sites can extend beyond 40 weeks.
Therapeutic Area Variation
Across therapeutic areas, oncology and cardiometabolic programs tend to show the tightest sourcing timelines, relative to the complexity of the data requested. This reflects the maturity of biobanking infrastructure in these areas — several decades of oncology biorepository investment means that many institutions have more streamlined governance for oncology data requests, existing template agreements, and larger data governance teams with oncology-specific expertise.
Neurology and neurodegeneration programs face longer timelines relative to the data complexity, in part because the phenotypic characterization requirements are more demanding. A cardiometabolic cohort can often be characterized using a relatively compact set of clinical variables — blood pressure, lipid panels, BMI, ICD codes for relevant diagnoses. A neurological phenotype cohort for a neurodegenerative disease program often requires cognitive assessment scores, imaging-derived biomarkers, and longitudinal motor or cognitive trajectory data that may have been collected inconsistently across sites or enrollment periods. Finding a collection with adequate phenotypic depth for a neurology program, and then navigating the sourcing process, consistently runs longer than the equivalent process for metabolic disease programs.
Rare disease programs represent the longest tail. When the required cohort size is limited by disease prevalence to begin with, and the phenotypic characterization requirements are specific, the sourcing process often requires outreach to multiple institutions sequentially rather than a single data access request. Programs that eventually succeed in assembling a rare disease cohort from biobank sources typically required 12 to 24 months of active sourcing effort — with timelines heavily dependent on the pre-existing network relationships of the research team.
Where the Time Actually Goes
Decomposing the sourcing timeline reveals that data access negotiations and legal review — not data preparation or delivery — account for the majority of elapsed time in most sourcing processes. In the programs we hear about, the pattern is roughly consistent: initial feasibility inquiry and informal discussion with a biobank data manager, 1–3 weeks; formal data access application and institutional review, 4–8 weeks; legal negotiation and DUA execution, 4–12 weeks; data preparation and quality review, 2–6 weeks; delivery and onboarding, 1–3 weeks.
The legal negotiation window is the most variable and the least predictable. It's also the window that is most often excluded from initial timeline estimates, because research teams frequently assume that if the scientific case for the data request is strong, the legal process will be relatively straightforward. In practice, legal review timelines are driven by institutional workload, the novelty of the transaction type, and the specific consent language in the collection — none of which are visible to the requesting team when they submit their initial access application.
What These Benchmarks Mean for Planning
We're not suggesting that all data sourcing timelines are fixable, or that program teams should simply adjust their expectations to accept 6-month delays as the new normal. The structural reasons for these timelines are real, and some of them — IRB governance, legal review, institutional DUA negotiation — represent necessary oversight that shouldn't be bypassed.
What is tractable is reducing the proportion of sourcing time spent on information-gathering that should have been available at the outset. When a research team can evaluate a dataset's consent scope, de-identification documentation, and phenotypic completeness before submitting a formal data access request, they can make a more informed decision about which collection to pursue — and they enter the legal review process with a clearer picture of what they're actually requesting. That doesn't compress the legal timeline dramatically, but it does eliminate the scenarios where a team invests two months in a sourcing process only to discover at the DUA negotiation stage that the consent language doesn't support the intended use.
The programs with the best sourcing track records tend to share one operational habit: they start sourcing earlier than feels necessary, they invest time in characterizing a dataset's fitness for purpose before initiating a formal request, and they maintain institutional relationships with biobank governance offices rather than treating each request as a cold transaction. None of that requires a platform — it requires the operational discipline to treat data sourcing as a long-lead-time activity that deserves the same proactive management as any other critical path item in a drug discovery program.