We started Network Bio in 2021 with a straightforward observation: the data that biopharma and academic research teams most need — consented, phenotypically characterized, longitudinally followed biobank collections — was sitting in repositories that weren't organized for the way research programs actually work. Today, we're sharing that we've closed an angel investment of $2 million to continue building the infrastructure layer that changes that.
This is a note about what the investment means for our work, what we're building toward, and why we think the problem we're working on matters enough to dedicate the next several years to it.
What We're Building — and Why It's Taken This Long
Network Bio is not a genomics tool, a clinical trial platform, or a patient records company. We are a data infrastructure company in the narrower sense: we work with biobank collections to curate, document, and package their data into a form that research teams can actually license and use without rebuilding the entire data readiness stack from scratch on their end.
The problem we're solving is unglamorous. It involves consent documentation, de-identification audit trails, phenotype harmonization across mismatched coding systems, and data format standardization. None of those tasks have a clean technical solution that eliminates the underlying institutional and organizational complexity. What we've built over the past four years is a curation methodology and tooling stack that handles the labor-intensive parts of this process systematically, so that the time between a biobank collection and a research-ready licensed dataset is measured in weeks rather than the months it typically takes when the work is done ad hoc.
It's taken time to build because the problem requires domain credibility on both sides of the transaction. Biobanks need to trust that the data infrastructure partner they work with handles consent documentation and de-identification with rigor — the reputational and regulatory stakes are too high for anything else. Research teams need to trust that a licensed dataset comes with the provenance documentation their legal teams will require. Building both forms of trust takes time that can't be compressed by moving faster technically.
What This Investment Allows Us to Do
The $2 million angel investment will be directed toward three priorities that have been waiting on capacity rather than knowledge.
First, expanding our curation capacity. Our current team of four has been handling the full pipeline from collection assessment through dataset delivery. That's a bottleneck we've been feeling acutely. The investment will let us bring on additional data scientists and curation specialists with biomedical domain depth — the kind of people who understand both the scientific requirements of a research-grade dataset and the regulatory framework that governs how that dataset can be used.
Second, deepening our de-identification pipeline. The current tooling handles Safe Harbor de-identification well. We've been working toward more robust Expert Determination capabilities — the statistical methodology, the documentation workflow, and the quality review process that allows Expert Determination analyses to be produced efficiently and defensibly across a range of collection types. This is important for collections where date precision and temporal granularity are scientifically critical, and it's an area where we want to be meaningfully better than the current state of the art in biobank data preparation.
Third, growing our dataset catalog. We currently have collections in three therapeutic areas — oncology, cardiometabolic, and neurology — with varying depth and follow-up across each. The investment will support the relationship and curation work required to add collections in additional therapeutic areas and to deepen the longitudinal coverage of existing collections. More datasets with better documentation means research teams have more options to evaluate before committing to a data access request, and it means more programs can find a fit within our catalog rather than having to start from scratch with a cold biobank outreach.
Why Angel Investment Is the Right Stage for This Work
We're sometimes asked why we haven't sought institutional funding at this stage. The honest answer is that the problem we're working on requires a longer proof-of-value arc than the timeline pressures of larger funding rounds typically accommodate.
Building genuine trust with biobank collections — the kind of trust that results in ongoing curation partnerships rather than one-time transactions — takes multiple years of consistent, careful execution. Demonstrating to research teams that our licensed datasets hold up through legal review and deliver the scientific utility they were selected for takes time to accumulate. The angel funding round gives us the resources to accelerate our core work without creating incentives to scale faster than the trust infrastructure that makes the work valuable can support.
We're building for the long arc. Biobank data licensing is not a problem that gets solved in a two-year product cycle, and we have no interest in pretending otherwise. The research programs that depend on cohort data are long-cycle programs — drug discovery timelines measure in years, not quarters — and the infrastructure they depend on needs to be built with that same orientation toward sustained, defensible quality.
A Note on What We're Not Changing
Investment announcements sometimes come with declarations of new directions or expanded scope. This isn't one of those. We are a Boston-based team doing a specific, difficult thing in biomedical data infrastructure, and the investment is allowing us to do that thing better and at greater scale — not to pivot toward an adjacent market or add product lines.
Our data standards approach — de-identification designed with HIPAA Safe Harbor and Expert Determination requirements, consent scope documentation, phenotype completeness characterization, format standardization — is not changing. Our orientation toward working only with collections where the consent architecture genuinely supports the intended use is not changing. What is changing is our capacity to execute on these standards across a larger collection portfolio and for a larger number of research programs.
If you're a research operations team with an upcoming cohort sourcing need, or a biobank interested in exploring what a curation partnership looks like, we'd like to hear from you. The problem we're working on is real, the work is meaningful, and we now have more capacity to take it on. Reach us at [email protected].