A researcher has built a prediction model on her institution's data. It performs well. Before it can mean anything, it needs external validation on a cohort she does not have.
She does the responsible thing. She searches PubMed for published cohorts with the right phenotype. She finds four promising papers. Each contains the sentence that has become the most quietly dishonest phrase in scientific publishing:
"Data are available from the corresponding author on reasonable request."
She writes four polite, specific, well-constructed emails.
She receives one reply, which is a courteous decline. The other three go unanswered.
Six months later, at a conference, she mentions the problem to someone in a hallway. That person says: "Oh, you should talk to Marcus, he has about 400 of those with five-year follow-up and biobanked samples."
Two weeks later she has the collaboration.
The formal channel produced nothing. The hallway produced everything. And that is not an anecdote about her bad luck. It is the documented behavior of the entire field.
The most-studied broken promise in science
The data-on-request statement has been examined rigorously, and the results are unambiguous.
A study in the Journal of Clinical Epidemiology examined 1,792 manuscripts whose authors had stated data were available on request. When the researchers actually requested the data:
- 93 percent did not respond or refused.
- 6.8 percent shared.
A companion analysis of 300 systematic reviews found that among those with data statements, only 13 percent had genuinely downloadable data, while 42 percent said "available upon request."
So the sentence that appears in an enormous share of the medical literature, and that satisfies journal policy, and that reviewers accept, has roughly a 7 percent chance of producing data.
This is well known among researchers and has produced a decade of policy response: journal mandates, funder requirements, and most recently the NIH Data Management and Sharing Policy. All of it aimed at making the formal channel work.
Meanwhile the informal channel works extremely well
Here is the finding that reframes the entire problem, and it deserves far more attention than it has received.
A 2025 study in Accountability in Research surveyed 285 cancer researchers:
- 45 percent had shared data from their latest paper.
- One third had shared privately, outside any formal mechanism.
- And among those who shared privately, 74 percent said it led to authorship or a future collaboration.
Read that alongside the 93 percent refusal rate and a coherent picture emerges.
Researchers are not unwilling to share. They are unwilling to share into a void.
A cold email from a stranger, requesting a dataset that took years and considerable funding to assemble, offering nothing, with no way to assess whether the requester is competent or will credit the work, is a request to hand over an asset to an unknown party for no return.
A request from someone known, or vouched for, that comes with an offer of collaboration, produces sharing three quarters of the time and generates a publication.
The variable is not policy. It is trust and reciprocity, and no mandate creates either.
Why federated networks do not solve it
The sophisticated response over the last decade has been federated research networks, and they are genuinely impressive infrastructure.
PCORnet covers data from more than 50 million people annually across eight clinical research networks, accessed through a centralized front door. TriNetX grew from 55 organizations in 2017 to more than 220 across 30 countries by 2022. N3C was built by more than 90 institutions and has supported publications carrying 1,589 authors.
These networks answer one question superbly: how many patients like X exist in the network?
That is a real and valuable question. It is not, however, the question the researcher above actually has.
Her question is: which verified colleague holds a well-phenotyped cohort of X, with the samples, the follow-up, and the consent scope I need, and would be willing to work with me?
And here is the structural irony. Federated networks anonymize the data holder by design. That anonymization is essential to their governance model and it removes precisely the person who could say yes.
There are further practical limits. Federated data is typically drawn from insured, academic, and acute care populations, with known gaps in adherence, lifestyle, and social determinants. Common data models frequently lack the specific variable a study needs, which researchers usually discover weeks into a data request process.
So the query infrastructure has been solved and the people-matching has not been touched.
What a cohort actually is, from the holder's side
To understand why sharing is difficult, you have to see the asset the way its owner does.
A well-phenotyped cohort is not a file. It represents years of recruitment, funding, consent processes, follow-up, sample handling, and data cleaning. It is frequently the basis of a research program, several careers, and future grant applications.
Sharing it carries real risks that the requester rarely appreciates:
Being scooped. Someone else publishes the obvious analysis first, on your data.
Credit disappearing. An acknowledgment where authorship was expected, or nothing at all.
Misuse. The data analyzed badly and published, with your cohort's name on it.
Governance exposure. Consent scope, data use agreements, and institutional review board conditions all constrain what can be shared with whom, and getting it wrong is a serious problem for the holder rather than the requester.
And pure workload. Preparing a dataset for external use is weeks of unfunded effort.
Against all of that, an unfunded cold email from a stranger offers nothing.
Now understand why the 74 percent figure is so high. A private, negotiated share with collaboration and authorship converts every one of those risks into a benefit. The holder gets credit, controls the analysis, and gains a collaborator.
Data does not move through repositories. It moves through people, on terms.
The related failure: the methodologist nobody can reach
There is a parallel version of this problem that affects an even larger population, and it is worth putting alongside it because the structure is identical.
Most physicians practise outside academic centers. When they attempt research, the wall they hit is not data. It is methodology.
From a community research network survey published in JABFM:
- 92 percent of community physicians cited lack of time.
- 80 percent cited inadequate research training.
- Focus groups added "lack of collaborators to sustain research."
- And most were unaware that CTSA-funded biostatistics consultation existed at all, despite it being funded partly to reach them.
A survey at a Dublin academic hospital found only about 45 percent felt staff were adequately supported for research, with over half lacking formal research training and requesting statistical assistance.
Meanwhile, on the supply side: a survey of 129 leaders representing 171 collaborative biostatistics units found a median of 9 staff, 86 percent NIH-funded, with leaders reporting too many projects and insufficient calendar time. Canada has launched a national training program in response to what its clinical trials community describes as a critical shortage of trial biostatisticians.
So there is genuine scarcity. But look at how the scarce resource is allocated: methodological capacity is institution-locked and internally rationed by department politics. A community physician has no institutional claim on it whatsoever.
The result is what you would expect. The clinician emails a former mentor, asks the quality department, which has analysts rather than methodologists, posts to a listserv, hires a freelance statistician with no clinical domain knowledge, or gives up and runs a t-test.
The shortage is real and partly an allocation illusion. Capacity is trapped inside institutions, rationed to internal faculty, while clinicians with better questions and more representative populations cannot reach it.
Why this matters more than it sounds
These two failures, cohort discovery and methodologist access, compound into something significant.
External validation is where most clinical prediction models go to die, and it requires someone else's cohort. The current mechanism for obtaining one has a 7 percent success rate through formal channels.
Community settings are where most care happens. Pragmatic trials, learning health systems, and implementation research all depend on exactly the clinicians who cannot get a statistician.
Underpowered single-center studies are a well-documented category of research waste, and the direct remedy, pooling with another site, is a people-finding problem.
And registries sit unused. Well-built cohorts exist that no one is using, because the people who could use them do not know they exist and the stewards cannot vet strangers.
What would work
Make the holder discoverable without making the data public. A registry of cohort metadata, not patient data: disease, approximate N, key variables, follow-up duration, sample availability, consent scope, and whether the holder is open to collaboration. That is publishable information that no institutional review board objects to, and it does not currently exist anywhere.
Verified identity on both sides. The holder's core problem is that they cannot assess a stranger. Verification of who the requester is, where they work, and what they have published converts an anonymous request into an evaluable one.
A collaboration covenant by default. Given that 74 percent of private shares produce authorship or collaboration, the default terms should reflect that: agreed authorship expectations, agreed credit, and a commitment to respond within a defined window rather than ignoring the request.
A response norm. The 93 percent non-response rate persists because there is no cost to silence. In a professional community where answering is expected and visible, a request gets a yes or a no, and a fast no is worth far more than months of silence.
Governance stays where it is. The introduction moves. The data does not. Institutional data use agreements and review board approvals remain fully in force, and any serious design must be explicit that it brokers the conversation rather than the dataset.
And the same structure serves methodology. A verified statistician or methodologist, reachable across institutional boundaries, with authorship or compensation attached, unlocks latent capacity including semi-retired methodologists who are currently invisible.
What you can do now
If you hold a cohort
Publish your metadata, even informally. A single page describing what you have, at what N, with what variables and follow-up, and that you are open to collaboration, put somewhere findable. You will be surprised who contacts you, and you retain complete control over what happens next.
Answer requests, even to decline. A fast no costs you two minutes and saves a colleague months. The 93 percent non-response rate is made up of individually reasonable decisions to deal with it later.
State your terms up front. Authorship expectations, analysis participation, and turnaround. Most requesters are happy to agree and simply do not know what to offer.
If you need a cohort
Offer collaboration, not a data request. The evidence is unambiguous that this is what works. Lead with what you bring and what the holder gets.
Use human introductions. A vouched introduction from a mutual connection converts a cold request into a warm one, and the difference in response rate is the difference between 7 percent and something far higher.
Ask the hallway question deliberately. "Who holds a good cohort of X" is a question senior colleagues can frequently answer instantly. Most researchers ask it accidentally at conferences rather than systematically.
If you need methodology
Ask whether your regional CTSA hub offers consultation. The survey evidence indicates most community clinicians do not know these services exist, and some are explicitly funded to serve beyond their own institution.
Offer authorship explicitly. A collaborative biostatistician's currency is publications, and a clinician offering genuine co-authorship on an interesting question is a far better proposition than an hourly engagement.
Ask before you collect data, not after. The most common way a community research project fails is that a methodologist is consulted after the data has been gathered in a form that cannot answer the question.
If you make research policy
Stop mandating statements and start measuring compliance. A policy with a documented 7 percent fulfillment rate is not a policy, it is a formality. Auditing actual response rates and publishing them by journal would change behavior more than another mandate.
Fund the introduction, not just the repository. Repositories have been funded generously for a decade. The evidence says sharing happens through relationships, and nobody has funded the relationship layer.
Frequently asked questions
Do researchers actually share data when they say it is available on request? Usually not. A study examining 1,792 manuscripts whose authors stated data were available on reasonable request found 93 percent did not respond or refused, with only 6.8 percent sharing. A separate analysis found that among systematic reviews with data statements, only 13 percent had genuinely downloadable data while 42 percent used the on-request formulation.
Why do researchers refuse to share data? Because a cohort represents years of work and continuing research value, and sharing carries real risks: being scooped, losing credit, seeing the data analyzed poorly, governance and consent complications, and substantial unfunded preparation effort. A cold request from an unknown party offers nothing against any of that.
Does private data sharing work better than formal channels? Substantially. Survey research found that among cancer researchers who had shared data privately, 74 percent said it led to authorship or a future collaboration. The variable is not policy but trust and reciprocity.
Do federated research networks solve cohort discovery? They solve a different problem. PCORnet, TriNetX, and N3C answer how many patients with a given profile exist across the network, which is valuable. They anonymize the data holder by design, which is essential to their governance and removes the specific person who could agree to collaborate.
Why can't community physicians get statistical support? Because methodological capacity is institution-locked. Survey data found 92 percent of community physicians cite lack of time and 80 percent inadequate research training, with most unaware that funded biostatistics consultation exists. Meanwhile academic biostatistics units, with a median of nine staff, report more projects than calendar time and allocate internally.
What would improve research data sharing? Making cohort holders discoverable through published metadata rather than making data public, verified identity so holders can evaluate requesters, default collaboration and authorship terms reflecting how sharing actually succeeds, and a professional norm of responding to requests even to decline. Institutional data governance remains unchanged; what moves is the introduction.
The bottom line
A decade of data sharing policy has produced journal mandates, funder requirements, and a standard sentence that appears throughout the medical literature and works 7 percent of the time.
The same period produced federated networks of genuine sophistication, covering tens of millions of patients, which answer how many patients exist and deliberately conceal who holds them.
Meanwhile the channel that actually works, private sharing between people who trust each other, produces authorship or collaboration in 74 percent of cases and has received essentially no infrastructure investment at all.
The researcher who needed a validation cohort sent four careful emails into a void and then solved her problem in a hallway conversation, because the hallway contains the one thing no repository does: a person who knows who holds what, and can vouch for her.
Data does not move through repositories. It never has. It moves through people, on terms, and we have spent ten years building everything except the part that carries it.
Part of a series on the missing professional infrastructure of healthcare. Previously: Testimony Without Peers
Evidence note: data sharing compliance figures come from Gabelica et al. and Page et al. in the Journal of Clinical Epidemiology (2022). Private sharing outcomes come from a 2025 survey of 285 cancer researchers published in Accountability in Research. Federated network figures are as published by PCORnet, TriNetX-related analyses, and NCATS. Community research barriers come from JABFM (2009) and PLOS One (2026). Biostatistics capacity figures come from a survey of 129 leaders representing 171 collaborative units (2023) and related work in Stat (2022). Some survey findings are single-field or single-country and may not generalize.