Here is a question that sounds simple and turns out to be nearly unanswerable.
Your patient needs a pancreaticoduodenectomy. A Whipple. It is one of the most technically demanding operations in general surgery, and the evidence that operator experience matters is about as strong as evidence gets in medicine.
How many did the surgeon you are about to refer to perform last year?
Try to answer it. Not for a hypothetical surgeon, for a real one, in your city, right now.
You cannot look it up. The hospital may publish institutional volume. The surgeon's bio will say they specialize in hepatobiliary surgery. Their board certification says general surgery, which is a fact about an examination. The registry that knows the answer will not tell you. And asking the surgeon directly feels rude, so almost nobody does.
So you refer on reputation, habit, and who returned your call, for an operation where the published data says the difference between operators is measured in deaths.
What the evidence actually says
The volume-outcome relationship is one of the most replicated findings in surgical health services research. It is worth stating precisely, because precision is what makes the transparency gap indefensible.
Birkmeyer and colleagues, New England Journal of Medicine, analyzing 474,108 Medicare patients:
- For pancreatic resection, low-volume surgeons carried an adjusted odds ratio of 3.61 for operative death compared to high-volume surgeons.
- Surgeon volume mediated 55 percent of the hospital-volume effect for pancreatectomy and 46 percent for esophagectomy.
That second finding is the one that reframes the whole issue and gets consistently overlooked. For decades, quality initiatives have steered patients toward high-volume hospitals. Birkmeyer's data indicates that roughly half of that hospital effect is actually the surgeon.
We have been measuring the building when a large share of the signal is the person.
Aquina and colleagues, Cancer (2021), 112,154 patients:
Among surgeons who all met Leapfrog volume standards, there was still a two-fold difference in adjusted complication rates between the best and worst performers.
Read that carefully, because it undercuts the easy fix. Volume thresholds help, and they are not sufficient. Even inside the group that clears the bar, outcomes vary by a factor of two. Volume is a useful proxy for something deeper, and clearing a minimum is not the same as being good.
Learning curves are real, steep, and specific.
A meta-analysis of robotic lobectomy put the learning curve at 25.3 plus or minus 12.6 cases. Across robotic procedures generally, proficiency estimates cluster in the 15 to 55 case range, with early-phase complication rates in single-surgeon series running several times higher than later-phase rates.
Now hold that next to the pace of technology change. Robotic platforms, transcatheter valves, left atrial appendage occluders, neuromodulation devices. A surgeon board-certified in 2009 may be using a device that did not exist in 2019.
"Board certified" tells you nothing about proficiency with the specific thing being implanted this year. It cannot, structurally. It was never designed to.
The data exists. It is simply not available to you.
This is the part that makes the situation genuinely absurd rather than merely unfortunate.
Operator-level volume data is meticulously collected. It sits in the Society of Thoracic Surgeons database, in the National Cardiovascular Data Registry, in NSQIP, in device manufacturer records, and in hospital operative logs. Somebody knows exactly how many of these your surgeon did last year.
Consider transcatheter aortic valve replacement. CMS national coverage determination requires that hospitals meet volume thresholds (at least 50 aortic valve replacements per year including at least 20 TAVR, or 100 over two years including 40 TAVR) and report operator-level data to a registry.
The operator data is collected as a condition of Medicare coverage. It is not public.
The Leapfrog Group publishes volume standards for procedures including bariatric surgery, scoring facility volume and surgeon volume thresholds, and reports at the hospital level only.
And there was one serious attempt to change this. ProPublica's Surgeon Scorecard, published in 2015, used Medicare data from 2009 to 2013 to publish complication rates for 16,827 individual surgeons.
It has not been updated since July 2015.
The methodological critiques were substantial and in significant part fair: risk adjustment on administrative data is genuinely difficult, complication attribution is contestable, and low case counts make individual estimates unstable. But the project also generated intense professional opposition, and the combination ended public operator-level reporting in the United States for a decade.
The obstacle to operator-level transparency has never been data availability. It is professional politics, and the politics are not entirely unreasonable, which is why the problem has stayed unsolved.
The case against transparency, taken seriously
An honest treatment has to engage the objections, because they are not merely self-interested.
Risk adjustment is genuinely hard. A surgeon who takes the sickest, most complex referrals will look worse on unadjusted outcomes and may be the best in the region. Publishing crude numbers actively punishes the people you most want doing hard cases. This is a real and serious problem.
Small numbers are unstable. For a procedure done 20 times a year, a single bad outcome swings the rate dramatically. Individual-level statistics on rare procedures are frequently noise.
Risk aversion is a documented response. Evidence from public cardiac surgery reporting suggests surgeons become more selective about accepting high-risk patients when individually reported, which harms exactly the patients with the fewest options.
Volume is a proxy, not the thing itself. As the Aquina data shows, two-fold outcome variation persists among surgeons who all clear volume standards.
Each of these is a genuine argument against public, risk-adjusted, outcome-based scorecards.
Notice that none of them is an argument against a referring physician being able to learn how many of a specific procedure an operator performed last year.
That is a count. It requires no risk adjustment. It is not an outcome measure. It cannot be gamed by patient selection. It is simply a fact, and it is the single most useful piece of information a referrer could have, and it is unobtainable.
The debate about scorecards has been conflated with the question of case counts, and the conflation has protected an information vacuum that serves nobody, including surgeons.
Why surgeons should want this more than anyone
The framing of volume transparency as adversarial is a strategic error, and it has cost the profession a great deal.
Consider who is harmed by the current opacity.
The high-volume specialist. A surgeon who does 60 of a complex procedure a year is competing for referrals against colleagues who do four, on entirely equal terms in the eyes of every directory. Their most valuable professional asset is invisible.
The developing surgeon. With learning curves running 15 to 55 cases, the honest question "who can proctor me through my first ten?" has no infrastructure. Proctoring relationships form through personal networks and device representatives, which is a genuinely strange way to distribute one of the most consequential forms of medical education.
The referring physician who genuinely wants to do right by their patient, and is guessing.
The patient, obviously.
And the surgeon who has quietly stopped doing a procedure, who is still receiving referrals for it because nothing anywhere recorded the change, and who now faces a difficult conversation or an uncomfortable case.
The only party served by opacity is the abstraction of professional solidarity, and it is being purchased at the price of an odds ratio of 3.61.
There is also a personal test that cuts through the abstraction quickly. Every surgeon who opposes volume transparency has, at some point, needed to find a surgeon for a family member. And in that moment, every one of them used exactly the information the public cannot get: they called colleagues and asked who actually does a lot of these and is good at them.
The profession already has this data. It is exchanged privately, constantly, on behalf of the families of people who happen to know surgeons.
What a workable system looks like
The failure of public scorecards tells you a great deal about the design constraints, and they point somewhere quite specific.
Counts, not outcomes. Volume and recency. No risk adjustment required, no attribution disputes, no patient-selection gaming. Enormously more useful than nothing and vastly less contentious than mortality rates.
Device and technique specific. "Cardiologist" is useless. "Has implanted 40 of this specific valve platform in the last 12 months" is decision-grade.
Attested by the operator, corroborated by peers. Self-report alone fails, since the Davis JAMA research shows self-assessment is least reliable among the least skilled. But a count corroborated by someone who was in the room, or who referred cases and saw outcomes, is checkable in a way that a rating never is.
Visible to verified peers, not the public. This is the design decision that makes the whole thing feasible. Almost every objection to transparency concerns public reporting: media misinterpretation, patient misunderstanding, risk aversion driven by publicity. A peer-only, verified system dissolves most of these while serving the referring physician, who is the person actually making the routing decision.
Recency-weighted. Twelve months matters far more than career total. A surgeon with 400 lifetime cases and none in three years is a different proposition from one with 90 cases in the last year.
With reciprocity built in. You get access to this data when you need a surgeon for your own family, and in exchange you attest honestly about your own practice. That is the exchange that makes honest contribution rational, and it is exactly how the informal version already works.
Paired with proctor matching. The learning curve data means the most valuable use of a case-exposure map is not only routing patients. It is connecting a surgeon at case eight to one at case four hundred. That is an unambiguous patient-safety win and it turns the whole system from a judgment mechanism into a teaching one.
What to do now
If you refer for procedures
Ask the question. "How many of these do you do a year?" It feels awkward exactly once. Surgeons who do a lot are generally pleased to answer, and the answer is more informative than anything in any directory. A hesitant or vague answer is itself information.
Ask about the specific device or technique. Especially for anything introduced in the last five years, where the learning curve literature says the first 15 to 55 cases carry elevated risk.
Keep your own list, with numbers and dates. Your personal referral record is the only volume database you will ever have access to. Make it explicit rather than leaving it in memory.
If you are an operator
Track and state your own numbers. Procedure, device, count, twelve-month window. If you do a lot, this is your most valuable and currently most invisible professional asset. If you do few of something, knowing it precisely is how you decide whether to keep doing it, refer it out, or get proctored.
Be honest about your learning curve position. The literature says 15 to 55 cases for robotic proficiency. A surgeon at case six who says so and arranges proctoring is behaving better than one who does not, and every experienced surgeon knows it.
Offer to proctor. If you are past the curve, you hold something that is scarce, valuable, and currently distributed by accident of who knows whom.
If you run a department or a credentialing committee
Look at your own operator-level data. You have it. Most departments have never systematically examined the distribution of case volume by operator by procedure. The distribution is usually more skewed than anyone expects.
Set a floor and a route, not just a floor. A volume threshold that says "you may not do this" without providing a proctoring path to competence is a rule rather than a system.
Separate certification from currency. Privileging that verifies a certificate but not recent volume is verifying the wrong variable, and everyone in the room knows it.
Frequently asked questions
Does surgeon volume affect outcomes? Substantially, for complex procedures. Birkmeyer's analysis of 474,108 Medicare patients found an adjusted odds ratio of 3.61 for operative death after pancreatic resection comparing low-volume to high-volume surgeons, and found surgeon volume mediated 55 percent of the hospital-volume effect for pancreatectomy and 46 percent for esophagectomy.
Is hospital volume or surgeon volume more important? Both matter and they are entangled. Birkmeyer's data indicates roughly half the apparent hospital-volume effect for pancreatectomy and esophagectomy is attributable to surgeon volume, which means volume-based quality initiatives focused only on institutions are capturing part of the signal.
Can I find out how many procedures my surgeon has done? Generally not through any public source. Operator-level volume is collected in clinical registries, including as a condition of Medicare coverage for procedures like TAVR, but is not published. Public reporting programs including Leapfrog report at the hospital level. The most reliable route is simply to ask the surgeon directly.
What happened to the ProPublica Surgeon Scorecard? It was published in 2015 using Medicare data from 2009 to 2013 covering 16,827 surgeons, and has not been updated since July 2015. It attracted substantial methodological criticism regarding risk adjustment and attribution, alongside significant professional opposition. No comparable public operator-level reporting has replaced it.
How many cases does it take to become proficient at a robotic procedure? It varies by procedure. A meta-analysis put robotic lobectomy at roughly 25.3 plus or minus 12.6 cases, and estimates across robotic procedures generally cluster between 15 and 55 cases, with elevated complication rates during the early phase in single-surgeon series.
If a surgeon meets volume standards, is that sufficient? Not entirely. Research in Cancer examining 112,154 patients found a two-fold difference in adjusted complication rates between the best and worst performing surgeons who all met Leapfrog volume standards. Volume is a useful proxy and clearing a threshold is not equivalent to excellence.
The bottom line
The relationship between operator volume and patient outcomes is among the best-established findings in surgery. The data is collected systematically, in national registries, sometimes as a legal condition of payment.
And a physician trying to route a patient to the right surgeon for a Whipple cannot obtain a single number.
The usual explanation is that transparency is complicated, and for outcome scorecards that is genuinely true: risk adjustment is hard, small numbers are unstable, and public reporting has documented side effects on patient selection.
None of that applies to a case count. A count needs no risk adjustment. It cannot be gamed by taking easier patients. It is a fact, it is already recorded, and it is the single most useful piece of information in the referring physician's decision.
Meanwhile, every surgeon in the country knows how to find this information when their own family needs an operation. They call three colleagues and ask who does a lot of these.
That informal network is medicine's real volume database. It works well, and it is available only to people who happen to know a surgeon.
Everyone else gets the directory, where the person who does sixty a year and the person who does four look exactly the same.
Part of a series on the missing professional infrastructure of healthcare. Previously: Your Directory Is Describing a Doctor Who No Longer Exists
Evidence note: sources include Birkmeyer et al. in the New England Journal of Medicine (2003) analyzing 474,108 Medicare patients; Aquina et al. in Cancer (2021) analyzing 112,154 patients; meta-analytic and single-institution robotic learning curve literature; the CMS national coverage determination for TAVR (CAG-00430R); Leapfrog Group survey standards; and the ProPublica Surgeon Scorecard (2015). Learning curve estimates vary considerably by procedure, surgeon, and definition of proficiency, and should be read as ranges rather than thresholds.