PROOF A causal field framework for measuring Microsoft Copilot in Program Management

Most Copilot programs measure whether program managers like the tool. Almost none can say whether the tool changed the program. This framework closes that gap — and the instrument below shows you exactly how wide it is.

Fig. 01 — The attribution gap Drag the sliders. Watch the claim collapse.
Observed delivery performance versus its counterfactual A line chart comparing what happened after Copilot rollout against what would have happened anyway. The shaded gap between the two lines is the only defensible causal effect.
Observed (Copilot cohort) Counterfactual Causal effect

Team maturity, process fixes, learning curve — gains you would have booked with no Copilot at all.

The PMO rewrite, the new tooling, the reorg — running in the same window and moving the same numbers.

The part that would vanish if you switched the licences off tomorrow.

Before / after claim +23%
Defensible causal effect +6%
Overstatement 3.8×

Origin — the source argument

The idea this is built on

In July 2026 the Harvard Business School AI Institute published Does Your AI Work? Test It Where It Lives, summarising work by David Arbour, Iavor Bojinov, Avi Feller and Tu Ni. Their argument, paraphrased: benchmark scores earn credibility in the lab but may not carry over to the population that actually uses the system, while field dashboards and satisfaction surveys sit close to real work yet cannot show that the AI caused anything. Their proposed target is the causal field evaluation — a randomised experiment run in the live setting, measured on outcomes that matter.

Source article: HBS AI Content & Learning, “Does Your AI Work? Test It Where It Lives,” Harvard Business School AI Institute, 14 July 2026. Underlying paper: Arbour, D., Bojinov, I., Feller, A., & Ni, T., “Toward Causal Field Evaluations of AI Systems,” Harvard Data Science Review 8(2), 11 May 2026. PROOF is my applied adaptation of that argument to enterprise program management; the underlying claims are theirs, the operating model is mine. — P.M.

Three consequences of that argument travel directly into program management, and they are what PROOF is designed to handle. First, the article's warning about self-selected adoption is the central measurement problem in every Copilot rollout I have seen: the program managers who take the licence early are not like the ones who don't, so any comparison between them measures the person, not the tool. Second, its caution about preference data indicts almost the whole standard Copilot scorecard — seats, prompts, satisfaction, self-reported hours saved. Third, its point that gains carry costs means a Copilot evaluation without guardrail metrics isn't an evaluation, it's a sales document.

Instrument 01 — diagnostic

Where does your current Copilot metric actually sit?

The source article's two axes — lab or field, preference or outcome — make a map. Tap any metric from a typical Copilot program dashboard and see which quadrant it falls into, and what it can and cannot support.

Outcome data  ←  Preference data

Lab × Preference

Bench & bake-off

Model comparisons, prompt bake-offs, “which draft reads better” panels. Useful for picking a tool. Says nothing about your programs.

Field × Preference

The default dashboard

Seats, prompts per week, CSAT, self-reported time saved. Where most Copilot reporting lives, and where it stops.

Lab × Outcome

Scripted task trial

Timed exercises on synthetic status reports. Clean and causal, but the population and the pressure are wrong.

Field × Outcome

Causal field evaluation — the target

Randomised or staggered licence assignment across real programs, read on schedule, rework, risk lead time and cost. This is the quadrant PROOF drives you to.

Lab  →  Field

Pick a metric from your deck

Awaiting selection

Choose a metric to diagnose it

Each one is scored on what it can legitimately support: tool selection, adoption tracking, mechanism explanation, or a causal claim about program performance.

The framework

PROOF — five layers, each answering one failure

Each layer exists because a specific claim in the source argument breaks a specific habit in enterprise program measurement. Run them in order; skipping a layer invalidates the ones after it.

PLayer 01

Place it in the field

Answers: lab results may not generalise to the deployment population.

  • Evaluate inside live programs, not in a Copilot sandbox or a training cohort.
  • Unit of analysis is the program or squad, not the individual PM — Copilot's effect travels through meetings and handovers, so person-level analysis leaks.
  • Freeze the deployment population up front: which portfolios, which delivery models, which client tiers.
RLayer 02

Randomise — or declare the distance

Answers: define the experiment you would run if you could, then measure how far your real design falls from it.

  • Write the ideal trial first, even if you cannot run it. It becomes the yardstick for every compromise.
  • Licences are scarce and staged anyway — that scarcity is a free randomisation instrument. Allocate wave 1 by lottery among eligible programs, not by enthusiasm.
  • If you cannot randomise, state the assumption your alternative rests on, in writing, in the steering pack.
OLayer 03

Outcomes over opinions

Answers: people may prefer output that looks better while performing worse downstream.

  • Demote seats, prompts and satisfaction to adoption telemetry. They are health checks, not evidence.
  • Promote one primary outcome and pre-commit to it: I default to risk detection lead time or status report cycle time, both machine-readable from the PPM tool.
  • Never let a self-reported hours-saved survey enter a benefits case unaccompanied.
OLayer 04

Offset the selection

Answers: when adoption is left to happen, the people who lean in differ from those who don't.

  • Analyse by intention to treat — keep programs in the arm they were assigned to, even if the PM never opened Copilot. Analysing only active users rebuilds the bias you removed.
  • Log the confounder register at baseline: concurrent initiatives, seniority mix, portfolio complexity, quarter-end seasonality, attrition.
  • Check parallel pre-trends before trusting any difference-in-differences read.
FLayer 05

Falsify, and face the trade-off

Answers: speed gains can carry quality costs; guardrails can cost advanced users their flexibility.

  • Pre-register the primary outcome, the analysis and the stopping rule before wave 1 goes live.
  • Run a negative control — a metric Copilot should not move. If it moves, your design is contaminated.
  • Report guardrails on the same slide as the benefit, never an appendix: artifact error rate, unchallenged-acceptance rate, senior override rate.

Instrument 02 — design selector

Define the ideal experiment, then find your nearest feasible one

Four questions about what your PMO can actually control. The output is the strongest design available to you, how far it sits from a clean randomised trial, and the specific threats you will have to defend at the steering committee.

Your constraints
Recommended design

Instrument 03 — estimation

Difference-in-differences, on the back of an envelope

Four numbers from your PPM tool. The calculator shows the two claims a steering committee will hear and the one you can defend. Example below uses status report cycle time in hours, where lower is better.

Before / after, treated only

–4.9 h

Treated vs control, after only

–2.1 h

Difference-in-differences

–2.7 h

Difference-in-differences is only valid if the two cohorts were drifting in parallel before rollout. Plot six or more pre-periods and eyeball it before you quote the number. It is a substitute for randomisation, not an equal to it.

How many programs do you need?

Underpowered pilots are the quiet failure mode: a real effect exists, the study cannot see it, and Copilot gets written off — or worse, noise gets promoted as success. Size the pilot before you run it.

Programs needed per arm

At 80% power, 5% significance, two-sided.

Smallest effect you can detect

With the arm size you entered.

Verdict

If the numbers say your portfolio is too small, that is a finding, not a blocker. Switch to a stepped-wedge rollout so every program contributes both a control period and a treated period, or lengthen the measurement window to reduce period-level noise.

Instrument 04 — the metric tree

Demote four metrics, promote four, and never drop the guardrails

The tiers below are ordered by evidential weight, not by ease of collection — which is exactly the inverse of how most Copilot dashboards get built. Tap to open each tier.

Instrument 05 — the interview kit

Four sessions that explain the number without pretending to prove it

Be honest about what interviews do. They cannot establish causation — that is the whole point of the source argument, and asking PMs whether Copilot helped produces exactly the preference data it warns against. What interviews can do is explain the mechanism behind an effect you measured, bound it when you couldn't measure it, and catch the effects your telemetry was never instrumented to see. Run them alongside the quantitative read, never instead of it. Highlighted lines are counterfactual probes: they force the respondent to reason about absence rather than benefit.

Operating model

Ninety days, five gates

The sequence matters more than the speed. Every gate below is a point where a program has historically been able to quietly convert itself back into a preference-data exercise.

References & attribution

Sources

  1. HBS AI Content & Learning. “Does Your AI Work? Test It Where It Lives.” Harvard Business School AI Institute, 14 July 2026. aiinstitute.hbs.edu/does-your-ai-work-test-it-where-it-lives — the source article for this framework.
  2. Arbour, David, Iavor Bojinov, Avi Feller, and Tu Ni. “Toward Causal Field Evaluations of AI Systems.” Harvard Data Science Review 8(2), 11 May 2026. — the underlying paper, and the origin of the lab/field × preference/outcome taxonomy, the causal field evaluation target, and the ideal-experiment yardstick used throughout PROOF.
  3. Rubin, Donald B. “Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies.” Journal of Educational Psychology 66(5), 1974. — potential outcomes; the formal basis for the counterfactual in Fig. 01.
  4. Campbell, Donald T., and Julian C. Stanley. Experimental and Quasi-Experimental Designs for Research. Rand McNally, 1963. — the quasi-experimental ladder behind Instrument 02.
  5. Card, David, and Alan B. Krueger. “Minimum Wages and Employment.” American Economic Review 84(4), 1994. — canonical difference-in-differences application.
  6. Abadie, Alberto, Alexis Diamond, and Jens Hainmueller. “Synthetic Control Methods for Comparative Case Studies.” Journal of the American Statistical Association 105(490), 2010.
  7. Angrist, Joshua D., and Jörn-Steffen Pischke. Mostly Harmless Econometrics. Princeton University Press, 2009.
  8. Flanagan, John C. “The Critical Incident Technique.” Psychological Bulletin 51(4), 1954. — basis for interview session B.
  9. Dalkey, Norman, and Olaf Helmer. “An Experimental Application of the Delphi Method to the Use of Experts.” Management Science 9(3), 1963. — basis for interview session D.
  10. Strathern, Marilyn. “Improving Ratings: Audit in the British University System.” European Review 5(3), 1997. — the widely cited formulation of Goodhart's law; the reason every metric in Instrument 04 carries a gaming-risk rating.

The P&G field research referenced in the source article is cited there as evidence that generative AI reshapes team structure and knowledge work itself, not merely individual task speed — the reason PROOF sets the unit of analysis at squad or program level rather than at the individual PM. Readers should consult the HDSR paper for that citation in full.