Most Copilot programs measure whether program managers like the tool. Almost none can say whether the tool changed the program. This framework closes that gap — and the instrument below shows you exactly how wide it is.
Origin — the source argument
In July 2026 the Harvard Business School AI Institute published Does Your AI Work? Test It Where It Lives, summarising work by David Arbour, Iavor Bojinov, Avi Feller and Tu Ni. Their argument, paraphrased: benchmark scores earn credibility in the lab but may not carry over to the population that actually uses the system, while field dashboards and satisfaction surveys sit close to real work yet cannot show that the AI caused anything. Their proposed target is the causal field evaluation — a randomised experiment run in the live setting, measured on outcomes that matter.
Source article: HBS AI Content & Learning, “Does Your AI Work? Test It Where It Lives,” Harvard Business School AI Institute, 14 July 2026. Underlying paper: Arbour, D., Bojinov, I., Feller, A., & Ni, T., “Toward Causal Field Evaluations of AI Systems,” Harvard Data Science Review 8(2), 11 May 2026. PROOF is my applied adaptation of that argument to enterprise program management; the underlying claims are theirs, the operating model is mine. — P.M.
Three consequences of that argument travel directly into program management, and they are what PROOF is designed to handle. First, the article's warning about self-selected adoption is the central measurement problem in every Copilot rollout I have seen: the program managers who take the licence early are not like the ones who don't, so any comparison between them measures the person, not the tool. Second, its caution about preference data indicts almost the whole standard Copilot scorecard — seats, prompts, satisfaction, self-reported hours saved. Third, its point that gains carry costs means a Copilot evaluation without guardrail metrics isn't an evaluation, it's a sales document.
Instrument 01 — diagnostic
The source article's two axes — lab or field, preference or outcome — make a map. Tap any metric from a typical Copilot program dashboard and see which quadrant it falls into, and what it can and cannot support.
Lab × Preference
Bench & bake-off
Model comparisons, prompt bake-offs, “which draft reads better” panels. Useful for picking a tool. Says nothing about your programs.
Field × Preference
The default dashboard
Seats, prompts per week, CSAT, self-reported time saved. Where most Copilot reporting lives, and where it stops.
Lab × Outcome
Scripted task trial
Timed exercises on synthetic status reports. Clean and causal, but the population and the pressure are wrong.
Field × Outcome
Causal field evaluation — the target
Randomised or staggered licence assignment across real programs, read on schedule, rework, risk lead time and cost. This is the quadrant PROOF drives you to.
Pick a metric from your deck
Awaiting selection
Each one is scored on what it can legitimately support: tool selection, adoption tracking, mechanism explanation, or a causal claim about program performance.
The framework
Each layer exists because a specific claim in the source argument breaks a specific habit in enterprise program measurement. Run them in order; skipping a layer invalidates the ones after it.
Answers: lab results may not generalise to the deployment population.
Answers: define the experiment you would run if you could, then measure how far your real design falls from it.
Answers: people may prefer output that looks better while performing worse downstream.
Answers: when adoption is left to happen, the people who lean in differ from those who don't.
Answers: speed gains can carry quality costs; guardrails can cost advanced users their flexibility.
Instrument 02 — design selector
Four questions about what your PMO can actually control. The output is the strongest design available to you, how far it sits from a clean randomised trial, and the specific threats you will have to defend at the steering committee.
Instrument 03 — estimation
Four numbers from your PPM tool. The calculator shows the two claims a steering committee will hear and the one you can defend. Example below uses status report cycle time in hours, where lower is better.
Before / after, treated only
–4.9 h
Treated vs control, after only
–2.1 h
Difference-in-differences
–2.7 h
Difference-in-differences is only valid if the two cohorts were drifting in parallel before rollout. Plot six or more pre-periods and eyeball it before you quote the number. It is a substitute for randomisation, not an equal to it.
Underpowered pilots are the quiet failure mode: a real effect exists, the study cannot see it, and Copilot gets written off — or worse, noise gets promoted as success. Size the pilot before you run it.
Programs needed per arm
—
At 80% power, 5% significance, two-sided.
Smallest effect you can detect
—
With the arm size you entered.
Verdict
—
If the numbers say your portfolio is too small, that is a finding, not a blocker. Switch to a stepped-wedge rollout so every program contributes both a control period and a treated period, or lengthen the measurement window to reduce period-level noise.
Instrument 04 — the metric tree
The tiers below are ordered by evidential weight, not by ease of collection — which is exactly the inverse of how most Copilot dashboards get built. Tap to open each tier.
Instrument 05 — the interview kit
Be honest about what interviews do. They cannot establish causation — that is the whole point of the source argument, and asking PMs whether Copilot helped produces exactly the preference data it warns against. What interviews can do is explain the mechanism behind an effect you measured, bound it when you couldn't measure it, and catch the effects your telemetry was never instrumented to see. Run them alongside the quantitative read, never instead of it. Highlighted lines are counterfactual probes: they force the respondent to reason about absence rather than benefit.
Operating model
The sequence matters more than the speed. Every gate below is a point where a program has historically been able to quietly convert itself back into a preference-data exercise.
References & attribution
The P&G field research referenced in the source article is cited there as evidence that generative AI reshapes team structure and knowledge work itself, not merely individual task speed — the reason PROOF sets the unit of analysis at squad or program level rather than at the individual PM. Readers should consult the HDSR paper for that citation in full.