Skip to content
CorpDev Wiki
4 min read

Measurement and Learning Across Acquisitions

A learning acquisition program compares what it expected with what happened, then changes a specific decision rule or operating practice. Measure investment outcomes, process performance, and AI quality separately so an improvement in one cannot conceal a failure in another.

Keep the original investment case and its evidence. Without that baseline, a revised forecast can make almost any outcome appear planned.

Build a scorecard with meaningful denominators

Choose measures that lead to a decision. Report the population, period, exclusions, and owner with each measure.

Layer Example measure Decision it supports
Strategy Outcomes by thesis and acquisition cohort Continue, revise or stop an acquisition theme
Investment Actual cash flows and value against the approved case Reallocate capital and revise assumptions
Integration Benefit delivery and dependency delays Change sequence, resources or scope
Sourcing Qualified coverage and sampled missed targets Improve research sources and screening rules
Workflow Time to reviewed output and rework rate Repair bottlenecks
AI quality Claim support, extraction errors and reviewer overrides Change retrieval, models or instructions
Economics Total cost per accepted output Decide whether automation earns its cost

Define “accepted output.” A memo returned for missing sources is not accepted just because it was generated successfully. Include the cost of review and correction alongside model and data-provider costs.

Evaluate the decision the workflow supports

Build a test set from permitted historical cases and synthetic edge cases. Have qualified reviewers label expected findings and acceptable evidence. Separate the development set from a held-out set used to assess changes.

For target screening, review both included and rejected companies. Precision asks how many flagged companies truly meet the agreed standard. Recall asks how many qualifying companies the system found within a labeled set. Recall across an unknown market universe cannot be measured simply by counting database results.

For diligence, test whether claims are supported by the cited passage, whether important contrary evidence is surfaced, and whether unknowns remain unknown. A citation can be real and still fail to support the sentence beside it.

For financial work, validate formulas, units, periods, and expected sensitivities. For actions, test authorization, duplicate handling, and cancellation at the actual write boundary.

Keep failures visible by severity

The following hypothetical evaluation contains 100 screening cases with a reviewed reference label. The results are illustrative, not acceptance targets.

Result Count
Qualifying cases in reference set 40
Cases flagged by system 35
Correctly flagged qualifying cases 28
Incorrectly flagged cases 7
Qualifying cases missed 12

Precision is 28 / 35 = 80%. Recall is 28 / 40 = 70%. The remaining 53 cases were correctly rejected. Whether this is acceptable depends on the cost of missed targets and the review capacity available.

Report results by segment and evidence quality. A strong overall score can hide poor results on private companies with sparse public information. Report permission failures and unauthorized actions individually; never average them away inside a quality score.

Turn outcomes into proposed changes

After closing, review results by deal age and thesis. Separate errors in strategy, target assessment, price, and execution. Include declined and abandoned opportunities where evidence is available, while acknowledging that their eventual counterfactual outcomes may remain unknown.

Propose a concrete change: a screening question, earlier diligence test, different integration sequence, or revised modeling assumption. Name its owner and rationale. Review it on historical and held-out examples before applying it to active deals.

Do not let the model rewrite the acquisition policy automatically based on a handful of outcomes. Changes to policy need approval; changes to workflow implementation need evaluation and release controls.

Explain overrides before changing the workflow

Reviewer edits are evidence to investigate. Some correct mistakes; others reflect new information, personal preference, or a change in strategy.

Classify a reviewed sample before choosing a repair.

Cause Appropriate repair
Parent-company financials attached to a subsidiary Fix entity resolution and test similar records
Reviewer applied a newer thesis Correct version handling and reassess affected targets
Source does not support the claim Improve retrieval or extraction and test claim support
Reviewer prefers shorter prose Change editorial instructions without changing the target score

Have the relevant owners validate the classification. Test a defined correction on held-out cases and monitor reviewed live outputs. Keep the previous version available for rollback. Automatically treating every edit as a correct training label can teach the system the wrong lesson.

Continue with the implementation playbook and program reporting.