Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

Advances in Quantitative Evaluation Models for Measuring Corporate Diversity and Inclusion Investments

Ngonadi Uchechi, Michael Ominyi, Cyril Chimelie Anichukwueze, Ngozi Samuel, Uzougbo.

Abstract

Corporate expenditure on diversity and inclusion (D&I) grew sharply between 2020 and 2022, yet the evaluation methods used to justify that expenditure lagged well behind it. Most organizations continued to report headcount ratios that describe a workforce state rather than estimate the return generated by an investment. This paper reviews the quantitative evaluation models available as of 2022, traces their development across three methodological generations, and assesses how well each supports causal inference about D&I investments. It examines nine families of technique: utility-analysis and human capital return on investment models, quasi-experimental econometric designs, latent-variable psychometrics, stochastic flow models of the talent pipeline, pay equity decomposition, organizational network analysis, causal machine learning, impact-weighted accounting, and real-options framing. The review finds that the principal advance of the 2020 to 2022 period was not a single new estimator but a shift in the unit of analysis, moving from stock measures of representation toward flow, perception, and monetized outcome measures. It also finds that this advance was accompanied by a widening validity gap, in which disclosure mandates and index providers standardized precisely the metrics with the weakest link to firm outcomes. Drawing on more than two hundred studies from adjacent applied literatures, including occupational safety climate, data governance, internal audit, marketing attribution, procurement, health equity, energy justice, condition monitoring, and education, it identifies six transferable methodological commitments that current D&I disclosure frameworks omit. The review then sets out a layered correspondence between evaluation questions, estimators, and identifying assumptions, running from committed input through activity, proximal outcome, and process outcome to terminal value, together with criteria for assessing whether a given metric is fit for the layer at which it is used. It concludes with a research agenda addressing endogeneity, intersectional small-cell problems, cross-jurisdictional data constraints, and metric gaming.

Keywords

diversity and inclusion; human capital analytics; return on investment; causal inference; ESG measurement; impact-weighted accounts; people analytics; corporate disclosure

References

d monetary value for workforce composition and quality. It does not identify the marginal effect of any particular investment, since the measure is a level rather than a difference. 4.9 Real options and portfolio framing A final and less developed family treats D&I investment as a portfolio of options rather than as a set of projects with point-estimate returns. The framing is appropriate because D&I interventions share the characteristics that make real options analysis apt: high uncertainty about payoff, substantial irreversibility in reputational commitment, sequential decision structure, and long, variable lags between expenditure and observable outcome. Under this framing, a pilot programme is priced as a call option on organizational learning, with value derived from the right to expand if early indicators are favorable and to abandon if they are not. This reframing resolves a practical evaluation problem, which is that a pilot with a negative measured ROI may nonetheless have positive value if it resolved uncertainty cheaply. It also provides an explicit rationale for staged investment and for the small-scale randomized designs that generate the information the option requires. Portfolio framing addresses a second practical problem. Firms typically fund a bundle of interventions and then attempt to evaluate the bundle. Treating the bundle as a portfolio with distinct risk and return profiles, correlated exposures, and a defined evaluation budget makes explicit the trade-off between funding many small unevaluable interventions and funding fewer interventions at a scale that permits evaluation. The approach remained largely conceptual as of 2022, with limited empirical application. Its practical value lies less in producing a number than in disciplining the sequencing of investment and evaluation. 5. Adjacent measurement literatures The estimators reviewed in Section 4 were not developed for diversity and inclusion. They were adapted to it from fields that had already confronted the same underlying problem, which is the attribution of outcomes to investments in organizational systems where the outcome of interest is latent, lagged, unevenly observed, and jointly determined by many causes. That problem is not distinctive to D&I. It recurs wherever an organization spends money on a capability rather than on a product, and the applied literatures that have wrestled with it longest offer transferable apparatus rather than merely rhetorical analogy. This section reviews those adjacent literatures. The organizing logic is methodological rather than topical: each subsection identifies a measurement problem that a neighbouring field has formalized, and states what the formalization contributes to the evaluation of D&I investment. The breadth is deliberate. A field that borrows only from within its own subject boundary tends to reproduce the limitations of its own data, and much of what is weakest in current D&I measurement, particularly the reliance on single-point stock indicators and on post hoc attribution, has been solved elsewhere under a different name. 5.1 Safety climate Occupational safety management is the nearest structural relative of D&I measurement. Both concern a latent organizational property that is unobservable directly, both are measured through a mixture of administrative records and perception surveys, both suffer acute lagging-indicator problems because the terminal outcome is rare and severe, and both are subject to strong incentives to under-report. The safety literature resolved these problems earlier and more thoroughly, and the resolution is instructive. The central move was the shift from lagging to leading indicators. Counting incidents after they occur is the safety equivalent of counting representation after the pipeline has produced it: accurate, auditable, and almost useless for management, because by the time the number moves the causal window has closed. Work on near-miss and hazard observation systems developed the alternative, treating the volume and quality of proactively reported precursor events as the primary management signal (Arumosoye & Obriki, 2019; Obogo et al., 2021a; Obriki & Arumosoye, 2018). The analytical subtlety is that a rise in reported near misses is ambiguous: it may indicate deteriorating conditions or improving willingness to report. This is precisely the interpretive problem that self-identification disclosure rates present in D&I measurement, discussed in Section 7, and the safety field’s answer, which is to model reporting propensity separately from event incidence rather than to treat the raw count as a performance metric, transfers directly. A second contribution is the formalization of causation pathways rather than correlation with outcomes. Conceptual models of human error causation in high-risk work locate failure in the interaction of task design, supervision, fatigue, and organizational pressure rather than in individual carelessness (Obriki & Arumosoye, 2020; Obriki & Arumosoye, 2022; Obogo et al., 2020b). Applied to inclusion, the parallel claim is that exclusionary outcomes are produced by allocation processes and decision architecture rather than by individual attitude, which is consistent with the finding that attitude-focused training produces little behavioral change. Work on leadership-driven culture transformation and on the frequency of management walkthroughs as a predictor of downstream safety performance provides an operational measure of leader behavior that has no established D&I equivalent (Obogo et al., 2019a; Obriki & Arumosoye, 2019). The frequency and quality of senior leader engagement in inclusion processes is measurable in exactly the same way and is rarely measured at all. Third, the safety literature developed maturity models that stage organizational capability rather than scoring it on a single scale, which addresses the composite index problem identified in Section 3.2. Organizational learning-based maturity models treat continuous improvement capability as the object of assessment (Arumosoye & Obriki, 2021), while governance-oriented models separate the accountability structure from the outcomes it produces (Arumosoye & Obriki, 2020; Obogo et al., 2022; Obogo et al., 2019b). This separation matters for D&I because the strongest available evidence on intervention efficacy concerns accountability structures specifically, and a maturity model measures structure directly rather than inferring it from outcomes. Fourth, contractor and extended-workforce measurement is well developed in safety and almost absent in D&I. Contractor safety performance frameworks address the problem that the entity accountable for outcomes does not employ the people exposed to risk (Obogo et al., 2021c; Obogo et al., 2019c; Obogo et al., 2020c). The D&I analogue is the contingent, agency, and outsourced workforce, which is systematically excluded from corporate diversity reporting even where it constitutes a substantial share of the people doing the work, and whose exclusion mechanically inflates reported composition when outsourced functions are demographically distinct from the direct workforce. Finally, the safety field treats training as a component of a control system rather than as an intervention in itself, with fidelity and dosage tracked as separate variables (Obriki et al., 2022a; Obogo et al., 2020a), and it maintains explicit models for the rare high-consequence event, including emergency preparedness and evacuation performance (Obriki et al., 2022b), lifting and rigging risk pathways (Arumosoye & Obriki, 2022; Obogo et al., 2020b), heat stress exposure (Arumosoye & Obriki, 2018), incident prevention modelling (Obogo et al., 2021b), and the institutionalization of practices so that they survive personnel turnover (Obriki & Arumosoye, 2021). The last of these is the most underdeveloped idea in D&I practice. Institutionalization, meaning the degree to which a practice persists when its sponsor departs, is measurable through practice-persistence rates and is a better indicator of durable change than any composition figure. 5.2 Data governance, privacy, and the legal envelope of demographic measurement Every D&I metric is constrained by what may lawfully be collected, stored, processed, and disclosed. This constraint is treated in most D&I writing as a footnote and in the data governance literature as the primary design problem, which is the correct ordering. Comparative analysis of data protection regimes establishes that demographic attributes fall into the most heavily restricted category in most modern frameworks, requiring a specific lawful basis rather than general consent, and that the restrictions differ materially across jurisdictions in ways that defeat naive global aggregation (Mbonu et al., 2019a; Annan, 2022). The practical consequence for a multinational is that a single global diversity figure is not merely difficult to compute but, in some configurations, unlawful to compute, and the frameworks for legal and ethical handling of personal data developed in this literature supply the governance apparatus that D&I programmes generally lack (Mbonu et al., 2018). Data protection impact assessment is the specific instrument of interest. It requires an ex ante analysis of processing purpose, necessity, proportionality, and residual risk before a new data flow is created (Mbonu et al., 2022a). Applied to D&I analytics, particularly to the organizational network analysis discussed in Section 4.6 and to any passive behavioral collection, this instrument converts a diffuse privacy concern into a documented decision with an owner. Its adoption would materially improve the defensibility of people analytics programmes, which frequently proceed on the assumption that internal data is unproblematic because it is already held. Identity and access management determines who inside the organization can see demographic data at individual granularity, and it is the operational control that makes small-cell analysis possible without re-identification exposure (Mbonu et al., 2020c; Mbonu et al., 2022b). Related work on risk-based business continuity and on infrastructure- as-code governance addresses the reproducibility question that assurance requires, since a metric computed by an analyst on a laptop is not auditable in the sense the CSRD contemplates (Mbonu et al., 2021a; Mbonu et al., 2020b). Enterprise log analytics and automated monitoring provide the audit trail that demonstrates who accessed sensitive workforce data and when (Mbonu et al., 2019b), and forensic analytics methods extend the same logic to communications data (Mbonu et al., 2021c). The literature on artificial intelligence in security and governance contexts contributes a further caution. Machine learning applied to sensitive personal data introduces failure modes that are not present in conventional analysis, including inference of protected attributes that were never collected, model inversion, and adversarial manipulation (Mbonu et al., 2021b; Mbonu et al., 2022c; Borah et al., 2022). A D&I analytics programme that avoids collecting ethnicity but deploys a model whose features permit its reconstruction has not avoided the legal or ethical exposure. Privacy-preserving data architectures for regulated industries offer the constructive response, holding sensitive attributes in a governed enclave and exposing only aggregate statistics to analytical consumers (Atakpa & Abolaji, 2022). Agile supply chain data governance work extends the same architecture to third-party data flows (Mbonu et al., 2020a), which is directly relevant to supplier diversity measurement discussed in Section 5.7. 5.3 Audit, assurance, and internal control The most consequential regulatory development described in Section 4.5 was the extension of assurance to workforce metrics. Assurance is an audit concept, and the internal audit literature specifies what a metric must satisfy to be assurable. Risk-based internal control models establish the basic architecture, in which control activities are allocated according to assessed risk rather than uniformly, and in which the design of the control is documented separately from its operating effectiveness (Akomolafe & Agu, 2019a; Akomolafe & Agu, 2018). Applied to D&I, this distinction separates two questions that practitioners routinely conflate: whether the firm has a pay equity review process, and whether that process operates as designed. A disclosure that reports the former while implying the latter is precisely the kind of claim that assurance is designed to test. Data-driven risk evaluation models for financial institutions in emerging markets contribute the observation that control quality degrades where data quality is weakest, which is a systematic rather than random pattern (Akomolafe & Agu, 2019b). The D&I parallel is exact: demographic data quality is lowest in the jurisdictions and employment categories where inequity risk is highest, so measurement precision and substantive risk are inversely correlated. Work on integrated governance and compliance as a route to financial resilience makes the further point that fragmented control ownership is itself a risk factor (Akomolafe & Agu, 2019c), which bears on the common arrangement in which D&I metrics are owned by human resources, reported by corporate affairs, and assured by no one. Public sector accounting research contributes the budget compliance perspective, in which the question is not whether an outcome improved but whether funds were spent as authorized and reported as spent (Dogbatsey & Ebhojie, 2018). This is the Layer 1 question in the framework of Section 8, and it is the layer most often skipped. Reconciliation control and financial workflow redesign work formalizes the tracing of a committed amount through to its ultimate application (Dogbatsey et al., 2019), and cross-border segregation of duties addresses the multinational case in which authorization and execution occur in different legal entities (Dogbatsey et al., 2021). A D&I programme that cannot reconcile its stated commitment to its actual disbursement cannot support any return calculation, and pledge-tracking analyses in the 2020 to 2022 period repeatedly found this reconciliation absent. Fraud and forensic analytics contribute the adversarial perspective that metric design requires. Predictive models for anti-money laundering detection are built on the assumption that the measured party is actively optimizing against the measure (Atakpa, 2021), and machine learning approaches to fraud detection formalize the resulting arms race (Atakpa & Abetoh, 2022). Financial forensics techniques address detection after the fact (Ekechi et al., 2022). This literature supplies the analytical vocabulary for the gaming problem set out in Section 6: a metric should be assessed not only on whether it measures the construct but on how it behaves when a rational agent optimizes for it. Work on the relationship between control investment and organizational outcome, including systematic assessment of cybersecurity investment returns, models the same evaluation problem this paper addresses in a domain where the counterfactual is equally unobservable (Ozowara et al., 2022). Comprehensive financial reporting models for compliance strengthening close the loop by specifying the disclosure architecture that assured metrics require (Medon & Oduleye, 2022). 5.4 Attribution modelling in marketing analytics Marketing analytics has confronted the attribution problem more directly than any other applied field, because marketing spend is large, continuous, multi-channel, and evaluated by finance functions that demand a return figure. The parallel to D&I investment is close enough that the estimators transfer with minimal modification.The core problem is identical. Multiple simultaneous interventions touch an individual over an extended period, the outcome is observed only at the end, and naive attribution assigns the entire effect to whichever touchpoint was last or most easily observed. Systematic treatment of attribution modelling methods catalogues the responses, from rule-based heuristics through to algorithmic and incrementality-based approaches, and documents the bias that each introduces (Sanni et al., 2022b). The finding most relevant here is that last-touch and first-touch heuristics, which are the easiest to compute, are also the most biased, and that the direction of bias is systematic rather than random. The D&I equivalent is the attribution of a promotion outcome to the most recently completed development programme, which is the modal practice in programme evaluation and which shares the bias structure exactly. Analytical models addressing measurement challenges in marketing effectiveness formalize the distinction between measured lift and true incrementality (Sanni et al., 2020a), which is the same distinction as that between an observed change in representation and the change attributable to intervention. Predictive audience segmentation work contributes the heterogeneity insight, namely that an intervention effective in aggregate may be ineffective or counterproductive in identifiable subgroups (Atima et al., 2022), which corresponds directly to the heterogeneous treatment effect estimation discussed in Section 4.7 and to the documented finding that diversity training effects differ sharply by participant group. Business intelligence dashboard frameworks address the executive visibility problem, and the caution they raise is pertinent (Sanni & Atima, 2021b). A dashboard that surfaces the most easily computed metrics tends to entrench them, because visibility creates management attention and management attention creates optimization pressure. Analytics-driven go-to-market frameworks that incorporate compliance and sustainability constraints demonstrate that non-financial objectives can be integrated into a performance measurement system without being reduced to a financial proxy (Sanni & Atima, 2021a), which is the design problem for D&I metrics in a finance-owned reporting environment. Data-driven brand positioning work contributes the reputational measurement dimension (Sanni et al., 2020b), relevant because a substantial part of the claimed return to D&I investment is reputational and is therefore measured, if at all, through instruments developed in marketing. Systematic review of product management strategy contributes the staged decision structure (Sanni et al., 2020c), and adaptive control models contribute the closed-loop formulation in which measurement feeds back into intervention design within the period rather than only at its end (Sanni et al., 2022a). The closed-loop formulation is the practical answer to the lag problem identified in Section 7. 5.5 Financial analytics and the capital allocation frame If D&I expenditure is to be evaluated as investment, it must eventually be expressed in terms that a capital allocation process recognizes. The corporate financial analytics literature specifies those terms. Conceptual decision models for capital allocation using financial analytics establish the basic requirement, which is that competing uses of funds be expressed in a common unit with a stated time horizon and a stated risk profile (Lawal & Oduleye, 2021a). D&I proposals typically fail this test not because the return is absent but because it is expressed in a unit the allocation process cannot compare. Work on aligning financial planning analytics with corporate strategy addresses the bridging problem, in which strategic priorities that resist quantification must nonetheless enter a quantified planning process (Lawal & Oduleye, 2021b), and enterprise value creation models supply the linkage from operational metric to valuation (Lawal & Oduleye, 2018a). Data-driven executive decision systems contribute the observation that the binding constraint is usually decision latency rather than analytical sophistication (Lawal & Oduleye, 2019b). A D&I metric produced annually cannot inform a quarterly resource decision regardless of its validity, which argues for the higher-frequency process metrics described in Section 8 rather than the annual composition disclosure. Transfer pricing risk assessment and cross-border tax governance work contributes the multi-entity perspective (Lawal & Oduleye, 2019a; Lawal & Oduleye, 2018b), relevant because workforce metrics in a multinational are produced by legal entities with different data regimes and consolidated by a parent that must attest to the result. Treasury and strategic finance research extend this to the operating level. Multinational cash flow consolidation and foreign exchange risk work formalizes the consolidation of heterogeneous local measures into a group figure with explicit treatment of the translation assumptions (Dada et al., 2021a), which is the exact structural problem of consolidating jurisdictionally incompatible demographic data. Chief financial officer-led strategic finance models in joint venture contexts address governance where control is shared (Isiekwu et al., 2021), which bears on D&I accountability in joint ventures and franchise structures. Customer relationship management-based forecasting work contributes the pipeline projection apparatus (Dada et al., 2021b), structurally identical to the talent pipeline projection of Section 4.4. Smart contract automation in cross-border payment systems illustrates the more general point that verifiability can be engineered into a process rather than assessed after it (Akomolafe et al., 2022), which is the direction assurance requirements are pushing workforce reporting. 5.6 Predictive analytics, forecasting, and the targeting-evaluation distinction Section 4.7 distinguished predictive from causal machine learning. The applied forecasting literature makes the same distinction operationally and shows what each is good for. Advances in demand forecasting demonstrate that flexible learners substantially outperform classical methods on predictive accuracy while offering no causal interpretation of their inputs (Tonoyan et al., 2021a). Predictive cash flow modelling makes the same point in a financial context (Tonoyan et al., 2021b), and algorithmic process optimization work shows how accurate prediction supports resource allocation without any claim about mechanism (Tonoyan et al., 2022b; Tonoyan et al., 2022c). This is exactly the correct posture for attrition-risk modelling in D&I: use the model to decide where to look and where to intervene, and use a different apparatus entirely to decide whether the intervention worked. Geospatial predictive analytics contributes the spatial dimension (Tonoyan et al., 2022a), directly relevant to the shift-share instruments discussed in Section 4.2 and to labour market benchmarking in impact-weighted accounts, where the availability comparison must be constructed from the actual geography of a firm’s establishments rather than from a national aggregate. Audience segmentation and forecasting work extends the same methods to population subgroups (Basnet et al., 2021), and resilient logistics frameworks for humanitarian supply chains demonstrate forecasting under conditions of poor data quality and high stakes (Anene & Clement, 2022), which describes the intersectional small-cell problem well. The most important contribution is the human-in-the-loop literature, which specifies how automated inference and human judgment should be combined when the decision affects people (Ladapo et al., 2022a). The relevant design principle is that the model proposes and the human disposes, with the human’s override recorded and analyzed. In a D&I context this is not merely an ethical safeguard: the record of overrides is itself a dataset revealing where human judgment systematically departs from the model, and its analysis is a direct measurement of decision bias. 5.7 Supplier diversity and procurement as an external investment channel Supplier diversity is the largest D&I investment channel that receives the least measurement attention. Spend directed to diverse-owned suppliers is straightforward to record, is material in magnitude, and produces effects outside the firm boundary that are entirely absent from workforce composition metrics. The procurement literature supplies the measurement apparatus that supplier diversity programmes generally lack. Strategic procurement optimization frameworks establish the basic structure of category strategy, sourcing decision, and performance management (Okonkwo et al., 2018a; Akinleye & Adeyoyin, 2022b), within which a diversity objective is a constraint or a weighted criterion rather than a separate programme. This framing matters because supplier diversity managed as a separate initiative is structurally vulnerable to being overridden by the cost objective in the main process. Data-driven vendor evaluation models make the trade- off explicit by scoring suppliers on multiple weighted criteria simultaneously (Ahiaeke Patrick et al., 2021), which is the mechanism by which a diversity criterion can survive contact with a procurement decision. Supply chain risk management models contribute the concentration analysis (Agbabiaka et al., 2019), relevant because diverse-owned suppliers are frequently smaller and the risk assessment that governs their qualification is calibrated to larger firms, producing systematic exclusion at the qualification gate rather than at the award decision. Supply chain resilience frameworks show that supplier base diversification and resilience are complementary rather than competing objectives (Ogunwole et al., 2021), which is the strongest available business case argument for supplier diversity and is rarely made quantitatively. Regulatory-compliant procurement frameworks address the documentation burden that qualification imposes (Okonkwo et al., 2021b), itself a barrier that falls disproportionately on smaller suppliers. The operational literature contributes measurement of the downstream consequences. Materials requirement and inventory availability models link supplier performance to plant uptime (Okonkwo et al., 2021a; Okonkwo et al., 2018b), demurrage elimination models quantify the cost of supply chain failure (Okonkwo et al., 2020), and cost reduction frameworks formalize the savings attribution problem in a way that parallels the D&I attribution problem closely (Okonkwo et al., 2019). Geographic information system-enabled enterprise resource planning frameworks address the location dimension of supplier development (Ahiaeke Patrick et al., 2020), relevant to local content and community investment obligations that frequently accompany diversity commitments in extractive and infrastructure sectors. Process automation in procurement contributes the transaction-level data that measurement requires (Akinleye & Adeyoyin, 2021), and negotiation optimization models quantify the cost consequence of constraints applied to the supplier set (Akinleye & Adeyoyin, 2022a), which is the honest way to state the trade-off where one exists. Sustainable procurement frameworks integrate social criteria into the sourcing decision directly (Efobi et al., 2022a), data-driven lean supply chain models address the efficiency consequence (Efobi et al., 2022b), and performance management frameworks specify the ongoing measurement (Efobi et al., 2021). Visualization work using standard analytical tools demonstrates that the reporting requirement is not technically demanding (Akin-Oluyomi et al., 2020), which locates the obstacle in governance rather than capability. Decision frameworks reconciling chemical safety with circular economy targets illustrate the general structure of multi-objective sourcing decisions where the objectives are not commensurable (Okojie & Abioye, 2020), which is the situation a diversity criterion creates. 5.8 ESG auditing and sustainable procurement disclosure The convergence of D&I reporting with environmental, social, and governance disclosure means that D&I metrics increasingly inherit the assurance and audit expectations developed for ESG data generally. Predictive approaches to supply chain and ESG auditing demonstrate the shift from sample-based to population-based testing that data availability permits (Aifuwa et al., 2020), which is directly applicable to pay equity testing where full population analysis is feasible and sampling is not necessary. Real-time risk assessment dashboards illustrate the continuous monitoring model that replaces periodic assessment (Filani et al., 2022), and work on ESG data quality and supplier disclosure addresses the upstream data reliability problem that any social metric spanning the value chain must confront (Ike et al., 2021; Ike et al., 2022). The last point is material for D&I because supplier workforce composition, where it is claimed at all, rests on self-reported data with no assurance and no consistent definition. Sustainable procurement governance work addresses the accountability structure for value chain social objectives (Nnabueze et al., 2021), and programmatic strategy for renewable energy integration in large infrastructure contributes the observation that social and local content objectives in major projects require measurement systems designed at the point of project structuring rather than retrofitted (Yeboah & Ike, 2020). The retrofitting problem is the single most common cause of unmeasurable D&I commitments: the commitment is announced, and only afterward is anyone asked how it will be evidenced. 5.9 Process standardization and lean operations Inclusion outcomes are produced by processes, and process design is the domain of operations management. The lean literature contributes the specific insight that variability and discretion in a process are the principal sources of inconsistent outcomes, which is the operational statement of the bias mechanism located in discretionary allocation decisions. Standard operating procedures treated as strategic instruments rather than as documentation formalize the reduction of unwarranted discretion (Eyetsemitan et al., 2022). Applied to hiring and promotion, this is the mechanism behind structured interviewing and calibrated review, and the operations literature supplies the measurement apparatus, in the form of process conformance rates and variability statistics, that these practices generally lack in D&I application. Lean Six Sigma applied in small enterprises demonstrates that the apparatus scales down to organizations without dedicated analytical functions (Ambali et al., 2021), which addresses the common objection that rigorous D&I measurement is available only to large firms. Cycle time reduction in client onboarding illustrates the flow-based measurement approach at a granularity that D&I measurement rarely reaches (Oyeleye et al., 2022). The analogue is time-to- promotion measured as a distribution rather than a mean, disaggregated by group, which reveals disparities that promotion rate ratios conceal. Multi-stakeholder governance alignment work addresses the coordination problem where process ownership is distributed (Eyetsemitan et al., 2020), and the translation of regulatory requirements into operating procedure addresses the compliance interface (Eyetsemitan et al., 2021), which is the mechanism by which a disclosure obligation becomes an actual change in how decisions are made rather than an additional reporting task. 5.10 Human capital formation and the upstream pipeline Firms frequently attribute weak representation to pipeline supply. The claim is testable and the education literature supplies the evidence base for testing it. Comparative analysis of educational attainment and access identifies where cohort attrition actually occurs in the formation pipeline (Boakye et al., 2021), which is the necessary input to any availability benchmark. A firm claiming supply constraint is making an implicit assertion about this distribution, and the assertion is frequently wrong at the relevant seniority and geography. Integration of information technology and mathematical pedagogy addresses the quantitative skills pipeline specifically (Boakye et al., 2020), relevant to technical role availability. Special education research contributes the most methodologically developed treatment of individualized measurement in this literature. Data-driven individualized education programme development formalizes the construction of individual-level goals with explicit measurement plans and periodic review against them (Yeboah & Oloto, 2022), which is a more rigorous individual development apparatus than most corporate talent processes employ. Compliance frameworks under disability education legislation demonstrate how a legal accommodation obligation is operationalized into documented, auditable practice (Oloto & Yeboah, 2022), which is directly transferable to workplace accommodation, an area of D&I practice that is almost entirely unmeasured despite carrying direct legal exposure. Pedagogical strategies for students with specific learning needs contribute the intervention design (Bobga et al., 2018), and student progress monitoring supplies the measurement cadence (Yeboah et al., 2022). Teacher professional development and competency research addresses the capability of the people delivering the intervention (Yeboah et al., 2019), which corresponds to manager capability in the D&I context and is a stronger predictor of outcomes than programme design. Parental collaboration research demonstrates measurement of a stakeholder relationship that resists direct observation (Ogbona et al., 2020b), and coaching and career transition work addresses the mobility pathways that structure advancement (Ogbona et al., 2020a). Cross-cultural communication research contributes the measurement of interaction quality across difference (Lilian et al., 2020), which is the construct that inclusion scales attempt to capture at the organizational level. Data- informed learning platform frameworks and real-time educational analytics contribute the instrumentation (Orise & Niniola, 2020; Orise & Niniola, 2022), demonstrating continuous rather than periodic measurement of a developmental process. 5.11 Health equity and access measurement Health services research has the most developed literature on equity measurement of any applied field, because access disparities are directly observable in utilization data and the outcome stakes are unambiguous. Infrastructure-driven expansion of diagnostic access demonstrates the measurement of an equity outcome through changes in the population able to reach a service (Aminu-Ibrahim et al., 2020), which is a genuine access metric rather than a proxy. The structural insight is that access is determined by facility siting, capacity, and operating model rather than by stated policy, and that measuring the policy tells you nothing about the access. The D&I parallel is the difference between having an inclusive policy and having a process people can actually use. Capital project delivery models for high-risk health facilities address the investment evaluation question in a setting where the return is partly social (Aminu-Ibrahim et al., 2019), and sustainable diagnostic laboratory infrastructure work addresses the durability of that investment (Aminu- Ibrahim et al., 2018). Laboratory spatial planning, regulatory-compliant design for molecular facilities, sustainable materials and energy efficiency, risk-managed construction, and design standards with operational planning collectively demonstrate how a social access objective is decomposed into engineering and operational specifications that are individually measurable (Ogbete et al., 2018; Ogbete et al., 2019; Ogbete et al., 2020; Ogbete et al., 2021; Ogbete et al., 2022). This decomposition is the discipline that D&I commitments most conspicuously lack. Health system financing research contributes the coverage and provider payment perspective (Ilodigwe & Adesemoye, 2019a), and the conceptual bridging from molecular science to health system outcomes illustrates the multi-level attribution problem in a field that has confronted it explicitly (Ilodigwe & Adesemoye, 2019b). Pharmacy-led delivery and access models demonstrate the measurement of a service redesign intended to reach underserved populations (Ilodigwe & Adesemoye, 2020), and governance and analytics work on the health system data backbone addresses the infrastructure prerequisite (Ilodigwe & Adesemoye, 2021), which corresponds to the human resource information system prerequisite for D&I analytics. Last-mile distribution research addresses the specific problem of reaching populations that standard delivery models systematically miss (Asiedu & Asiedu, 2022), a structure that recurs in workforce contexts wherever a programme designed for the centre fails to reach peripheral or non- standard employment groups. Drug interaction risk frameworks in antenatal care illustrate systematic risk assessment for a population defined by a protected characteristic (Asiedu, 2022). Drug shortage early warning systems and regulatory readiness maturity models contribute the anticipatory measurement approach (Eze et al., 2022a; Eze et al., 2022b), and integration of biomarker and imaging parameters for cardiovascular risk stratification demonstrates the construction of a composite index with validated components and a defined outcome, which is the standard that D&I composite indices should be held to and are not (Okwah, 2022). 5.12 Energy justice and distributional metrics The energy transition literature has developed explicit distributional measurement because the transition imposes costs and confers benefits unevenly, and because the unevenness is politically salient. This is the closest available model for measuring the distributive dimension of a corporate investment. Co-optimization of technical performance and community affordability formalizes the trade-off between an efficiency objective and a distributional objective within a single optimization rather than treating the distributional criterion as a constraint applied afterward (Komi & Adeniji, 2020). This is the methodological move that D&I evaluation most needs. A firm that optimizes for cost and then applies a diversity constraint will reach a different and generally worse solution than one that includes the distributional objective in the optimization directly. Wind power forecasting work that connects grid stability to market equity demonstrates that a technical performance metric and an equity metric can be derived from the same model rather than from separate systems (Komi, 2021). Interpretable machine learning for failure prediction addresses the explainability requirement that arises whenever a model output affects a protected interest (Komi & Adamolekun, 2021), which is the same requirement that governs algorithmic hiring and promotion tools. Systematic review of control architectures contributes the centralization question (Komi & Adamolekun, 2022), relevant to whether D&I accountability should sit centrally or with business units, a question usually decided by organizational preference rather than by evidence. Battery storage research in weak-grid regions addresses measurement where infrastructure quality is itself unequal (Komi & Ganiu, 2022), and smart grid architecture work under high renewable penetration contributes the system integration perspective (Adenuga, 2022). 5.13 Condition monitoring and the leading-indicator problem The engineering literature on condition monitoring addresses the same problem as leading indicators in safety, with a fully developed instrumentation apparatus, and it clarifies what a leading indicator actually is. The transition from reactive to predictive maintenance is the canonical case (Sunday et al., 2020). Reactive maintenance responds to failure, which is a lagging indicator; predictive maintenance monitors condition variables whose degradation precedes failure, and the analytical work consists of establishing which observable variables carry that predictive content. Smart fault detection using sensor-based monitoring demonstrates the instrumentation (Sunday & Omoegun, 2022), and thermodynamic efficiency and control strategy work shows how performance degradation is measured continuously rather than at failure (Sunday et al., 2019). The lesson for D&I measurement is that a leading indicator is not simply an earlier measurement of the same thing; it is a different variable, validated against the terminal outcome. Corporate D&I practice frequently mislabels an earlier composition measurement as a leading indicator, which it is not. Non-destructive testing research contributes the inspection design question of how much to inspect, where, and how often, given that inspection is costly and imperfect (Atta et al., 2021; Atta et al., 2020). This is the sampling design problem for pay equity review and for promotion decision audit, where full population review is expensive and risk- weighted sampling is the practical alternative. Welding inspection standards and quality assurance work formalizes the qualification of the inspection process itself (Atta et al., 2022), which corresponds to the validation of an audit method, and green corrosion inhibitor research illustrates the substitution problem in which an intervention chosen on one criterion must be evaluated against others (Atta et al., 2018). Solar integration and hybrid load distribution work contributes the design of measurement in systems with variable inputs (Sunday & Omoegun, 2018; Sunday & Omoegun, 2019). Remote sensing research contributes the coverage dimension. Unmanned aerial vehicle applications in transmission line inspection demonstrate the substitution of comprehensive automated observation for sampled manual observation (Dagodzo, 2018b; Dagodzo, 2018a), pipeline and corridor monitoring extends this to linear assets (Dagodzo & Ahiaeke Patrick, 2020), and integrated frameworks combining aerial survey, laser scanning, and geographic information systems demonstrate multi-source data fusion (Dagodzo & Ahiaeke Patrick, 2021b). Geographic information system applications in asset management contribute the spatial data architecture (Dagodzo & Ahiaeke Patrick, 2021a), encroachment detection illustrates automated exception identification (Dagodzo & Ahiaeke Patrick, 2022), and deep learning applied to vegetation classification demonstrates automated pattern recognition at scale (Dagodzo et al., 2022). The organizational network analysis of Section 4.6 is structurally the same move: comprehensive passive observation replacing sampled active inquiry, with the same gains in coverage and the same governance obligations. 5.14 Regulatory compliance benchmarking and oversight design Comply-or-explain disclosure regimes of the kind described in Section 4.2 rest on assumptions about how organizations respond to oversight. The regulatory compliance literature examines those assumptions. Predictive compliance monitoring frameworks demonstrate that inspection findings across an oversight population contain systemic signal that individual findings do not (Adeyelu, 2018), which is an argument for regulators to analyze the pattern of disclosure across firms rather than assessing each disclosure in isolation. Adaptive safety management system implementation models address the capability gap in resource-constrained settings (Adeyelu, 2019), which bears on the disproportionate compliance burden that standardized disclosure places on smaller firms. Comparative benchmarking models for certification compliance formalize cross- entity comparison where local conditions differ (Adeyelu & Dagodzo, 2022), the central difficulty in cross-jurisdictional D&I comparison. Risk quantification models for low-probability high-consequence events demonstrate estimation where the outcome is too rare to support direct measurement (Adeyelu, 2020), which is the situation for severe discrimination and harassment outcomes, where incidence is low, reporting is incomplete, and the standard response is to report the count without modelling the reporting process. Frameworks for managing encroachments and unauthorized structures illustrate enforcement against gradual boundary erosion (Adeyelu, 2021; Adeyelu, 2022), a structure that recurs in the slow drift of an organization away from a stated commitment. Cybersecurity research contributes the adversarial and architectural perspective. Security audit and enterprise risk assessment frameworks specify the assessment structure (Dosunmu & Ogundele, 2019), intrusion detection and prevention models address continuous monitoring (Dosunmu & Ogundele, 2020), incident response and forensics address the post-event process (Dosunmu & Ogundele, 2021), and threat intelligence integration demonstrates the use of external signal in internal risk assessment (Dosunmu & Ogundele, 2022). Operational technology security research contributes the convergence problem, in which two systems with different governance histories must be assessed under one framework (Adegbite et al., 2022; Adegbite et al., 2020), which describes the integration of human resource and financial reporting systems that assured workforce disclosure requires. Zero trust architecture contributes the design principle that access should be verified continuously rather than granted once (Ahmed et al., 2021), which applied to demographic data governance means that authorization to view sensitive workforce data should be evaluated per access rather than per role. Adversarial machine learning research establishes that a model deployed in a contested environment will be attacked through its inputs (Adebayo et al., 2022), which is the formal statement of the gaming problem for any metric that becomes consequential. Lifecycle risk assessment in offshore production contributes the long-horizon evaluation structure (Falegan & Aniebonam, 2022), relevant to the lag problem of Section 7. 5.15 Enterprise information systems as the data substrate Every claim in this paper about what firms should measure presumes a system of record capable of producing the measurement reproducibly. That presumption is frequently false, and the enterprise systems literature explains why. Platform governance frameworks establish the controls that keep a configurable business system consistent as it evolves (Badmus et al., 2019b), and version control and deployment pipeline research demonstrate the reproducibility apparatus (Badmus et al., 2022a; Badmus et al., 2018). The relevance is direct: a workforce metric is reproducible only if the definitions embedded in the reporting system are themselves versioned, and most human resource reporting environments cannot state which definition produced a figure published two years earlier. Extract, transform, and load design work addresses the integration of records from multiple sources into a consistent analytical layer (Badmus et al., 2019a), which is the practical obstacle to linking demographic, compensation, performance, and exit data into the person-period dataset that the estimators of Section 4 require. Comparative deployment framework analysis (Badmus et al., 2020) and enterprise service bus integration research (Badmus et al., 2022b) address the multi-system case. Directory services and identity infrastructure work supplies the authoritative person record on which all workforce analysis depends (Ladapo et al., 2019), and digital transformation practice research addresses the organizational change required to make such systems authoritative rather than nominal (Ladapo et al., 2022b). Offline assessment research contributes the observation that measurement systems must function where connectivity and infrastructure are limited (Ladapo et al., 2018), relevant to global workforce measurement in operations outside corporate centres. Capacity planning and resource utilization forecasting demonstrate the operational forecasting apparatus (Edivri & Oteri, 2022), and service delivery and security operations research addresses the sustained operation of the systems that produce the data (Ogbole et al., 2021; Okoruwa et al., 2020). 5.16 Rural and agricultural inclusion measurement Agricultural economics has measured inclusion in production systems for longer than corporate D&I has existed as a field, and under harder data conditions. Gender inclusion and equity across agricultural value chains is measured through participation, control of assets, and returns at each stage rather than through headcount at any single point (Michael & Ogunsola, 2022b). This value chain decomposition is precisely the flow-based logic of Section 4.4 applied to an economic rather than an organizational hierarchy, and it produces the same analytical benefit, which is the identification of the specific stage at which participation diverges. Access to finance research contributes the credit constraint analysis (Michael & Ogunsola, 2019a), a structural barrier with no direct workplace analogue but with a close parallel in access to sponsorship and to stretch assignments, which function as the internal capital of a career. Socioeconomic barriers to technology adoption among smallholders demonstrate the measurement of differential uptake of an ostensibly universal intervention (Michael & Ogunsola, 2022a), which is the dosage and coverage problem at Layer 2 of the framework in Section 8. Extension services research addresses the intermediated delivery of a capability (Michael & Ogunsola, 2021e), structurally identical to manager-mediated delivery of a workplace programme, where the intermediary’s own capability and incentives determine reach. Climate-smart practice research links adoption to household outcomes over multi-year horizons (Michael & Ogunsola, 2021d), an example of the long-horizon outcome measurement that D&I evaluation requires and rarely attempts. Data-driven policy models and policy alignment analysis contribute the framework linking intervention to system-level outcome (Michael & Ogunsola, 2021c; Michael & Ogunsola, 2021a), digital agriculture research addresses inclusive system design (Michael & Ogunsola, 2021b), and agribusiness education research addresses the skills pipeline (Michael & Ogunsola, 2019b). Related work in agricultural engineering and animal science illustrates the underlying measurement discipline. Factor analysis applied to machine performance demonstrates the decomposition of an aggregate outcome into weighted contributing variables (Bello et al., 2022b), the same technique underlying the confirmatory factor analysis of Section 4.3, and performance evaluation of purpose-built equipment illustrates the specification of measurement criteria before construction rather than after (Bello et al., 2022a). Controlled feeding trials with graded treatment levels demonstrate dose-response estimation under experimental control (Aye & Tawose, 2015; Aye & Tawose, 2016), which is the design that organizational research can rarely achieve and should approximate where it can. Phenotypic characterization across regions demonstrates systematic measurement of population heterogeneity (Ekeocha et al., 2021), the methodological ancestor of the dispersion indices discussed in Section 3.1. 5.17 Multi-objective weighting and the composite index problem The composite index problem identified in Section 3.2 is formally a multi-objective optimization problem, and the engineering optimization literature treats it as such. Multi-algorithm formulations of single-objective and multi-objective functions for electrical machine design make the weighting question explicit, showing that the optimal design depends entirely on the relative weight assigned to loss, weight, and cost objectives, and that no weighting is technically privileged (Amayo, 2017; Amayo & Popoola, 2017a; Amayo et al., 2015). The parallel to composite D&I indices is exact and unflattering. A provider that aggregates representation, policy existence, and disclosure completeness into one score has chosen a weighting, that weighting determines the ranking, and the choice is neither disclosed nor defended. The engineering treatment at least states the objective function. Voltage control research under distributed generation demonstrates system management where control is dispersed among many actors with individual objectives (Oshevire et al., 2017), which describes a devolved D&I accountability model. Telecommunications deployment research on service quality contributes the measurement of a service attribute experienced individually but delivered systemically (Amayo & Popoola, 2017b), which is the structural form of inclusion as a construct. 5.18 Policy alignment and institutional frameworks Corporate D&I commitments increasingly interact with national and supranational policy frameworks, and the policy alignment literature addresses the resulting coordination problem. Policy alignment modelling for climate diplomacy demonstrates the reconciliation of domestic objectives with international commitments where the two are measured differently (Liadi, 2022a), the structure that a multinational face in reconciling group-level commitments with jurisdictionally specific legal constraints. Continental integration frameworks address measurement across member entities with heterogeneous capacity (Liadi, 2022b), and peacebuilding effectiveness frameworks confront outcome measurement where the outcome is contested, long-horizon, and confounded by factors outside any actor’s control (Liadi, 2022c). The last of these is methodologically closest to D&I impact evaluation, and its central lesson is that effectiveness frameworks in such domains are most credible when they measure process fidelity and intermediate outcomes explicitly rather than claiming terminal effects they cannot identify. 5.19 Recurring methodological commitments Six things transfer from these literatures to D&I evaluation. The first is the leading indicator discipline of Sections 5.1 and 5.13, which requires that an early metric be a different variable validated against the terminal outcome rather than an earlier measurement of the same one. The second is the attribution apparatus of Section 5.4, which supplies estimators for multi-touch, multi- period interventions. The third is the assurance architecture of Sections 5.2, 5.3, and 5.15, which specifies what a metric must satisfy to be audited and what systems must exist to produce it. The fourth is the distributional optimization of Section 5.12, which includes the equity objective in the optimization rather than applying it as an afterward constraint. The fifth is the flow decomposition of Sections 5.9, 5.11, and 5.16, which locates disparity at a specific stage rather than reporting it in aggregate. The sixth is the adversarial framing of Sections 5.3 and 5.14, which assesses a metric by how it behaves when a rational agent optimizes against it. None of these is present in the disclosure frameworks that standardized D&I measurement between 2020 and 2022. That absence, rather than any shortage of statistical technique, is the substance of the field’s measurement problem. 6. Criteria for assessing metric quality Table 2 sets out criteria against which a D&I metric can be assessed. They are drawn from measurement theory, from audit practice, and from the failure modes documented in the preceding sections. Table 2. Evaluation criteria for D&I metrics Criterion Question it answers Common failure Construct validity Does the metric measure the construct it names? Representation percentages presented as measures of inclusion Measurement invariance Does it mean the same thing across the groups compared? Cross-group survey mean comparison without invariance testing Reliability Would repeated measurement give the same result? Single-item measures; small-cell estimates with wide confidence intervals Sensitivity to intervention Can it move within the evaluation horizon in response to action? Board composition as a programme metric, when board turnover is a multi- year process Attributability Can change in it be linked to a specific investment? Aggregate representation trends credited to a training programme Resistance to gaming Does optimizing the metric advance the goal? Junior-heavy hiring inflating aggregate representation while senior attrition worsens Assurability and privacy compliance Is it reproducible from auditable systems within legal constraints? Metrics requiring demographic categories that cannot lawfully be collected in some jurisdictions Resistance to gaming deserves particular emphasis. Goodhart’s law, in its familiar formulation, holds that when a measure becomes a target it ceases to be a good measure. D&I metrics are unusually exposed to this dynamic because the most visible metrics are the most easily manipulated. Aggregate workforce representation responds fastest to high-volume junior hiring, which is also the least expensive lever and the least indicative of equitable process. A firm can improve its headline number for several consecutive years while its promotion and senior attrition differentials deteriorate. Any evaluation framework that does not report flows alongside stocks is structurally vulnerable to this substitution. 7. Constraints on measurement Endogeneity. The core inferential obstacle is that firms select into D&I investment. Profitable, growing, well-governed firms in tight labor markets invest more and also perform better for unrelated reasons. Without exogenous variation, no amount of statistical sophistication recovers a causal estimate. Simultaneity and reverse causation. Inclusion perceptions and performance are jointly determined. Employees in successful units report better climate partly because success is pleasant. Panel data with unit fixed effects and lagged specifications mitigate this only partially. Long and variable lags. Composition responds to process change on a horizon measured in years, since it is the integral of flow changes. Evaluation windows of one to two years, which match budget cycles, are structurally mismatched to the phenomenon and systematically understate effects. This mismatch creates a bias toward funding interventions with fast-moving proxy metrics. Small cells and intersectionality. Intersectional analysis is where the substantive action lies and where statistical power is weakest. A cell containing a small number of employees produces estimates with confidence intervals too wide to support inference and cannot be published without re-identification risk. Suppression thresholds, typically requiring a minimum cell size before reporting, remove precisely the observations of greatest interest. Hierarchical Bayesian models with partial pooling across cells offer partial relief by borrowing strength across related groups, but this remained rare in corporate practice through 2022. Cross-jurisdictional data constraints. A multinational cannot construct a globally consistent metric. France restricts collection of ethnic and racial statistics on constitutional grounds. Germany’s data protection regime and works council structures constrain what may be collected and how it may be analyzed. Many jurisdictions have no analogue to the EEO-1 categories, which are in any case specific to United States social history and travel poorly. GDPR classifies data revealing racial or ethnic origin as a special category requiring a specific lawful basis. The practical consequence is that global aggregate diversity figures published by multinationals are typically constructed from inconsistent underlying definitions and heavy imputation, a fact rarely disclosed. Self-identification data quality. Voluntary self-identification produces non-random missingness. Disclosure rates for disability, sexual orientation, and gender identity are strongly conditional on perceived psychological safety, meaning that disclosure rates are themselves an inclusion signal and simultaneously a source of bias in every metric that depends on them. Rising measured representation may reflect rising disclosure willingness rather than changing composition, and these two interpretations have opposite implications for programme evaluation. Rater and instrument bias. Performance ratings, which serve as controls in most pay and promotion models, are themselves potentially biased. Conditioning on a biased mediator biases the estimated direct effect. This point has been made repeatedly in the context of expert testimony on workplace bias, and it remains the most consequential specification error in applied corporate pay equity work. The disclosure-validity paradox. The metrics that regulators mandate and index providers standardize are the demographic stock measures, because they are objectively verifiable and cross- firm comparable. The metrics with the strongest documented association with firm outcomes are perception-based and process-based measures, which are neither easily verified nor comparable across firms. Standardization has therefore concentrated on the least informative measures. Resolving this tension is the central open problem in the field. 8. Matching estimators to evaluation layers The preceding sections imply that no single model answers the evaluation question. What is required is an explicit chain in which each layer of the investment logic is matched to an estimator capable of identifying quantities at that layer, with the identifying assumptions stated at each step. Table 3 sets out the correspondence. Table 3. Correspondence between evaluation layers, estimators, and identifying assumptions Layer Question Primary estimator Identifying assumption Illustrative metric 1. Input What was committed? Activity-based costing including opportunity cost of participant and leader time Complete cost capture, including uncosted internal time Fully loaded cost per programme, per participant, per business unit 2. Activity What was delivered, to whom? Coverage and dosage analysis, treatment fidelity assessment Accurate exposure records Percentage of eligible population exposed; dosage distribution; delivery fidelity index 3. Proximal outcome Did experience and behavior change? Validated latent scales with measurement invariance; ONA structural metrics; pre- post designs with comparison units Invariance holds; comparison units satisfy parallel trends Inclusion scale scores by group; cross-group tie density; sponsorship access 4. Process outcome Did allocation become more equitable? Transition matrix estimation by group; Cox hazard models; Oaxaca-Blinder decomposition; causal ML on high- dimensional controls Unconfoundedness given the control set; controls are not themselves biased mediators Promotion rate ratio by group and grade; conditional attrition hazard ratio; adjusted and unadjusted pay gap 5. Terminal outcome What value was created or preserved? Monetized cost avoidance via utility analysis; impact- weighted employment accounts; quasi- experimental estimates where exogenous variation exists Effect size at Layer 4 is causally identified; monetization parameters are defensible Avoided replacement cost; monetized opportunity and diversity impact; steady-state composition projection versus target Three properties of this arrangement bear emphasis. First, each layer is independently informative, so that a firm unable to identify effects at Layer 5 can still establish at Layer 4 that its promotion rate ratio moved, which is a materially stronger claim than any representation percentage. Second, claims propagate upward only with stated assumptions: a Layer 5 monetary figure is valid conditional on a Layer 4 causal estimate, and reporting a return figure without the corresponding Layer 4 identification is the most common error in practitioner evaluation. Third, joint reporting constrains gaming, because a requirement to report flow metrics alongside any composition figure renders the junior-hiring substitution described in Section 6 immediately detectable. The arrangement carries a resource implication, since it is more demanding than a dashboard. The appropriate response is portfolio concentration, in the sense of Section 4.9: fund fewer interventions at sufficient scale to permit staggered rollout and credible comparison, rather than many interventions that can only ever be described. 9. Discussion For practice. The dominant corporate metric set as of 2022 measured the wrong layer. Firms should report flow metrics, specifically promotion rate ratios and conditional attrition hazards by group and grade, alongside every composition figure. Where interventions are deployed across multiple units, staggered rollout should be adopted by default, since the marginal cost is low and the inferential gain is large. Pay equity analysis should report adjusted and unadjusted gaps together, with the control set disclosed and justified. Inclusion surveys should undergo measurement invariance testing before any cross-group comparison is published. For investors and rating providers. Composite scores that aggregate disclosure quantity, demographic composition, and policy existence into one number have low construct validity and weak convergent validit