Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

The Diagnostic Reagent Value Index: A Two-Tier-Cost Composite Model for Evaluating Diagnostic Reagent Performance and Cost Efficiency in Hospital Procurement

Habeeb Damilola Yusuf, Kazeem Abdulrazaq, Sylviastella Favour Peteranaba, Helen, Ekwi Osinem, Florence Eribenne

Abstract

Hospital procurement teams routinely select in vitro diagnostic reagents on the basis of quoted unit price, a metric that captures only a fraction of the economic consequences of the decision and none of the clinical ones. Multi-criteria frameworks developed for medical device purchasing are either too general to accommodate the specific failure modes of consumable reagents or too demanding for a procurement committee working to a tender deadline. This paper develops the Diagnostic Reagent Value Index, a decision instrument that combines non-compensatory eligibility screening, a weighted performance score across five domains normalised against externally anchored acceptability limits, a two-tier cost model, and mandatory rank-stability testing. Costs are expressed as cost per reportable result rather than cost per test, which corrects for quality control consumption, calibration burden, open-vial waste, and repeat testing. Downstream consequence costs are attenuated by explicit propagation coefficients representing the fraction of analytical misclassifications that survive clinical filtering to reach a decision. The model was applied to a parameterised procurement of a high-sensitivity cardiac troponin I assay at 60,000 tests per year across three candidate suppliers with quoted prices spanning $2.55 to $4.10. The lowest-priced candidate carried the highest hidden-cost multiplier, 1.48 against 1.26 for the highest-priced candidate, because short onboard stability, a three-level quality control regimen, fortnightly recalibration, and a 3.2 percent repeat rate consumed 14.75 percent of purchased reagent on non-reportable use. Within the laboratory budget frame, the three candidates were nearly indistinguishable, with index values of 82.7, 80.3, and 79.3, which helps explain why price-led tenders appear defensible. At institutional scope the ranking reversed decisively in favour of the highest-priced candidate, 91.2 against 78.9 and 61.2, which returned $13.70 in avoided downstream cost for every additional dollar of laboratory spend. Rank reversal required propagation coefficients to fall to approximately 6 percent of base assumptions, and the lowest-priced candidate could not match the leader at institutional scope at any non-negative unit price. Unit price is therefore a poor proxy for the economic value of a diagnostic reagent when the consequence-to-direct-cost ratio is large, as in this high-consequence troponin case; a second, low-consequence application of the same model to a glycated haemoglobin tender shows the opposite pattern, with the price-based rule identifying the better candidate and the institutional- scope index costing $30,261 a year more, so it is the ratio between the two spreads, not the assay category, that determines when the correction changes the answer. The paper's recommendation is accordingly a low-cost screening step that uses this ratio to decide which tenders warrant full scoring, not a blanket replacement of price-based award. Where the correction does matter, the principal obstacle to acting on it is organisational rather than analytical, since the benefits accrue to budgets the laboratory does not control.

Keywords

diagnostic reagents; laboratory procurement; total cost of ownership; cost per reportable result; value-based procurement; multi-criteria decision analysis; in vitro diagnostics

References

G6 Supply continuity Documented business continuity provision, stated lead time, and confirmed shelf life on receipt Contractual undertaking Screens G3 and G4 depend on locally adopted analytical performance specifications, which should be set in advance from an accepted hierarchy, whether biological variation data, outcome-based criteria for the specific measurand, or state-of-the-art criteria, and recorded in the tender file. 3.2. Performance scoring Five domains are used, chosen to be exhaustive at the level of decision-relevant distinctions while remaining few enough to weight coherently. Analytical performance, at a default weight of 0.30, covers total imprecision at decision concentrations expressed against the allowable limit, trueness and method comparison bias against the incumbent or reference procedure, diagnostic sensitivity and specificity at the operative cut-off, limit of quantitation relative to clinical need, documented interference and cross-reactivity, and linearity across the measuring interval. Operational fit, at 0.25, covers time to first result and throughput against peak workload, onboard and open-vial stability, calibration frequency and hands-on time, the quality control architecture required, reflex and rerun handling, consumable and waste burden, and training requirement. Supply chain resilience, at 0.20, covers historical or contracted order fill rate, lead time and its variance, single versus multiple manufacturing sites, remaining shelf life guaranteed on delivery, lot allocation and reservation terms, geopolitical and logistics exposure, and substitution options if supply fails; these indicators are drawn from the sustainable and resilient procurement literature (Ekwunife et al., 2024; Filani et al., 2022, 2023; Ike et al., 2021, 2022, 2024a, 2024b). Quality and regulatory assurance, at 0.15, covers lot-to-lot consistency, external quality assessment performance, field safety notice and recall history, completeness of verification documentation, and post-market surveillance responsiveness. Supplier and lifecycle support, at 0.10, covers contractually backed service response times, applications support, product roadmap and discontinuation notice period, contract flexibility on volume variation, and implementation resource offered. Weights should be elicited from a panel including the laboratory director, the responsible scientist for the discipline, a clinical user of the assay, procurement, and finance, and fixed and documented before bid evaluation. Where the assay is used in a time-critical pathway the weight on operational fit will typically rise; where supply has previously failed, resilience will rise. Each indicator is converted to a sub-score on a 0 to 100 scale using anchored rather than within-tender normalisation. For a benefit-direction indicator with observed value x, a minimum acceptable anchor a, and an aspirational anchor b, and for a cost-direction indicator such as imprecision or time to result, the two expressions are: s = 100 × min[ 1 , max( 0 , (x − a) / (b − a) ) ] s = 100 × min[ 1 , max( 0 , (a − x) / (a − b) ) ] The anchors are absolute standards set before the tender opens. This matter, because under within- tender normalisation the best of three poor candidates scores 100, which flatters a weak field and makes scores incomparable between procurement cycles. Anchored normalisation preserves the meaning of a score across tenders and permits a legitimate finding that no candidate is good enough. The domain sub-score is the weighted mean of its indicator scores, and the composite is the weighted sum across domains: PCS = Σd=1..5 wd Sd , with Σd wd = 1 3.3. Cost model The elementary error in reagent costing is to multiply quoted unit price by expected patient test volume, which ignores every use of reagent that does not produce a reportable patient result. Let N be the annual number of reportable patient results required, r the repeat and rerun fraction, Q annual quality control test consumption, K annual calibration test consumption, and ω the fraction of loaded reagent lost to onboard expiry, open-vial instability, and pack changeover. Reagent consumption T, and quality control consumption for L control levels, R runs per day, and M analysers, are then: T = [ N (1 + r) + Q + K ] / (1 − ω) Q = L × R × M × 365 Direct laboratory cost follows, with p the negotiated unit price, E the number of calibration events and ccal their material cost, cQC the annual cost of control materials, H the annual technologist hours attributable to calibration and quality control review, τ the fully loaded hourly labour rate, and csvc the apportioned instrument, service, and maintenance cost attributable to the assay: Cdirect = T p + E ccal + cQC + H τ + csvc CPRR1 = Cdirect / N Analytical performance differences produce different numbers of misclassified patients. For a target condition of prevalence π among tested patients and an assay with sensitivity Se and specificity Sp, the annual counts of false negatives and false positives are N π (1 − Se) and N (1 − π) (1 − Sp) respectively. Not every misclassified result reaches a consequential decision, since serial sampling, confirmatory testing, clinical assessment, and physician override filter a substantial proportion. This is represented by propagation coefficients κFN and κFP in the interval from zero to one, giving the fraction of misclassifications that survive to alter management: Ccons = κFN FN cFN + κFP FP cFP Ctotal = Cdirect + Ccons CPRR2 = Ctotal / N The propagation coefficients depend on how results are actually used at the bedside and in the clinical pathway, and should therefore be informed by the service delivery and clinical literatures rather than set by the laboratory alone (Akinlolu et al., 2022, 2023, 2024; Fapohunda et al., 2023a, 2023b, 2023c, 2024; Igweonu et al., 2024). They are the most contestable parameters in the model and should be treated as such: they are not estimated with precision but declared, justified from the local care pathway, and then stress-tested. Their function is not to produce an exact figure but to prevent the consequence term from being either ignored entirely, which is the status quo, or applied without attenuation, which would overstate it. Two points on scope follow. The unit consequence costs should be expected values conditioned on the local pathway rather than worst-case figures, since a false negative does not always produce an adverse event and the coefficient should reflect the probability-weighted cost of the range of outcomes including any liability provision the institution actually carries. And where a false positive reflects a true elevation of the measurand arising from a condition other than the target, the misclassification is not attributable to the reagent, so only the differential between candidates should be treated as material. 3.4. Integration and robustness testing Cost is normalised reciprocally against the best performer in the field, and combined with the performance score under a trade-off parameter α with a default of 0.6: Ĉi = 100 × ( minj CPRRj ) / CPRRi DRVIi = α PCSi + (1 − α) Ĉi The reciprocal form is preferred to a linear minimum-maximum because it preserves proportionality: a candidate costing twice the cheapest scores 50, which is interpretable, whereas a linear form makes the score depend on the range of the field. The choice of α is a policy statement about the institution's willingness to pay for quality and should be set and recorded before bids are evaluated. A secondary statistic, cost per performance point, gives an efficiency reading independent of the field and is the more useful figure for tracking value across successive procurement cycles: CPPi = CPRRi / ( PCSi / 100 ) Four robustness analyses are then run before any recommendation is finalised. The trade-off sweep recomputes the index across α from 0.4 to 0.8 and reports whether the leading candidate changes. The propagation stress test scales both coefficients from zero to twice the base assumption and reports the multiplier at which the ranking reverses, so that where reversal requires an implausible value the recommendation is robust irrespective of disagreement about the coefficients. Volume sensitivity recomputes at plausible low and high demand, since fixed cost elements amortise differently. Weight perturbation varies each domain weight by plus or minus 0.10 with proportional rebalancing and records rank stability. Where the deterministic runs leave the outcome close, a probabilistic analysis should also be run, sampling the uncertain inputs jointly and reporting the proportion of draws in which each candidate leads, since one-at-a-time testing understates the combined effect of several parameters moving together and yields no measure of confidence. A recommendation that reverses within the plausible range of any parameter should be reported as such rather than presented as determinate, since honest reporting of indeterminacy is more useful to a decision committee than false precision and considerably more defensible under challenge. A reader auditing the scoring and cost formulas together will notice that several indicators enter both. Onboard and open-vial stability and calibration frequency and hands-on time are scored inside the Operational fit domain in Section 3.2 and also feed the waste fraction ω, the calibration count E, and the labour hours H that determine Cdirect and CPRR1 in Section 3.3; diagnostic sensitivity and specificity at the operative cut-off are scored inside the Analytical performance domain and also determine the false negative and false positive counts that drive Ccons and CPRR2. This overlap is real and is not symmetrical across indicators, since most PCS indicators, such as regulatory clearance history, lot-to-lot consistency, and lifecycle support, have no cost-model counterpart and are counted once, while a small number of operational and analytical indicators are counted twice. It is disclosed here rather than removed by restructuring the index, because the PCS sub-score and the corresponding cost term price different aspects of the same underlying attribute: the monetised terms capture the technologist hours, reagent waste, and misclassification cost that a given level of stability, calibration burden, or accuracy produces, while the PCS sub-score captures qualitative aspects the monetised terms do not, such as staff confidence in a predictable calibration schedule, the operational slack a long onboard life provides against unplanned workload, and a general standard of analytical trustworthiness a committee may reasonably value beyond the specific misclassification counts priced through Ccons. A committee that regards this double counting as unacceptable for a particular indicator can remove it from the domain sub-score, or down-weight the domain concerned, without altering the cost formulas, provided the change is fixed and documented before bids are opened. 3.5. Setting the consequence parameters Section 3.3 requires the propagation coefficients and unit consequence costs to be declared and justified, which invites the objection that a committee can set them wherever it likes and obtain whatever award it prefers. The objection is serious and the safeguard against it is procedural rather than mathematical: the coefficients are decomposed into steps that correspond to identifiable points in the local care pathway, each step is assigned a value by the clinicians who work at that point, and the decomposition is recorded before bids are opened. A single number pulled from the air is indefensible. A product of four stated conditional probabilities, each attributable to a named part of the pathway, can be argued with, corrected, and audited. For the false negative coefficient, the decomposition follows the sequence of filters that stand between an incorrect result and an adverse outcome. A missed elevation is first caught or not caught by the serial sampling protocol, then by clinical assessment and other investigations, then by the decision actually taken on the basis of the assembled evidence. Each filter is a conditional probability of the error surviving. For the false positive coefficient, the sequence is shorter, since a spurious elevation typically triggers investigation rather than being suppressed. Table 2 gives the decomposition used in the first application; the specific survival fractions shown are illustrative values constructed for this worked demonstration, not figures drawn from a real clinical audit. Table 2. Decomposition of the propagation coefficients used in the first application. Filter Question put to the clinical panel Survives Running value False negatives Serial sampling What fraction of missed elevations are detected on the repeat sample at the protocol interval? 0.30 0.30 Clinical assessment Of those still missed, what fraction are nonetheless investigated on clinical grounds? 0.50 0.15 Decision taken Of those not investigated, what fraction are discharged rather than held on other grounds? 0.80 0.12 False positives Confirmatory testing What fraction of spurious elevations are resolved by the repeat sample before action is taken? 0.60 0.60 Clinical context Of those not resolved, what fraction proceed to additional workup rather than being attributed to another cause? 0.50 0.30 Presented this way, the coefficient of 0.12 used in Section 3 is not an assumption but a claim that serial sampling catches seven missed elevations in ten, that half of the remainder are investigated anyway, and that four in five of what is left are discharged. Each of those three statements is contestable by someone who works in the department, and each can be revised without touching the rest of the model. A panel that disagrees can substitute its own values and rerun in seconds. What it cannot do is leave the number unexamined, which is the practice the decomposition is designed to displace. The unit consequence costs require a different treatment, since they are resource quantities rather than probabilities. They should be assembled as expected values across the distribution of outcomes that actually follows the error, not as the cost of the worst outcome, and the components should be drawn from sources the institution already maintains. For a false negative these typically comprise the incremental cost of later presentation at a more advanced stage, any additional length of stay, the difference in intervention cost between early and late treatment, and a probability-weighted allowance for the liability provision the institution carries. For a false positive they comprise the confirmatory workup, any observation or admission time attributable to the incorrect result, and the opportunity cost of the capacity consumed. Costs that the institution does not bear should be excluded from the institutional scope and, where the analysis is being prepared for a payer or a health system rather than a single hospital, stated separately rather than folded in. Three practices make the resulting figures more defensible. Values should be rounded generously, since precision beyond two significant figures implies knowledge the exercise does not possess and invites disputes that the robustness stage renders irrelevant. The panel should record its reasoning alongside the values, so that a subsequent tender can reuse or revise the parameters rather than starting again. And the same coefficients should be applied to every candidate, since the model relies on differences between candidates rather than on the absolute level of consequence cost, and a coefficient applied evenly cannot favour a particular supplier however it is set. 3.6. Setting and inputs The application below uses synthetic but internally consistent data constructed to represent a plausible procurement. It demonstrates the model's mechanics and its capacity to produce rank reversal; it is not evidence about any real product. A 600-bed acute hospital re-tenders its high-sensitivity cardiac troponin I assay. Annual demand is 60,000 reportable results across two analysers running two quality control cycles per day. Prevalence of the target condition among tested patients is 8 percent. The fully loaded technologist rate is $42 per hour. Expected downstream costs are set at $18,000 per propagated false negative and $900 per propagated false positive. Propagation coefficients are set at κFN = 0.12, reflecting the strong filtering effect of serial sampling and clinical assessment, and κFP = 0.30, reflecting the weaker filtering of a positive result that triggers further investigation. Three candidates pass all eligibility screens. Decision thresholds and the definition of the target condition follow the universal definition of myocardial infarction (Thygesen et al., 2018). Candidate characteristics are given in Table 3. Table 3. Candidate reagent characteristics. Parameter Vendor A (incumbent) Vendor B (premium) Vendor C (low cost) Quoted unit price $3.40 $4.10 $2.55 Onboard and open-vial loss 2.5% 1.2% 5.5% Calibration events per year 24 12 52 Calibration material cost per event $190 $210 $150 Quality control levels 2 2 3 Annual control material cost $3,100 $3,400 $4,200 Repeat and rerun rate 1.6% 0.9% 3.2% Apportioned service cost $28,000 $31,000 $22,000 Sensitivity 0.94 0.97 0.90 Specificity 0.91 0.93 0.88 Imprecision at cut-off (CV) 4.2% 3.1% 6.8% Time to first result 15 min 12 min 18 min 3.7. Performance scores and direct cost Domain sub-scores were derived from verification data, supplier documentation, and reference site enquiry, and are shown in Table 4. Vendor B's analytical and operational advantages are partly offset by weaker supply resilience arising from single-site manufacture, which is the kind of trade- off a checklist assessment tends to lose. Table 4. Domain sub-scores and Performance Composite Score. Domain Weight A B C Analytical performance 0.30 78 92 62 Operational fit 0.25 80 88 61 Supply chain resilience 0.20 85 72 68 Quality and regulatory 0.15 88 90 74 Supplier and lifecycle 0.10 82 78 70 Performance Composite Score 1.00 81.80 85.30 65.55 Reagent consumption and direct cost are given in Table 5. Table 6 sets the resulting figures against the naive calculation of price multiplied by volume. Table 5. Annual reagent consumption and direct laboratory cost. Component A B C Patient tests including repeats 60,960 60,540 61,920 Quality control tests 2,920 2,920 4,380 Calibration tests 96 48 208 Tests lost to instability and waste 1,640 771 3,871 Total reagent units consumed 65,616 64,279 70,379 Reagent cost $223,096 $263,545 $179,466 Calibration material $4,560 $2,520 $7,800 Control material $3,100 $3,400 $4,200 Attributable labour $10,556 (251 h) $10,388 (247 h) $12,992 (309 h) Service apportionment $28,000 $31,000 $22,000 Direct annual cost $269,312 $310,853 $226,458 CPRR, laboratory scope $4.49 $5.18 $3.77 Table 6. Quoted price against realised direct cost. Metric A B C Naive annual cost (price × volume) $204,000 $246,000 $153,000 Actual direct cost $269,312 $310,853 $226,458 Hidden cost multiplier 1.320 1.264 1.480 Reportable yield per unit consumed 91.4% 93.3% 85.3% Reagent consumed on non-reportable use 8.6% 6.7% 14.7% The cheapest reagent carries the largest multiplier. Vendor C's price advantage over Vendor A is 25 percent at the quotation stage and 16 percent at the direct cost stage. Roughly a third of the apparent saving is consumed by fortnightly recalibration, a three-level control architecture, a fourteen-day onboard stability limit, and a repeat rate twice the incumbents. 3.8. Institutional cost and the integrated index Table 7 adds the propagated consequence term. The consequence cost exceeds the direct cost in every case, but by a factor that varies substantially across candidates rather than by a uniform “roughly a factor of seven”: approximately 7.3x for Vendor A ($1,963,440 against $269,312), 4.4x for Vendor B ($1,354,320 against $310,853), and 12.5x for Vendor C ($2,825,280 against $226,458). That range is itself evidence for the candidate-dependent argument developed in Section 3.13, since it shows the consequence term does not scale direct cost uniformly but tracks each candidate's own analytical performance. None of this is an artefact of aggressive assumptions, since the propagation coefficients used are deliberately conservative and setting both to zero simply collapses the institutional scope back into the laboratory scope. Table 7. Propagated misclassification cost and total institutional cost. Component A B C False negatives per year 288 144 480 False positives per year 4,968 3,864 6,624 Propagated consequence cost $1,963,440 $1,354,320 $2,825,280 Consequence cost per reportable result $32.72 $22.57 $47.09 Total annual cost $2,232,752 $1,665,173 $3,051,738 CPRR, institutional scope $37.21 $27.75 $50.86 Relative to Vendor B, Vendor A saves $41,542 in direct laboratory cost and incurs $609,120 in additional propagated consequence cost, while Vendor C saves $84,395 in direct cost and incurs $1,470,960 in additional consequence cost. Table 8 gives the integrated index at both scope boundaries. Table 8. Integrated index at both cost scopes. Cand. PCS CPRR1 Ĉ1 DRVI CPP1 CPRR2 Ĉ2 DRVI (inst.) CPP2 A 81.80 $4.49 84.1 82.72 $5.49 $37.21 74.6 78.91 $45.49 B 85.30 $5.18 72.9 80.32 $6.07 $27.75 100.0 91.18 $32.54 C 65.55 $3.77 100.0 79.33 $5.76 $50.86 54.6 61.16 $77.59 This is the central result. At laboratory scope the three candidates fall within 3.4 index points of each other and the incumbent narrowly leads. A committee working in that frame would find the field effectively indistinguishable and would reasonably resolve the tie on price, awarding to Vendor C. At institutional scope the ordering is unambiguous and the spread is 30 index points. Awarding to Vendor B rather than Vendor A costs the laboratory an additional $41,542 per year and saves the institution $567,578, a return of $13.70 for every additional dollar of laboratory spend. 3.9. Robustness Vendor B leads at every value of α from 0.4 to 0.8 at institutional scope, with index values falling from 94.1 to 88.2 against Vendor A's 77.5 to 80.4, so the ranking is insensitive to the performance- cost trade-off setting. The propagation stress test is given in Table 9. Table 9. Propagation stress test, institutional cost per reportable result. κ multiplier A B C Leader 0.00 $4.49 $5.18 $3.77 C 0.25 $12.67 $10.82 $15.55 B 0.50 $20.85 $16.47 $27.32 B 1.00 $37.21 $27.75 $50.86 B 1.50 $53.57 $39.04 $74.41 B 2.00 $69.94 $50.32 $97.95 B Vendor B overtakes on total cost at a propagation multiplier of 0.068 against Vendor A and 0.057 against Vendor C. If as little as 6 percent of the assumed propagation actually occurs, corresponding to a false negative coefficient of roughly 0.008, the accuracy advantage already outweighs the price premium. The recommendation is therefore robust to very substantial disagreement about the consequence parameters, which is the practically important property, since the committee need not agree on the exact value of a missed diagnosis, only that it is not close to zero. Vendor C would need to offer a unit price of $3.75 to match Vendor B at laboratory scope. At institutional scope there is no non-negative price at which Vendor C matches Vendor B, the required figure being approximately negative $17 per test, so an analytical performance deficit of this size cannot be purchased away. Volume sensitivity is given in Table 10; rank order is stable across a factor of eight in volume under both scopes, with fixed elements amortising as expected but not altering the recommendation. Table 10. Volume sensitivity, cost per reportable result (laboratory / institutional). Annual volume A B C 15,000 $7.33 / $40.05 $8.16 / $30.73 $6.74 / $53.83 30,000 $5.43 / $38.16 $6.17 / $28.75 $4.76 / $51.85 60,000 $4.49 / $37.21 $5.18 / $27.75 $3.77 / $50.86 120,000 $4.02 / $36.74 $4.68 / $27.26 $3.28 / $50.37 3.10. Weight perturbation The domain weights are the most visible discretionary element of the model and the element most likely to be challenged by an unsuccessful bidder, since a losing candidate can always argue that a different weighting would have produced a different award. The perturbation test answers that argument directly rather than defending the chosen weights on their merits. Each domain weight in turn was moved by plus and minus 0.10, with the remaining four weights rebalanced proportionally so that the set continues to sum to one, and the full index was recomputed at both scope boundaries. Ten perturbations were run in total, and the results are given in Table 11. Table 11. Weight perturbation, plus or minus 0.10 with proportional rebalancing. Weight varied Change PCS (A / B / C) Index, laboratory (A / B / C) Lead Index, institutional (A / B / C) Lead Analytical performance +0.10 81.3 / 86.3 / 65.0 82.4 / 80.9 / 79.0 A 78.6 / 91.8 / 60.9 B Analytical performance −0.10 82.3 / 84.3 / 66.1 83.0 / 79.7 / 79.6 A 79.2 / 90.6 / 61.5 B Operational fit +0.10 81.6 / 85.7 / 64.9 82.6 / 80.5 / 79.0 A 78.8 / 91.4 / 60.8 B Operational fit −0.10 82.0 / 84.9 / 66.2 82.9 / 80.1 / 79.7 A 79.1 / 91.0 / 61.5 B Supply resilience +0.10 82.2 / 83.6 / 65.9 83.0 / 79.3 / 79.5 A 79.2 / 90.2 / 61.3 B Supply resilience −0.10 81.4 / 87.0 / 65.2 82.5 / 81.3 / 79.1 A 78.7 / 92.2 / 61.0 B Quality and regulatory +0.10 82.5 / 85.9 / 66.5 83.2 / 80.7 / 79.9 A 79.3 / 91.5 / 61.8 B Quality and regulatory −0.10 81.1 / 84.7 / 64.6 82.3 / 80.0 / 78.7 A 78.5 / 90.8 / 60.6 B Supplier and lifecycle +0.10 81.8 / 84.5 / 66.0 82.7 / 79.8 / 79.6 A 78.9 / 90.7 / 61.5 B Supplier and lifecycle −0.10 81.8 / 86.1 / 65.1 82.7 / 80.8 / 79.0 A 78.9 / 91.7 / 60.9 B The leading candidate does not change under any perturbation at either scope. At institutional scope Vendor B leads across a range of 90.2 to 92.2 index points while the nearest rival never exceeds 79.3, so the margin is roughly ten times the largest movement a single weight can produce. At laboratory scope the incumbent leads throughout, and its margin over Vendor C remains wide (3.5 points, 83.0 against 79.5) even when supply resilience is weighted more heavily; what narrows to 0.2 index points in that same row is the gap between Vendor B and Vendor C (79.3 against 79.5), with Vendor C fractionally overtaking Vendor B, which is a further indication that the laboratory-scope comparison is not resolving a real difference between the candidates. The asymmetry between the two scopes is itself informative. Weights govern only the performance half of the index, so their influence is bounded by α, and where the cost term separates candidates strongly the weighting debate becomes secondary. This is worth stating explicitly to an evaluation panel, since committees often spend most of their deliberation time on weights and very little on the cost model, when the sensitivity runs the other way. 3.11. Probabilistic sensitivity analysis The deterministic tests vary one parameter at a time, which understates the combined effect of several uncertain inputs moving together and gives no measure of how confident the recommendation is. A Monte Carlo analysis of 20,000 draws was therefore run over the parameters that a procurement team cannot pin down exactly. Sensitivity and specificity were drawn from normal distributions centred on the verification estimates with standard deviations of 0.015 and 0.010 respectively, reflecting the precision obtainable from a realistic verification sample. Waste fractions were scaled by a uniform factor between 0.75 and 1.25 and repeat rates by a uniform factor between 0.70 and 1.30, representing site-to-site variation of the kind discussed in Section 2.1. The propagation coefficients were drawn uniformly from 0.02 to 0.30 for false negatives and 0.10 to 0.50 for false positives, and the unit consequence costs from $8,000 to $30,000 and $400 to $1,600. Prices were held fixed, since they are contractual. Results are given in Table 12. Table 12. Monte Carlo analysis of the first application, 20,000 draws. Candidate Leads at laboratory scope Leads at institutional scope Lowest total institutional cost A (incumbent) 100.0% 2.2% 4.2% B (premium) 0.0% 97.8% 95.8% C (low cost) 0.0% 0.0% 0.0% Vendor B leads at institutional scope in 97.8 percent of draws and produces the lowest total cost in 95.8 percent, despite the wide intervals placed on the consequence parameters. The residual 2.2 percent consists of draws in which the propagation coefficients sit near the bottom of their ranges and Vendor A's sampled sensitivity happens to exceed its point estimate while Vendor B's falls below. Vendor C never leads at either scope and never produces the lowest total cost in any of the 20,000 draws, which is a stronger statement than the deterministic break-even analysis alone supports.The laboratory-scope column is worth reading alongside the institutional column rather than on its own. Vendor A leads there in every single draw, not because it is a better product but because the cost differences at that scope are small, stable, and driven almost entirely by contractual prices that were not varied. A rule that produces the same answer every time is not necessarily identifying a robust winner; it may simply be measuring a quantity on which the candidates barely differ. Reporting both columns prevents that stability from being mistaken for confidence. 3.12. A second application: a low-consequence assay The first application was chosen because troponin is a high-consequence assay in which the model's central argument is most visible. A method that only ever recommends the more expensive product would be of little use, so the model was applied to a deliberately different case: glycated haemoglobin, where the assay is high-volume, the decision threshold is not time-critical, results are routinely repeated at intervals, and the cost of acting on a wrong result is an order of magnitude smaller. The same hospital tenders its HbA1c assay at 90,000 reportable results per year across two analysers. Prevalence of the target condition among tested patients is 28 percent, reflecting a mixed screening and monitoring population. Expected downstream costs are set at $1,200 per propagated false negative, representing delayed treatment intensification, and $180 per propagated false positive, representing an unnecessary repeat and a counselling episode. Propagation coefficients are set at 0.10 and 0.20, lower than in the first case because HbA1c results are habitually confirmed on a subsequent sample. Three candidates pass eligibility screening, and their characteristics are given in Table 13. The analytical spread between them is deliberately narrow, as it typically is for a mature, standardised measurand. Table 13. Second application: candidate characteristics. Parameter Vendor P (incumbent) Vendor Q (premium) Vendor R (low cost) Quoted unit price $1.85 $2.30 $1.35 Onboard and open-vial loss 2.0% 1.0% 4.5% Calibration events per year 12 6 26 Calibration material cost per event $140 $165 $110 Quality control levels 2 2 3 Annual control material cost $2,600 $2,900 $3,400 Repeat and rerun rate 1.2% 0.7% 2.6% Apportioned service cost $24,000 $27,000 $19,000 Sensitivity 0.960 0.965 0.955 Specificity 0.940 0.945 0.935 Imprecision at decision point (CV) 2.4% 1.6% 3.6% This case is also used to illustrate how anchored normalisation converts raw indicator values into sub-scores, a step that was reported only in summary form for the first application. The analytical performance domain is built from four indicators: imprecision at the decision point, weighted 0.30 and normalised in the cost direction against a minimum acceptable anchor of 6.0 percent and an aspirational anchor of 1.5 percent; sensitivity, weighted 0.25 against anchors of 0.90 and 0.99; specificity, weighted 0.25 against anchors of 0.88 and 0.98; and a rubric score covering interference, cross-reactivity, and linearity, weighted 0.20. The anchors were set from the locally adopted performance specification before bids were opened. Table 14 shows the derivation. Table 14. Anchored normalisation of the analytical performance domain, second application. Indicator Weight Anchors (min / aspirational) P value → score Q value → score R value → score Imprecision (CV) 0.30 6.0% / 1.5% 2.4% → 80.0 1.6% → 97.8 3.6% → 53.3 Sensitivity 0.25 0.90 / 0.99 0.960 → 66.7 0.965 → 72.2 0.955 → 61.1 Specificity 0.25 0.88 / 0.98 0.940 → 60.0 0.945 → 65.0 0.935 → 55.0 Interference and linearity 0.20 Rubric 90 95 85 Domain sub-score 1.00 73.67 82.64 62.03 Two features of the derivation are worth noting. The imprecision indicator separates the candidates far more than sensitivity or specificity do, because the anchors on the accuracy indicators are wide relative to the observed spread, which is the correct behaviour: when three products are all comfortably inside the acceptable range on a criterion, that criterion should not drive the award. And Vendor R scores 53.3 on imprecision rather than zero, because anchoring is against an absolute acceptability limit rather than against the best candidate in the field. Under within-tender normalisation Vendor R would have scored zero on this indicator and Vendor Q one hundred, exaggerating a difference of two percentage points of coefficient of variation into the full width of the scale. Combining the derived analytical sub-score with the four remaining domains gives the composite scores in Table 15, and the cost analysis follows in Table 16. Table 15. Second application: domain sub-scores and Performance Composite Score. Domain Weight P Q R Analytical performance 0.30 73.67 82.64 62.03 Operational fit 0.25 82 90 64 Supply chain resilience 0.20 84 76 70 Quality and regulatory 0.15 86 89 76 Supplier and lifecycle 0.10 80 82 72 Performance Composite Score 1.00 80.30 84.04 67.21 Table 16. Second application: consumption, direct cost, and institutional cost. Component P Q R Patient tests including repeats 91,080 90,630 92,340 Component P Q R Quality control tests 2,920 2,920 4,380 Calibration tests 48 24 104 Tests lost to instability and waste 1,919 945 4,562 Total reagent units consumed 95,967 94,519 101,386 Reagent cost $177,540 $217,394 $136,872 Calibration material $1,680 $990 $2,860 Control material $2,600 $2,900 $3,400 Attributable labour $10,388 (247 h) $10,304 (245 h) $12,628 (301 h) Service apportionment $24,000 $27,000 $19,000 Direct annual cost $216,208 $258,588 $174,760 CPRR, laboratory scope $2.40 $2.87 $1.94 Hidden cost multiplier 1.299 1.249 1.438 False negatives / false positives per year 1,008 / 3,888 882 / 3,564 1,134 / 4,212 Propagated consequence cost $260,928 $234,144 $287,712 Total annual cost $477,136 $492,732 $462,472 CPRR, institutional scope $5.30 $5.47 $5.14 The pattern that held in the first application partly repeats and partly breaks down. The hidden cost multiplier again runs highest for the cheapest reagent, at 1.438 against 1.249 for the most expensive, and Vendor R again loses the largest share of purchased reagent to non-reportable use, at 11.2 percent against 4.8 percent for Vendor Q. The consumption correction is therefore a general finding rather than an artefact of the first case. But the consequence term behaves completely differently. It is $2.90 to $3.20 per reportable result rather than $22 to $47, it is of the same order as direct cost rather than seven times larger, and the spread between candidates is $0.30 rather than $23. Adding it changes the cost ranking hardly at all: Vendor R is cheapest at both scopes, and its margin over Vendor P narrows only from $0.46 to $0.16 per result. Table 17. Second application: integrated index at both scopes. Cand. PCS CPRR1 Ĉ1 DRVI CPP1 CPRR2 Ĉ2 DRVI (inst.) CPP2 P 80.30 $2.40 80.8 80.51 $2.99 $5.30 96.9 86.95 $6.60 Q 84.04 $2.87 67.6 77.46 $3.42 $5.47 93.9 87.97 $6.51 Cand. PCS CPRR1 Ĉ1 DRVI CPP1 CPRR2 Ĉ2 DRVI (inst.) CPP2 R 67.21 $1.94 100.0 80.32 $2.89 $5.14 100.0 80.32 $7.65 The result is indeterminate, and the model is required to say so. At laboratory scope Vendor P leads Vendor R by 0.19 index points, a margin far inside any reasonable estimate of input error. At institutional scope Vendor Q leads Vendor P by 1.02 points while Vendor R, which is genuinely the cheapest on total cost, ranks last because its performance deficit outweighs a cost advantage that the reciprocal normalisation compresses into a three-point range. The trade-off sweep confirms the instability: the leader at institutional scope is Vendor P at α = 0.4 and Vendor Q at every higher value, with the two separated by 0.4 index points at the crossover. The Monte Carlo analysis, run with narrower parameter ranges appropriate to this assay, gives Vendor P the institutional-scope lead in 35.4 percent of draws and Vendor Q in 63.8 percent, while the lowest total cost is achieved by Vendor R in 40.6 percent, Vendor P in 31.4 percent, and Vendor Q in 28.0 percent. No candidate commands a majority on both criteria. Presented honestly, this is a useful output rather than a failure. It tells the committee that the three products are close enough that the index cannot separate them, that the choice can legitimately be made on grounds the model does not capture, such as the disruption cost of changing platform or the value of consolidating suppliers, and that a price negotiation is likely to produce more value than further analysis. The alternative, reporting a point-estimate ranking of 87.97 against 86.95 as though it were a finding, would give a spurious mandate to a decision the data do not support. 3.13. Award rules compared Both applications can be used to ask what four candidate award rules would have chosen, and what each choice would have cost the institution. The comparison is given in Table 18, with regret measured as the difference between the total institutional cost of the chosen candidate and that of the best available candidate. Table 18. Award rules applied to both cases, with institutional cost and regret. Award rule Case 1 pick Case 1 cost Case 1 regret Case 2 pick Case 2 cost Case 2 regret Lowest unit price C $3,051,738 $1,386,565 R $462,472 $0 Lowest direct cost per result C $3,051,738 $1,386,565 R $462,472 $0 Index, laboratory scope A $2,232,752 $567,578 P $477,136 $14,664 Index, institutional scope B $1,665,173 $0 Q $492,732 $30,261 The two cases point in opposite directions, and reporting only the first would misrepresent the model. In the high-consequence case the institutional-scope index is the only rule that identifies the best candidate, and the two price-based rules are wrong by $1.39 million a year, which is more than five times the entire annual reagent budget for the assay. In the low-consequence case the ordering inverts: the price-based rules happen to select the best candidate, the laboratory-scope index costs $14,664 more, and the institutional-scope index costs $30,261 more. The model can select worse than lowest price when the candidates are genuinely close. What separates the two cases is not the assay but the ratio between the spread in consequence cost and the spread in direct cost across the candidate field. In the first case that ratio is roughly 15 to 1, and the consequence term dominates any plausible weighting. In the second it is roughly 0.6 to 1, so the consequence term cannot overturn the cost ranking and the index differences reduce to the performance score, which is measured with far less precision than cost. A useful screening step, computable before any scoring is done, is therefore to calculate the two spreads from the eligibility-stage data and proceed to full scoring only where the consequence spread is the larger. Where it is not, the honest recommendation is to verify that all candidates meet the analytical specification, award on cost per reportable result, and negotiate on price. Table 8 and Table 17 report cost per performance point at both scopes, and the figures are worth reading directly rather than left as an unused column. In the troponin application, institutional- scope CPP is $45.49 for Vendor A, $32.54 for Vendor B, and $77.59 for Vendor C: Vendor B is not only the leading candidate in this particular field, it is also the most efficient in absolute terms, converting each performance point into the lowest institutional cost of the three. This is a different statement from the DRVI ranking, because Ĉ is normalised reciprocally against whichever candidate is cheapest in the current tender, so a reagent with unchanged characteristics can score a different DRVI in a later tender simply because the field of bidders changed; CPP does not move for that reason, which is what makes it, rather than DRVI, the figure a committee should carry forward to compare against the next procurement cycle. The glycated haemoglobin application shows the complementary case. Vendor R has the lowest CPRR at both scopes, yet its institutional CPP of $7.65 is the highest of the three candidates, above Vendor P's $6.60 and Vendor Q's $6.51, because R's low performance score means its low absolute cost buys a comparatively poor return per performance point. Read alongside CPRR and PCS separately, a committee might still be drawn to Vendor R's low headline cost; CPP states directly that this cost is not cheap once performance is taken into account, without requiring the field-relative normalisation that makes DRVI itself dependent on who else bid. 4. Discussion and Conclusion The hidden cost multipliers of 1.26 to 1.48 confirm that unit price systematically understates reagent cost, and that it understates it unevenly across suppliers in a way that can invert an apparent price ranking. This alone justifies replacing cost per test with cost per reportable result in tender evaluation, and the required data are readily obtainable: onboard stability and calibration frequency are in the package insert, control architecture follows from laboratory policy, and repeat rates are recoverable from the incumbent's laboratory information system and obtainable from reference sites for others. But this correction, on its own, changed nothing in the worked example. Vendor C remained cheapest at laboratory scope, because the operational cost adjustments narrowed the gap without closing it. Procurement teams that adopt total cost of ownership methods and stop there will have improved the precision of their analysis without improving the quality of their decisions. What reversed the ranking was the consequence term, and its magnitude relative to direct cost is the finding with the widest applicability. For a high-consequence assay, the economic weight of a result is dominated by what happens when it is wrong, and the reagent price is close to a rounding error in the total. The savings from selecting the better-performing reagent surface as avoided admissions, avoided investigations, avoided readmissions, and avoided claims, distributed across clinical directorates and, in some cases, across financial years, rather than in the laboratory's own account. The laboratory that awards to Vendor B has increased its own costs by $41,542 and will be asked to explain why, an asymmetry the transfer arrangements set out below are designed to correct. Comparable misalignments between where cost is incurred and where benefit is realised are well documented in health financing and value-based contracting research (Ilodigwe & Adesemoye, 2019a, 2019b, 2020, 2021, 2024a, 2024b). The two-scope structure is designed to make this visible rather than to resolve it. Resolution is a matter of financial governance, not analysis, and requires one of a small number of specific institutional arrangements: evaluation of laboratory tenders at institutional scope with explicit sign-off from the finance function; a ring-fenced quality budget within the laboratory allocation, funded from the savings the index identifies; or an internal transfer mechanism recognising avoided downstream cost, such as a shared-savings agreement with the clinical directorates that avoid the admission or investigation, or a diagnosis-related-group-linked efficiency dividend returned to the laboratory budget. Where none exists, the model should still be run at both scopes and the divergence reported to the board, since documenting the gap is a precondition for closing it. On implementation, the model is intended to be run by a small group over a period of weeks rather than by a health economist over months. Before the tender opens, the evaluation panel fixes the screens, the domain weights, the normalisation anchors, the value of α, and the propagation coefficients and consequence costs, and records them in the tender file. This ordering is essential for defensibility, since parameters chosen after bids are seen will not survive challenge. During evaluation, analytical sub-scores are populated from supplier documentation and then confirmed for the shortlist by local verification, because verified performance regularly differs from claimed performance. Operational parameters are extracted from package inserts and confirmed with reference sites, and supply data are requested contractually rather than descriptively, since a stated fill rate that is not a contractual commitment should not score. Presentation matters as much as computation, since a committee that cannot read the output will fall back on price (Akokodaripon et al., 2023; Babatope et al., 2023); the output should be a single page giving the index at both scopes, the cost per performance point, the rank-stability findings, and an explicit statement of which budget bears which cost. Because the model is intended to survive procurement challenge, the order in which decisions are taken matters as much as the arithmetic. Every discretionary parameter has to be fixed and recorded before any bid is opened, and the roles have to be separated so that no single participant sets a parameter and then evaluates against it. Table 19 sets out a workable sequence and division of responsibility for a tender running to a normal timetable. Table 19. Sequence and division of responsibility for a reagent tender. Stage Decision or output Responsible Timing Specification Eligibility screens, analytical performance specification, decision thresholds Laboratory director with clinical lead Before issue Parameterisation Domain weights (via swing weighting), normalisation anchors, trade-off parameter Evaluation panel, minuted Before issue Consequence setting Propagation coefficient decomposition, unit consequence costs Clinical panel with finance Before issue Stage Decision or output Responsible Timing Scope decision Whether award is made at laboratory or institutional scope Finance director Before issue Data request Contractually binding supply, stability, and performance declarations Procurement With tender documents Desk scoring Provisional sub-scores from documentation and reference sites Responsible scientist On bid opening Verification Local confirmation of analytical claims for shortlisted candidates Laboratory, to standing protocols Shortlist stage Robustness Trade-off sweep, propagation stress test, volume and weight perturbation, probabilistic run Analyst, reported in full Before award Recommendation Single page: index at both scopes, cost per performance point, rank stability, budget incidence Evaluation panel to board At award The Parameterisation stage in Table 19 fixes the domain weights before bids are opened, but fixing them still requires a method for eliciting them from the panel, and none has been specified above beyond naming the participants. Swing weighting is recommended for this purpose: the panel first ranks the five domains by the value of moving each from its worst plausible level to its best, assigns the largest swing a reference value of 100, scores the remaining swings relative to it, and normalises the resulting scores to sum to one. Swing weighting is preferred here over direct point allocation or pairwise comparison methods such as the analytic hierarchy process, both discussed in the multi-criteria decision analysis literature cited in Section 2.4, because it forces the panel to weigh each domain against the actual range of variation observed across candidates rather than against an abstract notion of importance, which is the same anchoring discipline the anchored normalisation in Section 3.2 applies to the sub-scores themselves. The elicitation session should be minuted alongside the resulting weights, so that a challenge can examine not only the numbers but the reasoning that produced them. Two features of this sequence do most of the work. The scope decision is taken by the finance director rather than by the laboratory, because it determines whose budget bears the cost and is therefore not a technical question; leaving it to the laboratory guarantees that the institution-wide view is treated as advisory. And the robustness stage is reported in full rather than summarised, so that a committee sees the cases in which the recommendation is unstable. A panel that only ever receives point estimates will come to treat them as precise, which is the failure the second application illustrates. Section 2.3 notes that Directive 2014/24/EU requires award criteria to be objective, disclosed in advance, and applied consistently, and Section 3.10 defends the domain weights against exactly this kind of challenge. The propagation coefficients and unit consequence costs are a harder case, because they are declared by a clinical panel rather than measured, and in the troponin application it is this term, not the weights, that reverses the award. The defence available is the same procedural standard applied to a different parameter. A weight is defensible not because it is objectively correct but because it was fixed in advance, by a named panel, following a documented method, and applied identically to every candidate; the same is true of a propagation coefficient decomposed as in Table 2, signed off by the clinical panel before bids are opened, and stress-tested across a stated range rather than asserted as a single point value. What a challenge process actually scrutinises is whether a criterion was manipulated after the fact, not whether it could in principle have been set differently, and decomposition, panel sign-off, and pre-registration are what make that scrutiny possible for a parameter that cannot be measured directly. A coefficient that survived a documented, panel-sign-off, pre-registered process should be at least as defensible under the Directive 2014/24/EU standard as a weight arrived at by the same procedure, notwithstanding that one is a judgement about importance and the other a judgement about clinical probability. Verification deserves particular emphasis because it is the step most often skipped under time pressure. Declared analytical performance is a marketing figure produced under favourable conditions on a manufacturer's own instruments, and the parameters it feeds are among the most influential in the model. Local confirmation against the laboratory's standing protocols need only cover the shortlist and need only address the indicators that carry weight, principally imprecision at the decision points and agreement with the incumbent method. Where the timetable will not accommodate even that, the honest response is to widen the intervals used in the probabilistic run rather than to proceed as though the declared figures were verified. Contracting practice is where a scoring exercise becomes durable, a point emphasised across the procurement automation, negotiation, and supplier management literature (Akinleye & Adeyoyin, 2021, 2022a, 2022b; Akinleye et al., 2023). Because the model prices operational burden and waste, it creates the basis for shifting that risk to the supplier. Contracts structured as cost per reportable result rather than cost per unit shipped transfer the cost of calibration, control, waste, and repeat testing to the party best able to reduce it. Similar logic applies to guaranteed shelf life on receipt, contractual fill rates with defined remedies, and minimum notice periods for product discontinuation. Each converts a scoring criterion into an enforceable term. Access, affordability, and continuity constraints in lower-resource systems have been examined at length in the pharmaceutical access and health systems literature, and that work should govern any transfer of this model (Asiedu & Asiedu, 2022, 2023a, 2023b, 2024a, 2024b; Aminu-Ibrahim & Ogbete, 2023). The structure is portable but the parameters are not. Where reagent expenditure is a much larger share of total health spend, and where the downstream interventions that generate consequence cost may be unavailable or paid out of pocket, both unit consequence costs change in magnitude and in who bears them. Supply chain resilience typically merits a substantially higher weight, since stockout converts directly into absent diagnosis rather than into delay, and the consequence term should be interpreted with care where the counterfactual to a false negative is not delayed treatment but no treatment, which changes the cost from a resource figure to an outcome figure the model as specified does not capture. Setting the two applications side by side gives a sharper account of when the model earns its keep than either does alone. The determining quantity is the ratio between the spread in consequence cost and the spread in direct cost across the candidate field. Where analytical differences are clinically consequential and the field is genuinely separated on accuracy, as with troponin, that ratio is large, the consequence term dominates, and evaluating at laboratory scope alone gives the wrong answer by a margin that dwarfs the entire reagent budget. Where the measurand is standardised and mature, the field is tightly clustered on accuracy, and results are routinely confirmed on a later sample, as with glycated haemoglobin, the ratio is small and the elaborate machinery adds nothing that cost per reportable result did not already provide. The second application produced a result that does not flatter the model, and it is reported here for that reason. The institutional-scope index selected a candidate that was $30,261 a year worse than the one lowest price would have chosen. The failure is instructive rather than fatal. It occurs because the reciprocal cost normalisation compresses a narrow cost range into a narrow score range, at which point the index is effectively being decided by the performance score, and the performance score is the least precisely measured component of the model. Cost figures come from contracts, package inserts, and laboratory records; domain sub-scores come from expert judgement against anchors. Treating a one-point difference in a composite of the two as decisive is not defensible, which is why the probabilistic analysis and the trade-off sweep are specified as mandatory rather than optional. Both flagged the instability that the point estimate concealed. This suggests a screening step that costs nothing and should precede any scoring exercise. From eligibility-stage data alone, a team can compute the spread in quoted price across the field and a rough spread in expected misclassification cost using declared sensitivity and specificity, local prevalence, and an order-of-magnitude estimate of downstream cost. Where the consequence spread is clearly the larger, the full model is warranted and should be run at institutional scope. Where it is not, the appropriate response is to confirm that every candidate meets the analytical specification, award on cost per reportable result, and put the remaining effort into price negotiation and contract terms. Applying the full instrument to every tender would waste committee time on assays where it cannot change the answer, and would occasionally make the answer worse. The weight perturbation results carry a related lesson about where evaluation panels spend their attention. Moving any single domain weight by 0.10 shifted the institutional-scope index by at most 1.0 point, the largest single deviation being Vendor B's swing to 92.2 under a reduced supply-resilience weight in Table 11, against a lead of roughly twelve points for the winning candidate. The weighting debate that typically consumes most of a scoring meeting was, in the first application, incapable of changing the outcome, while the cost scope decision that is usually settled by default reversed it entirely. Panels would do better to spend their time agreeing the scope boundary and the propagation coefficients, and to treat the weights as a matter for documented judgement rather than extended negotiation. Several limitations qualify these findings. The application uses synthetic data: it demonstrates that the model can produce rank reversal between scopes and identifies the parameter ranges over which reversal occurs, but it is not empirical evidence about any product or about how often such reversals arise in practice. Prospective application across a series of real tenders, with post-award tracking of realised cost against predicted cost, is the necessary next step. The consequence module is the weakest component, since propagation coefficients are declared rather than estimated and no established method exists for deriving them for a specific assay in a specific pathway; prevalence, pathway structure, and clinical interpretation all vary by setting, and the surveillance and informatics literature shows how much that variation can matter (Anioke & Atima, 2023a, 2023b, 2023c). A concrete design for closing this gap is a retrospective chart-review audit: for a sample of flagged or discordant results on the assay in question, drawn from the laboratory information system over a defined period, a reviewer blind to the original propagation assumptions determines, from the clinical record, what fraction were acted upon by the treating team, what fraction were resolved by a repeat or confirmatory sample without further action, and what fraction were overridden on other clinical grounds; this operationalises the same decomposition set out in Table 2 with observed frequencies rather than panel judgement, and can be repeated periodically to track whether the coefficients drift as practice changes. Stress testing mitigates this, and the conclusion here held across a thirty-fold range, but an assay whose direct and consequence terms are closer in magnitude would be less forgiving. A further limitation is visible in the second application. The index reports a point estimate, and where candidates are closely matched that estimate can order them incorrectly relative to total institutional cost, because a small and imprecisely measured performance difference is allowed to outweigh a small and precisely measured cost difference. The model has no internal mechanism for detecting this and relies on the analyst running and reporting the robustness stage. A refinement worth pursuing would make the treatment of near-ties explicit, for example by declaring a minimum separation below which the index returns no ranking, though setting that threshold defensibly would itself require the kind of prospective data the model does not yet have. Attribution of false positives to the reagent is also imperfect for measurands where non-target elevation is common, high-sensitivity troponin being a clear example; restricting analysis to the differential between candidates partly addresses this but does not eliminate it. The model treats sensitivity and specificity as fixed properties of the reagent at a stated cut-off, whereas in practice they depend on the cut-off, the population, and the sampling protocol, and a laboratory may adjust the cut-off after implementation in ways that alter the trade-off the model priced. Weighted additive aggregation assumes preferential independence between domains, which is an approximation: poor onboard stability matters more at low volume, and supply fragility matters more where no validated alternative platform is available. Finally, the model prices what can be counted, and does not capture the value of clinician confidence in a result, the disruption cost of changing methods, or the effect of reference interval and cut-off changes on clinical behaviour during transition. A final observation concerns what the model does to the conversation rather than to the arithmetic. In both applications the decisive quantities turned out to be ones that a conventional tender never writes down: the fraction of purchased reagent that never becomes a reportable result, the number of technologist hours a control architecture consumes, the probability that an incorrect result survives to change a decision, and the identity of the budget that absorbs the consequence. None of these is difficult to estimate, and none is currently estimated. Making them explicit changes what a committee argues about, and in the first application it changed the award by an amount larger than the contract itself. That effect does not depend on accepting the particular index proposed here. A team that rejected the weighting scheme entirely, and simply computed cost per reportable result at both scopes, would have captured most of the value. 4.1. Conclusion Diagnostic reagent procurement is conducted almost everywhere in a frame that measures the wrong quantity. Cost per test is not cost per result, and cost to the laboratory is not cost to the institution. The model set out here corrects both errors within the practical constraints of a real tender, through screens that cannot be traded away, a transparently weighted performance score normalised against absolute rather than relative standards, a cost measure that counts every reagent unit consumed rather than every unit reported, and an explicit, attenuated, and stress-tested account of what wrong results cost. Applied to a parameterised high-sensitivity troponin procurement, the model showed that the cheapest reagent carried the highest hidden cost multiplier, that correcting for operational burden alone was insufficient to change the award, and that including downstream consequence reversed the recommendation decisively and robustly, with the higher-priced candidate returning $13.70 per additional dollar of laboratory spend. The lowest-priced candidate could not have matched the leader at any non-negative price. The analytical case is not difficult. The obstacle is that the finding is a statement about budget boundaries, and budget boundaries are not moved by arithmetic. What the model can do is make the transfer visible, quantify it, and place it in front of the people with authority to authorise it. That is a modest claim, but a necessary precondition for the rest. Declarations Ethics approval and consent to participate: Not applicable. This is a decision-model paper that does not involve human participants, human tissue, or identifiable patient data. Consent for publication: Not applicable. Availability of data and materials: Not applicable. The worked applications in this paper use parameterised illustrative inputs described in the manuscript; no primary datasets require deposition. Competing interests: The authors declare no competing interests.