Applying Termination Criteria by Project Stage
The strongest early warning signal in R5's data had fallen to seventh place by the end of development. A monitoring regime that fixes its criteria at the first gate is measuring the wrong things by the last one.
The four guidelines, and what each one costs to follow
R5's closing section turns its empirical results into four numbered instructions. The model and the twelve variables that feed them are on the companion page, making better project termination decisions; this page is about running the review.
R5's four operating guidelines
Different critical factors apply at different stages
Different variables have high and significant discriminating strength at each stage, so R&D project managers should emphasise individual factors dynamically through the monitoring process rather than holding one fixed factor set. The cost: your gate criteria stop being a single document and become three, and each one has to be justified separately.
Analyse stage by stage, and within one firm
Applying the discriminant procedure at each stage identifies stage-specific key factors positively and enhances the accuracy of results. R5 adds that it is better to collect data for R&D projects within one firm, to increase the chance of a correct termination decision for that firm. The cost: you need a classified project history before the method does anything for you.
Keep the framework multivariate and dynamic
The monitoring framework should measure the comprehensive effects of a number of factors and include dynamic identification functions for the changing R&D process, in order to avoid bias and faulty results. R5's empirical study indicates that integrating quantitative methods with qualitative ones is more helpful than relying only on personal operating experience.
Use a comprehensive risk measure
A mathematical measure of an ongoing project's overall risk, derived from the discriminant analysis, in place of indicator-by-indicator threshold comparison. Claimed to be simple to use and validated by case studies. The cost: the equations are not in the article, so what you build is your own implementation of the idea.
What the article actually establishes about stage differences
This is where a handbook has to be careful. The claim that critical factors differ by stage is central to the paper, but the published evidence for it is a small number of figures in running text. There are no numbered tables or figures in the article at all.
- Initial stagePriority placed on product quality relative to competitors carries a canonical discriminant coefficient of 0.676 — the highest of all coefficients at that stage. A study finding, on that sample.
- Middle stageThe same variable falls to 0.45 and remains the highest at that stage. R5's characterisation is that it has higher discriminating strength at all three stages, while its absolute strength declines.
- Final stageThe same variable is 0.269, ranking seventh among the eleven coefficients reported for that stage. Separately, the correlation between project output and degree of urgency at the final stage is 0.345, described as a larger positive value.
Two things follow from the final-stage line. First, the eleven-against-twelve count means not every surviving variable enters at every stage — the variable set is genuinely stage-specific, not merely re-weighted. Second, a criterion that dominated the first two gates has, by the last one, fallen to the middle of the pack.
The comprehensive-risk measure and the problem it removes
The fourth guideline is the paper's proposed replacement for threshold monitoring, and its stated advantage is worth isolating because it is a real operational pain rather than a statistical nicety.
R5'S STATED PROPERTIES OF THE COMPREHENSIVE-RISK APPROACH
| Claimed property | What it addresses | Status in the article |
|---|---|---|
| Simple to use | Adoption by project managers rather than analysts | Asserted, not demonstrated |
| Avoids having to set a threshold for every variable | Indicator-by-indicator monitoring requires a target or threshold per indicator — each of which must be set, defended and revised | The specific pain point the measure is offered to remove |
| Comprehensively measures the risk an ongoing project faces | The conflicting-signals problem: one number derived from the joint discriminating structure rather than a wall of separate readings | Asserted; follows from the discriminant derivation |
| Validated by case studies | Whether the measure's results are usable in practice | Claimed. The case studies are not reported in the article |
Properties as stated by R5. The mathematical details are not printed, so none of these can be independently checked from the published article.
The threshold point deserves attention even if you never build the measure. Every indicator added to a monitoring pack silently adds a governance obligation: someone has to decide what value counts as bad, and defend it when a project is close to the line. With a dozen indicators that is a dozen arguments, and they are usually settled by precedent rather than evidence.
Running the review
R5 gives decision rules rather than a meeting agenda. The sequence below assembles them into the order a review would use. The rules are the paper's; the sequencing is this library's, and the source does not prescribe it.
A continue-or-kill review, assembled from R5's rules
Establish which stage the project is in
Initial, middle or final within the development phase. This determines which factor set is in force. R5's scope is the development stage only — it claims nothing about pre-development or post-launch termination.
Take readings on the factor set for that stage, not the master list
Emphasise factors dynamically through the monitoring process rather than holding one fixed set. A variable that carried the first gate may not enter the model at the last.
Read the moving variables as movement
Expected probability of technical success is explicitly dynamic. If emerging problems are not resolved in the expected time frame and new ones arise, expect it to fall sharply — R5 treats that fall as a termination signal.
Combine, do not compare in isolation
Do not judge the project by comparing performance against target values of indicators taken one at a time. Success or failure depends on a combination of variables, so the assessment is of the combination.
Add the qualitative reading explicitly
R5's finding is that integrating quantitative methods with qualitative ones beats relying on personal operating experience alone. Record the qualitative judgement as a separate input, so it can be reviewed later rather than absorbed invisibly into the score.
Decide, and record what the decision rested on
The named failure mode is that leaders seldom make termination decisions for ongoing projects in time. A review that defers without recording why produces the same outcome as no review.
Rules R5 states in conditional form
Estimating it on your own project history
Guideline 2 is the one that turns this from reading into work. The requirements are not exotic, but they have to be in place before the first review that uses them, and most of them are records rather than analysis.
What you need before a local model is possible
- A history of completed development projects classified as successful or failed, with the classification rule written down
- Enough of both classes for a discriminant procedure to separate them — R5's own base was 152 successful and 65 failed projects
- Ratings captured at comparable points in development, not reconstructed afterwards from memory
- A stable definition of what initial, middle and final mean in your development process
- A single respondent type, or a stated rule for whose rating counts — R5 used project leaders and managers throughout
- Somewhere for the ratings to live between projects, so the history accumulates rather than being rebuilt each time
Conditions of applicability
Four boundaries govern how far any of this travels. They are the study's own, and they are the things to say out loud when someone proposes adopting the twelve variables as policy.
- Stage scope. Development only, chosen because that is where the considerable resources are consumed. A pre-development screen or a post-launch withdrawal decision is a different problem with different evidence.
- Population. Manufacturing firms in one metropolitan area of one country, projects from 1997 to 1999, across 17 industries. The paper describes its setting as one where the commercialisation ratio of R&D results is low and top leaders concentrate on selection rather than ongoing monitoring — a context that shapes what its respondents were doing.
- Data quality. 135 valid questionnaires from 375 distributed, with 217 projects nested inside those 135 responses, so the projects are not independent observations. Success and failure were classified by the responding managers, retrospectively.
- Disclosure. The risk model's mathematics are unpublished and its validation rests on unreported case studies. Anything you build is your own model informed by R5.
One last boundary is conceptual rather than methodological. R5 asks which factors separate projects that succeeded from projects that failed. It does not ask whether the uncertainty in a project is being resolved, which is the continuation test in sources of uncertainty in R&D projects, and it does not ask what the project turned out to be worth. A review can run all three questions, but it should know which one it is answering at any moment. Whatever the decision, the reasoning behind it is worth capturing while it is fresh — that is the material post-project reviews in R&D depend on.
What to carry forward
- Criteria sets should differ by stage. In this study the strongest early discriminator fell from first to seventh place by the final stage, and only eleven of the twelve variables entered there at all.
- The published evidence for stage differences is thinner than the claim. Adopt the structure — re-estimate by stage — rather than pretending to coefficients the article does not print.
- The comprehensive-risk idea earns its place by removing the need to set and defend a threshold for every indicator. The equations are not available, so treat any build as your own.
- Estimate locally. R5's own recommendation to collect data within one firm is the strongest caveat it places on its published ranking.
- Integrate quantitative screening with qualitative judgement, and record the qualitative input separately so it can be reviewed later.
- The failure mode being addressed is lateness, not error. A review that cannot decide because its indicators disagree is the thing this method exists to fix.
Frequently asked questions
How do I set the weights for each stage if the article does not print them?
You cannot take them from the article — only two variables' stage-level figures are published. What R5 supports is the practice of establishing a separate factor set per stage and re-estimating it rather than carrying one set forward. Until you have local data, use judgement, write down the reasoning, and label the weights as yours.
What counts as the initial, middle and final stage?
R5 places all three inside the development stage and does not define their boundaries further. That definition is yours to make and to hold stable, because a model estimated at three points is only meaningful if those points mean the same thing on every project.
Is the comprehensive-risk measure something we can implement?
Not from the article. Its mathematical details are not printed and readers are directed to contact the lead author. You can implement the idea — one combined measure derived from the variables that discriminate, in place of a threshold per indicator — but you would be building your own model, and it should be described that way.
Should the review kill a project on a single strong signal?
R5's position is that success or failure depends on a combination of variables, so no single reading is a kill trigger. The one signal it treats with particular weight is a sharp fall in expected technical success probability when problems overrun and new ones appear — and even that is described as a signal, not a rule.
How does this relate to stage gates we already run?
It challenges one specific feature of them: identical criteria and weights at every gate. R5's finding that the strongest early discriminator ranked seventh of eleven by the final stage is a direct argument for gate-specific criteria. The rest of your gate structure is untouched by the paper.
Where does qualitative judgement belong in this?
As a recorded input alongside the quantitative reading. R5's empirical study indicates that integrating quantitative methods with qualitative ones is more helpful than relying on personal operating experience alone — which is an argument for capturing judgement explicitly, not for excluding it.
References and source attribution
- R5 — making better project termination decisions. Practitioner-facing empirical short article in a journal for research and technology management, January–February 2002; 3 printed pages; 3 references; no numbered tables or figures, one sidebar box on method and one pull-quote. Sections used here: the four operating guidelines, the comprehensive-risk approach, the stated decision rules and the stated limitations.
- The case studies said to validate the comprehensive-risk measure, and the mathematical details of the measure itself, are not reported in R5; readers are directed to contact the lead author.
- Eleven copyrighted journal articles on R&D project management, supplied as a reading set for a literature review and profiled for this library. Front matter, abstracts, framework sections, tables and figures were read; article bodies were not reproduced, and all content here is paraphrase. The set is a reading list, not a systematic or representative survey of the field.
- Supplied teaching source for this library (research methods and research process materials). Used here for page conventions and voice only; it does not treat continue-or-kill reviews.
Suggested questions for Ask KEVOS
- Draft three stage-specific criteria sets for our development gates and mark which criteria are ours rather than the study's.
- Review our current gate pack and list every indicator that carries a threshold nobody can source.
- Design the record we would need to keep so a termination model could be estimated on our own history in three years.
- Write the section of a gate paper that reports a combined assessment instead of indicator-by-indicator comparison.
- How should the qualitative assessment be captured so it survives into the post-project review?
- Explain to a governance forum why the same criteria should not be weighted identically at every gate.
