KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesApplying Termination Criteria by Project StageProject Delivery · Research ProjectsLesson 130/139← PrevNext →
GuidePublished 16 Aug 202614 min readBy KEVOS Editorialcontinue or kill reviewstage specific termination criteriar&d gate review weightingcomprehensive risk measure
On this page

Ask about this page

KEVOS AIApplying Termination Criteria by Project Stage

KEVOS knowledge first · trusted web sources when needed

KEVOS/Project Delivery/Research Projects/R&D Management Papers
Project DeliveryResearch ProjectsAdvancedRd Project Management

Applying Termination Criteria by Project Stage

The strongest early warning signal in R5's data had fallen to seventh place by the end of development. A monitoring regime that fixes its criteria at the first gate is measuring the wrong things by the last one.

Reading time15 minutes
LevelAdvanced
Topic streamRd Project Management
Source materialR&D Management Papers
Updated2026-08-16

In brief

  • R5 closes with four operating guidelines. Together they say: vary your criteria by stage, estimate the model stage by stage and firm by firm, keep the framework multivariate and dynamic, and use a combined risk measure rather than a threshold per indicator.
  • The stage evidence in the article is real but thin. Full per-stage coefficient tables are not printed; two variables' stage behaviour is reported in prose, and the final stage carries eleven coefficients where twelve variables exist.
  • The comprehensive-risk measure is offered to remove a specific pain point — having to set and defend a target value for every indicator. Its mathematics are not published.
  • The guideline with the most immediate practical consequence is the one about data: estimate on your own firm's project history, not on cross-firm coefficients.
  • R5 is explicit that quantitative screening is to be integrated with qualitative judgement. Neither alone.

The four guidelines, and what each one costs to follow

R5's closing section turns its empirical results into four numbered instructions. The model and the twelve variables that feed them are on the companion page, making better project termination decisions; this page is about running the review.

R5's four operating guidelines

GUIDELINE 1

Different critical factors apply at different stages

Different variables have high and significant discriminating strength at each stage, so R&D project managers should emphasise individual factors dynamically through the monitoring process rather than holding one fixed factor set. The cost: your gate criteria stop being a single document and become three, and each one has to be justified separately.

GUIDELINE 2

Analyse stage by stage, and within one firm

Applying the discriminant procedure at each stage identifies stage-specific key factors positively and enhances the accuracy of results. R5 adds that it is better to collect data for R&D projects within one firm, to increase the chance of a correct termination decision for that firm. The cost: you need a classified project history before the method does anything for you.

GUIDELINE 3

Keep the framework multivariate and dynamic

The monitoring framework should measure the comprehensive effects of a number of factors and include dynamic identification functions for the changing R&D process, in order to avoid bias and faulty results. R5's empirical study indicates that integrating quantitative methods with qualitative ones is more helpful than relying only on personal operating experience.

GUIDELINE 4

Use a comprehensive risk measure

A mathematical measure of an ongoing project's overall risk, derived from the discriminant analysis, in place of indicator-by-indicator threshold comparison. Claimed to be simple to use and validated by case studies. The cost: the equations are not in the article, so what you build is your own implementation of the idea.

From the source

The guideline that is easiest to skip and hardest to substitute

R5 states that it is better to collect the data for R&D projects within one firm, to increase the chance of a correct termination decision for that firm.

Read plainly, that is the authors telling you their own published coefficients may not transfer to your organisation. The twelve variables are a place to start looking; the ranking is a result about 217 projects in 17 industries between 1997 and 1999. If you adopt the ranking as weights without local estimation, you have imported a finding and called it a standard.

What the article actually establishes about stage differences

This is where a handbook has to be careful. The claim that critical factors differ by stage is central to the paper, but the published evidence for it is a small number of figures in running text. There are no numbered tables or figures in the article at all.

  1. Initial stagePriority placed on product quality relative to competitors carries a canonical discriminant coefficient of 0.676 — the highest of all coefficients at that stage. A study finding, on that sample.
  2. Middle stageThe same variable falls to 0.45 and remains the highest at that stage. R5's characterisation is that it has higher discriminating strength at all three stages, while its absolute strength declines.
  3. Final stageThe same variable is 0.269, ranking seventh among the eleven coefficients reported for that stage. Separately, the correlation between project output and degree of urgency at the final stage is 0.345, described as a larger positive value.

Two things follow from the final-stage line. First, the eleven-against-twelve count means not every surviving variable enters at every stage — the variable set is genuinely stage-specific, not merely re-weighted. Second, a criterion that dominated the first two gates has, by the last one, fallen to the middle of the pack.

Source gap

What is not printed, and what that stops you doing

The article gives stage-level numbers for two of the twelve variables. There is no printed table of coefficients by variable by stage, and the mathematics of the comprehensive-risk measure are not published either — readers are directed to contact the lead author.

So you cannot reconstruct R5's stage-specific weightings from the article, and any weighting you apply is your own. What the paper does support is the structural claim — that the discriminating set changes between stages — and the instruction to re-estimate rather than carry a fixed set forward. Build on the structure; do not pretend to the numbers.

The comprehensive-risk measure and the problem it removes

The fourth guideline is the paper's proposed replacement for threshold monitoring, and its stated advantage is worth isolating because it is a real operational pain rather than a statistical nicety.

R5'S STATED PROPERTIES OF THE COMPREHENSIVE-RISK APPROACH

Claimed propertyWhat it addressesStatus in the article
Simple to useAdoption by project managers rather than analystsAsserted, not demonstrated
Avoids having to set a threshold for every variableIndicator-by-indicator monitoring requires a target or threshold per indicator — each of which must be set, defended and revisedThe specific pain point the measure is offered to remove
Comprehensively measures the risk an ongoing project facesThe conflicting-signals problem: one number derived from the joint discriminating structure rather than a wall of separate readingsAsserted; follows from the discriminant derivation
Validated by case studiesWhether the measure's results are usable in practiceClaimed. The case studies are not reported in the article

Properties as stated by R5. The mathematical details are not printed, so none of these can be independently checked from the published article.

The threshold point deserves attention even if you never build the measure. Every indicator added to a monitoring pack silently adds a governance obligation: someone has to decide what value counts as bad, and defend it when a project is close to the line. With a dozen indicators that is a dozen arguments, and they are usually settled by precedent rather than evidence.

Running the review

R5 gives decision rules rather than a meeting agenda. The sequence below assembles them into the order a review would use. The rules are the paper's; the sequencing is this library's, and the source does not prescribe it.

A continue-or-kill review, assembled from R5's rules

  1. Establish which stage the project is in

    Initial, middle or final within the development phase. This determines which factor set is in force. R5's scope is the development stage only — it claims nothing about pre-development or post-launch termination.

  2. Take readings on the factor set for that stage, not the master list

    Emphasise factors dynamically through the monitoring process rather than holding one fixed set. A variable that carried the first gate may not enter the model at the last.

  3. Read the moving variables as movement

    Expected probability of technical success is explicitly dynamic. If emerging problems are not resolved in the expected time frame and new ones arise, expect it to fall sharply — R5 treats that fall as a termination signal.

  4. Combine, do not compare in isolation

    Do not judge the project by comparing performance against target values of indicators taken one at a time. Success or failure depends on a combination of variables, so the assessment is of the combination.

  5. Add the qualitative reading explicitly

    R5's finding is that integrating quantitative methods with qualitative ones beats relying on personal operating experience alone. Record the qualitative judgement as a separate input, so it can be reviewed later rather than absorbed invisibly into the score.

  6. Decide, and record what the decision rested on

    The named failure mode is that leaders seldom make termination decisions for ongoing projects in time. A review that defers without recording why produces the same outcome as no review.

Rules R5 states in conditional form

IfDifferent indicators are giving conflicting signals
ThenDo not resolve the conflict by picking the indicator you trust. R5 identifies this as exactly the situation where isolated comparison produces faulty interpretation
IfYou are carrying the same criteria forward from the previous gate
ThenRe-estimate the discriminating factor set for the current stage instead. Do not carry a factor set forward unchanged
IfThe project is early in development
ThenMonitor the product-quality priority variable hardest — the single strongest discriminator at both the initial and the middle stage in this study
IfTechnical problems are overrunning and new ones are appearing
ThenTreat the fall in expected technical success probability as a termination signal, not as a schedule problem to be re-planned
IfYou are in the development phase and cannot list the expected uses of the output
ThenTreat that as a finding in itself. R5's result is that more identified uses during development correlates with a better chance of success
IfYou want a decision rule that is correct for your firm
ThenEstimate the model on your own firm's project history rather than on cross-firm coefficients

Estimating it on your own project history

Guideline 2 is the one that turns this from reading into work. The requirements are not exotic, but they have to be in place before the first review that uses them, and most of them are records rather than analysis.

What you need before a local model is possible

  • A history of completed development projects classified as successful or failed, with the classification rule written down
  • Enough of both classes for a discriminant procedure to separate them — R5's own base was 152 successful and 65 failed projects
  • Ratings captured at comparable points in development, not reconstructed afterwards from memory
  • A stable definition of what initial, middle and final mean in your development process
  • A single respondent type, or a stated rule for whose rating counts — R5 used project leaders and managers throughout
  • Somewhere for the ratings to live between projects, so the history accumulates rather than being rebuilt each time
Practice note

If you cannot estimate a model, you can still change the review

The source does not prescribe this, but three of R5's guidelines cost nothing statistical to adopt. Vary the criteria by stage using judgement and record why. Stop scoring indicators against thresholds one at a time. Write the qualitative assessment down as its own input rather than letting it arrive as the chair's summing-up.

Those changes do not deliver the model's accuracy. They do remove the failure mode the paper names — a review that cannot reach a decision because its indicators disagree — and they generate exactly the records a model would later need.

Conditions of applicability

Four boundaries govern how far any of this travels. They are the study's own, and they are the things to say out loud when someone proposes adopting the twelve variables as policy.

  • Stage scope. Development only, chosen because that is where the considerable resources are consumed. A pre-development screen or a post-launch withdrawal decision is a different problem with different evidence.
  • Population. Manufacturing firms in one metropolitan area of one country, projects from 1997 to 1999, across 17 industries. The paper describes its setting as one where the commercialisation ratio of R&D results is low and top leaders concentrate on selection rather than ongoing monitoring — a context that shapes what its respondents were doing.
  • Data quality. 135 valid questionnaires from 375 distributed, with 217 projects nested inside those 135 responses, so the projects are not independent observations. Success and failure were classified by the responding managers, retrospectively.
  • Disclosure. The risk model's mathematics are unpublished and its validation rests on unreported case studies. Anything you build is your own model informed by R5.
Check before you proceed

Before you take a stage-specific criteria set to a governance forum

Check that each stage's criteria set is written separately and that someone can say why a criterion appears at one gate and not another. Check that no criterion has a threshold attached that nobody can source. Check that the qualitative input has a named owner and a place on the form. And check that the paper you are citing is described as one empirical study of 217 projects, not as an industry practice — the same discipline the library applies on decision quality in R&D and governing R&D decisions organisationally.

One last boundary is conceptual rather than methodological. R5 asks which factors separate projects that succeeded from projects that failed. It does not ask whether the uncertainty in a project is being resolved, which is the continuation test in sources of uncertainty in R&D projects, and it does not ask what the project turned out to be worth. A review can run all three questions, but it should know which one it is answering at any moment. Whatever the decision, the reasoning behind it is worth capturing while it is fresh — that is the material post-project reviews in R&D depend on.

What to carry forward

  1. Criteria sets should differ by stage. In this study the strongest early discriminator fell from first to seventh place by the final stage, and only eleven of the twelve variables entered there at all.
  2. The published evidence for stage differences is thinner than the claim. Adopt the structure — re-estimate by stage — rather than pretending to coefficients the article does not print.
  3. The comprehensive-risk idea earns its place by removing the need to set and defend a threshold for every indicator. The equations are not available, so treat any build as your own.
  4. Estimate locally. R5's own recommendation to collect data within one firm is the strongest caveat it places on its published ranking.
  5. Integrate quantitative screening with qualitative judgement, and record the qualitative input separately so it can be reviewed later.
  6. The failure mode being addressed is lateness, not error. A review that cannot decide because its indicators disagree is the thing this method exists to fix.

Frequently asked questions

How do I set the weights for each stage if the article does not print them?

You cannot take them from the article — only two variables' stage-level figures are published. What R5 supports is the practice of establishing a separate factor set per stage and re-estimating it rather than carrying one set forward. Until you have local data, use judgement, write down the reasoning, and label the weights as yours.

What counts as the initial, middle and final stage?

R5 places all three inside the development stage and does not define their boundaries further. That definition is yours to make and to hold stable, because a model estimated at three points is only meaningful if those points mean the same thing on every project.

Is the comprehensive-risk measure something we can implement?

Not from the article. Its mathematical details are not printed and readers are directed to contact the lead author. You can implement the idea — one combined measure derived from the variables that discriminate, in place of a threshold per indicator — but you would be building your own model, and it should be described that way.

Should the review kill a project on a single strong signal?

R5's position is that success or failure depends on a combination of variables, so no single reading is a kill trigger. The one signal it treats with particular weight is a sharp fall in expected technical success probability when problems overrun and new ones appear — and even that is described as a signal, not a rule.

How does this relate to stage gates we already run?

It challenges one specific feature of them: identical criteria and weights at every gate. R5's finding that the strongest early discriminator ranked seventh of eleven by the final stage is a direct argument for gate-specific criteria. The rest of your gate structure is untouched by the paper.

Where does qualitative judgement belong in this?

As a recorded input alongside the quantitative reading. R5's empirical study indicates that integrating quantitative methods with qualitative ones is more helpful than relying on personal operating experience alone — which is an argument for capturing judgement explicitly, not for excluding it.

References and source attribution

  1. R5 — making better project termination decisions. Practitioner-facing empirical short article in a journal for research and technology management, January–February 2002; 3 printed pages; 3 references; no numbered tables or figures, one sidebar box on method and one pull-quote. Sections used here: the four operating guidelines, the comprehensive-risk approach, the stated decision rules and the stated limitations.
  2. The case studies said to validate the comprehensive-risk measure, and the mathematical details of the measure itself, are not reported in R5; readers are directed to contact the lead author.
  3. Eleven copyrighted journal articles on R&D project management, supplied as a reading set for a literature review and profiled for this library. Front matter, abstracts, framework sections, tables and figures were read; article bodies were not reproduced, and all content here is paraphrase. The set is a reading list, not a systematic or representative survey of the field.
  4. Supplied teaching source for this library (research methods and research process materials). Used here for page conventions and voice only; it does not treat continue-or-kill reviews.

Suggested questions for Ask KEVOS

  • Draft three stage-specific criteria sets for our development gates and mark which criteria are ours rather than the study's.
  • Review our current gate pack and list every indicator that carries a threshold nobody can source.
  • Design the record we would need to keep so a termination model could be estimated on our own history in three years.
  • Write the section of a gate paper that reports a combined assessment instead of indicator-by-indicator comparison.
  • How should the qualitative assessment be captured so it survives into the post-project review?
  • Explain to a governance forum why the same criteria should not be weighted identically at every gate.

Related KEVOS knowledge

Making Better Project Termination DecisionsCore · rd project managementSources of Uncertainty in R&D ProjectsCore · rd project managementWhat Distinguishes Successful R&D ProjectsCore · rd project managementGoverning R&D Decisions OrganisationallyAdvanced · rd project managementDecision Quality in R&D: Six DimensionsCore · rd project managementPost-Project Reviews in R&DCore · rd project management
KEVOS® · Project Delivery · Research Projects Page KVS-PM-RES-0130 · v1.0.0 · content 2026.08 Last reviewed 2026-08-16

Continue learning

Making Better Project Termination DecisionsGuide · Research ProjectsNEXT LESSON →Measuring New Product Success RatesGuide · Research ProjectsWhat Distinguishes Successful R&D ProjectsGuide · Research ProjectsFour Myths About New Product FailureGuide · Research Projects
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®