KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesSequential ProblemsProject Delivery · RiskLesson 28/31← PrevNext →
GuidePublished 6 Jul 2026Updated 13 Aug 20267 min readBy Kevin JoginMarkov gamesrepeated gamesfolk theoremmultiagent reinforcement learning
On this page

Ask about this page

KEVOS AISequential Problems

KEVOS knowledge first · trusted web sources when needed

KEVOS®

Project Management›Project Risk Management›Algorithms for Decision Making›Chapter 24

24Part V · Multiagent Systems

Sequential Problems

Strategic interaction plays out over time, and the other parties adapt as you do. Why long-term relationships cooperate where one-off deals defect.

Chapter 24 of 26 11 min read Original KEVOS® synthesis

Strategic decisions are rarely one-shot. They repeat, they unfold over a project's life, and the other parties learn and adapt just as you do. That single change — a future — transforms what's rational.

Extend the game of Chapter 23 across time and you get a Markov game (a stochastic game): a shared situation that evolves based on the joint actions of all agents, with each agent collecting its own rewards over the sequence. It is the MDP of Part II with more than one decision-maker — and the added twist that everyone is adapting simultaneously.

1The shadow of the future

The most important insight here is why repetition changes everything. In a one-off prisoner's dilemma, defection is rational. But when the interaction repeats — as nearly every real project relationship does — cooperation can become the rational choice, because defecting today invites retaliation tomorrow. This is the shadow of the future: the value of the ongoing relationship disciplines present behaviour. A well-known result confirms that a wide range of cooperative outcomes can be sustained in repeated play, held in place by the credible threat of future response.

Round t cooperate Round t+1 cooperate Round t+2 cooperate the threat of future punishment sustains cooperation today One-off deal → defect  ·  Ongoing relationship → cooperation becomes rational
Figure 1. Across repeated rounds, cooperation holds because betrayal now is punished later. The longer and more certain the future relationship, the stronger its grip on present behaviour — which is why enduring partnerships behave so differently from one-off transactions.

2Learning against a moving target

When agents don't know the game and must learn — multiagent reinforcement learning — a new difficulty appears. Each agent is learning while the others learn too, so from any one agent's view the environment keeps changing: the very thing it's adapting to is itself adapting. This non-stationarity is what makes multiagent learning genuinely hard, and it's why naïvely applying single-agent methods can chase its own tail. Techniques that anticipate others' adaptation, or that converge to equilibrium play, are needed to make progress.

Key idea

A future changes the game. Cooperation that's irrational in a one-shot encounter becomes rational when the relationship continues and defection can be punished. And when everyone is adapting at once, you're optimising against a moving target — the core challenge of learning among others.

What it means in practice

The single most useful lever in multi-party project work is lengthening the shadow of the future: make relationships ongoing, outcomes repeated, and reputations visible, so that cooperating pays and defecting costs. This is precisely why long-term framework partnerships tend to behave better than one-off, lowest-price contracts — the future keeps everyone honest. And when you're negotiating against a party that keeps shifting its approach, don't expect a fixed opponent; anticipate that they're adapting to you, and plan for a moving target.

Handbook application: from concept to controlled practice

Purpose. This expanded section turns the original page into a practical handbook. It preserves the supplied material and adds a repeatable way to apply, check and review Sequential Problems. It does not replace a contract, legislation, a controlled standard, competent engineering judgement or specialist advice.

The operating aim is to convert the subject into a governed decision, owned work, usable evidence and a reviewable outcome. Read the original explanation first, then use the workflow and checks below to convert knowledge into evidence.

Use Sequential Problems as a decision instrument rather than an administrative form. The subject terms—games, learning, markov, repeated, shadow—need an explicit connection to the project objective, business value and stakeholder commitments. Before completing the artefact, write one sentence stating who will use it, what decision it supports and when that decision is required.

Apply a disciplined information model. Separate facts supported by evidence, forecasts derived from a method, assumptions awaiting validation, constraints that limit choice, risks that may occur, issues that already exist and actions assigned to people. Each material entry should have an owner, date, status and next review point. Where probability or impact scores are used, define the scale so different reviewers interpret it consistently.

A baseline is useful only when changes are visible. Give the artefact an identifier, version, approval state and effective date. Define which changes require reapproval, how superseded versions are retained and where supporting evidence is stored. During reviews, focus on exceptions, decisions and trends rather than reading every field aloud. Record the decision and rationale, not merely that a meeting occurred.

Close the loop beyond delivery. Confirm acceptance criteria, unresolved items, transferred responsibilities and operational ownership. Where benefits are expected, identify the outcome measure, baseline, target, observation period and owner who remains accountable after the project team disbands. Lessons should describe the condition, consequence and reusable action; a generic statement such as “communicate better” cannot improve the next project.

Step-by-step operating method

  1. Clarify the decision. Name the outcome, sponsor, affected stakeholders and decision that this work must enable.
  2. Set boundaries. Record scope, assumptions, constraints, dependencies, tolerances and escalation conditions.
  3. Plan the evidence. Define deliverables, measures, owners, due dates and acceptance criteria before execution.
  4. Control delivery. Compare actual performance with the baseline, assess changes and manage risks and issues explicitly.
  5. Close the loop. Confirm acceptance, transfer ownership, capture lessons and track benefits beyond handover.

Completion and governance protocol

Start with a short drafting workshop involving the accountable owner and the people who hold the evidence. Complete high-consequence fields first: objective, scope, owner, baseline, acceptance, dependencies and escalation. Mark unknowns as assumptions or actions rather than hiding them behind vague prose. Circulate a review draft, resolve conflicting interpretations, baseline the approved version and place the next review date in an owned schedule.

Information typeMinimum useful contentReview test
OutcomeObservable change and intended recipientNot merely a deliverable or activity
MeasureDefinition, baseline, target, frequency and sourceTwo reviewers would calculate it the same way
OwnershipOne accountable role plus contributors and approverAuthority matches responsibility
UncertaintyAssumption, risk or issue with response and triggerStatus reflects current reality
ControlVersion, approval, review date and change ruleCurrent baseline is identifiable

Common failure modes and recovery actions

1. Watch for

Producing a document with no named decision or accountable owner.

Recovery: Return to the governing definition or requirement and restate the decision in one sentence.

2. Watch for

Mixing risks, current issues, assumptions and actions in one unstructured list.

Recovery: Separate evidence from assumption, assign an owner and set a date for validation.

3. Watch for

Measuring activity or output while leaving the intended outcome undefined.

Recovery: Run a small counterexample, boundary test, pilot or independent check before proceeding.

4. Watch for

Accepting changes without evaluating effects on value, scope, schedule, cost and risk.

Recovery: Record the consequence, decision and rationale, then update the controlled baseline.

5. Watch for

Closing the project at delivery even though benefit ownership has not transferred.

Recovery: Escalate when the issue affects safety, compliance, acceptance, material value or an agreed tolerance.

Review checklist

  • Which decision or commitment does this artefact support?
  • Who owns each action, risk, acceptance and post-project benefit?
  • What is the baseline and what variance triggers escalation?
  • Where is the evidence that the result was accepted and transferred?
  • Are mandatory requirements distinguished from recommendations and illustrative values?
  • Are sources, assumptions, units, dates and versions recorded closely enough to reproduce the decision?
  • Have safety, legal, ethical, stakeholder and operational consequences been considered at the appropriate level?
  • Is there a named owner and a trigger for review, escalation, change or retirement?

Questions for deeper application

What is the most important distinction a practitioner must preserve when applying Sequential Problems?

Answer with a fact or cited source where available. Where evidence is incomplete, record the assumption, consequence, responsible owner and next validation action.

Which assumption about games would change the result most if it proved false?

Answer with a fact or cited source where available. Where evidence is incomplete, record the assumption, consequence, responsible owner and next validation action.

What evidence would allow an independent reviewer to reproduce or challenge the conclusion?

Answer with a fact or cited source where available. Where evidence is incomplete, record the assumption, consequence, responsible owner and next validation action.

Which boundary, exception or failure case has not yet been tested?

Answer with a fact or cited source where available. Where evidence is incomplete, record the assumption, consequence, responsible owner and next validation action.

What must be handed over, monitored or reviewed after the immediate work is complete?

Answer with a fact or cited source where available. Where evidence is incomplete, record the assumption, consequence, responsible owner and next validation action.

Authoritative references and use notes

The sources below were selected as institutional or primary guidance for the broader practice. They support the handbook method; they do not imply that every statement or clause in a source applies to every project. Confirm the current edition, jurisdiction, contract and application before treating any requirement as mandatory.

  • PMI Standards and Publications — Project Management Institute. Used for project, program, portfolio and organisational project management. Accessed 2026-08-13.
  • ISO 31000 family — Risk management — International Organization for Standardization. Used for principles and guidance for enterprise risk management. Accessed 2026-08-13.
← PreviousCh 23 · Multiagent Reasoning Next →Ch 25 · State Uncertainty

↑ All 26 chapters — series overview

KEVOS® — Engineering & Project Consultancy © KEVOS®. All rights reserved.

Continue learning

Multiagent ReasoningGuide · RiskNEXT LESSON →State UncertaintyGuide · RiskController AbstractionsGuide · RiskCollaborative AgentsGuide · Risk
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®