Other decision-makers adapt
When outcomes depend on competitors, partners or counterparties, their choices cannot always be treated as random environmental noise.
A practical guide to strategic interaction among multiple decision-makers using normal-form games, best responses, equilibrium concepts, repeated adaptation and Markov games.
When outcomes depend on competitors, partners or counterparties, their choices cannot always be treated as random environmental noise.
A strategic equilibrium describes choices that are mutually stable under specified assumptions; it is not automatically socially optimal or uniquely predictive.
In Markov games, agents repeatedly choose actions while jointly affecting the evolving state and each other’s future opportunities.
A single-agent model assumes the environment responds probabilistically but not strategically to the decision-maker.
This is inadequate when another party observes, anticipates or adapts to the policy. Pricing, bidding, capacity competition, negotiation, cybersecurity, traffic and shared-resource allocation can all contain strategic interaction. The model must represent the other agents’ actions and objectives, not merely historical frequencies that may change once our behaviour changes.
Do not invoke game theory simply because several stakeholders exist. It is most useful when their choices materially affect one another and each has meaningful agency. If all parties share the same objective, collaborative optimisation may be more appropriate.
A normal-form game specifies players, available actions and each player’s payoff for every joint action.
The payoff matrix makes strategic dependence explicit. A best response is the action that maximises one player’s payoff given the other players’ strategies. Dominant strategies, when they exist, are best regardless of what others do. Many business games have no dominant action because the best choice depends on expected competitor behaviour.
Mixed strategies assign probabilities across actions. They are not merely indecision; they can make behaviour unpredictable in settings where predictability would be exploitable.
A Nash equilibrium is a strategy profile in which no player can improve unilaterally by deviating.
Equilibrium is a consistency condition. It does not mean the outcome is fair, efficient, stable under learning dynamics or inevitable. Games can have multiple equilibria or none in pure strategies. Which equilibrium becomes relevant can depend on history, conventions, communication and beliefs.
For strategy work, use equilibrium analysis to identify credible responses and strategic traps. Then test the assumptions about rationality, information and available actions. Real competitors have bounded information and organisational constraints; an equilibrium model is a structured scenario, not a guarantee.
Agents may learn by repeatedly responding to observed behaviour rather than solving equilibrium analytically.
A simple learning dynamic forms beliefs about the other agents from their historical actions and repeatedly chooses a best response to those beliefs. Such dynamics can converge in some games and cycle in others. The trajectory itself may matter because temporary behaviours can create real gains or losses before any stable pattern is reached.
Gradient-based multiagent learning changes each agent’s policy in response to its own payoff gradient. Simultaneous adaptation can create non-stationarity: while one agent learns, the environment is changing because others are learning too. Validation should therefore include adaptive opponents rather than only fixed historical behaviour.
In a zero-sum game, one player’s gain is the other’s loss; many business interactions are general-sum because cooperation and competition coexist.
Zero-sum structure supports minimax reasoning: choose a strategy that maximises the worst-case payoff against an adversary. This is useful in strongly adversarial security or contest settings. In general-sum interactions, there may be opportunities for coordination, negotiation or mutual gain that minimax reasoning would miss.
Clarify the payoff structure before selecting a solution concept. Treating a potentially cooperative supplier relationship as purely adversarial can destroy value; treating a strategic competitor as passive can expose the business to exploitation.
A Markov game extends an MDP to several agents whose joint actions determine state transitions and rewards.
At each state, agents choose actions, the environment transitions according to the joint action, and each agent receives its own reward. Policies can depend on the current state, and future strategic interaction affects present action value. This framework can represent repeated pricing, resource competition or autonomous systems sharing an environment.
Learning becomes harder than in a single-agent MDP because the transition experience includes changing opponent policies. Some algorithms extend value learning with equilibrium calculations in each state. Their practical success depends on the game structure, observability and stability of other agents.
Two generic providers decide whether to add capacity in a market with uncertain demand.
If one expands while the other does not, the expanding provider may gain share. If both expand, price and utilisation may fall. If neither expands and demand grows, both may face lost opportunities. A one-company forecast that assumes the competitor keeps current capacity can overvalue expansion.
A game model does not tell management exactly what the competitor will do. It identifies conditional payoffs and credible responses. Management can then explore strategic moves such as staged expansion, signalling, differentiated service or flexible capacity that reduce exposure to the competitor’s choice.
Multiagent models are sensitive to assumptions about objectives and information.
| Assumption | Question |
|---|---|
| Payoff | Have we represented what the other party actually values? |
| Actions | Are important strategic options missing? |
| Information | What can each agent observe when choosing? |
| Rationality | Will agents optimise consistently or use organisational heuristics? |
| Adaptation | How quickly can policies change after observing us? |
| Commitment | Are announcements or contracts credible and enforceable? |
Not automatically. It identifies mutually stable strategies under the model. Multiple equilibria, bounded rationality and incomplete information limit direct prediction.
No, but testing against strong or best-response competitors can reveal vulnerabilities. Scenario analysis can include realistic organisational behaviour as well as adversarial bounds.
A technically correct method still needs an auditable operating translation.
A useful strategic-game analysis should not stop at one assumed competitor strategy. Build a scenario set that includes a passive response, a plausible organisational response, a strong best response and a delayed response. Recalculate the preferred action under each and identify which assumptions cause the strategy to switch. Also examine second-round effects: a move that is attractive against the first response may provoke a later capacity, price or channel adjustment. Where the preferred action changes across credible scenarios, favour flexibility, staged commitment or information gathering rather than pretending one opponent forecast is certain. Document what real-world signals would indicate that the interaction is moving toward one scenario or another, and assign responsibility for monitoring them after the decision.