KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesMultiagent Strategy, Games and Sequential InteractionBusiness · StrategyLesson 13/14← PrevNext →
GuidePublished 13 Aug 20266 min readBy Kevin Jogingame theorymultiagent systemsNash equilibriumbest response
On this page

Ask about this page

KEVOS AIMultiagent Strategy, Games and Sequential Interaction

KEVOS knowledge first · trusted web sources when needed

Business · Strategy

Multiagent Strategy, Games and Sequential Interaction

A practical guide to strategic interaction among multiple decision-makers using normal-form games, best responses, equilibrium concepts, repeated adaptation and Markov games.

Handbook guide19 min readUpdated 2026-08-13

Other decision-makers adapt

When outcomes depend on competitors, partners or counterparties, their choices cannot always be treated as random environmental noise.

Equilibrium is a consistency concept

A strategic equilibrium describes choices that are mutually stable under specified assumptions; it is not automatically socially optimal or uniquely predictive.

Sequential games add state

In Markov games, agents repeatedly choose actions while jointly affecting the evolving state and each other’s future opportunities.

When a multiagent model is needed

A single-agent model assumes the environment responds probabilistically but not strategically to the decision-maker.

This is inadequate when another party observes, anticipates or adapts to the policy. Pricing, bidding, capacity competition, negotiation, cybersecurity, traffic and shared-resource allocation can all contain strategic interaction. The model must represent the other agents’ actions and objectives, not merely historical frequencies that may change once our behaviour changes.

Do not invoke game theory simply because several stakeholders exist. It is most useful when their choices materially affect one another and each has meaningful agency. If all parties share the same objective, collaborative optimisation may be more appropriate.

Normal-form games

A normal-form game specifies players, available actions and each player’s payoff for every joint action.

The payoff matrix makes strategic dependence explicit. A best response is the action that maximises one player’s payoff given the other players’ strategies. Dominant strategies, when they exist, are best regardless of what others do. Many business games have no dominant action because the best choice depends on expected competitor behaviour.

Best responseBRᵢ(π₋ᵢ) = arg max over πᵢ of expected utility for player i given the other players’ strategies π₋ᵢ.

Mixed strategies assign probabilities across actions. They are not merely indecision; they can make behaviour unpredictable in settings where predictability would be exploitable.

Nash equilibrium and interpretation

A Nash equilibrium is a strategy profile in which no player can improve unilaterally by deviating.

Equilibrium is a consistency condition. It does not mean the outcome is fair, efficient, stable under learning dynamics or inevitable. Games can have multiple equilibria or none in pure strategies. Which equilibrium becomes relevant can depend on history, conventions, communication and beliefs.

For strategy work, use equilibrium analysis to identify credible responses and strategic traps. Then test the assumptions about rationality, information and available actions. Real competitors have bounded information and organisational constraints; an equilibrium model is a structured scenario, not a guarantee.

Iterative response and fictitious play

Agents may learn by repeatedly responding to observed behaviour rather than solving equilibrium analytically.

A simple learning dynamic forms beliefs about the other agents from their historical actions and repeatedly chooses a best response to those beliefs. Such dynamics can converge in some games and cycle in others. The trajectory itself may matter because temporary behaviours can create real gains or losses before any stable pattern is reached.

Gradient-based multiagent learning changes each agent’s policy in response to its own payoff gradient. Simultaneous adaptation can create non-stationarity: while one agent learns, the environment is changing because others are learning too. Validation should therefore include adaptive opponents rather than only fixed historical behaviour.

Zero-sum and general-sum settings

In a zero-sum game, one player’s gain is the other’s loss; many business interactions are general-sum because cooperation and competition coexist.

Zero-sum structure supports minimax reasoning: choose a strategy that maximises the worst-case payoff against an adversary. This is useful in strongly adversarial security or contest settings. In general-sum interactions, there may be opportunities for coordination, negotiation or mutual gain that minimax reasoning would miss.

Clarify the payoff structure before selecting a solution concept. Treating a potentially cooperative supplier relationship as purely adversarial can destroy value; treating a strategic competitor as passive can expose the business to exploitation.

Sequential interaction and Markov games

A Markov game extends an MDP to several agents whose joint actions determine state transitions and rewards.

At each state, agents choose actions, the environment transitions according to the joint action, and each agent receives its own reward. Policies can depend on the current state, and future strategic interaction affects present action value. This framework can represent repeated pricing, resource competition or autonomous systems sharing an environment.

Learning becomes harder than in a single-agent MDP because the transition experience includes changing opponent policies. Some algorithms extend value learning with equilibrium calculations in each state. Their practical success depends on the game structure, observability and stability of other agents.

Worked example: capacity competition

Two generic providers decide whether to add capacity in a market with uncertain demand.

If one expands while the other does not, the expanding provider may gain share. If both expand, price and utilisation may fall. If neither expands and demand grows, both may face lost opportunities. A one-company forecast that assumes the competitor keeps current capacity can overvalue expansion.

A game model does not tell management exactly what the competitor will do. It identifies conditional payoffs and credible responses. Management can then explore strategic moves such as staged expansion, signalling, differentiated service or flexible capacity that reduce exposure to the competitor’s choice.

Governance and limits

Multiagent models are sensitive to assumptions about objectives and information.

AssumptionQuestion
PayoffHave we represented what the other party actually values?
ActionsAre important strategic options missing?
InformationWhat can each agent observe when choosing?
RationalityWill agents optimise consistently or use organisational heuristics?
AdaptationHow quickly can policies change after observing us?
CommitmentAre announcements or contracts credible and enforceable?

Does Nash equilibrium predict what will happen?

Not automatically. It identifies mutually stable strategies under the model. Multiple equilibria, bounded rationality and incomplete information limit direct prediction.

Should we always assume competitors act optimally?

No, but testing against strong or best-response competitors can reveal vulnerabilities. Scenario analysis can include realistic organisational behaviour as well as adversarial bounds.

Scenario set for strategic decisions

A technically correct method still needs an auditable operating translation.

A useful strategic-game analysis should not stop at one assumed competitor strategy. Build a scenario set that includes a passive response, a plausible organisational response, a strong best response and a delayed response. Recalculate the preferred action under each and identify which assumptions cause the strategy to switch. Also examine second-round effects: a move that is attractive against the first response may provoke a later capacity, price or channel adjustment. Where the preferred action changes across credible scenarios, favour flexibility, staged commitment or information gathering rather than pretending one opponent forecast is certain. Document what real-world signals would indicate that the interaction is moving toward one scenario or another, and assign responsibility for monitoring them after the decision.

Application checklist

  • Use a strategic model only when other agents materially adapt to choices.
  • Define each agent’s actions, information and payoff assumptions.
  • Identify best responses before relying on historical competitor behaviour.
  • Treat equilibrium as a consistency concept, not a guaranteed forecast.
  • Distinguish zero-sum from general-sum opportunities.
  • Use Markov games when strategic interaction evolves through state over time.
  • Validate against adaptive opponents and alternative payoff assumptions.
  • Seek robust strategies that remain acceptable across credible competitor responses.

Related KEVOS knowledge

Collaborative Agents and Decentralised Decision MakingPolicy Validation, Robustness and Rare EventsSequential Decisions and Markov Decision Processes
Source basis. Decision-analysis source set: probabilistic reasoning, sequential decisions, learning, state uncertainty and multiagent methods. This page is an original handbook synthesis of the supplied materials. Named people, organisations and identifying case details from the sources have been removed. Numerical examples are labelled as illustrative where used.

Continue learning

Belief-State Planning: Offline, Online and ControllersGuide · StrategyNEXT LESSON →Collaborative Agents and Decentralised Decision MakingGuide · StrategyState Uncertainty, Belief Updates and FiltersGuide · StrategyExploration, Exploitation and Model LearningGuide · Strategy
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®