Publication & research case study

CMA-MAPPO: combining evolutionary exploration with stable multi-agent learning.

I designed CMA-MAPPO to address a recurring problem in multi-agent reinforcement learning: agents must explore difficult, changing environments while still learning a coordinated policy from sparse or delayed rewards.

Read the published paper ↗View source code ↗
The research question

Why exploration breaks down

In sparse-reward environments, useful feedback may arrive only after a long sequence of actions. In multi-agent settings, the problem is harder: agents must discover promising regions, assign credit across teammates, and avoid collapsing into repetitive or locally optimal behaviors.

MAPPO provides a strong gradient-based baseline for coordinated policy learning, but its exploration can remain limited when the initial agent population does not visit useful parts of the search space. CMA-ES offers derivative-free distribution search, but does not by itself provide the stable shared-policy refinement of PPO.

CMA-MAPPO dual-loop framework showing MAPPO inner loop and CMA-ES outer curriculum loop
The dual-loop design: MAPPO refines the shared policy locally while CMA-ES adapts the initial agent-state curriculum globally.

What I designed

CMA-MAPPO is a hybrid framework with an inner PPO learning cycle and an outer CMA-ES curriculum cycle. The two loops solve complementary parts of the problem instead of forcing one optimizer to do both exploration and policy refinement.

  • Shared policy with agent-ID embeddings: agents use one neural policy for sample efficiency while embeddings allow specialized behaviors.
  • Dual-critic credit assignment: local critics guide individual execution and a centralized joint-state critic provides a global training view.
  • CMA-ES curriculum learning: the outer loop evolves a distribution over initial agent states, guiding agents toward more promising starting regions.
  • Periodic agent renewal: when renewal is triggered, the bottom 50% of agents are removed and reinitialized to preserve diversity and reduce stagnation.
  • Safeguarded evaluation: final performance is evaluated under the true task distribution rather than the curriculum, preventing optimistic reporting.

How one training cycle works

Agents first collect trajectories under the shared MAPPO policy. PPO updates the actor–critic networks using clipped objectives and generalized advantage estimation. After a fixed number of iterations, underperforming agents can be renewed. After a larger PPO interval, short validation rollouts evaluate candidate CMA-ES curriculum distributions; the best candidates update the mean and covariance of the starting-state distribution before the next cycle.

This “return, then explore” strategy lets the inner loop exploit useful behaviors while the outer loop keeps searching for high-potential regions of the environment.

Experimental evaluation

I evaluated the framework across challenging three-dimensional benchmark families, including extreme multi-agent landscapes, shifting optima, rotated Rastrigin, rapid peaks, deceptive landscapes, and noisy Schwefel dynamics. CMA-MAPPO was compared with MAPPO, IPPO, MAPPO-RND, ALP-GMM, GoExplore-Restart, MERL, CMA-ES, and differential evolution.

Normalized performance heatmap comparing CMA-MAPPO with baseline algorithms across six benchmark families
Normalized performance across benchmark families. CMA-MAPPO achieves the strongest overall pattern in the supplied evaluation summary.
Performance versus curriculum ratio across six CMA-MAPPO benchmark environments
Performance sensitivity to the default-versus-curriculum initialization ratio; the 70/30 reference is marked in red.

What the results show

The reported experiments show that CMA-MAPPO is especially effective when exploration is difficult, rewards are sparse, and the environment changes over time. The hybrid framework combines the stability and coordination of MAPPO with the broader search behavior of CMA-ES, while the evaluation safeguards keep the comparison fair.

My contribution covered the algorithm design, curriculum and renewal mechanisms, shared-policy/critic architecture, experimental implementation, benchmark evaluation, and analysis of the resulting performance patterns.