Why exploration breaks down
In sparse-reward environments, useful feedback may arrive only after a long sequence of actions. In multi-agent settings, the problem is harder: agents must discover promising regions, assign credit across teammates, and avoid collapsing into repetitive or locally optimal behaviors.
MAPPO provides a strong gradient-based baseline for coordinated policy learning, but its exploration can remain limited when the initial agent population does not visit useful parts of the search space. CMA-ES offers derivative-free distribution search, but does not by itself provide the stable shared-policy refinement of PPO.

What I designed
CMA-MAPPO is a hybrid framework with an inner PPO learning cycle and an outer CMA-ES curriculum cycle. The two loops solve complementary parts of the problem instead of forcing one optimizer to do both exploration and policy refinement.
- Shared policy with agent-ID embeddings: agents use one neural policy for sample efficiency while embeddings allow specialized behaviors.
- Dual-critic credit assignment: local critics guide individual execution and a centralized joint-state critic provides a global training view.
- CMA-ES curriculum learning: the outer loop evolves a distribution over initial agent states, guiding agents toward more promising starting regions.
- Periodic agent renewal: when renewal is triggered, the bottom 50% of agents are removed and reinitialized to preserve diversity and reduce stagnation.
- Safeguarded evaluation: final performance is evaluated under the true task distribution rather than the curriculum, preventing optimistic reporting.
How one training cycle works
Agents first collect trajectories under the shared MAPPO policy. PPO updates the actor–critic networks using clipped objectives and generalized advantage estimation. After a fixed number of iterations, underperforming agents can be renewed. After a larger PPO interval, short validation rollouts evaluate candidate CMA-ES curriculum distributions; the best candidates update the mean and covariance of the starting-state distribution before the next cycle.
This “return, then explore” strategy lets the inner loop exploit useful behaviors while the outer loop keeps searching for high-potential regions of the environment.
Experimental evaluation
I evaluated the framework across challenging three-dimensional benchmark families, including extreme multi-agent landscapes, shifting optima, rotated Rastrigin, rapid peaks, deceptive landscapes, and noisy Schwefel dynamics. CMA-MAPPO was compared with MAPPO, IPPO, MAPPO-RND, ALP-GMM, GoExplore-Restart, MERL, CMA-ES, and differential evolution.


What the results show
The reported experiments show that CMA-MAPPO is especially effective when exploration is difficult, rewards are sparse, and the environment changes over time. The hybrid framework combines the stability and coordination of MAPPO with the broader search behavior of CMA-ES, while the evaluation safeguards keep the comparison fair.
My contribution covered the algorithm design, curriculum and renewal mechanisms, shared-policy/critic architecture, experimental implementation, benchmark evaluation, and analysis of the resulting performance patterns.