advantage actor-critic

**A2C** (Advantage Actor-Critic) is a **synchronous policy gradient algorithm where an actor (policy) and critic (value function) are trained simultaneously** — the actor decides which actions to take, and the critic evaluates how good those actions are relative to expectation. **A2C Architecture** - **Actor**: Policy network $pi_ heta(a|s)$ — outputs action probabilities or continuous parameters. - **Critic**: Value network $V_phi(s)$ — estimates expected return from state $s$. - **Advantage**: $A_t = R_t - V_phi(s_t)$ — how much better the actual return was than expected. - **Loss**: Actor loss = $-log pi_ heta(a_t|s_t) A_t$; Critic loss = $(R_t - V_phi(s_t))^2$. **Why It Matters** - **Variance Reduction**: Using the baseline $V(s)$ reduces variance compared to REINFORCE — faster learning. - **Synchronous**: Unlike A3C, A2C is synchronous — deterministic, reproducible, and GPU-friendly. - **Foundation**: A2C is the building block for PPO, SAC, and most modern policy gradient methods. **A2C** is **the actor-critic baseline** — combining policy and value learning for stable, low-variance policy gradient training.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account