Back
VCGS keeps root actions until their policy probability reaches ρ, follows up to B routes each, samples K outcomes at information gates, and scores every route by the average value of its continuations at the turn boundary.