ray/rllib/evaluation/postprocessing.py

import numpy as np
import scipy.signal
from ray.rllib.policy.sample_batch import SampleBatch
from ray.rllib.utils.annotations import DeveloperAPI


def discount(x, gamma):
    return scipy.signal.lfilter([1], [1, -gamma], x[::-1], axis=0)[::-1]


class Postprocessing:
    """Constant definitions for postprocessing."""

    ADVANTAGES = "advantages"
    VALUE_TARGETS = "value_targets"


@DeveloperAPI
def compute_advantages(rollout, last_r, gamma=0.9, lambda_=1.0, use_gae=True):
    """
    Given a rollout, compute its value targets and the advantage.

    Args:
        rollout (SampleBatch): SampleBatch of a single trajectory
        last_r (float): Value estimation for last observation
        gamma (float): Discount factor.
        lambda_ (float): Parameter for GAE
        use_gae (bool): Using Generalized Advantage Estimation

    Returns:
        SampleBatch (SampleBatch): Object with experience from rollout and
            processed rewards.
    """

    traj = {}
    trajsize = len(rollout[SampleBatch.ACTIONS])
    for key in rollout:
        traj[key] = np.stack(rollout[key])

    if use_gae:
        assert SampleBatch.VF_PREDS in rollout, "Values not found!"
        vpred_t = np.concatenate(
            [rollout[SampleBatch.VF_PREDS],
             np.array([last_r])])
        delta_t = (
            traj[SampleBatch.REWARDS] + gamma * vpred_t[1:] - vpred_t[:-1])
        # This formula for the advantage comes from:
        # "Generalized Advantage Estimation": https://arxiv.org/abs/1506.02438
        traj[Postprocessing.ADVANTAGES] = discount(delta_t, gamma * lambda_)
        traj[Postprocessing.VALUE_TARGETS] = (
            traj[Postprocessing.ADVANTAGES] +
            traj[SampleBatch.VF_PREDS]).copy().astype(np.float32)
    else:
        rewards_plus_v = np.concatenate(
            [rollout[SampleBatch.REWARDS],
             np.array([last_r])])
        traj[Postprocessing.ADVANTAGES] = discount(rewards_plus_v, gamma)[:-1]
        # TODO(ekl): support using a critic without GAE
        traj[Postprocessing.VALUE_TARGETS] = np.zeros_like(
            traj[Postprocessing.ADVANTAGES])

    traj[Postprocessing.ADVANTAGES] = traj[
        Postprocessing.ADVANTAGES].copy().astype(np.float32)

    assert all(val.shape[0] == trajsize for val in traj.values()), \
        "Rollout stacked incorrectly!"
    return SampleBatch(traj)
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00			`import numpy as np`
			`import scipy.signal`
[rllib] Rename PolicyGraph => Policy, move from evaluation/ to policy/ (#4819) This implements some of the renames proposed in #4813 We leave behind backwards-compatibility aliases for *PolicyGraph and SampleBatch. 2019-05-20 16:46:05 -07:00			`from ray.rllib.policy.sample_batch import SampleBatch`
[rllib] annotate public vs developer vs private APIs (#3808) 2019-01-23 21:27:26 -08:00			`from ray.rllib.utils.annotations import DeveloperAPI`
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00

			`def discount(x, gamma):`
			`return scipy.signal.lfilter([1], [1, -gamma], x[::-1], axis=0)[::-1]`


Remove (object) from class declarations. (#6658) 2020-01-02 17:42:13 -08:00			`class Postprocessing:`
[rllib] Minor cleanups to TFPolicyGraph: add init args, constants for loss inputs (#4478) 2019-03-29 12:44:23 -07:00			`"""Constant definitions for postprocessing."""`

			`ADVANTAGES = "advantages"`
			`VALUE_TARGETS = "value_targets"`


[rllib] annotate public vs developer vs private APIs (#3808) 2019-01-23 21:27:26 -08:00			`@DeveloperAPI`
[rllib] Extra Changes for Usability (#2363) 2018-07-24 20:51:22 -07:00			`def compute_advantages(rollout, last_r, gamma=0.9, lambda_=1.0, use_gae=True):`
[RLlib] Implement PPO torch version. (#6826) 2020-01-21 08:06:50 +01:00			`"""`
			`Given a rollout, compute its value targets and the advantage.`
[rllib] A3C Configurations (#1370) * initial introduction of a3c configs * fix sample batch * flake but need to check save * save,resotre * fix * pickles * entropy * fix * moving ppo * results * jenkins 2017-12-24 12:25:13 -08:00
			`Args:`
[rllib] Extra Changes for Usability (#2363) 2018-07-24 20:51:22 -07:00			`rollout (SampleBatch): SampleBatch of a single trajectory`
[rllib] Refactor rllib to have a common sample collection pathway (#2149) 2018-06-09 00:21:35 -07:00			`last_r (float): Value estimation for last observation`
[rllib] Extra Changes for Usability (#2363) 2018-07-24 20:51:22 -07:00			`gamma (float): Discount factor.`
[rllib] Evaluators and Optimizers Refactoring (#1339) 2017-12-30 00:24:54 -08:00			`lambda_ (float): Parameter for GAE`
[tune] Reporter crash fix (#5426) Co-authored-by: Richard Liaw <rliaw@berkeley.edu> 2019-08-13 14:10:22 -07:00			`use_gae (bool): Using Generalized Advantage Estimation`
[rllib] A3C Configurations (#1370) * initial introduction of a3c configs * fix sample batch * flake but need to check save * save,resotre * fix * pickles * entropy * fix * moving ppo * results * jenkins 2017-12-24 12:25:13 -08:00
			`Returns:`
			`SampleBatch (SampleBatch): Object with experience from rollout and`
[rllib] Modularize Torch and TF policy graphs (#2294) * wip * cls * re * wip * wip * a3c working * torch support * pg works * lint * rm v2 * consumer id * clean up pg * clean up more * fix python 2.7 * tf session management * docs * dqn wip * fix compile * dqn * apex runs * up * impotrs * ddpg * quotes * fix tests * fix last r * fix tests * lint * pass checkpoint restore * kwar * nits * policy graph * fix yapf * com * class * pyt * vectorization * update * test cpe * unit test * fix ddpg2 * changes * wip * args * faster test * common * fix * add alg option * batch mode and policy serving * multi serving test * todo * wip * serving test * doc async env * num envs * comments * thread * remove init hook * update * fix ppo * comments1 * fix * updates * add jenkins tests * fix * fix pytorch * fix * fixes * fix a3c policy * fix squeeze * fix trunc on apex * fix squeezing for real * update * remove horizon test for now * multiagent wip * update * fix race condition * fix ma * t * doc * st * wip * example * wip * working * cartpole * wip * batch wip * fix bug * make other_batches None default * working * debug * nit * warn * comments * fix ppo * fix obs filter * update * wip * tf * update * fix * cleanup * cleanup * spacing * model * fix * dqn * fix ddpg * doc * keep names * update * fix * com * docs * clarify model outputs * Update torch_policy_graph.py * fix obs filter * pass thru worker index * fix * rename * vlad torch comments * fix log action * debug name * fix lstm * remove unused ddpg net * remove conv net * revert lstm * cast * clean up * fix lstm check * move to end * fix sphinx * fix cmd * remove bad doc * clarify * copy * async sa * fix 2018-06-26 13:17:15 -07:00			`processed rewards.`
			`"""`
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00
			`traj = {}`
[rllib] Minor cleanups to TFPolicyGraph: add init args, constants for loss inputs (#4478) 2019-03-29 12:44:23 -07:00			`trajsize = len(rollout[SampleBatch.ACTIONS])`
[rllib] Add magic methods for rollouts (#2024) 2018-05-16 22:59:46 -07:00			`for key in rollout:`
			`traj[key] = np.stack(rollout[key])`
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00
			`if use_gae:`
[rllib] Minor cleanups to TFPolicyGraph: add init args, constants for loss inputs (#4478) 2019-03-29 12:44:23 -07:00			`assert SampleBatch.VF_PREDS in rollout, "Values not found!"`
			`vpred_t = np.concatenate(`
			`[rollout[SampleBatch.VF_PREDS],`
			`np.array([last_r])])`
			`delta_t = (`
			`traj[SampleBatch.REWARDS] + gamma * vpred_t[1:] - vpred_t[:-1])`
[RLlib] Implement PPO torch version. (#6826) 2020-01-21 08:06:50 +01:00			`# This formula for the advantage comes from:`
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00			`# "Generalized Advantage Estimation": https://arxiv.org/abs/1506.02438`
[rllib] Minor cleanups to TFPolicyGraph: add init args, constants for loss inputs (#4478) 2019-03-29 12:44:23 -07:00			`traj[Postprocessing.ADVANTAGES] = discount(delta_t, gamma * lambda_)`
			`traj[Postprocessing.VALUE_TARGETS] = (`
			`traj[Postprocessing.ADVANTAGES] +`
			`traj[SampleBatch.VF_PREDS]).copy().astype(np.float32)`
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00			`else:`
[rllib] Refactor rllib to have a common sample collection pathway (#2149) 2018-06-09 00:21:35 -07:00			`rewards_plus_v = np.concatenate(`
[rllib] Minor cleanups to TFPolicyGraph: add init args, constants for loss inputs (#4478) 2019-03-29 12:44:23 -07:00			`[rollout[SampleBatch.REWARDS],`
			`np.array([last_r])])`
			`traj[Postprocessing.ADVANTAGES] = discount(rewards_plus_v, gamma)[:-1]`
[rllib] Add debug info back to PPO and fix optimizer compatibility (#2366) 2018-07-12 19:22:46 +02:00			`# TODO(ekl): support using a critic without GAE`
[rllib] Minor cleanups to TFPolicyGraph: add init args, constants for loss inputs (#4478) 2019-03-29 12:44:23 -07:00			`traj[Postprocessing.VALUE_TARGETS] = np.zeros_like(`
			`traj[Postprocessing.ADVANTAGES])`
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00
[rllib] Minor cleanups to TFPolicyGraph: add init args, constants for loss inputs (#4478) 2019-03-29 12:44:23 -07:00			`traj[Postprocessing.ADVANTAGES] = traj[`
			`Postprocessing.ADVANTAGES].copy().astype(np.float32)`
[rllib] A3C Configurations (#1370) * initial introduction of a3c configs * fix sample batch * flake but need to check save * save,resotre * fix * pickles * entropy * fix * moving ppo * results * jenkins 2017-12-24 12:25:13 -08:00
[rllib] PPO and A3C unification (#1253) 2017-12-14 01:08:23 -08:00			`assert all(val.shape[0] == trajsize for val in traj.values()), \`
			`"Rollout stacked incorrectly!"`
[rllib] A3C Configurations (#1370) * initial introduction of a3c configs * fix sample batch * flake but need to check save * save,resotre * fix * pickles * entropy * fix * moving ppo * results * jenkins 2017-12-24 12:25:13 -08:00			`return SampleBatch(traj)`