ray/rllib/agents/ppo/test/test.py

from __future__ import absolute_import
from __future__ import division
from __future__ import print_function

import unittest
import numpy as np
from numpy.testing import assert_allclose

from ray.rllib.models.tf.tf_action_dist import Categorical
from ray.rllib.agents.ppo.utils import flatten, concatenate
from ray.rllib.utils import try_import_tf

tf = try_import_tf()


# TODO(ekl): move to rllib/models dir
class DistributionsTest(unittest.TestCase):
    def testCategorical(self):
        num_samples = 100000
        logits = tf.placeholder(tf.float32, shape=(None, 10))
        z = 8 * (np.random.rand(10) - 0.5)
        data = np.tile(z, (num_samples, 1))
        c = Categorical(logits, {})  # dummy config dict
        sample_op = c.sample()
        sess = tf.Session()
        sess.run(tf.global_variables_initializer())
        samples = sess.run(sample_op, feed_dict={logits: data})
        counts = np.zeros(10)
        for sample in samples:
            counts[sample] += 1.0
        probs = np.exp(z) / np.sum(np.exp(z))
        self.assertTrue(np.sum(np.abs(probs - counts / num_samples)) <= 0.01)


class UtilsTest(unittest.TestCase):
    def testFlatten(self):
        d = {
            "s": np.array([[[1, -1], [2, -2]], [[3, -3], [4, -4]]]),
            "a": np.array([[[5], [-5]], [[6], [-6]]])
        }
        flat = flatten(d.copy(), start=0, stop=2)
        assert_allclose(d["s"][0][0][:], flat["s"][0][:])
        assert_allclose(d["s"][0][1][:], flat["s"][1][:])
        assert_allclose(d["s"][1][0][:], flat["s"][2][:])
        assert_allclose(d["s"][1][1][:], flat["s"][3][:])
        assert_allclose(d["a"][0][0], flat["a"][0])
        assert_allclose(d["a"][0][1], flat["a"][1])
        assert_allclose(d["a"][1][0], flat["a"][2])
        assert_allclose(d["a"][1][1], flat["a"][3])

    def testConcatenate(self):
        d1 = {"s": np.array([0, 1]), "a": np.array([2, 3])}
        d2 = {"s": np.array([4, 5]), "a": np.array([6, 7])}
        d = concatenate([d1, d2])
        assert_allclose(d["s"], np.array([0, 1, 4, 5]))
        assert_allclose(d["a"], np.array([2, 3, 6, 7]))

        D = concatenate([d])
        assert_allclose(D["s"], np.array([0, 1, 4, 5]))
        assert_allclose(D["a"], np.array([2, 3, 6, 7]))


if __name__ == "__main__":
    unittest.main(verbosity=2)
Add policy gradient example. (#344) * add policy gradient example * fix typos * Minor changes plus some documentation. * Minor fixes. 2017-03-07 23:42:44 -08:00			`from __future__ import absolute_import`
			`from __future__ import division`
			`from __future__ import print_function`

			`import unittest`
			`import numpy as np`
			`from numpy.testing import assert_allclose`

[rllib] Document ModelV2 and clean up the models/ directory (#5277) 2019-07-27 02:08:16 -07:00			`from ray.rllib.models.tf.tf_action_dist import Categorical`
[rllib] Document "v2" APIs (#2316) * re * wip * wip * a3c working * torch support * pg works * lint * rm v2 * consumer id * clean up pg * clean up more * fix python 2.7 * tf session management * docs * dqn wip * fix compile * dqn * apex runs * up * impotrs * ddpg * quotes * fix tests * fix last r * fix tests * lint * pass checkpoint restore * kwar * nits * policy graph * fix yapf * com * class * pyt * vectorization * update * test cpe * unit test * fix ddpg2 * changes * wip * args * faster test * common * fix * add alg option * batch mode and policy serving * multi serving test * todo * wip * serving test * doc async env * num envs * comments * thread * remove init hook * update * fix ppo * comments1 * fix * updates * add jenkins tests * fix * fix pytorch * fix * fixes * fix a3c policy * fix squeeze * fix trunc on apex * fix squeezing for real * update * remove horizon test for now * multiagent wip * update * fix race condition * fix ma * t * doc * st * wip * example * wip * working * cartpole * wip * batch wip * fix bug * make other_batches None default * working * debug * nit * warn * comments * fix ppo * fix obs filter * update * wip * tf * update * fix * cleanup * cleanup * spacing * model * fix * dqn * fix ddpg * doc * keep names * update * fix * com * docs * clarify model outputs * Update torch_policy_graph.py * fix obs filter * pass thru worker index * fix * rename * vlad torch comments * fix log action * debug name * fix lstm * remove unused ddpg net * remove conv net * revert lstm * wip * wip * cast * wip * works * fix a3c * works * lstm util test * doc * clean up * update * fix lstm check * move to end * fix sphinx * fix cmd * remove bad doc * envs * vec * doc prep * models * rl * alg * up * clarify * copy * async sa * fix * comments * fix a3c conf * tune lstm * fix reshape * fix * back to 16 * tuned a3c update * update * tuned * optional * merge * wip * fix up * move pg class * rename env * wip * update * tip * alg * readme * fix catalog * readme * doc * context * remove prep * comma * add env * link to paper * paper * update * rnn * update * wip * clean up ev creation * fix * fix * fix * fix lint * up * no comma * ma * Update run_multi_node_tests.sh * fix * sphinx is stupid * sphinx is stupid * clarify torch graph * no horizon * fix config * sb * Update test_optimizers.py 2018-07-01 00:05:08 -07:00			`from ray.rllib.agents.ppo.utils import flatten, concatenate`
[rllib] TensorFlow 2 compatibility (#4802) 2019-05-16 22:12:07 -07:00			`from ray.rllib.utils import try_import_tf`

			`tf = try_import_tf()`
Add policy gradient example. (#344) * add policy gradient example * fix typos * Minor changes plus some documentation. * Minor fixes. 2017-03-07 23:42:44 -08:00
Make example applications pep8 compliant. (#553) * Test examples for pep8 compliance. * Make rl_pong example pep8 compliant. * Make policy gradient example pep8 compliant. * Make lbfgs example pep8 compliant. * Make hyperopt example pep8 compliant. * Make a3c example pep8 compliant. * Make evolution strategies example pep8 compliant. * Make resnet example pep8 compliant. * Fix. 2017-05-16 14:12:18 -07:00
[rllib] Pull out shared models for evolution strategies and policy gradient. (#719) * wip * works with cartpole * lint * fix pg * comment * action dist rename * preprocessor * fix test * typo * fix the action[0] nonsense * revert * satisfy the lint * wip * works with cartpole * lint * fix pg * comment * action dist rename * preprocessor * fix test * typo * fix the action[0] nonsense * revert * satisfy the lint * Minor indentation changes. * fix merge * add humanoid * fix linting * more 4 space * fix * fix linT * oops * es parity 2017-07-17 01:58:54 -07:00			`# TODO(ekl): move to rllib/models dir`
			`class DistributionsTest(unittest.TestCase):`
Switch Python indentation from 2 spaces to 4 spaces. (#726) * 4 space indentation for actor.py. * 4 space indentation for worker.py. * 4 space indentation for more files. * 4 space indentation for some test files. * Check indentation in Travis. * 4 space indentation for some rl files. * Fix failure test. * Fix multi_node_test. * 4 space indentation for more files. * 4 space indentation for remaining files. * Fixes. 2017-07-13 14:53:57 -07:00			`def testCategorical(self):`
			`num_samples = 100000`
			`logits = tf.placeholder(tf.float32, shape=(None, 10))`
			`z = 8 * (np.random.rand(10) - 0.5)`
			`data = np.tile(z, (num_samples, 1))`
Custom action distributions (#5164) * custom action dist wip * Test case for custom action dist * ActionDistribution.get_parameter_shape_for_action_space pattern * Edit exception message to also suggest using a custom action distribution * Clean up ModelCatalog.get_action_dist * Pass model config to ActionDistribution constructors * Update custom action distribution test case * Name fix * Autoformatter * parameter shape static methods for torch distributions * Fix docstring * Generalize fake array for graph initialization * Fix action dist constructors * Correct parameter shape static methods for multicategorical and gaussian * Make suggested changes to custom action dist's * Correct instances of not passing model config to action dist * Autoformatter * fix tuple distribution constructor * bugfix 2019-08-06 18:13:16 +00:00			`c = Categorical(logits, {}) # dummy config dict`
Switch Python indentation from 2 spaces to 4 spaces. (#726) * 4 space indentation for actor.py. * 4 space indentation for worker.py. * 4 space indentation for more files. * 4 space indentation for some test files. * Check indentation in Travis. * 4 space indentation for some rl files. * Fix failure test. * Fix multi_node_test. * 4 space indentation for more files. * 4 space indentation for remaining files. * Fixes. 2017-07-13 14:53:57 -07:00			`sample_op = c.sample()`
			`sess = tf.Session()`
			`sess.run(tf.global_variables_initializer())`
			`samples = sess.run(sample_op, feed_dict={logits: data})`
			`counts = np.zeros(10)`
			`for sample in samples:`
			`counts[sample] += 1.0`
			`probs = np.exp(z) / np.sum(np.exp(z))`
			`self.assertTrue(np.sum(np.abs(probs - counts / num_samples)) <= 0.01)`
Add policy gradient example. (#344) * add policy gradient example * fix typos * Minor changes plus some documentation. * Minor fixes. 2017-03-07 23:42:44 -08:00
Make example applications pep8 compliant. (#553) * Test examples for pep8 compliance. * Make rl_pong example pep8 compliant. * Make policy gradient example pep8 compliant. * Make lbfgs example pep8 compliant. * Make hyperopt example pep8 compliant. * Make a3c example pep8 compliant. * Make evolution strategies example pep8 compliant. * Make resnet example pep8 compliant. * Fix. 2017-05-16 14:12:18 -07:00
Add policy gradient example. (#344) * add policy gradient example * fix typos * Minor changes plus some documentation. * Minor fixes. 2017-03-07 23:42:44 -08:00			`class UtilsTest(unittest.TestCase):`
Switch Python indentation from 2 spaces to 4 spaces. (#726) * 4 space indentation for actor.py. * 4 space indentation for worker.py. * 4 space indentation for more files. * 4 space indentation for some test files. * Check indentation in Travis. * 4 space indentation for some rl files. * Fix failure test. * Fix multi_node_test. * 4 space indentation for more files. * 4 space indentation for remaining files. * Fixes. 2017-07-13 14:53:57 -07:00			`def testFlatten(self):`
[rllib] format with yapf (#2427) * initial yapf * manual fix yapf bugs 2018-07-19 15:30:36 -07:00			`d = {`
			`"s": np.array([[[1, -1], [2, -2]], [[3, -3], [4, -4]]]),`
			`"a": np.array([[[5], [-5]], [[6], [-6]]])`
			`}`
Switch Python indentation from 2 spaces to 4 spaces. (#726) * 4 space indentation for actor.py. * 4 space indentation for worker.py. * 4 space indentation for more files. * 4 space indentation for some test files. * Check indentation in Travis. * 4 space indentation for some rl files. * Fix failure test. * Fix multi_node_test. * 4 space indentation for more files. * 4 space indentation for remaining files. * Fixes. 2017-07-13 14:53:57 -07:00			`flat = flatten(d.copy(), start=0, stop=2)`
			`assert_allclose(d["s"][0][0][:], flat["s"][0][:])`
			`assert_allclose(d["s"][0][1][:], flat["s"][1][:])`
			`assert_allclose(d["s"][1][0][:], flat["s"][2][:])`
			`assert_allclose(d["s"][1][1][:], flat["s"][3][:])`
			`assert_allclose(d["a"][0][0], flat["a"][0])`
			`assert_allclose(d["a"][0][1], flat["a"][1])`
			`assert_allclose(d["a"][1][0], flat["a"][2])`
			`assert_allclose(d["a"][1][1], flat["a"][3])`

			`def testConcatenate(self):`
			`d1 = {"s": np.array([0, 1]), "a": np.array([2, 3])}`
			`d2 = {"s": np.array([4, 5]), "a": np.array([6, 7])}`
			`d = concatenate([d1, d2])`
			`assert_allclose(d["s"], np.array([0, 1, 4, 5]))`
			`assert_allclose(d["a"], np.array([2, 3, 6, 7]))`

			`D = concatenate([d])`
			`assert_allclose(D["s"], np.array([0, 1, 4, 5]))`
			`assert_allclose(D["a"], np.array([2, 3, 6, 7]))`
Add policy gradient example. (#344) * add policy gradient example * fix typos * Minor changes plus some documentation. * Minor fixes. 2017-03-07 23:42:44 -08:00
Make example applications pep8 compliant. (#553) * Test examples for pep8 compliance. * Make rl_pong example pep8 compliant. * Make policy gradient example pep8 compliant. * Make lbfgs example pep8 compliant. * Make hyperopt example pep8 compliant. * Make a3c example pep8 compliant. * Make evolution strategies example pep8 compliant. * Make resnet example pep8 compliant. * Fix. 2017-05-16 14:12:18 -07:00
Add policy gradient example. (#344) * add policy gradient example * fix typos * Minor changes plus some documentation. * Minor fixes. 2017-03-07 23:42:44 -08:00			`if __name__ == "__main__":`
Switch Python indentation from 2 spaces to 4 spaces. (#726) * 4 space indentation for actor.py. * 4 space indentation for worker.py. * 4 space indentation for more files. * 4 space indentation for some test files. * Check indentation in Travis. * 4 space indentation for some rl files. * Fix failure test. * Fix multi_node_test. * 4 space indentation for more files. * 4 space indentation for remaining files. * Fixes. 2017-07-13 14:53:57 -07:00			`unittest.main(verbosity=2)`