Sven Mika
|
9c5a0cfd7a
|
[RLlib] Issue 14385: Policy.compute_actions_from_input_dict does not properly track accessed fields for Policy's view requirements. (#14386)
|
2021-04-11 18:20:04 +02:00 |
|
Sven Mika
|
95686a8fdd
|
[RLlib] Issue 14533: Tf-eager properly use tree.map_struct on value of type Repeated (RLlib-specific space) (#15015)
|
2021-03-30 19:28:45 +02:00 |
|
Raphael CHEN
|
93d4244d9c
|
[RLlib] Correctly get bytes size of SampleBatch (#14801)
|
2021-03-30 19:24:58 +02:00 |
|
Sven Mika
|
4f66309e19
|
[RLlib] Redo issue 14533 tf enable eager exec (#14984)
|
2021-03-29 20:07:44 +02:00 |
|
SangBin Cho
|
fa5f961d5e
|
Revert "[RLlib] Issue 14533: tf.enable_eager_execution() must be called at beginning. (#14737)" (#14918)
This reverts commit 3e389d5812 .
|
2021-03-25 00:42:01 -07:00 |
|
mvindiola1
|
5e350ceaa2
|
[RLlib] Issue 14119: Fix TD3 policy delay for torch. (#14840)
|
2021-03-24 16:26:22 +01:00 |
|
Sven Mika
|
3e389d5812
|
[RLlib] Issue 14533: tf.enable_eager_execution() must be called at beginning. (#14737)
|
2021-03-24 12:54:27 +01:00 |
|
Sven Mika
|
04bc0a9828
|
[RLlib] Remove all non-trajectory view API code. (#14860)
|
2021-03-23 09:50:18 -07:00 |
|
Sven Mika
|
69202c6a7d
|
[RLlib] Obsolete usage tracking dict via sample batch. (#13065)
|
2021-03-17 08:18:15 +01:00 |
|
Sven Mika
|
ee4b6e7e3b
|
[RLlib] Unity3D example broken due to change in ML-Agents API. Attention-net prev-n-a/r. Attention-wrapper works with images. (#14569)
|
2021-03-12 18:27:25 +01:00 |
|
Sven Mika
|
732197e23a
|
[RLlib] Multi-GPU for tf-DQN/PG/A2C. (#13393)
|
2021-03-08 15:41:27 +01:00 |
|
Sven Mika
|
8000258333
|
[RLlib] R2D2 Implementation. (#13933)
|
2021-02-25 12:18:11 +01:00 |
|
Sven Mika
|
95ef04b71a
|
[RLlib] Implement TorchPolicy.export_model . (#13989)
|
2021-02-22 17:09:40 +01:00 |
|
Sven Mika
|
4db86404ad
|
[RLlib] Issue #13507: Fix MB-MPO CartPole Env's reward function as well as MB-MPO running into a traj. view API related issue. (#14037)
|
2021-02-11 18:58:46 +01:00 |
|
Sven Mika
|
81e7434091
|
[RLlib] TFPolicy.export_model: Add timestep placeholder to model's signature, if needed. (#13988)
|
2021-02-10 15:21:46 +01:00 |
|
Sven Mika
|
d7301a51f4
|
[RLlib]: Trajectory View API: Keep env infos (e.g. for postprocessing callbacks), no matter what. (#13555)
|
2021-02-09 17:05:26 +01:00 |
|
Sven Mika
|
eb0038612f
|
[RLlib] Extend on_learn_on_batch callback to allow for custom metrics to be added. (#13584)
|
2021-02-08 15:02:19 +01:00 |
|
Sven Mika
|
0a0d9183fe
|
[RLlib] Trajectory view API example script (enhancements and tf2 support). (#13786)
|
2021-02-02 18:42:18 +01:00 |
|
Stanislav Chekmenev
|
b9c15a2551
|
[RLlib] Issue #13761: Fix get action shape (#13764)
|
2021-02-02 13:13:43 +01:00 |
|
Sven Mika
|
52c94b7ee9
|
[RLlib] Allow SAC to use custom models as Q- or policy nets and deprecate "state-preprocessor" for image spaces. (#13522)
|
2021-02-02 13:05:58 +01:00 |
|
Sven Mika
|
1f00f834ac
|
[RLlib] Solve PyTorch/TF-eager A3C async race condition between calling model and its value function. (#13467)
|
2021-01-18 10:29:03 -08:00 |
|
Sven Mika
|
d49c3fae0b
|
[RLlib] Trajectory View API: Atari framestacking. (#13315)
|
2021-01-13 08:53:34 +01:00 |
|
Maltimore
|
3a3e4aed86
|
[RLlib] Add __len__() method to SampleBatch (#13371)
|
2021-01-12 20:15:23 +01:00 |
|
Sven Mika
|
6f342a2221
|
[RLlib] Preparatory PR for: Documentation on Model Building. (#13260)
|
2021-01-08 10:56:09 +01:00 |
|
Sven Mika
|
a5b39ef8e2
|
[RLlib] Fix missing "info_batch" arg (None) in compute_actions calls. (#13237)
|
2021-01-07 21:25:02 +01:00 |
|
Sven Mika
|
391cdfae8c
|
[RLlib] Trajectory view API docs. (#12718)
|
2020-12-30 17:32:21 -08:00 |
|
Sven Mika
|
c524f86785
|
[RLlib] BC/MARWIL/recurrent nets minor cleanups and bug fixes. (#13064)
|
2020-12-27 09:46:03 -05:00 |
|
Sven Mika
|
a5318961de
|
[RLlib] Preprocessor fixes (multi-discrete) and tests. (#13083)
|
2020-12-26 20:14:36 -05:00 |
|
Sven Mika
|
99ae7bae05
|
[RLlib] JAXPolicy prep. PR #1. (#13077)
|
2020-12-26 20:14:18 -05:00 |
|
Sven Mika
|
d5604eaba3
|
[RLlib] Attention nets PyTorch support and cleanup (using traj. view API). (#12029)
|
2020-12-21 18:38:34 -08:00 |
|
Sven Mika
|
b2bcab711d
|
[RLlib] Attention Nets: tf (#12753)
|
2020-12-20 20:22:32 -05:00 |
|
Sven Mika
|
74c98ac38e
|
[RLlib] Issue 12244: Unable to restore multi-agent PPOTFPolicy's Model (from exported). (#12786)
|
2020-12-11 16:13:38 +01:00 |
|
Sven Mika
|
a082ea18b8
|
[RLlib] Issue 12212: "TFEagerPolicy has no attribute action_sampler_fn.
|
2020-12-11 12:57:33 +01:00 |
|
Sven Mika
|
28108c905b
|
[RLlib] Tf-eager policy bug fix: Duplicate model call in compute_gradients. (#12682)
|
2020-12-09 08:03:58 +01:00 |
|
Sven Mika
|
e40b14d255
|
[RLlib] Batch-size for truncate_episode batch_mode should be confgurable in agent-steps (rather than env-steps), if needed. (#12420)
|
2020-12-08 16:41:45 -08:00 |
|
Sven Mika
|
99c81c6795
|
[RLlib] Attention Net prep PR #3. (#12450)
|
2020-12-07 13:08:17 +01:00 |
|
Sven Mika
|
19c8033df2
|
[RLlib] Fix most remaining RLlib algos for running with trajectory view API. (#12366)
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* WIP.
* LINT and fixes.
MB-MPO and MAML not working yet.
* wip
* update
* update
* rmeove
* remove dep
* higher
* Update requirements_rllib.txt
* Update requirements_rllib.txt
* relpos
* no mbmpo
Co-authored-by: Eric Liang <ekhliang@gmail.com>
|
2020-12-01 17:41:10 -08:00 |
|
Sven Mika
|
9021f15b2a
|
[RLlib] Fix setup-dev.py error when creating a softlink for new_dashboard. (#12442)
|
2020-12-01 11:46:59 +01:00 |
|
Sven Mika
|
3ad9365e1d
|
[RLlib] Attention Net prep PR #2: Smaller cleanups. (#12449)
|
2020-12-01 08:21:45 +01:00 |
|
Sven Mika
|
fb318addcb
|
[RLlib] Curiosity exploration module: tf/tf2.x/tf-eager support. (#11945)
|
2020-11-29 12:31:24 +01:00 |
|
Sven Mika
|
0df55a139c
|
[RLlib] Attention Net prep PR #1: Smaller cleanups. (#12447)
* WIP.
* Fix.
* Fix.
* Fix.
|
2020-11-27 16:25:47 -08:00 |
|
Sven Mika
|
6475297bd3
|
[RLlib] Torch LR schedule not working. Fix and added test case. (#12396)
|
2020-11-26 13:14:11 +01:00 |
|
Sven Mika
|
6da4342822
|
[RLlib] Add on_learn_on_batch (Policy) callback to DefaultCallbacks. (#12070)
|
2020-11-18 15:39:23 +01:00 |
|
Sven Mika
|
b6b54f1c81
|
[RLlib] Trajectory view API: enable by default for SAC, DDPG, DQN, SimpleQ (#11827)
|
2020-11-16 10:54:35 -08:00 |
|
Sven Mika
|
62c7ab5182
|
[RLlib] Trajectory view API: Enable by default for PPO, IMPALA, PG, A3C (tf and torch). (#11747)
|
2020-11-12 16:27:34 +01:00 |
|
mvindiola1
|
4518fe790f
|
[RLLIB] Convert torch state arrays to tensors during compute log likelihoods (#11708)
|
2020-11-04 09:33:56 +01:00 |
|
Sven Mika
|
5b788ccb13
|
[RLlib] Trajectory view API (prep PR for switching on by default across all RLlib; plumbing only) (#11717)
|
2020-11-03 12:53:34 -08:00 |
|
Sven Mika
|
54d85a6c2a
|
[RLlib] Fix RNN learning for tf-eager/tf2.x. (#11720)
|
2020-11-02 11:18:41 +01:00 |
|
Sven Mika
|
d9f1874e34
|
[RLlib] Minor fixes (torch GPU bugs + some cleanup). (#11609)
|
2020-10-27 10:00:24 +01:00 |
|
Sven Mika
|
0c0f67c14d
|
[RLlib] ARS/ES eval workers not working: Issue 9933. (#11308)
|
2020-10-12 13:49:48 -07:00 |
|