Sven Mika
|
bdda73e2dd
|
[RLlib] Torch multi-GPU bug fixes (discussion 1755). (#15421)
Thanks a lot @Bam4d for raising this and your help on fixing the worker GPU issue for torch!
|
2021-04-22 11:29:42 +02:00 |
|
Sven Mika
|
41968512ca
|
[RLlib] Partial GPU examples (for learner and workers). (#15334)
|
2021-04-20 08:46:05 +02:00 |
|
Sven Mika
|
cecfc3b43b
|
[RLlib] Multi-GPU support for Torch algorithms. (#14709)
|
2021-04-16 09:16:24 +02:00 |
|
Sven Mika
|
8b3554e37e
|
[RLlib] Remove all (already soft-deprecated) SampleBatch.data from code. (#15335)
|
2021-04-15 19:19:51 +02:00 |
|
mvindiola1
|
5e350ceaa2
|
[RLlib] Issue 14119: Fix TD3 policy delay for torch. (#14840)
|
2021-03-24 16:26:22 +01:00 |
|
Sven Mika
|
69202c6a7d
|
[RLlib] Obsolete usage tracking dict via sample batch. (#13065)
|
2021-03-17 08:18:15 +01:00 |
|
Sven Mika
|
8000258333
|
[RLlib] R2D2 Implementation. (#13933)
|
2021-02-25 12:18:11 +01:00 |
|
Sven Mika
|
95ef04b71a
|
[RLlib] Implement TorchPolicy.export_model . (#13989)
|
2021-02-22 17:09:40 +01:00 |
|
Sven Mika
|
eb0038612f
|
[RLlib] Extend on_learn_on_batch callback to allow for custom metrics to be added. (#13584)
|
2021-02-08 15:02:19 +01:00 |
|
Sven Mika
|
0a0d9183fe
|
[RLlib] Trajectory view API example script (enhancements and tf2 support). (#13786)
|
2021-02-02 18:42:18 +01:00 |
|
Sven Mika
|
1f00f834ac
|
[RLlib] Solve PyTorch/TF-eager A3C async race condition between calling model and its value function. (#13467)
|
2021-01-18 10:29:03 -08:00 |
|
Sven Mika
|
391cdfae8c
|
[RLlib] Trajectory view API docs. (#12718)
|
2020-12-30 17:32:21 -08:00 |
|
Sven Mika
|
a5318961de
|
[RLlib] Preprocessor fixes (multi-discrete) and tests. (#13083)
|
2020-12-26 20:14:36 -05:00 |
|
Sven Mika
|
d5604eaba3
|
[RLlib] Attention nets PyTorch support and cleanup (using traj. view API). (#12029)
|
2020-12-21 18:38:34 -08:00 |
|
Sven Mika
|
b2bcab711d
|
[RLlib] Attention Nets: tf (#12753)
|
2020-12-20 20:22:32 -05:00 |
|
Sven Mika
|
0df55a139c
|
[RLlib] Attention Net prep PR #1: Smaller cleanups. (#12447)
* WIP.
* Fix.
* Fix.
* Fix.
|
2020-11-27 16:25:47 -08:00 |
|
Sven Mika
|
6475297bd3
|
[RLlib] Torch LR schedule not working. Fix and added test case. (#12396)
|
2020-11-26 13:14:11 +01:00 |
|
Sven Mika
|
6da4342822
|
[RLlib] Add on_learn_on_batch (Policy) callback to DefaultCallbacks. (#12070)
|
2020-11-18 15:39:23 +01:00 |
|
Sven Mika
|
62c7ab5182
|
[RLlib] Trajectory view API: Enable by default for PPO, IMPALA, PG, A3C (tf and torch). (#11747)
|
2020-11-12 16:27:34 +01:00 |
|
mvindiola1
|
4518fe790f
|
[RLLIB] Convert torch state arrays to tensors during compute log likelihoods (#11708)
|
2020-11-04 09:33:56 +01:00 |
|
Sven Mika
|
5b788ccb13
|
[RLlib] Trajectory view API (prep PR for switching on by default across all RLlib; plumbing only) (#11717)
|
2020-11-03 12:53:34 -08:00 |
|
Sven Mika
|
54d85a6c2a
|
[RLlib] Fix RNN learning for tf-eager/tf2.x. (#11720)
|
2020-11-02 11:18:41 +01:00 |
|
Sven Mika
|
ce96b03b07
|
[RLlib] MB-MPO cleanup (comments, docstrings, type annotations). (#11033)
|
2020-10-06 20:28:16 +02:00 |
|
Philsik Chang
|
2b26d2ca1b
|
[rllib] Fix for Torch checkpoint taken on GPU fails to deserialize on CPU (#11071) (#11208)
|
2020-10-05 22:01:55 -07:00 |
|
Sven Mika
|
c17169dc11
|
[RLlib] Fix all example scripts to run on GPUs. (#11105)
|
2020-10-02 23:07:44 +02:00 |
|
Sven Mika
|
36bda8432b
|
[RLlib] Trajectory view API: Simple List Collector (on by default for PPO); LSTM-agnostic (#11056)
|
2020-10-01 16:57:10 +02:00 |
|
Sven Mika
|
d7c42d6d92
|
[RLlib] Unity blogpost final fixes. (#10894)
|
2020-09-20 14:13:20 +02:00 |
|
Sven Mika
|
5c7b35d694
|
[RLlib] Issue 10833 TorchPolicy GPU. (#10834)
|
2020-09-17 09:04:46 +02:00 |
|
Michael Luo
|
4e9888ce2f
|
[RLlib] Dreamer (#10172)
|
2020-08-26 13:24:05 +02:00 |
|
Sven Mika
|
e968b52cb7
|
[RLlib] Trajectory view API - 03 Fast LSTM + prev actions/rewards (#9950)
|
2020-08-21 12:35:16 +02:00 |
|
Sven Mika
|
2cbe29a7fa
|
[RLlib] Curiosity minor fixes, do-overs, and testing. (#10143)
|
2020-08-19 17:49:50 +02:00 |
|
Olli Huotari
|
9ff599cbb8
|
torch policy now includes model.metrics (#10121)
* torch policy now includes model.metrics
* Fixed tests to work with custom metrics
* Forgot to run format.sh
|
2020-08-15 10:43:11 -07:00 |
|
Sven Mika
|
2256047876
|
[RLlib] Rename rllib.utils.types into typing to match built-in python module's name. (#10114)
|
2020-08-15 13:24:22 +02:00 |
|
Tanay Wakhare
|
1826b29757
|
[RLlib] Curiosity (intrinsic motivation) Exploration module. (#9912)
|
2020-08-13 20:14:16 +02:00 |
|
yncxcw
|
32cd94b750
|
[Core] Do not convert gpu id to int (#9744)
Co-authored-by: Richard Liaw <rliaw@berkeley.edu>
|
2020-08-11 12:09:46 -07:00 |
|
Sven Mika
|
57690a3a9f
|
[RLlib] Trajectory view API - 02 actual API scaffold (#9753)
|
2020-08-06 10:54:20 +02:00 |
|
Eric Liang
|
5acd3e66dd
|
[rllib] Fix torch TD error, IMPALA LR updates (#9477)
* update
* add test
* lint
* fix super call
* speed es test up
|
2020-07-23 12:50:25 -07:00 |
|
Sven Mika
|
8204717eed
|
[RLlib] Issue 9218: PyTorch Policy places Model on GPU even with num_gpus=0 (#9516)
|
2020-07-17 05:53:25 +02:00 |
|
Sven Mika
|
935d8308fb
|
[RLlib] Issue #9437 (PyTorch converts to CPU tensor, even if on GPU). (#9497)
|
2020-07-16 14:55:50 +02:00 |
|
Sven Mika
|
03ab86567f
|
[RLlib] Layout of Trajectory View API (new class: Trajectory; not used yet). (#9269)
|
2020-07-14 04:27:49 +02:00 |
|
Sven Mika
|
f43d934817
|
[RLlib] Type annotations for policy. (#9248)
|
2020-07-05 13:09:51 +02:00 |
|
Sven Mika
|
0d37103f84
|
[RLlib] Prototype: Model Trajectory View API, part 0 (#9171)
|
2020-06-30 05:33:19 +02:00 |
|
Sven Mika
|
5c6d5d4ab1
|
This PR fixes the currently broken lstm_use_prev_action_reward flag for default lstm models (model.use_lstm=True). (#8970)
|
2020-06-27 20:50:01 +02:00 |
|
Sven Mika
|
af1203b9df
|
[RLlib] Issue 8507 (PyTorch does not support custom loss). (#9142)
|
2020-06-26 09:52:22 +02:00 |
|
Sven Mika
|
14405b90d5
|
[RLlib] Prototype of a DynaTrainer (for env dynamics learning in upcoming MBMPO algo). (#8860)
|
2020-06-16 09:01:20 +02:00 |
|
Sven Mika
|
25c0974543
|
[RLlib] Issue 8412 (Adam vars not stored in ModelV2). (#8480)
|
2020-06-05 21:07:02 +02:00 |
|
Sven Mika
|
d8a081a185
|
[RLlib] Unity3D integration (n Unity3D clients vs learning server). (#8590)
|
2020-05-30 22:48:34 +02:00 |
|
Sven Mika
|
5f278c6411
|
[RLlib] Examples folder restructuring (models) part 1 (#8353)
|
2020-05-08 08:20:18 +02:00 |
|
Sven Mika
|
a00144f746
|
[RLlib] Fix issue 8135 (DDPG inf actions when using [-inf,inf] action space). (#8302)
|
2020-05-04 22:27:30 +02:00 |
|
Tomasz Wrona
|
b508166419
|
Copy initial state of an RNN to a CPU before converting it to a NumPy array (#8097)
|
2020-04-25 18:49:09 -07:00 |
|