hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-06 02:21:39 -05:00

Author	SHA1	Message	Date
Sven Mika	36bda8432b	[RLlib] Trajectory view API: Simple List Collector (on by default for PPO); LSTM-agnostic (#11056 )	2020-10-01 16:57:10 +02:00
Eric Liang	daa03ba6e6	[rllib] Add execution module to package ref (#10941 ) * add init * add * update	2020-09-21 23:03:06 -07:00
Sven Mika	d7c42d6d92	[RLlib] Unity blogpost final fixes. (#10894 )	2020-09-20 14:13:20 +02:00
Sven Mika	805dad3bc4	[RLlib] SAC algo cleanup. (#10825 )	2020-09-20 11:27:02 +02:00
Benjamin Black	f2408b719c	Fixed PettingZooEnv (#10847 )	2020-09-17 11:28:42 -07:00
Sven Mika	4b278c36fc	[RLlib] Behavioral Cloning (from MARWIL). (#10619 )	2020-09-09 17:33:21 +02:00
Michael Luo	8e613652af	[RLLib] MBMPO Fixes (#10296 )	2020-09-09 09:34:34 +02:00
Sven Mika	28ab797cf5	[RLlib] Deprecate old classes, methods, functions, config keys (in prep for RLlib 1.0). (#10544 )	2020-09-06 10:58:00 +02:00
Richard Liaw	551c597312	[tune] API revamp fix (#10518 )	2020-09-05 15:34:53 -07:00
Sven Mika	244aafdcf8	[RLlib] Curiosity enhancements. (#10373 )	2020-09-05 13:14:24 +02:00
Justin Terry	352718610d	Multi-agent Algorithm Documentation Updates (#9722 )	2020-09-03 22:37:46 -07:00
Sven Mika	715ee8dfc9	[RLlib] Issue 10469: Callbacks should receive env idx ... (#10477 )	2020-09-03 17:27:05 +02:00
Sven Mika	ef18893fb5	[RLlib] PPO, APPO, and DD-PPO code cleanup. (#10420 )	2020-09-02 14:03:01 +02:00
Michael Luo	4e9888ce2f	[RLlib] Dreamer (#10172 )	2020-08-26 13:24:05 +02:00
Eric Liang	deea1861ab	[rllib] Try fixing torch GPU and masking errors (#10168 )	2020-08-25 18:34:19 -07:00
Benjamin Black	2689fb439c	Fixed pettingzoo env example (#9973 )	2020-08-25 13:22:25 +02:00
Michael Luo	48a39d7cb9	[RLlib] Deepmind Control Suite Examples (#9751 )	2020-08-23 12:53:08 +02:00
Sven Mika	e968b52cb7	[RLlib] Trajectory view API - 03 Fast LSTM + prev actions/rewards (#9950 )	2020-08-21 12:35:16 +02:00
Eric Liang	ca133e2699	[rllib] Remove extra model config kwargs passed incorrectly for Torch models (#10055 )	2020-08-17 11:12:20 -07:00
Olli Huotari	9ff599cbb8	torch policy now includes model.metrics (#10121 ) * torch policy now includes model.metrics * Fixed tests to work with custom metrics * Forgot to run format.sh	2020-08-15 10:43:11 -07:00
Barak Michener	8e76796fd0	ci: Redo `format.sh --all` script & backfill lint fixes (#9956 )	2020-08-07 16:49:49 -07:00
Michael Luo	4d7bd8c892	[RLlib] Implementation of "Model-based Meta Policy Optimization" (MB MPO) (#9409 )	2020-08-02 18:12:09 +02:00
Sven Mika	b0b0463161	[RLlib] Trajectory View API (preparatory cleanup and enhancements). (#9678 )	2020-07-29 21:15:09 +02:00
Sven Mika	e6ea33a03c	[RLlib] Enhance reward clipping test; add action_clipping tests. (#9684 )	2020-07-28 10:44:54 +02:00
Sven Mika	5dc4b6686e	[RLlib] Implement DQN PyTorch distributional head. (#9589 )	2020-07-25 09:29:24 +02:00
Sven Mika	78dfed2683	[RLlib] Issue 8384: QMIX doesn't learn anything. (#9527 )	2020-07-17 12:14:34 +02:00
Sven Mika	935d8308fb	[RLlib] Issue #9437 (PyTorch converts to CPU tensor, even if on GPU). (#9497 )	2020-07-16 14:55:50 +02:00
Sven Mika	617eb8f279	[RLlib] Issue 9402 MARWIL producing nan rewards. (#9429 )	2020-07-14 05:07:16 +02:00
Sven Mika	14160ca58c	[RLlib] Issue #9366 (DQN w/o dueling produces invalid actions). (#9386 )	2020-07-10 12:43:03 +02:00
Benjamin Black	1425cdf834	Pettingzoo environment support (#9271 ) * added pettingzoo wrapper env and example * added docs, examples for pettingzoo env support * fixed pettingzoo env flake8, added test * fixed pettingzoo env import * fixed pettingzoo env import * fixed pettingzoo import issue * fixed pettingzoo test * fixed linting problem * fixed bad quotes * future proofed pettingzoo dependency * fixed ray init in pettingzoo env * lint * manual lint Co-authored-by: Eric Liang <ekhliang@gmail.com>	2020-07-06 21:32:26 -07:00
Sven Mika	f43d934817	[RLlib] Type annotations for policy. (#9248 )	2020-07-05 13:09:51 +02:00
Michael Luo	851d02463b	[Doc] RLlib Algorithms Documentation: MAML + PyTorch MAML (#9189 )	2020-07-03 11:05:15 -07:00
Sven Mika	5b2a97597b	[RLlib] Retire `try_import_tree` (should be installed along with other requirements). (#9211 ) - Retire try_import_tree. - Stabilize test_supported_multi_agent.py.	2020-07-02 13:06:34 +02:00
Sven Mika	43043ee4d5	[RLlib] Tf2x preparation; part 2 (upgrading `try_import_tf()`). (#9136 ) * WIP. * Fixes. * LINT. * WIP. * WIP. * Fixes. * Fixes. * Fixes. * Fixes. * WIP. * Fixes. * Test * Fix. * Fixes and LINT. * Fixes and LINT. * LINT.	2020-06-30 10:13:20 +02:00
Sven Mika	5c6d5d4ab1	This PR fixes the currently broken lstm_use_prev_action_reward flag for default lstm models (model.use_lstm=True). (#8970 )	2020-06-27 20:50:01 +02:00
Sven Mika	af1203b9df	[RLlib] Issue 8507 (PyTorch does not support custom loss). (#9142 )	2020-06-26 09:52:22 +02:00
Sven Mika	4fd8977eaf	[RLlib] Minor cleanup in preparation to tf2.x support. (#9130 ) * WIP. * Fixes. * LINT. * Fixes. * Fixes and LINT. * WIP.	2020-06-25 19:01:32 +02:00
Michael Luo	cf0894d396	[rllib] MAML Agent (#8862 ) * Halfway done with transferring MAML to new Ray * MAML Beta Out * Debugging MAML atm * Distributed Execution * Pendulum Mass Working * All experiments complete * Cleaned up codebase * Travis CI * Travis CI * Tests * Merged conflicts * Fixed variance bug conflict * Comment resolved * Apply suggestions from code review fixed test_maml * Update rllib/agents/maml/tests/test_maml.py * asdf * Fix testing Co-authored-by: Sven Mika <sven@anyscale.io>	2020-06-23 09:48:23 -07:00
Sven Mika	14405b90d5	[RLlib] Prototype of a DynaTrainer (for env dynamics learning in upcoming MBMPO algo). (#8860 )	2020-06-16 09:01:20 +02:00
Sven Mika	7008902cff	[RLlib] Minor `rllib.utils` cleanup. (#8932 )	2020-06-16 08:52:20 +02:00
Sven Mika	a90cd0fcbb	[RLlib] Unity3d soccer benchmarks (#8834 )	2020-06-11 14:29:57 +02:00
Sven Mika	0ba7472da9	[Testing] Fix LINT/sphinx errors. (#8874 )	2020-06-10 15:41:59 +02:00
Eric Liang	be26a7b1b0	[rllib] Support for complex / variable-length observation spaces (#8393 )	2020-06-06 12:22:19 +02:00
Sven Mika	c74dc58f8b	[RLlib] Fix `use_lstm` flag for ModelV2 (w/o ModelV1 wrapping) and add it for PyTorch. (#8734 )	2020-06-05 15:40:30 +02:00
Sven Mika	d8a081a185	[RLlib] Unity3D integration (n Unity3D clients vs learning server). (#8590 )	2020-05-30 22:48:34 +02:00
Sven Mika	2746fc0476	[RLlib] Auto-framework, retire `use_pytorch` in favor of `framework=...` (#8520 )	2020-05-27 16:19:13 +02:00
Sven Mika	6d196197bc	[RLlib] utils/spaces ... (#8608 )	2020-05-27 10:21:30 +02:00
Sven Mika	0422e9c5a8	[RLlib] Add 2 Transformer learning test cases on StatelessCartPole (PPO and IMPALA). (#8624 )	2020-05-27 10:19:47 +02:00
Eric Liang	9a83908c46	[rllib] Deprecate policy optimizers (#8345 )	2020-05-21 10:16:18 -07:00
Sven Mika	796a834c48	[RLlib] Attention Net integration into ModelV2 and learning RL example. (#8371 )	2020-05-18 17:26:40 +02:00

1 2 3

104 commits