mwtian
1dd8b3d2bc
[Build] Remove debug info from Ray libraries. ( #20389 )
...
## Why are these changes needed?
Ray wheel size limit is still at 100MB. Removing debug symbols would decrease Ray Linux wheel sizes.
## Related issue number
## Checks
2021-11-15 16:40:48 -08:00
Amog Kamsetty
90dc5460d4
Revert "[RLlib] POC: Deprecate build_policy
(policy template) for torch only; PPOTorchPolicy ( #20061 )" ( #20399 )
...
This reverts commit 5b1c8e46e1
.
2021-11-15 16:11:35 -08:00
iasoon
171ad62e30
Link to the documentation on contributing from CONTRIBUTING.rst ( #19396 )
...
Co-authored-by: Richard Liaw <rliaw@berkeley.edu>
2021-11-15 15:34:18 -08:00
Antoni Baum
ec81f52061
[Docs] Fix typo in C++ Placement Group example ( #20386 )
2021-11-16 08:19:09 +09:00
matthewdeng
35dc3cf21b
[train] fix Train/Tune integration on Client ( #20351 )
...
* [train] fix Train/Tune integration on Client
* remove force_on_current_node
2021-11-15 14:36:33 -08:00
Alex Wu
884bb3de33
[Dataset] Bump numpy >=1.20
dependency ( #20374 )
...
* done?
* .
Co-authored-by: Alex Wu <alex@anyscale.com>
2021-11-15 14:10:00 -08:00
Kai Fricke
d191ad2de8
[ci/release] Return exit codes based on different errors ( #20289 )
2021-11-15 19:41:00 +00:00
Simon Mo
72ae22e82b
[CI] Fix frontend build issue ( #20375 )
2021-11-15 10:12:43 -08:00
Kai Fricke
91920f1d02
[release/xgboost] xgboost release test fixes via app config ( #20325 )
...
* [xgboost] Fix release test app configs
* Revert full app config
* Update base docker image
* Only change cpu base image
* default
* Pin xgboost to 1.5. in cpu tests
* Remove numpy hack
* Revert one line
Co-authored-by: Amog Kamsetty <amogkamsetty@yahoo.com>
2021-11-15 10:03:21 -08:00
Amog Kamsetty
ef7967476c
[Train] Torch data transfer automatic conversion ( #20333 )
...
* update
* formatting
* fix failures
* fix session tests
* address comments
* add to api docs
* package refactor
* wip
* wip
* wip
* finish
* finish
* fix
* comment
* fix
* install horovod for docs
* address comment
* Update python/ray/train/session.py
Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>
* Update python/ray/train/torch.py
Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>
* address comments
* try fix docs
* fix doc build failure
* wip
* fix
* fix
* fix
* try fix doc highlighting
* fix docs
* finish
* formatting
* address comments and fix tests
* address comments and fix test
Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>
2021-11-15 09:14:12 -08:00
Lixin Wei
85dbda8cf1
[Core] Fix Crash in Debug Log ( #20322 )
...
* fix crash in debug log
* fix
* fix
* fix
2021-11-16 00:45:00 +09:00
Lixin Wei
b7e35acf14
[RuntimeEnv] Raise RuntimeEnvSetupError when Actor Creation Failed due to It ( #19888 )
...
* ray_pkg passed
* fix
* fix typo
* fix test
* fix test
* fix test
* fix
* draft
* compile OK
* lint
* fix
* lint
* fix ci
* Update src/ray/gcs/gcs_server/gcs_actor_manager.cc
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
* remove comment
* rename
* resolve conflict
* use unique ownership
* use DestroyActor instead of ReconstructActor
* fix sigment fault
* fix crash in debug log
* Revert "fix crash in debug log"
This reverts commit 8f0e3d37f062b664d8d0e07c6c1a9a715b8ba1ee.
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
2021-11-15 07:43:35 -08:00
Will Drevo
fa878e2d4d
Added example to user guide for cloud checkpointing ( #20045 )
...
Co-authored-by: will <will@anyscale.com>
Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>
Co-authored-by: Kai Fricke <kai@anyscale.com>
2021-11-15 15:43:06 +00:00
Sven Mika
6ff4061f3a
[RLlib] Issue 20269: Offline RL example not working due to new_obs not being written to file. ( #20366 )
...
* wip.
* Apply suggestions from code review
2021-11-15 16:41:08 +01:00
Amog Kamsetty
a74cf7ff1c
[Train] Torch Prepare utilities ( #20254 )
...
* update
* formatting
* fix failures
* fix session tests
* address comments
* add to api docs
* package refactor
* wip
* wip
* wip
* finish
* finish
* fix
* comment
* fix
* install horovod for docs
* address comment
* Update python/ray/train/session.py
Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>
* Update python/ray/train/torch.py
Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>
* address comments
* try fix docs
* fix doc build failure
* fix
* fix
* fix
* try fix doc highlighting
* fix docs
Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>
2021-11-15 07:34:17 -08:00
matthewdeng
ed3cbe48f5
[train][xgboost][release] fix ml_user_tests using ray client ( #20345 )
2021-11-15 15:24:23 +00:00
Kai Fricke
4300039d01
[ci/release] Display commit hash in buildkite overview ( #20323 )
2021-11-15 10:09:04 +00:00
Sven Mika
5b1c8e46e1
[RLlib] POC: Deprecate build_policy
(policy template) for torch only; PPOTorchPolicy ( #20061 )
2021-11-15 10:41:54 +01:00
Qing Wang
1172195571
[Java] Remove global named actor and global pg ( #20135 )
...
This PR removes global named actor and global PGs.
I believe these APIs are not used widely in OSS.
CPP part is not included in this PR.
@kfstorm @clay4444 @raulchen Please take a look if this change is reasonable.
IMPORTANT NOTE: This is a Java API change and will lead backward incompatibility in Java global named actor and global PG usage.
CPP part is not included in this PR.
INCLUDES:
Remove setGlobalName() and getGlobalActor() APIs.
Remove getGlobalPlacementGroup() and setGlobalPG
Add getActor(name, namespace) API
Add getPlacementGroup(name, namespace) API
Update doc pages.
2021-11-15 16:28:53 +08:00
SangBin Cho
a4f72c6606
[nightly] Fix pg stress test ( #20362 )
...
## Why are these changes needed?
This was mistakenly added to the nightly. Fixing it.
## Related issue number
2021-11-15 00:17:18 -08:00
SangBin Cho
477b6265d9
[Core] Fix get_actor consistency on ray.kill ( #20178 )
...
* Improve race condition on ray.kill
* Fix a bug.
* Fix a bug
* fix core worker test
* done
2021-11-14 23:29:47 -08:00
Jiajun Yao
61778a952d
Only grant or reject spillback lease request ( #20050 )
2021-11-14 21:34:28 -08:00
Philipp Moritz
440da92263
Fix manylinux2014 build scripts ( #20347 )
2021-11-14 19:42:23 -08:00
SangBin Cho
475e4dbf76
[Core] Fix pg stats not imported ( #20288 )
...
* Fix pg stats not imported
* Fix metrics are not exported to the test
2021-11-14 19:05:57 -08:00
Chen Shen
6777a31751
[Core][actor out-of-order execution 2/n] create abstraction for the queuing logic on the client/actor submission. ( #20149 )
...
This second PR in the stack that supports out or order execution for threaded/async actors. Previous PR #20148 Next PR #20150
At a high level, threaded actor/async actor already don't guarantee execution order, and the current "sequential" order implementation has caused some confusion and inconvenience. Please refer to #19822 for detailed discussion.
This PR we further separate out the logic for ordering actor requests on the client side. In the next PR, we will implement a different type of queue that supports out of order execution.
2021-11-14 18:50:54 -08:00
Stephanie Wang
0f57a9a105
Revert "Revert "[core] Fail objects when pull/reconstruction hangs ( #19789 )" ( #19904 )" ( #20120 )
...
* Revert "Revert "[core] Fail objects when pull/reconstruction hangs (#19789 )" (#19904 )"
This reverts commit 630a8cacb3
.
* debug
* fix/
* lint
* x
* fix
* x
* test
* x
2021-11-14 14:24:02 -08:00
Chen Shen
e49251d82e
[Core][actor out-of-order execution 4/n] refactor the actor receiver code
...
This is part of stack that enable out-of-order execution for actors. Previous PR #20150 Next PR #20176
Refactor the actor receiver code, by separating classes into their own header/cc files. specifically:
scheduling_queue.h for ScheduleQueue interface;
actor_scheduling_util.h for InBountRequest/DependencyWaiter/DependencyWaiterImpl
actor_scheduling_queue.h for ActorScheudlingQueue (the sequential execution queue)
normal_scheduling_queue.h for NormalSchedulingQueue (the task execution queue)
fiber_state_manager.h for FiberStateManager
thread_pool_manager.h for PoolManager and BoundedExecutor
2021-11-14 14:01:33 -08:00
Edward Oakes
2d5d499f67
[job submission] Support specifying runtime_env to job submission CLI ( #20339 )
2021-11-14 13:52:47 -08:00
Yi Cheng
87fa56def4
[gcs] Make gcs client in python able to auto reconnect ( #20299 )
...
## Why are these changes needed?
Since we are using gcs client as kv backend, we need to make it auto-reconnect in case of a failure. This PR adds this feature.
This PR adds auto_reconnect decorator to gcs-utils and in case of a failure it'll try to reconnect to gcs until it succeeds.
This feature right now support redis which should be deleted later once we finished bootstrap since kv will always go to gcs.
## Related issue number
2021-11-14 11:27:49 -08:00
Jin Hyung Ahn
722f935f9a
[docs] Fix broken links in README ( #20326 )
2021-11-14 10:23:55 -08:00
iasoon
4cbe8f4c9c
Remove confusing max_calls examples from documentation ( #19395 )
2021-11-14 10:16:41 -08:00
shrekris-anyscale
c0aeb4a236
[runtime_env] Support working_dir and py_modules from HTTPS and Google Cloud Storage ( #20280 )
2021-11-14 02:16:45 -08:00
Edward Oakes
6c3bad52b6
[job submission] Better validation + tests for input types, refactor API ( #20332 )
2021-11-13 22:54:01 -08:00
Edward Oakes
07add6f7f2
Revert "Revert "[job submission] Use ray.init format addresses for Jo… ( #20328 )
2021-11-13 16:24:02 -08:00
SangBin Cho
6cc493079b
[Core] Add Placement group performance test ( #20218 )
...
* in progress
* ip
* Fix issues
* done
* Address code review.
2021-11-14 09:17:54 +09:00
matthewdeng
e22632dabc
[train] wrap BackendExecutor in ray.remote() ( #20123 )
...
* [train] wrap BackendExecutor in ray.remote()
* wip
* fix trainer tests
* move CheckpointManager to Trainer
* [tune] move force_on_current_node to ml_utils
* fix import
* force on head node
* init ray
* split test files
* update example
* move tests to ray client
* address comments
* move comment
* address comments
2021-11-13 15:30:44 -08:00
Stephanie Wang
9e2bd508d7
[core] Move test_reconstruction to large test suite ( #20306 )
2021-11-13 13:51:18 -08:00
Sven Mika
e5ead6a4b0
[RLlib; Documentation] Minor fixes "rllib in 60s" and per-feature sigils. ( #20248 )
2021-11-13 22:10:47 +01:00
mwtian
df8042c576
[Client] connect to localhost via 127.0.0.1 ( #20274 )
...
Some Ray client users are likely seeing an issue similar to #7084 . Inside a container, connecting to localhost: fails but connecting to 127.0.0.1: succeeds. Changing Ray client to use 127.0.0.1 for localhost connection / serving should fix the issue.
2021-11-13 11:12:55 -08:00
architkulkarni
96de740cd2
[runtime env] Enable multinode tests ( #20264 )
2021-11-13 11:08:29 -08:00
Amog Kamsetty
65a17da2ec
[Train] Refactor Backends ( #20312 )
...
* wip
* finish
* comment
* fix
* install horovod for docs
* address comment
* fix doc build failure
2021-11-13 11:05:53 -08:00
Amog Kamsetty
4396419a64
[Release] Fix tune_rllib connect test ( #20321 )
...
* [Release] Fix tune_rllib connect test
* use canonical app config
2021-11-13 10:11:20 -08:00
xwjiang2010
f13c2a5350
[Tune] Revert "remove pg caching" ( #20308 )
...
This reverts commit 5f14eb3ee4
.
2021-11-13 16:25:22 +00:00
Antoni Baum
1b867520e6
[docs]Add pyarrow as a dependency ( #20320 )
2021-11-13 16:00:58 +00:00
mwtian
875b0aea0a
fallback to grpc.experimental.aio when importing grpc.aio ( #20287 )
2021-11-13 15:59:57 +09:00
mwtian
cdadc2b7d2
Change owner ( #20313 )
2021-11-12 21:23:36 -08:00
Eric Liang
567e955810
Revert "[job submission] Use ray.init format addresses for JobSubmissionClient ( #20245 )" ( #20314 )
...
This reverts commit adc15a0fb0
.
2021-11-12 21:11:24 -08:00
gjoliver
7fe42341ed
[release] Switch many_ppo test to use the canonical rllib app cfg as well. ( #20310 )
2021-11-12 20:51:28 -08:00
Jiajun Yao
f8b738d029
[scheduler] Add scheduling debug log ( #20302 )
2021-11-12 18:48:05 -08:00
matthewdeng
e77cc926be
[train] minor doc updates ( #20271 )
2021-11-12 17:20:23 -08:00