hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-06 10:31:39 -05:00

Author	SHA1	Message	Date
SangBin Cho	a4f72c6606	[nightly] Fix pg stress test (#20362 ) ## Why are these changes needed? This was mistakenly added to the nightly. Fixing it. ## Related issue number	2021-11-15 00:17:18 -08:00
SangBin Cho	477b6265d9	[Core] Fix get_actor consistency on ray.kill (#20178 ) * Improve race condition on ray.kill * Fix a bug. * Fix a bug * fix core worker test * done	2021-11-14 23:29:47 -08:00
Jiajun Yao	61778a952d	Only grant or reject spillback lease request (#20050 )	2021-11-14 21:34:28 -08:00
Philipp Moritz	440da92263	Fix manylinux2014 build scripts (#20347 )	2021-11-14 19:42:23 -08:00
SangBin Cho	475e4dbf76	[Core] Fix pg stats not imported (#20288 ) * Fix pg stats not imported * Fix metrics are not exported to the test	2021-11-14 19:05:57 -08:00
Chen Shen	6777a31751	[Core][actor out-of-order execution 2/n] create abstraction for the queuing logic on the client/actor submission. (#20149 ) This second PR in the stack that supports out or order execution for threaded/async actors. Previous PR #20148 Next PR #20150 At a high level, threaded actor/async actor already don't guarantee execution order, and the current "sequential" order implementation has caused some confusion and inconvenience. Please refer to #19822 for detailed discussion. This PR we further separate out the logic for ordering actor requests on the client side. In the next PR, we will implement a different type of queue that supports out of order execution.	2021-11-14 18:50:54 -08:00
Stephanie Wang	0f57a9a105	Revert "Revert "[core] Fail objects when pull/reconstruction hangs (#19789 )" (#19904 )" (#20120 ) * Revert "Revert "[core] Fail objects when pull/reconstruction hangs (#19789)" (#19904)" This reverts commit `630a8cacb3`. * debug * fix/ * lint * x * fix * x * test * x	2021-11-14 14:24:02 -08:00
Chen Shen	e49251d82e	[Core][actor out-of-order execution 4/n] refactor the actor receiver code This is part of stack that enable out-of-order execution for actors. Previous PR #20150 Next PR #20176 Refactor the actor receiver code, by separating classes into their own header/cc files. specifically: scheduling_queue.h for ScheduleQueue interface; actor_scheduling_util.h for InBountRequest/DependencyWaiter/DependencyWaiterImpl actor_scheduling_queue.h for ActorScheudlingQueue (the sequential execution queue) normal_scheduling_queue.h for NormalSchedulingQueue (the task execution queue) fiber_state_manager.h for FiberStateManager thread_pool_manager.h for PoolManager and BoundedExecutor	2021-11-14 14:01:33 -08:00
Edward Oakes	2d5d499f67	[job submission] Support specifying runtime_env to job submission CLI (#20339 )	2021-11-14 13:52:47 -08:00
Yi Cheng	87fa56def4	[gcs] Make gcs client in python able to auto reconnect (#20299 ) ## Why are these changes needed? Since we are using gcs client as kv backend, we need to make it auto-reconnect in case of a failure. This PR adds this feature. This PR adds auto_reconnect decorator to gcs-utils and in case of a failure it'll try to reconnect to gcs until it succeeds. This feature right now support redis which should be deleted later once we finished bootstrap since kv will always go to gcs. ## Related issue number	2021-11-14 11:27:49 -08:00
Jin Hyung Ahn	722f935f9a	[docs] Fix broken links in README (#20326 )	2021-11-14 10:23:55 -08:00
iasoon	4cbe8f4c9c	Remove confusing max_calls examples from documentation (#19395 )	2021-11-14 10:16:41 -08:00
shrekris-anyscale	c0aeb4a236	[runtime_env] Support working_dir and py_modules from HTTPS and Google Cloud Storage (#20280 )	2021-11-14 02:16:45 -08:00
Edward Oakes	6c3bad52b6	[job submission] Better validation + tests for input types, refactor API (#20332 )	2021-11-13 22:54:01 -08:00
Edward Oakes	07add6f7f2	Revert "Revert "[job submission] Use ray.init format addresses for Jo… (#20328 )	2021-11-13 16:24:02 -08:00
SangBin Cho	6cc493079b	[Core] Add Placement group performance test (#20218 ) * in progress * ip * Fix issues * done * Address code review.	2021-11-14 09:17:54 +09:00
matthewdeng	e22632dabc	[train] wrap BackendExecutor in ray.remote() (#20123 ) * [train] wrap BackendExecutor in ray.remote() * wip * fix trainer tests * move CheckpointManager to Trainer * [tune] move force_on_current_node to ml_utils * fix import * force on head node * init ray * split test files * update example * move tests to ray client * address comments * move comment * address comments	2021-11-13 15:30:44 -08:00
Stephanie Wang	9e2bd508d7	[core] Move test_reconstruction to large test suite (#20306 )	2021-11-13 13:51:18 -08:00
Sven Mika	e5ead6a4b0	[RLlib; Documentation] Minor fixes "rllib in 60s" and per-feature sigils. (#20248 )	2021-11-13 22:10:47 +01:00
mwtian	df8042c576	[Client] connect to localhost via 127.0.0.1 (#20274 ) Some Ray client users are likely seeing an issue similar to #7084. Inside a container, connecting to localhost: fails but connecting to 127.0.0.1: succeeds. Changing Ray client to use 127.0.0.1 for localhost connection / serving should fix the issue.	2021-11-13 11:12:55 -08:00
architkulkarni	96de740cd2	[runtime env] Enable multinode tests (#20264 )	2021-11-13 11:08:29 -08:00
Amog Kamsetty	65a17da2ec	[Train] Refactor Backends (#20312 ) * wip * finish * comment * fix * install horovod for docs * address comment * fix doc build failure	2021-11-13 11:05:53 -08:00
Amog Kamsetty	4396419a64	[Release] Fix tune_rllib connect test (#20321 ) * [Release] Fix tune_rllib connect test * use canonical app config	2021-11-13 10:11:20 -08:00
xwjiang2010	f13c2a5350	[Tune] Revert "remove pg caching" (#20308 ) This reverts commit `5f14eb3ee4`.	2021-11-13 16:25:22 +00:00
Antoni Baum	1b867520e6	[docs]Add pyarrow as a dependency (#20320 )	2021-11-13 16:00:58 +00:00
mwtian	875b0aea0a	fallback to grpc.experimental.aio when importing grpc.aio (#20287 )	2021-11-13 15:59:57 +09:00
mwtian	cdadc2b7d2	Change owner (#20313 )	2021-11-12 21:23:36 -08:00
Eric Liang	567e955810	Revert "[job submission] Use ray.init format addresses for JobSubmissionClient (#20245 )" (#20314 ) This reverts commit `adc15a0fb0`.	2021-11-12 21:11:24 -08:00
gjoliver	7fe42341ed	[release] Switch many_ppo test to use the canonical rllib app cfg as well. (#20310 )	2021-11-12 20:51:28 -08:00
Jiajun Yao	f8b738d029	[scheduler] Add scheduling debug log (#20302 )	2021-11-12 18:48:05 -08:00
matthewdeng	e77cc926be	[train] minor doc updates (#20271 )	2021-11-12 17:20:23 -08:00
Clark Zinzow	918a215442	[Datasets] Multi-aggregations [2/3]: Add groupby multi-column/multi-lambda aggregation (#20074 )	2021-11-12 15:53:58 -08:00
mwtian	a39fd74674	disable //python/ray/tests:test_autoscaler_drain_node_api in HA GCS build (#20296 )	2021-11-12 15:47:42 -08:00
Tricia Fu	e59c14117f	[Doc] [Serve] Add summary sub header to each page (#20231 )	2021-11-12 14:18:42 -08:00
Nikita Vemuri	adc15a0fb0	[job submission] Use ray.init format addresses for JobSubmissionClient (#20245 )	2021-11-12 13:52:43 -08:00
xwjiang2010	cdf70c2900	[Tune] Remove legacy resources implementations in Runner and Executor. (#19773 )	2021-11-12 12:33:39 -08:00
Edward Oakes	73e570c426	Fix windows build (don't skip test_job_manager.py) (#20294 )	2021-11-12 11:13:15 -08:00
architkulkarni	138ec75246	[runtime env] Revert reference counting for per-actor URIs (#20281 )	2021-11-12 11:09:38 -08:00
Matti Picus	1e80a2a83a	[WINDOWS] unskip tests (#20212 )	2021-11-12 10:11:11 -08:00
Chen Shen	a617cb8813	[Core][actor out-of-order execution 1/n] Move CoreWorkerDirectActorTaskSubmitter into a separate file This is the stack of PRs that supports out or order execution for threaded/async actors. Next PR #20149 At a high level, threaded actor/async actor already don't guarantee execution order, and the current "sequential" order implementation has caused some confusion and inconvenience. Please refer to #19822 for detailed discussion. The major changes of this stack is introduce OutOfOrderActorSubmitQueue ([Core][actor out-of-order execution 3/n] Introducing out-of-order actor submit queue #20150) Specifically, we have a per-client task submit queue, which guarantees the sequential order of task submission. In [Core][actor out-of-order execution 3/n] Introducing out-of-order actor submit queue #20150 we introduce OutOfOrderActorSubmitQueue, which relaxes the guarantee; it send the task over the network as soon as its dependency is resolved. -- there are 2 PRs ([Core][actor out-of-order execution 1/n] Move CoreWorkerDirectActorTaskSubmitter into a separate file #20148 and [Core][actor out-of-order execution 2/n] create abstraction for the queuing logic on the client/actor submission. #20149) precedes this PR, which refactor the actor submission logic to make the abstraction possible. OutOfOrderActorSchedullingQueue ([Core][actor out-of-order execution 5/n] implement out-of-order scheduling queue #20176) Similarly, we also have a per-client task scheduling queue on the actor to ensure tasks are executed according to the submission order (sequence_no). OutOfOrderActorSchedullingQueue relaxes the guarantee by enqueuing the task as soon as all their dependencies are resolved. -- there is one PR ([Core][actor out-of-order execution 4/n] refactor the actor receiver code #20160) precedes this PR, which refactor the actor scheduling logic to make it the code easier to read; however this one is optional. plumbing PR ([Core][actor out-of-order execution 6/n] plumbing work to make it work e2e #20177) This PR enables the out of order execution by introducing options(execute_out_of_order=True); and create actor client/server components according to the configuration. There are something the PR hasn't touched upon is the restart/retry guarantees, which might need some discussion. This PR we separate client from server code for Actor task submission. This makes the follow up change easier.	2021-11-12 09:28:58 -08:00
Siyuan (Ryans) Zhuang	3b62388a9a	[Workflow] Workflow tail recursion optimization (#19928 ) * tail recursion optimization	2021-11-12 09:13:40 -08:00
Simon Mo	b6bd4fd5f3	[Serve] Don't recover from current state checkpoint (#19998 )	2021-11-12 09:02:27 -08:00
xwjiang2010	ce8504b0b2	[CI] Rebalance Tune tests a bit. (#20263 )	2021-11-12 15:30:18 +00:00
xwjiang2010	5f14eb3ee4	[Tune] Remove PG caching. (#19515 ) Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>	2021-11-12 14:36:04 +00:00
Sven Mika	38c456b6f4	[RLlib; Tune] Fix rllib/train.py script after tune.Experiment c'tor change. (#20283 )	2021-11-12 15:25:50 +01:00
Kai Fricke	246787cdd9	Revert "[RLlib] POC: `PGTrainer` class that works by sub-classing, not `trainer_template.py`. (#20055 )" (#20284 ) This reverts commit `6f85af435f`.	2021-11-12 13:09:43 +00:00
Kai Fricke	d88fdd6e38	[tune] refactor SyncConfig (#20155 )	2021-11-12 09:36:15 +00:00
SangBin Cho	7132f91789	[Core] Reduce the frequency of retry messages (#20175 ) * Reduce the frequency of retry messages * done	2021-11-11 23:52:37 -08:00
Sven Mika	70fe25055a	[RLlib] Issue: Get single step input dict incorrect. (#20217 )	2021-11-12 08:38:51 +01:00
Edward Oakes	ee4e4f4036	[runtime_env] Support specifying the runtime_resources directory for testing (#20257 )	2021-11-11 21:50:42 -08:00

1 2 3 4 5 ...

10474 commits