hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-06 10:31:39 -05:00

Author	SHA1	Message	Date
Qing Wang	6504ad6bb2	[xlang] Add named actor xlang tests. (#20368 ) We add named actor xlang tests, including both getting java named actor in python and get python named actor in Java. Related issue number #19794	2021-11-16 21:42:05 +08:00
SangBin Cho	5ec63ccc5f	[Regresion test] Placement group long running test (#20251 ) Why are these changes needed? In the past, there was a regression the placement group creation time gets slower as time goes. I believe the issue is fixed in the master, but this PR verifies if that's actually fixed. This PR adds a long running test for the placement group. There are 2 purposes of the test. Make sure the placement group creation / removal doesn't get slower as time goes. The test basically measure the first 20 iteration P50 creation time and run very long iteration. After all iteration, it checks if the p50 creation time is not too slow compared to the initial round. Make sure placement group removal / creation works consistently for a long time without an issue. Q: Should we make it a real long running test? (that runs for a day?)	2021-11-16 04:21:18 -08:00
SangBin Cho	137aec04c0	[Core] Better logs job message failure (#20363 ) <!-- Please add a reviewer to the assignee section when you create a PR. If you don't have the access to it, we will shortly find a reviewer and assign them to your PR. --> ## Why are these changes needed? There's one user who has an issue that one of raylets cannot schedule tasks anymore because `num_worker_not_started_by_job_config_not_exist ` > 0. This PR adds better log messages to figure out if the root cause is the job information is not properly propagated from GCS to raylet through Redis pubsub. ## Related issue number <!-- For example: "Closes #1234" --> ## Checks - [ ] I've run `scripts/format.sh` to lint the changes in this PR. - [ ] I've included any doc changes needed for https://docs.ray.io/en/master/. - [ ] I've made sure the tests are passing. Note that there might be a few flaky tests, see the recent failures at https://flakey-tests.ray.io/ - Testing Strategy - [ ] Unit tests - [ ] Release tests - [ ] This PR is not tested :(	2021-11-16 04:20:49 -08:00
Stefan Schneider	2b3d0c691f	[RLlib] Document and extend action mask example. (#20390 ) Co-authored-by: Richard Liaw <rliaw@berkeley.edu> Co-authored-by: Sven Mika <sven@anyscale.io> Co-authored-by: sven1977 <svenmika1977@gmail.com>	2021-11-16 13:20:41 +01:00
Kai Fricke	3e6ba5d6d2	Revert "Revert [RLlib] POC: `PGTrainer` class that works by sub-classing, not `trainer_template.py`." (#20285 ) * Revert "Revert "[RLlib] POC: `PGTrainer` class that works by sub-classing, not `trainer_template.py`. (#20055)" (#20284)" This reverts commit `246787cdd9`. Co-authored-by: sven1977 <svenmika1977@gmail.com>	2021-11-16 12:26:47 +01:00
Eric Liang	a1d78088e6	[Hotfix] Fix flaky test_basic_workflows_2 This test seems very flaky after the nested tasks PR merged. I think it's since the test is broken, and we made one of the branches more likely.	2021-11-16 01:11:36 -08:00
Yi Cheng	a4e187c0e7	[gcs] Update function table to use internal kv (#20152 ) ## Why are these changes needed? This is a part of redis removal. This PR remove redis kv in function table. rpush related code is not updated in this PR. ## Related issue number	2021-11-15 23:34:41 -08:00
Eric Liang	460cf86858	Split blocks automatically into 500MB chunks on file read and transformation (#20235 ) This PR adds support for automatic block splitting on read and map transforms, to keep block size bounded to ~500MiB. This avoids potential OOM situations where a map task may consume too much intermediate Python heap memory, or too much object store shared memory for one block.	2021-11-15 22:25:11 -08:00
Siyuan (Ryans) Zhuang	3e9cd4248e	[workflow] Refactoring workflow to make it easier to follow the logic (#20349 ) * update * cleanup	2021-11-15 21:02:33 -08:00
Yiran Wang	f4e8319eaa	Remove .boto files that are no longer needed during docker build (#20407 ) ## Why are these changes needed? The .boto files are already added to the base image and ACL'ed to root, adding them again during app config build causes permission issues. ## Related issue number	2021-11-15 20:49:33 -08:00
Stephanie Wang	31eb385426	Revert "Revert "Revert "[core] Fail objects when pull/reconstruction hangs (#19789 )" (#19904 )" (#20120 )" (#20406 ) This reverts commit `0f57a9a105`.	2021-11-15 20:36:22 -08:00
Alex Wu	75f421a3fd	[core] Nested task support via task depth + backpressure (#17887 ) * needs depth * depth * . * . * . * lint * . * lint * fix tests * . * . * . * . * cleanup * . * tests * . * more tests * fix rest(?) of tests * cleanup * . * . * . * . * lint * fix test basic * fix ref counting? * cleanup * lint * . * pass dataset pipeline test * . * stephanie's comments + fix tests * cleanup * cleanup * minor cleanup, then fix merge conflict * lint * cast * feature flag * lint * lint * refactor * needs cleanup * should pass * lint * . * . * . * work? * . * works? * lint * work? * . * fix cpp tests * . * . * split test * fix windows? * fix windows? * fix test + check * . * all passing * tests * lint * cleanup * . * most stephanie ocmments * lint * remove timer * . * allowed - capacity * . * everything except barrier * addd guard * works * lint * works? * debug string * last comment? * short comments * most comments * lint * done? * done? * . * . * . * . * done? * done? * update * lint * fix last test * . * . * . * . * . * . * debug * . * . * . * . * fix type * . * . * cleanup Co-authored-by: Alex Wu <alex@anyscale.com>	2021-11-15 17:39:50 -08:00
Edward Oakes	48bc1af2da	[job submission] Remove DOES_NOT_EXIST status (#20354 )	2021-11-15 16:57:32 -08:00
mwtian	1dd8b3d2bc	[Build] Remove debug info from Ray libraries. (#20389 ) ## Why are these changes needed? Ray wheel size limit is still at 100MB. Removing debug symbols would decrease Ray Linux wheel sizes. ## Related issue number ## Checks	2021-11-15 16:40:48 -08:00
Amog Kamsetty	90dc5460d4	Revert "[RLlib] POC: Deprecate `build_policy` (policy template) for torch only; PPOTorchPolicy (#20061 )" (#20399 ) This reverts commit `5b1c8e46e1`.	2021-11-15 16:11:35 -08:00
iasoon	171ad62e30	Link to the documentation on contributing from CONTRIBUTING.rst (#19396 ) Co-authored-by: Richard Liaw <rliaw@berkeley.edu>	2021-11-15 15:34:18 -08:00
Antoni Baum	ec81f52061	[Docs] Fix typo in C++ Placement Group example (#20386 )	2021-11-16 08:19:09 +09:00
matthewdeng	35dc3cf21b	[train] fix Train/Tune integration on Client (#20351 ) * [train] fix Train/Tune integration on Client * remove force_on_current_node	2021-11-15 14:36:33 -08:00
Alex Wu	884bb3de33	[Dataset] Bump `numpy >=1.20` dependency (#20374 ) * done? * . Co-authored-by: Alex Wu <alex@anyscale.com>	2021-11-15 14:10:00 -08:00
Kai Fricke	d191ad2de8	[ci/release] Return exit codes based on different errors (#20289 )	2021-11-15 19:41:00 +00:00
Simon Mo	72ae22e82b	[CI] Fix frontend build issue (#20375 )	2021-11-15 10:12:43 -08:00
Kai Fricke	91920f1d02	[release/xgboost] xgboost release test fixes via app config (#20325 ) * [xgboost] Fix release test app configs * Revert full app config * Update base docker image * Only change cpu base image * default * Pin xgboost to 1.5. in cpu tests * Remove numpy hack * Revert one line Co-authored-by: Amog Kamsetty <amogkamsetty@yahoo.com>	2021-11-15 10:03:21 -08:00
Amog Kamsetty	ef7967476c	[Train] Torch data transfer automatic conversion (#20333 ) * update * formatting * fix failures * fix session tests * address comments * add to api docs * package refactor * wip * wip * wip * finish * finish * fix * comment * fix * install horovod for docs * address comment * Update python/ray/train/session.py Co-authored-by: matthewdeng <matthew.j.deng@gmail.com> * Update python/ray/train/torch.py Co-authored-by: matthewdeng <matthew.j.deng@gmail.com> * address comments * try fix docs * fix doc build failure * wip * fix * fix * fix * try fix doc highlighting * fix docs * finish * formatting * address comments and fix tests * address comments and fix test Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>	2021-11-15 09:14:12 -08:00
Lixin Wei	85dbda8cf1	[Core] Fix Crash in Debug Log (#20322 ) * fix crash in debug log * fix * fix * fix	2021-11-16 00:45:00 +09:00
Lixin Wei	b7e35acf14	[RuntimeEnv] Raise RuntimeEnvSetupError when Actor Creation Failed due to It (#19888 ) * ray_pkg passed * fix * fix typo * fix test * fix test * fix test * fix * draft * compile OK * lint * fix * lint * fix ci * Update src/ray/gcs/gcs_server/gcs_actor_manager.cc Co-authored-by: SangBin Cho <rkooo567@gmail.com> * remove comment * rename * resolve conflict * use unique ownership * use DestroyActor instead of ReconstructActor * fix sigment fault * fix crash in debug log * Revert "fix crash in debug log" This reverts commit 8f0e3d37f062b664d8d0e07c6c1a9a715b8ba1ee. Co-authored-by: SangBin Cho <rkooo567@gmail.com>	2021-11-15 07:43:35 -08:00
Will Drevo	fa878e2d4d	Added example to user guide for cloud checkpointing (#20045 ) Co-authored-by: will <will@anyscale.com> Co-authored-by: Antoni Baum <antoni.baum@protonmail.com> Co-authored-by: Kai Fricke <kai@anyscale.com>	2021-11-15 15:43:06 +00:00
Sven Mika	6ff4061f3a	[RLlib] Issue 20269: Offline RL example not working due to new_obs not being written to file. (#20366 ) * wip. * Apply suggestions from code review	2021-11-15 16:41:08 +01:00
Amog Kamsetty	a74cf7ff1c	[Train] Torch Prepare utilities (#20254 ) * update * formatting * fix failures * fix session tests * address comments * add to api docs * package refactor * wip * wip * wip * finish * finish * fix * comment * fix * install horovod for docs * address comment * Update python/ray/train/session.py Co-authored-by: matthewdeng <matthew.j.deng@gmail.com> * Update python/ray/train/torch.py Co-authored-by: matthewdeng <matthew.j.deng@gmail.com> * address comments * try fix docs * fix doc build failure * fix * fix * fix * try fix doc highlighting * fix docs Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>	2021-11-15 07:34:17 -08:00
matthewdeng	ed3cbe48f5	[train][xgboost][release] fix ml_user_tests using ray client (#20345 )	2021-11-15 15:24:23 +00:00
Kai Fricke	4300039d01	[ci/release] Display commit hash in buildkite overview (#20323 )	2021-11-15 10:09:04 +00:00
Sven Mika	5b1c8e46e1	[RLlib] POC: Deprecate `build_policy` (policy template) for torch only; PPOTorchPolicy (#20061 )	2021-11-15 10:41:54 +01:00
Qing Wang	1172195571	[Java] Remove global named actor and global pg (#20135 ) This PR removes global named actor and global PGs. I believe these APIs are not used widely in OSS. CPP part is not included in this PR. @kfstorm @clay4444 @raulchen Please take a look if this change is reasonable. IMPORTANT NOTE: This is a Java API change and will lead backward incompatibility in Java global named actor and global PG usage. CPP part is not included in this PR. INCLUDES: Remove setGlobalName() and getGlobalActor() APIs. Remove getGlobalPlacementGroup() and setGlobalPG Add getActor(name, namespace) API Add getPlacementGroup(name, namespace) API Update doc pages.	2021-11-15 16:28:53 +08:00
SangBin Cho	a4f72c6606	[nightly] Fix pg stress test (#20362 ) ## Why are these changes needed? This was mistakenly added to the nightly. Fixing it. ## Related issue number	2021-11-15 00:17:18 -08:00
SangBin Cho	477b6265d9	[Core] Fix get_actor consistency on ray.kill (#20178 ) * Improve race condition on ray.kill * Fix a bug. * Fix a bug * fix core worker test * done	2021-11-14 23:29:47 -08:00
Jiajun Yao	61778a952d	Only grant or reject spillback lease request (#20050 )	2021-11-14 21:34:28 -08:00
Philipp Moritz	440da92263	Fix manylinux2014 build scripts (#20347 )	2021-11-14 19:42:23 -08:00
SangBin Cho	475e4dbf76	[Core] Fix pg stats not imported (#20288 ) * Fix pg stats not imported * Fix metrics are not exported to the test	2021-11-14 19:05:57 -08:00
Chen Shen	6777a31751	[Core][actor out-of-order execution 2/n] create abstraction for the queuing logic on the client/actor submission. (#20149 ) This second PR in the stack that supports out or order execution for threaded/async actors. Previous PR #20148 Next PR #20150 At a high level, threaded actor/async actor already don't guarantee execution order, and the current "sequential" order implementation has caused some confusion and inconvenience. Please refer to #19822 for detailed discussion. This PR we further separate out the logic for ordering actor requests on the client side. In the next PR, we will implement a different type of queue that supports out of order execution.	2021-11-14 18:50:54 -08:00
Stephanie Wang	0f57a9a105	Revert "Revert "[core] Fail objects when pull/reconstruction hangs (#19789 )" (#19904 )" (#20120 ) * Revert "Revert "[core] Fail objects when pull/reconstruction hangs (#19789)" (#19904)" This reverts commit `630a8cacb3`. * debug * fix/ * lint * x * fix * x * test * x	2021-11-14 14:24:02 -08:00
Chen Shen	e49251d82e	[Core][actor out-of-order execution 4/n] refactor the actor receiver code This is part of stack that enable out-of-order execution for actors. Previous PR #20150 Next PR #20176 Refactor the actor receiver code, by separating classes into their own header/cc files. specifically: scheduling_queue.h for ScheduleQueue interface; actor_scheduling_util.h for InBountRequest/DependencyWaiter/DependencyWaiterImpl actor_scheduling_queue.h for ActorScheudlingQueue (the sequential execution queue) normal_scheduling_queue.h for NormalSchedulingQueue (the task execution queue) fiber_state_manager.h for FiberStateManager thread_pool_manager.h for PoolManager and BoundedExecutor	2021-11-14 14:01:33 -08:00
Edward Oakes	2d5d499f67	[job submission] Support specifying runtime_env to job submission CLI (#20339 )	2021-11-14 13:52:47 -08:00
Yi Cheng	87fa56def4	[gcs] Make gcs client in python able to auto reconnect (#20299 ) ## Why are these changes needed? Since we are using gcs client as kv backend, we need to make it auto-reconnect in case of a failure. This PR adds this feature. This PR adds auto_reconnect decorator to gcs-utils and in case of a failure it'll try to reconnect to gcs until it succeeds. This feature right now support redis which should be deleted later once we finished bootstrap since kv will always go to gcs. ## Related issue number	2021-11-14 11:27:49 -08:00
Jin Hyung Ahn	722f935f9a	[docs] Fix broken links in README (#20326 )	2021-11-14 10:23:55 -08:00
iasoon	4cbe8f4c9c	Remove confusing max_calls examples from documentation (#19395 )	2021-11-14 10:16:41 -08:00
shrekris-anyscale	c0aeb4a236	[runtime_env] Support working_dir and py_modules from HTTPS and Google Cloud Storage (#20280 )	2021-11-14 02:16:45 -08:00
Edward Oakes	6c3bad52b6	[job submission] Better validation + tests for input types, refactor API (#20332 )	2021-11-13 22:54:01 -08:00
Edward Oakes	07add6f7f2	Revert "Revert "[job submission] Use ray.init format addresses for Jo… (#20328 )	2021-11-13 16:24:02 -08:00
SangBin Cho	6cc493079b	[Core] Add Placement group performance test (#20218 ) * in progress * ip * Fix issues * done * Address code review.	2021-11-14 09:17:54 +09:00
matthewdeng	e22632dabc	[train] wrap BackendExecutor in ray.remote() (#20123 ) * [train] wrap BackendExecutor in ray.remote() * wip * fix trainer tests * move CheckpointManager to Trainer * [tune] move force_on_current_node to ml_utils * fix import * force on head node * init ray * split test files * update example * move tests to ray client * address comments * move comment * address comments	2021-11-13 15:30:44 -08:00
Stephanie Wang	9e2bd508d7	[core] Move test_reconstruction to large test suite (#20306 )	2021-11-13 13:51:18 -08:00

1 2 3 4 5 ...

10406 commits