hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-09 21:06:39 -04:00

Author	SHA1	Message	Date
Amog Kamsetty	0b8c21922b	[Train] Improvements to fault tolerance (#22511 ) Various improvements to Ray Train fault tolerance. Add more log statements for better debugging of Ray Train failure handling. Fixes [Bug] [Train] Cannot reproduce fault-tolerance, script hangs upon any node shutdown #22349. Simplifies fault tolerance by removing backend specific handle_failure. If any workers have failed, all workers will be restarted and training will continue from the last checkpoint. Also adds a test for fault tolerance with an actual torch example. When testing locally, the test hangs before the fix, but passes after.	2022-03-29 15:36:46 -07:00
Stephanie Wang	da7901f3fc	[core] Filter out self node from the list of object locations during pull (#23539 ) Running Datasets shuffle with 1TB data and 2k partitions sometimes times out due to a failed object fetch. This happens because the object directory notifies the PullManager that the object is already on the local node, even though it isn't. This seems to be a bug in the object directory. To work around this on the PullManager side, this PR filters out the current node from the list of locations provided by the object directory. @jjyao confirmed that this fixes the issue for Datasets shuffle.	2022-03-29 15:18:14 -07:00
Kai Fricke	922367d158	[ci/release] Fix smoke test compute templates (#23561 ) The smoke test definitions of a few tests were faulty for compute template override. Core tests @rkooo567: https://buildkite.com/ray-project/release-tests-branch/builds/294	2022-03-29 13:48:09 -07:00
Gagandeep Singh	3856011267	[Serve] [Test] Bytecode check to verify imported function correctness in `test_pipeline_driver::test_loading_check` (#23552 )	2022-03-29 13:18:25 -07:00
Linsong Chu	2a6cbc5202	[workflow]align the behavior of workflow's options() with remote function's options() (#23469 ) The current behavior of workflow's `.options()` is to completely rewrite all the options rather than update options, this is less intuitive and inconsistent with the behavior of `.options()` in remote functions. For example: ``` # Remote Function @ray.remote(num_cpus=2, max_retries=2) f.options(num_cpus=1) ``` `options()` here updated num_cpus while the rest options are untouched, i.e. max_retires is still 2. This is the expected behavior and more intuitive. ``` # Workflow Step @workflow.step(num_cpus=2, max_retries=2) f.options(num_cpus=1) ``` `options()` here completely drop all existing options and only set num_cpus, i.e. previous value of max_retires (2) is dropped and reverted to default (3). This will also drop other fields like `name` and `metadata` if name and metadata are given in the decorator but not in the options().	2022-03-29 12:35:04 -07:00
Yi Cheng	61c9186b59	[2][cleanup][gcs] Cleanup GCS client options. (#23519 ) This PR cleanup GCS client options.	2022-03-29 12:01:58 -07:00
Simon Mo	cb1919b8d0	[Doc][Serve] Add minimal docs for model wrappers and http adapters (#23536 )	2022-03-29 11:33:14 -07:00
Kai Fricke	afd287eb93	[ci] linkcheck should soft fail (#23559 ) Linkcheck failures should not break the build.	2022-03-29 10:57:03 -07:00
Philipp Moritz	005ea36850	[linkcheck] Remove flaky url (#23549 )	2022-03-29 08:36:54 -07:00
Artur Niederfahrenhorst	9a64bd4e9b	[RLlib] Simple-Q uses training iteration fn (instead of execution_plan); ReplayBuffer API for Simple-Q (#22842 )	2022-03-29 14:44:40 +02:00
Jun Gong	a7e5aa8c6a	[RLlib] Delete some unused confusing logics. (#23513 )	2022-03-29 13:45:13 +02:00
Hao Chen	b7d32df8b0	Refactor scheduler data structures (#22854 ) This is the first PR to refactor scheduler data structures (See #22850). Major changes: - Hid the implementation details in the `ResourceRequest` and `TaskResourceInstnaces` classes, which expose public methods such as algebra operators and comparison operators. - Hid the differences between "predefined" and "custom" resources inside these 2 classes. Call sites can simply use the resource ID to access the resource, no matter it is predefined or custom. - The predefined_resources vector now always has the full length. So no more "resize"s are needed. - Removed the `ResourceCapacity` class. Now "total" and "available" resources are stored in separate fields in "NodeResources". - Moved helper functions for FixedPoint vectors from "cluster_resource_data.h" to "fixed_point.h" - "ResourceID" now has static methods to get the resource ids of predefined resources, e.g. "ResourceID::CPU()". - Encapsulated unit-instance resource logic to "ResourceID" Other planned changes that are not included in this PR: - Rename ResourceRequest to ResourceSet, and move it to its own file. - Remove the predefined vectors and always use maps. Co-authored-by: Chong-Li <lc300133@antgroup.com>	2022-03-29 19:44:59 +08:00
Artur Niederfahrenhorst	32ad6c6ef1	[RLlib] Replay Buffer capacity check (#23523 )	2022-03-29 12:06:27 +02:00
Eric Liang	990b0ec934	Move linkcheck into a separate CI build Why are these changes needed? Linkcheck is inherently flaky, so separate it from the normal LINT build which is never flaky. This also separates the verbose linkcheck logs, making it easier to read the LINT output.	2022-03-29 01:08:53 -07:00
Matti Picus	0cb2847e2f	WINDOWS: make default node startup timeout longer (#23462 ) Timeouts when starting nodes rank high on https://flakey-tests.ray.io/#owner=core. The default timeout should be longer on windows. For instance, [here](https://buildkite.com/ray-project/ray-builders-branch/builds/6720#d4cf497e-13d5-4b6b-9354-2dd8828bd0e7/2835-3259) is one such error.	2022-03-29 01:01:43 -07:00
Matti Picus	84026ef55d	WINDOWS: make timeout longer in test_metrics (#23461 ) `test_metrics` scales quite high on https://flakey-tests.ray.io/#owner=core. This test is often hitting the timeout limit. Making it larger should help the test pass.	2022-03-29 00:59:13 -07:00
Matti Picus	77c4c1e48e	WINDOWS: enable and fix failures in test_runtime_env_complicated (#22449 )	2022-03-29 00:56:42 -07:00
Chen Shen	44114c8422	[CI] pin click version to fix broken test. #23544	2022-03-29 00:44:48 -07:00
Chen Shen	1d0fe1e1c3	[doc/linter] fix broken deepmind link #23542	2022-03-28 22:35:53 -07:00
Yi Cheng	7de751dbab	[1][core][cleanup] remove enable gcs bootstrap in cpp. (#23518 ) This PR remove enable_gcs_bootstrap flag in cpp.	2022-03-28 21:37:24 -07:00
Jian Xiao	cc0db8b92a	Fix Dataset zip for pandas (#23532 ) Dataset zip cannot work for Pandas.	2022-03-28 17:58:31 -07:00
Kai Fricke	62414525f9	[tune] Optuna should ignore additional results after trial termination (#23495 ) In rare cases (#19274) (and possibly old versions of Ray), buffered results can lead to calling on_trial_complete multiple times with the same trial ID. In these cases, Optuna should gracefully handle this case and discard the results.	2022-03-28 20:07:41 +01:00
Kai Fricke	262d6121bb	[rllib] Fix error messages and example for dataset writer (#23419 ) Currently the error message and example refer to a field type that is actually format.	2022-03-28 19:53:12 +01:00
Chen Shen	51bdefc2c8	[scheduler][monitoring] dump detailed spilling metrics (#23321 ) Dump the detailed spilling metrics in scheduler.	2022-03-28 10:49:04 -07:00
shrekris-anyscale	aae144d7f9	[serve] Make Serve CLI and serve.build() non-public (#23504 ) This change makes `serve.build()` non-public and hides the following Serve CLI commands: * `deploy` * `config` * `delete` * `build`	2022-03-28 10:40:57 -07:00
Kai Fricke	1465eaa306	[tune] Use new Checkpoint interface internally (#22801 ) Follow up from #22741, also use the new checkpoint interface internally. This PR is low friction and just replaces some internal bookkeeping methods. With the new Checkpoint interface, there is no need to revamp the save/restore APIs completely. Instead, we will focus on the bookkeeping part, which takes place in the Ray Tune's and Ray Train's checkpoint managers. These will be consolidated in a future PR.	2022-03-28 18:33:40 +01:00
Chen Shen	c3e04ab275	[nighly-test] try out spot instances for chaos test #23507	2022-03-27 20:10:21 -07:00
mwtian	d1ef498638	[Python Worker] load actor dependency without importer thread (#23383 ) Import actor dependency when not found, so actor dependencies can be imported without the importer thread. Remaining blockers to remove importer thread are to support running a function on all workers `run_function_on_all_workers()`, and raising a warning when the same function / class is exported too many times.	2022-03-27 15:09:08 -07:00
Siyuan (Ryans) Zhuang	6b1b25168f	[workflow][doc] Doc for workflow checkpointing (#23510 )	2022-03-27 12:18:14 -07:00
shrekris-anyscale	65d72dbd91	[serve] Make `serve.shutdown()` shut down remote Serve applications (#23476 )	2022-03-25 18:27:34 -05:00
Amog Kamsetty	7fd7efc8d9	[AIR] Do not deepcopy RunConfig (#23499 ) RunConfig is not a tunable hyperparameter, so we do not need to deep copy it when merging parameters with Ray Tune's param_space.	2022-03-25 13:12:17 -07:00
Edward Oakes	cf7b4e65c2	[serve] Implement `serve.build` (#23232 ) The Serve REST API relies on YAML config files to specify and deploy deployments. This change introduces `serve.build()` and `serve build`, which translate Pipelines to YAML files. Co-authored-by: Shreyas Krishnaswamy <shrekris@anyscale.com>	2022-03-25 13:36:59 -05:00
shrekris-anyscale	be216a0e8c	[serve] Raise error in `test_local_store_recovery` (#23444 )	2022-03-25 13:36:51 -05:00
shrekris-anyscale	891301ff54	[serve] [docs] Add tip about `serve status` (#23481 ) The `serve status` command allows users to get their deployments' status info through the CLI. This change adds a tip to the health-checking documentation to inform users about `serve status`.	2022-03-25 13:36:15 -05:00
dependabot[bot]	e69f7f33ee	[tune](deps): Bump optuna in /python/requirements/ml (#19669 ) Bumps [optuna](https://github.com/optuna/optuna) from 2.9.1 to 2.10.0. - [Release notes](https://github.com/optuna/optuna/releases) - [Commits](https://github.com/optuna/optuna/compare/v2.9.1...v2.10.0) --- updated-dependencies: - dependency-name: optuna dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Kai Fricke <kai@anyscale.com>	2022-03-25 17:58:22 +00:00
Sven Mika	7cb86acce2	[RLlib] trainer_template.py: hard deprecation (error when used). (#23488 )	2022-03-25 18:25:51 +01:00
Jan Weßling	f78404da4a	[serve] Add ensemble model example to docs (#22771 ) Added ensemble model examples to the Documentation. That was needed, due to a user request and there was no methodology outlining the creation of higher level ensemble models. Co-authored-by: Jiao Dong <sophchess@gmail.com>	2022-03-25 11:17:54 -05:00
ddelange	e109c13b83	[ci] Clean up ray-ml requirements (#23325 ) In https://github.com/ray-project/ray/blob/ray-1.11.0/docker/ray-ml/Dockerfile, the order of pip install commands currently matters (potentially a lot). It would be good to run one big pip install command to avoid ending up with a broken env. Co-authored-by: Kai Fricke <krfricke@users.noreply.github.com>	2022-03-25 15:59:54 +00:00
Philipp Moritz	46c1b98b2f	[ci/lint] Fix linkcheck flakiness (#23482 ) As seen in https://buildkite.com/ray-project/ray-builders-branch/builds/6736#de2e23c9-3ec1-4c1b-83cb-41ae658ef1f8	2022-03-25 15:58:24 +00:00
Maxim Egorushkin	3e7ef04203	Don't rsync checkpoint_tmp directories. (#18434 ) checkpoint_tmpxxxxxx directories must not be synced from the worker nodes to the head node. Co-authored-by: Maxim Egorushkin <maxim.egorushkin@gmail.com> Co-authored-by: Kai Fricke <kai@anyscale.com> Co-authored-by: Kai Fricke <krfricke@users.noreply.github.com>	2022-03-25 15:50:38 +00:00
Kai Fricke	940c028540	[ci] Clean up artifacts before/after jobs (#23463 ) We sometimes end up with stale wheel uploads from previous runs of a Buildkite agent. The result is that commit wheels are being overwritten from old build jobs - effectively breaking the wheel build logic. Example: This Agent: https://buildkite.com/organizations/ray-project/agents/4b955117-2f6c-4849-b703-3457daf69f89 - builds wheels (in post-wheels tests) for a35ebc945b - and then runs both the Ray CPP worker and the Train + Tune tests in 6746e9f - Usually these two tests shouldn't provide artifacts at all, but they do - these are the wheels from a35ebc945b though! Meaning these are uncleaned leftovers from the first build task. - See here for proof of artifact upload: https://buildkite.com/ray-project/ray-builders-pr/builds/27622#d11bc514-ebd8-4e0c-a2ce-826b9bad27de The solution is thus to always clean up the artifacts directory in the worker, i.e. `rm -rf /artifact-mount/*` This PR adds two of such clean up instructions - once before commands are run and once after artifacts are uploaded. We can probably just do either, but it doesn't hurt to have both.	2022-03-25 13:07:20 +00:00
Brett Göhre	f5e492ea8a	[Docs] optuna notebook (#23477 )	2022-03-25 09:04:53 +01:00
mwtian	c2404cce62	avoid adding gpu utilization when unavailable (#23468 ) From #22954, GPU utilization can be unavailable for consumer hardware. So dashboard should not assume the value cannot be None. There might be a better way to represent "not reported". But currently utilizations are summed up which makes using non-zero to represent "not reported" hard to do.	2022-03-24 22:48:17 -07:00
Antoni Baum	ebb592b2ca	[Tune/Train] Make `MLflowLoggerUtil` copyable (#23333 ) Makes sure mlflow module is not saved as an attribute in MLflowLoggerUtil, which was causing an exception when running deepcopy on the callback.	2022-03-24 17:48:02 -07:00
Max Pumperla	60054995e6	[docs] fix doctests and activate CI (#23418 )	2022-03-24 17:04:02 -07:00
Siyuan (Ryans) Zhuang	d39cef5725	[workflow] Deprecate "workflow.step" [Part 1 - most common cases] (#23456 ) * replacing workflow decorator * replacing workflow decorator	2022-03-24 13:11:46 -07:00
Stephanie Wang	d67a4f5c88	[datasets] Fix missing arg in RandomIntRowDatasource (#23255 ) Previously failed with ``` E ray.exceptions.RayTaskError(TypeError): ray::_prepare_read() (pid=166631, ip=10.103.212.102) E File "/home/swang/ray/python/ray/data/read_api.py", line 902, in _prepare_read E return ds.prepare_read(parallelism, **kwargs) E File "/home/swang/ray/python/ray/data/datasource/datasource.py", line 331, in prepare_read E input_files=None, E TypeError: __init__() missing 1 required keyword-only argument: 'exec_stats' ``` This PR adds the missing arg.	2022-03-24 13:05:01 -07:00
Sven Mika	22c9c4aa39	[RLlib] Slate-Q +GPU torch bug fix. (#23464 )	2022-03-24 17:39:33 +01:00
Dmitri Gekhtman	9ce221f514	Disable KubeRay tests on windows. (#23453 ) This PR disables KubeRay tests on windows, because they're not relevant there.	2022-03-24 08:11:17 -07:00
xwjiang2010	4f34b53e83	[AIR] Add tuner test. (#23364 ) Add tuner tests. These tests are mainly focusing on non ray client mode, including successful runs, and failures in both driver and trainer side and resume. One issue surfaced through writing the tests (which probably means the API is not quite right) is whether RunConfig should be supplied in Tuner.init v.s. Tuner.fit(). At least for some fields in RunConfig, we want to be able to change it across runs (e.g. callbacks). Plus with current impl, it's not possible to checkpoint "stateful" callbacks, which could confuse our users. cc @ericl for API inputs. See "test_tuner_with_xgboost_trainer_driver_fail_and_resume" (search for hack). The PR also cleans up some API docs. Fixes some bugs in loading trial from checkpoint, namely get_default_resource (which probably is not necessary given self.placement_group_factory is already set anyways) is called with an empty config, as self.config is only loaded through __setstate__, which happens later than get_default_resource. Remove the call to get_default_resource when loading trials from checkpoint.	2022-03-24 14:54:21 +00:00

1 2 3 4 5 ...

12011 commits