hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-06 18:41:40 -05:00

Author	SHA1	Message	Date
Eric Liang	f2401a14d9	[air] Explicitly list out the args for BatchPredictor.predict_pipelined (#26551 ) Signed-off-by: Eric Liang <ekhliang@gmail.com>	2022-07-13 22:30:32 -07:00
Scott Cheng	1bc44c13fb	Update Python3.10 in docs (#26463 ) Make it clear to users that ray supports Python 3.10	2022-07-13 20:08:56 -07:00
Stephanie Wang	6ef26cd8ff	[core] Cancel pending dependency resolution before failing a task (#26267 ) Actor tasks are sometimes failed while their dependencies are still being resolved. This can cause hanging or crashes when we resolve the dependencies for a task that has already been canceled. It can lead to a crash from the ref counter when, for the same actor, actor task 1 depends on actor task 2. The sequence is: Actor tasks 1 and 2 queued, 1 depends on 2. Fail actor task 1. We clear its refs, including its dependency on 2. Fail actor task 2. We store an error as its return value. Since task 1 depends on it, we inline the dependency and try to clear task 1's refs again, causing a ref counting error because we already cleared them in step 2. This PR fixes the issue by canceling dependency resolution for tasks before failing them. This involves some refactoring of the LocalDependencyResolver. Most of the changes are for testing (split out the unit tests for LocalDependencyResolver into their own suite). Related issue number Closes #18908.	2022-07-13 14:39:11 -07:00
Sihan Wang	b606169cb5	[Serve] Promote autoscaling feature (#26393 ) 1. get rid of the private attribute 2. fix unit test 3. docs and workflows	2022-07-13 14:38:38 -05:00
Sven Mika	ab10890e90	Revert "Bump pytest from 5.4.3 to 7.0.1" (breaks lots of RLlib tests for unknown reasons) (#26517 )	2022-07-13 11:19:30 -07:00
Antoni Baum	cc7115f6a2	[Tune/CI] Fix tune-sklearn notebook example (#26470 ) Fixes the tune-sklearn notebook example as found in #26410 Signed-off-by: Antoni Baum <antoni.baum@protonmail.com>	2022-07-13 18:14:36 +01:00
Sihan Wang	e2cac0b324	[Serve][Part1] Update the tests to use graph deploy (#26310 )	2022-07-13 09:53:51 -07:00
Ricky Xu	365ffe21e5	[Core \| State Observability] Implement API Server (Dashboard) HTTP Requests Throttling (#26257 ) This is to limit the max number of HTTP requests the dashboard (API server) will accept before rejecting more requests. This will make sure the observability requests do not overload the downstream systems (raylet/gcs) when delegating too many concurrent state observability requests to the cluster.	2022-07-13 09:05:26 -07:00
Antoni Baum	ddb5572040	[Tune/CI] Fix Hyperopt notebook example (#26469 ) Fixes failing hyperopt notebook in CI (as found in #26410). The cause was a mismatch between keys in points to evaluate and the search space - now, an informative exception will be raised. Signed-off-by: Antoni Baum <antoni.baum@protonmail.com>	2022-07-13 16:50:11 +01:00
Amog Kamsetty	8ca5584b9f	Annotate more api (#26501 )	2022-07-12 22:29:14 -07:00
clarng	b2cdd45e7c	Update import sorting blacklist, enable sorting for experimental dir (#26101 ) Why are these changes needed? There are directories that we don't lint / format. Ensure they are also the case for the import sorting tool. Enable sorting for python/experimental to show case how to enable sorting for a directory as we convert more of the directories to be automatically sorted by the tool.	2022-07-12 21:25:58 -07:00
Riatre	2cdb76789e	Bump pytest from 5.4.3 to 7.0.1 (#26334 ) See #23676 for context. This is another attempt at that as I figured out what's going wrong in `bazel test`. Supersedes #24828. Now that there are Python 3.10 wheels for Ray 1.13 and this is no longer a blocker for supporting Python 3.10, I still want to make `bazel test //python/ray/tests/...` work for developing in a 3.10 env, and make it easier to add Python 3.10 tests to CI in future. The change contains three commits with rather descriptive commit message, which I repeat here: Pass deps to py_test in py_test_module_list Bazel macro py_test_module_list takes a `deps` argument, but completely ignores it instead of passes it to `native.py_test`. Fixing that as we are going to use deps of py_test_module_list in BUILD in later changes. cpp/BUILD.bazel depends on the broken behaviour: it deps-on a cc_library from a py_test, which isn't working, see upstream issue: https://github.com/bazelbuild/bazel/issues/701. This is fixed by simply removing the (non-working) deps. Depend on conftest and data files in Python tests BUILD files Bazel requires that all the files used in a test run should be represented in the transitive dependencies specified for the test target. For py_test, it means srcs, deps and data. Bazel enforces this constraint by creating a "runfiles" directory, symbolic links files in the dependency closure and run the test in the "runfiles" directory, so that the test shouldn't see files not in the dependency graph. Unfortunately, the constraint does not apply for a large number of Python tests, due to pytest (>=3.9.0, <6.0) resolving these symbolic links during test collection and effectively "breaks out" of the runfiles tree. pytest >= 6.0 introduces a breaking change and removed the symbolic link resolving behaviour, see pytest pull request https://github.com/pytest-dev/pytest/pull/6523 for more context. Currently, we are underspecifying dependencies in a lot of BUILD files and thus blocking us from updating to newer pytest (for Python 3.10 support). This change hopefully fixes all of them, and at least those in CI, by adding data or source dependencies (mostly for conftest.py-s) where needed. Bump pytest version from 5.4.3 to 7.0.1 We want at least pytest 6.2.5 for Python 3.10 support, but not past 7.1.0 since it drops Python 3.6 support (which Ray still supports), thus the version constraint is set to <7.1. Updating pytest, combined with earlier BUILD fixes, changed the ground truth of a few error message based unit test, these tests are updated to reflect the change. There are also two small drive-by changes for making test_traceback and test_cli pass under Python 3.10. These are discovered while debugging CI failures (on earlier Python) with a Python 3.10 install locally. Expect more such issues when adding Python 3.10 to CI.	2022-07-12 21:14:35 -07:00
Eric Liang	9de1add073	[Datasets] Autodetect dataset parallelism based on available resources and data size (#25883 ) This PR defaults the parallelism of Dataset reads to `-1`. The parallelism is determined according to the following rule in this case: - The number of available CPUs is estimated. If in a placement group, the number of CPUs in the cluster is scaled by the size of the placement group compared to the cluster size. If not in a placement group, this is the number of CPUs in the cluster. If the estimated CPUs is less than 8, it is set to 8. - The parallelism is set to the estimated number of CPUs multiplied by 2. - The in-memory data size is estimated. If the parallelism would create in-memory blocks larger than the target block size (512MiB), the parallelism is increased until the blocks are < 512MiB in size. These rules fix two common user problems: 1. Insufficient parallelism in a large cluster, or too much parallelism on a small cluster. 2. Overly large block sizes leading to OOMs when processing a single block. TODO: - [x] Unit tests - [x] Docs update Supercedes part of: https://github.com/ray-project/ray/pull/25708 Co-authored-by: Ubuntu <ubuntu@ip-172-31-32-136.us-west-2.compute.internal>	2022-07-12 21:08:49 -07:00
Clark Zinzow	12ea100527	Revert "Object GC for block splitting inside the dataset splitting (#26196 )" (#26495 ) This reverts commit `45ba0e3cac`. Failures in the Train GPU job started popping up involving lost references around when this PR was merged; there was an ongoing failure that was reverted that overlaps this PR, but this PR is the most likely culprit for this particular lost reference issue, so we should try reverting the PR. - Flakey test tracker: https://flakey-tests.ray.io/ - Example failure: https://buildkite.com/ray-project/ray-builders-branch/builds/8585#0181f423-0fe2-42b5-9dd8-47d2c7f9efa7	2022-07-12 18:44:51 -07:00
brucez-anyscale	57258335bd	[Serve] Fix test_cli flakiness (#26471 )	2022-07-12 17:57:08 -07:00
Amog Kamsetty	e6c04031fd	Revert "[Train] Add support for handling multiple batch data types for prepare_data_loader (#26386 )" (#26483 ) This reverts commit `36229d1234`.	2022-07-12 17:18:46 -07:00
truelegion47	980a59477d	[Serve] [AIR] Adding reconfigure method to model deployment (#26026 )	2022-07-12 17:06:33 -07:00
Jian Xiao	45ba0e3cac	Object GC for block splitting inside the dataset splitting (#26196 ) The pipeline will spill objects when splitting the dataset into multiple equal parts. Co-authored-by: Ubuntu <ubuntu@ip-172-31-32-136.us-west-2.compute.internal>	2022-07-12 11:34:52 -07:00
Philipp Moritz	b155bc4a54	Add ray/widgets/templates/ files to wheel (fix #26452 ) (#26457 ) Add the html template files in `ray/widgets/templates/` to the wheel to make sure the Jupyter widget that is displayed in `ray.init()` works for the Ray wheels.	2022-07-12 11:23:57 -07:00
Kai Fricke	ae7e30ddc8	[air/lightgbm] Hotfix lightgbm predictor for categoricals (#26467 ) #26442 didn't trigger doc tests (fixed with #26466). The PR broke lightgbm prediction with categorical variables - this PR fixes this. In a follow-up we should specifically test prediction with categorical variables. Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-12 18:19:58 +01:00
Vishnu Deva	36229d1234	[Train] Add support for handling multiple batch data types for prepare_data_loader (#26386 ) When working with Ray Train, using the ray.train.torch.prepare_data_loader method with a dataset that returns a dictionary instead of a tuple from its __getitem__ method causes issues. Co-authored-by: matthewdeng <matthew.j.deng@gmail.com>	2022-07-12 10:16:09 -07:00
Antoni Baum	8bb67427c1	[AIR] Discard returns of train loops in Trainers (#26448 ) Discards returns of user defined train loop functions to prevent deser issues with eg. torch models. Those returns are not used anywhere in AIR, so there is no loss of functionality.	2022-07-12 10:14:06 -07:00
Guyang Song	781c2a7834	[runtime env] plugin refactor[3/n]: support strong type by @dataclass (#26296 )	2022-07-13 00:40:42 +08:00
Antoni Baum	b3878e26d7	[AIR] Fix `ResourceChangingScheduler` not working with AIR (#26307 ) This PR ensures that the new trial resources set by `ResourceChangingScheduler` are respected by the train loop logic by modifying the scaling config to match. Previously, even though trials had their resources updated, the scaling config was not modified which lead to eg. new workers not being spawned in the `DataParallelTrainer` even though resources were available. In order to accomplish this, `ScalingConfigDataClass` is updated to allow equality comparisons with other `ScalingConfigDataClass`es (using the underlying PGF) and to create a `ScalingConfigDataClass` from a PGF. Please note that this is an internal only change intended to actually make `ResourceChangingScheduler` work. In the future, `ResourceChangingScheduler` should be updated to operate on `ScalingConfigDataClass`es instead of PGFs as it is now. That will require a deprecation cycle.	2022-07-12 17:36:42 +01:00
Sihan Wang	f5c5215fe6	[Serve] Add deprecated warnings (#26374 )	2022-07-12 09:35:16 -07:00
Guyang Song	22dfd1f1f3	Revert "Revert "[runtime env] plugin refactor[2/n]: support json sche… (#26255 )	2022-07-12 23:58:18 +08:00
Kai Fricke	adfdc26dd3	[air] Test predictors with all data batch types (#26442 ) This adds a parameterized `test_predict` test for all predictors to test prediction with all data batch types. Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-12 13:58:36 +01:00
Dmitri Gekhtman	8f8f036957	[autoscaler][kuberay] Deflake KubeRay autoscaling test (#26411 ) Improves stability of KubeRay autoscaling test.	2022-07-12 00:56:36 -07:00
Archit Kulkarni	0914e5602d	[Serve] [runtime_env] [CI] Skip flaky Ray Client test (#26400 )	2022-07-12 14:39:48 +08:00
Richard Liaw	92efc85b3b	[air/docs] checkpoints (#25901 )	2022-07-11 20:40:23 -07:00
Richard Liaw	191921f4ec	[docs] Fix pytest and add stacklevel (#26340 )	2022-07-11 19:43:37 -07:00
Tao Wang	1de0d35cda	[core][c++ worker]Add namespace support for c++ worker (#26327 )	2022-07-12 09:58:26 +08:00
Antoni Baum	65ea710e30	[Docs] Update Train user guide to use the new APIs (#26091 )	2022-07-11 15:10:10 -07:00
Chen Shen	2c5c0f6cee	[Core] ensure uniqueness in spilled file name (#26420 ) There are cases that same object is being spilled twice due to failures. This made two spill worker overwrites the same file and causing corruption. The fix is as simple as ensure the uniqueness of the file. close #26395	2022-07-11 14:39:44 -07:00
Jiao	d95dc2f2e5	[AIR][GPU Batch Prediction] Add basic support for GPU batch prediction (#26251 ) This PR adds GPU support for pytorch and tensorflow predictor, as well as automatic setting `use_gpu` flag in `BatchPredictor`. Notable changes: - Added `use_gpu` flag in the constructor of `TorchPredictor` and `TensorflowPredictor` (note it's slightly different from our latest design doc that puts this flag at `predict()` call) - Added `use_gpu` flag to `SklearnPredictor` so its interface is compatible with `BatchPredictor` - Code to move both model weights and input tensor to default visible GPU at index 0 if flag is set - parametrized existing predictor tests to use GPU for both CPU & GPU coverage - Changed BUILD CI tests with an added `gpu` tag (I'm not 100% sure if that's a right way tho) Follow ups: https://github.com/ray-project/ray/issues/26249 is created in case our host has multiple GPU devices. It's a bit out of scope for this PR, but for GPU batch inference ideally we should be able to evenly use all GPU devices on host where CPU & DRAM are busy with pre-fetching + data movement to GPU. We might approximately do the same by scheduling same # of Predictor instances on the host, but that's worth verifying once benchmarks are set.	2022-07-11 13:04:15 -07:00
Kai Fricke	753f5feaf4	[tune] Remove TrialCheckpoint class (#25406 ) The old user-facing TrialCheckpoint class has been deprecated in favor of `ray.ml.Checkpoint` and will be removed with this PR. The main change in this PR is to delete the old `TrialCheckpoint` class and replace remaining API calls (e.g. `checkpoint.local_path`) with the correct AIR equivalents. One issue that comes up is that with Ray client usage, checkpoint directories are not available on the local node (the client). Thus, we can't construct `Checkpoint` objects easily. (Previously, the TrialCheckpoint object held a reference to the location, even if it is not locally available). There are ongoing discussions on how to resolve this in the future. For now, we print an error when such a checkpoint is requested. Depends on #25805 Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-11 20:08:10 +01:00
Philipp Moritz	dae4ec2f23	Fix dashboard link in HTML reprs for `ClientContext` and `WorkerContext` (#26431 ) This fixes the dashboard link in https://github.com/ray-project/ray/pull/25730 -- without this I'm getting <img width="1378" alt="Screen Shot 2022-07-09 at 8 08 06 PM" src="https://user-images.githubusercontent.com/113316/178129698-7ef19ee3-d577-4fd9-a4d5-0cee1ca35f5f.png"> because Jupyter is interpreting the URL as relative to the notebook URL.	2022-07-09 23:30:40 -07:00
Richard Liaw	5892a76a44	[air/tune] Documentation testing fixes (#26409 )	2022-07-09 19:47:21 -07:00
Yi Cheng	39cb1e5f97	[core][1/2] Improve liveness check in GCS (#26405 ) CheckAlive in GCS is only for checking GCS's liveness. But we also need to check the liveness for raylet. In KubeRay, we can check the liveness directly by monitoring the raylet's liveness. But it's not good enough given that raylet's process liveness is not directly related to raylet's liveness. For example, during a network partition, raylet is not able to connect to GCS. GCS mark raylet as dead. So for the cluster, although raylet process is still alive, it can't be treated as alive because GCS has told all the nodes that it's dead. So for KubeRay, it also needs to talk with GCS to check whether it's alive or not. This PR extends the CheckAlive API to include raylet address so that we can query GCS to get the cluster status directly.	2022-07-09 16:32:31 +00:00
Siyuan (Ryans) Zhuang	7fcf0adebb	[Workflow] Minor refactoring of workflow exceptions (#26398 ) * minor refactoring	2022-07-09 00:46:43 -07:00
Siyuan (Ryans) Zhuang	b0e913fd07	[workflow] Workflow queue (#24697 ) * implement workflow queue	2022-07-08 17:24:45 -07:00
Amog Kamsetty	cc43bcccb4	[AIR] Update TensorflowPredictor to new API (#26215 ) Updates TensorflowPredictor to use the new _predict_pandas API. Also as agreed upon offline, removes the extra configurations from TensorflowPredictor (column selection, concatenation) in favor of having this be done via a Preprocessor.	2022-07-08 13:04:49 -07:00
Nikita Vemuri	56716a1c1b	[dashboard] Add `RAY_CLUSTER_ACTIVITY_HOOK` to `/api/component_activities` (#26297 ) Add external hook to /api/component_activities endpoint in dashboard snapshot router Change is_active field of RayActivityResponse to take an enum RayActivityStatus instead of bool. This is a backward incompatible change, but should be ok because [dashboard] Add component_activities API #25996 wasn't included in any branch cuts. RayActivityResponse now supports informing when there was an error getting the activity observation and the reason.	2022-07-08 10:51:59 -07:00
Kai Fricke	e1a7efe148	[tune] Use `Checkpoint.to_bytes()` for store_to_object (#25805 ) We currently use our own serialization to ship checkpoints as objects. Instead we should use the Checkpoint class. This PR also adds support to create results from checkpoints pointing to object references. Depends on #26351 Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-08 18:01:20 +01:00
Antoni Baum	0e259ff844	[tune] Fix `SyncerCallback` having a size limit (#26371 ) #25655 refactored syncing but it introduced a regression - previously, dirs of any size could have been synced, but now only dirs below the default limit of 1 GB can be. This PR fixes this regression allowing dirs of any size to be synced.	2022-07-08 17:58:41 +01:00
Kai Fricke	86b9b4b7a5	[air] Serialize additional files in dict checkpoints turned dir checkpoints (#26351 ) With this PR, files put into directory checkpoints that were dict checkpoints will be serialized and retained when a subsequent to_dict() is called. This is to enable storing additional files, as e.g. needed by Ray Tune. Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-08 10:03:16 +01:00
Jiajun Yao	743e2f403a	Set RAY_USAGE_STATS_EXTRA_TAGS for release tests (#26366 ) - Record the test name for the usage stats. - Change the cluster name to indicate if it's smoke test or not.	2022-07-07 21:17:34 -07:00
Cheng Su	4e674b6ad3	[Datasets] Update docs for drop_columns and fix typos (#26317 ) We added drop_columns() API to datasets in #26200, so updating documentation here to use the new API - doc/source/data/examples/nyc_taxi_basic_processing.ipynb. In addition, fixing some minor typos after proofreading the datasets documentation.	2022-07-07 17:17:33 -07:00
Antoni Baum	ea94cda1f3	[AIR] Replace `train.` with `session.` (#26303 ) This PR replaces legacy API calls to `train.` with AIR `session.` in Train code, examples and docs. Depends on https://github.com/ray-project/ray/pull/25735	2022-07-07 16:29:04 -07:00
Yi Cheng	f2f1086868	[serve] Add healthz endpoint for HttpProxy (#26347 )	2022-07-07 14:01:42 -07:00

1 2 3 4 5 ...

7184 commits