hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-07 02:51:39 -05:00

Author	SHA1	Message	Date
Guyang Song	ad56b9b432	[runtime env] redefine runtime env to protobuf (#19511 )	2021-11-20 16:54:42 +08:00
Jiajun Yao	255bdc8fb1	Make fake node provider thread safe (#20591 ) We may have multiple NodeLauncher threads access the same node provider so it should be thread safe.	2021-11-19 18:59:38 -08:00
Jiao	12c11894e8	[Jobs] Add documentation for ray job submission (#20530 )	2021-11-19 16:59:05 -08:00
Amog Kamsetty	e1cd7b0016	[Train] Propagate env vars to `BackendExecutor` (#20523 ) Propagates environment variables to BackendExecutor actor using runtime envs. Also actually run test_callbacks in CI. Note that there is an issue with runtime envs: #20587. But this only happens if you shutdown Ray and start a new session again.	2021-11-19 15:36:03 -08:00
Alex Wu	4cc225e9d4	Revert "Revert "[core] Nested task support via task depth + backpressure" (#20438 )" (#20443 ) This PR reverts the previous revert with the following minor changes. Worker capping is off by default. The cap feature flag is on the for the tests that explicitely require it.	2021-11-19 15:22:35 -08:00
Antoni Baum	c385756817	[datasets] Add an env var for progress bar behavior (#20586 ) Adds a RAY_DATA_DISABLE_PROGRESS_BARS env var to control the default progress bar behavior. The default value is "0". Setting it to "1" disables progress bars, unless they are reenabled again by the set_progress_bars method.	2021-11-19 15:09:58 -08:00
Jiao	1a00964902	[job submission] Fix job sdk's lazy import of requests that led to minimal build failure (#20577 )	2021-11-19 17:04:22 -06:00
Mark	379732a181	Bump abseil-cpp LTS 20211102 for clang-13 build (#20565 ) Abseil LTS 20210324.2 will fail the compilation with clang-13. After this version bump, ray can be successfully built with clang-13.	2021-11-19 14:09:40 -08:00
architkulkarni	42085fd3d5	[runtime env] [Doc] Add concepts and basic workflows (#20222 ) Address followup comments from https://github.com/ray-project/ray/pull/19863 - Add short "Concepts" section - Add more section headings to break up the text - Add "Workflow: Local Files" example - Add "Workflow: Library development" example	2021-11-19 13:58:50 -08:00
mwtian	da79f24e8c	[Core][Pubsub] Refactor to prepare for migrating logging to Ray pubsub (#20560 ) ## Why are these changes needed? Publisher and subscriber for logs, in driver, dashboard and tests are refactored to make it easier to support using Ray pubsub for logs. Actual support of Ray pubsub for logs will be added later in #20492. This PR does not intend to introduce any behavior change. ## Related issue number	2021-11-19 12:28:37 -08:00
Chen Shen	77a8723bba	[Core][actor out-of-order execution 6/n] plumbing work to make it work e2e (#20177 ) This PR is the last PR that enables out of order execution. Previous PR: #20176 In this PR specifically, we added an execute_out_of_order option to .options call, which creates the actor with both out_of_order_submit_queue and out_of_order_scheduling queue. this PR also added @simon-mo original case for testing.	2021-11-19 11:05:18 -08:00
Dmitri Gekhtman	41af24c2ea	Head pod identity can change (#20566 )	2021-11-19 09:57:48 -08:00
Eric Liang	79911510d3	Raise better error message when workers are killed with SIGTERM in k8s (#20557 ) In k8s, sigterm almost always means the pod was killed due to memory limits. Raise a better error message there.	2021-11-19 09:36:37 -08:00
Chen Shen	f0e8d66a85	[Core][Refactor CoreWorker 1/n] move CoreWorkerOptions to its own file #19675 Why are these changes needed? This is a serial of PRs to make CoreWorkerProcess thread-safe and CoreWorker Code easy to read. [#19675 #19677 #19678 #19679] Move CoreWorkerOptions out of core_worker.h; makes the code easier to read. Next PR: #19677	2021-11-19 09:24:30 -08:00
Alex Wu	24f27203ba	[hotfix] Fix inference nightly test by upgrading numpy (#20546 ) The ray-ml image depends on numpy ~=1.19.2 via the tensorflow==2.6 requirement. Unfortunately that's incompatible with Dataset (see here #20258 (comment)). This PR upgrades the numpy dependency only for the nightly test.	2021-11-19 08:15:23 -08:00
shrekris-anyscale	b910d7e9e1	[runtime_env] Remove deprecated username-password GitHub use case from doc (#20558 )	2021-11-19 10:03:44 -06:00
Artur Niederfahrenhorst	d07e50e957	[RLlib] Replay buffer API (cleanups; docstrings; renames; move into `rllib/execution/buffers` dir) (#20552 )	2021-11-19 11:57:37 +01:00
gjoliver	18862f9f44	[RLlib] Add a comment in the doc string of `on_learn_on_batch` callback function. (#20456 )	2021-11-19 10:49:07 +01:00
Ameer Haj Ali	3c308667f1	[Tune] Fix checkpointing error message on K8s (#20559 ) This commit improves the error message to guide users to setup cloud checkpointing if trial checkpoint syncing failed.	2021-11-19 09:17:38 +00:00
Sven Mika	9d5c4a9d21	[RLlib] API reference pages: `rllib/env` package only. (#20486 )	2021-11-19 10:06:40 +01:00
Avnish Narayan	b6077a36d4	[RLlib; Pre-checks/better failure behavior]: Env Checker for Gym Environments (#20481 )	2021-11-19 09:41:03 +01:00
Alex Wu	88266a6fce	Revert "Revert "[Docs] More detailed M1 Mac installation instructions"" (#20549 ) Reverts ray-project/ray#20547	2021-11-18 20:18:37 -08:00
Eric Liang	65a8698e82	Raise the dataset block size limit to 2GiB (#20551 ) The default block size of 500MiB seems too low for some common workloads, e.g. shuffling 500GB. This creates 1000 blocks which means 1 million intermediate shuffle objects until we implement #20500.	2021-11-18 19:36:10 -08:00
Clark Zinzow	2d50bf1302	[Datasets] Bump NumPy version to >= 1.19.0 for Python 3.6. (#20542 ) Datasets groupby boundary sampling requires `numpy>=1.19.0` otherwise it fails to concatenate the Arrow table columns.	2021-11-18 17:33:06 -08:00
Clark Zinzow	462e389791	[Datasets] Fix empty Dataset.iter_batches() when trying to prefetch more blocks than exist in the dataset (#20480 ) Before this PR, `ds.iter_batches()` would yield no batches if `prefetch_blocks > ds.num_blocks()` was given, since the sliding window semantics were to return no windows if `window_size > len(iterable)`. This PR tweaks the sliding window implementation to always return at least one window, even if the one window is smaller than the given window size.	2021-11-18 17:02:54 -08:00
Simon Mo	add2450b92	[CI] [Hotfix] Skip test_standalone (#20556 )	2021-11-18 16:47:18 -08:00
Richard Liaw	c964455642	Revert "[Docs] More detailed M1 Mac installation instructions" (#20547 ) Reverts ray-project/ray#20512 due to lint errors.	2021-11-18 12:06:57 -08:00
Alex Wu	a811b2b6d7	[hotfix] Fix stress_test_many_tasks cluster environment (#20519 ) This should fix the long running release tests that are failing to build their app configs. It seems like pip install ray[all] now downgrades the ray version. It's unclear why, but most likely, a dependency has pinned the ray version now. This PR explicitely install the version of Ray that we want after the pip install ray[all] to fix the problem.	2021-11-18 11:51:46 -08:00
Amog Kamsetty	3f1092fb3d	[Release] Revert impala app config (#20397 )	2021-11-18 11:24:22 -08:00
Antoni Baum	0b14f38ac7	[tune] Multi-objective support for Optuna (#20489 ) This PR adds multi-objective support for Optuna searchers, including a test and example. Co-authored-by: gjoliver <jungong@anyscale.com>	2021-11-18 18:47:29 +00:00
Simon Mo	7143d5d494	[Serve] Bump timeout for test_standalone to fix windows (#20543 )	2021-11-18 10:00:23 -08:00
Alex Wu	540c9e35d1	[Docs] More detailed M1 Mac installation instructions (#20512 ) This PR adds more detail the M1 mac installation instructions following the bug bash.	2021-11-18 09:35:43 -08:00
Sven Mika	7a585fb275	[RLlib; Documentation] RLlib README overhaul. (#20249 )	2021-11-18 18:08:40 +01:00
Edward Oakes	d26c9e67e8	[job submission] Add a `message` to the JobStatus to return more detailed errors (#20491 )	2021-11-18 10:15:23 -06:00
shrekris-anyscale	a91ddbdeb9	Add `smart_open` dependency to `ray[default]` (#20420 )	2021-11-18 10:00:30 -06:00
Chen Shen	2012b469f6	fix gcs client hang (#20531 )	2021-11-18 07:28:15 -08:00
qicosmos	a49c1d5f55	[C++] Deprecated global named actor and global PGs. (#20468 ) Why are these changes needed? This PR removes global named actor and global PGs. Related issue number #20460	2021-11-18 23:21:59 +08:00
Simon Mo	d7f208dea4	[Releaes] Make e2e.py link clickable on buildkite (#20436 ) Adds log formatting to output clickable links to buildkite console logs	2021-11-18 12:45:59 +00:00
SangBin Cho	140a180ebb	[xgboost] Fix flaky train_small test (#20529 ) Xgboosts train_small timed out because of a CPU borrowing feature related to placement groups. The root bug will be fixed in the coming weeks, but this PR makes the release test consistently pass by requesting 0 CPUs for the remote wrapper script.	2021-11-18 10:20:08 +00:00
shrekris-anyscale	65a023ef71	[runtime_env][docs] Add documentation on using remote URIs for runtime environments (#20352 )	2021-11-17 23:17:48 -06:00
Edward Oakes	eae523159f	[job submission] Prefix job ID with `raysubmit_` and pass `job_name` metadata (#20490 )	2021-11-17 21:48:22 -06:00
Amog Kamsetty	9796ae56d5	[Train][Data] Change usages of `iter_datasets` to `iter_epochs` (#20487 )	2021-11-17 18:05:51 -08:00
Gagandeep Singh	33b4245df2	Fix race condition when starting redis (#19836 ) Co-authored-by: Philipp Moritz <pcmoritz@gmail.com>	2021-11-17 17:43:35 -08:00
Simon Mo	c85e9e69b3	[Serve] Change multi_deployment_1k_noop_replica threshold (#20514 )	2021-11-17 17:25:54 -08:00
Yi Cheng	cbf5826040	[workflow] Fix workflow event doc typo (#20465 ) In the example, it says `after_checkpoint`, but this should be `event_checkpointed`	2021-11-17 16:18:20 -08:00
Amog Kamsetty	4cbcb11458	[Docker] Add commit as label (#20504 ) Adds the Ray commit sha as a label for the docker image.	2021-11-17 15:20:41 -08:00
Richard Liaw	1cadd61917	Fix horovod failing tests by pinning down (#20484 )	2021-11-17 13:54:25 -08:00
Sven Mika	56619b955e	[RLlib; Documentation] Some docstring cleanups; Rename RemoteVectorEnv into RemoteBaseEnv for clarity. (#20250 )	2021-11-17 21:40:16 +01:00
gjoliver	724a140795	[rllib] Make sure json can serialize result dict (#20439 ) We may have fields in the result dict that are or None. Make sure our results are json serializable.	2021-11-17 10:27:00 -08:00
xwjiang2010	03aec4e04a	[Tune] Remove `runner` argument in start_trial. (#20464 ) This internal legacy argument was not used by any code.	2021-11-17 16:59:57 +00:00

... 3 4 5 6 7 ...

10685 commits