hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-06 10:31:39 -05:00

Author	SHA1	Message	Date
Jian Xiao	30cf449807	Add data ingest benchmark (#27533 ) Make sure Dataset/DatasetPipeline work performantly for data ingestion.	2022-08-05 12:31:06 -07:00
Avnish Narayan	6a31b61580	[RLlib] CQL change hparams and data reading strategy (#27451 )	2022-08-04 18:55:32 -07:00
Avnish Narayan	55209692ee	[RLlib] Deflake MARWIL and BC and remove memory leak from torch MARWIL policy (#27406 )	2022-08-03 16:53:12 -07:00
Simon Mo	8ac6d02502	[Serve][Nightly] Environment for Nightly K8s Tests (#27126 )	2022-08-02 23:05:47 -07:00
Simon Mo	8beb887bbe	[Serve] Remove release tests for checkpoint_path (#27194 )	2022-07-28 12:30:30 -07:00
Avnish Narayan	f5a9a44b9c	[RLlib] Revert Revert Fix apex long running test (#26928 )	2022-07-26 15:10:25 -07:00
Chen Shen	acbab51d3e	[Nightly] fix microbenchmark scripts (#26947 ) Signed-off-by: scv119 scv119@gmail.com Why are these changes needed? microbenchmarks failed complaining raise ValueError(f"Malformed address: {address}") ValueError: Malformed address: this is due to `55a0f7b` and fix it by set RAY_ADDRESS="local"	2022-07-24 14:16:43 -07:00
Avnish Narayan	a50a81a13a	Revert "[RLlib] Fix apex breakout release test performance. (#26867 )" (#26927 )	2022-07-23 17:27:50 +02:00
Avnish Narayan	2cfd6c2e97	[RLlib] Fix apex breakout release test performance. (#26867 )	2022-07-23 13:53:03 +02:00
Richard Liaw	96e8027c7e	[air] large tune/torch benchmark (#26763 ) Co-authored-by: Kai Fricke <krfricke@users.noreply.github.com>	2022-07-23 01:17:25 -07:00
Jiao	840b0478aa	[AIR CUJ] Add wait_for_nodes for 4x4 gpu test	2022-07-22 16:04:54 -07:00
Avnish Narayan	82395c4646	[RLlib] Put learning test into own folders (#26862 ) Co-authored-by: Artur Niederfahrenhorst <artur@anyscale.com>	2022-07-22 11:20:47 -07:00
Archit Kulkarni	e043f49957	[Serve] [CI] Increase instance size and add debug log for `autoscaling_multi_deployment` release test (#26732 )	2022-07-20 16:13:36 -07:00
Kai Fricke	2e35d47bd2	[air/train/benchmark] Add TF GPU 4x4 benchmark (#26776 )	2022-07-20 14:07:51 -07:00
Jiajun Yao	2603aea4c9	[CI] Chaos tests for dataset random shuffle 1tb (#26738 ) - Add chaos tests for dataset random shuffle 1tb: both simple shuffle and push-based shuffle - Mark dataset_shuffle_push_based_random_shuffle_1tb as stable	2022-07-19 15:16:51 -07:00
xwjiang2010	75027eb479	[air/benchmarks] train/tune benchmark (#26564 ) Making sure that tuning multiple trials in parallel is not significantly slower than training each individual trials. Some overhead is expected. Signed-off-by: Xiaowei Jiang <xwjiang2010@gmail.com> Signed-off-by: Richard Liaw <rliaw@berkeley.edu> Signed-off-by: Kai Fricke <kai@anyscale.com> Co-authored-by: Jimmy Yao <jiahaoyao.math@gmail.com> Co-authored-by: Richard Liaw <rliaw@berkeley.edu> Co-authored-by: Kai Fricke <kai@anyscale.com>	2022-07-19 18:24:39 +01:00
Richard Liaw	7e62e1187c	[air/benchmark] Torch benchmarks for 4x4 (#26692 ) Add benchmark data for 4x4 GPU setup. Signed-off-by: Richard Liaw <rliaw@berkeley.edu> Co-authored-by: Jimmy Yao <jiahaoyao.math@gmail.com> Co-authored-by: Kai Fricke <kai@anyscale.com>	2022-07-19 17:06:37 +01:00
Jiajun Yao	40a4777bc0	Mark chaos_dataset_shuffle_push_based_sort_1tb and chaos_dataset_shuffle_sort_1tb stable (#26677 ) They passed for the past 7 runs.	2022-07-18 14:34:08 -07:00
Kai Fricke	00947fd949	[air/benchmarks] Add 4x1 GPU benchmark for Torch (#26562 )	2022-07-18 12:14:10 -07:00
Jiao	98a07920d3	[AIR][CUJ] Make distributing training benchmark at silver tier (#26640 )	2022-07-17 22:07:09 -07:00
Eric Liang	0855bcb77e	[air] Use SPREAD strategy by default and don't special case it in benchmarks (#26633 )	2022-07-16 17:37:06 -07:00
Jiao	196e52ad7c	[AIR][CUJ] E2E Pytorch training (#26621 )	2022-07-16 08:23:19 -07:00
Jiao	988ffd494b	[AIR][CUJ] Add GPU bench prediction benchmark (#26614 )	2022-07-16 08:22:37 -07:00
matthewdeng	e3a096f412	[air] add bulk ingest benchmarks (#26618 )	2022-07-15 22:01:23 -07:00
xwjiang2010	a241e6a0f5	[air] Add xgboost release test for silver tier(10-node case). (#26460 ) Co-authored-by: Antoni Baum <antoni.baum@protonmail.com> Co-authored-by: Richard Liaw <rliaw@berkeley.edu>	2022-07-15 13:21:10 -07:00
Kai Fricke	213a96e239	[air/benchmarks] Add distributed Tensorflow benchmarks (CPU only) (#26519 ) Following up from #26436, this PR adds a distributed benchmark test for Tensorflow FashionMNIST training. It compares training with Ray AIR with training with vanilla PyTorch. Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-14 22:08:43 +01:00
Kai Fricke	cd95569b01	[tune/release] Add up/down scaling release test (#25392 ) This adds a nightly release test that asserts that autoscaling a cluster up and down in a Ray Tune run works. Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-13 22:57:24 +01:00
Kai Fricke	cf75cf7232	[air] Add AIR distributed training benchmark for Torch FashionMNIST (#26436 ) This PR adds a distributed benchmark test for Pytorch MNIST training. It compares training with Ray AIR with training with vanilla PyTorch. In both cases, the same training loop is used. For Ray AIR, we use a TorchTrainer with 4 CPU workers. For vanilla PyTorch, we upload a training script and kick it off (using Ray tasks) in subprocesses on each node. In both cases, we collect the end to end runtime. Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-07-13 10:53:24 +01:00
Jian Xiao	923209895d	Pipelined training test: change num of windows; log the ingestion perf (#26429 ) Why are these changes needed? Improve test perf Log the perf stats With 2 windows there are a lot of spilling, slowing down the throughput.	2022-07-11 11:03:35 -07:00
Stephanie Wang	dcc913073f	[testing] Run 100TB shuffle test nightly (#26306 ) Run this test nightly to collect more datapoints on stability and performance of 100TB shuffle.	2022-07-07 09:59:54 -07:00
Stephanie Wang	a90e53b76f	[core] Add weekly test for 100TB random shuffle (#25908 ) Adds a CI test for 100TB shuffle. There is a custom config for this nightly test to: (1) make sure each node gets 4TB of storage, (2) head node has 0 CPUs, (3) worker nodes have half their actual vCPU count. Related issue number Closes #24480.	2022-07-01 13:30:07 -07:00
Kai Fricke	e2d8e7a6ae	[ci/release/ml] Run ML release tests on staging (#26168 ) This moves all ML release tests to staging. Signed-off-by: Kai Fricke <kai@anyscale.com>	2022-06-30 13:24:28 -07:00
Kai Fricke	f9e787115f	[ci/release/core] Run `many_nodes` test on staging (#26164 ) This moves many_nodes to anyscale staging.	2022-06-29 11:07:32 -07:00
Kai Fricke	7091a32fe1	[ci/release] Support running tests on staging (#25889 ) This adds "environments" to the release package that can be used to configure some environment variables. These variables will be loaded either by an `--env` argument or a `env` definition in the test definition and can be used to e.g. run release tests on staging.	2022-06-28 10:14:01 -07:00
Jun Gong	c026374acb	[RLlib] Fix the 2 failing RLlib release tests. (#25603 )	2022-06-14 14:51:08 +02:00
Kai Fricke	c3b608f757	[tune] Fix cloud tests, mark as stable (#25583 ) #25063 broke release tests, but they've been consistently stable before. This PR fixes the tests and marks tune cloud tests as stable.	2022-06-08 17:47:54 +01:00
Stephanie Wang	f7692e4602	[core] Remove more expensive shuffle tests (#25165 ) Now that the "smaller_instances" versions of these tests are stable, we can stop running the version that uses bigger instances.	2022-05-24 18:05:18 -07:00
Jiajun Yao	00cdd8dce5	Add chaos test for dataset shuffle (#25161 ) Add chaos tests for dataset shuffle: both push-based and non-push-based.	2022-05-24 15:12:20 -07:00
Jiajun Yao	b825a839f9	Mark dataset_shuffle_push_based_sort_1tb as stable (#25162 ) dataset_shuffle_push_based_sort_1tb is consistently passing for weeks.	2022-05-24 11:07:27 -07:00
mwtian	36d6d59169	[Datasets] mark nightly test dataset_shuffle_random_shuffle_1tb_small_instances stable #24861	2022-05-17 09:56:05 -07:00
Kai Fricke	6c5229295e	[ci/release] Support running tests with different python versions (#24843 ) OSS release tests currently run with hardcoded Python 3.7 base. In the future we will want to run tests on different python versions. This PR adds support for a new `python` field in the test configuration. The python field will determine both the base image used in the Buildkite runner docker container (for Ray client compatibility) and the base image for the Anyscale cluster environments. Note that in Buildkite, we will still only wait for the python 3.7 base image before kicking off tests. That is acceptable, as we can assume that most wheels finish in a similar time, so even if we wait for the 3.7 image and kick off a 3.8 test, that runner will wait maybe for 5-10 more minutes.	2022-05-17 17:03:12 +01:00
Sven Mika	0cd7bc4054	[RLlib] Re-establish dashboard performance tests. (#24728 )	2022-05-16 13:13:49 +02:00
Kai Fricke	de69b0d6d6	[train/release] Fix horovod user test master app config (#24734 )	2022-05-14 21:20:45 -07:00
Amog Kamsetty	a36e2a8f51	[Tune] Deprecate DistributedTrainableCreator (#24453 ) Fully deprecate DistributedTrainableCreator for Ray 2.0 Closes #24453	2022-05-10 11:06:43 -07:00
mwtian	918d3601c6	[Datasets] mark nightly test dataset_shuffle_sort_1tb_small_instances stable (#24481 )	2022-05-06 15:55:59 -07:00
SangBin Cho	295b4436b3	[Nightly tests] Increase wait for nodes timeout (#24457 ) Although there's enough quota, it is possible the AWS doesn't have enough capacity to start up new nodes. According to @allenyin55, the current wait for node timeout is too short. This PR increases the timeout to 3000 seconds (50 minutes) from 600 seconds. Let's see if this can resolve the issue. If it makes things worse, I will revert it quickly (I will closely monitor the infra failure rate)	2022-05-04 19:42:21 -07:00
Sihan Wang	3f5da8af7a	[Serve] Add serve handle graph workload nightly tests (#24435 )	2022-05-04 09:07:50 -07:00
Stephanie Wang	fbbc9c33d6	Add nightly tests for push-based shuffle (#24352 ) Adds 1TB tests for push-based random shuffle and sort. Initially marked unstable.	2022-05-02 11:35:14 -07:00
Jiao	ba7cc1803a	[Deployment Graph] Add release test for long chain & wide fanout pattern (#24246 )	2022-04-29 17:03:33 -07:00
SangBin Cho	46cd7f1830	Make large multi tests to nightly + remove k8s tests (#24302 ) As discussed, to reduce backlog for large tests, we will (1) remove k8s tests (2) make large multi daily tests to nightly tests	2022-04-29 03:40:12 -07:00

1 2

99 commits