hiro/ray - Forgejo: Beyond coding. We Forge.

hiro/ray

mirror of https://github.com/vale981/ray synced 2025-03-07 02:51:39 -05:00

Author	SHA1	Message	Date
Robert Nishihara	232601f90d	Change all table calls to use default retry behavior. (#312 ) * Change all table calls to use default retry behavior and change default retry behavior. * Add warning for table retries.	2017-02-24 12:41:32 -08:00
Robert Nishihara	7f5be96683	Remove object table tests that are failing. (#310 )	2017-02-23 13:39:59 -08:00
Alexey Tumanov	3159a78ad7	terminate photon task dispatch early when workers or resources are unavailable (#311 ) * terminate photon task dispatch early when no workers or resources available * style	2017-02-23 00:05:16 -08:00
Robert Nishihara	a6bf16f6a9	Make global scheduler periodically resubmit tasks that can't be sched… (#306 ) * Make global scheduler periodically resubmit tasks that can't be scheduled because their resource requirements are not met. * Address comments and fix bug. * Rename impossible_tasks -> pending_tasks. * Fix formatting.	2017-02-21 23:15:46 -08:00
Robert Nishihara	e399f57e6b	Let actors use GPUs. (#302 ) * Add num_cpus and num_gpus to actor decorator. * Assign GPU IDs to actors. * Add additional actor test. * Remove duplicated line. * Factor out local scheduler selection method. * Add test and simplify local scheduler selection.	2017-02-21 01:13:04 -08:00
Robert Nishihara	3e67d28922	Address numbuf compiler warnings. (#300 )	2017-02-20 22:42:03 -08:00
Stephanie Wang	334aed9fa9	Fetch the object after requesting reconstruction during ray.get (#301 ) * Fetch the object after requesting reconstruction during ray.get * revert * Fix documentation and memory leak * Fix hanging reconstruction bug * Fix for python3	2017-02-20 21:41:34 -08:00
Stephanie Wang	67c591c33b	Retry connections in photon connect, consolidate code in io.c (#294 )	2017-02-17 23:41:21 -08:00
Philipp Moritz	9973a6e37c	fix bug in numbuf serialization (#296 )	2017-02-17 23:35:41 -08:00
Stephanie Wang	a0dd3a44c0	Dynamically grow worker pool to partially solve hanging workloads (#286 ) * First pass at a policy to solve deadlock * Address Robert's comments * stress test * unit test * Fix test cases * Fix test for python3 * add more logging * White space.	2017-02-17 17:08:52 -08:00
Philipp Moritz	dd7e8d9105	Avoid segfaults in arrow if data is too large (#287 ) * arrow limits * more logging * set the right limit * update * simplify * fix * account for subsequences * fixes and deactivate arrow limit tests in travis * fixes * Minor formatting. * Add a couple more tests.	2017-02-16 15:16:20 -08:00
Robert Nishihara	88a5b4e77b	Simplify imports and exports and provide driver isolation for remote functions. (#288 ) * Remove import counter and export counter. * Provide isolation between drivers for remote functions. * Add test for driver function isolation. * Hash source code into function ID to reduce likelihood of collisions. * Fix failure test example. * Replace assertTrue with assertIn to improve failure messages in tests. * Fix failure test.	2017-02-16 11:30:35 -08:00
Philipp Moritz	12a68e84d2	Implement a first pass at actors in the API. (#242 ) * Implement actor field for tasks * Implement actor management in local scheduler. * initial python frontend for actors * import actors on worker * IPython code completion and tests * prepare creating actors through local schedulers * add actor id to PyTask * submit actor calls to local scheduler * starting to integrate * simple fix * Fixes from rebasing. * more work on python actors * Improve local scheduler actor handlers. * Pass actor ID to local scheduler when connecting a client. * first working version of actors * fixing actors * fix creating two copies of the same actor * fix actors * remove sleep * get rid of export synchronization * update * insert actor methods into the queue in the right order * remove print statements * make it compile again after rebase * Minor updates. * fix python actor ids * Pass actor_id to start_worker. * add test * Minor changes. * Update actor tests. * Temporary plan for import counter. * Temporarily fix import counters. * Fix some tests. * Fixes. * Make actor creation non-blocking. * Fix test? * Fix actors on Python 2. * fix rare case. * Fix python 2 test. * More tests. * Small fixes. * Linting. * Revert tensorflow version to 0.12.0 temporarily. * Small fix. * Enhance inheritance test.	2017-02-15 00:10:05 -08:00
Stephanie Wang	2b8e6485e3	Start and clean up workers from the local scheduler. (#250 ) * Start and clean up workers from the local scheduler Ability to kill workers in photon scheduler Test for old method of starting workers Common codepath for killing workers Common codepath for killing workers Photon test case for starting and killing workers fix build Fix component failure test Register a worker's pid as part of initial connection Address comments and revert photon_connect Set PATH during travis install Fix * Fix photon test case to accept clients on plasma manager fd	2017-02-10 12:46:23 -08:00
Alexey Tumanov	dfb6107b22	General attribute-based heterogeneity support with hard and soft constraints (#248 ) * attribute-based heterogeneity-awareness in global scheduler and photon * minor post-rebase fix * photon: enforce dynamic capacity constraint on task dispatch * globalsched: cap the number of times we try to schedule a task in round robin * propagating ability to specify resource capacity to ray.init * adding resources to remote function export and fetch/register * globalsched: remove unused functions; update cached photon resource capacity (until next photon heartbeat) * Add some integration tests. * globalsched: cleanup + factor out constraint checking * lots of style * task_spec_required_resource: global refactor * clang format * clang format + comment update in photon * clang format photon comment * valgrind * reduce verbosity for Travis * Add test for scheduler load balancing. * addressing comments * refactoring global scheduler algorithm * Minor cleanups. * Linting. * Fix array_test.py and linting. * valgrind fix for photon tests * Attempt to fix stress tests. * fix hashmap free * fix hashmap free comment * memset photon resource vectors to 0 in case they get used before the first heartbeat * More whitespace changes. * Undo whitespace error I introduced.	2017-02-09 01:34:14 -08:00
Philipp Moritz	fefc7d9b49	fix segfault in photon.Task (#253 )	2017-02-07 11:17:11 -08:00
Robert Nishihara	2d1c980ad7	Refactor local scheduler to remove worker indices. (#245 ) * Refactor local scheduler to remove worker indices. * Change scheduling state enum to int in all function signatures. * Bug fix, don't use pointers into a resizable array. * Remove total_num_workers. * Fix tests.	2017-02-05 14:52:28 -08:00
Philipp Moritz	ca254b8689	Fix stack overflow if many objects are fetched. (#237 ) * fix stack overflow if many objects are fetched * fix other stack allocations * add tests and fix linting * address stephanie's comments * fix linting * fix tests	2017-02-04 16:49:36 -08:00
Stephanie Wang	241b539ff8	Reconstruction for evicted objects (#181 ) * First pass at reconstruction in the worker Modify reconstruction stress testing to start Plasma service before rest of Ray cluster TODO about reconstructing ray.puts Fix ray.put error for double creates Distinguish between empty entry and no entry in object table Fix test case Fix Python test Fix tests * Only call reconstruct on objects we have not yet received * Address review comments * Fix reconstruction for Python3 * remove unused code * Address Robert's comments, stress tests are crashing * Test and update the task's scheduling state to suppress duplicate reconstruction requests. * Split result table into two lookups, one for task ID and the other as a test-and-set for the task state * Fix object table tests * Fix redis module result_table_lookup test case * Multinode reconstruction tests * Fix python3 test case * rename * Use new start_redis * Remove unused code * lint * indent * Address Robert's comments * Use start_redis from ray.services in state table tests * Remove unnecessary memset	2017-02-01 19:18:46 -08:00
Robert Nishihara	f69d4aaaa7	Change fetch requests in plasma manager to use a single timer. (#234 ) * Change fetch requests in plasma manager to use a single timer. * Fix manager tests, other cleanups.	2017-02-01 12:21:52 -08:00
Robert Nishihara	6703f7be6f	Provide functionality for local scheduler to start new workers. (#230 ) * Provide functionality for local scheduler to start new workers. * Pass full command for starting new worker in to local scheduler. * Separate out configuration state of local scheduler.	2017-01-27 01:28:48 -08:00
Stephanie Wang	a5c8f28f33	Plasma subscribe (#227 ) * Use object_info as notification, not just the object_id * Add a regression test for plasma managers connecting to store after some objects have been created * Send notifications for existing objects to new plasma subscribers * Continuously try the request to the plasma manager instead of setting a timeout in the test case * Use ray.services to start Redis in plasma test cases * fix test case	2017-01-25 22:57:15 -08:00
Robert Nishihara	ab8c3432f7	Add driver ID to task spec and add driver ID to Python error handling. (#225 ) * Add driver ID to task spec and add driver ID to Python error handling. * Make constants global variables. * Add test for error isolation.	2017-01-25 22:53:48 -08:00
Stephanie Wang	3c6686db08	Photon optimizations (#219 ) * Optimizations: - Track mapping of missing object to dependent tasks to avoid iterating over task queue - Perform all fetch requests for missing objects using the same timer * Fix bug and add regression test * Record task dependencies and active fetch requests in the same hash table * fix typo * Fix memory leak and add test cases for scheduling when dependencies are evicted * Fix python3 test case * Minor details.	2017-01-23 19:44:15 -08:00
Robert Nishihara	b98a63fd3a	Change get to take a timeout and multiple object IDs. (#212 ) * Change plasma_get to take a timeout and an array of object IDs. * Address comments. * Bug fix related to computing object hashes. * Add test. * Fix file descriptor leak. * Fix valgrind. * Formatting. * Remove call to plasma_contains from the plasma client. Use timeout internally in ray.get. * small fixes	2017-01-19 12:21:12 -08:00
Robert Nishihara	677a019cbd	Remove unnecessary bookkeepping in utlist in plasma client. (#215 )	2017-01-18 23:03:08 -08:00
Stephanie Wang	f1987cdc16	Split local scheduler task queue (#211 ) * Split local scheduler task queue into waiting and dispatch queue * Fix memory leak * Add a new task scheduling status for when a task has been queued locally * Fix global scheduler test case and add task status doc * Documentation * Address Philipp's comments * Move tasks back to the waiting queue if their dependencies become unavailable * Update existing task table entries instead of overwriting	2017-01-18 20:27:40 -08:00
Robert Nishihara	303d0fed3e	Prevent plasma store and manager from dying when a client dies. (#203 ) * Prevent plasma store and manager from dying when a worker dies. * Check errno inside of warn_if_sigpipe. Passing in errno doesn't work because the arguments to warn_if_sigpipe can be evaluated out of order.	2017-01-17 20:34:31 -08:00
Philipp Moritz	7f329db4b2	wait until kill operation was successful (#210 )	2017-01-17 20:15:48 -08:00
Philipp Moritz	a708e36225	Switch build system to use CMake completely. (#200 ) * switch to CMake completely ... * cleanup * Run C tests, update installation instructions.	2017-01-17 16:56:40 -08:00
Philipp Moritz	ab3448a9b4	Plasma Optimizations (#190 ) * bypass python when storing objects into the object store * clang-format * Bug fixes. * fix include paths * Fixes. * fix bug * clang-format * fix * fix release after disconnect	2017-01-09 20:15:54 -08:00
Robert Nishihara	973716d310	Use cloudpickle 0.2.2. (#189 )	2017-01-08 17:30:06 -08:00
Alexey Tumanov	674ec3a3cb	generate pytask from string and string from pytask (#188 ) * pytask creation from bytestring: saving work * pytask now works * documentation and tests * linting * Lint and fix test case	2017-01-08 02:16:40 -08:00
Stephanie Wang	c13d73b4c9	Suppress duplicate transfer requests (#185 )	2017-01-06 22:14:51 -08:00
Robert Nishihara	651aa6007a	Log profiling information from worker. (#178 ) * Log timing events on workers. * Have workers log to the event log through the local scheduler. * Fixes and address comments. * bug fix * styling	2017-01-05 16:47:16 -08:00
Johann Schleier-Smith	b1e76e582e	Check /dev/shm on Linux (#174 ) * check available shared memory when starting object store * exit with error if not enough shared memory available for object store * Some comments and formatting.	2017-01-03 12:33:29 -08:00
Stephanie Wang	6828d694ae	Test object notifications from Plasma store (#141 ) * Object notification test for Photon, and turn on valgrind for Photon C tests * Test object notification handler in the plasma manager * Fix hanging test case	2016-12-29 23:10:38 -08:00
Robert Nishihara	acf1703afd	Implement naive scheduling algorithm using local scheduler load. (#164 ) * Implement naive scheduling algorithm using local scheduler load. * Have the global scheduler estimate load on local schedulers better. * Fixes.	2016-12-28 22:33:20 -08:00
Robert Nishihara	baf835efcd	Throw Python exception if plasma store cannot create new object. (#162 ) * Propagate error messages through plasma create. * Use custom exception types instead of exception messages.	2016-12-28 11:56:16 -08:00
Robert Nishihara	10e067e5e5	Delay releasing a maximum number of bytes in the plasma client. (#160 ) * Send message from plasma client to get plasma store capacity. * Release objects from plasma client if they are too large. * Use doubly-linked list instead of ring buffer for plasma client release history. * Address comments. * Fix problem with slicing PlasmaBuffer objects. * Fix crash in plasma manager during transfer. * Formatting. * Make plasma client cache larger and make caching test not throw exceptions on Travis.	2016-12-27 19:51:26 -08:00
Robert Nishihara	26941e02aa	Attempt to free up to 20% of the plasma store capacity during eviction. (#159 )	2016-12-27 12:12:33 -08:00
Robert Nishihara	985c424172	Use redismodules for task table and result table. (#156 ) * Switch to using redis modules for task table. * Switch to using redis modules for the task table. * Fix some tests. * Fix naming and remove code duplication. * Remove duplication in redis modules and add more cleanups. * Address comments.	2016-12-25 23:57:05 -08:00
Philipp Moritz	d6695c867a	fix wait test (#158 )	2016-12-25 23:43:01 -08:00
Philipp Moritz	8309e3f355	Redis string formatting (#157 ) * redis string formatting * fixes * add documentation * fixes	2016-12-25 22:43:07 -08:00
Robert Nishihara	3d697c7ed2	Introduce local scheduler heartbeats which carry load information. (#155 ) * Introduce local scheduler heartbeats which carry load information.	2016-12-24 20:02:25 -08:00
Robert Nishihara	9bb9f8cb54	Fix bug in ray.wait. (#153 ) * Fix bug in wait implementation. * Add test that exposes previous bug.	2016-12-23 16:22:41 -08:00
Robert Nishihara	86b211f5c2	Give run_function_on_all_workers to take a worker_info dictionary including a counter. (#149 ) * Suppress Redis warnings and remove some global scheduler logging. * Pass a counter into run_function_on_all_workers indicating how many workers have begun executing this function.	2016-12-22 22:05:58 -08:00
Alexey Tumanov	46a887039e	Global scheduler - per-task transfer-aware policy (#145 ) * global scheduler with object transfer cost awareness -- upstream rebase * debugging global scheduler: multiple subscriptions * global scheduler: utarray push bug fix; tasks change state to SCHEDULED * change global scheduler test to be an integraton test * unit and integration tests are passing for global scheduler * improve global scheduler test: break up into several * global scheduler checkpoint: fix photon object id bug in test * test with timesync between object and task notifications; TODO: handle OoO object+task notifications in GS * fallback to base policy if no object dependencies are cached (may happen due to OoO object+task notification arrivals * clean up printfs; handle a missing LS in LS cache * Minor changes to Python test and factor out some common code. * refactoring handle task waiting * addressing comments * log_info -> log_debug * Change object ID printing. * PRId64 merge * Python 3 fix. * PRId64. * Python 3 fix. * resurrect differentiation between no args and missing object info; spacing * Valgrind fix. * Run all global scheduler tests in valgrind. * clang format * Comments and documentation changes. * Minor cleanups. * fix whitespace * Fix. * Documentation fix.	2016-12-22 03:11:46 -08:00
Robert Nishihara	6cd02d71f8	Fixes and cleanups for the multinode setting. (#143 ) * Add function for driver to get address info from Redis. * Use Redis address instead of Redis port. * Configure Redis to run in unprotected mode. * Add method for starting Ray processes on non-head node. * Pass in correct node ip address to start_plasma_manager. * Script for starting Ray processes. * Handle the case where an object already exists in the store. Maybe this should also compare the object hashes. * Have driver get info from Redis when start_ray_local=False. * Fix. * Script for killing ray processes. * Catch some errors when the main_loop in a worker throws an exception. * Allow redirecting stdout and stderr to /dev/null. * Wrap start_ray.py in a shell script. * More helpful error messages. * Fixes. * Wait for redis server to start up before configuring it. * Allow seeding of deterministic object ID generation. * Small change.	2016-12-21 18:53:12 -08:00
Robert Nishihara	c9c1b3e6af	Change db_connect to allow different arguments from different processes. (#142 ) * Allow db_connect to take a variable number of arguments. * Fix tests. * Fixes. * Formatting. * Fixes. * Simplifications. * Fix typo.	2016-12-20 20:21:35 -08:00

... 5 6 7 8 9 ...

474 commits