server: fix LRU hang on multiple requests same model (#28539)
* server: fix LRU hang on multiple requests same model * server: keep a queued model out of the victim pool until its waiters leave A waiter that gave up while its model was still loading left the model idle with no request behind it, and nothing recounted the free slots, so a second request queued behind it stayed queued forever. tick() was only driven by requests: join, claim and the end of a proxied request. Keep the queue entry alive after a successful claim so the model coming up is never picked as a victim before its waiters use it, and recount the slots on every status change and whenever a waiter abandons the queue. The model is then evicted as soon as it comes up with nobody left to serve. --------- Co-authored-by: Pascal <admin@serveurperso.com>
This commit is contained in:
co-authored by
Pascal
parent
dbeb37548e
commit
160bd031b2
@@ -297,6 +297,26 @@ def test_router_queue_is_fifo():
|
||||
assert first.done_at < second.done_at, "queue was not served in arrival order"
|
||||
|
||||
|
||||
def test_router_queue_two_waiters_share_one_eviction():
|
||||
"""two requests that both find the same idle model must both be served in the end"""
|
||||
global server
|
||||
server.models_max = 1
|
||||
server.start()
|
||||
|
||||
_load_model_and_wait(MODEL_A, timeout=120)
|
||||
|
||||
# both arrive while MODEL_A is idle, so both want its slot; only one eviction can happen
|
||||
first = _Bg(lambda: _tokenize(MODEL_B)).start()
|
||||
second = _Bg(lambda: _tokenize(MODEL_C)).start()
|
||||
|
||||
first.join(90)
|
||||
second.join(90)
|
||||
|
||||
first.assert_ok("first queued request")
|
||||
second.assert_ok("second queued request")
|
||||
assert _get_model_status(MODEL_A) == "unloaded"
|
||||
|
||||
|
||||
def test_router_no_models_autoload():
|
||||
global server
|
||||
server.no_models_autoload = True
|
||||
|
||||
Reference in New Issue
Block a user