fix: offload flush to thread pool, add QwenVL stream_options assert, document single-worker assumption
Fix 1 (bootstrap.py): wrap store.flush() in asyncio.to_thread() inside the periodic _flush_loop() to avoid blocking the async event loop every 60s. Synchronous signatures of _start/_stop_model_usage_persistence() and the one-time seed/shutdown flushes are left unchanged per review scope. Fix 2 (test_stream_chat_usage_capture.py): add the two-line stream_options assertion to test_qwen_vl_stream_chat_returns_usage_from_trailing_chunk, matching the identical check already present in the DeepSeek and Qwen tests. Fix 3 (design doc): note the single-worker assumption in A3's write strategy section — multi-worker deployments get last-writer-wins per-row semantics. Tests: 69 passed, 0 failed (python -m pytest backend/tests -q) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
@@ -464,7 +464,7 @@ def _start_model_usage_persistence() -> None:
|
||||
while True:
|
||||
await asyncio.sleep(60)
|
||||
try:
|
||||
store.flush(tracker.snapshot())
|
||||
await asyncio.to_thread(store.flush, tracker.snapshot())
|
||||
except Exception as exc: # noqa: BLE001 - one bad cycle must not kill the loop
|
||||
logger.warning("Failed to flush model usage stats: {}", exc)
|
||||
|
||||
|
||||
@@ -112,3 +112,5 @@ def test_qwen_vl_stream_chat_returns_usage_from_trailing_chunk():
|
||||
|
||||
assert chunks == ["Describing image"]
|
||||
assert returned_usage == usage
|
||||
sent_payload = client._client.stream.call_args.kwargs["json"]
|
||||
assert sent_payload["stream_options"] == {"include_usage": True}
|
||||
|
||||
Reference in New Issue
Block a user