fix: offload flush to thread pool, add QwenVL stream_options assert, document single-worker assumption

Fix 1 (bootstrap.py): wrap store.flush() in asyncio.to_thread() inside the
periodic _flush_loop() to avoid blocking the async event loop every 60s.
Synchronous signatures of _start/_stop_model_usage_persistence() and the
one-time seed/shutdown flushes are left unchanged per review scope.

Fix 2 (test_stream_chat_usage_capture.py): add the two-line stream_options
assertion to test_qwen_vl_stream_chat_returns_usage_from_trailing_chunk,
matching the identical check already present in the DeepSeek and Qwen tests.

Fix 3 (design doc): note the single-worker assumption in A3's write strategy
section — multi-worker deployments get last-writer-wins per-row semantics.

Tests: 69 passed, 0 failed (python -m pytest backend/tests -q)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
wangwei
2026-07-23 15:30:21 +08:00
co-authored by Copilot
parent 29f79d7434
commit 0907470a2d
3 changed files with 142 additions and 1 deletions
+1 -1
View File
@@ -464,7 +464,7 @@ def _start_model_usage_persistence() -> None:
while True:
await asyncio.sleep(60)
try:
store.flush(tracker.snapshot())
await asyncio.to_thread(store.flush, tracker.snapshot())
except Exception as exc: # noqa: BLE001 - one bad cycle must not kill the loop
logger.warning("Failed to flush model usage stats: {}", exc)