Files

274 lines
17 KiB
Markdown
Raw Permalink Normal View History

# System Status — Connected AI Models & Token Usage Design
**Date:** 2026-07-02
**Scope:** Extend the existing System Status module with a new "AI Models" panel showing which LLM/Embedding/Reranker models are configured, their connection status, and cumulative token consumption.
**Relationship to existing roadmap:** This is a lightweight, self-contained first slice of the "P0-A observability" priority already identified in `AI_Agent_优化分析报告_2026-06-18.md` (full Langfuse tracing + Ragas evaluation remains a separate, larger future effort — see Out of Scope).
---
## Goals
1. Show all "connected" AI models in one place: main answer-generation LLM, the dedicated HyDE query-expansion LLM, the embedding model, and the reranker (even when disabled).
2. Show connection status per model, derived passively from real traffic (no extra cost), plus an optional manual "test connection" action for an on-demand active check.
3. Show cumulative token consumption per model since process start (in-memory; resets on restart — no new database table).
4. Guarantee accuracy by instrumenting the single shared LLM client factory, so intermediate Agentic RAG steps, HyDE, regulation-perception analysis, compliance review, and document summarization are all captured — not just the final chat answer.
## Non-Goals (see "Out of Scope" at the end)
- Persistent/historical token usage (DB-backed, survives restart) — deferred.
- Cost/spend estimation in currency — deferred (no reliable public pricing for the internal gateway).
- Per-session or per-user token breakdown — deferred.
- Accurate token counting for **streaming** chat responses — deferred (see Known Limitations).
- Full distributed tracing / LLM-as-judge faithfulness scoring (Langfuse + Ragas, `P0-A` in the existing roadmap) — this feature is a lightweight precursor, not a replacement.
---
## Architecture Overview
### Layering (must not be violated — per `docs/architecture/backend-project-architecture.md`)
```
api/routes/status.py → thin handlers, reads tracker + settings, no business logic
shared/model_usage_tracker.py → cross-cutting support (same tier as shared/bootstrap.py)
services/llm/llm_factory.py → wraps clients with TrackedLLMClient at creation time
infrastructure/embedding/… → direct instrumentation (single implementation)
infrastructure/vectorstore/cross_encoder_reranker.py → direct instrumentation (single implementation)
```
No new business orchestration is added to `services/*` or `workflows/*`. The tracker is passive, cross-cutting infrastructure support, consistent with how `shared/bootstrap.py` and `shared/errors.py` are described in the backend README as "composition root 与横切支撑".
### Data Model
`ModelUsageTracker` keys its internal state by **`f"{provider}:{model}"`**, not by business role. This is more robust than keying by role: if a future Agentic sub-step uses a different provider/model, it is still captured under its own key rather than being silently dropped because no role mapping exists for it. "Role" (`main_llm` / `hyde_llm` / `embedding` / `reranker`) is purely a **presentation-layer label**, resolved at read time in the `/status/models` handler by looking up the current `settings` (`llm_provider`/`llm_model`, `hyde_llm_provider`/`hyde_llm_model` with its existing "empty means reuse main" fallback, `embedding_model`, `reranker_model`).
```python
# backend/app/shared/model_usage_tracker.py
@dataclass
class ModelUsageEntry:
"""Represent accumulated usage/connection state for one provider+model pair."""
provider: str
model: str
total_tokens: int = 0
prompt_tokens: int = 0
completion_tokens: int = 0
call_count_ok: int = 0
call_count_error: int = 0
last_called_at: datetime | None = None
last_latency_ms: int | None = None
last_error: str | None = None
@property
def status(self) -> str:
"""Derive display status from call history: never_called | ok | error.
Note: this only reflects the tracker's own history. The route handler
(not this class) overrides the value to "disabled" for the reranker role
when settings.reranker_enabled is False — config always wins over any
stale historical data, e.g. if the reranker was enabled in the past and
later turned off in .env.
"""
if self.last_called_at is None:
return "never_called"
return "error" if self.last_error else "ok"
class ModelUsageTracker:
"""Thread-safe in-memory registry of per-model call/usage stats.
Never raises: a bug here must not break a real user-facing LLM call.
"""
def __init__(self) -> None:
self._entries: dict[str, ModelUsageEntry] = {}
self._lock = threading.Lock()
def record(
self,
*,
provider: str,
model: str,
success: bool,
usage: dict | None = None,
latency_ms: int | None = None,
error: str | None = None,
) -> None:
"""Record the outcome of one call to provider/model. Safe to call from any thread."""
...
def snapshot(self) -> dict[str, ModelUsageEntry]:
"""Return a shallow copy of all tracked entries, safe to iterate without the lock."""
...
@lru_cache
def get_model_usage_tracker() -> ModelUsageTracker:
"""Return the process-wide singleton tracker (mirrors get_settings()/get_llm_factory() pattern)."""
return ModelUsageTracker()
```
All `record()` bodies are wrapped in `try/except Exception: logger.warning(...)` internally — tracking failures are logged and swallowed, never propagated.
### LLM Instrumentation — `TrackedLLMClient` Wrapper
Every LLM call in the codebase goes through `get_llm_client()` in `backend/app/services/llm/llm_factory.py` (confirmed call sites: `agentic_service.py`, `hyde_expander.py`, `perception/services.py`, `perception/llm_pipeline.py`, `api/routes/compliance.py` ×2, `infrastructure/llm/openai_compatible_answer_generator.py` ×2, `services/llm/document_summarizer.py`). `LLMFactory.create()` wraps the concrete client (`DeepSeekClient`/`QwenClient`/`QwenVLClient`) in `TrackedLLMClient` before caching it, so every current and future call site is covered automatically with **one** change point.
```python
# backend/app/services/llm/tracked_client.py
class TrackedLLMClient:
"""Transparent decorator that records usage/latency into ModelUsageTracker.
Deliberately does NOT subclass BaseLLMClient: that ABC declares abstract
methods (_init_client, get_available_models) which would have to be stubbed
out, defeating the point of __getattr__ delegation and instantiation would
fail with "Can't instantiate abstract class" before __getattr__ ever runs.
Plain composition + __getattr__ forwarding is sufficient since callers only
ever use duck-typed access (.chat(), .stream_chat(), .get_available_models(), .close()).
"""
def __init__(self, inner: BaseLLMClient, tracker: ModelUsageTracker) -> None:
self._inner = inner
self._tracker = tracker
def chat(self, messages, max_tokens=None, temperature=None, tools=None, **kwargs) -> LLMResponse:
"""Delegate to the wrapped client's chat(), then record usage/latency/outcome."""
start = time.time()
response = self._inner.chat(messages, max_tokens, temperature, tools, **kwargs)
self._tracker.record(
provider=self._inner.config.provider.value,
model=response.model or self._inner.config.model,
success=response.is_success,
usage=response.usage,
latency_ms=int((time.time() - start) * 1000),
error=response.error,
)
return response
def stream_chat(self, messages, *args, **kwargs):
"""Delegate to stream_chat(); records call success/latency only (no token usage — see Known Limitations)."""
...
def __getattr__(self, name):
"""Forward any other attribute/method access to the wrapped client."""
return getattr(self._inner, name)
```
### Embedding & Reranker Instrumentation
Both have a single concrete implementation today, so they are instrumented directly (no wrapper needed):
- `OpenAICompatibleEmbeddingProvider._request()` — additionally reads `data.get("usage", {})` from the OpenAI-compatible embeddings response and calls `get_model_usage_tracker().record(provider="embedding", model=self.model, ...)`.
- `OpenAICompatibleReranker._call_reranker()` / `rerank()` — records call success/failure + latency only. TEI/Cohere-style rerank responses do not include token usage, so `total_tokens` for the reranker role will always show as unavailable (`—`), which is factually correct, not a bug to fix later.
---
## API
Both endpoints are added to the existing `backend/app/api/routes/status.py` (no new router file), returning plain dicts — matching the existing convention in this file and in `perception.py` (no Pydantic response models for these "reporting" endpoints).
### `GET /status/models`
Passive read: no outbound network calls, just tracker snapshot + settings resolution.
```json
{
"models": [
{
"role": "main_llm",
"role_label": "主问答 LLM",
"provider": "deepseek",
"model": "deepseek-v4-flash",
"enabled": true,
"status": "ok",
"total_tokens": 12345,
"call_count_ok": 42,
"call_count_error": 1,
"last_called_at": "2026-07-02T10:00:00+08:00",
"last_latency_ms": 350,
"last_error": null,
"shares_usage_with": null
}
]
}
```
Always returns exactly 4 entries in a fixed order: `main_llm`, `hyde_llm`, `embedding`, `reranker` — even if a model has never been called (`status: "never_called"`, all counters zero) or is disabled (`reranker.enabled: false` when `settings.reranker_enabled` is `False`). When `hyde_llm_provider`/`hyde_llm_model` are empty (config falls back to the main LLM), `hyde_llm.shares_usage_with` is set to `"main_llm"` and both rows naturally show identical numbers because they resolve to the same tracker key.
`status` precedence (resolved by the route handler, not by `ModelUsageEntry` itself): if the role is disabled by config (`reranker` only, when `reranker_enabled=False`) the handler always reports `"disabled"`, regardless of any historical call data the tracker may still hold from when it was previously enabled. Otherwise it passes through the tracker's own `ok` / `error` / `never_called`.
### `POST /status/models/ping`
Active check, run only for `enabled` models, in parallel (`asyncio.gather` over `run_in_threadpool`, since the underlying clients are synchronous `httpx`):
- `main_llm` / `hyde_llm`: `chat([{"role": "user", "content": "ping"}], max_tokens=1)`
- `embedding`: `embed_query("ping")`
- `reranker`: `rerank("ping", [one placeholder chunk], top_k=1)` — only when `reranker_enabled=True`
Each ping is wrapped independently so one timeout doesn't block the others. Ping calls go through the same instrumented code paths, so they naturally (and honestly) add a small amount to the token counters — this is not hidden or special-cased. Response shape is identical to `GET /status/models`, reflecting the fresh post-ping state.
---
## Frontend
### New Card: "AI Models" in `frontend/src/pages/Status/StatusPage.tsx`
Placed in `panel-left`, directly after the existing "System Health" card (conceptually related — both are live connectivity views).
- Card header: title + a "Test Connection" button (`POST /status/models/ping`, disabled + spinner while in flight).
- Body: 4 rows reusing the existing `StatusIcon` + `service-row` styling, extended with a right-aligned token count column (monospace, `toLocaleString()`, matching `ConfigRow`'s number formatting) and a small last-called relative-time hint.
- `never_called` and `disabled` map to the existing muted/info badge styles already used elsewhere on this page (no new visual language needed).
- **Cleanup**: the existing "Runtime" card (`panel-right`) currently shows a single Reranker enabled/model line — this is removed from that card since the new "AI Models" card now shows it with richer detail (status + tokens), avoiding duplicate information on the page.
### Data & Types
- `frontend/src/api/status.ts`: add `getModelUsage()` (`GET /status/models`) and `pingModelConnections()` (`POST /status/models/ping`).
- `frontend/src/api/index.ts`: add `ModelUsageEntry` / `ModelUsageResponse` types alongside the existing `SystemStats`/`SystemConfig`/`SystemHealth`.
- `StatusPage.tsx`: extend the existing `Promise.allSettled([...])` fetch-on-mount/refresh with a 4th parallel call for model usage, following the same "partial failure doesn't crash the page" pattern already used for stats/health/config.
- i18n: add new keys under the existing `t.status.*` namespace in both `frontend/src/locales/en.ts` and `zh.ts` (card title, role labels, status labels, button label, "shares usage with main LLM" note).
- Desktop-first, no responsive/mobile work, per `AGENTS.md`.
### Files Changed
| File | Action |
|---|---|
| `backend/app/shared/model_usage_tracker.py` | New — `ModelUsageEntry`, `ModelUsageTracker`, `get_model_usage_tracker()` |
| `backend/app/services/llm/tracked_client.py` | New — `TrackedLLMClient` wrapper |
| `backend/app/services/llm/llm_factory.py` | Wrap client with `TrackedLLMClient` in `LLMFactory.create()` before caching |
| `backend/app/infrastructure/embedding/openai_compatible_embedding_provider.py` | Capture `usage` from embeddings response, record to tracker |
| `backend/app/infrastructure/vectorstore/cross_encoder_reranker.py` | Record call success/failure + latency to tracker |
| `backend/app/api/routes/status.py` | Add `GET /status/models`, `POST /status/models/ping` |
| `frontend/src/api/status.ts` | Add `getModelUsage()`, `pingModelConnections()` |
| `frontend/src/api/index.ts` | Add `ModelUsageEntry`/`ModelUsageResponse` types |
| `frontend/src/pages/Status/StatusPage.tsx` | Add "AI Models" card; remove duplicate reranker line from "Runtime" card |
| `frontend/src/locales/en.ts`, `zh.ts` | Add new `status.*` keys |
---
## Error Handling
- Tracker `record()` never raises — internal `try/except Exception: logger.warning(...)`, so a bug in observability code cannot break a real RAG answer, HyDE expansion, or compliance review call.
- `GET /status/models` mirrors the existing per-service try/except pattern already used in `/status/health` — a failure resolving one role's config falls back to a safe "unknown" entry rather than a 500 for the whole endpoint.
- `POST /status/models/ping`: each per-model ping is wrapped individually (`asyncio.gather(..., return_exceptions=True)` or equivalent per-task try/except); one model's timeout/error does not prevent the other three from completing and being reported.
- Frontend: ping failures surface as inline text on that row (existing `service-row` already supports a muted "detail" slot); page-level fetch failures already degrade gracefully via the existing `Promise.allSettled` fallback pattern.
## Testing
Backend (existing `pytest` setup, `backend/tests/`):
- `backend/tests/shared/test_model_usage_tracker.py` (new) — accumulation across multiple `record()` calls, status transitions (`never_called``ok``error`), basic concurrent-write safety.
- Test for `TrackedLLMClient` — verifies it delegates `chat()` faithfully (return value unchanged) while recording usage, and that a wrapped-client exception still propagates correctly.
- `backend/tests/api/test_status_models_routes.py` (new) — `GET /status/models` returns exactly 4 roles with correct defaults when nothing has been called yet (including `reranker.enabled == settings.reranker_enabled`); `POST /status/models/ping` with mocked clients (no real network calls in tests), verifying partial-failure handling.
Frontend: no test framework exists in this repo today (`frontend/package.json` has no test script, no vitest/jest config) — per project convention, this feature does not introduce one. Verification is `npm --prefix frontend run lint` + `npm --prefix frontend run build`, plus manual visual check of the new card.
## Known Limitations
- **Streaming token gap**: `stream_chat()` implementations in `DeepSeekClient`/`QwenClient` currently only yield content deltas and do not parse a trailing `usage` chunk (would require requesting `stream_options: {include_usage: true}` from the gateway and handling the final SSE chunk). This means token counts from streamed chat (the main RAG chat UI's default interaction mode) are **not** captured in this iteration — only call count/latency/success are recorded for streaming calls. Non-streaming calls (HyDE, agentic intent/plan/grounding steps, compliance review, document summarization, perception analysis) are fully captured. This gap is called out explicitly rather than silently under-counting without explanation, and is a natural follow-up.
- In-memory only: counters reset on every backend restart/redeploy; acceptable per explicit product decision in this design (no new DB table).
## Out of Scope (deferred to future iterations)
- Persistent historical token usage (Postgres-backed, time-windowed charts).
- Cost/spend estimation in currency.
- Per-session/per-user attribution.
- Parsing streaming `usage` chunks for exact streaming token counts.
- Full Langfuse distributed tracing + Ragas/LLM-as-Judge faithfulness scoring (existing roadmap `P0-A` remains the larger follow-on effort; this feature's tracker data model is intentionally simple and would need to coexist with, not replace, a future Langfuse integration).