Files
siemens_ragas/docs/superpowers/plans/2026-07-01-chinese-judge-prompt.md

1261 lines
48 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 中文评判 Prompt 适配 Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** 让 RAGAS 的 6 个 LLM 评判指标可切换为中文 prompt,通过场景 YAML `judge_language: zh` 与 score API 可选字段 `judge_language` 控制,默认 `en` 完全向后兼容。
**Architecture:** 用 RAGAS 原生 `BasePrompt.adapt("chinese", llm, adapt_instruction=True)` 一次性生成中文 prompt,序列化为提交入库的缓存 JSON。运行时一个共享本地化器把缓存覆盖到指标实例的 prompt 属性上,两条路径(YAML 场景 `factory.build_metric_pipeline` / score API `inline_scorer`)共用它。缺失或漂移的缓存优雅降级回英文,绝不中断评分。
**Tech Stack:** Python 3.12、RAGAS 0.4.3`ragas.metrics.collections`)、Pydantic v2、pydantic-settings、FastAPI、pytest。
## Global Constraints
- Python 3.12+PEP 8,4 空格缩进;公共函数加类型注解;每个函数加函数注释,非显然逻辑加行内注释(AGENTS.md)。
- 测试用 pytest,放在 `tests/`,命名 `test_*.py`;**确定性、Mock 外部调用,不依赖在线模型 API**AGENTS.md)。
- 运行单测命令:`python -m pytest tests/<file>::<test> -v`(本机 `python` = C:\software\Python312\python.exepytest 9.0.3)。
- 提交信息用简短祈使句(如 `Add ...`),每次提交聚焦一个逻辑变更;提交尾部加 `Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>`
- 需本地化的指标→prompt 属性映射(已核实 RAGAS 0.4.3 源码,权威,勿改):
- `faithfulness`: `statement_generator_prompt`, `nli_statement_prompt`
- `answer_relevancy`: `prompt`
- `context_recall`: `prompt`
- `context_precision`: `prompt`
- `noise_sensitivity`: 函数式 `to_string()`,无 `instruction`/`examples`**无法 adapt,跳过**
- `factual_correctness`: `prompt`, `nli_prompt`
- `semantic_similarity`: 无(跳过)
- 语言优先级(两条路径一致):显式值(YAML/请求)> `settings.ragas_judge_language`(默认 `"en"`);仅规范化后等于 `"zh"` 触发本地化,其余按英文。
- 缓存目录:`configs/judge_prompts/<language>/<metric>__<attr>.json`(如 `configs/judge_prompts/zh/faithfulness__nli_statement_prompt.json`)。
- 设计文档:`docs/superpowers/specs/2026-07-01-chinese-judge-prompt-design.md`
---
## File Structure
**新增:**
- `rag_eval/metrics/judge_prompts.py` — 本地化器:映射表、`LocalizationReport``prompt_source_hash`、缓存加载(内存 memo)、`apply_localized_prompt``localize_pipeline_prompts``reset_cache`
- `scripts/build_judge_prompt_cache.py` — 一次性引导脚本,用 `adapt()` 生成缓存 JSON。
- `configs/judge_prompts/zh/*.json` — 提交入库的中文 prompt 缓存(Task 8 生成)。
- `tests/test_judge_prompt_localizer.py` — 本地化器单测。
- `tests/test_judge_prompt_cache_builder.py` — 引导脚本单测(mock LLM)。
- `tests/test_judge_language_config.py` — settings/Scenario/ScoreRequest 配置单测。
**修改:**
- `rag_eval/settings.py` — 新增 `ragas_judge_language`
- `rag_eval/config/schema.py``ScenarioModel.judge_language`
- `rag_eval/config/loader.py` — 透传 `judge_language`
- `rag_eval/shared/models.py``Scenario.judge_language` 字段。
- `rag_eval/metrics/factory.py` — 新增 `build_metric_registry()``build_metric_pipeline()` 接入本地化。
- `webapp/models.py``ScoreRequest.judge_language`
- `webapp/services/inline_scorer.py``score()`/`_build_metric_instances()` +参数+本地化,复用 `build_metric_registry`
- `webapp/api/score.py``webapp/services/score_job_manager.py``webapp/services/session_score_manager.py` — 透传 `judge_language`
- `scenarios/siemens_build/siemens-pdf-question-bank-online.yaml`(或现有 siemens 评估场景)+ 一个 offline 示例 — 增加 `judge_language: zh`
- `README.md` — 简述机制与重新生成缓存的命令。
---
### Task 1: 配置管道(settings + Scenario 字段)
新增 `judge_language` 配置面,暂不产生行为,只让值能被读取与校验。
**Files:**
- Modify: `rag_eval/settings.py:24`
- Modify: `rag_eval/config/schema.py:44-59`
- Modify: `rag_eval/config/loader.py:48-67`
- Modify: `rag_eval/shared/models.py:66-81`
- Test: `tests/test_judge_language_config.py`
**Interfaces:**
- Produces:
- `EvaluationSettings().ragas_judge_language: str`(默认 `"en"`env `RAGAS_JUDGE_LANGUAGE`
- `ScenarioModel.judge_language: Literal["en", "zh"] | None`(默认 `None`
- `Scenario.judge_language: str | None`(默认 `None`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_language_config.py`
```python
"""Tests for judge_language plumbing across settings, scenario schema, and loader."""
from pathlib import Path
import pytest
from rag_eval.settings import EvaluationSettings
from rag_eval.config.loader import load_scenario
def test_settings_default_judge_language_is_en():
"""ragas_judge_language defaults to 'en' when the env var is absent."""
settings = EvaluationSettings(_env_file=None)
assert settings.ragas_judge_language == "en"
def _write_scenario(tmp_path: Path, extra: str) -> Path:
"""Write a minimal valid offline scenario YAML plus the given extra line(s)."""
dataset = tmp_path / "data.csv"
dataset.write_text("sample_id,question,answer,contexts,ground_truth\n", encoding="utf-8")
text = (
"scenario_name: t\n"
"mode: offline\n"
f"dataset: {dataset.name}\n"
"judge_model: gpt-5\n"
"embedding_model: text-embedding-3-small\n"
"metrics: [faithfulness]\n"
"output_dir: out\n"
f"{extra}"
)
path = tmp_path / "s.yaml"
path.write_text(text, encoding="utf-8")
return path
def test_scenario_loads_judge_language_zh(tmp_path):
"""A scenario may declare judge_language: zh and it lands on the dataclass."""
path = _write_scenario(tmp_path, "judge_language: zh\n")
scenario = load_scenario(path)
assert scenario.judge_language == "zh"
def test_scenario_defaults_judge_language_none(tmp_path):
"""Omitting judge_language leaves it None so the factory can apply the settings default."""
path = _write_scenario(tmp_path, "")
scenario = load_scenario(path)
assert scenario.judge_language is None
def test_scenario_rejects_invalid_judge_language(tmp_path):
"""An unsupported judge_language value is rejected at schema validation."""
path = _write_scenario(tmp_path, "judge_language: fr\n")
with pytest.raises(Exception):
load_scenario(path)
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_language_config.py -v`
Expected: FAIL`AttributeError: ragas_judge_language` 或 Scenario 无 `judge_language`)。
- [ ] **Step 3: 实现配置字段**
`rag_eval/settings.py` 第 24 行 `ragas_judge_model` 之后新增:
```python
ragas_judge_language: str = Field(default="en", alias="RAGAS_JUDGE_LANGUAGE")
```
`rag_eval/config/schema.py``ScenarioModel` 中(`embedding_model` 字段之后)新增:
```python
judge_language: Literal["en", "zh"] | None = None
```
`Literal` 已在该文件顶部 `from typing import Any, Literal` 导入,无需改导入。)
`rag_eval/config/loader.py``Scenario(...)` 构造里(`embedding_model=model.embedding_model,` 之后)新增:
```python
judge_language=model.judge_language,
```
`rag_eval/shared/models.py``Scenario` dataclass 中(`doc_weights` 之后,作为带默认值字段)新增:
```python
judge_language: str | None = None
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_language_config.py -v`
Expected: 4 passed。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/settings.py rag_eval/config/schema.py rag_eval/config/loader.py rag_eval/shared/models.py tests/test_judge_language_config.py
git commit -m "Add judge_language config plumbing (settings + scenario)" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 2: ScoreRequest 增加 judge_language 字段
让 score 系列 API 能接收可选 `judge_language``SessionScoreRequest` 继承自动获得。
**Files:**
- Modify: `webapp/models.py:472-475`
- Test: `tests/test_judge_language_config.py`(追加)
**Interfaces:**
- Produces: `ScoreRequest.judge_language: str | None`(默认 `None`),`SessionScoreRequest` 继承。
- [ ] **Step 1: 写失败测试**
`tests/test_judge_language_config.py` 末尾追加:
```python
def test_score_request_judge_language_defaults_none():
"""ScoreRequest exposes an optional judge_language defaulting to None."""
from webapp.models import ScoreRequest
req = ScoreRequest(question="q", answer="a")
assert req.judge_language is None
req_zh = ScoreRequest(question="q", answer="a", judge_language="zh")
assert req_zh.judge_language == "zh"
def test_session_score_request_inherits_judge_language():
"""SessionScoreRequest inherits the judge_language field from ScoreRequest."""
from webapp.models import SessionScoreRequest
req = SessionScoreRequest(session_id="s1", question="q", answer="a", judge_language="zh")
assert req.judge_language == "zh"
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_language_config.py -k judge_language -v`
Expected: FAIL`ScoreRequest``judge_language`)。
- [ ] **Step 3: 实现字段**
`webapp/models.py``ScoreRequest` 中,`embedding_model` 字段(第 472-475 行)之后新增:
```python
judge_language: str | None = Field(
default=None,
description="评判 prompt 语言;'zh' 启用中文评判,为 null 时使用 RAGAS_JUDGE_LANGUAGE(默认 en)。",
)
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_language_config.py -k judge_language -v`
Expected: 2 passed(若 `SessionScoreRequest` 必填字段名不同,按其真实必填字段调整测试构造参数后再跑)。
- [ ] **Step 5: 提交**
```bash
git add webapp/models.py tests/test_judge_language_config.py
git commit -m "Add optional judge_language field to ScoreRequest" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 3: 指标 registry 复用(DRY 重构)
抽出 `build_metric_registry()`,供 factory、inline_scorer、引导脚本共用,消除三处重复。
**Files:**
- Modify: `rag_eval/metrics/factory.py:114-131`
- Modify: `webapp/services/inline_scorer.py:34-45`
- Test: `tests/test_metric_registry.py`
**Interfaces:**
- Produces: `rag_eval.metrics.factory.build_metric_registry(llm: Any, embeddings: Any) -> dict[str, Any]`,返回全部 7 个指标实例。
- [ ] **Step 1: 写失败测试**
创建 `tests/test_metric_registry.py`
```python
"""Tests for the shared metric registry factory."""
from rag_eval.metrics.factory import build_metric_registry
def test_build_metric_registry_has_all_seven_metrics():
"""The registry exposes every supported metric keyed by its canonical name."""
registry = build_metric_registry(llm=object(), embeddings=object())
assert set(registry) == {
"faithfulness", "answer_relevancy", "context_recall", "context_precision",
"noise_sensitivity", "factual_correctness", "semantic_similarity",
}
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_metric_registry.py -v`
Expected: FAIL`ImportError: cannot import name 'build_metric_registry'`)。
- [ ] **Step 3: 实现 registry 工厂并复用**
`rag_eval/metrics/factory.py` 新增函数(放在 `build_metric_pipeline` 之前):
```python
def build_metric_registry(llm: Any, embeddings: Any) -> dict[str, Any]:
"""Instantiate the full set of supported RAGAS metrics keyed by canonical name.
Shared by the scenario pipeline, the inline scorer, and the prompt-cache
bootstrap so the metric set is defined in exactly one place.
"""
return {
"faithfulness": Faithfulness(llm=llm),
"answer_relevancy": AnswerRelevancy(llm=llm, embeddings=embeddings),
"context_recall": ContextRecall(llm=llm),
"context_precision": ContextPrecision(llm=llm),
"noise_sensitivity": NoiseSensitivity(llm=llm),
"factual_correctness": FactualCorrectness(llm=llm),
"semantic_similarity": SemanticSimilarity(embeddings=embeddings),
}
```
`build_metric_pipeline` 中内联的 `registry = { ... }`(第 114-127 行)替换为:
```python
registry = build_metric_registry(llm, embeddings)
```
`webapp/services/inline_scorer.py``_build_metric_instances`(第 34-45 行)改为复用:
```python
def _build_metric_instances(metrics: list[str], llm: Any, embeddings: Any) -> dict[str, Any]:
"""Instantiate only the RAGAS metric objects requested."""
from rag_eval.metrics.factory import build_metric_registry
registry = build_metric_registry(llm, embeddings)
return {name: registry[name] for name in metrics if name in registry}
```
并删除该文件顶部不再使用的 `from ragas.metrics.collections import (...)` 导入块(第 23-31 行)。
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_metric_registry.py tests/test_score_api.py -v`
Expected: registry 测试 passed;既有 score API 测试保持 passed(若个别用例需真实 LLM 则会 skip/无关,重点是不新增失败)。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/metrics/factory.py webapp/services/inline_scorer.py tests/test_metric_registry.py
git commit -m "Extract shared build_metric_registry factory" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 4: 本地化器核心 `judge_prompts.py`
纯逻辑模块:加载缓存、覆盖 prompt、漂移检测、内存 memo、优雅降级。用假 prompt 对象测试,不触碰 RAGAS。
**Files:**
- Create: `rag_eval/metrics/judge_prompts.py`
- Test: `tests/test_judge_prompt_localizer.py`
**Interfaces:**
- Produces:
- `METRIC_PROMPT_ATTRS: dict[str, tuple[str, ...]]`
- `CACHE_ROOT: Path`
- `LocalizationReport(language, applied, skipped, warnings)`dataclass
- `prompt_source_hash(prompt) -> str`
- `apply_localized_prompt(prompt, data: dict) -> None`
- `localize_pipeline_prompts(registry: dict[str, Any], language: str) -> LocalizationReport`
- `reset_cache() -> None`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_prompt_localizer.py`
```python
"""Tests for the judge-prompt localizer (no RAGAS dependency; uses fake prompts)."""
import json
from pydantic import BaseModel
from rag_eval.metrics import judge_prompts as jp
class _In(BaseModel):
question: str
class _Out(BaseModel):
statements: list[str]
class _FakePrompt:
"""Minimal stand-in for a RAGAS BasePrompt with overridable attributes."""
def __init__(self):
self.input_model = _In
self.output_model = _Out
self.instruction = "English instruction."
self.examples = [(_In(question="q"), _Out(statements=["s"]))]
self.language = "english"
class _FakeMetric:
def __init__(self):
self.prompt = _FakePrompt()
def _cache_dict(prompt):
"""Build a valid cache dict for the given fake prompt."""
return {
"metric": "context_recall",
"prompt_attr": "prompt",
"language": "chinese",
"ragas_version": "0.4.3",
"source_hash": jp.prompt_source_hash(prompt),
"instruction": "中文指令。",
"examples": [{"input": {"question": "问题"}, "output": {"statements": ["陈述"]}}],
}
def setup_function(_):
jp.reset_cache()
def test_apply_localized_prompt_overrides_instruction_and_examples():
"""apply_localized_prompt swaps instruction/examples and rebuilds example models."""
prompt = _FakePrompt()
jp.apply_localized_prompt(prompt, _cache_dict(prompt))
assert prompt.instruction == "中文指令。"
assert prompt.examples[0][0].question == "问题"
assert prompt.examples[0][1].statements == ["陈述"]
assert prompt.language == "chinese"
def test_localize_english_is_noop():
"""language='en' leaves the registry untouched."""
metric = _FakeMetric()
report = jp.localize_pipeline_prompts({"context_recall": metric}, "en")
assert report.applied == []
assert metric.prompt.instruction == "English instruction."
def test_localize_applies_from_cache_file(tmp_path, monkeypatch):
"""localize reads <root>/zh/context_recall__prompt.json and applies it."""
metric = _FakeMetric()
root = tmp_path / "configs" / "judge_prompts"
(root / "zh").mkdir(parents=True)
(root / "zh" / "context_recall__prompt.json").write_text(
json.dumps(_cache_dict(metric.prompt), ensure_ascii=False), encoding="utf-8"
)
monkeypatch.setattr(jp, "CACHE_ROOT", root)
report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh")
assert "context_recall.prompt" in report.applied
assert metric.prompt.instruction == "中文指令。"
def test_localize_missing_cache_keeps_english(tmp_path, monkeypatch):
"""A missing cache file degrades gracefully to the English prompt with a warning."""
metric = _FakeMetric()
monkeypatch.setattr(jp, "CACHE_ROOT", tmp_path / "empty")
report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh")
assert metric.prompt.instruction == "English instruction."
assert "context_recall.prompt" in report.skipped
assert report.warnings
def test_localize_stale_hash_warns_but_applies(tmp_path, monkeypatch):
"""A source_hash mismatch still applies Chinese but records a stale warning."""
metric = _FakeMetric()
data = _cache_dict(metric.prompt)
data["source_hash"] = "deadbeef"
root = tmp_path / "configs" / "judge_prompts"
(root / "zh").mkdir(parents=True)
(root / "zh" / "context_recall__prompt.json").write_text(
json.dumps(data, ensure_ascii=False), encoding="utf-8"
)
monkeypatch.setattr(jp, "CACHE_ROOT", root)
report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh")
assert metric.prompt.instruction == "中文指令。"
assert any("stale" in w for w in report.warnings)
def test_cache_memoized(tmp_path, monkeypatch):
"""A second localize call does not re-read the file (in-memory memo)."""
metric = _FakeMetric()
root = tmp_path / "configs" / "judge_prompts"
(root / "zh").mkdir(parents=True)
path = root / "zh" / "context_recall__prompt.json"
path.write_text(json.dumps(_cache_dict(metric.prompt), ensure_ascii=False), encoding="utf-8")
monkeypatch.setattr(jp, "CACHE_ROOT", root)
jp.localize_pipeline_prompts({"context_recall": _FakeMetric()}, "zh")
path.unlink() # delete file; memo should still serve the parsed data
metric2 = _FakeMetric()
report = jp.localize_pipeline_prompts({"context_recall": metric2}, "zh")
assert "context_recall.prompt" in report.applied
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_prompt_localizer.py -v`
Expected: FAIL`ModuleNotFoundError` / 函数缺失)。
- [ ] **Step 3: 实现本地化器**
创建 `rag_eval/metrics/judge_prompts.py`
```python
"""Localize RAGAS collections judge prompts to a target language (e.g. Chinese).
Loads committed, pre-translated prompt cache files from
configs/judge_prompts/<language>/<metric>__<attr>.json and overrides each
metric's prompt instance attributes in place. Missing, corrupt, or schema-drifted
cache entries degrade gracefully to the built-in English prompt so scoring never
breaks.
"""
from __future__ import annotations
import hashlib
import json
import logging
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
logger = logging.getLogger("rag_eval.metrics.judge_prompts")
_REPO_ROOT = Path(__file__).resolve().parents[2]
CACHE_ROOT = _REPO_ROOT / "configs" / "judge_prompts"
# Metric name -> prompt instance attribute names holding a BasePrompt.
# Verified against RAGAS 0.4.3 collections source; semantic_similarity has none.
METRIC_PROMPT_ATTRS: dict[str, tuple[str, ...]] = {
"faithfulness": ("statement_generator_prompt", "nli_statement_prompt"),
"answer_relevancy": ("prompt",),
"context_recall": ("prompt",),
"context_precision": ("prompt",),
"noise_sensitivity": ("statement_prompt", "faithfulness_prompt"),
"factual_correctness": ("prompt", "nli_prompt"),
}
# In-memory memoization of parsed cache files, keyed by (language, metric, attr).
_MEMO: dict[tuple[str, str, str], dict | None] = {}
@dataclass
class LocalizationReport:
"""Outcome of a localize_pipeline_prompts() call, for logging and tests."""
language: str
applied: list[str] = field(default_factory=list)
skipped: list[str] = field(default_factory=list)
warnings: list[str] = field(default_factory=list)
def reset_cache() -> None:
"""Clear the in-memory parsed-cache memo (used by tests)."""
_MEMO.clear()
def prompt_source_hash(prompt: Any) -> str:
"""Return a stable sha256 of a prompt's English instruction + examples."""
examples = [
{"input": inp.model_dump(), "output": out.model_dump()}
for inp, out in getattr(prompt, "examples", [])
]
payload = json.dumps(
{"instruction": getattr(prompt, "instruction", ""), "examples": examples},
ensure_ascii=False,
sort_keys=True,
)
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
def _cache_path(language: str, metric: str, attr: str) -> Path:
"""Resolve the cache file path for one (language, metric, attr) triple."""
return CACHE_ROOT / language / f"{metric}__{attr}.json"
def _load_cache_file(language: str, metric: str, attr: str) -> dict | None:
"""Load and memoize a cache file; return None if absent or unreadable."""
key = (language, metric, attr)
if key in _MEMO:
return _MEMO[key]
path = _cache_path(language, metric, attr)
data: dict | None
try:
data = json.loads(path.read_text(encoding="utf-8"))
except FileNotFoundError:
data = None
except (OSError, json.JSONDecodeError) as exc: # corrupt file: degrade to English
logger.warning("[judge_prompts] cache read failed %s: %s", path, exc)
data = None
_MEMO[key] = data
return data
def apply_localized_prompt(prompt: Any, data: dict) -> None:
"""Override a live prompt's instruction/examples/language from a cache dict.
Examples are rebuilt using the live prompt's input/output models, so an
upstream schema change raises here and is caught by the caller (which then
keeps the English prompt).
"""
examples = [
(prompt.input_model(**ex["input"]), prompt.output_model(**ex["output"]))
for ex in data.get("examples", [])
]
prompt.instruction = data["instruction"]
prompt.examples = examples
prompt.language = data.get("language", "chinese")
def localize_pipeline_prompts(registry: dict[str, Any], language: str) -> LocalizationReport:
"""Override judge prompts in `registry` with cached `language` translations.
`registry` maps metric name -> RAGAS metric instance. Only metrics in
METRIC_PROMPT_ATTRS are touched; unknown metrics and semantic_similarity are
left untouched. English ("en"/"english"/empty) is a no-op.
"""
report = LocalizationReport(language=language)
normalized = (language or "en").strip().lower()
if normalized in ("", "en", "english"):
return report
for metric_name, attrs in METRIC_PROMPT_ATTRS.items():
metric = registry.get(metric_name)
if metric is None:
continue
for attr in attrs:
tag = f"{metric_name}.{attr}"
prompt = getattr(metric, attr, None)
if prompt is None:
report.skipped.append(tag)
continue
data = _load_cache_file(normalized, metric_name, attr)
if data is None:
report.skipped.append(tag)
report.warnings.append(f"missing cache for {tag}")
continue
# Drift detection: warn if the English source changed since caching.
if data.get("source_hash") and data["source_hash"] != prompt_source_hash(prompt):
report.warnings.append(f"stale cache for {tag} (regenerate)")
try:
apply_localized_prompt(prompt, data)
report.applied.append(tag)
except Exception as exc: # noqa: BLE001 schema drift -> keep English
report.warnings.append(f"apply failed for {tag}: {exc}; kept english")
if report.warnings:
logger.warning(
"[judge_prompts] language=%s applied=%d skipped=%d warnings=%s",
normalized, len(report.applied), len(report.skipped), report.warnings,
)
else:
logger.info(
"[judge_prompts] language=%s applied=%d", normalized, len(report.applied)
)
return report
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_prompt_localizer.py -v`
Expected: 6 passed。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/metrics/judge_prompts.py tests/test_judge_prompt_localizer.py
git commit -m "Add judge-prompt localizer with graceful fallback and drift detection" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 5: 接入 factory 与 inline_scorer
让两条运行时路径按语言实际调用本地化器。
**Files:**
- Modify: `rag_eval/metrics/factory.py:96-131`
- Modify: `webapp/services/inline_scorer.py:80-114`
- Test: `tests/test_judge_language_wiring.py`
**Interfaces:**
- Consumes: `build_metric_registry`Task 3)、`localize_pipeline_prompts`Task 4)、`Scenario.judge_language`Task 1)、`settings.ragas_judge_language`Task 1)。
- Produces:
- `build_metric_pipeline(scenario, settings, llm=None, embeddings=None)` 在语言解析为 `zh` 时本地化。
- `InlineScorer.score(..., judge_language: str = "en")``_build_metric_instances(metrics, llm, embeddings, judge_language="en")`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_language_wiring.py`
```python
"""Tests that the factory and inline scorer invoke the localizer per language."""
import rag_eval.metrics.factory as factory
import webapp.services.inline_scorer as inline_mod
def test_build_pipeline_localizes_when_zh(monkeypatch):
"""build_metric_pipeline calls localize_pipeline_prompts with the resolved language."""
calls = []
monkeypatch.setattr(
factory, "localize_pipeline_prompts",
lambda registry, language: calls.append(language),
)
from rag_eval.shared.models import DatasetConfig, Scenario
from rag_eval.settings import EvaluationSettings
from pathlib import Path
scenario = Scenario(
scenario_name="t", mode="offline",
dataset=DatasetConfig(path=Path("x.csv")),
judge_model="gpt-5", embedding_model="text-embedding-3-small",
metrics=["faithfulness"], output_dir=Path("out"),
judge_language="zh",
)
factory.build_metric_pipeline(scenario, EvaluationSettings(_env_file=None),
llm=object(), embeddings=object())
assert calls == ["zh"]
def test_build_pipeline_falls_back_to_settings_default(monkeypatch):
"""When scenario.judge_language is None the settings default is used."""
calls = []
monkeypatch.setattr(
factory, "localize_pipeline_prompts",
lambda registry, language: calls.append(language),
)
from rag_eval.shared.models import DatasetConfig, Scenario
from rag_eval.settings import EvaluationSettings
from pathlib import Path
scenario = Scenario(
scenario_name="t", mode="offline",
dataset=DatasetConfig(path=Path("x.csv")),
judge_model="gpt-5", embedding_model="text-embedding-3-small",
metrics=["faithfulness"], output_dir=Path("out"),
judge_language=None,
)
settings = EvaluationSettings(_env_file=None)
settings.ragas_judge_language = "zh"
factory.build_metric_pipeline(scenario, settings, llm=object(), embeddings=object())
assert calls == ["zh"]
def test_inline_score_threads_judge_language(monkeypatch):
"""InlineScorer.score forwards judge_language into _build_metric_instances."""
seen = {}
monkeypatch.setattr(inline_mod.InlineScorer, "_get_models",
lambda self, j, e, s: (object(), object()))
monkeypatch.setattr(
inline_mod, "_build_metric_instances",
lambda metrics, llm, embeddings, judge_language="en": seen.setdefault("lang", judge_language) or {},
)
class _Pipe:
def __init__(self, *a, **k):
pass
monkeypatch.setattr(inline_mod, "MetricPipeline", _Pipe)
monkeypatch.setattr(inline_mod.asyncio, "run", lambda coro: type("R", (), {"metrics": {}})())
scorer = inline_mod.InlineScorer()
scorer.score(question="q", answer="a", contexts=[], ground_truth=None,
metrics=["faithfulness"], judge_model="gpt-5",
embedding_model="e", settings=object(), judge_language="zh")
assert seen["lang"] == "zh"
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_language_wiring.py -v`
Expected: FAILfactory 未导入/调用 localizer`score()``judge_language` 参数)。
- [ ] **Step 3: 实现接入**
`rag_eval/metrics/factory.py` 顶部(第 27 行 `from .pipeline import MetricPipeline` 之后)新增导入:
```python
from .judge_prompts import localize_pipeline_prompts
```
`build_metric_pipeline` 结尾(构建 registry 之后、`return MetricPipeline(...)` 处)改为先切片再本地化:
```python
registry = build_metric_registry(llm, embeddings)
selected = {name: registry[name] for name in scenario.metrics}
language = (scenario.judge_language or settings.ragas_judge_language or "en")
localize_pipeline_prompts(selected, language)
return MetricPipeline(
metrics=selected,
metric_timeout_seconds=settings.ragas_metric_timeout_seconds,
)
```
`webapp/services/inline_scorer.py`:把 `_build_metric_instances` 增加 `judge_language` 形参并本地化:
```python
def _build_metric_instances(
metrics: list[str], llm: Any, embeddings: Any, judge_language: str = "en"
) -> dict[str, Any]:
"""Instantiate only the RAGAS metric objects requested, localized if needed."""
from rag_eval.metrics.factory import build_metric_registry
from rag_eval.metrics.judge_prompts import localize_pipeline_prompts
registry = build_metric_registry(llm, embeddings)
selected = {name: registry[name] for name in metrics if name in registry}
localize_pipeline_prompts(selected, judge_language)
return selected
```
`InlineScorer.score` 签名(第 80-90 行)末尾加参数并在调用处透传:
```python
def score(
self,
question: str,
answer: str,
contexts: list[str],
ground_truth: str | None,
metrics: list[str],
judge_model: str,
embedding_model: str,
settings: EvaluationSettings,
judge_language: str = "en",
) -> dict[str, float | None]:
"""Score one sample synchronously and return {metric_name: score | None}."""
llm, embeddings = self._get_models(judge_model, embedding_model, settings)
metric_instances = _build_metric_instances(metrics, llm, embeddings, judge_language)
```
(其余函数体不变。)
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_language_wiring.py -v`
Expected: 3 passed。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/metrics/factory.py webapp/services/inline_scorer.py tests/test_judge_language_wiring.py
git commit -m "Wire judge-prompt localization into factory and inline scorer" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 6: 三个 score 端点透传 judge_language
**Files:**
- Modify: `webapp/api/score.py:117-142`
- Modify: `webapp/services/score_job_manager.py:120-141`
- Modify: `webapp/services/session_score_manager.py:205-226`
- Test: `tests/test_score_endpoint_language.py`
**Interfaces:**
- Consumes: `ScoreRequest.judge_language`Task 2)、`settings.ragas_judge_language``InlineScorer.score(..., judge_language=...)`Task 5)。
- [ ] **Step 1: 写失败测试**
创建 `tests/test_score_endpoint_language.py`
```python
"""The /api/score route forwards the resolved judge_language to the scorer."""
from fastapi.testclient import TestClient
import webapp.api.score as score_mod
from webapp.server import create_app
def test_score_route_forwards_judge_language(monkeypatch):
"""A request with judge_language='zh' reaches inline_scorer.score."""
captured = {}
def fake_score(**kwargs):
captured.update(kwargs)
return {"faithfulness": 0.9}
monkeypatch.setattr(score_mod.inline_scorer, "score", lambda **kw: fake_score(**kw))
client = TestClient(create_app())
resp = client.post("/api/score", json={
"question": "q", "answer": "a", "contexts": "c",
"ground_truth": "g", "metrics": ["faithfulness"], "judge_language": "zh",
})
assert resp.status_code == 200
assert captured.get("judge_language") == "zh"
def test_score_route_defaults_language_from_settings(monkeypatch):
"""Omitting judge_language falls back to settings.ragas_judge_language."""
captured = {}
monkeypatch.setattr(score_mod.inline_scorer, "score",
lambda **kw: captured.update(kw) or {"faithfulness": 0.9})
client = TestClient(create_app())
resp = client.post("/api/score", json={
"question": "q", "answer": "a", "contexts": "c",
"ground_truth": "g", "metrics": ["faithfulness"],
})
assert resp.status_code == 200
assert captured.get("judge_language") == "en"
```
(若 `create_app` 的导入路径不同,先 `grep -n "def create_app" webapp/server.py` 校正 import。)
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_score_endpoint_language.py -v`
Expected: FAIL`score()` 未收到 `judge_language`)。
- [ ] **Step 3: 实现透传**
`webapp/api/score.py`:在 `judge_model = request.judge_model or settings.ragas_judge_model`(第 117 行)附近新增解析,并在 `inline_scorer.score(...)` 调用(第 133-142 行)加参数:
```python
judge_model = request.judge_model or settings.ragas_judge_model
embedding_model = request.embedding_model or settings.ragas_embedding_model
judge_language = request.judge_language or settings.ragas_judge_language
```
```python
raw_scores = inline_scorer.score(
question=request.question,
answer=request.answer,
contexts=request.contexts_as_list(),
ground_truth=request.ground_truth,
metrics=effective,
judge_model=judge_model,
embedding_model=embedding_model,
settings=settings,
judge_language=judge_language,
)
```
`webapp/services/score_job_manager.py`:在第 121-122 行之后新增,并在第 132-141 行的 `inline_scorer.score(...)` 加同名参数:
```python
judge_language = request.judge_language or settings.ragas_judge_language
```
```python
judge_language=judge_language,
```
`webapp/services/session_score_manager.py`:在第 206-207 行之后新增,并在第 217-226 行的 `inline_scorer.score(...)` 加同名参数:
```python
judge_language = request.judge_language or settings.ragas_judge_language
```
```python
judge_language=judge_language,
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_score_endpoint_language.py -v`
Expected: 2 passed。
- [ ] **Step 5: 提交**
```bash
git add webapp/api/score.py webapp/services/score_job_manager.py webapp/services/session_score_manager.py tests/test_score_endpoint_language.py
git commit -m "Forward judge_language through score, async, and session endpoints" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 7: 引导脚本 `build_judge_prompt_cache.py`
用 RAGAS 原生 `adapt()` 生成缓存 JSON。用 mock LLM/prompt 测试,不触发真实网络。
**Files:**
- Create: `scripts/build_judge_prompt_cache.py`
- Test: `tests/test_judge_prompt_cache_builder.py`
**Interfaces:**
- Consumes: `build_metric_registry`Task 3)、`METRIC_PROMPT_ATTRS`/`CACHE_ROOT`/`prompt_source_hash`Task 4)、`build_models`(既有)。
- Produces:
- `serialize_prompt(adapted, source_hash, metric, attr, language, ragas_version) -> dict`
- `async build_cache(language, judge_model, embedding_model, settings) -> list[Path]`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_prompt_cache_builder.py`
```python
"""Tests for the prompt-cache bootstrap script (mocked adapt/LLM)."""
import asyncio
import json
from pydantic import BaseModel
import scripts.build_judge_prompt_cache as builder
class _In(BaseModel):
question: str
class _Out(BaseModel):
statements: list[str]
class _FakeAdapted:
def __init__(self):
self.instruction = "中文指令。"
self.language = "chinese"
self.examples = [(_In(question="问题"), _Out(statements=["陈述"]))]
class _FakePrompt:
def __init__(self):
self.input_model = _In
self.output_model = _Out
self.instruction = "English."
self.examples = [(_In(question="q"), _Out(statements=["s"]))]
self.language = "english"
async def adapt(self, target_language, llm, adapt_instruction=False):
assert target_language == "chinese"
assert adapt_instruction is True
return _FakeAdapted()
def test_serialize_prompt_shape():
"""serialize_prompt emits the committed cache schema."""
data = builder.serialize_prompt(_FakeAdapted(), "hash123", "context_recall", "prompt", "zh", "0.4.3")
assert data["metric"] == "context_recall"
assert data["prompt_attr"] == "prompt"
assert data["language"] == "chinese"
assert data["source_hash"] == "hash123"
assert data["instruction"] == "中文指令。"
assert data["examples"] == [{"input": {"question": "问题"}, "output": {"statements": ["陈述"]}}]
def test_build_cache_writes_all_files(tmp_path, monkeypatch):
"""build_cache writes one JSON per (metric, attr) into CACHE_ROOT/<language>/."""
fake_registry = {name: type("M", (), {})() for name in builder.METRIC_PROMPT_ATTRS}
for name, attrs in builder.METRIC_PROMPT_ATTRS.items():
for attr in attrs:
setattr(fake_registry[name], attr, _FakePrompt())
monkeypatch.setattr(builder, "build_models", lambda j, e, s: (object(), object()))
monkeypatch.setattr(builder, "build_metric_registry", lambda llm, emb: fake_registry)
monkeypatch.setattr(builder, "CACHE_ROOT", tmp_path)
written = asyncio.run(builder.build_cache("zh", "gpt-5", "emb", object()))
expected = sum(len(a) for a in builder.METRIC_PROMPT_ATTRS.values())
assert len(written) == expected
sample = tmp_path / "zh" / "context_recall__prompt.json"
assert sample.exists()
data = json.loads(sample.read_text(encoding="utf-8"))
assert data["instruction"] == "中文指令。"
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_prompt_cache_builder.py -v`
Expected: FAIL(模块不存在)。
- [ ] **Step 3: 实现脚本**
创建 `scripts/build_judge_prompt_cache.py`
```python
"""One-off bootstrap: generate committed Chinese judge-prompt cache via RAGAS adapt().
Usage:
python -m scripts.build_judge_prompt_cache --language zh
Runs BasePrompt.adapt("chinese", llm, adapt_instruction=True) for every judge
prompt of every LLM-scored metric and writes the result to
configs/judge_prompts/<language>/<metric>__<attr>.json. Re-run after a RAGAS
upgrade to refresh the cache. Requires a working judge LLM.
"""
from __future__ import annotations
import argparse
import asyncio
import json
import logging
from pathlib import Path
from typing import Any
from rag_eval.compat import ensure_ragas_import_compat
from rag_eval.metrics.factory import build_metric_registry, build_models
from rag_eval.metrics.judge_prompts import CACHE_ROOT, METRIC_PROMPT_ATTRS, prompt_source_hash
from rag_eval.settings import EvaluationSettings
ensure_ragas_import_compat()
logger = logging.getLogger("scripts.build_judge_prompt_cache")
# judge_language short code -> RAGAS adapt() natural-language name.
_ADAPT_LANGUAGE = {"zh": "chinese"}
def serialize_prompt(
adapted: Any, source_hash: str, metric: str, attr: str, language: str, ragas_version: str
) -> dict:
"""Serialize an adapted prompt into the committed cache JSON schema."""
return {
"metric": metric,
"prompt_attr": attr,
"language": getattr(adapted, "language", language),
"ragas_version": ragas_version,
"source_hash": source_hash,
"instruction": adapted.instruction,
"examples": [
{"input": inp.model_dump(), "output": out.model_dump()}
for inp, out in adapted.examples
],
}
async def build_cache(
language: str, judge_model: str, embedding_model: str, settings: EvaluationSettings
) -> list[Path]:
"""Generate and write the full prompt cache for one language; return written paths."""
import ragas
target = _ADAPT_LANGUAGE.get(language, language)
llm, embeddings = build_models(judge_model, embedding_model, settings)
registry = build_metric_registry(llm, embeddings)
out_dir = CACHE_ROOT / language
out_dir.mkdir(parents=True, exist_ok=True)
written: list[Path] = []
for metric, attrs in METRIC_PROMPT_ATTRS.items():
for attr in attrs:
prompt = getattr(registry[metric], attr)
source_hash = prompt_source_hash(prompt)
adapted = await prompt.adapt(target, llm, adapt_instruction=True)
data = serialize_prompt(adapted, source_hash, metric, attr, language, ragas.__version__)
path = out_dir / f"{metric}__{attr}.json"
path.write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")
written.append(path)
logger.info("wrote %s", path)
return written
def main() -> None:
"""CLI entry point: parse args, resolve models, and build the cache."""
parser = argparse.ArgumentParser(description="Build the Chinese judge-prompt cache.")
parser.add_argument("--language", default="zh")
parser.add_argument("--judge-model", default=None)
parser.add_argument("--embedding-model", default=None)
args = parser.parse_args()
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
settings = EvaluationSettings()
judge_model = args.judge_model or settings.ragas_judge_model
embedding_model = args.embedding_model or settings.ragas_embedding_model
paths = asyncio.run(build_cache(args.language, judge_model, embedding_model, settings))
logger.info("done: %d files written", len(paths))
if __name__ == "__main__":
main()
```
同时确保 `scripts/` 可作为包导入:若 `scripts/__init__.py` 不存在则创建空文件(测试用 `import scripts.build_judge_prompt_cache` 需要)。
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_prompt_cache_builder.py -v`
Expected: 2 passed。
- [ ] **Step 5: 提交**
```bash
git add scripts/build_judge_prompt_cache.py scripts/__init__.py tests/test_judge_prompt_cache_builder.py
git commit -m "Add judge-prompt cache bootstrap script" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 8: 生成中文缓存 + 示例场景 + 文档
产出真实中文缓存、暴露配置示例、补文档。这是让功能端到端可用的收尾任务。
**Files:**
- Create: `configs/judge_prompts/zh/*.json`10 个文件)
- Modify: 一个 siemens 评估场景 YAML + 一个 offline 示例 YAML(增加 `judge_language: zh`
- Modify: `README.md`
- [ ] **Step 1: 生成中文缓存**
有可用评判 LLM 时运行(生产做法):
```bash
python -m scripts.build_judge_prompt_cache --language zh
```
预期在 `configs/judge_prompts/zh/` 生成 10 个文件:
`faithfulness__statement_generator_prompt.json``faithfulness__nli_statement_prompt.json``answer_relevancy__prompt.json``context_recall__prompt.json``context_precision__prompt.json``noise_sensitivity__statement_prompt.json``noise_sensitivity__faithfulness_prompt.json``factual_correctness__prompt.json``factual_correctness__nli_prompt.json`(注意 `answer_relevancy` 与部分指标各一个)。
**无 LLM 访问时的兜底**:按 Task 4 缓存 schema 手工撰写等价中文 JSON`instruction` 译为中文、`examples` 内字符串译为中文、`source_hash` 用当前英文源经 `prompt_source_hash` 计算填入;结构字段名/枚举/数字保持不变)。功能与测试不依赖翻译质量,可后续用脚本刷新。
- [ ] **Step 2: 验证缓存被正确加载**
写一个临时验证(跑完删除或保留为集成测试 `tests/test_zh_cache_integration.py`):
```python
"""Integration: committed zh cache loads onto a real metric instance."""
from rag_eval.metrics.judge_prompts import CACHE_ROOT, localize_pipeline_prompts
def test_zh_cache_files_present():
"""All 10 expected zh cache files are committed."""
expected = [
"faithfulness__statement_generator_prompt.json",
"faithfulness__nli_statement_prompt.json",
"answer_relevancy__prompt.json",
"context_recall__prompt.json",
"context_precision__prompt.json",
"noise_sensitivity__statement_prompt.json",
"noise_sensitivity__faithfulness_prompt.json",
"factual_correctness__prompt.json",
"factual_correctness__nli_prompt.json",
]
for name in expected:
assert (CACHE_ROOT / "zh" / name).exists(), f"missing {name}"
```
Run: `python -m pytest tests/test_zh_cache_integration.py -v`
Expected: PASS(确认 9 个必需文件在位;`answer_relevancy` 只有 1 个 prompt,总计与实际 `METRIC_PROMPT_ATTRS` 展开数一致)。
> 注:`METRIC_PROMPT_ATTRS` 展开后文件总数 = 2+1+1+1+2+2 = 9。若脚本因某指标含多 prompt 而不同,以 `METRIC_PROMPT_ATTRS` 为准同步该测试清单。
- [ ] **Step 3: 场景 YAML 增加 judge_language**
在一个 siemens 评估场景与一个 offline 示例的顶层加入:
```yaml
judge_language: zh
```
(用 `grep -rl "mode: offline" scenarios/` 找到目标文件;仅改评估场景,勿改 dataset_build 场景。)
- [ ] **Step 4: 补文档**
`README.md` 增加一节「中文评判 Prompt 适配」,说明:
- 场景 `judge_language: zh` 与 score API `judge_language` 字段、`RAGAS_JUDGE_LANGUAGE` 全局默认。
- 重新生成缓存命令:`python -m scripts.build_judge_prompt_cache --language zh`
- RAGAS 升级后需重跑脚本(漂移检测会在日志告警)。
- [ ] **Step 5: 运行相关测试并提交**
Run: `python -m pytest tests/test_judge_prompt_localizer.py tests/test_zh_cache_integration.py -v`
Expected: 全部 passed。
```bash
git add configs/judge_prompts/zh/*.json scenarios/ README.md tests/test_zh_cache_integration.py
git commit -m "Add committed zh judge-prompt cache and enable it in sample scenarios" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
## 最终回归
- [ ] 全量单测:`python -m pytest tests/ -q`
- Expected: 本计划新增测试全部通过;不引入新的失败。已知的 6 个历史失败(见设计背景,与本功能无关)保持原状,勿在本计划内处理。
## Self-Review 记录
- **Spec 覆盖**:§4.4 配置面→Task1/2;§4.3 本地化器→Task4;§4.5 集成→Task5/6;§4.1 引导脚本→Task7;§4.2 缓存文件→Task8;§5 错误处理/漂移→Task4missing/stale/apply-fail 测试);§6 测试→各 Task;§2 指标映射→Task4 常量;§8 兼容性→默认 `en` no-opTask4/5 覆盖)。
- **占位符扫描**:无 TBD/TODO;所有步骤含真实代码与命令。
- **类型一致**`localize_pipeline_prompts(registry, language)``build_metric_registry(llm, embeddings)``InlineScorer.score(..., judge_language="en")``serialize_prompt(...)``build_cache(...)` 在定义与调用处签名一致。