# 中文评判 Prompt 适配 Implementation Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** 让 RAGAS 的 6 个 LLM 评判指标可切换为中文 prompt,通过场景 YAML `judge_language: zh` 与 score API 可选字段 `judge_language` 控制,默认 `en` 完全向后兼容。 **Architecture:** 用 RAGAS 原生 `BasePrompt.adapt("chinese", llm, adapt_instruction=True)` 一次性生成中文 prompt,序列化为提交入库的缓存 JSON。运行时一个共享本地化器把缓存覆盖到指标实例的 prompt 属性上,两条路径(YAML 场景 `factory.build_metric_pipeline` / score API `inline_scorer`)共用它。缺失或漂移的缓存优雅降级回英文,绝不中断评分。 **Tech Stack:** Python 3.12、RAGAS 0.4.3(`ragas.metrics.collections`)、Pydantic v2、pydantic-settings、FastAPI、pytest。 ## Global Constraints - Python 3.12+,PEP 8,4 空格缩进;公共函数加类型注解;每个函数加函数注释,非显然逻辑加行内注释(AGENTS.md)。 - 测试用 pytest,放在 `tests/`,命名 `test_*.py`;**确定性、Mock 外部调用,不依赖在线模型 API**(AGENTS.md)。 - 运行单测命令:`python -m pytest tests/:: -v`(本机 `python` = C:\software\Python312\python.exe,pytest 9.0.3)。 - 提交信息用简短祈使句(如 `Add ...`),每次提交聚焦一个逻辑变更;提交尾部加 `Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>`。 - 需本地化的指标→prompt 属性映射(已核实 RAGAS 0.4.3 源码,权威,勿改): - `faithfulness`: `statement_generator_prompt`, `nli_statement_prompt` - `answer_relevancy`: `prompt` - `context_recall`: `prompt` - `context_precision`: `prompt` - `noise_sensitivity`: 函数式 `to_string()`,无 `instruction`/`examples`,**无法 adapt,跳过** - `factual_correctness`: `prompt`, `nli_prompt` - `semantic_similarity`: 无(跳过) - 语言优先级(两条路径一致):显式值(YAML/请求)> `settings.ragas_judge_language`(默认 `"en"`);仅规范化后等于 `"zh"` 触发本地化,其余按英文。 - 缓存目录:`configs/judge_prompts//__.json`(如 `configs/judge_prompts/zh/faithfulness__nli_statement_prompt.json`)。 - 设计文档:`docs/superpowers/specs/2026-07-01-chinese-judge-prompt-design.md`。 --- ## File Structure **新增:** - `rag_eval/metrics/judge_prompts.py` — 本地化器:映射表、`LocalizationReport`、`prompt_source_hash`、缓存加载(内存 memo)、`apply_localized_prompt`、`localize_pipeline_prompts`、`reset_cache`。 - `scripts/build_judge_prompt_cache.py` — 一次性引导脚本,用 `adapt()` 生成缓存 JSON。 - `configs/judge_prompts/zh/*.json` — 提交入库的中文 prompt 缓存(Task 8 生成)。 - `tests/test_judge_prompt_localizer.py` — 本地化器单测。 - `tests/test_judge_prompt_cache_builder.py` — 引导脚本单测(mock LLM)。 - `tests/test_judge_language_config.py` — settings/Scenario/ScoreRequest 配置单测。 **修改:** - `rag_eval/settings.py` — 新增 `ragas_judge_language`。 - `rag_eval/config/schema.py` — `ScenarioModel.judge_language`。 - `rag_eval/config/loader.py` — 透传 `judge_language`。 - `rag_eval/shared/models.py` — `Scenario.judge_language` 字段。 - `rag_eval/metrics/factory.py` — 新增 `build_metric_registry()`;`build_metric_pipeline()` 接入本地化。 - `webapp/models.py` — `ScoreRequest.judge_language`。 - `webapp/services/inline_scorer.py` — `score()`/`_build_metric_instances()` +参数+本地化,复用 `build_metric_registry`。 - `webapp/api/score.py`、`webapp/services/score_job_manager.py`、`webapp/services/session_score_manager.py` — 透传 `judge_language`。 - `scenarios/siemens_build/siemens-pdf-question-bank-online.yaml`(或现有 siemens 评估场景)+ 一个 offline 示例 — 增加 `judge_language: zh`。 - `README.md` — 简述机制与重新生成缓存的命令。 --- ### Task 1: 配置管道(settings + Scenario 字段) 新增 `judge_language` 配置面,暂不产生行为,只让值能被读取与校验。 **Files:** - Modify: `rag_eval/settings.py:24` - Modify: `rag_eval/config/schema.py:44-59` - Modify: `rag_eval/config/loader.py:48-67` - Modify: `rag_eval/shared/models.py:66-81` - Test: `tests/test_judge_language_config.py` **Interfaces:** - Produces: - `EvaluationSettings().ragas_judge_language: str`(默认 `"en"`,env `RAGAS_JUDGE_LANGUAGE`) - `ScenarioModel.judge_language: Literal["en", "zh"] | None`(默认 `None`) - `Scenario.judge_language: str | None`(默认 `None`) - [ ] **Step 1: 写失败测试** 创建 `tests/test_judge_language_config.py`: ```python """Tests for judge_language plumbing across settings, scenario schema, and loader.""" from pathlib import Path import pytest from rag_eval.settings import EvaluationSettings from rag_eval.config.loader import load_scenario def test_settings_default_judge_language_is_en(): """ragas_judge_language defaults to 'en' when the env var is absent.""" settings = EvaluationSettings(_env_file=None) assert settings.ragas_judge_language == "en" def _write_scenario(tmp_path: Path, extra: str) -> Path: """Write a minimal valid offline scenario YAML plus the given extra line(s).""" dataset = tmp_path / "data.csv" dataset.write_text("sample_id,question,answer,contexts,ground_truth\n", encoding="utf-8") text = ( "scenario_name: t\n" "mode: offline\n" f"dataset: {dataset.name}\n" "judge_model: gpt-5\n" "embedding_model: text-embedding-3-small\n" "metrics: [faithfulness]\n" "output_dir: out\n" f"{extra}" ) path = tmp_path / "s.yaml" path.write_text(text, encoding="utf-8") return path def test_scenario_loads_judge_language_zh(tmp_path): """A scenario may declare judge_language: zh and it lands on the dataclass.""" path = _write_scenario(tmp_path, "judge_language: zh\n") scenario = load_scenario(path) assert scenario.judge_language == "zh" def test_scenario_defaults_judge_language_none(tmp_path): """Omitting judge_language leaves it None so the factory can apply the settings default.""" path = _write_scenario(tmp_path, "") scenario = load_scenario(path) assert scenario.judge_language is None def test_scenario_rejects_invalid_judge_language(tmp_path): """An unsupported judge_language value is rejected at schema validation.""" path = _write_scenario(tmp_path, "judge_language: fr\n") with pytest.raises(Exception): load_scenario(path) ``` - [ ] **Step 2: 运行测试确认失败** Run: `python -m pytest tests/test_judge_language_config.py -v` Expected: FAIL(`AttributeError: ragas_judge_language` 或 Scenario 无 `judge_language`)。 - [ ] **Step 3: 实现配置字段** 在 `rag_eval/settings.py` 第 24 行 `ragas_judge_model` 之后新增: ```python ragas_judge_language: str = Field(default="en", alias="RAGAS_JUDGE_LANGUAGE") ``` 在 `rag_eval/config/schema.py` 的 `ScenarioModel` 中(`embedding_model` 字段之后)新增: ```python judge_language: Literal["en", "zh"] | None = None ``` (`Literal` 已在该文件顶部 `from typing import Any, Literal` 导入,无需改导入。) 在 `rag_eval/config/loader.py` 的 `Scenario(...)` 构造里(`embedding_model=model.embedding_model,` 之后)新增: ```python judge_language=model.judge_language, ``` 在 `rag_eval/shared/models.py` 的 `Scenario` dataclass 中(`doc_weights` 之后,作为带默认值字段)新增: ```python judge_language: str | None = None ``` - [ ] **Step 4: 运行测试确认通过** Run: `python -m pytest tests/test_judge_language_config.py -v` Expected: 4 passed。 - [ ] **Step 5: 提交** ```bash git add rag_eval/settings.py rag_eval/config/schema.py rag_eval/config/loader.py rag_eval/shared/models.py tests/test_judge_language_config.py git commit -m "Add judge_language config plumbing (settings + scenario)" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ### Task 2: ScoreRequest 增加 judge_language 字段 让 score 系列 API 能接收可选 `judge_language`;`SessionScoreRequest` 继承自动获得。 **Files:** - Modify: `webapp/models.py:472-475` - Test: `tests/test_judge_language_config.py`(追加) **Interfaces:** - Produces: `ScoreRequest.judge_language: str | None`(默认 `None`),`SessionScoreRequest` 继承。 - [ ] **Step 1: 写失败测试** 在 `tests/test_judge_language_config.py` 末尾追加: ```python def test_score_request_judge_language_defaults_none(): """ScoreRequest exposes an optional judge_language defaulting to None.""" from webapp.models import ScoreRequest req = ScoreRequest(question="q", answer="a") assert req.judge_language is None req_zh = ScoreRequest(question="q", answer="a", judge_language="zh") assert req_zh.judge_language == "zh" def test_session_score_request_inherits_judge_language(): """SessionScoreRequest inherits the judge_language field from ScoreRequest.""" from webapp.models import SessionScoreRequest req = SessionScoreRequest(session_id="s1", question="q", answer="a", judge_language="zh") assert req.judge_language == "zh" ``` - [ ] **Step 2: 运行测试确认失败** Run: `python -m pytest tests/test_judge_language_config.py -k judge_language -v` Expected: FAIL(`ScoreRequest` 无 `judge_language`)。 - [ ] **Step 3: 实现字段** 在 `webapp/models.py` 的 `ScoreRequest` 中,`embedding_model` 字段(第 472-475 行)之后新增: ```python judge_language: str | None = Field( default=None, description="评判 prompt 语言;'zh' 启用中文评判,为 null 时使用 RAGAS_JUDGE_LANGUAGE(默认 en)。", ) ``` - [ ] **Step 4: 运行测试确认通过** Run: `python -m pytest tests/test_judge_language_config.py -k judge_language -v` Expected: 2 passed(若 `SessionScoreRequest` 必填字段名不同,按其真实必填字段调整测试构造参数后再跑)。 - [ ] **Step 5: 提交** ```bash git add webapp/models.py tests/test_judge_language_config.py git commit -m "Add optional judge_language field to ScoreRequest" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ### Task 3: 指标 registry 复用(DRY 重构) 抽出 `build_metric_registry()`,供 factory、inline_scorer、引导脚本共用,消除三处重复。 **Files:** - Modify: `rag_eval/metrics/factory.py:114-131` - Modify: `webapp/services/inline_scorer.py:34-45` - Test: `tests/test_metric_registry.py` **Interfaces:** - Produces: `rag_eval.metrics.factory.build_metric_registry(llm: Any, embeddings: Any) -> dict[str, Any]`,返回全部 7 个指标实例。 - [ ] **Step 1: 写失败测试** 创建 `tests/test_metric_registry.py`: ```python """Tests for the shared metric registry factory.""" from rag_eval.metrics.factory import build_metric_registry def test_build_metric_registry_has_all_seven_metrics(): """The registry exposes every supported metric keyed by its canonical name.""" registry = build_metric_registry(llm=object(), embeddings=object()) assert set(registry) == { "faithfulness", "answer_relevancy", "context_recall", "context_precision", "noise_sensitivity", "factual_correctness", "semantic_similarity", } ``` - [ ] **Step 2: 运行测试确认失败** Run: `python -m pytest tests/test_metric_registry.py -v` Expected: FAIL(`ImportError: cannot import name 'build_metric_registry'`)。 - [ ] **Step 3: 实现 registry 工厂并复用** 在 `rag_eval/metrics/factory.py` 新增函数(放在 `build_metric_pipeline` 之前): ```python def build_metric_registry(llm: Any, embeddings: Any) -> dict[str, Any]: """Instantiate the full set of supported RAGAS metrics keyed by canonical name. Shared by the scenario pipeline, the inline scorer, and the prompt-cache bootstrap so the metric set is defined in exactly one place. """ return { "faithfulness": Faithfulness(llm=llm), "answer_relevancy": AnswerRelevancy(llm=llm, embeddings=embeddings), "context_recall": ContextRecall(llm=llm), "context_precision": ContextPrecision(llm=llm), "noise_sensitivity": NoiseSensitivity(llm=llm), "factual_correctness": FactualCorrectness(llm=llm), "semantic_similarity": SemanticSimilarity(embeddings=embeddings), } ``` 把 `build_metric_pipeline` 中内联的 `registry = { ... }`(第 114-127 行)替换为: ```python registry = build_metric_registry(llm, embeddings) ``` 在 `webapp/services/inline_scorer.py` 把 `_build_metric_instances`(第 34-45 行)改为复用: ```python def _build_metric_instances(metrics: list[str], llm: Any, embeddings: Any) -> dict[str, Any]: """Instantiate only the RAGAS metric objects requested.""" from rag_eval.metrics.factory import build_metric_registry registry = build_metric_registry(llm, embeddings) return {name: registry[name] for name in metrics if name in registry} ``` 并删除该文件顶部不再使用的 `from ragas.metrics.collections import (...)` 导入块(第 23-31 行)。 - [ ] **Step 4: 运行测试确认通过** Run: `python -m pytest tests/test_metric_registry.py tests/test_score_api.py -v` Expected: registry 测试 passed;既有 score API 测试保持 passed(若个别用例需真实 LLM 则会 skip/无关,重点是不新增失败)。 - [ ] **Step 5: 提交** ```bash git add rag_eval/metrics/factory.py webapp/services/inline_scorer.py tests/test_metric_registry.py git commit -m "Extract shared build_metric_registry factory" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ### Task 4: 本地化器核心 `judge_prompts.py` 纯逻辑模块:加载缓存、覆盖 prompt、漂移检测、内存 memo、优雅降级。用假 prompt 对象测试,不触碰 RAGAS。 **Files:** - Create: `rag_eval/metrics/judge_prompts.py` - Test: `tests/test_judge_prompt_localizer.py` **Interfaces:** - Produces: - `METRIC_PROMPT_ATTRS: dict[str, tuple[str, ...]]` - `CACHE_ROOT: Path` - `LocalizationReport(language, applied, skipped, warnings)`(dataclass) - `prompt_source_hash(prompt) -> str` - `apply_localized_prompt(prompt, data: dict) -> None` - `localize_pipeline_prompts(registry: dict[str, Any], language: str) -> LocalizationReport` - `reset_cache() -> None` - [ ] **Step 1: 写失败测试** 创建 `tests/test_judge_prompt_localizer.py`: ```python """Tests for the judge-prompt localizer (no RAGAS dependency; uses fake prompts).""" import json from pydantic import BaseModel from rag_eval.metrics import judge_prompts as jp class _In(BaseModel): question: str class _Out(BaseModel): statements: list[str] class _FakePrompt: """Minimal stand-in for a RAGAS BasePrompt with overridable attributes.""" def __init__(self): self.input_model = _In self.output_model = _Out self.instruction = "English instruction." self.examples = [(_In(question="q"), _Out(statements=["s"]))] self.language = "english" class _FakeMetric: def __init__(self): self.prompt = _FakePrompt() def _cache_dict(prompt): """Build a valid cache dict for the given fake prompt.""" return { "metric": "context_recall", "prompt_attr": "prompt", "language": "chinese", "ragas_version": "0.4.3", "source_hash": jp.prompt_source_hash(prompt), "instruction": "中文指令。", "examples": [{"input": {"question": "问题"}, "output": {"statements": ["陈述"]}}], } def setup_function(_): jp.reset_cache() def test_apply_localized_prompt_overrides_instruction_and_examples(): """apply_localized_prompt swaps instruction/examples and rebuilds example models.""" prompt = _FakePrompt() jp.apply_localized_prompt(prompt, _cache_dict(prompt)) assert prompt.instruction == "中文指令。" assert prompt.examples[0][0].question == "问题" assert prompt.examples[0][1].statements == ["陈述"] assert prompt.language == "chinese" def test_localize_english_is_noop(): """language='en' leaves the registry untouched.""" metric = _FakeMetric() report = jp.localize_pipeline_prompts({"context_recall": metric}, "en") assert report.applied == [] assert metric.prompt.instruction == "English instruction." def test_localize_applies_from_cache_file(tmp_path, monkeypatch): """localize reads /zh/context_recall__prompt.json and applies it.""" metric = _FakeMetric() root = tmp_path / "configs" / "judge_prompts" (root / "zh").mkdir(parents=True) (root / "zh" / "context_recall__prompt.json").write_text( json.dumps(_cache_dict(metric.prompt), ensure_ascii=False), encoding="utf-8" ) monkeypatch.setattr(jp, "CACHE_ROOT", root) report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh") assert "context_recall.prompt" in report.applied assert metric.prompt.instruction == "中文指令。" def test_localize_missing_cache_keeps_english(tmp_path, monkeypatch): """A missing cache file degrades gracefully to the English prompt with a warning.""" metric = _FakeMetric() monkeypatch.setattr(jp, "CACHE_ROOT", tmp_path / "empty") report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh") assert metric.prompt.instruction == "English instruction." assert "context_recall.prompt" in report.skipped assert report.warnings def test_localize_stale_hash_warns_but_applies(tmp_path, monkeypatch): """A source_hash mismatch still applies Chinese but records a stale warning.""" metric = _FakeMetric() data = _cache_dict(metric.prompt) data["source_hash"] = "deadbeef" root = tmp_path / "configs" / "judge_prompts" (root / "zh").mkdir(parents=True) (root / "zh" / "context_recall__prompt.json").write_text( json.dumps(data, ensure_ascii=False), encoding="utf-8" ) monkeypatch.setattr(jp, "CACHE_ROOT", root) report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh") assert metric.prompt.instruction == "中文指令。" assert any("stale" in w for w in report.warnings) def test_cache_memoized(tmp_path, monkeypatch): """A second localize call does not re-read the file (in-memory memo).""" metric = _FakeMetric() root = tmp_path / "configs" / "judge_prompts" (root / "zh").mkdir(parents=True) path = root / "zh" / "context_recall__prompt.json" path.write_text(json.dumps(_cache_dict(metric.prompt), ensure_ascii=False), encoding="utf-8") monkeypatch.setattr(jp, "CACHE_ROOT", root) jp.localize_pipeline_prompts({"context_recall": _FakeMetric()}, "zh") path.unlink() # delete file; memo should still serve the parsed data metric2 = _FakeMetric() report = jp.localize_pipeline_prompts({"context_recall": metric2}, "zh") assert "context_recall.prompt" in report.applied ``` - [ ] **Step 2: 运行测试确认失败** Run: `python -m pytest tests/test_judge_prompt_localizer.py -v` Expected: FAIL(`ModuleNotFoundError` / 函数缺失)。 - [ ] **Step 3: 实现本地化器** 创建 `rag_eval/metrics/judge_prompts.py`: ```python """Localize RAGAS collections judge prompts to a target language (e.g. Chinese). Loads committed, pre-translated prompt cache files from configs/judge_prompts//__.json and overrides each metric's prompt instance attributes in place. Missing, corrupt, or schema-drifted cache entries degrade gracefully to the built-in English prompt so scoring never breaks. """ from __future__ import annotations import hashlib import json import logging from dataclasses import dataclass, field from pathlib import Path from typing import Any logger = logging.getLogger("rag_eval.metrics.judge_prompts") _REPO_ROOT = Path(__file__).resolve().parents[2] CACHE_ROOT = _REPO_ROOT / "configs" / "judge_prompts" # Metric name -> prompt instance attribute names holding a BasePrompt. # Verified against RAGAS 0.4.3 collections source; semantic_similarity has none. METRIC_PROMPT_ATTRS: dict[str, tuple[str, ...]] = { "faithfulness": ("statement_generator_prompt", "nli_statement_prompt"), "answer_relevancy": ("prompt",), "context_recall": ("prompt",), "context_precision": ("prompt",), "noise_sensitivity": ("statement_prompt", "faithfulness_prompt"), "factual_correctness": ("prompt", "nli_prompt"), } # In-memory memoization of parsed cache files, keyed by (language, metric, attr). _MEMO: dict[tuple[str, str, str], dict | None] = {} @dataclass class LocalizationReport: """Outcome of a localize_pipeline_prompts() call, for logging and tests.""" language: str applied: list[str] = field(default_factory=list) skipped: list[str] = field(default_factory=list) warnings: list[str] = field(default_factory=list) def reset_cache() -> None: """Clear the in-memory parsed-cache memo (used by tests).""" _MEMO.clear() def prompt_source_hash(prompt: Any) -> str: """Return a stable sha256 of a prompt's English instruction + examples.""" examples = [ {"input": inp.model_dump(), "output": out.model_dump()} for inp, out in getattr(prompt, "examples", []) ] payload = json.dumps( {"instruction": getattr(prompt, "instruction", ""), "examples": examples}, ensure_ascii=False, sort_keys=True, ) return hashlib.sha256(payload.encode("utf-8")).hexdigest() def _cache_path(language: str, metric: str, attr: str) -> Path: """Resolve the cache file path for one (language, metric, attr) triple.""" return CACHE_ROOT / language / f"{metric}__{attr}.json" def _load_cache_file(language: str, metric: str, attr: str) -> dict | None: """Load and memoize a cache file; return None if absent or unreadable.""" key = (language, metric, attr) if key in _MEMO: return _MEMO[key] path = _cache_path(language, metric, attr) data: dict | None try: data = json.loads(path.read_text(encoding="utf-8")) except FileNotFoundError: data = None except (OSError, json.JSONDecodeError) as exc: # corrupt file: degrade to English logger.warning("[judge_prompts] cache read failed %s: %s", path, exc) data = None _MEMO[key] = data return data def apply_localized_prompt(prompt: Any, data: dict) -> None: """Override a live prompt's instruction/examples/language from a cache dict. Examples are rebuilt using the live prompt's input/output models, so an upstream schema change raises here and is caught by the caller (which then keeps the English prompt). """ examples = [ (prompt.input_model(**ex["input"]), prompt.output_model(**ex["output"])) for ex in data.get("examples", []) ] prompt.instruction = data["instruction"] prompt.examples = examples prompt.language = data.get("language", "chinese") def localize_pipeline_prompts(registry: dict[str, Any], language: str) -> LocalizationReport: """Override judge prompts in `registry` with cached `language` translations. `registry` maps metric name -> RAGAS metric instance. Only metrics in METRIC_PROMPT_ATTRS are touched; unknown metrics and semantic_similarity are left untouched. English ("en"/"english"/empty) is a no-op. """ report = LocalizationReport(language=language) normalized = (language or "en").strip().lower() if normalized in ("", "en", "english"): return report for metric_name, attrs in METRIC_PROMPT_ATTRS.items(): metric = registry.get(metric_name) if metric is None: continue for attr in attrs: tag = f"{metric_name}.{attr}" prompt = getattr(metric, attr, None) if prompt is None: report.skipped.append(tag) continue data = _load_cache_file(normalized, metric_name, attr) if data is None: report.skipped.append(tag) report.warnings.append(f"missing cache for {tag}") continue # Drift detection: warn if the English source changed since caching. if data.get("source_hash") and data["source_hash"] != prompt_source_hash(prompt): report.warnings.append(f"stale cache for {tag} (regenerate)") try: apply_localized_prompt(prompt, data) report.applied.append(tag) except Exception as exc: # noqa: BLE001 schema drift -> keep English report.warnings.append(f"apply failed for {tag}: {exc}; kept english") if report.warnings: logger.warning( "[judge_prompts] language=%s applied=%d skipped=%d warnings=%s", normalized, len(report.applied), len(report.skipped), report.warnings, ) else: logger.info( "[judge_prompts] language=%s applied=%d", normalized, len(report.applied) ) return report ``` - [ ] **Step 4: 运行测试确认通过** Run: `python -m pytest tests/test_judge_prompt_localizer.py -v` Expected: 6 passed。 - [ ] **Step 5: 提交** ```bash git add rag_eval/metrics/judge_prompts.py tests/test_judge_prompt_localizer.py git commit -m "Add judge-prompt localizer with graceful fallback and drift detection" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ### Task 5: 接入 factory 与 inline_scorer 让两条运行时路径按语言实际调用本地化器。 **Files:** - Modify: `rag_eval/metrics/factory.py:96-131` - Modify: `webapp/services/inline_scorer.py:80-114` - Test: `tests/test_judge_language_wiring.py` **Interfaces:** - Consumes: `build_metric_registry`(Task 3)、`localize_pipeline_prompts`(Task 4)、`Scenario.judge_language`(Task 1)、`settings.ragas_judge_language`(Task 1)。 - Produces: - `build_metric_pipeline(scenario, settings, llm=None, embeddings=None)` 在语言解析为 `zh` 时本地化。 - `InlineScorer.score(..., judge_language: str = "en")`;`_build_metric_instances(metrics, llm, embeddings, judge_language="en")`。 - [ ] **Step 1: 写失败测试** 创建 `tests/test_judge_language_wiring.py`: ```python """Tests that the factory and inline scorer invoke the localizer per language.""" import rag_eval.metrics.factory as factory import webapp.services.inline_scorer as inline_mod def test_build_pipeline_localizes_when_zh(monkeypatch): """build_metric_pipeline calls localize_pipeline_prompts with the resolved language.""" calls = [] monkeypatch.setattr( factory, "localize_pipeline_prompts", lambda registry, language: calls.append(language), ) from rag_eval.shared.models import DatasetConfig, Scenario from rag_eval.settings import EvaluationSettings from pathlib import Path scenario = Scenario( scenario_name="t", mode="offline", dataset=DatasetConfig(path=Path("x.csv")), judge_model="gpt-5", embedding_model="text-embedding-3-small", metrics=["faithfulness"], output_dir=Path("out"), judge_language="zh", ) factory.build_metric_pipeline(scenario, EvaluationSettings(_env_file=None), llm=object(), embeddings=object()) assert calls == ["zh"] def test_build_pipeline_falls_back_to_settings_default(monkeypatch): """When scenario.judge_language is None the settings default is used.""" calls = [] monkeypatch.setattr( factory, "localize_pipeline_prompts", lambda registry, language: calls.append(language), ) from rag_eval.shared.models import DatasetConfig, Scenario from rag_eval.settings import EvaluationSettings from pathlib import Path scenario = Scenario( scenario_name="t", mode="offline", dataset=DatasetConfig(path=Path("x.csv")), judge_model="gpt-5", embedding_model="text-embedding-3-small", metrics=["faithfulness"], output_dir=Path("out"), judge_language=None, ) settings = EvaluationSettings(_env_file=None) settings.ragas_judge_language = "zh" factory.build_metric_pipeline(scenario, settings, llm=object(), embeddings=object()) assert calls == ["zh"] def test_inline_score_threads_judge_language(monkeypatch): """InlineScorer.score forwards judge_language into _build_metric_instances.""" seen = {} monkeypatch.setattr(inline_mod.InlineScorer, "_get_models", lambda self, j, e, s: (object(), object())) monkeypatch.setattr( inline_mod, "_build_metric_instances", lambda metrics, llm, embeddings, judge_language="en": seen.setdefault("lang", judge_language) or {}, ) class _Pipe: def __init__(self, *a, **k): pass monkeypatch.setattr(inline_mod, "MetricPipeline", _Pipe) monkeypatch.setattr(inline_mod.asyncio, "run", lambda coro: type("R", (), {"metrics": {}})()) scorer = inline_mod.InlineScorer() scorer.score(question="q", answer="a", contexts=[], ground_truth=None, metrics=["faithfulness"], judge_model="gpt-5", embedding_model="e", settings=object(), judge_language="zh") assert seen["lang"] == "zh" ``` - [ ] **Step 2: 运行测试确认失败** Run: `python -m pytest tests/test_judge_language_wiring.py -v` Expected: FAIL(factory 未导入/调用 localizer;`score()` 无 `judge_language` 参数)。 - [ ] **Step 3: 实现接入** 在 `rag_eval/metrics/factory.py` 顶部(第 27 行 `from .pipeline import MetricPipeline` 之后)新增导入: ```python from .judge_prompts import localize_pipeline_prompts ``` 把 `build_metric_pipeline` 结尾(构建 registry 之后、`return MetricPipeline(...)` 处)改为先切片再本地化: ```python registry = build_metric_registry(llm, embeddings) selected = {name: registry[name] for name in scenario.metrics} language = (scenario.judge_language or settings.ragas_judge_language or "en") localize_pipeline_prompts(selected, language) return MetricPipeline( metrics=selected, metric_timeout_seconds=settings.ragas_metric_timeout_seconds, ) ``` 在 `webapp/services/inline_scorer.py`:把 `_build_metric_instances` 增加 `judge_language` 形参并本地化: ```python def _build_metric_instances( metrics: list[str], llm: Any, embeddings: Any, judge_language: str = "en" ) -> dict[str, Any]: """Instantiate only the RAGAS metric objects requested, localized if needed.""" from rag_eval.metrics.factory import build_metric_registry from rag_eval.metrics.judge_prompts import localize_pipeline_prompts registry = build_metric_registry(llm, embeddings) selected = {name: registry[name] for name in metrics if name in registry} localize_pipeline_prompts(selected, judge_language) return selected ``` 把 `InlineScorer.score` 签名(第 80-90 行)末尾加参数并在调用处透传: ```python def score( self, question: str, answer: str, contexts: list[str], ground_truth: str | None, metrics: list[str], judge_model: str, embedding_model: str, settings: EvaluationSettings, judge_language: str = "en", ) -> dict[str, float | None]: """Score one sample synchronously and return {metric_name: score | None}.""" llm, embeddings = self._get_models(judge_model, embedding_model, settings) metric_instances = _build_metric_instances(metrics, llm, embeddings, judge_language) ``` (其余函数体不变。) - [ ] **Step 4: 运行测试确认通过** Run: `python -m pytest tests/test_judge_language_wiring.py -v` Expected: 3 passed。 - [ ] **Step 5: 提交** ```bash git add rag_eval/metrics/factory.py webapp/services/inline_scorer.py tests/test_judge_language_wiring.py git commit -m "Wire judge-prompt localization into factory and inline scorer" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ### Task 6: 三个 score 端点透传 judge_language **Files:** - Modify: `webapp/api/score.py:117-142` - Modify: `webapp/services/score_job_manager.py:120-141` - Modify: `webapp/services/session_score_manager.py:205-226` - Test: `tests/test_score_endpoint_language.py` **Interfaces:** - Consumes: `ScoreRequest.judge_language`(Task 2)、`settings.ragas_judge_language`、`InlineScorer.score(..., judge_language=...)`(Task 5)。 - [ ] **Step 1: 写失败测试** 创建 `tests/test_score_endpoint_language.py`: ```python """The /api/score route forwards the resolved judge_language to the scorer.""" from fastapi.testclient import TestClient import webapp.api.score as score_mod from webapp.server import create_app def test_score_route_forwards_judge_language(monkeypatch): """A request with judge_language='zh' reaches inline_scorer.score.""" captured = {} def fake_score(**kwargs): captured.update(kwargs) return {"faithfulness": 0.9} monkeypatch.setattr(score_mod.inline_scorer, "score", lambda **kw: fake_score(**kw)) client = TestClient(create_app()) resp = client.post("/api/score", json={ "question": "q", "answer": "a", "contexts": "c", "ground_truth": "g", "metrics": ["faithfulness"], "judge_language": "zh", }) assert resp.status_code == 200 assert captured.get("judge_language") == "zh" def test_score_route_defaults_language_from_settings(monkeypatch): """Omitting judge_language falls back to settings.ragas_judge_language.""" captured = {} monkeypatch.setattr(score_mod.inline_scorer, "score", lambda **kw: captured.update(kw) or {"faithfulness": 0.9}) client = TestClient(create_app()) resp = client.post("/api/score", json={ "question": "q", "answer": "a", "contexts": "c", "ground_truth": "g", "metrics": ["faithfulness"], }) assert resp.status_code == 200 assert captured.get("judge_language") == "en" ``` (若 `create_app` 的导入路径不同,先 `grep -n "def create_app" webapp/server.py` 校正 import。) - [ ] **Step 2: 运行测试确认失败** Run: `python -m pytest tests/test_score_endpoint_language.py -v` Expected: FAIL(`score()` 未收到 `judge_language`)。 - [ ] **Step 3: 实现透传** `webapp/api/score.py`:在 `judge_model = request.judge_model or settings.ragas_judge_model`(第 117 行)附近新增解析,并在 `inline_scorer.score(...)` 调用(第 133-142 行)加参数: ```python judge_model = request.judge_model or settings.ragas_judge_model embedding_model = request.embedding_model or settings.ragas_embedding_model judge_language = request.judge_language or settings.ragas_judge_language ``` ```python raw_scores = inline_scorer.score( question=request.question, answer=request.answer, contexts=request.contexts_as_list(), ground_truth=request.ground_truth, metrics=effective, judge_model=judge_model, embedding_model=embedding_model, settings=settings, judge_language=judge_language, ) ``` `webapp/services/score_job_manager.py`:在第 121-122 行之后新增,并在第 132-141 行的 `inline_scorer.score(...)` 加同名参数: ```python judge_language = request.judge_language or settings.ragas_judge_language ``` ```python judge_language=judge_language, ``` `webapp/services/session_score_manager.py`:在第 206-207 行之后新增,并在第 217-226 行的 `inline_scorer.score(...)` 加同名参数: ```python judge_language = request.judge_language or settings.ragas_judge_language ``` ```python judge_language=judge_language, ``` - [ ] **Step 4: 运行测试确认通过** Run: `python -m pytest tests/test_score_endpoint_language.py -v` Expected: 2 passed。 - [ ] **Step 5: 提交** ```bash git add webapp/api/score.py webapp/services/score_job_manager.py webapp/services/session_score_manager.py tests/test_score_endpoint_language.py git commit -m "Forward judge_language through score, async, and session endpoints" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ### Task 7: 引导脚本 `build_judge_prompt_cache.py` 用 RAGAS 原生 `adapt()` 生成缓存 JSON。用 mock LLM/prompt 测试,不触发真实网络。 **Files:** - Create: `scripts/build_judge_prompt_cache.py` - Test: `tests/test_judge_prompt_cache_builder.py` **Interfaces:** - Consumes: `build_metric_registry`(Task 3)、`METRIC_PROMPT_ATTRS`/`CACHE_ROOT`/`prompt_source_hash`(Task 4)、`build_models`(既有)。 - Produces: - `serialize_prompt(adapted, source_hash, metric, attr, language, ragas_version) -> dict` - `async build_cache(language, judge_model, embedding_model, settings) -> list[Path]` - [ ] **Step 1: 写失败测试** 创建 `tests/test_judge_prompt_cache_builder.py`: ```python """Tests for the prompt-cache bootstrap script (mocked adapt/LLM).""" import asyncio import json from pydantic import BaseModel import scripts.build_judge_prompt_cache as builder class _In(BaseModel): question: str class _Out(BaseModel): statements: list[str] class _FakeAdapted: def __init__(self): self.instruction = "中文指令。" self.language = "chinese" self.examples = [(_In(question="问题"), _Out(statements=["陈述"]))] class _FakePrompt: def __init__(self): self.input_model = _In self.output_model = _Out self.instruction = "English." self.examples = [(_In(question="q"), _Out(statements=["s"]))] self.language = "english" async def adapt(self, target_language, llm, adapt_instruction=False): assert target_language == "chinese" assert adapt_instruction is True return _FakeAdapted() def test_serialize_prompt_shape(): """serialize_prompt emits the committed cache schema.""" data = builder.serialize_prompt(_FakeAdapted(), "hash123", "context_recall", "prompt", "zh", "0.4.3") assert data["metric"] == "context_recall" assert data["prompt_attr"] == "prompt" assert data["language"] == "chinese" assert data["source_hash"] == "hash123" assert data["instruction"] == "中文指令。" assert data["examples"] == [{"input": {"question": "问题"}, "output": {"statements": ["陈述"]}}] def test_build_cache_writes_all_files(tmp_path, monkeypatch): """build_cache writes one JSON per (metric, attr) into CACHE_ROOT//.""" fake_registry = {name: type("M", (), {})() for name in builder.METRIC_PROMPT_ATTRS} for name, attrs in builder.METRIC_PROMPT_ATTRS.items(): for attr in attrs: setattr(fake_registry[name], attr, _FakePrompt()) monkeypatch.setattr(builder, "build_models", lambda j, e, s: (object(), object())) monkeypatch.setattr(builder, "build_metric_registry", lambda llm, emb: fake_registry) monkeypatch.setattr(builder, "CACHE_ROOT", tmp_path) written = asyncio.run(builder.build_cache("zh", "gpt-5", "emb", object())) expected = sum(len(a) for a in builder.METRIC_PROMPT_ATTRS.values()) assert len(written) == expected sample = tmp_path / "zh" / "context_recall__prompt.json" assert sample.exists() data = json.loads(sample.read_text(encoding="utf-8")) assert data["instruction"] == "中文指令。" ``` - [ ] **Step 2: 运行测试确认失败** Run: `python -m pytest tests/test_judge_prompt_cache_builder.py -v` Expected: FAIL(模块不存在)。 - [ ] **Step 3: 实现脚本** 创建 `scripts/build_judge_prompt_cache.py`: ```python """One-off bootstrap: generate committed Chinese judge-prompt cache via RAGAS adapt(). Usage: python -m scripts.build_judge_prompt_cache --language zh Runs BasePrompt.adapt("chinese", llm, adapt_instruction=True) for every judge prompt of every LLM-scored metric and writes the result to configs/judge_prompts//__.json. Re-run after a RAGAS upgrade to refresh the cache. Requires a working judge LLM. """ from __future__ import annotations import argparse import asyncio import json import logging from pathlib import Path from typing import Any from rag_eval.compat import ensure_ragas_import_compat from rag_eval.metrics.factory import build_metric_registry, build_models from rag_eval.metrics.judge_prompts import CACHE_ROOT, METRIC_PROMPT_ATTRS, prompt_source_hash from rag_eval.settings import EvaluationSettings ensure_ragas_import_compat() logger = logging.getLogger("scripts.build_judge_prompt_cache") # judge_language short code -> RAGAS adapt() natural-language name. _ADAPT_LANGUAGE = {"zh": "chinese"} def serialize_prompt( adapted: Any, source_hash: str, metric: str, attr: str, language: str, ragas_version: str ) -> dict: """Serialize an adapted prompt into the committed cache JSON schema.""" return { "metric": metric, "prompt_attr": attr, "language": getattr(adapted, "language", language), "ragas_version": ragas_version, "source_hash": source_hash, "instruction": adapted.instruction, "examples": [ {"input": inp.model_dump(), "output": out.model_dump()} for inp, out in adapted.examples ], } async def build_cache( language: str, judge_model: str, embedding_model: str, settings: EvaluationSettings ) -> list[Path]: """Generate and write the full prompt cache for one language; return written paths.""" import ragas target = _ADAPT_LANGUAGE.get(language, language) llm, embeddings = build_models(judge_model, embedding_model, settings) registry = build_metric_registry(llm, embeddings) out_dir = CACHE_ROOT / language out_dir.mkdir(parents=True, exist_ok=True) written: list[Path] = [] for metric, attrs in METRIC_PROMPT_ATTRS.items(): for attr in attrs: prompt = getattr(registry[metric], attr) source_hash = prompt_source_hash(prompt) adapted = await prompt.adapt(target, llm, adapt_instruction=True) data = serialize_prompt(adapted, source_hash, metric, attr, language, ragas.__version__) path = out_dir / f"{metric}__{attr}.json" path.write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8") written.append(path) logger.info("wrote %s", path) return written def main() -> None: """CLI entry point: parse args, resolve models, and build the cache.""" parser = argparse.ArgumentParser(description="Build the Chinese judge-prompt cache.") parser.add_argument("--language", default="zh") parser.add_argument("--judge-model", default=None) parser.add_argument("--embedding-model", default=None) args = parser.parse_args() logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s") settings = EvaluationSettings() judge_model = args.judge_model or settings.ragas_judge_model embedding_model = args.embedding_model or settings.ragas_embedding_model paths = asyncio.run(build_cache(args.language, judge_model, embedding_model, settings)) logger.info("done: %d files written", len(paths)) if __name__ == "__main__": main() ``` 同时确保 `scripts/` 可作为包导入:若 `scripts/__init__.py` 不存在则创建空文件(测试用 `import scripts.build_judge_prompt_cache` 需要)。 - [ ] **Step 4: 运行测试确认通过** Run: `python -m pytest tests/test_judge_prompt_cache_builder.py -v` Expected: 2 passed。 - [ ] **Step 5: 提交** ```bash git add scripts/build_judge_prompt_cache.py scripts/__init__.py tests/test_judge_prompt_cache_builder.py git commit -m "Add judge-prompt cache bootstrap script" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ### Task 8: 生成中文缓存 + 示例场景 + 文档 产出真实中文缓存、暴露配置示例、补文档。这是让功能端到端可用的收尾任务。 **Files:** - Create: `configs/judge_prompts/zh/*.json`(10 个文件) - Modify: 一个 siemens 评估场景 YAML + 一个 offline 示例 YAML(增加 `judge_language: zh`) - Modify: `README.md` - [ ] **Step 1: 生成中文缓存** 有可用评判 LLM 时运行(生产做法): ```bash python -m scripts.build_judge_prompt_cache --language zh ``` 预期在 `configs/judge_prompts/zh/` 生成 10 个文件: `faithfulness__statement_generator_prompt.json`、`faithfulness__nli_statement_prompt.json`、`answer_relevancy__prompt.json`、`context_recall__prompt.json`、`context_precision__prompt.json`、`noise_sensitivity__statement_prompt.json`、`noise_sensitivity__faithfulness_prompt.json`、`factual_correctness__prompt.json`、`factual_correctness__nli_prompt.json`(注意 `answer_relevancy` 与部分指标各一个)。 **无 LLM 访问时的兜底**:按 Task 4 缓存 schema 手工撰写等价中文 JSON(`instruction` 译为中文、`examples` 内字符串译为中文、`source_hash` 用当前英文源经 `prompt_source_hash` 计算填入;结构字段名/枚举/数字保持不变)。功能与测试不依赖翻译质量,可后续用脚本刷新。 - [ ] **Step 2: 验证缓存被正确加载** 写一个临时验证(跑完删除或保留为集成测试 `tests/test_zh_cache_integration.py`): ```python """Integration: committed zh cache loads onto a real metric instance.""" from rag_eval.metrics.judge_prompts import CACHE_ROOT, localize_pipeline_prompts def test_zh_cache_files_present(): """All 10 expected zh cache files are committed.""" expected = [ "faithfulness__statement_generator_prompt.json", "faithfulness__nli_statement_prompt.json", "answer_relevancy__prompt.json", "context_recall__prompt.json", "context_precision__prompt.json", "noise_sensitivity__statement_prompt.json", "noise_sensitivity__faithfulness_prompt.json", "factual_correctness__prompt.json", "factual_correctness__nli_prompt.json", ] for name in expected: assert (CACHE_ROOT / "zh" / name).exists(), f"missing {name}" ``` Run: `python -m pytest tests/test_zh_cache_integration.py -v` Expected: PASS(确认 9 个必需文件在位;`answer_relevancy` 只有 1 个 prompt,总计与实际 `METRIC_PROMPT_ATTRS` 展开数一致)。 > 注:`METRIC_PROMPT_ATTRS` 展开后文件总数 = 2+1+1+1+2+2 = 9。若脚本因某指标含多 prompt 而不同,以 `METRIC_PROMPT_ATTRS` 为准同步该测试清单。 - [ ] **Step 3: 场景 YAML 增加 judge_language** 在一个 siemens 评估场景与一个 offline 示例的顶层加入: ```yaml judge_language: zh ``` (用 `grep -rl "mode: offline" scenarios/` 找到目标文件;仅改评估场景,勿改 dataset_build 场景。) - [ ] **Step 4: 补文档** 在 `README.md` 增加一节「中文评判 Prompt 适配」,说明: - 场景 `judge_language: zh` 与 score API `judge_language` 字段、`RAGAS_JUDGE_LANGUAGE` 全局默认。 - 重新生成缓存命令:`python -m scripts.build_judge_prompt_cache --language zh`。 - RAGAS 升级后需重跑脚本(漂移检测会在日志告警)。 - [ ] **Step 5: 运行相关测试并提交** Run: `python -m pytest tests/test_judge_prompt_localizer.py tests/test_zh_cache_integration.py -v` Expected: 全部 passed。 ```bash git add configs/judge_prompts/zh/*.json scenarios/ README.md tests/test_zh_cache_integration.py git commit -m "Add committed zh judge-prompt cache and enable it in sample scenarios" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>" ``` --- ## 最终回归 - [ ] 全量单测:`python -m pytest tests/ -q` - Expected: 本计划新增测试全部通过;不引入新的失败。已知的 6 个历史失败(见设计背景,与本功能无关)保持原状,勿在本计划内处理。 ## Self-Review 记录 - **Spec 覆盖**:§4.4 配置面→Task1/2;§4.3 本地化器→Task4;§4.5 集成→Task5/6;§4.1 引导脚本→Task7;§4.2 缓存文件→Task8;§5 错误处理/漂移→Task4(missing/stale/apply-fail 测试);§6 测试→各 Task;§2 指标映射→Task4 常量;§8 兼容性→默认 `en` no-op(Task4/5 覆盖)。 - **占位符扫描**:无 TBD/TODO;所有步骤含真实代码与命令。 - **类型一致**:`localize_pipeline_prompts(registry, language)`、`build_metric_registry(llm, embeddings)`、`InlineScorer.score(..., judge_language="en")`、`serialize_prompt(...)`、`build_cache(...)` 在定义与调用处签名一致。