Files

1261 lines
48 KiB
Markdown
Raw Permalink Normal View History

# 中文评判 Prompt 适配 Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** 让 RAGAS 的 6 个 LLM 评判指标可切换为中文 prompt,通过场景 YAML `judge_language: zh` 与 score API 可选字段 `judge_language` 控制,默认 `en` 完全向后兼容。
**Architecture:** 用 RAGAS 原生 `BasePrompt.adapt("chinese", llm, adapt_instruction=True)` 一次性生成中文 prompt,序列化为提交入库的缓存 JSON。运行时一个共享本地化器把缓存覆盖到指标实例的 prompt 属性上,两条路径(YAML 场景 `factory.build_metric_pipeline` / score API `inline_scorer`)共用它。缺失或漂移的缓存优雅降级回英文,绝不中断评分。
**Tech Stack:** Python 3.12、RAGAS 0.4.3`ragas.metrics.collections`)、Pydantic v2、pydantic-settings、FastAPI、pytest。
## Global Constraints
- Python 3.12+PEP 8,4 空格缩进;公共函数加类型注解;每个函数加函数注释,非显然逻辑加行内注释(AGENTS.md)。
- 测试用 pytest,放在 `tests/`,命名 `test_*.py`;**确定性、Mock 外部调用,不依赖在线模型 API**AGENTS.md)。
- 运行单测命令:`python -m pytest tests/<file>::<test> -v`(本机 `python` = C:\software\Python312\python.exepytest 9.0.3)。
- 提交信息用简短祈使句(如 `Add ...`),每次提交聚焦一个逻辑变更;提交尾部加 `Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>`
- 需本地化的指标→prompt 属性映射(已核实 RAGAS 0.4.3 源码,权威,勿改):
- `faithfulness`: `statement_generator_prompt`, `nli_statement_prompt`
- `answer_relevancy`: `prompt`
- `context_recall`: `prompt`
- `context_precision`: `prompt`
- `noise_sensitivity`: 函数式 `to_string()`,无 `instruction`/`examples`**无法 adapt,跳过**
- `factual_correctness`: `prompt`, `nli_prompt`
- `semantic_similarity`: 无(跳过)
- 语言优先级(两条路径一致):显式值(YAML/请求)> `settings.ragas_judge_language`(默认 `"en"`);仅规范化后等于 `"zh"` 触发本地化,其余按英文。
- 缓存目录:`configs/judge_prompts/<language>/<metric>__<attr>.json`(如 `configs/judge_prompts/zh/faithfulness__nli_statement_prompt.json`)。
- 设计文档:`docs/superpowers/specs/2026-07-01-chinese-judge-prompt-design.md`
---
## File Structure
**新增:**
- `rag_eval/metrics/judge_prompts.py` — 本地化器:映射表、`LocalizationReport``prompt_source_hash`、缓存加载(内存 memo)、`apply_localized_prompt``localize_pipeline_prompts``reset_cache`
- `scripts/build_judge_prompt_cache.py` — 一次性引导脚本,用 `adapt()` 生成缓存 JSON。
- `configs/judge_prompts/zh/*.json` — 提交入库的中文 prompt 缓存(Task 8 生成)。
- `tests/test_judge_prompt_localizer.py` — 本地化器单测。
- `tests/test_judge_prompt_cache_builder.py` — 引导脚本单测(mock LLM)。
- `tests/test_judge_language_config.py` — settings/Scenario/ScoreRequest 配置单测。
**修改:**
- `rag_eval/settings.py` — 新增 `ragas_judge_language`
- `rag_eval/config/schema.py``ScenarioModel.judge_language`
- `rag_eval/config/loader.py` — 透传 `judge_language`
- `rag_eval/shared/models.py``Scenario.judge_language` 字段。
- `rag_eval/metrics/factory.py` — 新增 `build_metric_registry()``build_metric_pipeline()` 接入本地化。
- `webapp/models.py``ScoreRequest.judge_language`
- `webapp/services/inline_scorer.py``score()`/`_build_metric_instances()` +参数+本地化,复用 `build_metric_registry`
- `webapp/api/score.py``webapp/services/score_job_manager.py``webapp/services/session_score_manager.py` — 透传 `judge_language`
- `scenarios/siemens_build/siemens-pdf-question-bank-online.yaml`(或现有 siemens 评估场景)+ 一个 offline 示例 — 增加 `judge_language: zh`
- `README.md` — 简述机制与重新生成缓存的命令。
---
### Task 1: 配置管道(settings + Scenario 字段)
新增 `judge_language` 配置面,暂不产生行为,只让值能被读取与校验。
**Files:**
- Modify: `rag_eval/settings.py:24`
- Modify: `rag_eval/config/schema.py:44-59`
- Modify: `rag_eval/config/loader.py:48-67`
- Modify: `rag_eval/shared/models.py:66-81`
- Test: `tests/test_judge_language_config.py`
**Interfaces:**
- Produces:
- `EvaluationSettings().ragas_judge_language: str`(默认 `"en"`env `RAGAS_JUDGE_LANGUAGE`
- `ScenarioModel.judge_language: Literal["en", "zh"] | None`(默认 `None`
- `Scenario.judge_language: str | None`(默认 `None`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_language_config.py`
```python
"""Tests for judge_language plumbing across settings, scenario schema, and loader."""
from pathlib import Path
import pytest
from rag_eval.settings import EvaluationSettings
from rag_eval.config.loader import load_scenario
def test_settings_default_judge_language_is_en():
"""ragas_judge_language defaults to 'en' when the env var is absent."""
settings = EvaluationSettings(_env_file=None)
assert settings.ragas_judge_language == "en"
def _write_scenario(tmp_path: Path, extra: str) -> Path:
"""Write a minimal valid offline scenario YAML plus the given extra line(s)."""
dataset = tmp_path / "data.csv"
dataset.write_text("sample_id,question,answer,contexts,ground_truth\n", encoding="utf-8")
text = (
"scenario_name: t\n"
"mode: offline\n"
f"dataset: {dataset.name}\n"
"judge_model: gpt-5\n"
"embedding_model: text-embedding-3-small\n"
"metrics: [faithfulness]\n"
"output_dir: out\n"
f"{extra}"
)
path = tmp_path / "s.yaml"
path.write_text(text, encoding="utf-8")
return path
def test_scenario_loads_judge_language_zh(tmp_path):
"""A scenario may declare judge_language: zh and it lands on the dataclass."""
path = _write_scenario(tmp_path, "judge_language: zh\n")
scenario = load_scenario(path)
assert scenario.judge_language == "zh"
def test_scenario_defaults_judge_language_none(tmp_path):
"""Omitting judge_language leaves it None so the factory can apply the settings default."""
path = _write_scenario(tmp_path, "")
scenario = load_scenario(path)
assert scenario.judge_language is None
def test_scenario_rejects_invalid_judge_language(tmp_path):
"""An unsupported judge_language value is rejected at schema validation."""
path = _write_scenario(tmp_path, "judge_language: fr\n")
with pytest.raises(Exception):
load_scenario(path)
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_language_config.py -v`
Expected: FAIL`AttributeError: ragas_judge_language` 或 Scenario 无 `judge_language`)。
- [ ] **Step 3: 实现配置字段**
`rag_eval/settings.py` 第 24 行 `ragas_judge_model` 之后新增:
```python
ragas_judge_language: str = Field(default="en", alias="RAGAS_JUDGE_LANGUAGE")
```
`rag_eval/config/schema.py``ScenarioModel` 中(`embedding_model` 字段之后)新增:
```python
judge_language: Literal["en", "zh"] | None = None
```
`Literal` 已在该文件顶部 `from typing import Any, Literal` 导入,无需改导入。)
`rag_eval/config/loader.py``Scenario(...)` 构造里(`embedding_model=model.embedding_model,` 之后)新增:
```python
judge_language=model.judge_language,
```
`rag_eval/shared/models.py``Scenario` dataclass 中(`doc_weights` 之后,作为带默认值字段)新增:
```python
judge_language: str | None = None
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_language_config.py -v`
Expected: 4 passed。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/settings.py rag_eval/config/schema.py rag_eval/config/loader.py rag_eval/shared/models.py tests/test_judge_language_config.py
git commit -m "Add judge_language config plumbing (settings + scenario)" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 2: ScoreRequest 增加 judge_language 字段
让 score 系列 API 能接收可选 `judge_language``SessionScoreRequest` 继承自动获得。
**Files:**
- Modify: `webapp/models.py:472-475`
- Test: `tests/test_judge_language_config.py`(追加)
**Interfaces:**
- Produces: `ScoreRequest.judge_language: str | None`(默认 `None`),`SessionScoreRequest` 继承。
- [ ] **Step 1: 写失败测试**
`tests/test_judge_language_config.py` 末尾追加:
```python
def test_score_request_judge_language_defaults_none():
"""ScoreRequest exposes an optional judge_language defaulting to None."""
from webapp.models import ScoreRequest
req = ScoreRequest(question="q", answer="a")
assert req.judge_language is None
req_zh = ScoreRequest(question="q", answer="a", judge_language="zh")
assert req_zh.judge_language == "zh"
def test_session_score_request_inherits_judge_language():
"""SessionScoreRequest inherits the judge_language field from ScoreRequest."""
from webapp.models import SessionScoreRequest
req = SessionScoreRequest(session_id="s1", question="q", answer="a", judge_language="zh")
assert req.judge_language == "zh"
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_language_config.py -k judge_language -v`
Expected: FAIL`ScoreRequest``judge_language`)。
- [ ] **Step 3: 实现字段**
`webapp/models.py``ScoreRequest` 中,`embedding_model` 字段(第 472-475 行)之后新增:
```python
judge_language: str | None = Field(
default=None,
description="评判 prompt 语言;'zh' 启用中文评判,为 null 时使用 RAGAS_JUDGE_LANGUAGE(默认 en)。",
)
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_language_config.py -k judge_language -v`
Expected: 2 passed(若 `SessionScoreRequest` 必填字段名不同,按其真实必填字段调整测试构造参数后再跑)。
- [ ] **Step 5: 提交**
```bash
git add webapp/models.py tests/test_judge_language_config.py
git commit -m "Add optional judge_language field to ScoreRequest" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 3: 指标 registry 复用(DRY 重构)
抽出 `build_metric_registry()`,供 factory、inline_scorer、引导脚本共用,消除三处重复。
**Files:**
- Modify: `rag_eval/metrics/factory.py:114-131`
- Modify: `webapp/services/inline_scorer.py:34-45`
- Test: `tests/test_metric_registry.py`
**Interfaces:**
- Produces: `rag_eval.metrics.factory.build_metric_registry(llm: Any, embeddings: Any) -> dict[str, Any]`,返回全部 7 个指标实例。
- [ ] **Step 1: 写失败测试**
创建 `tests/test_metric_registry.py`
```python
"""Tests for the shared metric registry factory."""
from rag_eval.metrics.factory import build_metric_registry
def test_build_metric_registry_has_all_seven_metrics():
"""The registry exposes every supported metric keyed by its canonical name."""
registry = build_metric_registry(llm=object(), embeddings=object())
assert set(registry) == {
"faithfulness", "answer_relevancy", "context_recall", "context_precision",
"noise_sensitivity", "factual_correctness", "semantic_similarity",
}
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_metric_registry.py -v`
Expected: FAIL`ImportError: cannot import name 'build_metric_registry'`)。
- [ ] **Step 3: 实现 registry 工厂并复用**
`rag_eval/metrics/factory.py` 新增函数(放在 `build_metric_pipeline` 之前):
```python
def build_metric_registry(llm: Any, embeddings: Any) -> dict[str, Any]:
"""Instantiate the full set of supported RAGAS metrics keyed by canonical name.
Shared by the scenario pipeline, the inline scorer, and the prompt-cache
bootstrap so the metric set is defined in exactly one place.
"""
return {
"faithfulness": Faithfulness(llm=llm),
"answer_relevancy": AnswerRelevancy(llm=llm, embeddings=embeddings),
"context_recall": ContextRecall(llm=llm),
"context_precision": ContextPrecision(llm=llm),
"noise_sensitivity": NoiseSensitivity(llm=llm),
"factual_correctness": FactualCorrectness(llm=llm),
"semantic_similarity": SemanticSimilarity(embeddings=embeddings),
}
```
`build_metric_pipeline` 中内联的 `registry = { ... }`(第 114-127 行)替换为:
```python
registry = build_metric_registry(llm, embeddings)
```
`webapp/services/inline_scorer.py``_build_metric_instances`(第 34-45 行)改为复用:
```python
def _build_metric_instances(metrics: list[str], llm: Any, embeddings: Any) -> dict[str, Any]:
"""Instantiate only the RAGAS metric objects requested."""
from rag_eval.metrics.factory import build_metric_registry
registry = build_metric_registry(llm, embeddings)
return {name: registry[name] for name in metrics if name in registry}
```
并删除该文件顶部不再使用的 `from ragas.metrics.collections import (...)` 导入块(第 23-31 行)。
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_metric_registry.py tests/test_score_api.py -v`
Expected: registry 测试 passed;既有 score API 测试保持 passed(若个别用例需真实 LLM 则会 skip/无关,重点是不新增失败)。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/metrics/factory.py webapp/services/inline_scorer.py tests/test_metric_registry.py
git commit -m "Extract shared build_metric_registry factory" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 4: 本地化器核心 `judge_prompts.py`
纯逻辑模块:加载缓存、覆盖 prompt、漂移检测、内存 memo、优雅降级。用假 prompt 对象测试,不触碰 RAGAS。
**Files:**
- Create: `rag_eval/metrics/judge_prompts.py`
- Test: `tests/test_judge_prompt_localizer.py`
**Interfaces:**
- Produces:
- `METRIC_PROMPT_ATTRS: dict[str, tuple[str, ...]]`
- `CACHE_ROOT: Path`
- `LocalizationReport(language, applied, skipped, warnings)`dataclass
- `prompt_source_hash(prompt) -> str`
- `apply_localized_prompt(prompt, data: dict) -> None`
- `localize_pipeline_prompts(registry: dict[str, Any], language: str) -> LocalizationReport`
- `reset_cache() -> None`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_prompt_localizer.py`
```python
"""Tests for the judge-prompt localizer (no RAGAS dependency; uses fake prompts)."""
import json
from pydantic import BaseModel
from rag_eval.metrics import judge_prompts as jp
class _In(BaseModel):
question: str
class _Out(BaseModel):
statements: list[str]
class _FakePrompt:
"""Minimal stand-in for a RAGAS BasePrompt with overridable attributes."""
def __init__(self):
self.input_model = _In
self.output_model = _Out
self.instruction = "English instruction."
self.examples = [(_In(question="q"), _Out(statements=["s"]))]
self.language = "english"
class _FakeMetric:
def __init__(self):
self.prompt = _FakePrompt()
def _cache_dict(prompt):
"""Build a valid cache dict for the given fake prompt."""
return {
"metric": "context_recall",
"prompt_attr": "prompt",
"language": "chinese",
"ragas_version": "0.4.3",
"source_hash": jp.prompt_source_hash(prompt),
"instruction": "中文指令。",
"examples": [{"input": {"question": "问题"}, "output": {"statements": ["陈述"]}}],
}
def setup_function(_):
jp.reset_cache()
def test_apply_localized_prompt_overrides_instruction_and_examples():
"""apply_localized_prompt swaps instruction/examples and rebuilds example models."""
prompt = _FakePrompt()
jp.apply_localized_prompt(prompt, _cache_dict(prompt))
assert prompt.instruction == "中文指令。"
assert prompt.examples[0][0].question == "问题"
assert prompt.examples[0][1].statements == ["陈述"]
assert prompt.language == "chinese"
def test_localize_english_is_noop():
"""language='en' leaves the registry untouched."""
metric = _FakeMetric()
report = jp.localize_pipeline_prompts({"context_recall": metric}, "en")
assert report.applied == []
assert metric.prompt.instruction == "English instruction."
def test_localize_applies_from_cache_file(tmp_path, monkeypatch):
"""localize reads <root>/zh/context_recall__prompt.json and applies it."""
metric = _FakeMetric()
root = tmp_path / "configs" / "judge_prompts"
(root / "zh").mkdir(parents=True)
(root / "zh" / "context_recall__prompt.json").write_text(
json.dumps(_cache_dict(metric.prompt), ensure_ascii=False), encoding="utf-8"
)
monkeypatch.setattr(jp, "CACHE_ROOT", root)
report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh")
assert "context_recall.prompt" in report.applied
assert metric.prompt.instruction == "中文指令。"
def test_localize_missing_cache_keeps_english(tmp_path, monkeypatch):
"""A missing cache file degrades gracefully to the English prompt with a warning."""
metric = _FakeMetric()
monkeypatch.setattr(jp, "CACHE_ROOT", tmp_path / "empty")
report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh")
assert metric.prompt.instruction == "English instruction."
assert "context_recall.prompt" in report.skipped
assert report.warnings
def test_localize_stale_hash_warns_but_applies(tmp_path, monkeypatch):
"""A source_hash mismatch still applies Chinese but records a stale warning."""
metric = _FakeMetric()
data = _cache_dict(metric.prompt)
data["source_hash"] = "deadbeef"
root = tmp_path / "configs" / "judge_prompts"
(root / "zh").mkdir(parents=True)
(root / "zh" / "context_recall__prompt.json").write_text(
json.dumps(data, ensure_ascii=False), encoding="utf-8"
)
monkeypatch.setattr(jp, "CACHE_ROOT", root)
report = jp.localize_pipeline_prompts({"context_recall": metric}, "zh")
assert metric.prompt.instruction == "中文指令。"
assert any("stale" in w for w in report.warnings)
def test_cache_memoized(tmp_path, monkeypatch):
"""A second localize call does not re-read the file (in-memory memo)."""
metric = _FakeMetric()
root = tmp_path / "configs" / "judge_prompts"
(root / "zh").mkdir(parents=True)
path = root / "zh" / "context_recall__prompt.json"
path.write_text(json.dumps(_cache_dict(metric.prompt), ensure_ascii=False), encoding="utf-8")
monkeypatch.setattr(jp, "CACHE_ROOT", root)
jp.localize_pipeline_prompts({"context_recall": _FakeMetric()}, "zh")
path.unlink() # delete file; memo should still serve the parsed data
metric2 = _FakeMetric()
report = jp.localize_pipeline_prompts({"context_recall": metric2}, "zh")
assert "context_recall.prompt" in report.applied
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_prompt_localizer.py -v`
Expected: FAIL`ModuleNotFoundError` / 函数缺失)。
- [ ] **Step 3: 实现本地化器**
创建 `rag_eval/metrics/judge_prompts.py`
```python
"""Localize RAGAS collections judge prompts to a target language (e.g. Chinese).
Loads committed, pre-translated prompt cache files from
configs/judge_prompts/<language>/<metric>__<attr>.json and overrides each
metric's prompt instance attributes in place. Missing, corrupt, or schema-drifted
cache entries degrade gracefully to the built-in English prompt so scoring never
breaks.
"""
from __future__ import annotations
import hashlib
import json
import logging
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
logger = logging.getLogger("rag_eval.metrics.judge_prompts")
_REPO_ROOT = Path(__file__).resolve().parents[2]
CACHE_ROOT = _REPO_ROOT / "configs" / "judge_prompts"
# Metric name -> prompt instance attribute names holding a BasePrompt.
# Verified against RAGAS 0.4.3 collections source; semantic_similarity has none.
METRIC_PROMPT_ATTRS: dict[str, tuple[str, ...]] = {
"faithfulness": ("statement_generator_prompt", "nli_statement_prompt"),
"answer_relevancy": ("prompt",),
"context_recall": ("prompt",),
"context_precision": ("prompt",),
"noise_sensitivity": ("statement_prompt", "faithfulness_prompt"),
"factual_correctness": ("prompt", "nli_prompt"),
}
# In-memory memoization of parsed cache files, keyed by (language, metric, attr).
_MEMO: dict[tuple[str, str, str], dict | None] = {}
@dataclass
class LocalizationReport:
"""Outcome of a localize_pipeline_prompts() call, for logging and tests."""
language: str
applied: list[str] = field(default_factory=list)
skipped: list[str] = field(default_factory=list)
warnings: list[str] = field(default_factory=list)
def reset_cache() -> None:
"""Clear the in-memory parsed-cache memo (used by tests)."""
_MEMO.clear()
def prompt_source_hash(prompt: Any) -> str:
"""Return a stable sha256 of a prompt's English instruction + examples."""
examples = [
{"input": inp.model_dump(), "output": out.model_dump()}
for inp, out in getattr(prompt, "examples", [])
]
payload = json.dumps(
{"instruction": getattr(prompt, "instruction", ""), "examples": examples},
ensure_ascii=False,
sort_keys=True,
)
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
def _cache_path(language: str, metric: str, attr: str) -> Path:
"""Resolve the cache file path for one (language, metric, attr) triple."""
return CACHE_ROOT / language / f"{metric}__{attr}.json"
def _load_cache_file(language: str, metric: str, attr: str) -> dict | None:
"""Load and memoize a cache file; return None if absent or unreadable."""
key = (language, metric, attr)
if key in _MEMO:
return _MEMO[key]
path = _cache_path(language, metric, attr)
data: dict | None
try:
data = json.loads(path.read_text(encoding="utf-8"))
except FileNotFoundError:
data = None
except (OSError, json.JSONDecodeError) as exc: # corrupt file: degrade to English
logger.warning("[judge_prompts] cache read failed %s: %s", path, exc)
data = None
_MEMO[key] = data
return data
def apply_localized_prompt(prompt: Any, data: dict) -> None:
"""Override a live prompt's instruction/examples/language from a cache dict.
Examples are rebuilt using the live prompt's input/output models, so an
upstream schema change raises here and is caught by the caller (which then
keeps the English prompt).
"""
examples = [
(prompt.input_model(**ex["input"]), prompt.output_model(**ex["output"]))
for ex in data.get("examples", [])
]
prompt.instruction = data["instruction"]
prompt.examples = examples
prompt.language = data.get("language", "chinese")
def localize_pipeline_prompts(registry: dict[str, Any], language: str) -> LocalizationReport:
"""Override judge prompts in `registry` with cached `language` translations.
`registry` maps metric name -> RAGAS metric instance. Only metrics in
METRIC_PROMPT_ATTRS are touched; unknown metrics and semantic_similarity are
left untouched. English ("en"/"english"/empty) is a no-op.
"""
report = LocalizationReport(language=language)
normalized = (language or "en").strip().lower()
if normalized in ("", "en", "english"):
return report
for metric_name, attrs in METRIC_PROMPT_ATTRS.items():
metric = registry.get(metric_name)
if metric is None:
continue
for attr in attrs:
tag = f"{metric_name}.{attr}"
prompt = getattr(metric, attr, None)
if prompt is None:
report.skipped.append(tag)
continue
data = _load_cache_file(normalized, metric_name, attr)
if data is None:
report.skipped.append(tag)
report.warnings.append(f"missing cache for {tag}")
continue
# Drift detection: warn if the English source changed since caching.
if data.get("source_hash") and data["source_hash"] != prompt_source_hash(prompt):
report.warnings.append(f"stale cache for {tag} (regenerate)")
try:
apply_localized_prompt(prompt, data)
report.applied.append(tag)
except Exception as exc: # noqa: BLE001 schema drift -> keep English
report.warnings.append(f"apply failed for {tag}: {exc}; kept english")
if report.warnings:
logger.warning(
"[judge_prompts] language=%s applied=%d skipped=%d warnings=%s",
normalized, len(report.applied), len(report.skipped), report.warnings,
)
else:
logger.info(
"[judge_prompts] language=%s applied=%d", normalized, len(report.applied)
)
return report
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_prompt_localizer.py -v`
Expected: 6 passed。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/metrics/judge_prompts.py tests/test_judge_prompt_localizer.py
git commit -m "Add judge-prompt localizer with graceful fallback and drift detection" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 5: 接入 factory 与 inline_scorer
让两条运行时路径按语言实际调用本地化器。
**Files:**
- Modify: `rag_eval/metrics/factory.py:96-131`
- Modify: `webapp/services/inline_scorer.py:80-114`
- Test: `tests/test_judge_language_wiring.py`
**Interfaces:**
- Consumes: `build_metric_registry`Task 3)、`localize_pipeline_prompts`Task 4)、`Scenario.judge_language`Task 1)、`settings.ragas_judge_language`Task 1)。
- Produces:
- `build_metric_pipeline(scenario, settings, llm=None, embeddings=None)` 在语言解析为 `zh` 时本地化。
- `InlineScorer.score(..., judge_language: str = "en")``_build_metric_instances(metrics, llm, embeddings, judge_language="en")`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_language_wiring.py`
```python
"""Tests that the factory and inline scorer invoke the localizer per language."""
import rag_eval.metrics.factory as factory
import webapp.services.inline_scorer as inline_mod
def test_build_pipeline_localizes_when_zh(monkeypatch):
"""build_metric_pipeline calls localize_pipeline_prompts with the resolved language."""
calls = []
monkeypatch.setattr(
factory, "localize_pipeline_prompts",
lambda registry, language: calls.append(language),
)
from rag_eval.shared.models import DatasetConfig, Scenario
from rag_eval.settings import EvaluationSettings
from pathlib import Path
scenario = Scenario(
scenario_name="t", mode="offline",
dataset=DatasetConfig(path=Path("x.csv")),
judge_model="gpt-5", embedding_model="text-embedding-3-small",
metrics=["faithfulness"], output_dir=Path("out"),
judge_language="zh",
)
factory.build_metric_pipeline(scenario, EvaluationSettings(_env_file=None),
llm=object(), embeddings=object())
assert calls == ["zh"]
def test_build_pipeline_falls_back_to_settings_default(monkeypatch):
"""When scenario.judge_language is None the settings default is used."""
calls = []
monkeypatch.setattr(
factory, "localize_pipeline_prompts",
lambda registry, language: calls.append(language),
)
from rag_eval.shared.models import DatasetConfig, Scenario
from rag_eval.settings import EvaluationSettings
from pathlib import Path
scenario = Scenario(
scenario_name="t", mode="offline",
dataset=DatasetConfig(path=Path("x.csv")),
judge_model="gpt-5", embedding_model="text-embedding-3-small",
metrics=["faithfulness"], output_dir=Path("out"),
judge_language=None,
)
settings = EvaluationSettings(_env_file=None)
settings.ragas_judge_language = "zh"
factory.build_metric_pipeline(scenario, settings, llm=object(), embeddings=object())
assert calls == ["zh"]
def test_inline_score_threads_judge_language(monkeypatch):
"""InlineScorer.score forwards judge_language into _build_metric_instances."""
seen = {}
monkeypatch.setattr(inline_mod.InlineScorer, "_get_models",
lambda self, j, e, s: (object(), object()))
monkeypatch.setattr(
inline_mod, "_build_metric_instances",
lambda metrics, llm, embeddings, judge_language="en": seen.setdefault("lang", judge_language) or {},
)
class _Pipe:
def __init__(self, *a, **k):
pass
monkeypatch.setattr(inline_mod, "MetricPipeline", _Pipe)
monkeypatch.setattr(inline_mod.asyncio, "run", lambda coro: type("R", (), {"metrics": {}})())
scorer = inline_mod.InlineScorer()
scorer.score(question="q", answer="a", contexts=[], ground_truth=None,
metrics=["faithfulness"], judge_model="gpt-5",
embedding_model="e", settings=object(), judge_language="zh")
assert seen["lang"] == "zh"
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_language_wiring.py -v`
Expected: FAILfactory 未导入/调用 localizer`score()``judge_language` 参数)。
- [ ] **Step 3: 实现接入**
`rag_eval/metrics/factory.py` 顶部(第 27 行 `from .pipeline import MetricPipeline` 之后)新增导入:
```python
from .judge_prompts import localize_pipeline_prompts
```
`build_metric_pipeline` 结尾(构建 registry 之后、`return MetricPipeline(...)` 处)改为先切片再本地化:
```python
registry = build_metric_registry(llm, embeddings)
selected = {name: registry[name] for name in scenario.metrics}
language = (scenario.judge_language or settings.ragas_judge_language or "en")
localize_pipeline_prompts(selected, language)
return MetricPipeline(
metrics=selected,
metric_timeout_seconds=settings.ragas_metric_timeout_seconds,
)
```
`webapp/services/inline_scorer.py`:把 `_build_metric_instances` 增加 `judge_language` 形参并本地化:
```python
def _build_metric_instances(
metrics: list[str], llm: Any, embeddings: Any, judge_language: str = "en"
) -> dict[str, Any]:
"""Instantiate only the RAGAS metric objects requested, localized if needed."""
from rag_eval.metrics.factory import build_metric_registry
from rag_eval.metrics.judge_prompts import localize_pipeline_prompts
registry = build_metric_registry(llm, embeddings)
selected = {name: registry[name] for name in metrics if name in registry}
localize_pipeline_prompts(selected, judge_language)
return selected
```
`InlineScorer.score` 签名(第 80-90 行)末尾加参数并在调用处透传:
```python
def score(
self,
question: str,
answer: str,
contexts: list[str],
ground_truth: str | None,
metrics: list[str],
judge_model: str,
embedding_model: str,
settings: EvaluationSettings,
judge_language: str = "en",
) -> dict[str, float | None]:
"""Score one sample synchronously and return {metric_name: score | None}."""
llm, embeddings = self._get_models(judge_model, embedding_model, settings)
metric_instances = _build_metric_instances(metrics, llm, embeddings, judge_language)
```
(其余函数体不变。)
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_language_wiring.py -v`
Expected: 3 passed。
- [ ] **Step 5: 提交**
```bash
git add rag_eval/metrics/factory.py webapp/services/inline_scorer.py tests/test_judge_language_wiring.py
git commit -m "Wire judge-prompt localization into factory and inline scorer" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 6: 三个 score 端点透传 judge_language
**Files:**
- Modify: `webapp/api/score.py:117-142`
- Modify: `webapp/services/score_job_manager.py:120-141`
- Modify: `webapp/services/session_score_manager.py:205-226`
- Test: `tests/test_score_endpoint_language.py`
**Interfaces:**
- Consumes: `ScoreRequest.judge_language`Task 2)、`settings.ragas_judge_language``InlineScorer.score(..., judge_language=...)`Task 5)。
- [ ] **Step 1: 写失败测试**
创建 `tests/test_score_endpoint_language.py`
```python
"""The /api/score route forwards the resolved judge_language to the scorer."""
from fastapi.testclient import TestClient
import webapp.api.score as score_mod
from webapp.server import create_app
def test_score_route_forwards_judge_language(monkeypatch):
"""A request with judge_language='zh' reaches inline_scorer.score."""
captured = {}
def fake_score(**kwargs):
captured.update(kwargs)
return {"faithfulness": 0.9}
monkeypatch.setattr(score_mod.inline_scorer, "score", lambda **kw: fake_score(**kw))
client = TestClient(create_app())
resp = client.post("/api/score", json={
"question": "q", "answer": "a", "contexts": "c",
"ground_truth": "g", "metrics": ["faithfulness"], "judge_language": "zh",
})
assert resp.status_code == 200
assert captured.get("judge_language") == "zh"
def test_score_route_defaults_language_from_settings(monkeypatch):
"""Omitting judge_language falls back to settings.ragas_judge_language."""
captured = {}
monkeypatch.setattr(score_mod.inline_scorer, "score",
lambda **kw: captured.update(kw) or {"faithfulness": 0.9})
client = TestClient(create_app())
resp = client.post("/api/score", json={
"question": "q", "answer": "a", "contexts": "c",
"ground_truth": "g", "metrics": ["faithfulness"],
})
assert resp.status_code == 200
assert captured.get("judge_language") == "en"
```
(若 `create_app` 的导入路径不同,先 `grep -n "def create_app" webapp/server.py` 校正 import。)
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_score_endpoint_language.py -v`
Expected: FAIL`score()` 未收到 `judge_language`)。
- [ ] **Step 3: 实现透传**
`webapp/api/score.py`:在 `judge_model = request.judge_model or settings.ragas_judge_model`(第 117 行)附近新增解析,并在 `inline_scorer.score(...)` 调用(第 133-142 行)加参数:
```python
judge_model = request.judge_model or settings.ragas_judge_model
embedding_model = request.embedding_model or settings.ragas_embedding_model
judge_language = request.judge_language or settings.ragas_judge_language
```
```python
raw_scores = inline_scorer.score(
question=request.question,
answer=request.answer,
contexts=request.contexts_as_list(),
ground_truth=request.ground_truth,
metrics=effective,
judge_model=judge_model,
embedding_model=embedding_model,
settings=settings,
judge_language=judge_language,
)
```
`webapp/services/score_job_manager.py`:在第 121-122 行之后新增,并在第 132-141 行的 `inline_scorer.score(...)` 加同名参数:
```python
judge_language = request.judge_language or settings.ragas_judge_language
```
```python
judge_language=judge_language,
```
`webapp/services/session_score_manager.py`:在第 206-207 行之后新增,并在第 217-226 行的 `inline_scorer.score(...)` 加同名参数:
```python
judge_language = request.judge_language or settings.ragas_judge_language
```
```python
judge_language=judge_language,
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_score_endpoint_language.py -v`
Expected: 2 passed。
- [ ] **Step 5: 提交**
```bash
git add webapp/api/score.py webapp/services/score_job_manager.py webapp/services/session_score_manager.py tests/test_score_endpoint_language.py
git commit -m "Forward judge_language through score, async, and session endpoints" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 7: 引导脚本 `build_judge_prompt_cache.py`
用 RAGAS 原生 `adapt()` 生成缓存 JSON。用 mock LLM/prompt 测试,不触发真实网络。
**Files:**
- Create: `scripts/build_judge_prompt_cache.py`
- Test: `tests/test_judge_prompt_cache_builder.py`
**Interfaces:**
- Consumes: `build_metric_registry`Task 3)、`METRIC_PROMPT_ATTRS`/`CACHE_ROOT`/`prompt_source_hash`Task 4)、`build_models`(既有)。
- Produces:
- `serialize_prompt(adapted, source_hash, metric, attr, language, ragas_version) -> dict`
- `async build_cache(language, judge_model, embedding_model, settings) -> list[Path]`
- [ ] **Step 1: 写失败测试**
创建 `tests/test_judge_prompt_cache_builder.py`
```python
"""Tests for the prompt-cache bootstrap script (mocked adapt/LLM)."""
import asyncio
import json
from pydantic import BaseModel
import scripts.build_judge_prompt_cache as builder
class _In(BaseModel):
question: str
class _Out(BaseModel):
statements: list[str]
class _FakeAdapted:
def __init__(self):
self.instruction = "中文指令。"
self.language = "chinese"
self.examples = [(_In(question="问题"), _Out(statements=["陈述"]))]
class _FakePrompt:
def __init__(self):
self.input_model = _In
self.output_model = _Out
self.instruction = "English."
self.examples = [(_In(question="q"), _Out(statements=["s"]))]
self.language = "english"
async def adapt(self, target_language, llm, adapt_instruction=False):
assert target_language == "chinese"
assert adapt_instruction is True
return _FakeAdapted()
def test_serialize_prompt_shape():
"""serialize_prompt emits the committed cache schema."""
data = builder.serialize_prompt(_FakeAdapted(), "hash123", "context_recall", "prompt", "zh", "0.4.3")
assert data["metric"] == "context_recall"
assert data["prompt_attr"] == "prompt"
assert data["language"] == "chinese"
assert data["source_hash"] == "hash123"
assert data["instruction"] == "中文指令。"
assert data["examples"] == [{"input": {"question": "问题"}, "output": {"statements": ["陈述"]}}]
def test_build_cache_writes_all_files(tmp_path, monkeypatch):
"""build_cache writes one JSON per (metric, attr) into CACHE_ROOT/<language>/."""
fake_registry = {name: type("M", (), {})() for name in builder.METRIC_PROMPT_ATTRS}
for name, attrs in builder.METRIC_PROMPT_ATTRS.items():
for attr in attrs:
setattr(fake_registry[name], attr, _FakePrompt())
monkeypatch.setattr(builder, "build_models", lambda j, e, s: (object(), object()))
monkeypatch.setattr(builder, "build_metric_registry", lambda llm, emb: fake_registry)
monkeypatch.setattr(builder, "CACHE_ROOT", tmp_path)
written = asyncio.run(builder.build_cache("zh", "gpt-5", "emb", object()))
expected = sum(len(a) for a in builder.METRIC_PROMPT_ATTRS.values())
assert len(written) == expected
sample = tmp_path / "zh" / "context_recall__prompt.json"
assert sample.exists()
data = json.loads(sample.read_text(encoding="utf-8"))
assert data["instruction"] == "中文指令。"
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_judge_prompt_cache_builder.py -v`
Expected: FAIL(模块不存在)。
- [ ] **Step 3: 实现脚本**
创建 `scripts/build_judge_prompt_cache.py`
```python
"""One-off bootstrap: generate committed Chinese judge-prompt cache via RAGAS adapt().
Usage:
python -m scripts.build_judge_prompt_cache --language zh
Runs BasePrompt.adapt("chinese", llm, adapt_instruction=True) for every judge
prompt of every LLM-scored metric and writes the result to
configs/judge_prompts/<language>/<metric>__<attr>.json. Re-run after a RAGAS
upgrade to refresh the cache. Requires a working judge LLM.
"""
from __future__ import annotations
import argparse
import asyncio
import json
import logging
from pathlib import Path
from typing import Any
from rag_eval.compat import ensure_ragas_import_compat
from rag_eval.metrics.factory import build_metric_registry, build_models
from rag_eval.metrics.judge_prompts import CACHE_ROOT, METRIC_PROMPT_ATTRS, prompt_source_hash
from rag_eval.settings import EvaluationSettings
ensure_ragas_import_compat()
logger = logging.getLogger("scripts.build_judge_prompt_cache")
# judge_language short code -> RAGAS adapt() natural-language name.
_ADAPT_LANGUAGE = {"zh": "chinese"}
def serialize_prompt(
adapted: Any, source_hash: str, metric: str, attr: str, language: str, ragas_version: str
) -> dict:
"""Serialize an adapted prompt into the committed cache JSON schema."""
return {
"metric": metric,
"prompt_attr": attr,
"language": getattr(adapted, "language", language),
"ragas_version": ragas_version,
"source_hash": source_hash,
"instruction": adapted.instruction,
"examples": [
{"input": inp.model_dump(), "output": out.model_dump()}
for inp, out in adapted.examples
],
}
async def build_cache(
language: str, judge_model: str, embedding_model: str, settings: EvaluationSettings
) -> list[Path]:
"""Generate and write the full prompt cache for one language; return written paths."""
import ragas
target = _ADAPT_LANGUAGE.get(language, language)
llm, embeddings = build_models(judge_model, embedding_model, settings)
registry = build_metric_registry(llm, embeddings)
out_dir = CACHE_ROOT / language
out_dir.mkdir(parents=True, exist_ok=True)
written: list[Path] = []
for metric, attrs in METRIC_PROMPT_ATTRS.items():
for attr in attrs:
prompt = getattr(registry[metric], attr)
source_hash = prompt_source_hash(prompt)
adapted = await prompt.adapt(target, llm, adapt_instruction=True)
data = serialize_prompt(adapted, source_hash, metric, attr, language, ragas.__version__)
path = out_dir / f"{metric}__{attr}.json"
path.write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")
written.append(path)
logger.info("wrote %s", path)
return written
def main() -> None:
"""CLI entry point: parse args, resolve models, and build the cache."""
parser = argparse.ArgumentParser(description="Build the Chinese judge-prompt cache.")
parser.add_argument("--language", default="zh")
parser.add_argument("--judge-model", default=None)
parser.add_argument("--embedding-model", default=None)
args = parser.parse_args()
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
settings = EvaluationSettings()
judge_model = args.judge_model or settings.ragas_judge_model
embedding_model = args.embedding_model or settings.ragas_embedding_model
paths = asyncio.run(build_cache(args.language, judge_model, embedding_model, settings))
logger.info("done: %d files written", len(paths))
if __name__ == "__main__":
main()
```
同时确保 `scripts/` 可作为包导入:若 `scripts/__init__.py` 不存在则创建空文件(测试用 `import scripts.build_judge_prompt_cache` 需要)。
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_judge_prompt_cache_builder.py -v`
Expected: 2 passed。
- [ ] **Step 5: 提交**
```bash
git add scripts/build_judge_prompt_cache.py scripts/__init__.py tests/test_judge_prompt_cache_builder.py
git commit -m "Add judge-prompt cache bootstrap script" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
### Task 8: 生成中文缓存 + 示例场景 + 文档
产出真实中文缓存、暴露配置示例、补文档。这是让功能端到端可用的收尾任务。
**Files:**
- Create: `configs/judge_prompts/zh/*.json`10 个文件)
- Modify: 一个 siemens 评估场景 YAML + 一个 offline 示例 YAML(增加 `judge_language: zh`
- Modify: `README.md`
- [ ] **Step 1: 生成中文缓存**
有可用评判 LLM 时运行(生产做法):
```bash
python -m scripts.build_judge_prompt_cache --language zh
```
预期在 `configs/judge_prompts/zh/` 生成 10 个文件:
`faithfulness__statement_generator_prompt.json``faithfulness__nli_statement_prompt.json``answer_relevancy__prompt.json``context_recall__prompt.json``context_precision__prompt.json``noise_sensitivity__statement_prompt.json``noise_sensitivity__faithfulness_prompt.json``factual_correctness__prompt.json``factual_correctness__nli_prompt.json`(注意 `answer_relevancy` 与部分指标各一个)。
**无 LLM 访问时的兜底**:按 Task 4 缓存 schema 手工撰写等价中文 JSON`instruction` 译为中文、`examples` 内字符串译为中文、`source_hash` 用当前英文源经 `prompt_source_hash` 计算填入;结构字段名/枚举/数字保持不变)。功能与测试不依赖翻译质量,可后续用脚本刷新。
- [ ] **Step 2: 验证缓存被正确加载**
写一个临时验证(跑完删除或保留为集成测试 `tests/test_zh_cache_integration.py`):
```python
"""Integration: committed zh cache loads onto a real metric instance."""
from rag_eval.metrics.judge_prompts import CACHE_ROOT, localize_pipeline_prompts
def test_zh_cache_files_present():
"""All 10 expected zh cache files are committed."""
expected = [
"faithfulness__statement_generator_prompt.json",
"faithfulness__nli_statement_prompt.json",
"answer_relevancy__prompt.json",
"context_recall__prompt.json",
"context_precision__prompt.json",
"noise_sensitivity__statement_prompt.json",
"noise_sensitivity__faithfulness_prompt.json",
"factual_correctness__prompt.json",
"factual_correctness__nli_prompt.json",
]
for name in expected:
assert (CACHE_ROOT / "zh" / name).exists(), f"missing {name}"
```
Run: `python -m pytest tests/test_zh_cache_integration.py -v`
Expected: PASS(确认 9 个必需文件在位;`answer_relevancy` 只有 1 个 prompt,总计与实际 `METRIC_PROMPT_ATTRS` 展开数一致)。
> 注:`METRIC_PROMPT_ATTRS` 展开后文件总数 = 2+1+1+1+2+2 = 9。若脚本因某指标含多 prompt 而不同,以 `METRIC_PROMPT_ATTRS` 为准同步该测试清单。
- [ ] **Step 3: 场景 YAML 增加 judge_language**
在一个 siemens 评估场景与一个 offline 示例的顶层加入:
```yaml
judge_language: zh
```
(用 `grep -rl "mode: offline" scenarios/` 找到目标文件;仅改评估场景,勿改 dataset_build 场景。)
- [ ] **Step 4: 补文档**
`README.md` 增加一节「中文评判 Prompt 适配」,说明:
- 场景 `judge_language: zh` 与 score API `judge_language` 字段、`RAGAS_JUDGE_LANGUAGE` 全局默认。
- 重新生成缓存命令:`python -m scripts.build_judge_prompt_cache --language zh`
- RAGAS 升级后需重跑脚本(漂移检测会在日志告警)。
- [ ] **Step 5: 运行相关测试并提交**
Run: `python -m pytest tests/test_judge_prompt_localizer.py tests/test_zh_cache_integration.py -v`
Expected: 全部 passed。
```bash
git add configs/judge_prompts/zh/*.json scenarios/ README.md tests/test_zh_cache_integration.py
git commit -m "Add committed zh judge-prompt cache and enable it in sample scenarios" -m "Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>"
```
---
## 最终回归
- [ ] 全量单测:`python -m pytest tests/ -q`
- Expected: 本计划新增测试全部通过;不引入新的失败。已知的 6 个历史失败(见设计背景,与本功能无关)保持原状,勿在本计划内处理。
## Self-Review 记录
- **Spec 覆盖**:§4.4 配置面→Task1/2;§4.3 本地化器→Task4;§4.5 集成→Task5/6;§4.1 引导脚本→Task7;§4.2 缓存文件→Task8;§5 错误处理/漂移→Task4missing/stale/apply-fail 测试);§6 测试→各 Task;§2 指标映射→Task4 常量;§8 兼容性→默认 `en` no-opTask4/5 覆盖)。
- **占位符扫描**:无 TBD/TODO;所有步骤含真实代码与命令。
- **类型一致**`localize_pipeline_prompts(registry, language)``build_metric_registry(llm, embeddings)``InlineScorer.score(..., judge_language="en")``serialize_prompt(...)``build_cache(...)` 在定义与调用处签名一致。