wangwei and Copilot
24a8688a34
Add committed zh judge-prompt cache + enable in Siemens scenarios + docs
...
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-07-01 18:15:53 +08:00
wangwei and Copilot
bd5658c3ac
Wire judge-prompt localization into factory and inline scorer
...
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-07-01 17:58:43 +08:00
wangwei and Copilot
555328cb3b
Add judge-prompt localizer with graceful fallback and drift detection
...
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-07-01 17:55:26 +08:00
wangwei and Copilot
2bb804b059
Extract shared build_metric_registry factory (DRY)
...
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-07-01 17:52:55 +08:00
wangwei
9828b1d44c
update
2026-06-27 14:31:45 +08:00
wangwei and Copilot
1df4010acc
fix(llm): resolve score runtime config from saved profiles
...
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-06-26 20:34:01 +08:00
wangwei and Copilot
e0b064587f
feat: add metric/doc weight computation module (weights.py)
...
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-06-18 16:47:47 +08:00
wangwei
24956bbf75
更新
2026-06-16 18:12:33 +08:00
wangwei and Claude
f5c2dce64a
feat(advisor): add optimization advisor module
...
- rag_eval/advisor/: new package with rules engine, LLM analyzer, writer
- rules.py: 7-metric diagnostic rules (warning/critical thresholds, top-3 low samples)
- llm_analyzer.py: Chinese optimization report via judge_model, graceful fallback
- writer.py: writes optimization_advice.md + log summary
- __init__.py: run_advisor() entry point (no-op when optimization_advisor=False)
- Scenario.optimization_advisor: new bool field (default False)
- ScenarioModel: same field added, loader.py透传
- RunArtifactPaths.advice_md: new path field
- factory.py: build_models() now public; build_metric_pipeline() accepts pre-built llm/embeddings
- runner.py: lifts llm, passes to pipeline and advisor; calls run_advisor() at end
- siemens online YAML: optimization_advisor: true enabled
- tests: 9 rules tests + 6 writer tests, all pass
- docs: advisor section added to engine-flow.md and architecture.md
Co-Authored-By: Claude <noreply@anthropic.com >
2026-06-16 17:06:19 +08:00
wangwei and Claude
629304aa6d
feat(logging): add structured evaluation logs for metric-level debugging
...
- pipeline.py: log each metric score/timeout/error with sample_id,
elapsed time, and score value; log NaN list per sample; progress
counter N/total after each sample completes
- evaluator.py: log eval start, dataset counts, adapter enrichment
progress (per-sample OK/FAIL with elapsed), metric scoring summary,
and per-metric NaN rate at end of run
- runner.py: _setup_logging() helper writes to stderr + optional file;
ragas/httpx/openai noisy loggers throttled to WARNING
- main.py: add --log-file and --log-level CLI flags
Usage:
python main.py --scenario scenarios/online/... --log-file logs/eval.log --log-level DEBUG
Co-Authored-By: Claude <noreply@anthropic.com >
2026-06-16 10:48:41 +08:00
Guangfei.Zhao
9cbdc1d95d
first commit
2026-06-12 14:02:15 +08:00