feat(advisor): add 0.85 advisory threshold triggering LLM suggestions
- Add advisory_threshold=0.85 field to MetricRule (higher-is-better metrics) - diagnose() now emits severity='low' for scores in (warning_threshold, 0.85) - noise_sensitivity (lower-is-better) keeps its existing two-tier thresholds - writer.py: severity labels mapped to Chinese (严重/警告/待优化) - llm_analyzer.py: prompt explains low/warning/critical tiers in Chinese - Tests: 5 new cases for 'low' severity, updated log summary assertions Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
@@ -22,22 +22,31 @@ _PROMPT_TEMPLATE = """\
|
||||
|
||||
## 报告要求
|
||||
|
||||
1. 按指标分节(## 指标名 [severity]),先解释"为什么低"(结合低分样本具体分析),再给出"具体怎么改"
|
||||
2. "具体怎么改"要结合低分样本的实际内容,而不只是泛泛建议
|
||||
3. 最后写一节 **## 优先优化次序**,按性价比排序(不增加 LLM 调用次数的优化优先)
|
||||
4. 语言简洁,面向工程师,不要废话,不要重复列表内容
|
||||
1. 按指标分节(## 指标名 [严重程度]),先解释"为什么低"(结合低分样本具体分析),再给出"具体怎么改"
|
||||
2. 严重程度说明:critical=严重(<阈值50%),warning=警告(<阈值70%),low=待优化(低于0.85,有提升空间)
|
||||
3. "具体怎么改"要结合低分样本的实际内容,而不只是泛泛建议
|
||||
4. 最后写一节 **## 优先优化次序**,按性价比排序(不增加 LLM 调用次数的优化优先),critical 和 warning 项优先于 low 项
|
||||
5. 语言简洁,面向工程师,不要废话,不要重复列表内容
|
||||
|
||||
只输出 Markdown 报告正文,不要任何前置说明。
|
||||
"""
|
||||
|
||||
|
||||
_SEVERITY_LABEL_ZH: dict[str, str] = {
|
||||
"critical": "严重",
|
||||
"warning": "警告",
|
||||
"low": "待优化",
|
||||
}
|
||||
|
||||
|
||||
def _build_diagnosis_summary(diagnoses: list[Diagnosis]) -> str:
|
||||
lines = []
|
||||
for d in diagnoses:
|
||||
direction = "(越低越好)" if d.metric == "noise_sensitivity" else ""
|
||||
label = _SEVERITY_LABEL_ZH.get(d.severity, d.severity)
|
||||
lines.append(
|
||||
f"- **{d.metric}** {direction} 均值={d.mean_score:.4f},"
|
||||
f"阈值={d.threshold},严重程度={d.severity}"
|
||||
f"阈值={d.threshold},严重程度={label}"
|
||||
)
|
||||
lines.append(f" - 可能原因:{'; '.join(d.root_causes)}")
|
||||
lines.append(f" - 建议动作:{'; '.join(d.suggested_actions)}")
|
||||
|
||||
Reference in New Issue
Block a user