siemens_ragas

Author	SHA1	Message	Date
wangwei	e0b064587f	feat: add metric/doc weight computation module (weights.py) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-06-18 16:47:47 +08:00
wangwei	078097af00	docs: add metric/doc weights implementation plan	2026-06-18 16:43:08 +08:00
wangwei	ca586bf9bb	docs: add metric and doc weights feature design spec Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-06-18 16:37:18 +08:00
wangwei	9ad2daff73	feat: restore API文档 nav item (iframe /docs) without touching other 4 modules	2026-06-17 11:24:16 +08:00
wangwei	e8af5b906c	chore: remove API docs iframe nav item, rename title to RAGAS 评估控制台	2026-06-17 11:18:01 +08:00
wangwei	8ea2b9c7d2	feat: add API文档 nav item with embedded Swagger UI iframe	2026-06-17 11:09:55 +08:00
wangwei	074800b741	feat: add history report switcher dropdown in report detail view	2026-06-17 10:35:56 +08:00
wangwei	3019390592	feat: add export-to-PDF via browser print with @media print CSS	2026-06-17 10:28:01 +08:00
wangwei	24956bbf75	更新	2026-06-16 18:12:33 +08:00
wangwei	ca01e44ad2	feat(webapp): add session persistence via URL hash routing + sessionStorage - app.js: hash-based router (#runs / #new / #profiles / #report/{runId}) - navigate() pushes history entries for back/forward support - _restoreSession() reads hash on load and popstate - sessionStorage fallback for same-tab refreshes - run-card highlights selected run (.run-card.selected) - runner.js: use App.navigate() for report redirect; persist lastRunId to sessionStorage - index.html: report nav button starts disabled (enabled on run select/restore) - app.css: .run-card.selected with petrol border + ring Co-Authored-By: Claude <noreply@anthropic.com>	2026-06-16 17:55:07 +08:00
wangwei	1a2cc534b8	feat(webapp): add optimization advice section to report UI - index.html: add section ⑤ advice block (hidden by default, shown when advice_markdown present) - report.js: add renderAdvice() called in render(), simple Markdown→HTML converter - app.js: add noise_sensitivity / factual_correctness / semantic_similarity to shortMetric map - app.css: add .advice-panel, .advice-badge, .advice-md styles (purple left-border theme) Co-Authored-By: Claude <noreply@anthropic.com>	2026-06-16 17:26:37 +08:00
wangwei	91c0dab4f9	fix(advisor): fix LLM API call, wire advice_markdown to webapp, update .env.example timeouts - llm_analyzer.py: use llm.langchain_llm.ainvoke() (correct RAGAS 0.4.3 API) - webapp/models.py: add advice_markdown field to ReportData - webapp/services/run_reader.py: add read_advice_markdown() reading optimization_advice.md - webapp/services/report_builder.py: pass advice_markdown into ReportData - .env.example: OPENAI_TIMEOUT_SECONDS 30→180, RAGAS_METRIC_TIMEOUT_SECONDS 45→300 Co-Authored-By: Claude <noreply@anthropic.com>	2026-06-16 17:12:32 +08:00
wangwei	f5c2dce64a	feat(advisor): add optimization advisor module - rag_eval/advisor/: new package with rules engine, LLM analyzer, writer - rules.py: 7-metric diagnostic rules (warning/critical thresholds, top-3 low samples) - llm_analyzer.py: Chinese optimization report via judge_model, graceful fallback - writer.py: writes optimization_advice.md + log summary - __init__.py: run_advisor() entry point (no-op when optimization_advisor=False) - Scenario.optimization_advisor: new bool field (default False) - ScenarioModel: same field added, loader.py透传 - RunArtifactPaths.advice_md: new path field - factory.py: build_models() now public; build_metric_pipeline() accepts pre-built llm/embeddings - runner.py: lifts llm, passes to pipeline and advisor; calls run_advisor() at end - siemens online YAML: optimization_advisor: true enabled - tests: 9 rules tests + 6 writer tests, all pass - docs: advisor section added to engine-flow.md and architecture.md Co-Authored-By: Claude <noreply@anthropic.com>	2026-06-16 17:06:19 +08:00
wangwei	d68399d39b	chore: update startup scripts and .env.example for LLM profile feature	2026-06-16 17:03:25 +08:00
wangwei	719c3b4ca4	test: ensure test package structure and all webapp tests pass	2026-06-16 16:27:54 +08:00
wangwei	5b60ed12ea	feat: add LLM role-assignment panel to 新建评估 view	2026-06-16 16:27:00 +08:00
wangwei	dc8baf8662	feat: add LLM配置 management page (profiles view)	2026-06-16 16:25:20 +08:00
wangwei	e329f59139	feat: add yaml_patcher service to apply LLM profiles to scenario YAML	2026-06-16 16:21:19 +08:00
wangwei	b19054bd66	feat: add /api/llm-profiles CRUD router	2026-06-16 16:18:40 +08:00
wangwei	5d09deb420	feat: add ProfileManager service with JSON persistence	2026-06-16 16:14:31 +08:00
wangwei	b98af29449	feat: add LLMProfile pydantic models	2026-06-16 16:10:37 +08:00
wangwei	4173a40d93	feat(scripts): add run_eval.bat / run_eval.ps1 evaluation launcher scripts Both scripts support: - Shortcut args: online (default), offline, or any custom .yaml path - Second arg: log level (DEBUG/INFO/WARNING/ERROR), default INFO - Auto-timestamped log file saved to logs\eval_<date>_<time>.log - Sets PYTHONIOENCODING=utf-8 and PYTHONPATH=. automatically - Friendly error/success banners with log file path Usage: run_eval.bat # online eval run_eval.bat offline DEBUG # offline eval with DEBUG logs .\run_eval.ps1 online DEBUG # PowerShell equivalent Co-Authored-By: Claude <noreply@anthropic.com>	2026-06-16 11:16:53 +08:00
wangwei	629304aa6d	feat(logging): add structured evaluation logs for metric-level debugging - pipeline.py: log each metric score/timeout/error with sample_id, elapsed time, and score value; log NaN list per sample; progress counter N/total after each sample completes - evaluator.py: log eval start, dataset counts, adapter enrichment progress (per-sample OK/FAIL with elapsed), metric scoring summary, and per-metric NaN rate at end of run - runner.py: _setup_logging() helper writes to stderr + optional file; ragas/httpx/openai noisy loggers throttled to WARNING - main.py: add --log-file and --log-level CLI flags Usage: python main.py --scenario scenarios/online/... --log-file logs/eval.log --log-level DEBUG Co-Authored-By: Claude <noreply@anthropic.com>	2026-06-16 10:48:41 +08:00
wangwei	1ff4a3943a	feat(dataset-builder): add retry logic and ASCII-safe logging for Siemens PDF pipeline - question_generator.py: add max_retries=3/retry_delay=5s loop with exponential backoff on LLM timeout or server errors; encode filenames with ascii/replace before printing to avoid UnicodeEncodeError on Windows cp1252 consoles - runner.py: encode PDF filenames ASCII-safe for progress messages; catch generation failures per-document and skip (or re-raise) based on failure_mode, preventing one bad doc from aborting the whole build Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>	2026-06-15 23:06:33 +08:00
wangwei	75ae7927ad	Add Siemens CT document evaluation scenario (three-step pipeline) - scenarios/siemens_build/siemens-pdf-build.yaml: dataset build for all 17 Siemens medical-imaging PDFs (aliyun_docmind parser, 10 questions/doc, failure_mode=skip, ~170 question total) - scenarios/offline/siemens-pdf-offline-smoke.yaml: offline evaluation using source chunks as contexts and ground_truth as answer (up to 30 samples) - scenarios/online/siemens-pdf-question-bank-online.yaml: online evaluation calling siemens_pdf_qa adapter, batch_size=4, up to 50 samples - apps/siemens_pdf_qa/adapter.py: Siemens-specific adapter with bilingual (zh/en) system prompt and strict evidence-grounding for CT domain - scripts/build_siemens_offline_smoke.py: helper to derive offline smoke CSV from completed dataset build artifacts (run after dataset build step) - docs/superpowers/specs/2026-06-15-siemens-scenario-design.md: design spec All three scenarios are automatically discovered by the web console. Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>	2026-06-15 17:00:52 +08:00
wangwei	1288a366d1	Fix start.bat (ASCII-only, guaranteed window stays open) + add start.ps1 - start.bat: remove all Chinese characters (caused silent failure when Windows batch parser ran before chcp 65001 took effect); add :error label so window stays open with pause on any failure - start.ps1: PowerShell alternative launcher with coloured output, works without worrying about cmd.exe encoding issues Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>	2026-06-15 16:14:53 +08:00
wangwei	e89695e490	Add RAGAS evaluation web console (FastAPI + vanilla JS) - webapp/: FastAPI backend with runs/scenarios/evaluations API routers; services for run_reader, report_builder, scenario_scanner, task_manager (lazy ragas import — server boots even without ragas); Pydantic models - webapp/static/: single-page console (layout A: left-nav + main area); report detail with metric cards, Chart.js distribution histogram, grouping table, lowest-score sample review; trigger evaluation + log polling - webmain.py: uvicorn entry point (alongside existing main.py CLI) - start.bat: Windows one-click launcher with env checks and auto-browser open - rag_eval/datasets/: implement missing loader + normalizer modules (load_dataset_records, normalize_records) required by evaluator - scripts/seed_sample_run.py: generate realistic demo run artifacts - .gitignore: exclude datasets/ data files but keep rag_eval/datasets/ source Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>	2026-06-15 15:53:57 +08:00
Guangfei.Zhao	9cbdc1d95d	first commit	2026-06-12 14:02:15 +08:00

28 Commits