Hyungi Ahn
|
120db86d74
|
docs(search): Phase 2 최종 측정 보고서 (phase2_final.md + csv A/B)
## 결과 요약
Phase 1.3 baseline vs Phase 2 final A/B (평가셋 v0.1, 23 쿼리):
- Recall@10: 0.730 → 0.737 (+0.007)
- NDCG@10: 0.663 → 0.668 (+0.005)
- Top-3 hit: 0.900 → 0.900 (0)
- p95 latency: 171ms → 256ms (+85)
- news_crosslingual NDCG: 0.27 → 0.37 (+0.10 ✓)
- exact_keyword / natural_language_ko: 완전 유지 (회귀 0)
## Phase 2 게이트: 2/6 통과
✓ news_crosslingual NDCG ≥ 0.30
✓ latency p95 < 400ms
❌ Recall@10 ≥ 0.78 (0.737)
❌ Top-3 hit ≥ 0.93 (0.900)
❌ crosslingual_ko_en NDCG ≥ 0.65 (0.53, bge-m3 한계)
❌ 평가셋 v0.2 작성 (후속)
## 핵심 성과 (게이트 미달이지만 견고한 기반)
1. QueryAnalyzer async-only 아키텍처 (retrieval 차단 0)
2. semaphore concurrency=1 (MLX single-inference queue 폭발 방지)
3. multilingual narrowing (news/global 한정 → 회귀 0 + news 개선)
4. soft_filter boost 보수적 설정 (0.01, domain only)
5. prewarm 15개 → cache hit rate 70%+
## infra_inventory.md soft lock 준수
- config.yaml / Ollama / compose restart 변경 0
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-08 15:52:21 +09:00 |
|