paper-follower/app/scoring.py(244 行)这个引擎强制要求"每条推荐都能说出人话理由,写不出理由就不许进日报"。
六维加权(权重在文件开头明确列出):
"""
权重沿用任务卡:
score = 0.35×主题匹配 + 0.20×引用势能 + 0.15×种子关联
+ 0.15×作者权重 + 0.10×可复现性 + 0.05×新颖度
"""
WEIGHTS = {"topic": 0.35, "citation": 0.20, "seed": 0.15,
"author": 0.15, "repro": 0.10, "novelty": 0.05}
主流程(第 177-244 行,节选):
def score_paper(paper, profile, seed_ids=None, cross_pairs=None, low_trust_venues=None, ...):
topic, hit_title, hit_abs = _topic_score(paper, profile)
citation = _citation_score(paper)
seed, seed_reasons = _seed_score(paper, profile, seed_ids)
...
parts = {"topic": topic, "citation": citation, "seed": seed,
"author": author, "repro": repro, "novelty": novelty}
score = sum(WEIGHTS[k] * v for k, v in parts.items())
trust, trust_note = venue_trust(paper, low_trust_venues, low_trust_factor)
score *= trust # 无同行评审的平台打折
...
reasons = []
if hit_title: reasons.append(f"标题命中主题词:{hit_title[0]}")
if hit_abs: reasons.append(f"摘要命中主题词:{hit_abs[0]}")
...
return {"score": round(score, 4), "score_detail": {...}, "reasons": reasons, ...}
来源可信度折扣的设计说明(第 162-174 行注释)写得很克制:
"""来源可信度:无同行评审的自存档平台打折。
不打折到 0 是因为 arXiv / bioRxiv 上也有真正重要的东西;
也不直接丢掉,是因为"忽略"比"标注后降权"更容易漏掉好东西。"""
score_detail 里否则降级为"主题关联弱,需人工确认"(防止宽词误伤)
interests.Profile / CrossPair 数据结构(轻)tests/test_scoring.py(136 行)explainable-scorer:定义一个 Scorer 基类,子类实现若干 dimension(paper) -> (value, reason),框架负责加权与理由汇总
| 方案 | 可解释 | 成本 | 稳定 |
|---|---|---|---|
| 规则加权(本功能) | 高 | 零 | 高 |
| 机器学习排序 | 低 | 需训练 | 中 |
| LLM 打分 | 中 | 高 | 低(随机性) |
第 17 课(数据流水线)