paper-follower/app/normalize.py(223 行)字段名不同、摘要格式不同、标题大小写标点不同。不去重就会在日报里出现两遍。
三级唯一键(第 217-223 行):
def paper_uid(doi: str = "", arxiv_id: str = "", title_norm: str = "") -> str:
"""内部唯一键:DOI > arXiv id > 标题指纹"""
if doi:
return f"doi:{norm_doi(doi)}"
if arxiv_id:
return f"arxiv:{norm_arxiv_id(arxiv_id)}"
return f"title:{title_norm[:120]}"
倒排索引还原(第 204-214 行)——OpenAlex 为了省流量,摘要是"词→位置":
def abstract_from_inverted_index(inv: dict | None) -> str:
"""OpenAlex 的摘要是倒排索引 {词: [位置...]},还原成正常文本"""
slots: list[str] = []
for word, positions in inv.items():
for pos in positions:
while len(slots) <= pos:
slots.append("")
slots[pos] = word
return " ".join(w for w in slots if w).strip()
两套匹配规则——这是全文件最有价值的工程判断(第 150-201 行):
def term_hits(text, term):
"""主题词命中判定(按词干比对)。规则:
- 词组里有高特异性词时,命中任一个就算数(摘要常写 bioelectrical 而非完整词组)
- 全是泛词时要求全部命中(否则单凭 potential 就把核聚变论文放进来)
这是**召回**用的宽松规则,故意比 term_all_hits 宽。"""
def term_all_hits(text, term):
"""交叉监控用的**严格**判定:词组里每个实词都必须命中。
...实测切到宽松规则后,"ion channel expression × development"
立刻捞进了《阳光暴晒对日灼病的影响》。交叉监控宁可漏也不能错。"""
配套 stem() / squash_doubles() / same_stem() 处理形态差异:signalling ↔ signaling、regenerative ↔ regeneration 要认;generation ↔ regeneration 不能误合。
re)_STOP/SPECIFIC_LEN 等常量可配置)tests/test_normalize.py(105 行)biblionorm:DOI/arXiv 归一、标题指纹、词干匹配、倒排摘要还原doi-regex / arxiv 等库做单点归一;你的实现是成套的、带取舍说明的nltk/snowball,但你的"极简 stem + 双写辅音压缩"零依赖且够用第 17 课(数据流水线)