本文介绍一种简洁可靠的 Pandas 方法,通过字符串分割与集合成员判断,准确统计每行 original 文本中出现在对应 people 列(逗号分隔)中的人名词频,自动处理空值与边界匹配问题。
本文介绍一种简洁可靠的 pandas 方法,通过字符串分割与集合成员判断,准确统计每行 `original` 文本中出现在对应 `people` 列(逗号分隔)中的人名词频,自动处理空值与边界匹配问题。
在文本分析任务中,常需判断某段自然语言中是否包含指定关键词(如人名、标签等),并统计其出现次数。与模糊匹配或正则搜索不同,本场景强调精确单词匹配——即仅当 original 中独立单词(非子串)完全等于 people 中某一拆分项时才计数。例如 "Bond" 应匹配 "Bond just met the Bond girl" 中的两个 "Bond",但 "Mary" 不应匹配 "All Marys are here" 中的 "Marys"(因非完全相等)。
以下为推荐实现方案,兼顾可读性、性能与鲁棒性:
import pandas as pd
# 构建示例数据
df = pd.DataFrame({
"original": [
"John is a good friend",
"Mary and Peter are going to marry",
"Bond just met the Bond girl",
"Chris is having dinner",
"All Marys are here"
],
"people": ["John, Mary", "Peter, Mary", "Bond", None, "Mary"]
})
# 关键步骤:填充空值 + 拆分 + 逐词比对 + 求和
df["people"] = df["people"].fillna("") # 防止 str.split 报错
df["result"] = [
sum(word in person_list for word in text.split())
for text, person_list in zip(
df["original"],
df["people"].str.split(", ") # 注意:split 分隔符含空格,更健壮
)
]
print(df)
输出结果:
original people result 0 John is a good friend John, Mary 1 1 Mary and Peter are going to marry Peter, Mary 2 2 Bond just met the Bond girl Bond 2 3 Chris is having dinner 0 4 All Marys are here Mary 0
✅ 为什么此方法更优?
- 避免正则陷阱:原尝试使用 re.search(f'\b{p}\b', o) 易受特殊字符(如括号、点号)干扰,且 o 是非法转义(导致报错);本方案无需正则,规避语法风险。
- 精准单词匹配:word in person_list 依赖 Python 原生字符串相等判断,天然满足“全词匹配”语义(如 "Mary" ≠ "Marys")。
- 空值安全:fillna("") 确保 str.split() 在 NaN 上返回空列表 [],使 sum(...) 自然得 0。
- 高效简洁:纯 Python 列表推导式 + Pandas 向量化拆分,无嵌套循环,代码易维护。
⚠️ 注意事项:
- 若 people 列存在多余空格(如 "John , Mary "),建议预处理:df["people"].str.replace(r's*,s*', ', ').str.strip();
- 如需忽略大小写匹配,可统一转小写:word.lower() in [p.lower() for p in person_list];
- 对超大数据集,可考虑用 apply + lambda 替代列表推导以提升可读性,性能差异通常可忽略。
该方法直击需求本质——以最小代价实现可靠、可解释的单词交集计数,是 Pandas 文本列间匹配任务的实践范式。










