
本文详解如何在 Selenium 中处理部分容器缺失 标签导致 NoSuchElementException 的常见爬虫问题,通过 XPath 条件过滤、异常捕获与容错设计,确保 for 循环完整执行并稳定提取标题、副标题及 href 链接。
本文详解如何在 selenium 中处理部分容器缺失 `` 标签导致 `nosuchelementexception` 的常见爬虫问题,通过 xpath 条件过滤、异常捕获与容错设计,确保 for 循环完整执行并稳定提取标题、副标题及 href 链接。
在使用 Selenium 解析动态新闻列表(如 The Sun 体育页)时,一个典型陷阱是:看似结构统一的卡片容器(如 div.teaser__copy-container),实际并非全部包含 标签——部分可能仅用于广告、占位或富媒体嵌入。原始代码中直接调用 container.find_element(By.XPATH, './/a') 会在遇到首个无链接的容器时立即抛出 NoSuchElementException,导致循环中断,数据截断。
根本原因在于:XPath 表达式 //div[@class="teaser__copy-container"] 匹配了 68 个容器,但其中仅 67 个内嵌 标签。这意味着存在一个“孤儿容器”,它不含可点击链接,却仍被纳入循环范围。
✅ 推荐解决方案(双重保障)
方案一:XPath 前置过滤(推荐首选)
在定位容器阶段即排除无链接项,从源头避免异常:
# ✅ 仅选取内部包含 <a> 标签的容器
containers = browser.find_elements(
By.XPATH,
'//div[@class="teaser__copy-container" and .//a]'
)</a>
该表达式中 and .//a 是关键:它要求每个匹配的 div 必须至少有一个后代 元素。实测后 containers 数量将精确为 67,与真实链接数一致,后续循环无需额外判断。
方案二:异常捕获 + 默认值(增强鲁棒性)
若需保留所有容器(例如某些卡片需记录“无链接”状态),则改用 find_elements()(复数形式)配合条件判断:
for container in containers:
try:
title = container.find_element(By.CSS_SELECTOR, 'span').get_attribute("textContent").strip()
sub_title = container.find_element(By.CSS_SELECTOR, 'h3').get_attribute("textContent").strip()
# 使用 find_elements 并取首项,避免异常
link_elements = container.find_elements(By.XPATH, './/a')
link = link_elements[0].get_attribute("href") if link_elements else ""
except Exception as e:
print(f"Warning: Skipped container due to {type(e).__name__}")
title = sub_title = link = ""
titles.append(title)
sub_titles.append(sub_title)
links.append(link)
⚠️ 注意事项:
- 永远避免在循环中对不稳定子元素使用 find_element()(单数),优先用 find_elements() + 列表判空;
- get_attribute("textContent") 可能返回空白或换行符,建议 .strip() 清洗;
- 若页面含懒加载内容,需滚动到底部并等待新容器渲染(browser.execute_script("window.scrollTo(0, document.body.scrollHeight);") + time.sleep(2));
- 生产环境应添加显式等待(WebDriverWait)替代 sleep,提升稳定性。
最终生成的 DataFrame 将完整覆盖所有有效新闻条目,且 links 列中缺失项明确为空字符串,便于后续分析或清洗。此模式适用于任何存在“非强制子元素”的网页结构,是 Selenium 爬虫健壮性的核心实践之一。










