
本文介绍如何利用 python 的 beautifulsoup 库,根据动态 html 文档中语义明确的标题文本(如“notes to unaudited condensed...”和“item 2.”),精准定位并提取二者之间的所有 html 元素,适用于财报、年报等结构松散但标题稳定的文档解析场景。
本文介绍如何利用 python 的 beautifulsoup 库,根据动态 html 文档中语义明确的标题文本(如“notes to unaudited condensed...”和“item 2.”),精准定位并提取二者之间的所有 html 元素,适用于财报、年报等结构松散但标题稳定的文档解析场景。
在处理非标准化的 HTML 报告(如 SEC 文件、企业财报)时,依赖固定标签层级或 class/id 属性往往不可靠——因为格式常随版本变化。更稳健的策略是基于可读性强、语义稳定的内容文本进行锚点定位。BeautifulSoup 提供了灵活的查找机制,配合状态机式遍历,即可实现“从 A 标题到 B 标题之间所有内容”的精确截取。
以下是一个生产就绪的解析方案,已针对原始需求优化:
from bs4 import BeautifulSoup
def extract_section_by_heading(html_content: str, start_text: str, end_text: str) -> list:
"""
从 HTML 字符串中提取 start_text 所在标签之后、end_text 所在标签之前的所有同级兄弟元素。
注意:此方法假设起始与结束标题位于同一 DOM 深度(如均为 <a><span> 内文本),且目标内容位于二者之间(非嵌套子树内)。
"""
soup = BeautifulSoup(html_content, "html.parser")
# 查找包含关键词的最内层文本节点(支持 span/a 等嵌套)
def find_heading_tag(text_hint):
for tag in soup.find_all(['a', 'span', 'h1', 'h2', 'h3', 'p', 'div']):
if text_hint in (tag.get_text(strip=True) or ""):
return tag
return None
tag_start = find_heading_tag(start_text)
tag_end = find_heading_tag(end_text)
if not tag_start:
raise ValueError(f"未找到起始标题:'{start_text}'")
if not tag_end:
raise ValueError(f"未找到结束标题:'{end_text}'")
# 获取共同父容器(确保在同一层级上下文中遍历)
parent = tag_start.parent if tag_start.parent == tag_end.parent else soup.body or soup.html
# 遍历父容器的直接子节点,收集 start 之后、end 之前的元素
result = []
capture = False
for child in parent.children:
if not hasattr(child, 'name') or not child.name: # 跳过 NavigableString(如换行、空格)
continue
if child is tag_start:
capture = True
continue
if child is tag_end:
break
if capture:
result.append(child)
return result
# 使用示例
html_sample = """
<div>无关内容1</div>
<div><a href="#a1NatureofOperations_790426"><span>Notes to Unaudited Condensed Consolidated Financial Statements</span></a></div>
<div><p>附注1:现金及现金等价物包括...</p><div class="aritcle_card flexRow artxards">
<div class="artcardd flexRow">
<a class="aritcle_card_img" rel="nofollow" href="/xiazai/skill7761" title="Article To Html"><img
src="https://img.php.cn/upload/skill/000/000/081/179168408018805.jpg" alt="Article To Html" onerror="this.onerror='';this.src='/static/lhimages/moren/morentu.png'" ></a>
<div class="aritcle_card_info flexColumn">
<a rel="nofollow" href="/xiazai/skill7761" title="Article To Html" class="overflowclass">Article To Html</a>
<p class="overflowclass">文章转信息图。将文章/笔记转化为手机可读的 HTML 信息图,自动匹配视觉风格。触发场景:文章转图、笔记转图、信息图、转小红书图、做张图、可视化这篇文章、文生图。</p>
</div>
<a rel="nofollow" href="/xiazai/skill7761" title="Article To Html" class="aritcle_card_btn flexRow flexcenter"><b></b><span>下载</span>
</a>
</div>
</div></div>
<div><table><tr><td>金额单位:百万美元</td></tr></table></div>
<div><a href="#ITEM2MANAGEMENTSDISCUSSIONANDANALYSIS_77"><span>Item 2.</span></a></div>
<div>无关内容2</div>
"""
section_elements = extract_section_by_heading(
html_sample,
start_text="Notes to Unaudited Condensed Consolidated Financial Statements",
end_text="Item 2."
)
for elem in section_elements:
print(elem.prettify())</span></a>
✅ 关键优势说明:
- ✅ 不依赖 ID 或 class:规避 HTML 动态生成导致的属性不可预测性;
- ✅ 容错文本匹配:使用 get_text(strip=True) 处理换行、空格、内联样式干扰;
- ✅ 层级鲁棒性:自动回溯至共同父级,避免因 和 嵌套深度不同而漏匹配;
- ✅ 可扩展性强:只需修改 start_text / end_text 即可适配其他章节(如 "Item 1.", "Significant Accounting Policies")。
⚠️ 注意事项:
- 若目标内容深度嵌套在起始/结束标签内部(例如 ...),上述方法仅捕获
- ...
- 需提取内容
- 整体;此时应改用 .find_next_siblings() 或递归提取子树;
- 对超大 HTML(>100MB),建议配合 lxml 解析器提升性能:BeautifulSoup(html_content, "lxml");
- 生产环境务必添加异常处理(如编码错误、空文档、标题缺失),并做日志记录便于调试。
掌握这种“语义锚点 + 状态流遍历”的模式,你将能高效应对各类非结构化 HTML 文档的定向抽取任务——无需正则硬解析,也无需强依赖 CSS 选择器,真正实现以内容为中心的智能提取。
前端入门到VUE实战笔记:立即使用
在学习笔记中,你将探索 前端 的入门与实战技巧!










