How to use python beautifulsoup4 module-Python Tutorial-php.cn

Home

Backend Development

Python Tutorial

How to use python beautifulsoup4 module

王林

May 11, 2023 pm 10:31 PM

pythonbeautifulsoup4

1. Basic knowledge supplement of BeautifulSoup4

BeautifulSoup4 is a python parsing library, mainly used to parse HTML and XML. There will be more parsing of HTML in the crawler knowledge system,

The library installation command is as follows:

pip install beautifulsoup4

BeautifulSoup When parsing data, you need to rely on a third-party parser. Commonly used parsers and The advantages are as follows:

python standard library html.parser: python has a built-in standard library with strong fault tolerance;
lxml parser: fast, fault-tolerant;
html5lib: the most fault-tolerant, parsing method and browsing The device is consistent.

Next, use a custom HTML code to demonstrate the basic use of the beautifulsoup4 library. The test code is as follows:

<html>
  <head>
    <title>测试bs4模块脚本</title>
  </head>
  <body>
    <h2 id="橡皮擦的爬虫课">橡皮擦的爬虫课</h2>
    <p>用一段自定义的 HTML 代码来演示</p>
  </body>
</html>

Use BeautifulSoup Perform simple operations on it, including instantiating BS objects, outputting page tags, etc.

from bs4 import BeautifulSoup
text_str = """<html>
	<head>
		<title>测试bs4模块脚本</title>
	</head>
	<body>
		<h2 id="橡皮擦的爬虫课">橡皮擦的爬虫课</h2>
		<p>用1段自定义的 HTML 代码来演示</p>
		<p>用2段自定义的 HTML 代码来演示</p>
	</body>
</html>
"""
# 实例化 Beautiful Soup 对象
soup = BeautifulSoup(text_str, "html.parser")
# 上述是将字符串格式化为 Beautiful Soup 对象，你可以从一个文件进行格式化
# soup = BeautifulSoup(open(&#39;test.html&#39;))
print(soup)
# 输入网页标题 title 标签
print(soup.title)
# 输入网页 head 标签
print(soup.head)

# 测试输入段落标签 p
print(soup.p) # 默认获取第一个

We can directly call the web page tag through the BeautifulSoup object. There is a problem here. Calling the tag through the BS object can only get the tag ranked first. As in the above code, only one is obtained p tag, if you want to get more content, please continue reading.

To learn this, we need to understand the 4 built-in objects in BeautifulSoup:

BeautifulSoup: basic object, The entire HTML object can generally be viewed as a Tag object;
Tag: tag object, tags are each node in the web page, such as title, head, p;
NavigableString: tag internal string;
Comment: comment object, inside the crawler There are not many usage scenarios.

The following code demonstrates for you the scenarios in which these objects appear. Pay attention to the relevant comments in the code:

from bs4 import BeautifulSoup
text_str = """<html>
	<head>
		<title>测试bs4模块脚本</title>
	</head>
	<body>
		<h2 id="橡皮擦的爬虫课">橡皮擦的爬虫课</h2>
		<p>用1段自定义的 HTML 代码来演示</p>
		<p>用2段自定义的 HTML 代码来演示</p>
	</body>
</html>
"""
# 实例化 Beautiful Soup 对象
soup = BeautifulSoup(text_str, "html.parser")
# 上述是将字符串格式化为 Beautiful Soup 对象，你可以从一个文件进行格式化
# soup = BeautifulSoup(open(&#39;test.html&#39;))
print(soup)
print(type(soup))  # <class &#39;bs4.BeautifulSoup&#39;>
# 输入网页标题 title 标签
print(soup.title)
print(type(soup.title)) # <class &#39;bs4.element.Tag&#39;>
print(type(soup.title.string)) # <class &#39;bs4.element.NavigableString&#39;>
# 输入网页 head 标签
print(soup.head)

For Tag object has two important attributes, which are name and attrs

from bs4 import BeautifulSoup
text_str = """<html>
	<head>
		<title>测试bs4模块脚本</title>
	</head>
	<body>
		<h2 id="橡皮擦的爬虫课">橡皮擦的爬虫课</h2>
		<p>用1段自定义的 HTML 代码来演示</p>
		<p>用2段自定义的 HTML 代码来演示</p>
		<a href="http://www.csdn.net" rel="external nofollow"  rel="external nofollow" >CSDN 网站</a>
	</body>
</html>
"""
# 实例化 Beautiful Soup 对象
soup = BeautifulSoup(text_str, "html.parser")
print(soup.name) # [document]
print(soup.title.name) # 获取标签名 title
print(soup.html.body.a) # 可以通过标签层级获取下层标签
print(soup.body.a) # html 作为一个特殊的根标签，可以省略
print(soup.p.a) # 无法获取到 a 标签
print(soup.a.attrs) # 获取属性

The above code demonstrates obtaining Usage of name attribute and attrs attribute. The attrs attribute is a dictionary, and the corresponding value can be obtained by key.

Get the attribute value of the tag. In BeautifulSoup, you can also use the following method:

print(soup.a["href"])
print(soup.a.get("href"))

Get NavigableString Object After getting the web page tag, To get the text within the label, use the following code.

print(soup.a.string)

In addition, you can also use the text attribute and the get_text() method to get the tag content.

print(soup.a.string)
print(soup.a.text)
print(soup.a.get_text())

You can also get all the text in the tag by using strings and stripped_strings.

print(list(soup.body.strings)) # 获取到空格或者换行
print(list(soup.body.stripped_strings)) # 去除空格或者换行

Extended tag/node selector to traverse the document tree

Direct child node

The direct child element of the tag (Tag) object can be used contents and children attributes are obtained.

from bs4 import BeautifulSoup
text_str = """<html>
	<head>
		<title>测试bs4模块脚本</title>
	</head>
	<body>
		<div id="content">
			<h2 id="橡皮擦的爬虫课-span-最棒-span">橡皮擦的爬虫课<span>最棒</span></h2>
            <p>用1段自定义的 HTML 代码来演示</p>
            <p>用2段自定义的 HTML 代码来演示</p>
            <a href="http://www.csdn.net" rel="external nofollow"  rel="external nofollow" >CSDN 网站</a>
		</div>
        <ul class="nav">
            <li>首页</li>
            <li>博客</li>
            <li>专栏课程</li>
        </ul>

	</body>
</html>
"""
# 实例化 Beautiful Soup 对象
soup = BeautifulSoup(text_str, "html.parser")
# contents 属性获取节点的直接子节点，以列表的形式返回内容
print(soup.div.contents) # 返回列表
# children 属性获取的也是节点的直接子节点，以生成器的类型返回
print(soup.div.children) # 返回 <list_iterator object at 0x00000111EE9B6340>

Please note that the above two attributes obtain direct child nodes, such as the descendant tag span within the h2 tag, which will not obtained separately.

If you want to get all tags, use the descendants attribute, which returns a generator, and all tags including the text within the tags will be fetched separately.

print(list(soup.div.descendants))

Acquisition of other nodes (just understand it, check it and use it immediately)

parent and parents: directly Parent node and all parent nodes;
next_sibling, next_siblings, previous_sibling, previous_siblings : Represents the next sibling node, all sibling nodes below, the previous sibling node, and all sibling nodes above. Since the newline character is also a node, when using these attributes, pay attention to the newline character;
next_element, next_elements, previous_element, previous_elements: These attributes represent the previous node or the next node respectively. A node, note that they are not hierarchical, but for all nodes. For example, the next node of the div node in the above code is h2, and the div node The sibling node is ul.

Document tree search related functions

The first function to learn is the find_all() function, The prototype is as follows:

find_all(name,attrs,recursive,text,limit=None,**kwargs)

name: This parameter is the name of the tag tag, for example find_all('p') is to find all p tags, and can accept tag name strings, regular expressions and lists;
attrs：传入的属性，该参数可以字典的形式传入，例如 attrs={'class': 'nav'}，返回的结果是 tag 类型的列表；

上述两个参数的用法示例如下：

print(soup.find_all(&#39;li&#39;)) # 获取所有的 li
print(soup.find_all(attrs={&#39;class&#39;: &#39;nav&#39;})) # 传入 attrs 属性
print(soup.find_all(re.compile("p"))) # 传递正则，实测效果不理想
print(soup.find_all([&#39;a&#39;,&#39;p&#39;])) # 传递列表

recursive：调用 find_all () 方法时，BeautifulSoup 会检索当前 tag 的所有子孙节点，如果只想搜索 tag 的直接子节点，可以使用参数 recursive=False，测试代码如下：

print(soup.body.div.find_all([&#39;a&#39;,&#39;p&#39;],recursive=False)) # 传递列表

text：可以检索文档中的文本字符串内容，与 name 参数的可选值一样，text 参数接受标签名字符串、正则表达式、列表；

print(soup.find_all(text=&#39;首页&#39;)) # [&#39;首页&#39;]
print(soup.find_all(text=re.compile("^首"))) # [&#39;首页&#39;]
print(soup.find_all(text=["首页",re.compile(&#39;课&#39;)])) # [&#39;橡皮擦的爬虫课&#39;, &#39;首页&#39;, &#39;专栏课程&#39;]

limit：可以用来限制返回结果的数量；
kwargs：如果一个指定名字的参数不是搜索内置的参数名，搜索时会把该参数当作 tag 的属性来搜索。这里要按 class 属性搜索，因为 class 是 python 的保留字，需要写作 class_，按 class_ 查找时，只要一个 CSS 类名满足即可，如需多个 CSS 名称，填写顺序需要与标签一致。

print(soup.find_all(class_ = &#39;nav&#39;))
print(soup.find_all(class_ = &#39;nav li&#39;))

还需要注意网页节点中，有些属性在搜索中不能作为kwargs参数使用，比如html5 中的 data-*属性，需要通过attrs参数进行匹配。

与 find_all()方法用户基本一致的其它方法清单如下：

find()：函数原型find( name , attrs , recursive , text , **kwargs )，返回一个匹配元素；
find_parents()，find_parent()：函数原型 find_parent(self, name=None, attrs={}, **kwargs)，返回当前节点的父级节点；
find_next_siblings()，find_next_sibling()：函数原型 find_next_sibling(self, name=None, attrs={}, text=None, **kwargs)，返回当前节点的下一兄弟节点；
find_previous_siblings()，find_previous_sibling()：同上，返回当前的节点的上一兄弟节点；
find_all_next()，find_next()，find_all_previous () ，find_previous ()：函数原型 find_all_next(self, name=None, attrs={}, text=None, limit=None, **kwargs)，检索当前节点的后代节点。

CSS 选择器 该小节的知识点与pyquery有点撞车，核心使用select()方法即可实现，返回数据是列表元组。

通过标签名查找，soup.select("title")；
通过类名查找，soup.select(".nav")；
通过 id 名查找，soup.select("#content")；
通过组合查找，soup.select("div#content")；
通过属性查找，soup.select("div[id='content'")，soup.select("a[href]")；

在通过属性查找时，还有一些技巧可以使用，例如：

^=：可以获取以 XX 开头的节点：

print(soup.select(&#39;ul[class^="na"]&#39;))

*=：获取属性包含指定字符的节点：

print(soup.select(&#39;ul[class*="li"]&#39;))

二、爬虫案例

BeautifulSoup 的基础知识掌握之后，在进行爬虫案例的编写，就非常简单了，本次要采集的目标网站，该目标网站有大量的艺术二维码，可以供设计大哥做参考。

How to use python beautifulsoup4 module

下述应用到了 BeautifulSoup 模块的标签检索与属性检索，完整代码如下：

from bs4 import BeautifulSoup
import requests
import logging
logging.basicConfig(level=logging.NOTSET)
def get_html(url, headers) -> None:
    try:
        res = requests.get(url=url, headers=headers, timeout=3)
    except Exception as e:
        logging.debug("采集异常", e)

    if res is not None:
        html_str = res.text
        soup = BeautifulSoup(html_str, "html.parser")
        imgs = soup.find_all(attrs={&#39;class&#39;: &#39;lazy&#39;})
        print("获取到的数据量是", len(imgs))
        datas = []
        for item in imgs:
            name = item.get(&#39;alt&#39;)
            src = item["src"]
            logging.info(f"{name},{src}")
            # 获取拼接数据
            datas.append((name, src))
        save(datas, headers)
def save(datas, headers) -> None:
    if datas is not None:
        for item in datas:
            try:
                # 抓取图片
                res = requests.get(url=item[1], headers=headers, timeout=5)
            except Exception as e:
                logging.debug(e)

            if res is not None:
                img_data = res.content
                with open("./imgs/{}.jpg".format(item[0]), "wb+") as f:
                    f.write(img_data)
    else:
        return None
if __name__ == &#39;__main__&#39;:
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.82 Safari/537.36"
    }
    url_format = "http://www.9thws.com/#p{}"
    urls = [url_format.format(i) for i in range(1, 2)]
    get_html(urls[0], headers)

本次代码测试输出采用的 logging 模块实现，效果如下图所示。测试仅采集了 1 页数据，如需扩大采集范围，只需要修改 main 函数内页码规则即可。 ==代码编写过程中，发现数据请求是类型是 POST，数据返回格式是 JSON，所以本案例仅作为 BeautifulSoup 的上手案例吧==

How to use python beautifulsoup4 module

The above is the detailed content of How to use python beautifulsoup4 module. For more information, please follow other related articles on the PHP Chinese website!

Statement

This article is reproduced at:亿速云. If there is any infringement, please contact admin@php.cn delete

Python: A Deep Dive into Compilation and InterpretationMay 12, 2025 am 12:14 AM

Pythonusesahybridmodelofcompilationandinterpretation:1)ThePythoninterpretercompilessourcecodeintoplatform-independentbytecode.2)ThePythonVirtualMachine(PVM)thenexecutesthisbytecode,balancingeaseofusewithperformance.

Is Python an interpreted or a compiled language, and why does it matter?May 12, 2025 am 12:09 AM

Pythonisbothinterpretedandcompiled.1)It'scompiledtobytecodeforportabilityacrossplatforms.2)Thebytecodeistheninterpreted,allowingfordynamictypingandrapiddevelopment,thoughitmaybeslowerthanfullycompiledlanguages.

For Loop vs While Loop in Python: Key Differences ExplainedMay 12, 2025 am 12:08 AM

Forloopsareidealwhenyouknowthenumberofiterationsinadvance,whilewhileloopsarebetterforsituationswhereyouneedtoloopuntilaconditionismet.Forloopsaremoreefficientandreadable,suitableforiteratingoversequences,whereaswhileloopsoffermorecontrolandareusefulf

For and While loops: a practical guideMay 12, 2025 am 12:07 AM

Forloopsareusedwhenthenumberofiterationsisknowninadvance,whilewhileloopsareusedwhentheiterationsdependonacondition.1)Forloopsareidealforiteratingoversequenceslikelistsorarrays.2)Whileloopsaresuitableforscenarioswheretheloopcontinuesuntilaspecificcond

Python: Is it Truly Interpreted? Debunking the MythsMay 12, 2025 am 12:05 AM

Pythonisnotpurelyinterpreted;itusesahybridapproachofbytecodecompilationandruntimeinterpretation.1)Pythoncompilessourcecodeintobytecode,whichisthenexecutedbythePythonVirtualMachine(PVM).2)Thisprocessallowsforrapiddevelopmentbutcanimpactperformance,req

Python concatenate lists with same elementMay 11, 2025 am 12:08 AM

ToconcatenatelistsinPythonwiththesameelements,use:1)the operatortokeepduplicates,2)asettoremoveduplicates,or3)listcomprehensionforcontroloverduplicates,eachmethodhasdifferentperformanceandorderimplications.

Interpreted vs Compiled Languages: Python's PlaceMay 11, 2025 am 12:07 AM

Pythonisaninterpretedlanguage,offeringeaseofuseandflexibilitybutfacingperformancelimitationsincriticalapplications.1)InterpretedlanguageslikePythonexecuteline-by-line,allowingimmediatefeedbackandrapidprototyping.2)CompiledlanguageslikeC/C transformt

For and While loops: when do you use each in python?May 11, 2025 am 12:05 AM

Useforloopswhenthenumberofiterationsisknowninadvance,andwhileloopswheniterationsdependonacondition.1)Forloopsareidealforsequenceslikelistsorranges.2)Whileloopssuitscenarioswheretheloopcontinuesuntilaspecificconditionismet,usefulforuserinputsoralgorit

See all articles

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

Video Face Swap

Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

Roblox: Grow A Garden - Complete Mutation Guide

3 weeks agoByDDD

Roblox: Bubble Gum Simulator Infinity - How To Get And Use Royal Keys

3 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

How to fix KB5055612 fails to install in Windows 10?

3 weeks agoByDDD

Nordhold: Fusion System, Explained

3 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Mandragora: Whispers Of The Witch Tree - How To Unlock The Grappling Hook

3 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Hot Tools

SublimeText3 Chinese version

Chinese version, very easy to use

mPDF

mPDF is a PHP library that can generate PDF files from UTF-8 encoded HTML. The original author, Ian Back, wrote mPDF to output PDF files "on the fly" from his website and handle different languages. It is slower than original scripts like HTML2FPDF and produces larger files when using Unicode fonts, but supports CSS styles etc. and has a lot of enhancements. Supports almost all languages, including RTL (Arabic and Hebrew) and CJK (Chinese, Japanese and Korean). Supports nested block-level elements (such as P, DIV),

SecLists

SecLists is the ultimate security tester's companion. It is a collection of various types of lists that are frequently used during security assessments, all in one place. SecLists helps make security testing more efficient and productive by conveniently providing all the lists a security tester might need. List types include usernames, passwords, URLs, fuzzing payloads, sensitive data patterns, web shells, and more. The tester can simply pull this repository onto a new test machine and he will have access to every type of list he needs.

MinGW - Minimalist GNU for Windows

This project is in the process of being migrated to osdn.net/projects/mingw, you can continue to follow us there. MinGW: A native Windows port of the GNU Compiler Collection (GCC), freely distributable import libraries and header files for building native Windows applications; includes extensions to the MSVC runtime to support C99 functionality. All MinGW software can run on 64-bit Windows platforms.

SAP NetWeaver Server Adapter for Eclipse

Integrate Eclipse with SAP NetWeaver application server.

Hot Topics

1665

1424

1321

1269

1249