詳解Python利用Beautiful Soup模組搜尋內容方法-Python教學-PHP中文網

首頁

後端開發

Python教學

詳解Python利用Beautiful Soup模組搜尋內容方法

高洛峰

Mar 31, 2017 am 09:47 AM

這篇文章主要為大家介紹了python中 Beautiful Soup 模組的搜尋方法函數。方法不同類型的過濾參數能夠進行不同的過濾，得到想要的結果。文中介紹的非常詳細，對大家有一定的參考價值，需要的朋友們下面來一起看看吧。

前言

我們將利用Beautiful Soup 模組的搜尋功能，根據標籤名稱、標籤屬性、文件文字和正規表示式來搜尋。

搜尋方法

Beautiful Soup 內建的搜尋方法如下：

find()
find_all()
find_parent()
find_parents()
find_next_sibling()
find_next_siblings()
find_previous_sibling()
#find_previous_siblings()
find_previous()
find_all_previous()
find_next()
find_all_next()

# #首先還是需要建立一個HTML 檔案用來做測試。

<html>
<body>
<p class="ecopyramid">
 <ul id="producers">
 <li class="producerlist">
  <p class="name">plants</p>
  <p class="number">100000</p>
 </li>
 <li class="producerlist">
  <p class="name">algae</p>
  <p class="number">100000</p>
 </li>
 </ul>
 <ul id="primaryconsumers">
 <li class="primaryconsumerlist">
  <p class="name">deer</p>
  <p class="number">1000</p>
 </li>
 <li class="primaryconsumerlist">
  <p class="name">rabbit</p>
  <p class="number">2000</p>
 </li>
 </ul>
 <ul id="secondaryconsumers">
 <li class="secondaryconsumerlist">
  <p class="name">fox</p>
  <p class="number">100</p>
 </li>
 <li class="secondaryconsumerlist">
  <p class="name">bear</p>
  <p class="number">100</p>
 </li>
 </ul>
 <ul id="tertiaryconsumers">
 <li class="tertiaryconsumerlist">
  <p class="name">lion</p>
  <p class="number">80</p>
 </li>
 <li class="tertiaryconsumerlist">
  <p class="name">tiger</p>
  <p class="number">50</p>
 </li>
 </ul>
</p>
</body>
</html>

我們可以透過

find()

方法來獲得

標籤，預設還是會得到第一個出現的，接著取得
標籤，透過輸出內容來驗證是否取得了第一個出現的標籤。
```
from bs4 import BeautifulSoup
with open(&#39;search.html&#39;,&#39;r&#39;) as filename:
 soup = BeautifulSoup(filename,&#39;lxml&#39;)
first_ul_entries = soup.find(&#39;ul&#39;)
print first_ul_entries.li.p.string
```
find() 方法具體如下：
```
find(name,attrs,recursive,text,**kwargs)
```
如上程式碼所示，
find( )
方法接受五個參數：name、attrs、recursive、text 和**kwargs 。 name 、attrs 和 text 參數都可以在
find()
方法充當過濾器，提高匹配結果的精確度。
搜尋標籤

除了上面程式碼的搜尋
- 標籤，回傳結果也是回傳出現的第一個符合內容。
```
tag_li = soup.find(&#39;li&#39;)
# tag_li = soup.find(name = "li")
print type(tag_li)
print tag_li.p.string
```
 搜尋文字
 
 #如果我們只想根據文字內容來搜尋的話，我們可以只傳入文字參數：
```
search_for_text = soup.find(text=&#39;plants&#39;)
print type(search_for_text)
<class &#39;bs4.element.NavigableString&#39;>
```
 傳回的結果也是NavigableString 物件。
 
 根據正規表示式搜尋
 
 如下的一段HTML 文字內容
```
The below HTML has the information that has email ids.
 abc@example.com 
xyz@example.com 
 foo@example.com
```
 - 在使用正規表示式進行比對時，如果有多個符合項，也是先傳回第一個。
 可以透過標籤的屬性值來搜尋：
```
search_for_attribute = soup.find(id=&#39;primaryconsumers&#39;)
print search_for_attribute.li.p.string
```
 根據標籤屬性值來搜尋對大多數屬性都是可用的，例如：id、style 和title 。
 
 但對以下兩種情況會有不同：
 
 自訂屬性
 
 類別( class )屬性
 
 我們不能再直接使用屬性值來搜尋了，而是得使用attrs 參數來傳遞給find()
 函數。
 
 根據自訂屬性來搜尋
 
 在 HTML5 中是可以為標籤新增自訂屬性的，例如為標籤新增屬性。
 
 如下程式碼所示，如果我們再像搜尋 id 那樣進行操作的話，會報錯的，Python 的變數不能包括 - 符號。
```
customattr = """
 custom attribute example
 """
customsoup = BeautifulSoup(customattr,&#39;lxml&#39;)
customsoup.find(data-custom="custom")
# SyntaxError: keyword can&#39;t be an expression
```
 這個時候使用attrs 屬性值來傳遞一個字典類型作為參數進行搜尋：
```
using_attrs = customsoup.find(attrs={&#39;data-custom&#39;:&#39;custom&#39;})
print using_attrs
```
 #基於CSS 中的類別進行搜尋
 
 對於CSS 的類別屬性，由於在Python 中class 是個關鍵字，所以是不能當做標籤屬性參數傳遞的，這種情況下，就和自定義屬性一樣進行搜尋。也是使用 attrs 屬性，傳遞一個字典來配對。
 
 除了使用 attrs 屬性之外，還可以使用 class_ 屬性進行傳遞，這與 class 區別開了，也不會導致錯誤。
```
css_class = soup.find(attrs={&#39;class&#39;:&#39;producerlist&#39;})
css_class2 = soup.find(class_ = "producerlist")
print css_class
print css_class2
```
 ###使用自訂的函數搜尋#########可以給###find() ###方法傳遞一個函數，這樣就會根據函數定義的條件進行搜尋。 #########函數應該會傳回 true 或是 false 值。 ############
```
def is_producers(tag):
 return tag.has_attr(&#39;id&#39;) and tag.get(&#39;id&#39;) == &#39;producers&#39;
tag_producers = soup.find(is_producers)
print tag_producers.li.p.string
```
 ###程式碼中定義了一個 is_producers 函數，它將檢查標籤是否具體 id 屬性以及屬性值是否等於 producers，如果符合條件則傳回 true ，否則傳回 false 。 #########聯合使用各種搜尋方法#########Beautiful Soup 提供了各種搜尋方法，同樣，我們也可以聯合使用這些方法來進行匹配，提高搜尋的準確度。 ###
```
combine_html = """
 
 Example of p tag with class identical
 
 
 Example of p tag with class identical
 
 """
combine_soup = BeautifulSoup(combine_html,&#39;lxml&#39;)
identical_p = combine_soup.find("p",class_="identical")
print identical_p
```
 使用 find_all() 方法搜索
 
 使用 find() 方法会从搜索结果中返回第一个匹配的内容，而 find_all() 方法则会返回所有匹配的项。
 
 在 find() 方法中用到的过滤项，同样可以用在 find_all() 方法中。事实上，它们可以用到任何搜索方法中，例如：find_parents() 和 find_siblings() 中。
```
# 搜索所有 class 属性等于 tertiaryconsumerlist 的标签。
all_tertiaryconsumers = soup.find_all(class_=&#39;tertiaryconsumerlist&#39;)
print type(all_tertiaryconsumers)
for tertiaryconsumers in all_tertiaryconsumers:
 print tertiaryconsumers.p.string
```
 find_all() 方法为：
```
find_all(name,attrs,recursive,text,limit,**kwargs)
```
 它的参数和 find() 方法有些类似，多个了 limit 参数。limit 参数是用来限制结果数量的。而 find() 方法的 limit 就是 1 了。
 
 同时，我们也能传递一个字符串列表的参数来搜索标签、标签属性值、自定义属性值和 CSS 类。
```
# 搜索所有的 p 和 li 标签
p_li_tags = soup.find_all(["p","li"])
print p_li_tags
print
# 搜索所有类属性是 producerlist 和 primaryconsumerlist 的标签
all_css_class = soup.find_all(class_=["producerlist","primaryconsumerlist"])
print all_css_class
print
```
 搜索相关标签
 
 一般情况下，我们可以使用 find() 和 find_all() 方法来搜索指定的标签，同时也能搜索其他与这些标签相关的感兴趣的标签。
 
 搜索父标签
 
 可以使用 find_parent() 或者 find_parents() 方法来搜索标签的父标签。
 
 find_parent() 方法将返回第一个匹配的内容，而 find_parents() 将返回所有匹配的内容，这一点与 find() 和 find_all() 方法类似。
```
# 搜索 父标签
primaryconsumers = soup.find_all(class_=&#39;primaryconsumerlist&#39;)
print len(primaryconsumers)
# 取父标签的第一个
primaryconsumer = primaryconsumers[0]
# 搜索所有 ul 的父标签
parent_ul = primaryconsumer.find_parents(&#39;ul&#39;)
print len(parent_ul)
# 结果将包含父标签的所有内容
print parent_ul
print
# 搜索,取第一个出现的父标签.有两种操作
immediateprimary_consumer_parent = primaryconsumer.find_parent()
# immediateprimary_consumer_parent = primaryconsumer.find_parent(&#39;ul&#39;)
print immediateprimary_consumer_parent
```
 搜索同级标签
 
 Beautiful Soup 还提供了搜索同级标签的功能。
 
 使用函数 find_next_siblings() 函数能够搜索同一级的下一个所有标签，而 find_next_sibling() 函数能够搜索同一级的下一个标签。
```
producers = soup.find(id=&#39;producers&#39;)
next_siblings = producers.find_next_siblings()
print next_siblings
```
 同样，也可以使用 find_previous_siblings() 和 find_previous_sibling() 方法来搜索上一个同级的标签。
 
 搜索下一个标签
 
 使用 find_next() 方法将搜索下一个标签中第一个出现的，而 find_next_all() 将会返回所有下级的标签项。
```
# 搜索下一级标签
first_p = soup.p
all_li_tags = first_p.find_all_next("li")
print all_li_tags
```
 搜索上一个标签
 
 与搜索下一个标签类似，使用 find_previous() 和 find_all_previous() 方法来搜索上一个标签。

以上是詳解Python利用Beautiful Soup模組搜尋內容方法的詳細內容。更多資訊請關注PHP中文網其他相關文章！

陳述

本文內容由網友自願投稿，版權歸原作者所有。本站不承擔相應的法律責任。如發現涉嫌抄襲或侵權的內容，請聯絡admin@php.cn

Python中的合併列表：選擇正確的方法May 14, 2025 am 12:11 AM

Tomergelistsinpython，YouCanusethe操作員，estextMethod，ListComprehension，Oritertools

如何在Python 3中加入兩個列表？May 14, 2025 am 12:09 AM

在Python3中，可以通過多種方法連接兩個列表：1)使用運算符，適用於小列表，但對大列表效率低；2)使用extend方法，適用於大列表，內存效率高，但會修改原列表；3)使用*運算符，適用於合併多個列表，不修改原列表；4)使用itertools.chain，適用於大數據集，內存效率高。

Python串聯列表字符串May 14, 2025 am 12:08 AM

使用join()方法是Python中從列表連接字符串最有效的方法。 1)使用join()方法高效且易讀。 2)循環使用運算符對大列表效率低。 3)列表推導式與join()結合適用於需要轉換的場景。 4)reduce()方法適用於其他類型歸約，但對字符串連接效率低。完整句子結束。

Python執行，那是什麼？May 14, 2025 am 12:06 AM

pythonexecutionistheprocessoftransformingpypythoncodeintoExecutablestructions.1）InternterPreterReadSthecode，ConvertingTingitIntObyTecode，whepythonvirtualmachine（pvm）theglobalinterpreterpreterpreterpreterlock（gil）the thepythonvirtualmachine（pvm）

Python：關鍵功能是什麼May 14, 2025 am 12:02 AM

Python的關鍵特性包括：1.語法簡潔易懂，適合初學者；2.動態類型系統，提高開發速度；3.豐富的標準庫，支持多種任務；4.強大的社區和生態系統，提供廣泛支持；5.解釋性，適合腳本和快速原型開發；6.多範式支持，適用於各種編程風格。

Python：編譯器還是解釋器？May 13, 2025 am 12:10 AM

Python是解釋型語言，但也包含編譯過程。 1）Python代碼先編譯成字節碼。 2）字節碼由Python虛擬機解釋執行。 3）這種混合機制使Python既靈活又高效，但執行速度不如完全編譯型語言。

python用於循環與循環時：何時使用哪個？May 13, 2025 am 12:07 AM

UseeAforloopWheniteratingOveraseQuenceOrforAspecificnumberoftimes; useAwhiLeLoopWhenconTinuingUntilAcIntiment.forloopsareIdealForkNownsences，而WhileLeleLeleLeleLeleLoopSituationSituationsItuationsItuationSuationSituationswithUndEtermentersitations。