Beautiful Soup
表示
| 作者 | Leonard Richardson |
|---|---|
| 初版 | 2004年 |
| 最新版 | |
| リポジトリ | |
| プログラミング 言語 | Python |
| プラットフォーム | Python |
| 種別 | HTMLパーサ、ウェブスクレイピング |
| ライセンス |
|
| 公式サイト |
www |
Beautiful Soup は、HTMLやXMLといったマークアップ言語の文書を構文解析するためのPythonパッケージである。ドキュメントから構築された解析木はウェブスクレイピングに有用である[3][4]。
Beautiful Soupはレオナルド・リチャードソンによって開始された。彼はプロジェクトへの貢献を続けている[5]。また、オープンソースソフトウェアを管理するTideliftによっても支援されている[6]。
コード例
[編集]Beautiful Soupは、通常のPythonのループで探索及び反復処理が可能な木構造として、パースしたデータを表現する[7]。
以下の例では、Pythonの標準ライブラリである urllib[8] を用いて、ウィキペディアのメインページを読み込み、Beautiful Soupで構文解析し、全てのハイパーリンクを得る。
#!/usr/bin/env python3
# HTML文書からのハイパーリンクの抽出
from bs4 import BeautifulSoup
from urllib.request import urlopen
with urlopen('https://en.wikipedia.org/wiki/Main_Page') as response:
soup = BeautifulSoup(response, 'html.parser')
for anchor in soup.find_all('a'):
print(anchor.get('href', '/'))
次の例では、PythonのRequestsライブラリ[9]を使用してページを取得し、全div要素を抽出している。
import requests
from bs4 import BeautifulSoup
url = "https://wikipedia.org"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
headings = soup.find_all("div")
for heading in headings:
print(heading.text.strip())
歴史
[編集]Beautiful Soupは不思議の国のアリス[10]の詩と構造が粗末なHTMLコードを意味するtag soup[11]の両方にちなんで名づけられた。
2006年4月から2012年3月までは、Beautiful Soup 3 がリリースされていた。 最新版の Beautiful Soup 4.x はpip install beautifulsoup4からインストールできる。
2021年に、Python 2.7 のサポートが終了し、 Beautiful Soup 4.9.3 がPython 2.7をサポートする最後のバージョンとなった[12]。
脚注
[編集]- ↑ https://git.launchpad.net/beautifulsoup/tree/CHANGELOG.
{{cite web2}}:|title=は必須です。 (説明); Cite webテンプレートでは|access-date=引数が必須です。 (説明) - ↑ “Beautiful Soup website”. 2012年4月18日閲覧。 “Beautiful Soup is licensed under the same terms as Python itself”
- ↑ Hajba, Gábor László (2018), Hajba, Gábor László, ed., “Using Beautiful Soup” (英語), Website Scraping with Python: Using BeautifulSoup and Scrapy (Apress): 41–96, doi:10.1007/978-1-4842-3925-4_3, ISBN 978-1-4842-3925-4
- ↑ Python. “Beautiful Soup: Build a Web Scraper With Python – Real Python” (英語). realpython.com. 2023年6月1日閲覧。
- ↑ “Code : Leonard Richardson” (英語). Launchpad. 2020年9月19日閲覧。
- ↑ Tidelift. “beautifulsoup4 | pypi via the Tidelift Subscription” (英語). tidelift.com. 2020年9月19日閲覧。
- ↑ “How To Scrape Web Pages with Beautiful Soup and Python 3 | DigitalOcean” (英語). www.digitalocean.com. 2023年6月1日閲覧。
- ↑ Python. “Python's urllib.request for HTTP Requests – Real Python” (英語). realpython.com. 2023年6月1日閲覧。
- ↑ Blog, SerpApi (2024年3月5日). “Beautiful Soup: Web Scraping with Python” (英語). serpapi.com. 2024年6月27日閲覧。
- ↑ makcorps (2022年12月13日). “BeautifulSoup tutorial: Let's Scrape Web Pages with Python” (英語). 2024年1月24日閲覧。
- ↑ “Python Web Scraping” (英語). Udacity (2021年2月11日). 2024年1月24日閲覧。
- ↑ Richardson (2021年9月7日). “Beautiful Soup 4.10.0” (英語). beautifulsoup. Google Groups. 2022年9月27日閲覧。