Scrapy 教程
在本教程中,我们假设 Scrapy 已安装在您的系统上。如果尚未安装,请参阅 安装指南。
我们将抓取 quotes.toscrape.com,这是一个列出著名作家引言的网站。
本教程将引导您完成以下任务
创建一个新的 Scrapy 项目
编写一个 spider 来爬取网站并提取数据
使用命令行导出抓取到的数据
修改 spider 以递归跟踪链接
使用 spider 参数
Scrapy 是用 Python 编写的。您对 Python 了解得越多,就能从 Scrapy 中获得越多。
如果您已经熟悉其他语言并想快速学习 Python,Python 教程是一个很好的资源。
如果您是编程新手并想从 Python 开始,以下书籍可能对您有用
您还可以查阅这份面向非程序员的 Python 资源列表,以及 learnpython-subreddit 中推荐的资源。
创建项目
在开始抓取之前,您需要设置一个新的 Scrapy 项目。进入您希望存储代码的目录并运行
scrapy startproject tutorial
这将创建一个包含以下内容的 tutorial 目录
tutorial/
scrapy.cfg # deploy configuration file
tutorial/ # project's Python module, you'll import your code from here
__init__.py
items.py # project items definition file
middlewares.py # project middlewares file
pipelines.py # project pipelines file
settings.py # project settings file
spiders/ # a directory where you'll later put your spiders
__init__.py
我们的第一个 Spider
Spider 是您定义的类,Scrapy 用它们从网站(或一组网站)抓取信息。它们必须继承 Spider 并定义要发出的初始请求,还可以选择定义如何跟踪页面中的链接以及解析下载的页面内容以提取数据。
这是我们第一个 Spider 的代码。将其保存为名为 quotes_spider.py 的文件,放在您项目中的 tutorial/spiders 目录下
from pathlib import Path
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
async def start(self):
urls = [
"https://quotes.toscrape.com/page/1/",
"https://quotes.toscrape.com/page/2/",
]
for url in urls:
yield scrapy.Request(url=url, callback=self.parse)
def parse(self, response):
page = response.url.split("/")[-2]
filename = f"quotes-{page}.html"
Path(filename).write_bytes(response.body)
self.log(f"Saved file {filename}")
如您所见,我们的 Spider 继承自 scrapy.Spider 并定义了一些属性和方法
name:标识 Spider。它在项目中必须是唯一的,也就是说,您不能为不同的 Spider 设置相同的名称。start():必须是一个异步生成器,用于生成请求(以及可选的 item),供 spider 开始爬取。后续请求将从这些初始请求中陆续生成。parse():一个方法,用于处理为每个发出的请求下载的响应。`response` 参数是TextResponse的一个实例,它包含页面内容并具有更多有用的处理方法。parse()方法通常会解析响应,将抓取到的数据提取为字典,并找到要跟踪的新 URL 并从中创建新的请求(Request)。
如何运行我们的 spider
要让我们的 spider 工作,请进入项目的顶级目录并运行
scrapy crawl quotes
此命令运行我们刚添加的名为 quotes 的 spider,它将向 quotes.toscrape.com 域发送一些请求。您将看到类似于以下的输出
... (omitted for brevity)
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Spider opened
2016-12-16 21:24:05 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2016-12-16 21:24:05 [scrapy.extensions.telnet] DEBUG: Telnet console listening on 127.0.0.1:6023
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (404) <GET https://quotes.toscrape.com/robots.txt> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/2/> (referer: None)
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-1.html
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-2.html
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Closing spider (finished)
...
现在,检查当前目录中的文件。您应该注意到已创建了两个新文件:quotes-1.html 和 quotes-2.html,其中包含相应 URL 的内容,正如我们的 parse 方法所指示的。
注意
如果您想知道为什么我们还没有解析 HTML,请稍等,我们很快就会讲到。
底层发生了什么?
Scrapy 发送由 start() spider 方法生成的第一个 scrapy.Request 对象。在收到每个请求的响应后,Scrapy 会调用与该请求关联的回调方法(在此示例中为 parse 方法),并传入一个 Response 对象。
对 start 方法的快捷方式
除了实现一个 start() 方法来从 URL 生成 Request 对象,您还可以定义一个 start_urls 类属性,其中包含一个 URL 列表。然后,这个列表将被 start() 的默认实现用来为您的 spider 创建初始请求。
from pathlib import Path
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
"https://quotes.toscrape.com/page/2/",
]
def parse(self, response):
page = response.url.split("/")[-2]
filename = f"quotes-{page}.html"
Path(filename).write_bytes(response.body)
即使我们没有明确告诉 Scrapy,parse() 方法也会被调用来处理这些 URL 的每个请求。之所以会这样,是因为 parse() 是 Scrapy 的默认回调方法,它会在没有明确指定回调的请求中被调用。
提取数据
学习如何使用 Scrapy 提取数据的最佳方法是使用 Scrapy shell 尝试选择器。运行
scrapy shell 'https://quotes.toscrape.com/page/1/'
注意
请记住,从命令行运行 Scrapy shell 时,务必将 URL 用引号括起来,否则包含参数(即 & 字符)的 URL 将无法工作。
在 Windows 上,请改用双引号
scrapy shell "https://quotes.toscrape.com/page/1/"
您将看到类似以下内容
[ ... Scrapy log here ... ]
2016-09-19 12:09:27 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
[s] Available Scrapy objects:
[s] scrapy scrapy module (contains scrapy.Request, scrapy.Selector, etc)
[s] crawler <scrapy.crawler.Crawler object at 0x7fa91d888c90>
[s] item {}
[s] request <GET https://quotes.toscrape.com/page/1/>
[s] response <200 https://quotes.toscrape.com/page/1/>
[s] settings <scrapy.settings.Settings object at 0x7fa91d888c10>
[s] spider <DefaultSpider 'default' at 0x7fa91c8af990>
[s] Useful shortcuts:
[s] shelp() Shell help (print this help)
[s] fetch(req_or_url) Fetch request (or URL) and update local objects
[s] view(response) View response in a browser
使用 shell,您可以尝试使用 CSS 和 `response` 对象选择元素
>>> response.css("title")
[<Selector query='descendant-or-self::title' data='<title>Quotes to Scrape</title>'>]
运行 response.css('title') 的结果是一个类似列表的对象,称为 SelectorList,它表示一系列 Selector 对象,这些对象封装了 XML/HTML 元素,允许您运行进一步的查询来优化选择或提取数据。
要从上面的标题中提取文本,您可以这样做
>>> response.css("title::text").getall()
['Quotes to Scrape']
这里有两点需要注意:一是我们在 CSS 查询中添加了 ::text,表示我们只想选择 <title> 元素内部的文本元素。如果我们不指定 ::text,我们将得到完整的标题元素,包括其标签
>>> response.css("title").getall()
['<title>Quotes to Scrape</title>']
另一点是,调用 .getall() 的结果是一个列表:选择器可能返回多个结果,因此我们提取所有结果。当您知道只需要第一个结果时,就像本例中一样,您可以这样做
>>> response.css("title::text").get()
'Quotes to Scrape'
作为替代,您可以这样写
>>> response.css("title::text")[0].get()
'Quotes to Scrape'
如果 SelectorList 实例没有结果,访问其索引将引发 IndexError 异常
>>> response.css("noelement")[0].get()
Traceback (most recent call last):
...
IndexError: list index out of range
您可能希望直接在 SelectorList 实例上使用 .get(),如果没有任何结果,它将返回 None
>>> response.css("noelement").get()
这里有一个教训:对于大多数抓取代码,您希望它能够弹性地处理页面上未找到内容导致的错误,这样即使某些部分抓取失败,您也至少可以获取到一些数据。
除了 getall() 和 get() 方法之外,您还可以使用 re() 方法来使用 正则表达式 进行提取
>>> response.css("title::text").re(r"Quotes.*")
['Quotes to Scrape']
>>> response.css("title::text").re(r"Q\w+")
['Quotes']
>>> response.css("title::text").re(r"(\w+) to (\w+)")
['Quotes', 'Scrape']
为了找到合适的 CSS 选择器,您可能会发现使用 view(response) 从 shell 在您的网络浏览器中打开响应页面很有用。您可以使用浏览器的开发者工具检查 HTML 并找出选择器(参见 使用浏览器开发者工具进行抓取)。
Selector Gadget 也是一个很好的工具,可以快速找到视觉选择元素的 CSS 选择器,它在许多浏览器中都适用。
XPath:简介
除了 CSS,Scrapy 选择器还支持使用 XPath 表达式
>>> response.xpath("//title")
[<Selector query='//title' data='<title>Quotes to Scrape</title>'>]
>>> response.xpath("//title/text()").get()
'Quotes to Scrape'
XPath 表达式非常强大,是 Scrapy 选择器的基础。实际上,CSS 选择器在底层会被转换为 XPath。如果您仔细阅读 shell 中选择器对象的文本表示,您就会发现这一点。
虽然 XPath 表达式可能不像 CSS 选择器那样流行,但它们提供了更大的能力,因为除了导航结构之外,它还可以查看内容。使用 XPath,您可以选择诸如:包含文本“下一页”的链接。这使得 XPath 非常适合抓取任务,我们鼓励您学习 XPath,即使您已经知道如何构建 CSS 选择器,它也会让抓取变得更容易。
我们在这里不会过多介绍 XPath,但您可以在此处阅读更多关于在 Scrapy 选择器中使用 XPath 的信息。要了解更多关于 XPath 的信息,我们推荐 这个通过示例学习 XPath 的教程,以及 这个学习“如何用 XPath 思考”的教程。
在我们的 spider 中提取数据
让我们回到我们的 spider。到目前为止,它还没有提取任何特定数据,只是将整个 HTML 页面保存到本地文件。现在让我们将上述提取逻辑集成到我们的 spider 中。
Scrapy spider 通常会生成许多包含从页面中提取的数据的字典。为此,我们可以在回调中使用 Python 关键字 yield,如下所示
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
"https://quotes.toscrape.com/page/2/",
]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}
要运行此 spider,请通过输入以下内容退出 scrapy shell
quit()
然后,运行
scrapy crawl quotes
现在,它应该会输出带有日志的提取数据
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['life', 'love'], 'author': 'André Gide', 'text': '“It is better to be hated for what you are than to be loved for what you are not.”'}
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['edison', 'failure', 'inspirational', 'paraphrased'], 'author': 'Thomas A. Edison', 'text': "“I have not failed. I've just found 10,000 ways that won't work.”"}
存储抓取到的数据
存储抓取到的数据最简单的方法是使用 Feed exports,命令如下
scrapy crawl quotes -O quotes.json
这将生成一个 quotes.json 文件,其中包含所有抓取到的项目,以 JSON 格式序列化。
-O 命令行开关会覆盖任何现有文件;请改用 -o 将新内容追加到任何现有文件。但是,追加到 JSON 文件会使文件内容变为无效 JSON。当追加到文件时,请考虑使用不同的序列化格式,例如 JSON Lines
scrapy crawl quotes -o quotes.jsonl
JSON Lines 格式很有用,因为它像流一样,您可以轻松地向其中追加新记录。当您运行两次时,它不会出现与 JSON 相同的问题。此外,由于每个记录都是单独的一行,您可以处理大文件而无需将所有内容都加载到内存中,还有 JQ 等工具可以帮助您在命令行上完成此操作。
在小型项目(例如本教程中的项目)中,这应该足够了。但是,如果您想对抓取到的项目执行更复杂的操作,您可以编写一个 Item Pipeline。当项目创建时,已经在 tutorial/pipelines.py 中为您设置了一个 Item Pipelines 占位文件。不过,如果您只是想存储抓取到的项目,则无需实现任何 item pipeline。
跟踪链接
假设您不想只从 https://quotes.toscrape.com 的前两个页面抓取内容,而是想从网站的所有页面抓取引言。
既然您已经知道如何从页面中提取数据,那么接下来让我们看看如何跟踪页面中的链接。
首先要做的是提取我们想要跟踪的页面的链接。检查我们的页面,我们可以看到有一个指向下一页的链接,其标记如下
<ul class="pager">
<li class="next">
<a href="/page/2/">Next <span aria-hidden="true">→</span></a>
</li>
</ul>
我们可以在 shell 中尝试提取它
>>> response.css('li.next a').get()
'<a href="/page/2/">Next <span aria-hidden="true">→</span></a>'
这获取了 `anchor` 元素,但我们想要的是属性 href。为此,Scrapy 支持一个 CSS 扩展,允许您选择属性内容,如下所示
>>> response.css("li.next a::attr(href)").get()
'/page/2/'
还有一个 attrib 属性可用(更多信息请参阅 选择元素属性)
>>> response.css("li.next a").attrib["href"]
'/page/2/'
现在让我们看看我们的 spider,它经过修改后可以递归跟踪到下一页的链接,并从中提取数据
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page is not None:
next_page = response.urljoin(next_page)
yield scrapy.Request(next_page, callback=self.parse)
现在,提取数据后,parse() 方法会查找指向下一页的链接,使用 urljoin() 方法构建一个完整的绝对 URL(因为链接可以是相对的),并向下一页生成一个新的请求,将自身注册为回调函数以处理下一页的数据提取,并使爬取在所有页面中持续进行。
您在这里看到的是 Scrapy 跟踪链接的机制:当您在回调方法中生成一个 Request 时,Scrapy 将调度该请求发送,并注册一个回调方法,以便在该请求完成时执行。
利用这一点,您可以构建复杂的爬虫,它们根据您定义的规则跟踪链接,并根据正在访问的页面提取不同类型的数据。
在我们的示例中,它创建了一个循环,跟踪所有指向下一页的链接,直到找不到为止——这对于爬取博客、论坛和其他带有分页的网站非常方便。
创建请求的快捷方式
作为创建 Request 对象的快捷方式,您可以使用 response.follow
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("span small::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page is not None:
yield response.follow(next_page, callback=self.parse)
与 scrapy.Request 不同,response.follow 直接支持相对 URL,无需调用 urljoin。请注意,response.follow 仅返回一个 Request 实例;您仍然需要生成(yield)此 Request。
您也可以向 response.follow 传入一个选择器而不是字符串;这个选择器应该提取必要的属性
for href in response.css("ul.pager a::attr(href)"):
yield response.follow(href, callback=self.parse)
对于 <a> 元素有一个快捷方式:response.follow 会自动使用它们的 href 属性。因此代码可以进一步缩短
for a in response.css("ul.pager a"):
yield response.follow(a, callback=self.parse)
要从可迭代对象创建多个请求,您可以改用 response.follow_all
anchors = response.css("ul.pager a")
yield from response.follow_all(anchors, callback=self.parse)
或者,进一步缩短它
yield from response.follow_all(css="ul.pager a", callback=self.parse)
更多示例和模式
这里是另一个 spider,它演示了回调和跟踪链接,这次是为了抓取作者信息
import scrapy
class AuthorSpider(scrapy.Spider):
name = "author"
start_urls = ["https://quotes.toscrape.com/"]
def parse(self, response):
author_page_links = response.css(".author + a")
yield from response.follow_all(author_page_links, self.parse_author)
pagination_links = response.css("li.next a")
yield from response.follow_all(pagination_links, self.parse)
def parse_author(self, response):
def extract_with_css(query):
return response.css(query).get(default="").strip()
yield {
"name": extract_with_css("h3.author-title::text"),
"birthdate": extract_with_css(".author-born-date::text"),
"bio": extract_with_css(".author-description::text"),
}
这个 spider 将从主页开始,它将跟踪所有指向作者页面的链接,并为每个链接调用 parse_author 回调,同时也会像我们之前看到的那样,使用 parse 回调处理分页链接。
这里我们以位置参数的形式将回调传递给 response.follow_all,以使代码更简洁;它也适用于 Request。
parse_author 回调定义了一个辅助函数,用于从 CSS 查询中提取和清理数据,并生成包含作者数据的 Python 字典。
这个 spider 演示的另一个有趣之处是,即使有许多来自同一作者的引言,我们也不必担心多次访问同一个作者页面。默认情况下,Scrapy 会过滤掉已访问 URL 的重复请求,避免因编程错误而过多地访问服务器。这可以在 DUPEFILTER_CLASS 设置中进行配置。
希望到现在您已经对如何在 Scrapy 中使用跟踪链接和回调机制有了很好的理解。
作为另一个利用跟踪链接机制的 spider 示例,请查看 CrawlSpider 类,它是一个通用的 spider,实现了一个小型规则引擎,您可以在其之上编写您的爬虫。
此外,一种常见的模式是使用 将额外数据传递给回调的技巧,从多个页面构建一个 item。
使用 spider 参数
您可以在运行 spider 时使用 -a 选项为其提供命令行参数
scrapy crawl quotes -O quotes-humor.json -a tag=humor
这些参数会传递给 Spider 的 __init__ 方法,并默认成为 spider 的属性。
在此示例中,为 tag 参数提供的值将通过 self.tag 获得。您可以使用此功能让您的 spider 只抓取带有特定标签的引言,并根据参数构建 URL
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
async def start(self):
url = "https://quotes.toscrape.com/"
tag = getattr(self, "tag", None)
if tag is not None:
url = url + "tag/" + tag
yield scrapy.Request(url, self.parse)
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page is not None:
yield response.follow(next_page, self.parse)
如果您将 tag=humor 参数传递给此 spider,您会注意到它将只访问来自 humor 标签的 URL,例如 https://quotes.toscrape.com/tag/humor。
后续步骤
本教程仅涵盖了 Scrapy 的基础知识,但还有许多其他功能此处未提及。请查看 还有什么? 部分在Scrapy 概览章节中,以快速了解最重要的功能。
您可以从 基本概念 部分继续学习,以了解更多关于命令行工具、spider、选择器以及本教程未涵盖的其他内容,例如抓取数据的建模。如果您更喜欢使用示例项目,请查看 示例 部分。