Scrapy 教程

在本教程中,我们假设 Scrapy 已安装在您的系统上。如果尚未安装,请参阅 安装指南

我们将抓取 quotes.toscrape.com,这是一个列出著名作家引言的网站。

本教程将引导您完成以下任务

  1. 创建一个新的 Scrapy 项目

  2. 编写一个 spider 来爬取网站并提取数据

  3. 使用命令行导出抓取到的数据

  4. 修改 spider 以递归跟踪链接

  5. 使用 spider 参数

Scrapy 是用 Python 编写的。您对 Python 了解得越多,就能从 Scrapy 中获得越多。

如果您已经熟悉其他语言并想快速学习 Python,Python 教程是一个很好的资源。

如果您是编程新手并想从 Python 开始,以下书籍可能对您有用

您还可以查阅这份面向非程序员的 Python 资源列表,以及 learnpython-subreddit 中推荐的资源

创建项目

在开始抓取之前,您需要设置一个新的 Scrapy 项目。进入您希望存储代码的目录并运行

scrapy startproject tutorial

这将创建一个包含以下内容的 tutorial 目录

tutorial/
    scrapy.cfg            # deploy configuration file

    tutorial/             # project's Python module, you'll import your code from here
        __init__.py

        items.py          # project items definition file

        middlewares.py    # project middlewares file

        pipelines.py      # project pipelines file

        settings.py       # project settings file

        spiders/          # a directory where you'll later put your spiders
            __init__.py

我们的第一个 Spider

Spider 是您定义的类,Scrapy 用它们从网站(或一组网站)抓取信息。它们必须继承 Spider 并定义要发出的初始请求,还可以选择定义如何跟踪页面中的链接以及解析下载的页面内容以提取数据。

这是我们第一个 Spider 的代码。将其保存为名为 quotes_spider.py 的文件,放在您项目中的 tutorial/spiders 目录下

from pathlib import Path

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"

    async def start(self):
        urls = [
            "https://quotes.toscrape.com/page/1/",
            "https://quotes.toscrape.com/page/2/",
        ]
        for url in urls:
            yield scrapy.Request(url=url, callback=self.parse)

    def parse(self, response):
        page = response.url.split("/")[-2]
        filename = f"quotes-{page}.html"
        Path(filename).write_bytes(response.body)
        self.log(f"Saved file {filename}")

如您所见,我们的 Spider 继承自 scrapy.Spider 并定义了一些属性和方法

  • name:标识 Spider。它在项目中必须是唯一的,也就是说,您不能为不同的 Spider 设置相同的名称。

  • start():必须是一个异步生成器,用于生成请求(以及可选的 item),供 spider 开始爬取。后续请求将从这些初始请求中陆续生成。

  • parse():一个方法,用于处理为每个发出的请求下载的响应。`response` 参数是 TextResponse 的一个实例,它包含页面内容并具有更多有用的处理方法。

    parse() 方法通常会解析响应,将抓取到的数据提取为字典,并找到要跟踪的新 URL 并从中创建新的请求(Request)。

如何运行我们的 spider

要让我们的 spider 工作,请进入项目的顶级目录并运行

scrapy crawl quotes

此命令运行我们刚添加的名为 quotes 的 spider,它将向 quotes.toscrape.com 域发送一些请求。您将看到类似于以下的输出

... (omitted for brevity)
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Spider opened
2016-12-16 21:24:05 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2016-12-16 21:24:05 [scrapy.extensions.telnet] DEBUG: Telnet console listening on 127.0.0.1:6023
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (404) <GET https://quotes.toscrape.com/robots.txt> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/2/> (referer: None)
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-1.html
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-2.html
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Closing spider (finished)
...

现在,检查当前目录中的文件。您应该注意到已创建了两个新文件:quotes-1.htmlquotes-2.html,其中包含相应 URL 的内容,正如我们的 parse 方法所指示的。

注意

如果您想知道为什么我们还没有解析 HTML,请稍等,我们很快就会讲到。

底层发生了什么?

Scrapy 发送由 start() spider 方法生成的第一个 scrapy.Request 对象。在收到每个请求的响应后,Scrapy 会调用与该请求关联的回调方法(在此示例中为 parse 方法),并传入一个 Response 对象。

start 方法的快捷方式

除了实现一个 start() 方法来从 URL 生成 Request 对象,您还可以定义一个 start_urls 类属性,其中包含一个 URL 列表。然后,这个列表将被 start() 的默认实现用来为您的 spider 创建初始请求。

from pathlib import Path

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    start_urls = [
        "https://quotes.toscrape.com/page/1/",
        "https://quotes.toscrape.com/page/2/",
    ]

    def parse(self, response):
        page = response.url.split("/")[-2]
        filename = f"quotes-{page}.html"
        Path(filename).write_bytes(response.body)

即使我们没有明确告诉 Scrapy,parse() 方法也会被调用来处理这些 URL 的每个请求。之所以会这样,是因为 parse() 是 Scrapy 的默认回调方法,它会在没有明确指定回调的请求中被调用。

提取数据

学习如何使用 Scrapy 提取数据的最佳方法是使用 Scrapy shell 尝试选择器。运行

scrapy shell 'https://quotes.toscrape.com/page/1/'

注意

请记住,从命令行运行 Scrapy shell 时,务必将 URL 用引号括起来,否则包含参数(即 & 字符)的 URL 将无法工作。

在 Windows 上,请改用双引号

scrapy shell "https://quotes.toscrape.com/page/1/"

您将看到类似以下内容

[ ... Scrapy log here ... ]
2016-09-19 12:09:27 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
[s] Available Scrapy objects:
[s]   scrapy     scrapy module (contains scrapy.Request, scrapy.Selector, etc)
[s]   crawler    <scrapy.crawler.Crawler object at 0x7fa91d888c90>
[s]   item       {}
[s]   request    <GET https://quotes.toscrape.com/page/1/>
[s]   response   <200 https://quotes.toscrape.com/page/1/>
[s]   settings   <scrapy.settings.Settings object at 0x7fa91d888c10>
[s]   spider     <DefaultSpider 'default' at 0x7fa91c8af990>
[s] Useful shortcuts:
[s]   shelp()           Shell help (print this help)
[s]   fetch(req_or_url) Fetch request (or URL) and update local objects
[s]   view(response)    View response in a browser

使用 shell,您可以尝试使用 CSS 和 `response` 对象选择元素

>>> response.css("title")
[<Selector query='descendant-or-self::title' data='<title>Quotes to Scrape</title>'>]

运行 response.css('title') 的结果是一个类似列表的对象,称为 SelectorList,它表示一系列 Selector 对象,这些对象封装了 XML/HTML 元素,允许您运行进一步的查询来优化选择或提取数据。

要从上面的标题中提取文本,您可以这样做

>>> response.css("title::text").getall()
['Quotes to Scrape']

这里有两点需要注意:一是我们在 CSS 查询中添加了 ::text,表示我们只想选择 <title> 元素内部的文本元素。如果我们不指定 ::text,我们将得到完整的标题元素,包括其标签

>>> response.css("title").getall()
['<title>Quotes to Scrape</title>']

另一点是,调用 .getall() 的结果是一个列表:选择器可能返回多个结果,因此我们提取所有结果。当您知道只需要第一个结果时,就像本例中一样,您可以这样做

>>> response.css("title::text").get()
'Quotes to Scrape'

作为替代,您可以这样写

>>> response.css("title::text")[0].get()
'Quotes to Scrape'

如果 SelectorList 实例没有结果,访问其索引将引发 IndexError 异常

>>> response.css("noelement")[0].get()
Traceback (most recent call last):
...
IndexError: list index out of range

您可能希望直接在 SelectorList 实例上使用 .get(),如果没有任何结果,它将返回 None

>>> response.css("noelement").get()

这里有一个教训:对于大多数抓取代码,您希望它能够弹性地处理页面上未找到内容导致的错误,这样即使某些部分抓取失败,您也至少可以获取到一些数据。

除了 getall()get() 方法之外,您还可以使用 re() 方法来使用 正则表达式 进行提取

>>> response.css("title::text").re(r"Quotes.*")
['Quotes to Scrape']
>>> response.css("title::text").re(r"Q\w+")
['Quotes']
>>> response.css("title::text").re(r"(\w+) to (\w+)")
['Quotes', 'Scrape']

为了找到合适的 CSS 选择器,您可能会发现使用 view(response) 从 shell 在您的网络浏览器中打开响应页面很有用。您可以使用浏览器的开发者工具检查 HTML 并找出选择器(参见 使用浏览器开发者工具进行抓取)。

Selector Gadget 也是一个很好的工具,可以快速找到视觉选择元素的 CSS 选择器,它在许多浏览器中都适用。

XPath:简介

除了 CSS,Scrapy 选择器还支持使用 XPath 表达式

>>> response.xpath("//title")
[<Selector query='//title' data='<title>Quotes to Scrape</title>'>]
>>> response.xpath("//title/text()").get()
'Quotes to Scrape'

XPath 表达式非常强大,是 Scrapy 选择器的基础。实际上,CSS 选择器在底层会被转换为 XPath。如果您仔细阅读 shell 中选择器对象的文本表示,您就会发现这一点。

虽然 XPath 表达式可能不像 CSS 选择器那样流行,但它们提供了更大的能力,因为除了导航结构之外,它还可以查看内容。使用 XPath,您可以选择诸如:包含文本“下一页”的链接。这使得 XPath 非常适合抓取任务,我们鼓励您学习 XPath,即使您已经知道如何构建 CSS 选择器,它也会让抓取变得更容易。

我们在这里不会过多介绍 XPath,但您可以在此处阅读更多关于在 Scrapy 选择器中使用 XPath 的信息。要了解更多关于 XPath 的信息,我们推荐 这个通过示例学习 XPath 的教程,以及 这个学习“如何用 XPath 思考”的教程

提取引言和作者

现在您对选择和提取有了一些了解,接下来让我们通过编写代码从网页中提取引言来完善我们的 spider。

https://quotes.toscrape.com 中的每个引言都由类似以下内容的 HTML 元素表示

<div class="quote">
    <span class="text">“The world as we have created it is a process of our
    thinking. It cannot be changed without changing our thinking.”</span>
    <span>
        by <small class="author">Albert Einstein</small>
        <a href="/author/Albert-Einstein">(about)</a>
    </span>
    <div class="tags">
        Tags:
        <a class="tag" href="/tag/change/page/1/">change</a>
        <a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a>
        <a class="tag" href="/tag/thinking/page/1/">thinking</a>
        <a class="tag" href="/tag/world/page/1/">world</a>
    </div>
</div>

让我们打开 scrapy shell 试一下,看看如何提取我们想要的数据

scrapy shell 'https://quotes.toscrape.com'

我们通过以下方式获取引言 HTML 元素的选择器列表

>>> response.css("div.quote")
[<Selector query="descendant-or-self::div[@class and contains(concat(' ', normalize-space(@class), ' '), ' quote ')]" data='<div class="quote" itemscope itemtype...'>,
<Selector query="descendant-or-self::div[@class and contains(concat(' ', normalize-space(@class), ' '), ' quote ')]" data='<div class="quote" itemscope itemtype...'>,
...]

上述查询返回的每个选择器都允许我们对其子元素运行进一步的查询。让我们将第一个选择器赋给一个变量,这样我们就可以直接在特定的引言上运行我们的 CSS 选择器

>>> quote = response.css("div.quote")[0]

现在,让我们使用刚刚创建的 quote 对象从该引言中提取 textauthortags

>>> text = quote.css("span.text::text").get()
>>> text
'“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”'
>>> author = quote.css("small.author::text").get()
>>> author
'Albert Einstein'

由于标签是一个字符串列表,我们可以使用 .getall() 方法来获取所有标签

>>> tags = quote.css("div.tags a.tag::text").getall()
>>> tags
['change', 'deep-thoughts', 'thinking', 'world']

弄清楚如何提取每个部分后,我们现在可以遍历所有引言元素并将它们组合成一个 Python 字典

>>> for quote in response.css("div.quote"):
...     text = quote.css("span.text::text").get()
...     author = quote.css("small.author::text").get()
...     tags = quote.css("div.tags a.tag::text").getall()
...     print(dict(text=text, author=author, tags=tags))
...
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']}
{'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']}
...

在我们的 spider 中提取数据

让我们回到我们的 spider。到目前为止,它还没有提取任何特定数据,只是将整个 HTML 页面保存到本地文件。现在让我们将上述提取逻辑集成到我们的 spider 中。

Scrapy spider 通常会生成许多包含从页面中提取的数据的字典。为此,我们可以在回调中使用 Python 关键字 yield,如下所示

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    start_urls = [
        "https://quotes.toscrape.com/page/1/",
        "https://quotes.toscrape.com/page/2/",
    ]

    def parse(self, response):
        for quote in response.css("div.quote"):
            yield {
                "text": quote.css("span.text::text").get(),
                "author": quote.css("small.author::text").get(),
                "tags": quote.css("div.tags a.tag::text").getall(),
            }

要运行此 spider,请通过输入以下内容退出 scrapy shell

quit()

然后,运行

scrapy crawl quotes

现在,它应该会输出带有日志的提取数据

2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['life', 'love'], 'author': 'André Gide', 'text': '“It is better to be hated for what you are than to be loved for what you are not.”'}
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['edison', 'failure', 'inspirational', 'paraphrased'], 'author': 'Thomas A. Edison', 'text': "“I have not failed. I've just found 10,000 ways that won't work.”"}

存储抓取到的数据

存储抓取到的数据最简单的方法是使用 Feed exports,命令如下

scrapy crawl quotes -O quotes.json

这将生成一个 quotes.json 文件,其中包含所有抓取到的项目,以 JSON 格式序列化。

-O 命令行开关会覆盖任何现有文件;请改用 -o 将新内容追加到任何现有文件。但是,追加到 JSON 文件会使文件内容变为无效 JSON。当追加到文件时,请考虑使用不同的序列化格式,例如 JSON Lines

scrapy crawl quotes -o quotes.jsonl

JSON Lines 格式很有用,因为它像流一样,您可以轻松地向其中追加新记录。当您运行两次时,它不会出现与 JSON 相同的问题。此外,由于每个记录都是单独的一行,您可以处理大文件而无需将所有内容都加载到内存中,还有 JQ 等工具可以帮助您在命令行上完成此操作。

在小型项目(例如本教程中的项目)中,这应该足够了。但是,如果您想对抓取到的项目执行更复杂的操作,您可以编写一个 Item Pipeline。当项目创建时,已经在 tutorial/pipelines.py 中为您设置了一个 Item Pipelines 占位文件。不过,如果您只是想存储抓取到的项目,则无需实现任何 item pipeline。

使用 spider 参数

您可以在运行 spider 时使用 -a 选项为其提供命令行参数

scrapy crawl quotes -O quotes-humor.json -a tag=humor

这些参数会传递给 Spider 的 __init__ 方法,并默认成为 spider 的属性。

在此示例中,为 tag 参数提供的值将通过 self.tag 获得。您可以使用此功能让您的 spider 只抓取带有特定标签的引言,并根据参数构建 URL

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"

    async def start(self):
        url = "https://quotes.toscrape.com/"
        tag = getattr(self, "tag", None)
        if tag is not None:
            url = url + "tag/" + tag
        yield scrapy.Request(url, self.parse)

    def parse(self, response):
        for quote in response.css("div.quote"):
            yield {
                "text": quote.css("span.text::text").get(),
                "author": quote.css("small.author::text").get(),
            }

        next_page = response.css("li.next a::attr(href)").get()
        if next_page is not None:
            yield response.follow(next_page, self.parse)

如果您将 tag=humor 参数传递给此 spider,您会注意到它将只访问来自 humor 标签的 URL,例如 https://quotes.toscrape.com/tag/humor

您可以在此处了解更多关于处理 spider 参数的信息

后续步骤

本教程仅涵盖了 Scrapy 的基础知识,但还有许多其他功能此处未提及。请查看 还有什么? 部分在Scrapy 概览章节中,以快速了解最重要的功能。

您可以从 基本概念 部分继续学习,以了解更多关于命令行工具、spider、选择器以及本教程未涵盖的其他内容,例如抓取数据的建模。如果您更喜欢使用示例项目,请查看 示例 部分。