4

response.replace() を使用して、Google の検索結果ページの検索結果ブロックの応答本文を置き換えようとしていますが、エンコードの問題に直面しています。

scrapy  shell "http://www.google.de/search?q=Zuckerccc"

>>> srb = hxs.select("//li[@class='g']").extract()
>>> body = '<html><body>' + srb[0] + '</body></html>'    # get only 1st search result block
>>> b = response.replace(body = body)
Traceback (most recent call last):
  File "<console>", line 1, in <module>
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/text.py", line 54, in replace
    return Response.replace(self, *args, **kwargs)
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/__init__.py", line 77, in replace
    return cls(*args, **kwargs)
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/text.py", line 31, in __init__
    super(TextResponse, self).__init__(*args, **kwargs)
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/__init__.py", line 19, in __init__
    self._set_body(body)
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/text.py", line 48, in _set_body
    self._body = body.encode(self._encoding)
  File "../local_1/Linux-2.6c2.5-x86_64/Python/Python-147.0-0/lib/python2.6/encodings/cp1252.py", line 12, in encode
    return codecs.charmap_encode(input,errors,encoding_table)
UnicodeEncodeError: 'charmap' codec can't encode character u'\u0131' in position 529: character maps to <undefined>

自分なりの回答も作ってみましたが、

>>> x = HtmlResponse("http://www.google.de/search?q=Zuckerccc", body = body, encoding = response.encoding)
Traceback (most recent call last):
  File "<console>", line 1, in <module>
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/text.py", line 31, in __init__
    super(TextResponse, self).__init__(*args, **kwargs)
    self._set_body(body)
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/text.py", line 48, in _set_body
    self._body = body.encode(self._encoding)
  File "../local_1/Linux-2.6c2.5-x86_64/Python/Python-147.0-0/lib/python2.6/encodings/cp1252.py", line 12, in encode
    return codecs.charmap_encode(input,errors,encoding_table)
UnicodeEncodeError: 'charmap' codec can't encode character u'\u0131' in position 529: character maps to <undefined>
  File "scrapy/lib/python2.6/site-packages/scrapy/http/response/__init__.py", line 19, in __init__

また、replace() 関数でのエンコーディングに _body_declared_encoding() を使用すると、機能します。

replace(body = body, encoding = response._body_declared_encoding())

response._body_declared_encoding() と response.encoding が異なる理由がわかりません。誰でもこれに光を当ててください。

それで、これを修正する良い方法は何ですか?

4

2 に答える 2