python - urllib2.urlopen() vs urllib.urlopen() - urllib2 は動作中に 404 をスローします! なぜ？

Question

import urllib

print urllib.urlopen('http://www.reefgeek.com/equipment/Controllers_&_Monitors/Neptune_Systems_AquaController/Apex_Controller_&_Accessories/').read()

上記のスクリプトは機能し、次の場合に期待される結果を返します。

import urllib2

print urllib2.urlopen('http://www.reefgeek.com/equipment/Controllers_&_Monitors/Neptune_Systems_AquaController/Apex_Controller_&_Accessories/').read()

次のエラーがスローされます。

Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/usr/lib/python2.5/urllib2.py", line 124, in urlopen
    return _opener.open(url, data)
  File "/usr/lib/python2.5/urllib2.py", line 387, in open
    response = meth(req, response)
  File "/usr/lib/python2.5/urllib2.py", line 498, in http_response
    'http', request, response, code, msg, hdrs)
  File "/usr/lib/python2.5/urllib2.py", line 425, in error
    return self._call_chain(*args)
  File "/usr/lib/python2.5/urllib2.py", line 360, in _call_chain
    result = func(*args)
  File "/usr/lib/python2.5/urllib2.py", line 506, in http_error_default
    raise HTTPError(req.get_full_url(), code, msg, hdrs, fp)
urllib2.HTTPError: HTTP Error 404: Not Found

これがなぜなのか誰か知っていますか？私はプロキシ設定なしでホームネットワーク上のラップトップからこれを実行しています.ラップトップからルーターへ、そしてwww.

score 35 · Accepted Answer

その URL は確かに 404 になりますが、多くの HTML コンテンツが含まれています。urllib2エラー状態として（正しく）処理しています。次のように、そのサイトの 404 ページのコンテンツを復元できます。

import urllib2
try:
    print urllib2.urlopen('http://www.reefgeek.com/equipment/Controllers_&_Monitors/Neptune_Systems_AquaController/Apex_Controller_&_Accessories/').read()
except urllib2.HTTPError, e:
    print e.code
    print e.msg
    print e.headers
    print e.fp.read()

python - urllib2.urlopen() vs urllib.urlopen() - urllib2 は動作中に 404 をスローします! なぜ？

1 に答える 1

Related

Reference