• linkedu视频
  • 平面设计
  • 电脑入门
  • 操作系统
  • 办公应用
  • 电脑硬件
  • 动画设计
  • 3D设计
  • 网页设计
  • CAD设计
  • 影音处理
  • 数据库
  • 程序设计
  • 认证考试
  • 信息管理
  • 信息安全
菜单
linkedu.com
  • 网页制作
  • 数据库
  • 程序设计
  • 操作系统
  • CMS教程
  • 游戏攻略
  • 脚本语言
  • 平面设计
  • 软件教程
  • 网络安全
  • 电脑知识
  • 服务器
  • 视频教程
  • vbs
  • DOS/BAT
  • hta/htc
  • python
  • perl
  • VBA
  • ColdFusion
  • ruby
  • PowerShell
  • Lua
  • Golang
  • linux shell
您的位置:首页 > 脚本语言 >python > Python3中使用urllib的方法详解(header,代理,超时,认证,异常处理)

Python3中使用urllib的方法详解(header,代理,超时,认证,异常处理)

作者: 字体:[增加 减小] 来源:互联网

通过本文主要向大家介绍了python3 urllib,python3 urllib模块,python3 urllib2,python3 urllib.quote,python3 urllib下载等相关知识,希望对您有所帮助,也希望大家支持linkedu.com www.linkedu.com

我们可以利用urllib来抓取远程的数据进行保存哦,以下是python3 抓取网页资源的多种方法,有需要的可以参考借鉴。

1、最简单

import urllib.request
response = urllib.request.urlopen('http://python.org/')
html = response.read()
</div>

2、使用 Request

import urllib.request
req = urllib.request.Request('http://python.org/')
response = urllib.request.urlopen(req)
the_page = response.read()
</div>

3、发送数据

#! /usr/bin/env python3
import urllib.parse
import urllib.request
url = 'http://localhost/login.php'
user_agent = 'Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)'
values = {
'act' : 'login',
'login[email]' : 'yzhang@i9i8.com',
'login[password]' : '123456'
}
data = urllib.parse.urlencode(values)
req = urllib.request.Request(url, data)
req.add_header('Referer', 'http://www.python.org/')
response = urllib.request.urlopen(req)
the_page = response.read()
print(the_page.decode("utf8"))
</div>

4、发送数据和header

#! /usr/bin/env python3
import urllib.parse
import urllib.request
url = 'http://localhost/login.php'
user_agent = 'Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)'
values = {
'act' : 'login',
'login[email]' : 'yzhang@i9i8.com',
'login[password]' : '123456'
}
headers = { 'User-Agent' : user_agent }
data = urllib.parse.urlencode(values)
req = urllib.request.Request(url, data, headers)
response = urllib.request.urlopen(req)
the_page = response.read()
print(the_page.decode("utf8"))
</div>

5、http 错误

#! /usr/bin/env python3
import urllib.request
req = urllib.request.Request('http://www.jb51.net ')
try:
urllib.request.urlopen(req)
except urllib.error.HTTPError as e:
print(e.code)
print(e.read().decode("utf8"))
</div>

6、异常处理1

#! /usr/bin/env python3
from urllib.request import Request, urlopen
from urllib.error import URLError, HTTPError
req = Request("http://www.jb51.net /")
try:
response = urlopen(req)
except HTTPError as e:
print('The server couldn't fulfill the request.')
print('Error code: ', e.code)
except URLError as e:
print('We failed to reach a server.')
print('Reason: ', e.reason)
else:
print("good!")
print(response.read().decode("utf8"))
</div>

7、异常处理2

#! /usr/bin/env python3
from urllib.request import Request, urlopen
from urllib.error import URLError
req = Request("http://www.jb51.net /")
try:
response = urlopen(req)
except URLError as e:
if hasattr(e, 'reason'):
print('We failed to reach a server.')
print('Reason: ', e.reason)
elif hasattr(e, 'code'):
print('The server couldn't fulfill the request.')
print('Error code: ', e.code)
else:
print("good!")
print(response.read().decode("utf8"))
</div>

8、HTTP 认证

#! /usr/bin/env python3
import urllib.request
# create a password manager
password_mgr = urllib.request.HTTPPasswordMgrWithDefaultRealm()
# Add the username and password.
# If we knew the realm, we could use it instead of None.
top_level_url = "https://www.jb51.net /"
password_mgr.add_password(None, top_level_url, 'rekfan', 'xxxxxx')
handler = urllib.request.HTTPBasicAuthHandler(password_mgr)
# create "opener" (OpenerDirector instance)
opener = urllib.request.build_opener(handler)
# use the opener to fetch a URL
a_url = "https://www.jb51.net /"
x = opener.open(a_url)
print(x.read())
# Install the opener.
# Now all calls to urllib.request.urlopen use our opener.
urllib.request.install_opener(opener)
a = urllib.request.urlopen(a_url).read().decode('utf8')
print(a)
</div>

9、使用代理

#! /usr/bin/env python3
import urllib.request
proxy_support = urllib.request.ProxyHandler({'sock5': 'localhost:1080'})
opener = urllib.request.build_opener(proxy_support)
urllib.request.install_opener(opener)

a = urllib.request.urlopen("http://www.jb51.net ").read().decode("utf8")
print(a)
</div>

10、超时

#! /usr/bin/env python3
import socket
import urllib.request
# timeout in seconds
timeout = 2
socket.setdefaulttimeout(timeout)
# this call to urllib.request.urlopen now uses the default timeout
# we have set in the socket module
req = urllib.request.Request('http://www.jb51.net /')
a = urllib.request.urlopen(req).read()
print(a)
</div>

总结

以上就是这篇文章的全部内容,希望本文的内容对大家学习或使用python能有所帮助,如果有疑问大家可以留言交流。

</div>

您可能想查找下面的文章:

  • 深入理解Python3中的http.client模块
  • 【Python】Python的urllib模块、urllib2模块批量进行网页下载文件
  • Python3中使用urllib的方法详解(header,代理,超时,认证,异常处理)
  • Python网络编程中urllib2模块的用法总结
  • Python使用urllib2模块抓取HTML页面资源的实例分享
  • Python使用urllib2模块抓取HTML页面资源的实例分享
  • 使用Python的urllib2模块处理url和图片的技巧两则
  • 使用Python的urllib2模块处理url和图片的技巧两则
  • Python中的urllib模块使用详解
  • python使用urllib2提交http post请求的方法

相关文章

  • python中的编码知识整理汇总
  • 详解Python中用于计算指数的exp()方法
  • Python模拟登录12306的方法
  • Python的__builtin__模块中的一些要点知识
  • 数据挖掘之Apriori算法详解和Python实现代码分享
  • Python中的map()函数和reduce()函数的用法
  • Python只用40行代码编写的计算器实例
  • 基于Python的接口测试框架实例
  • Python 'takes exactly 1 argument (2 given)' Python error
  • Python爬虫利用cookie实现模拟登陆实例详解

文章分类

  • vbs
  • DOS/BAT
  • hta/htc
  • python
  • perl
  • VBA
  • ColdFusion
  • ruby
  • PowerShell
  • Lua
  • Golang
  • linux shell

最近更新的内容

    • Django的信号机制详解
    • Python最基本的输入输出详解
    • Django中处理出错页面的方法
    • python实现查找两个字符串中相同字符并输出的方法
    • 用Python实现协同过滤的教程
    • 在Python中使用next()方法操作文件的教程
    • python 3.5下xadmin的使用及修复源码bug
    • Python 执行字符串表达式函数(eval exec execfile)
    • python简单获取本机计算机名和IP地址的方法
    • Python实现的简单文件传输服务器和客户端

关于我们 - 联系我们 - 免责声明 - 网站地图

©2020-2025 All Rights Reserved. linkedu.com 版权所有