• linkedu视频
  • 平面设计
  • 电脑入门
  • 操作系统
  • 办公应用
  • 电脑硬件
  • 动画设计
  • 3D设计
  • 网页设计
  • CAD设计
  • 影音处理
  • 数据库
  • 程序设计
  • 认证考试
  • 信息管理
  • 信息安全
菜单
linkedu.com
导航菜单
  • 网页制作
  • 数据库
  • 程序设计
  • 操作系统
  • CMS教程
  • 游戏攻略
  • 脚本语言
  • 平面设计
  • 软件教程
  • 网络安全
  • 电脑知识
  • 服务器
  • 视频教程
  • windows
  • 服务器硬件
  • 服务器运维
  • 云计算
  • 虚拟化
  • IIS教程
  • Linux
  • Apache
  • Ftp
  • DNS
  • Nginx
您的位置:首页 > 服务器 >windows > Mailbox:日支撑过亿信息数据库的性能调优及集群迁移

Mailbox:日支撑过亿信息数据库的性能调优及集群迁移

作者:网友 字体:[增加 减小] 来源:互联网

本文主要包含mailbox,mailbox是什么意思,163 mailbox,mailbox下载,qq mailbox等服务器相关知识,网友希望可以进行参考

在Mailbox快速扩展过程中,其中一个性能问题就是MongoDB的数据库级别写锁,在锁等待过程中耗费的时间,直接反应到用户使用服务过程中的延时。为了解决这个长期存在的问题,我们决定将一个常用的MongoDB集合储存了邮件相关数据)迁移到独立的集群上。根据我们推断,这将减少50%的锁等待时间;同时,我们还可以添加更多的分片,我们还期望可以独立的优化及管理不同类型数据。


我们首先从MongoDB文档开始,很快的就发现了 cloneCollection命令。然而随后悲剧的发现,它不可以在分片集合中使用;同样, renameCollection也不能在分片集合中使用。在否定了其它可能性之后基于性能问题),我们编写了一个Python脚本用以复制数据,和另一个用于比较原始和目标数据的脚本。在这个过程中,我们还发现了许多有意思的事情,比如 gevent及 pymongo复制大数据集的时间是 mongodumpC++编写)的一半,即使MongoDB客户端和服务器在同台主机上。通过最终努力,我们开发了 Hydra,用于MongoDB迁移的工具集,现已开源。首先,我们建立了MongoDB集合的原始快照。


问题1:悲剧的性能


早期我做了一个实验以测试MongoDB API运作所能达到的极限速度——启用一个简单的使用MongoDB C++ 软件开发工具包的速度。一方面对C++ 感觉厌烦,一方面希望我大多数熟练使用Python的同事可以在其他用途上使用或适应这种代码,我没有更进一步的探索C++的使用,而是发现,如果是针对少量数据,在处理相同任务上,简单的C++应用速度是简单Python应用的5-10倍。


所以,我的研究方向回到了Python,这个Dropbox默认语言。此外,进行了诸如对mongod查询等的一系列远程网络请求时,客户端往往需要耗费大量时间等待服务器响应;似乎也没有很多copy_collection.py 我的MongoDB集合复制工具)需要的CPU密集型操作部分)。initialcopy_collection.py占很少的CPU使用率也证实了这一点。


然后,MongoDB请求到copy_collection.py.。最初的工作线程实验结果并不理想。但接下来,我们通过Python Queue对象来实现工作线程通信。这样的性能依旧不是很好,因为IPC上的开销让并发带来的提升黯然失色。使用Pipes和其他IPC机制也并没有多大帮助。


接下来,我们尝试了使用单线程Python进行MongoDB异步查询,看看可以有多少性能结余。其中Gevent是实现这个途径常用库之一,我们对它进行了尝试。Gevent 修改了标准Python模块以实现异步操作,比如socket。比较好的一点是,你可以简单的编写异步读取代码,就像同步代码一样。




def copy_documents(source_collection, destination_collection, _ids, callback):


"""


Given a list of _id's (MongoDB's unique identifier field for each document),


copies the corresponding documents from the source collection to the destination


collection


"""


def _copy_documents_callback(...):

if error_detected():

callback(error)



# copy documents, passing a callback function that will handle errors and


# other notifications

for _id in _ids:

copy_document(source_collection, destination_collection, _id,

_copy_documents_callback)

http://www.dianjiare.net


# more error handling omitted for brevity

callback(None)


def copy_document(source_collection, destination_collection, _id, callback):


"""


Copies document corresponding to the given _id from the source to the


destination.


"""

def _insert_doc(doc):


"""http://www.kongzhixitong.net


callback that takes the document read from the source collection


and inserts it into destination collection


"""

if error_detected():

callback(error)

destination_collection.insert(doc, callback)

# another MongoDB operation



# find the specified document asynchronously, passing a callback to receive


# the retrieved data

source_collection.find_one({'$id': _id}, callback=_insert_doc)

有了gevent,这些代码不再需要使用callback:http://www.dianjiarequan.net


1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

21

22

23

24

import gevent

gevent.monkey.patch_all()


def copy_documents(source_collection, destination_collection, _ids):


"""


Given a list of _id's (MongoDB's unique identifier field for each document),


copies the corresponding documents from the source collection to the destination


collectio n http://www.dantouguan.net


"""



# copies each document using a separate greenlet; optimizations are certainly


# possible but omitted in this example

for _id in _ids:

gevent.spawn(copy_document, source_collection, destination_collection, _id)


def copy_document(source_collection, destination_collection, _id):


"""


Copies document corresponding to the given _id from the source to the


destination.http://www.yetijiare.com


"""


# both of the following function calls block without gevent; with gevent they


# simply cede control to another greenlet while waiting for Mongo to respond

source_doc = source_collection.find_one({'$id': _id})

destination_collection.insert(source_doc)

# another MongoDB operation

这种简单的代码可以根据它们的_idfields,从MongoDB源集合拷取代码到目标位置,它们的_idfields是每个MongoDB文档的唯一标识符。opy_documents 会产委派greenlets运行runcopy_document()做文档复制。当greenlets执行一项阻塞操作,比如对MongoDB的任何需求,它会将控制放给其它准备执行的greenlet。因为所有greenlets都在相同的线程和进程中执行,你一般不需要任何形式的内部锁定。

http://www.wujindianqi.net

有了gevent,就能够找到比工作者线程池或工作者进程池更快的方法。下面总结了每种方法的性能:


ApproachPerformance (higher is better)

single process, no gevent520 documents/sec

thread worker pool652 documents/sec

process worker pool670 documents/sec

single process, with gevent2,381 documents/sec

综合gevent和工作者进程每个分片一个)可以在性能上得到一个线性提升。有效使用工作进程的关键是尽可能使用更少的IPC。


问题2:快照后的复制修改http://www.sujiaojixie.net


因为MongoDB不支持事务,如果你对正在执行修改的大数据集进行读取,你得到的结果可能会因时而异。举个例子,你使用MongoDB find()进行整个数据集上的读取,你的结果集可能是:


ncluded: document saved before your find()

included: document saved before your find()

included: document saved before your find()

included: document inserted after your find() began

此外,为了在Mailbox后端指向新副本集时能最小化故障时间,尽可能减少从源集群应用到新集群过程中所耗费的时间则至关重要。


类似多数的异步复制存储,MongoDB使用了操作日志oplog记录下了mongod实例上发生的增、

分享到:QQ空间新浪微博腾讯微博微信百度贴吧QQ好友复制网址打印

您可能想查找下面的文章:

  • Mailbox:日支撑过亿信息数据库的性能调优及集群迁移

相关文章

  • 通过shell脚本防止端口扫描
  • win2003/2008 IIS服务器.flv文件不能访问解决办法
  • LNMP一键安装
  • 解决dell服务器无法识别4G内存办法
  • UAC窗体应用程序默认配置信息读写的问题
  • windows下安装vmware tools图文详解
  • 正则表达式和grep的基本用法9
  • Mac OS X/windows下启用Mod Rewrite和.htaccess
  • apache和nignx中禁止目录访问安装配置方法
  • Apache日志中“指定的网络名不再可用”解决办法

文章分类

  • windows
  • 服务器硬件
  • 服务器运维
  • 云计算
  • 虚拟化
  • IIS教程
  • Linux
  • Apache
  • Ftp
  • DNS
  • Nginx

最近更新的内容

    • iis与apache取消目录脚本执行权限方法
    • windows(32位 64位)下python安装mysqldb模块
    • windows中关闭135危险端口方法
    • Tomcat 404 500错误页面上的版本号隐藏或修改
    • 让虚拟化的风暴 来得再猛烈些吧
    • Hyper-V第1代虚拟机和第2代虚拟机特性对照表
    • Windows Server 2008 的十种新特性
    • liunx基本命令(进程的学习)
    • 64 位硬件和软件的优势(Office SharePoint Server 2007)
    • Hyper-V 虚拟网络技术之二

关于我们 - 联系我们 - 免责声明 - 网站地图

©2020-2025 All Rights Reserved. linkedu.com 版权所有