“丝路通:分布式爬虫任务分配”的版本间的差异

来自CloudWiki
跳转至: 导航搜索
任务表建立
第120行: 第120行:
  
 
==任务表建立==
 
==任务表建立==
===安装mysql和django===
+
===安装mysql===
 
*[[Centos7 安装python3]],本项目安装python3.6
 
*[[Centos7 安装python3]],本项目安装python3.6
 
*[[Centos7 安装MySQL]]
 
*[[Centos7 安装MySQL]]
*[[Django安装与启动]]
 
  
 +
===建立数据表===
 
===model设计和资源导入 ===
 
===model设计和资源导入 ===
*[[项目初始化]]、补充:[[MySQL中的character set和collation]]
+
 
*[[user models设计]]
+
<nowiki>MariaDB [crawler]> CREATE TABLE IF NOT EXISTS `task`(
*[[goods modeles设计]]、补充:[[django中choice的使用]]
+
    ->    `id` INT UNSIGNED AUTO_INCREMENT,
*[[trade交易的model设计]]
+
    ->    `site_title` VARCHAR(100) NOT NULL,
*[[用户操作的model设计]]、补充:[[django中unique together使用]]
+
    ->    `task_name` VARCHAR(40) NOT NULL,
*[[丝路通_爬虫:项目初始化]]、补充:[[MySQL中的character set和collation]]
+
    ->    `task_status` INT UNSIGNED NOT NULL,
*[[丝路通_爬虫:task models设计]]
+
    ->    `availability zones` VARCHAR(60) NOT NULL,
 +
    ->    `start_date` DATE,
 +
    ->    PRIMARY KEY ( `id` ))ENGINE=InnoDB DEFAULT CHARSET=utf8;</nowiki>

2020年9月17日 (四) 14:31的版本

任务切割

将原始的待爬目录表 分割成许多小份,当作许多小任务去完成。

敦煌网

import time

task_header ='../../task/dh_task/dh_task_' #header
def assign_task():
    task_content = ""  # 创建类别网址列表
    fp = open('dh_sub_category.csv', "rt")  # 打开csv文件
    
    count= 0
    num =0
    #类别名  类目级别  父类目级别

    s =set()#储存已有的类别
    for line in fp:  # 文件对象可以直接迭代
        count +=1
        task_content +=line
        
        if count%100 ==0:
            num += 1
            fw = open(task_header+str(num)+".csv","w",encoding="utf-8")
            fw.write(task_content)
            fw.close()
            task_content =""
        
    
    fw = open(task_header+str(num)+".csv","a",encoding="utf-8")
    fw.write(task_content)
    fw.close()
    task_content =""    
    fp.close()
    

if __name__ == '__main__':
    assign_task()


阿里巴巴

import time

task_header ='../../task/ali_task/ali_task_' #header
def assign_task():
    task_content = ""  # 创建类别网址列表
    fp = open('alibaba_categary.csv', "rt")  # 打开csv文件
    
    count= 0
    num =0
    #类别名  类目级别  父类目级别

    
    for line in fp:  # 文件对象可以直接迭代
        count +=1
        task_content +=line
        
        if count%100 == 0:
            num += 1
            fw = open(task_header+str(num)+".csv","w",encoding="utf-8")
            fw.write(task_content)
            fw.close()
            task_content =""
        
    if num <= 50:
        fw = open(task_header+str(num)+".csv","a",encoding="utf-8")
    else:
        fw = open(task_header+str(num+1)+".csv","a",encoding="utf-8")
        
    fw.write(task_content)
    fw.close()    
    task_content =""    
    fp.close()
    

if __name__ == '__main__':
    assign_task()
    


中国制造网

import time

task_header ='../../task/mc_task/mc_task_' #header
def assign_task():
    task_content = ""  # 创建类别网址列表
    fp = open('made_in_china_sub_cat.csv', "rt")  # 打开csv文件
    
    count= 0
    num =0
 
    for line in fp:  # 文件对象可以直接迭代
        count +=1
        task_content +=line
        
        if count%200 ==0:
            num += 1
            fw = open(task_header+str(num)+".csv","w",encoding="utf-8")
            fw.write(task_content)
            fw.close()
            task_content =""
        
    if num <= 100:
        fw = open(task_header+str(num)+".csv","a",encoding="utf-8")
    else:
        fw = open(task_header+str(num+1)+".csv","a",encoding="utf-8")
    
    fw.write(task_content)
    fw.close()
    task_content =""    
    fp.close()
    

if __name__ == '__main__':
    assign_task()
    


任务表建立

安装mysql

建立数据表

model设计和资源导入

MariaDB [crawler]> CREATE TABLE IF NOT EXISTS `task`(
    ->    `id` INT UNSIGNED AUTO_INCREMENT,
    ->    `site_title` VARCHAR(100) NOT NULL,
    ->    `task_name` VARCHAR(40) NOT NULL,
    ->    `task_status` INT UNSIGNED NOT NULL,
    ->    `availability zones` VARCHAR(60) NOT NULL,
    ->    `start_date` DATE,
    ->    PRIMARY KEY ( `id` ))ENGINE=InnoDB DEFAULT CHARSET=utf8;