首頁 >後端開發 >Python教學 >Python開發MapReduce系列之WordCount Demo

Python開發MapReduce系列之WordCount Demo

ringa_lee原創: 2017-09-17 09:28:381828瀏覽

我們知道MapReduce是hadoop這隻大象的核心，Hadoop中，資料處理核心就是 MapReduce 程式設計模型。一個Map/Reduce 通常會把輸入的資料集切分為若干獨立的資料區塊，由 map任務（task）以完全並行的方式處理它們。框架會對map的輸出先進行排序，然後把結果輸入給reduce任務。通常作業的輸入和輸出都會儲存在檔案系統中。因此，我們的程式設計中心主要是 mapper階段和reducer階段。

下面來從零開發一個MapReduce程序，並在hadoop叢集上運行。
mapper程式碼map.py：

 import sys    
    for line in sys.stdin:
        word_list = line.strip().split(&#39; &#39;)    
        for word in word_list:            print &#39;\t&#39;.join([word.strip(), str(1)])

#View Code

reducer程式碼reduce.py：

 import sys
    
    cur_word = None
    sum = 0    
    for line in sys.stdin:
        ss = line.strip().split(&#39;\t&#39;)        
        if len(ss) < 2:            continue
    
        word = ss[0].strip()
        count = ss[1].strip()    
        if cur_word == None:
            cur_word = word    
        if cur_word != word:            print &#39;\t&#39;.join([cur_word, str(sum)])
            cur_word = word
            sum = 0
        
        sum += int(count)    
    print &#39;\t&#39;.join([cur_word, str(sum)])
    sum = 0

View Code

資源檔src.txt（測試用，在叢集中跑時，記得上傳到hdfs）：

hello    
    ni hao ni haoni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ao ni haoni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni haoao ni haoni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao ni hao
    Dad would get out his mandolin and play for the family
    Dad loved to play the mandolin for his family he knew we enjoyed singing
    I had to mature into a man and have children of my own before I realized how much he had sacrificed
    I had to,mature into a man and,have children of my own before.I realized how much he had sacrificed

View Code

首先本地偵錯查看結果是否正確，輸入命令以下：

cat src.txt | python map.py | sort -k 1 | python reduce.py

命令列中輸出的結果：

a    2
    and    2
    and,have    1
    ao    1
    before    1
    before.I    1
    children    2
    Dad    2
    enjoyed    1
    family    2
    for    2
    get    1
    had    4
    hao    33
    haoao    1
    haoni    3
    have    1
    he    3
    hello    1
    his    2
    how    2
    I    3
    into    2
    knew    1
    loved    1
    man    2
    mandolin    2
    mature    1
    much    2
    my    2
    ni    34
    of    2
    out    1
    own    2
    play    2
    realized    2
    sacrificed    2
    singing    1
    the    2
    to    2
    to,mature    1
    we    1
    would    1

View Code

透過調試發現本地調試，程式碼是OK的。下面扔到集群上面跑。為了方便，專門寫了一個腳本 run.sh，解放勞動力嘛。

HADOOP_CMD="/home/hadoop/hadoop/bin/hadoop"
    STREAM_JAR_PATH="/home/hadoop/hadoop/contrib/streaming/hadoop-streaming-1.2.1.jar"
    
    INPUT_FILE_PATH="/home/input/src.txt"
    OUTPUT_PATH="/home/output"
    
    $HADOOP_CMD fs -rmr  $OUTPUT_PATH 
    
    $HADOOP_CMD jar $STREAM_JAR_PATH \        -input $INPUT_FILE_PATH \        -output $OUTPUT_PATH \        
    -mapper "python map.py" \        -reducer "python reduce.py" \        -file ./map.py \        -file ./reduce.py

下面解析下腳本：

　HADOOP_CMD： hadoop的bin的路径
    STREAM_JAR_PATH：streaming jar包的路径
    INPUT_FILE_PATH：hadoop集群上的资源输入路径
    OUTPUT_PATH：hadoop集群上的结果输出路径。（注意：这个目录不应该存在的，因此在脚本加了先删除这个目录。**注意****注意****注意**：若是第一次执行，没有这个目录，会报错的。可以先手动新建一个新的output目录。）
    $HADOOP_CMD fs -rmr  $OUTPUT_PATH
    
    $HADOOP_CMD jar $STREAM_JAR_PATH \        -input $INPUT_FILE_PATH \        -output $OUTPUT_PATH \       
     -mapper "python map.py" \        -reducer "python reduce.py" \       
      -file ./map.py \        -file ./reduce.py                 
      #这里固定格式，指定输入，输出的路径；指定mapper，reducer的文件；
      #并分发mapper，reducer角色的我们用户写的代码文件，因为集群其他的节点还没有mapper、reducer的可执行文件。

# 輸入以下指令查看經過reduce階段後輸出的記錄：

cat src.txt | python map.py | sort -k 1 | python reduce.py | wc -l
命令行中输出：43

在瀏覽器輸入：master：50030 查看任務的詳細情況。

Kind    % Complete    Num Tasks    Pending    Running    Complete    Killed     Failed/Killed Task Attempts
map       100.00%        2            0        0        2            0            0 / 0
reduce    100.00%        1            0        0        1            0            0 / 0

Map-Reduce Framework中看到這個。

Counter                    　　Map    Reduce    Total
Reduce output records    0    　　0       　　 43

證明整個過程成功。第一個hadoop程式開發結束。

以上是Python開發MapReduce系列之WordCount Demo的詳細內容。更多資訊請關注PHP中文網其他相關文章！

陳述：

本文內容由網友自願投稿，版權歸原作者所有。本站不承擔相應的法律責任。如發現涉嫌抄襲或侵權的內容，請聯絡admin@php.cn

上一篇：基於Python3.4實作簡單抓取爬蟲功能詳細介紹下一篇：基於Python3.4實作簡單抓取爬蟲功能詳細介紹

看更多