2016-10-30 3 views
1

저는 PySpark를 사용하여 Eratosthenes의 체를 구현하려고합니다.PySpark RDD 필터링 된 요소가 다시 나타납니다

이 경우 나는 많은 RDD에 filter을 적용하려고 시도하고 있지만, 각 반복마다 이전 반복에서 필터링 된 내용이 계속 되돌려지고 그 이유가 궁금합니다.

from math import ceil 
from math import sqrt 

min_number = 2 
max_number = 101 

rdd = sc.parallelize(range(min_number, max_number), 4) 
pivot = min_number 

max_pivot = ceil(sqrt(max_number)) 

while pivot <= max_pivot: 
    print "RDD for pivot = " + str(pivot) + ":" 
    rdd = rdd.filter(lambda x: x <= pivot or x % pivot != 0) 
    pivot = rdd.filter(lambda x: x > pivot).reduce(min) 
    rdd.collect() 

그리고 출력 : 여기에

코드의 당신이 볼 수 있듯이, 각 반복에, 현재의 피벗의 배수가 필터링되고

Pivot = 2 
[2, 3, 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19, 20, 22, 23, 25, 26, 28, 29, 31, 32, 34, 35, 37, 38, 40, 41, 43, 44, 46, 47, 49, 50, 52, 53, 55, 56, 58, 59, 61, 62, 64, 65, 67, 68, 70, 71, 73, 74, 76, 77, 79, 80, 82, 83, 85, 86, 88, 89, 91, 92, 94, 95, 97, 98, 100] 
Pivot = 3 
[2, 3, 4, 5, 6, 7, 9, 10, 11, 13, 14, 15, 17, 18, 19, 21, 22, 23, 25, 26, 27, 29, 30, 31, 33, 34, 35, 37, 38, 39, 41, 42, 43, 45, 46, 47, 49, 50, 51, 53, 54, 55, 57, 58, 59, 61, 62, 63, 65, 66, 67, 69, 70, 71, 73, 74, 75, 77, 78, 79, 81, 82, 83, 85, 86, 87, 89, 90, 91, 93, 94, 95, 97, 98, 99] 
Pivot = 4 
[2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 14, 16, 17, 18, 19, 21, 22, 23, 24, 26, 27, 28, 29, 31, 32, 33, 34, 36, 37, 38, 39, 41, 42, 43, 44, 46, 47, 48, 49, 51, 52, 53, 54, 56, 57, 58, 59, 61, 62, 63, 64, 66, 67, 68, 69, 71, 72, 73, 74, 76, 77, 78, 79, 81, 82, 83, 84, 86, 87, 88, 89, 91, 92, 93, 94, 96, 97, 98, 99] 
Pivot = 5 
[2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 25, 26, 27, 28, 29, 31, 32, 33, 34, 35, 37, 38, 39, 40, 41, 43, 44, 45, 46, 47, 49, 50, 51, 52, 53, 55, 56, 57, 58, 59, 61, 62, 63, 64, 65, 67, 68, 69, 70, 71, 73, 74, 75, 76, 77, 79, 80, 81, 82, 83, 85, 86, 87, 88, 89, 91, 92, 93, 94, 95, 97, 98, 99, 100] 
Pivot = 6 
[2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, 26, 27, 29, 30, 31, 32, 33, 34, 36, 37, 38, 39, 40, 41, 43, 44, 45, 46, 47, 48, 50, 51, 52, 53, 54, 55, 57, 58, 59, 60, 61, 62, 64, 65, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 78, 79, 80, 81, 82, 83, 85, 86, 87, 88, 89, 90, 92, 93, 94, 95, 96, 97, 99, 100] 
Pivot = 7 
[2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 18, 19, 20, 21, 22, 23, 25, 26, 27, 28, 29, 30, 31, 33, 34, 35, 36, 37, 38, 39, 41, 42, 43, 44, 45, 46, 47, 49, 50, 51, 52, 53, 54, 55, 57, 58, 59, 60, 61, 62, 63, 65, 66, 67, 68, 69, 70, 71, 73, 74, 75, 76, 77, 78, 79, 81, 82, 83, 84, 85, 86, 87, 89, 90, 91, 92, 93, 94, 95, 97, 98, 99, 100] 
Pivot = 8 
[2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 28, 29, 30, 31, 32, 33, 34, 35, 37, 38, 39, 40, 41, 42, 43, 44, 46, 47, 48, 49, 50, 51, 52, 53, 55, 56, 57, 58, 59, 60, 61, 62, 64, 65, 66, 67, 68, 69, 70, 71, 73, 74, 75, 76, 77, 78, 79, 80, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 93, 94, 95, 96, 97, 98, 100] 
Pivot = 9 
[2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28, 29, 31, 32, 33, 34, 35, 36, 37, 38, 39, 41, 42, 43, 44, 45, 46, 47, 48, 49, 51, 52, 53, 54, 55, 56, 57, 58, 59, 61, 62, 63, 64, 65, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 77, 78, 79, 81, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 93, 94, 95, 96, 97, 98, 99] 
Pivot = 10 
[2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 100] 
Pivot = 11 
[2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 97, 98, 99, 100] 

만했다 번호 각 반복에서 rdd 참조를 바꿀 때도 이미 필터링되어 다시 필터링됩니다.

도움이 필요한 경우 Mac 용 Python 2.7.10에서 PySpark 2.0.1을 실행하고 있습니다.

감사합니다.

+0

그래서, 당신은 피벗을 사용하여 값을 제외하려고? –

+0

@AlbertoBonsanto 네, 맞습니다. –

답변

3

파이썬 클로저는 함수가 호출 될 때 평가되며, 작성 될 때 평가되지 않습니다 (late binding). 두번째로

(sc.parallelize(range(min_number, max_number), 4) 
    .filter(lambda x: x <= 2 or x % 2 != 0)) 

: 첫번째 반복 rdd의 결과

같이 평가 세번째의

(sc.parallelize(range(min_number, max_number), 4) 
    .filter(lambda x: x <= 3 or x % 3 != 0) 
    .filter(lambda x: x <= 3 or x % 3 != 0)) 

:

(sc.parallelize(range(min_number, max_number), 4) 
    .filter(lambda x: x <= 4 or x % 4 != 0) 
    .filter(lambda x: x <= 4 or x % 4 != 0) 
    .filter(lambda x: x <= 4 or x % 4 != 0)) 

매번 pivot은 현재 범위에서 해결됩니다.

올바른 구현 :

while pivot <= max_pivot: 
    def f(x, pivot=pivot): 
     return x <= pivot or x % pivot != 0 

    rdd = rdd.filter(f) 
    pivot = rdd.filter(lambda x: x > pivot).min() 
관련 문제