2014-02-07 3 views
3

ValueError 메시지가 표시되고 잘못된 작업을하거나 Python 설치에 오류가 있는지 확실하지 않습니다. 나는 문서가 픽션인지 논픽션인지를 결정하기위한 테스트를 개발하려고합니다. 내 코드는 다음과 같습니다 {'contains(girls)': True, 'contains(farm)': True, 'contains(new)': True, 'contains(left)': True, 'contains(days)': True, 'contains(work)': True, 'contains(stood)': True, 'contains("")': True, 'contains(subject)': True, 'contains(might)': True, 'contains(mrs)': False, 'contains(like)': True, 'contains(father)': True, 'contains(said)': True, 'contains(taken)': True, 'contains(little)': True, 'contains(every)': True, 'contains(first)': True, 'contains(."")': True, 'contains(uncle)': False, 'contains(close)': True, 'contains(week)': True, 'contains(women)': True, 'contains(interest)': True, 'contains(sally)': False, 'contains(body)': True, 'contains(life)': True, 'contains(home)': True, 'contains(nonfiction)': True, 'contains(spite)': True, 'contains(read)': True, 'contains(done)': True, 'contains(travis)': False, 'contains(place)': True, 'contains(woman)': True, 'contains(!"")': True, 'contains(old)': True, 'contains(boy)': True, 'contains(know)': True, 'contains(made)': True, 'contains(together)': True, 'contains(farmer)': True, 'contains(make)': True, 'contains(great)': True, 'contains(upon)': True, 'contains(men)': True, 'contains(hand)': True, 'contains(time)': True, 'contains(always)': True, 'contains(fiction)': True, 'contains(back)': True, 'contains(two)': True, 'contains(mother)': True, 'contains(would)': True, 'contains(country)': True, 'contains(put)': True, 'contains(,"")': True, 'contains(never)': True, 'contains(.")': True, 'contains(well)': True, 'contains(think)': True, 'contains(living)': True, 'contains(man)': True, 'contains(came)': True, 'contains(fruit)': True, 'contains(year)': True, 'contains(state)': True, 'contains(years)': True, 'contains(may)': True, 'contains(something)': True, 'contains(\x97)': True, 'contains(esther)': False, 'contains(,")': True, 'contains(get)': True, 'contains(children)': True, 'contains(many)': True, 'contains(better)': True, 'contains(away)': True, 'contains(spring)': True, 'contains(last)': True, 'contains(long)': True, 'contains(food)': True, 'contains(summer)': True, 'contains(girl)': True, 'contains(paper)': True, 'contains(city)': True, 'contains(could)': True, 'contains(come)': True, 'contains(part)': True, 'contains(see)': True, 'contains(wife)': True, 'contains(keep)': True, 'contains(along)': True, 'contains(even)': True, 'contains(people)': True, 'contains(best)': True, 'contains(good)': True, 'contains(day)': True, 'contains(season)': True, 'contains(one)': True} Traceback (most recent call last): File "fiction.py", line 44, in <module> classifier = nltk.NaiveBayesClassifier.train(train_set) File "/usr/local/lib/python2.6/dist-packages/nltk/classify/naivebayes.py", line 214, in train label_probdist = estimator(label_freqdist) File "/usr/local/lib/python2.6/dist-packages/nltk/probability.py", line 898, in __init__ LidstoneProbDist.__init__(self, freqdist, 0.5, bins) File "/usr/local/lib/python2.6/dist-packages/nltk/probability.py", line 782, in __init__ 'must have at least one bin.') ValueError: A ELE probability distribution must have at least one bin.NLTK 분류를 사용하는 ValueError

내가 각 카테고리에서 가장 많이 사용되는 단어에 대한 확률을 수신해야하지만 오류를 받고 있어요 :

import nltk, re, string 
from nltk.corpus import CategorizedPlaintextCorpusReader 
corpus_root = './nltk_data/corpora/fiction' 

fiction = CategorizedPlaintextCorpusReader(corpus_root, r'(\w+)/*.txt', cat_file='cat.txt') 
fiction.categories() 
['fic', 'nonfic'] 

documents = [(list(fiction.words(fileid)), category) 
    for category in fiction.categories() 
    for fileid in fiction.fileids(category)] 

all_words=nltk.FreqDist(
    w.lower() 
    for w in fiction.words() 
    if w.lower() not in nltk.corpus.stopwords.words('english') and w.lower() not in string.punctuation) 
word_features = all_words.keys()[:100] 

def document_features(document): # [_document-classify-extractor] 
    document_words = set(document) # [_document-classify-set] 
    features = {} 
    for word in word_features: 
     features['contains(%s)' % word] = (word in document_words) 
    return features 
#print document_features(fiction.words('fic/11.txt')) 

featuresets = [(document_features(d), c) for (d,c) in documents] 
train_set, test_set = featuresets[100:], featuresets[:100] 
classifier = nltk.NaiveBayesClassifier.train(train_set) 

print nltk.classify.accuracy(classifier, test_set) 
classifier.show_most_informative_features(5) 

나는 반환에 다음을 얻고있다.

+0

어디에서 '허구'코퍼스를 얻었습니까? – alvas

답변

2

당신이 물어 본 지 꽤 오래되었습니다. 그러나, 나는 당신의 'featuresets'이 100 개 미만의 레코드라고 확신합니다. test_set이 아마도 빈 목록 일 것이기 때문에 이것을 얻을 것이다.