
本文详解如何修正NLTK上下文无关文法(CFG)中因变量未正确插入导致的ValueError: Grammar does not cover some of the input words错误,通过动态构建CFG、合理设计词性规则与生成逻辑,实现真正可解析的随机诗歌生成。
本文详解如何修正nltk上下文无关文法(cfg)中因变量未正确插入导致的`valueerror: grammar does not cover some of the input words`错误,通过动态构建cfg、合理设计词性规则与生成逻辑,实现真正可解析的随机诗歌生成。
你在使用 NLTK 的 CFG 生成随机句子时遇到的报错:
ValueError: Grammar does not cover some of the input words: "'unimprovedness'"
根本原因在于:你试图用 parser.parse(word_tokenize(chosen_word)) 去解析一个孤立单词(如 'unimprovedness'),但当前 CFG 中并未将该词定义为任何词性(如 N、V 等),且 'chosen_word' 被当作字面字符串写入了语法规则,而非其实际值。
例如,原始代码中这行:
N -> 'cat' | 'dog' | ... | 'chosen_word' | 'Fabrikoid'
实际注册的是字面量 'chosen_word'(即字符 c-h-o-s-e-n-_-w-o-r-d),而不是变量 chosen_word 的值(比如 'serendipity')。因此当 chosen_word = 'unimprovedness' 时,文法里根本没有 'unimprovedness' 这个终结符,自然无法覆盖,解析失败。
WorkBuddy 4.22.10 是腾讯推出的 AI 智能体桌面工作台,主打“下载即用”的零部署体验。它能通过自然语言指令,自主规划并执行文档处理、数据分析、PPT 生成等复杂任务。支持多模型切换与技能扩展,可像同事一样帮你完成重复性工作,大幅提升办公效率。
✅ 正确做法是:动态构建 CFG 字符串,将 chosen_word 的实际值以合法终结符形式注入到对应非终结符(如 N)的规则中。同时,需确保所选单词在语法中具有明确的词性归类(此处我们统一视作名词 N)。
以下是修复后的完整可运行教程代码:
import nltk
from nltk.corpus import words
from nltk.tokenize import word_tokenize
from nltk.grammar import CFG
from nltk.parse import ChartParser
from random import choice
# 下载必要资源(首次运行需启用)
nltk.download('words')
nltk.download('punkt')
# 随机选取一个英文名词(简化:取长度3–8、全小写的常见词)
word_list = [w.lower() for w in words.words() if w.isalpha() and 3 NP VP
NP -> Det N | N
VP -> V NP | V
Det -> 'the' | 'a' | 'an'
N -> 'cat' | 'dog' | 'bird' | 'tree' | 'flower' | '{chosen_word}' | 'Fabrikoid'
V -> 'sings' | 'walks' | 'flies' | 'grows' | 'blooms' | 'dances' | 'whispers'
"""
grammar = CFG.fromstring(grammar_str)
parser = ChartParser(grammar)
# ✅ 改用生成式方法:不依赖 parse() 反向推导,而是正向递归展开树(更稳定)
def generate_sentence(tree):
if tree.height() == 2: # 叶子节点(终结符)
return tree.leaves()[0]
else:
return ' '.join(generate_sentence(subtree) for subtree in tree)
# 尝试生成最多10次,避免死循环(某些词可能因语法限制难生成完整S)
sentence = ""
for _ in range(10):
try:
# 从起始符号 S 开始随机生成一棵合法句法树
trees = list(parser.parse(['the', 'cat', 'sings'])) # 占位输入,仅用于触发生成器(不推荐)
# 更健壮的做法:使用 nltk.parse.generate(需 NLTK ≥ 3.8.1)
from nltk.parse.generate import generate
for sent in generate(grammar, n=1, depth=5):
sentence = ' '.join(sent).capitalize() + '.'
break
if sentence:
break
except Exception:
continue
if not sentence:
# 保底方案:手工组合一句符合语法的诗行
det = choice(['The', 'A', 'An'])
noun = choice(['cat', 'dog', 'bird', chosen_word, 'Fabrikoid'])
verb = choice(['sings', 'walks', 'flies', 'grows', 'blooms', 'dances', 'whispers'])
sentence = f"{det} {noun} {verb}."
print("Here is your poem:")
print(sentence)
? 关键修正点与注意事项:
- 不要对单个词调用 parser.parse():ChartParser.parse() 期望输入是符合文法结构的词序列(如 ['the', 'cat', 'sings']),而非单个未标注词性、无上下文的单词。原逻辑 parser.parse(word_tokenize(chosen_word)) 是根本性误用。
- 用 f-string 动态注入词汇:确保 '{chosen_word}' 在字符串中被正确展开为实际单词(如 'serendipity'),并置于 N -> ... 规则内,使其成为合法名词终结符。
- 增强语法鲁棒性:扩展 NP 和 VP 规则(如 NP -> N 允许无冠词名词短语;VP -> V 允许不及物动词),避免因规则过严导致生成失败。
- 推荐使用 nltk.parse.generate:它是专为“从CFG正向生成句子”设计的接口,比逆向解析更直观可靠(需确认 NLTK 版本 ≥ 3.8.1)。
- 添加降级策略:当自动生成失败时,提供简洁的手工组合方案,保障程序始终有输出。
? 小结:NLTK 的 CFG 是静态规则系统,无法自动理解词汇语义或词性。要让随机词“合法”,必须显式将其纳入对应词性规则中,并通过生成(而非解析)方式构造句子。掌握 f-string + generate() 组合,是构建可控随机语言生成器的核心技巧。










