Summary
Pl196xCorpusReader still parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.
Details
- Vulnerability type: Regular-expression denial of service
- Affected component:
nltk.corpus.reader.pl196x.TEICorpusView.read_block and Pl196xCorpusReader public methods
- Affected versions: Published
3.9.4 and current source v3.10.0-rc2 both reproduced.
- Patched versions: Not yet patched
- Root cause: Lazy
.*? whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.
The parser uses regexes for paragraphs, sentences, and word tags across the whole <text> block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed <p> tags doubled, through normal public calls such as words() and tagged_words().
PoC
Preconditions
- The application parses attacker-influenced PL196X or TEI-like corpus files through public reader APIs.
Steps
- Create a corpus file with a valid header followed by a
<text> block that contains many opening tags and no matching closing tags.
- Instantiate
Pl196xCorpusReader on that corpus.
- Call
words() or tagged_words() and measure elapsed time as the malformed tag count doubles.
- Observe near quadratic growth instead of near-linear behavior.
Minimal reproducible excerpt
size=1000 0.014s
size=2000 0.057s
size=4000 0.231s
size=8000 0.927s
Impact
A consumer that accepts attacker-influenced corpus files can be forced into heavy CPU use and parser-thread stalling before the application concludes the input contains no valid content.
Remediation
Replace the whole-block lazy-regex parser with a linear parser or bounded tokenizer, and add regression tests that assert near-linear behavior on malformed inputs with many unmatched tags.
References
Summary
Pl196xCorpusReaderstill parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.Details
nltk.corpus.reader.pl196x.TEICorpusView.read_blockandPl196xCorpusReaderpublic methods3.9.4and current sourcev3.10.0-rc2both reproduced..*?whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.The parser uses regexes for paragraphs, sentences, and word tags across the whole
<text>block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed<p>tags doubled, through normal public calls such aswords()andtagged_words().PoC
Preconditions
Steps
<text>block that contains many opening tags and no matching closing tags.Pl196xCorpusReaderon that corpus.words()ortagged_words()and measure elapsed time as the malformed tag count doubles.Minimal reproducible excerpt
Impact
A consumer that accepts attacker-influenced corpus files can be forced into heavy CPU use and parser-thread stalling before the application concludes the input contains no valid content.
Remediation
Replace the whole-block lazy-regex parser with a linear parser or bounded tokenizer, and add regression tests that assert near-linear behavior on malformed inputs with many unmatched tags.
References