GHSA-8mpw-7fpc-4gqjMedium

NLTK: Pl196xCorpusReader has quadratic ReDoS on malformed TEI blocks

Published
September 8, 2026
Last Modified
September 8, 2026

🔗 CVE IDs covered (1)

📋 Description

Summary

Pl196xCorpusReader still parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.

Details

  • Vulnerability type: Regular-expression denial of service
  • Affected component: nltk.corpus.reader.pl196x.TEICorpusView.read_block and Pl196xCorpusReader public methods
  • Affected versions: Published 3.9.4 and current source v3.10.0-rc2 both reproduced.
  • Patched versions: Not yet patched
  • Root cause: Lazy .*? whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.

The parser uses regexes for paragraphs, sentences, and word tags across the whole <text> block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed <p> tags doubled, through normal public calls such as words() and tagged_words().

PoC

Preconditions

  • The application parses attacker-influenced PL196X or TEI-like corpus files through public reader APIs.

Steps

  1. Create a corpus file with a valid header followed by a <text> block that contains many opening tags and no matching closing tags.
  2. Instantiate Pl196xCorpusReader on that corpus.
  3. Call words() or tagged_words() and measure elapsed time as the malformed tag count doubles.
  4. Observe near quadratic growth instead of near-linear behavior.

Minimal reproducible excerpt

size=1000 0.014s
size=2000 0.057s
size=4000 0.231s
size=8000 0.927s

Impact

A consumer that accepts attacker-influenced corpus files can be forced into heavy CPU use and parser-thread stalling before the application concludes the input contains no valid content.

Remediation

Replace the whole-block lazy-regex parser with a linear parser or bounded tokenizer, and add regression tests that assert near-linear behavior on malformed inputs with many unmatched tags.

🎯 Affected products1

  • pip/nltk:<= 3.10.2

🔗 References (7)