Some checks are pending
Security Scan / security-scan (push) Waiting to run
Dynamic XFA forms — real-estate forms, mortgage forms, government forms —
have been a silent failure mode for every tool in this server: the layout
and fields live in an XFA program that only Adobe's runtime executes, so
PyMuPDF/pdfium/MuPDF only see the "Open in Adobe Reader" placeholder page.
extract_form_data returned "document closed", convert_to_images returned
the placeholder, analyze_pdf_health reported total_pages=1 — all
correctly per the visible PDF, but all misleading about what the form
actually contains.
New capabilities:
- src/mcp_pdf/xfa.py — XFA detection, packet extraction, field parsing,
classification. No new deps (pypdf + stdlib ElementTree). Lifted from
a working prototype with attribution preserved; parameterized for
producer profiles.
- is_xfa_pdf — MCP tool to detect XFA presence + classify dynamic vs
static. Use for branching BEFORE extract_form_data or convert_to_images.
- extract_xfa_fields — MCP tool that parses the XFA template for field
names, captions, UI types. Splits fields into shared (cross-form
canonical Global_Info-* vocabulary), positional (opaque codes like
p01tf022), and plumbing (producer internals, dropped). Defaults to
the zipForm producer profile; pass profile="generic" + custom regex
patterns for other producers. canonical_separator selects _ / . / -.
UX fixes for existing tools (no more cryptic failures on dynamic XFA):
- extract_form_data — diagnoses dynamic XFA and returns
{error, hint: "extract_xfa_fields"} instead of "document closed"
- convert_to_images — still produces the rendered image (caller may
want the placeholder), but now warns that it's not the real form
- analyze_pdf_health — surfaces is_xfa + xfa_type in document_stats,
adds a warning for dynamic XFA
Cross-tool alignment:
- extract_form_data field-type strings aligned to the same six-term
vocabulary as XFA (text/checkbox/radio/dropdown/date/signature +
button/unknown edge categories). listbox + combobox both collapse
to "dropdown" — the widget-hover distinction wasn't semantic.
Tests: 31/31 passing against the synthetic XFA fixture
(tests/fixtures/xfa/synthetic_dynamic_xfa.pdf). The fixture is
hand-built, license-clean, ~2 KB, and exercises all three classification
categories including the "denylist beats shared-prefix" invariant
(Global_Info-Invisibind-Test drops despite the shared prefix).
Build hygiene:
- .gitignore exception !tests/fixtures/**/*.pdf so fixtures survive
the global *.pdf rule
- pyproject.toml force-include for the fixture (uses hatchling's
force-include not include — the latter is restrictive, not additive)
Coordinated via the agent-thread protocol; full design history is
in docs/agent-threads/xfa-form-support/ (excluded from the sdist).
XFA test fixtures
synthetic_dynamic_xfa.pdf
A minimal, license-clean dynamic XFA form for regression-testing XFA
detection and extract_xfa_fields. Hand-built; contains no copyrighted form
content. ~2 KB.
Structure:
- 1 static "Please wait... requires Adobe Reader 8 or higher" placeholder page (what a real dynamic XFA shows to non-Adobe readers — exercises detection).
- Catalog
/NeedsRendering true— the canonical dynamic-XFA marker. - XFA packet in array form
[template, datasets](exercises the name/stream-pair parsing path, not the single-stream path). - An 8-field XFA
<template>covering all three categories:
| field name | category | exercises |
|---|---|---|
Global_Info-Test-Property-Address |
shared | canonicalization -> test_property_address; textEdit -> text; in-field <caption> harvest ("Property Address") |
Global_Info-Test-Buyer-Name |
shared | shared + checkButton -> checkbox |
p01tf001 |
positional | opaque code, page-1 text |
p01cb002 |
positional | opaque code, page-1 checkbox |
p02te001 |
positional | different page prefix — regex must not over-match |
tbd1 |
plumbing (dropped) | ^tbd\d+$ denylist |
folder_id |
plumbing (dropped) | exact-name denylist |
Global_Info-Invisibind-Test |
plumbing (dropped) | denylist precedence — dropped despite the Global_Info- shared prefix |
Expected classify() result: 2 shared, 3 positional, 3 plumbing dropped, 1
caption harvested. (Verified with the prototype xfa_parse.py + field_naming.py.)