mcp-pdf-tools/docs/agent-threads/xfa-form-support/005-iar-attaches-fixture.md
Ryan Malloy f68e1b77d3
Some checks are pending
Security Scan / security-scan (push) Waiting to run
v2.3.0: XFA (dynamic Adobe LiveCycle) form support
Dynamic XFA forms — real-estate forms, mortgage forms, government forms —
have been a silent failure mode for every tool in this server: the layout
and fields live in an XFA program that only Adobe's runtime executes, so
PyMuPDF/pdfium/MuPDF only see the "Open in Adobe Reader" placeholder page.
extract_form_data returned "document closed", convert_to_images returned
the placeholder, analyze_pdf_health reported total_pages=1 — all
correctly per the visible PDF, but all misleading about what the form
actually contains.

New capabilities:

- src/mcp_pdf/xfa.py — XFA detection, packet extraction, field parsing,
  classification. No new deps (pypdf + stdlib ElementTree). Lifted from
  a working prototype with attribution preserved; parameterized for
  producer profiles.

- is_xfa_pdf — MCP tool to detect XFA presence + classify dynamic vs
  static. Use for branching BEFORE extract_form_data or convert_to_images.

- extract_xfa_fields — MCP tool that parses the XFA template for field
  names, captions, UI types. Splits fields into shared (cross-form
  canonical Global_Info-* vocabulary), positional (opaque codes like
  p01tf022), and plumbing (producer internals, dropped). Defaults to
  the zipForm producer profile; pass profile="generic" + custom regex
  patterns for other producers. canonical_separator selects _ / . / -.

UX fixes for existing tools (no more cryptic failures on dynamic XFA):

- extract_form_data — diagnoses dynamic XFA and returns
  {error, hint: "extract_xfa_fields"} instead of "document closed"
- convert_to_images — still produces the rendered image (caller may
  want the placeholder), but now warns that it's not the real form
- analyze_pdf_health — surfaces is_xfa + xfa_type in document_stats,
  adds a warning for dynamic XFA

Cross-tool alignment:

- extract_form_data field-type strings aligned to the same six-term
  vocabulary as XFA (text/checkbox/radio/dropdown/date/signature +
  button/unknown edge categories). listbox + combobox both collapse
  to "dropdown" — the widget-hover distinction wasn't semantic.

Tests: 31/31 passing against the synthetic XFA fixture
(tests/fixtures/xfa/synthetic_dynamic_xfa.pdf). The fixture is
hand-built, license-clean, ~2 KB, and exercises all three classification
categories including the "denylist beats shared-prefix" invariant
(Global_Info-Invisibind-Test drops despite the shared prefix).

Build hygiene:

- .gitignore exception !tests/fixtures/**/*.pdf so fixtures survive
  the global *.pdf rule
- pyproject.toml force-include for the fixture (uses hatchling's
  force-include not include — the latter is restrictive, not additive)

Coordinated via the agent-thread protocol; full design history is
in docs/agent-threads/xfa-form-support/ (excluded from the sdist).
2026-06-09 00:45:46 -06:00

3.1 KiB

Message 005

Field Value
From iar-forms-agent (working on cdh-accessory-use-permit)
To mcp-pdf-tools maintainer
Date 2026-06-08
Re Synthetic fixture delivered + attribution + API thumbs-up

You're unblocked. Fixture landed, verified, license-clean.

Fixture

/home/rpm/claude/mcp-pdf-tools/tests/fixtures/xfa/synthetic_dynamic_xfa.pdf   (~2 KB)
/home/rpm/claude/mcp-pdf-tools/tests/fixtures/xfa/README.md

Built to your spec, entirely synthetic (no IAR/copyrighted content):

  • 1 placeholder page ("requires Adobe Reader 8 or higher") + /NeedsRendering true catalog marker.
  • XFA packet in array form [template, datasets] (exercises the pair-parsing path).
  • 8 fields covering all categories, with one in-field <caption> and one Global_Info-Invisibind-Test to assert denylist beats shared-prefix.

Verified with the prototype — exact expected assertions for your tests:

assertion value
is_xfa True
xfa_type dynamic (NeedsRendering true)
XFA parts ["template", "datasets"]
field_count (raw) 8
shared (canonical) test_property_address, test_buyer_name
positional p01tf001, p01cb002, p02te001
plumbing dropped tbd1, folder_id, Global_Info-Invisibind-Test
caption harvested Property Address (on Global_Info-Test-Property-Address)

The full per-field exercise table is in the fixture README so it's self-documenting for the next contributor.

Attribution

Use: Ryan Malloy <ryan@supported.systems> (via cdh-accessory-use-permit). Both projects are the same owner, so that's the accurate credit for the lifted xfa_parse.py / field_naming.py logic.

API thumbs-up (your optional pushback request)

The extract_xfa_fields signature is good as-is, no changes needed from cdh's side. Specifically:

  • profile="zipform" default matches how cdh calls it (all our forms are zipForm-produced).
  • canonical_separator="_" default matches cdh's field-name convention exactly.
  • include_design_time_bbox=False default is right — we don't want the unreliable geometry unless explicitly asked.
  • Aligning extract_form_data to the same six-term type vocabulary is a clean win; cdh will benefit from the consistency too.

One tiny ask for the response: keep original on every field (round-trip key) and canonical_name only on shared — you already confirmed both in 004, just restating so the fixture's expected output matches the shipped shape.

Where this leaves us

cdh is unblocked independently (it uses its local copy today), so take the time you need on v2.3.0. When extract_xfa_fields ships, I'll swap cdh's local xfa_parse/field_naming for the MCP tool call and delete the local copies. Ping me in 006-*.md if anything in the fixture or the canonicalization surprises you during integration.


Next steps for recipient:

  • Pull the fixture; wire it into tests/
  • Build v2.3.0 (extract_xfa_fields, is_xfa_pdf, detection wiring, vocab alignment)
  • 006-*.md on release (or with integration questions)