
XML publishing separates content from its presentation. A useful pipeline validates source documents, transforms their structure, renders the chosen outputs, and checks what readers receive. The same content can serve a web page and a printed document while each format has its own layout rules.
Separate the publishing stages
Treat the source XML, schema, transformation, and renderer as separate versioned inputs. XML names describe your content model. XSLT selects and rearranges that content into an output structure. HTML and CSS present it on the web; XSL Formatting Objects (XSL-FO) describe paged output for a formatter.
The W3C XSL 1.1 recommendation specifies formatting objects. Apache FOP’s running guide documents a formatter that can consume FO and create PDF. A stylesheet that produces FO completes the transformation stage; the formatter still has to lay out the pages.
Publish the fictional catalog
Extract the XML practice pack. With Python 3.10 or later and lxml installed, run python publish_catalog.py. It validates catalog.xml, applies two supplied XSLT 1.0 stylesheets, and creates catalog.html and catalog.fo. Open the HTML in a browser. It should show Desk & Lamp with quantity 2 and Notebook with quantity 3.
If Apache FOP is installed and available on your command path, render the FO with:
fop -fo catalog.fo -pdf catalog.pdfUse fop.bat on Windows when that is the installed launcher. Follow FOP’s installation requirements for its Java runtime. The exercise provides FO for this optional step; it does not bundle FOP. Its primary reproducible output is the HTML, and browser Print to PDF is another way to inspect that layout.
The stylesheets use explicit namespace bindings and value extraction. Literal text is escaped by the serializer. Keep stylesheets trusted: transformation languages can load resources, and the sample disables file/network access during transformation.
Check the reader-facing result
| Area | Inspect | Useful failure signal |
|---|---|---|
| Completeness | Titles, rows, figures, notes and references | A source count differs from the delivered output. |
| Screen reading | Heading order, link labels, table headers | Content needs visual position alone to make sense. |
| Print layout | Page boundaries, repeated headers, long values | Clipped text, detached headings, or split identifiers. |
| Assets | Image resolution, fonts and licenses | Missing graphics or unexpected font substitution. |
| Navigation | Internal anchors and external links | Reference points to a removed or renamed target. |
Open every intended format. A valid XML result may contain all the words while arranging them badly. Test a long item name, many rows, an empty optional section, and non-ASCII text. Check search/copy behavior and reading order if accessible PDF is a requirement; selecting a PDF output format alone does not establish accessibility or archival conformance.
Release a reproducible build
Save source and template versions, tool versions, build commands, and diagnostics with each release. Use an output directory separate from editable sources. Compare generated output after a template change even when the content is unchanged.
For a large collection, validate sources first, then build into staging. Stop or quarantine failed documents according to a documented policy. Reconcile expected and produced filenames, inspect representative difficult pages, and publish assets before the pages that link to them.
When a screen layout and a print layout disagree, trace the first stage at which the content diverges. A missing node suggests selection or validation; a present but clipped node suggests layout. That distinction narrows the repair and prevents unnecessary changes to source content.