Page images and preview¶
The core writes PDF in managed code, with no native dependency. Two optional packages add what needs a
rasteriser: Rustaveli.Pdf.Raster draws pages as images, SVG and XPS through SkiaSharp, and recompresses images for
PDF export; Rustaveli.Pdf.Preview shows those images in the browser while a document is being written. This page
explains how both work. The guides show how to use them: Output and
Preview and debugging.
One layout, several surfaces¶
Layout never depends on how its result is drawn (How layout works). The typesetter plans and renders every page through one drawing seam, and a page image is simply a different surface behind that seam:
flowchart LR
doc["Document"] --> ts["Typesetter<br/>same passes, same measurer"]
ts --> pdf["PDF surface"] --> pdfout["PDF"]
ts --> skia["Skia surface"]
skia --> raster["Raster target"] --> img["PNG, JPEG or WebP"]
skia --> svg["SVG target"] --> svgout["One SVG per page"]
skia --> xps["XPS target"] --> xpsout["One XPS document (Windows)"]
An image export measures text with the same shaper, built from the same
TypefaceLibrary,
as a PDF export does. Lines therefore break at the same words and pages at the same places, and the counting passes
that settle "page 3 of 12" run the same way. Given the same typefaces, a page image shows the PDF's layout, not an
approximation of it.
The Skia surface — the internal SkiaRasterSurface — draws onto one of three targets: pixels, an SVG canvas or an
XPS document. The surface does the drawing; the target decides what a page becomes when it ends.
Text: the same glyphs, not the same string¶
Skia could shape text itself, but then the image would show Skia's idea of the text rather than the layout's. The surface never gives Skia a string. It walks the same shaped glyphs the layout measured and the PDF surface writes — glyph numbers, advances, kerning and offsets — and places each glyph where the layout put it, as a positioned run, one run per typeface. Each face is loaded into Skia from the very bytes the core parsed, so Skia draws the same font file, not whatever the system has of the same name.
Glyphs are drawn without hinting, with linear metrics and subpixel positioning, so their outlines sit exactly where
the layout placed them instead of snapping to the pixel grid. A character no typeface has is drawn as the
missing-glyph box, or refused with
MissingGlyphException
when RequireEveryGlyph is on. How glyphs are chosen and shaped is described in Text and
Fonts.
Pixels¶
The size of a page in pixels is its size in points times the resolution over 72, rounded to the nearest pixel, and at least one. The drawing is then scaled across and down separately so the page fills that grid exactly, as a viewer rendering the PDF at the same resolution fills it: an A4 page at 96 pixels per inch is 794 pixels wide. Scaling by the resolution alone would let the rounding drift across the page.
The rest follows the PDF surface as closely as pixels allow:
| In a page image | |
|---|---|
| Background | White for JPEG; transparent for PNG and WebP where no paper is set |
| Colour | Every ink converted to RGB: process colours without a profile, spot inks shown in their fallback (Colour) |
| Dashes | A pattern of odd length repeated twice over, as PDF repeats it |
| Rounded corners | Fitted to the box as the PDF surface fits them |
| Shadows | Blurred by Skia at the shadow's deviation |
| Gradients | Drawn as a linear gradient between the same inks |
| Images | Decoded as stored and turned upright by their EXIF orientation with the same placement the PDF surface uses; each decoded once per export |
| Images made for their box | Asked for at the page image's own resolution |
| Links, anchors, bookmarks, tags | Left out: an image has nothing to follow or read |
Turning images by their orientation in the surface, rather than letting the decoder do it, keeps the two surfaces in step: both start from the stored pixels and apply the same turn. An image whose data is cut short is drawn as far as it goes. The export is at 144 pixels per inch unless told otherwise, and JPEG and WebP at quality 90.
SVG pages¶
ExportSvg draws each page onto Skia's SVG canvas, one unit to the point. Text is drawn as the outlines of its
glyphs, so the page looks the same on a machine without its fonts; the price is that its text cannot be selected or
searched. Images are carried inside the document as data: URIs. Images made for their box are asked for at the
export's ImageResolution, 288 pixels per inch unless set. An exported page can be read back by Artwork.FromSvg,
at the same size.
XPS¶
ExportXps writes one XPS document, the fixed-layout format Windows prints through. Skia writes XPS only through
Windows' own XPS support, so elsewhere the export raises PlatformNotSupportedException. The result is a ZIP
package with a part for each page; its text stays text, with the fonts it is set in carried in the package.
Recompressing images for PDF¶
The core embeds images as they were encoded (Images). Scaling a photograph down to the resolution it
is shown at, or recompressing it at a quality, needs a decoder and an encoder, which the managed core does not have.
So the core asks for an
IImageProcessor
when an image or the export asks for a quality or a maximum resolution, or when PDF/A needs a CMYK image without a
profile made RGB. The Raster package supplies
SkiaImageProcessor.
For each image it is given, it decodes the image, turns it upright by its EXIF orientation, scales it to the pixels
asked for with a cubic (Mitchell) filter, and encodes it: as JPEG at the quality asked for, or as PNG where the image
has transparency or no quality was asked. An image with an alpha channel whose pixels are all opaque is still
compressed as JPEG. An image is never enlarged: a maximum resolution only ever lowers the pixel count. The processor
keeps no state, so the single SkiaImageProcessor.Instance serves every export.
How the rendering is checked¶
Two surfaces that must agree are tested against each other. For each of the fourteen specimen documents of the conformance tests, the PDF is rendered by PDFium, a renderer that shares no code with this library, and the same document is exported as page images through Skia, both at 96 pixels per inch. The test then requires:
- the same number of pages, each the same size in pixels;
- after both are laid over white, at most 0.5% of each page's pixels may differ, a pixel counting as different when any of its red, green or blue values is more than 96 apart.
Any two rasterisers smooth edges differently, so a pixel-exact comparison would fail on anti-aliasing alone. A glyph in the wrong place, an image turned the wrong way or a fill left out changes far more than half a percent of a page, and fails.
The PDF side is pinned too: every page of every specimen, rendered by PDFium, must match an approved snapshot image, with a tighter tolerance — a channel may be 24 apart, and 0.2% of pixels may differ. Other tests check the page images directly: the resolution and format asked for, fills where layout put them, images upright whatever their orientation. Embedded images are also rendered by PDFium and compared with Skia's decoding of the original file. See How it is tested.
The preview¶
Rustaveli.Pdf.Preview shows a document in the browser as its code changes. ADR 0010
chose a browser interface and hot reload through dotnet watch. It speaks of a dotnet tool; what ships is a
package whose
DocumentPreview
the program itself calls, with the same browser interface.
A small web server¶
StartPreview creates a
PreviewSession,
which serves on http://localhost only, at the port
PreviewOptions
gives or at a free one. It listens on a background thread of its own and answers each request on the thread pool:
| Address | Answer |
|---|---|
/ |
The preview page: HTML, CSS and script in one |
/state |
The version, each page's size on screen, and any failure |
/pages/N |
Page N as PNG |
/frames/N |
The frames drawn on page N, as nested JSON |
Nothing is cached by the browser. Preview starts a session, opens the browser unless told not to, and waits for
Ctrl+C.
Drawing when the browser looks¶
A session counts versions. Refresh — or a hot reload — only moves the count on; nothing is drawn yet. The next time
the browser asks, the session sees the count has moved, calls the compose function, and exports the document as page
images at the preview's resolution, recording its frames as it goes. A burst of saves therefore costs one drawing.
sequenceDiagram
participant Code as dotnet watch
participant Session as PreviewSession
participant Browser
Code->>Session: code changed (version + 1)
loop every 600 ms
Browser->>Session: GET /state
end
Session->>Session: compose, lay out, draw pages
Session-->>Browser: new version, page sizes
Browser->>Session: GET /pages/1, /pages/2 ...
The page asks for /state every 600 milliseconds. When the version it gets differs from the one it shows, it loads
the pages again, with the version in each address so no old image is reused. Pages are shown at 96 CSS pixels to the
inch, whatever resolution they were drawn at, so a page appears at its nominal size unless the window is too narrow
for it.
Hot reload¶
Under dotnet watch, the .NET runtime applies code changes to the running program and then tells every type named
by a MetadataUpdateHandler attribute. The preview package names one; when told, it moves on the version of every
open session. Sessions are held by weak references and removed when disposed, so a forgotten session does not live
on. On netstandard2.0, which does not declare the attribute, the package declares it itself: the runtime finds it
by name.
This is why Preview takes a function returning a document rather than a document. A document already composed
holds the frames the old code built; only calling the composing code again, after the update, builds what the new
code says.
Failures on the page¶
Whatever the compose function or the layout throws is caught. The session keeps no pages, and the browser lays the failure over them: the type and message of the exception and of each inner one, and the status reads "Cannot be set". The next change that fixes the code brings the pages back.
For content that cannot be set, the message traces the way down to it. When layout fails with
OversetException,
the typesetter plans the failing part again with tracing switched on, recording each plan within the one that asked
for it. Planning changes nothing, so doing it twice is safe, and the trace costs nothing when layout
succeeds. The trace lists each frame from the page down with the room it was offered, and the last says why it could
not fit; the guide shows an example.
The inspector¶
While the preview draws its pages, every frame rendered is recorded, each within the frame that drew it: its name,
its top left on the page in points, and the room it was given. Frames are named by what they are — Stack, Inset,
Text — or by the name Named gave them, in quotation marks. The browser fetches a page's frames only when the
inspector needs them.
Moving the pointer over a page finds the innermost frame under it, searching the frames drawn last first, and outlines it. Clicking selects it, opens its place in the tree, and shows where it lies and its size.
Each frame also knows the line of code that made it. While the preview calls the compose function — and only then —
each block made records the first place on the call stack outside this library's own assemblies and .NET's that
has a file name. That is the line in the document's code, and the inspector links to it with a vscode://file/
address, which opens it in Visual Studio Code. Reading the stack is slow, so an ordinary export never does it.
A file name and line are known only where the calling code was built with its symbols, as it is by default. Code built without them, such as a helper package, is passed over to the code that called it. Blocks made outside the compose function record no line.
ShowFrameEdges and Named¶
Both are frame modifiers in the core
(FrameModifiers),
so they work in every export, not only in the preview.
ShowFrameEdges draws its content, then the edges of the room the frame was given: a dashed outline half a point
wide, three points on and two off, in the ink given or a strong red. A label, if given, sits in the top corner in
6-point white type on a tab of the same ink. It is real content, drawn into the PDF as into the image, so it is a
tool for finding where frames lie, to be removed before a document is shipped.
Named draws nothing. It gives a frame a name, which the inspector shows and a layout failure uses, so that a trace
reads "Totals" rather than one more Stack among many.