Adapter between
@awacloud/fonts/embed-pdfand PDF dicts — ISO 32000-2 §9.6 / §9.7 / §9.10.
Module pdfFontEmbed | Source packages/front/office/pdf/src/font/embed.js | Deps pdfErrors, embedSubsetForPdf, embedFontDescriptor, embedCidSystemInfo, embedToUnicodeBuilder | Worker-safe yes
The only file in @awacloud/pdf that knows about @awacloud/fonts/embed-pdf. It takes an already-parsed font (the output of fonts.read(bytes)) plus a set of code points, and produces the PDF dicts needed to embed a subset of that font — either simple (/TrueType + /WinAnsiEncoding) or composite (/Type0 + /Identity-H + /CIDFontType2). The subsetting itself is delegated to subsetForPdf; this module exists so the writer's call sites don't re-derive the wiring.
Resolve
const emb = runtime.resolve('pdfFontEmbed');
// Returns: { embedSimple, embedCid }
API
| Method | Signature | Returns |
|---|---|---|
embedSimple |
(parsedFont, codePoints, opts?) => SimpleEmbed |
/TrueType + /WinAnsiEncoding. Every code point must be CP1252-representable. |
embedCid |
(parsedFont, codePoints, opts?) => CidEmbed |
/Type0 + /Identity-H with a /CIDFontType2 descendant. Any Unicode code point. |
parsedFont must be a real parsed Font — unicodeMap, glyphIndexForCodePoint, advanceWidth and a numeric unitsPerEm are all read. codePoints is any iterable of integers; it is de-duplicated and sorted, and the sorted array is echoed back as codePoints. opts is forwarded to subsetForPdf (embedCid adds { cid: true }).
The factory validates at construction that its four injected @awacloud/fonts/embed-pdf helpers (subsetForPdf, buildFontDescriptor, buildCidSystemInfo, embedBuildToUnicode) are functions, throwing ContractError otherwise.
What the adapter consumes
subsetForPdf(font, codePoints, opts?) returns exactly eight keys — subsetBytes, gidMap, glyphMap, encoding, widths, toUnicodeCmap, postScriptName, fontDescriptor. The adapter reads all of them except encoding (always null: under Identity-H the encoding is the cmap). A unit-tier contract-shape guard (src/font/embed.test.js) resolves the REAL embedSubsetForPdf through a ModuleRuntime and asserts that key set, so the stub can never drift from the package again.
Two conversions happen here and nowhere else:
- Widths.
subset.widthsis indexed by NEW gid and expressed in font units. Everything this module emits (/Widths,/W,widthOf) isround(w * 1000 / unitsPerEm). /ToUnicodesource. A/ToUnicodeCMap is keyed by character code. ForembedCidthe code IS the CID (= the subset's new gid), which is exactly howsubset.toUnicodeCmapis keyed — it is used verbatim. ForembedSimplethe codes are WinAnsi bytes, so the gid-keyed CMap would be wrong; the adapter builds the byte-keyed map and passes it toembedBuildToUnicodewith{ codeBytes: 1 }, so the CMap declares a one-byte codespace (<00> <FF>) and 2-hex-digit codes;embedCidkeeps the 2-byte default. Never both for one route.
The consumer's contract — the descriptor carries no font program
buildFontDescriptor puts the raw subset bytes in FontFile2 (TrueType) or FontFile3 (CFF). The adapter lifts them out: the returned descriptor dict has no font-program key, and the bytes come back as fontFile with the key name in fontFileKey. The consumer must:
- allocate the font program as a stream indirect whose dict carries
/Length1 = fontFile.length(the serializer adds/Length), and setdescriptor.entries[fontFileKey]to that ref; - allocate
toUnicodeStreamas a stream indirect too and set the font dict'sToUnicodeto that ref —pdfSerializerrefuses an inline stream (pdf/serializer/inline-stream).
The font dict itself may be written inline or as an indirect; the descriptor nests fine inside it.
Do this wiring in new dicts — spread the result's entries into a fresh
obj.dict({ … }) and add the references there — rather than by assigning into
the result's own dicts. The embed result is then left untouched and can be
reused for another document (or another page tree) as is.
pdfBuilder's addFont({ name, embedded }) does
exactly this wiring for you: it copies the entries into new dicts, never
mutates the result, and allocates its indirects per document.
Shape SimpleEmbed
{
subtype: 'TrueType',
fontDict: typed dict (Type/Subtype/BaseFont/Encoding/FirstChar/LastChar/Widths/FontDescriptor/ToUnicode),
descriptor: typed dict — no FontFile2/FontFile3 key,
toUnicodeStream: typed stream (byte-keyed CMap),
fontFile: Uint8Array — the subset font program,
fontFileKey: 'FontFile2' | 'FontFile3',
encode(text): Uint8Array — one WinAnsi byte per code point,
widthOf(cp): number — 1000/em, 0 when the cp is not embedded,
codePoints: number[] — de-duplicated, ascending
}
BaseFont is the subsetter's postScriptName (already ABCDEF+Family). FirstChar/LastChar are the min/max WinAnsi bytes of the embedded set, and Widths covers [FirstChar..LastChar] with 0 in the unused slots.
Shape CidEmbed
{
type0Dict: typed dict (Type/Subtype:Type0/BaseFont/Encoding:Identity-H/DescendantFonts/ToUnicode),
cidFontDict: typed dict (Subtype:CIDFontType2/CIDSystemInfo/FontDescriptor/DW/W/CIDToGIDMap:Identity),
descriptor: typed dict — no FontFile2/FontFile3 key,
toUnicodeStream: typed stream (CID-keyed CMap),
fontFile: Uint8Array,
fontFileKey: 'FontFile2' | 'FontFile3',
encode(text): Uint8Array — 2 big-endian bytes (the CID) per code point;
an unknown cp becomes CID 0 (.notdef) and bumps `encode.missing`,
widthOf(cp): number — 1000/em, 0 when the cp is not embedded,
codePoints: number[]
}
The subset is renumbered, so CID === new gid and /CIDToGIDMap is /Identity. /DW is 1000; /W is the compact [c [w …] …] form over the subset's gids.
Examples
Simple embed (/TrueType + /WinAnsiEncoding)
const emb = runtime.resolve('pdfFontEmbed');
const { obj } = runtime.resolve('pdfParserObj');
const e = emb.embedSimple(parsedFont, [...'Hello'].map(c => c.codePointAt(0)));
// Clone, don't mutate: new dicts carry the two stream references,
// `e` itself stays reusable for another document.
const descriptor = obj.dict({ ...e.descriptor.entries, [e.fontFileKey]: obj.ref(6, 0) });
const fontDict = obj.dict({ ...e.fontDict.entries, FontDescriptor: descriptor, ToUnicode: obj.ref(7, 0) });
const indirects = [
/* … catalog, pages, page, contents … */
{ num: 5, gen: 0, value: fontDict },
{ num: 6, gen: 0, value: obj.stream(obj.dict({ Length1: obj.int(e.fontFile.length) }), e.fontFile) },
{ num: 7, gen: 0, value: obj.stream(obj.dict({}), e.toUnicodeStream.raw) }
];
// Content stream: the show-string is what `encode` produced.
e.encode('Hello'); // Uint8Array [0x48, 0x65, 0x6C, 0x6C, 0x6F]
e.widthOf(0x48); // advance width of 'H' in 1000/em, never font units
CID embed (composite /Type0)
const e = emb.embedCid(parsedFont, [...'Uni é fi'].map(c => c.codePointAt(0)));
const descriptor = obj.dict({ ...e.descriptor.entries, [e.fontFileKey]: obj.ref(6, 0) });
const cidFont = obj.dict({ ...e.cidFontDict.entries, FontDescriptor: descriptor });
const type0 = obj.dict({ ...e.type0Dict.entries,
DescendantFonts: obj.array([cidFont]),
ToUnicode: obj.ref(7, 0) });
// font program at 6 and e.toUnicodeStream at 7, as in the simple route
const codes = e.encode('Uni é fi'); // 2 bytes per code point, big-endian CIDs
e.encode.missing; // 0 — every cp was in the subset
Errors
| Code | Class | When |
|---|---|---|
pdf/embed/missing-fonts-embed |
ContractError |
The injected @awacloud/fonts/embed-pdf helpers are absent or don't supply the 4 required functions (thrown at factory time). |
pdf/embed/bad-font |
ContractError |
parsedFont is not a parsed Font (missing unicodeMap / glyphIndexForCodePoint / advanceWidth / numeric unitsPerEm). |
pdf/embed/bad-codepoints |
ContractError |
codePoints is not iterable, is empty, or holds a non-integer / out-of-range value (context.cp). |
pdf/embed/not-winansi |
ContractError |
embedSimple was given a code point outside CP1252 (context.cp) — use embedCid. Also thrown by SimpleEmbed.encode for a character outside the embedded set. |
See also
pdfFont— read-side typing.pdfFontEncoding—/Encodingresolution.pdfWriter— consumes the payload.parser-obj—obj.*constructors.tests/font-embed-real.integration.test.js— both routes written and read back against a REAL parsed font.