Per-font character decoder for the tier-1 pdf reader — code → Unicode text, and code → glyph width.
Module oconvPdfFontDecoder (oconvPdfFontDecoder) | Source packages/front/office/oconv/src/read/pdf/font-decoder.js | Deps pdfFont, pdfFontEncoding, pdfFilterDispatch, cmapToUnicode, encodingLookup, encodingAgl, standard14Lookup | Worker-safe yes
Turns a PDF Font dict into a decoder honouring the tier-1 decode paths in priority order: a /ToUnicode CMap (authoritative), then an /Encoding table + the AGL glyph-name hop for simple fonts, then — for composite (Type0) fonts only — the predefined CMap the /Encoding names, when that CMap's input codes are Unicode. The same decoder also gives each code's glyph width, which oconvPdfTextExtract uses to place the end of every text piece. Composes only documented public exports of @awacloud/pdf and @awacloud/fonts — no internal reach.
Resolve
import { ModuleRuntime } from '@awacloud/fw/core/runtime.js';
import { fw_require, modules } from '@awacloud/oconv';
const runtime = new ModuleRuntime();
runtime.registerAll(fw_require);
runtime.registerAll(modules);
const { buildDecoder, decodeShow } = runtime.resolve('oconvPdfFontDecoder');
API
| Member | Signature | Returns | Throws |
|---|---|---|---|
buildDecoder |
(fontDict: object, resolve: (ref) => object) => {subtype, cidBytes, decode, width, vertical, widthApproximated} |
a decoder bound to one resolved Font dict; decode(code: number) => string|null; width(code: number) => number (glyph width w0, thousandths of text space); vertical (true for a …-V composite CMap); widthApproximated() => boolean |
— |
decodeShow |
(decoder: {cidBytes, decode}, strBytes: Uint8Array) => {text, decoded, undecodable} |
splits strBytes into 1- or 2-byte codes per decoder.cidBytes and decodes each |
— |
Examples
Decode a show-operator string through a ToUnicode CMap
const pdfParserObj = runtime.resolve('pdfParser').obj;
const cmapToUnicode = runtime.resolve('cmapToUnicode');
const enc = (s) => new TextEncoder().encode(s);
const toUniSrc = cmapToUnicode.buildToUnicode(new Map([[0x41, 'A'], [0x42, 'B']]));
const fontDict = pdfParserObj.dict({
Type: pdfParserObj.name('Font'), Subtype: pdfParserObj.name('Type1'),
BaseFont: pdfParserObj.name('Helvetica'),
ToUnicode: pdfParserObj.stream(pdfParserObj.dict({}), enc(toUniSrc))
});
const decoder = buildDecoder(fontDict, (ref) => ref);
decodeShow(decoder, new Uint8Array([0x41, 0x42]));
// { text: 'AB', decoded: 2, undecodable: 0 }
Notes
- Decode priority:
/ToUnicodefirst; when absent (or it has no entry for a code), the/Encodingtable +encodingAgl.glyphNameToUnicodehop is tried for simple (non-Type0) fonts only. - Standard CMap: a
Type0font whose/Encodingnames aUni{GB,CNS,JIS,KS}-{UCS2,UTF16}[-HW]-{H,V}predefined CMap (ISO 32000-2 §9.7.5.2) decodes a code without a ToUnicode entry as the UTF-16 code unit it is; a surrogate half is counted undecodable.Identity-H/Identity-V(codes are CIDs, not Unicode) and the legacy RKSJ/EUC/… CMaps (tables not bundled) decode nothing here — a composite font in that position with no ToUnicode stays an honesttext/undecodableloss (CID→GID→Unicode font-program walking is out of tier-1 scope). - A code neither path resolves is counted undecodable and omitted from the text — never silently kept as a wrong glyph, never dropped without a count (
decodeShow's own return shape). - Glyph widths (
width), by font kind:- simple fonts:
/Widthsindexed from/FirstChar. A code outside the array takes the descriptor's/MissingWidth(0 when absent, the spec default). A Type 3 font's widths are scaled by its/FontMatrix. - composite (
Type0) fonts: the descendant CIDFont's/Warray, in both forms (c [w1 w2 …]andcFirst cLast w), with/DW(default 1000) for CIDs it does not list. Codes are read as CIDs only underIdentity-H/Identity-V. Under any other CMap the code → CID map is unknown here, so/DWis used. - Standard 14 fonts with no
/Widths: the Adobe Core 14 AFM widths shipped by@awacloud/fonts(standard14Lookup), looked up by the glyph name the font's encoding gives the code.SymbolandZapfDingbatsare looked up by code.
- simple fonts:
- When a width cannot be read from the font, a declared fallback is used: 500 thousandths for a simple font, 1000 for a composite one.
widthApproximated()then turnstrue, andoconvPdfTextExtractrecordstext/width-approximated. The cases are:/Widthsabsent on a font that is not one of the Standard 14, a malformed width array or entry, a missing descendant font, a/Wtable under a non-identity CMap, and a named Standard 14 glyph missing from the AFM table. A malformed width structure never throws. cidBytesis2forType0fonts,1otherwise —decodeShowuses it to splitstrBytesinto codes.
See also
- docs/api/README.md — full module index +
exportsboundary note oconvPdfTextExtract— the sole consumer of this decoderoconvPdfToIr— the reader composing the pdf pipeline- loss matrix