Plain-text extraction helpers for a typed
docxdocument tree — the module behinddocx.toText.
Module docxText | Source packages/front/office/ooxml/src/docx/docx-text.js | Deps none | Worker-safe yes
Pure functions over the typed model returned by docx.read (result.document): no runtime dependency, no state, no side effect. The docx orchestrator depends on this module and re-exposes its toText as docx.toText — the two are the same function, so most callers never resolve docxText directly.
Resolve
const text = runtime.resolve('docxText');
// Returns: { toText, textOfParagraph, textOfRun, textOfTable }
API
| Method | Signature | Returns |
|---|---|---|
toText |
(doc: { body? }) => string |
The text of the whole document: one line per top-level paragraph or table, joined by \n. |
textOfParagraph |
(p: paragraph) => string |
The concatenated text of the paragraph's runs, including the runs of hyperlink and ins children. |
textOfRun |
(r: run) => string |
The text of one run. |
textOfTable |
(t: table) => string |
Rows joined by \n, cells by \t. |
Rules
| Node | Contributes |
|---|---|
text |
its value |
tab |
\t |
break |
\n |
noBreakHyphen |
U+2011 (non-breaking hyphen) |
hyperlink, ins |
the text of their runs |
del |
nothing — deleted text is not visible text |
| table cell | its paragraphs (and nested tables) joined by a single space |
Examples
const text = runtime.resolve('docxText');
const doc = {
body: [
{ type: 'paragraph', children: [
{ type: 'run', children: [{ type: 'text', value: 'Hello' }, { type: 'tab' }, { type: 'text', value: 'world' }] },
{ type: 'del', children: [{ type: 'run', children: [{ type: 'text', value: 'gone' }] }] }
] },
{ type: 'table', rows: [{ cells: [
{ children: [{ type: 'paragraph', children: [{ type: 'run', children: [{ type: 'text', value: 'A1' }] }] }] },
{ children: [{ type: 'paragraph', children: [{ type: 'run', children: [{ type: 'text', value: 'B1' }] }] }] }
] }] }
]
};
console.log(JSON.stringify(text.toText(doc))); // "Hello\tworld\nA1\tB1"
On a read document
const word = runtime.resolve('docx');
const read = word.read(word.write(word.fromText(['First', 'Second'])));
console.log(word.toText(read.document)); // "First\nSecond"
console.log(word.toText === runtime.resolve('docxText').toText); // true
Notes
- Only top-level
paragraphandtablenodes ofdoc.bodyare visited. Other block-level nodes, such as ablockSdt(a structured document tag wrapping paragraphs), contribute nothing, and neither does an inlinesdtinside a paragraph (see templating for the SDT model). - Headers, footers, footnotes, endnotes and comments are separate parts of the read result, not part of
doc.body:toTextdoes not reach them. - The output is a flat reading of the text, not a layout: formatting, fields, drawings and list numbers are dropped.
See also
- docx — orchestrator (
toText). - docx-walker — the other module extracted from the orchestrator.
- docx-structure — the node types (
paragraph,run,table, …).