Orchestrator parsing/rendering the body of
<office:text>(paragraphs, headings, lists, sections, tables, soft page breaks).
Module textContent | Source packages/front/office/odf/src/text/content.js | Deps xml, textParagraph, textHeading, textList, textSection, tableTable | Worker-safe yes
Each body node has a type:
| type | element |
|---|---|
paragraph |
<text:p> |
heading |
<text:h> |
list |
<text:list> |
section |
<text:section> |
table |
<table:table> |
soft-page-break |
<text:soft-page-break/> |
unknown |
(preserved as-is) |
Table cells are typed here — unlike standalone tableCell
tableCell carries a cell's children as raw XML
element nodes verbatim — it has no notion of paragraphs, lists, or nested
tables. Within textContent, once a <table:table> is
reached via parseNode/parseBody, every non-covered cell's children
array is replaced with the same TYPED body nodes used everywhere else in
this module (paragraphs, headings, lists, nested tables, or unknown for
anything unrecognised) — produced by recursing each raw cell child through
parseNode. Covered cells (covered: true) always end up with
children: [], regardless of what was in the source XML.
The reverse happens on render: renderNode/renderBody deep-copy the table
(and its rows and cells) and replace each cell's typed children with
rendered raw XML nodes (via renderNode) before delegating to
tableTable.renderTable. The typed model passed in is
never mutated — a node handed to renderBody compares deep-equal to
itself afterwards (odt.write must not alter the caller's doc).
Everything else about a table node — name, styleName, columns,
headerRows, and per-cell styleName/valueType/value/currency/
formula/repeated/colSpan/rowSpan/covered/_extras — is exactly the
tableTable/tableRow/
tableCell model shape; only children changes meaning
within textContent — plus the semantic grid field below.
Grid tables — the grid field
A body table node may carry grid?: boolean, a textContent-level field
(it is not part of tableTable's model and is never
serialised as an attribute). It means "draw every cell with a visible
border, align the table to the page margins", expressed through the
textStyleRegistry seam:
- Render — when
node.grid === trueandctxexposescellStyle, the rendered table getstable:style-name = node.styleName || ctx.tableStyle({align: 'margins'})(awa-tb-m) and every non-covered cell getstable:style-name = cell.styleName || ctx.cellStyle({bordered: true})(awa-c-b). An explicitstyleNamealways wins — the same precedence rule as paragraphs and lists. Covered cells are untouched. Withoutctx, or withoutgrid, the output is byte-identical to a seam-free render. - Parse — when
ctxexposescellBorders, the table has at least one non-covered cell and every non-covered cell'sstyleNameresolvesbordered, the table getsgrid: true. Each cell style resolved withsource: 'auto'is dropped from its cell (a named style is kept — "keep and gain"), and the table's ownstyleNameis dropped too whentableAlignresolves it withsource: 'auto'. - Consume-and-drop — the auto resolutions land in the resolver's
consumedset, soodt.readdrops the matching automatic styles from the surfacedautoStyles. If any non-covered cell fails, the table is left exactly as without the seam (nogrid, every name kept) and any automatic style the attempt consumed is released again, so it stays surfaced.
Invariant: for T = { type: 'table', grid: true, columns: [{repeated: 3}], rows: [...] } with no cell styleName, parseBody(renderBody([T], registry), resolver) deep-equals [T], and odt.read(odt.write({ body: [T] })).autoStyles
is undefined.
Resolve
const content = runtime.resolve('textContent');
// → { parseNode, renderNode, parseBody, renderBody, bodyText }
API
| Method | Description |
|---|---|
parseBody(officeTextEl, ctx?) |
Converts the contents of an <office:text> into an array of typed nodes. |
renderBody(nodes, ctx?) |
Converts an array of typed nodes into XML elements. |
parseNode(el, ctx?) / renderNode(node, ctx?) |
Works node by node; used internally to delegate section/list/table children. |
bodyText(nodes) |
One line per paragraph or heading (lists, sections and table cells included), followed by the text of any text box (draw:frame / draw:text-box) anchored in it; image frames and alternative text contribute nothing; spacing inside a text:span (text:s, text:tab, text:line-break) is honoured, and the text of a text box anchored inside a span stays inline with the span. |
The ctx parameter
Every parse/render entry point accepts an optional trailing ctx — the style
seam object built by odt: a
textStyleRegistry resolver on read, a registry
on write. textContent never interprets it; it passes it straight through as
an extra trailing argument:
| dispatch | call |
|---|---|
text:p |
parseParagraph(el, ctx) / renderParagraph(node, ctx) |
text:h |
parseHeading(el, ctx) / renderHeading(node, ctx) |
text:list |
parseList(el, hooks, ctx) / renderList(node, hooks, ctx) |
text:section |
parseSection(el, hooks) / renderSection(node, hooks) — no ctx |
table:table |
cell children go through parseNode(el, ctx) / renderNode(node, ctx) directly; the table itself consumes ctx for the grid field |
The child hooks are built per call, so nested content keeps the ctx of the
enclosing call. textSection receives none: its children reach it through
those hooks, which already carry it. Table cells are not routed through
tableTable/tableRow/tableCell for their content — textContent types
and untypes cell children itself, calling parseNode/renderNode
recursively (so a table nested inside a table, or inside a list item, keeps
the same ctx all the way down).
All three delegate pairs consume ctx: textParagraph/textHeading for
typed run emphasis and link runs, textList for ordered/bullet list styles —
and body tables consume it for grid (above).
ctx === undefined reproduces the previous behaviour exactly, byte for byte.
Examples
const content = runtime.resolve('textContent');
const officeText = xml.parse('<office:text><text:p>Hello</text:p></office:text>');
const nodes = content.parseBody(officeText);
content.bodyText(nodes); // 'Hello'
content.renderBody(nodes);
// With the style seam threaded through (as `odt` does):
const resolver = styleRegistry.createResolver(autoStyles, fontFaces, namedStyles);
content.parseBody(officeText, resolver);
A table cell's content, once parsed, is a typed body-node array like any other:
const withTable = xml.parse(
'<office:text><table:table><table:table-row><table:table-cell>' +
'<text:p>cell text</text:p></table:table-cell></table:table-row></table:table></office:text>'
);
const [table] = content.parseBody(withTable);
table.type; // 'table'
table.rows[0].cells[0].children[0]; // { type: 'paragraph', runs: [...] }
content.bodyText([table]); // 'cell text'
A grid table through the registry/resolver seam:
const styleRegistry = runtime.resolve('textStyleRegistry');
const para = runtime.resolve('textParagraph');
const grid = { type: 'table', grid: true, columns: [{ repeated: 2 }], rows: [
{ type: 'row', cells: [
{ type: 'cell', children: [para.paragraph('a')] },
{ type: 'cell', children: [para.paragraph('b')] }
] }
] };
const registry = styleRegistry.createRegistry();
const rendered = content.renderBody([grid], registry);
rendered[0].attrs; // { 'table:style-name': 'awa-tb-m' }
registry.toAutomaticStyles().styles.map(s => s.name); // ['awa-c-b', 'awa-tb-m']
const resolver = styleRegistry.createResolver(registry.toAutomaticStyles(), [], null);
const [back] = content.parseBody(xml.el('office:text', {}, rendered), resolver);
back.grid; // true — cell/table style names dropped
Notes
- Orchestrates
textParagraph,textHeading,textList,textSection,tableTable— unrecognised element names become{ type: 'unknown', element: el }and round-trip verbatim, including an{ type: 'unknown', element: <table:table ...> }node produced by an older reader that predates thetablenode type: it still renders back to its original XML verbatim. bodyTextrecurses into list items, section children, and every row/cell of a table. It does not readunknownnodes as text, except to find text boxes: the text of everydraw:text-boxkept raw by the reader (in a paragraph's or heading's raw extras, in a link run's, or in a body-levelunknownelement) is emitted on its own line after the line of the element that holds it, and boxes nested in a box are reached once, in document order. The read model itself is unchanged — frames stay raw.bodyTexttakes noctx: it reads only what is already in the typed nodes.- Covered cells (
table:covered-table-cell,covered: true) always carrychildren: []after parsing — any raw children present in the source XML are dropped, matching the ODF semantics that a covered cell has no content of its own.
See also
- text/paragraph
- text/heading
- text/list
- text/section
- text/style-registry — what a
ctxis - table/table / table/row / table/cell — the cell model whose raw
childrentextContentreplaces with typed nodes - odt/odt — builds the
ctxand threads it in