Skip to content
SilktideHelp

PDF text mapping

Silktide checks that the text in a PDF can be recovered as text. A PDF page does not store letters, it stores instructions to draw shapes, and a separate mapping inside the file records which character each shape stands for. Where that mapping is missing, the words are visible but unreadable to software.

Those shapes are , and the mapping connects each one to its character. Without it, text cannot be read aloud, searched, or copied, however crisp it looks on screen.

Why this matters

This failure is unusually deceptive, because the page is perfect. The text is sharp, it is selectable, it prints beautifully, and nothing about opening the document suggests a problem. Then someone copies a paragraph and pastes gibberish, or a reads a page of nonsense syllables, or announces nothing at all.

Because it is invisible to the person publishing the document, it tends to be found by the person relying on it. A reader who hears a stream of noise cannot tell whether the document is broken or whether their software is, and either way the content is gone.

The consequences reach past accessibility. Text that cannot be extracted cannot be indexed by a search engine, cannot be found by your own site search, cannot be translated, and cannot be quoted by anyone. It is why an otherwise well-made document sometimes returns no results for a phrase that is plainly printed inside it.

How to fix it

  1. Re-export from the source with fonts embedded properly, which resolves nearly all of these. The mapping is written by the export process, so exporting again with a current tool is usually enough.
  2. Avoid printing to PDF. Print drivers are the most common cause, because printing describes a page as marks to put on paper and discards the text behind them. Use the application's own export or "Save as PDF" instead.
  3. Watch for icon and symbol fonts. Decorative and dingbat fonts often have no sensible character behind each shape. Where a symbol carries meaning, use a real character or a described image; where it is decoration, mark it as an .
  4. Have the mapping rebuilt. Silktide remediation reconstructs missing mappings where it can, and gives correct spoken text to stretches it cannot map. See How PDF remediation works.
  5. Recognise the text if the page is really an image. A scan has no text to map, so the fix is rather than repair. See Scanned documents.

How Silktide tests this

  1. Find the PDF documents linked from your website, and analyse each unique file once however many pages link to it.
  2. Read the text-drawing instructions on each page and follow every glyph back to the character it claims to represent.
  3. Report the document when a glyph has no valid mapping to a Unicode character.
  4. Report each document once, however much text is affected.

Text that draws a font's missing-character placeholder is reported by PDF fonts instead, because the cause there is the font rather than the mapping.

Troubleshooting

I can select and copy the text, so what is wrong?

Selecting text and recovering it are different operations. The viewer lets you select the region you clicked on; what lands on the clipboard is whatever the mapping says those glyphs mean. Paste it into a plain text editor and compare it with the page - if it comes back garbled, you are seeing exactly what a screen reader receives.

Only a few characters are affected

That is common and still worth fixing, because the affected characters are rarely random. It is usually one font used for one purpose, so the damage lands on all the bullets, all the mathematical symbols, or every fraction in the document.

The document is a form

Forms often mix fonts from several sources, and field labels drawn in a symbol font are a frequent cause. Check the labels and any custom checkbox or radio characters. See Interactive forms.

Learn more

Last updated

Was this page helpful?

PDF text mapping | Silktide Help