Skip to content
NoWatermark

Invisible Unicode character detector and inspection guide

Detect hidden zero-width spaces, bidi overrides, and invisible Unicode characters in your text locally in your browser with verified re-scanning.

By NoWatermark · Published

An invisible Unicode character detector identifies non-printing code points, zero-width spaces, and hidden formatting controls that sit undetected within digital text. Finding these characters requires inspecting the raw string byte by byte rather than relying on visual inspection in a standard text editor.

You can inspect text directly in your browser using the hidden character checker. All scanning runs entirely client-side, meaning your pasted text is never transmitted to an external server or stored in a remote database.

Why invisible characters appear in normal text

Invisible characters are not inherently malicious, nor do they indicate that a document has been tampered with. Unicode contains thousands of defined code points designed for layout formatting, script directionality, and typographic control.

A detector is an inspection tool, not an accusation. Invisible code points have numerous legitimate uses across standard digital publishing:

  • Typographic layout: Soft hyphens specify where a word processing engine may insert a line break, while word joiners prevent line breaks across specific character boundaries.
  • Bidirectional scripts: Languages written from right to left, such as Arabic and Hebrew, rely on directional marks and isolates to interleave correctly with left-to-right numbers and Latin phrases.
  • Copy-paste artefacts: Copying text from web pages, PDF documents, or desktop word processors often pulls along layout markers, non-breaking spaces, and formatting isolates that remain invisible on screen.

Invisible characters do not prove that text originated from any specific AI system. While discussion around hidden text markers often focuses on automated generation, everyday publishing workflows routinely introduce zero-width characters into human-written documents.

Code points detected by NoWatermark

NoWatermark scans text strings against an explicit registry of zero-width markers, bidirectional controls, layout formatters, and unusual whitespace characters. Each detected character is assigned a risk classification based on its potential to alter rendering or conceal arbitrary payload data.

Individual code points

The scanner checks for specific single code points across zero-width, bidirectional, and non-standard whitespace categories:

Code point Name Category Risk
U+00AD Soft hyphen zero-width medium
U+061C Arabic letter mark bidi control medium
U+180E Mongolian vowel separator zero-width medium
U+200B Zero-width space zero-width high
U+200C Zero-width non-joiner zero-width high
U+200D Zero-width joiner zero-width high
U+200E Left-to-right mark bidi control medium
U+200F Right-to-left mark bidi control medium
U+2028 Line separator control low
U+2029 Paragraph separator control low
U+2060 Word joiner zero-width high
U+2061 Function application zero-width medium
U+2062 Invisible times zero-width medium
U+2063 Invisible separator zero-width medium
U+2064 Invisible plus zero-width medium
U+2800 Braille pattern blank unusual space medium
U+3164 Hangul filler unusual space high
U+FEFF Zero-width no-break space (BOM) zero-width high
U+FFA0 Halfwidth Hangul filler unusual space high

Detected character ranges

In addition to individual code points, the detector monitors structured ranges containing legacy control codes, formatting overrides, and variation selectors:

Range Name Category Risk
U+0000–U+0008 Control character control medium
U+000B–U+000C Control character control low
U+000E–U+001F Control character control medium
U+007F–U+009F Control character control medium
U+202A–U+202E Bidirectional override bidi control high
U+2066–U+2069 Bidirectional isolate bidi control high
U+FE00–U+FE0F Variation selector variation selector medium
U+E0000–U+E007F Tag character tag character high
U+E0100–U+E01EF Variation selector supplement variation selector medium

Non-standard space characters

Standard body text uses the basic ASCII space (U+0020). However, formatting engines and typography software frequently insert alternative whitespace code points that alter spacing widths:

  • U+00A0: Non-breaking space
  • U+2000–U+200A: Fixed-width typographical spaces (en quad, em quad, en space, em space, three-per-em, four-per-em, six-per-em, figure space, punctuation space, thin space, and hair space)
  • U+202F: Narrow no-break space
  • U+205F: Medium mathematical space
  • U+3000: Ideographic space

While these characters render visually, they differ from standard spaces and can cause unexpected behaviour in parsers, code interpreters, and strict text-matching pipelines.

The false-positive challenge with emoji

The central technical difficulty in detecting and removing invisible characters is avoiding false positives. A naive script that strips every zero-width code point will corrupt valid text containing modern emoji.

Two specific code points are structurally load-bearing inside emoji sequences:

  1. U+200D (Zero-width joiner): Used to combine multiple separate emoji glyphs into a single composite sequence. For example, a family emoji is composed of individual person glyphs fused together with U+200D characters. If a cleaner indiscriminately removes U+200D, the single family emoji disintegrates into three or four distinct individual people.
  2. U+FE0F (Variation selector-16): Dictates whether a character should render as a full-colour graphical emoji rather than a black-and-white typographic symbol. Stripping U+FE0F causes supported symbols to fall back to monochrome glyphs.

Because these code points serve essential rendering functions inside emoji sequences, NoWatermark inspects the surrounding context. It classifies contextual emoji joiners and variation selectors as legitimate and leaves them intact by default.

Cleaning text correctly requires identifying which invisible characters serve a functional rendering purpose and which are isolated or unattached. For more background on these specific code points, read our guide on hidden Unicode characters.

High-risk characters and Unicode tags

Certain invisible characters carry a higher risk profile because they can alter text direction or conceal encoded binary payloads.

Tag characters (U+E0000–U+E007F)

Tag characters represent the highest-concern category detected by the tool. Originally defined in Unicode for language tagging, these code points mirror standard ASCII characters but remain completely invisible when rendered by modern text engines.

Because the tag character block corresponds directly to the standard ASCII table, a sequence of tag characters can invisibly mirror an entire plain-text string or binary payload inside otherwise ordinary prose. The text appears completely normal on screen, while the underlying byte sequence carries hidden data.

Bidirectional overrides (U+202A–U+202E)

Bidirectional overrides force the rendering engine to display subsequent characters in a specific direction (left-to-right or right-to-left), overriding the inherent directionality of the characters themselves. When misplaced or deliberately injected, these code points cause visual text to render in an order that contradicts the physical sequence of bytes in the file.

Fillers and zero-width spaces

Code points such as the Hangul filler (U+3164), halfwidth Hangul filler (U+FFA0), and zero-width space (U+200B) occupy space in the byte stream without producing a visible glyph. In structural parsing contexts, these characters can disrupt search indexing, string matching, and automated form processing.

Character data versus statistical watermarks

When analysing digital text, it is critical to distinguish between three entirely different concepts:

  1. Character data and metadata: Raw code points, EXIF blocks, and XMP packets exist physically in the file or string. They can be detected directly, removed programmatically, and their removal can be confirmed by re-scanning the output.
  2. Statistical text watermarks: Some generative systems alter token selection probabilities during text generation (such as SynthID). No local tool can detect these or confirm their absence; NoWatermark reports statistical text watermarks as unable to verify in all cases.
  3. Server-side provenance: A platform or service may maintain internal database records associating generated text with a user account or query timestamp. Local editing cannot affect server-side records.

Removing hidden Unicode characters from a string strips specific, non-printing byte sequences. It does not alter statistical token distributions, does not certify human authorship, and does not guarantee that text will pass statistical AI detectors. NoWatermark has not measured any effect on detector output from removing them — which is a statement about the limits of what we have tested, not a claim about every detector. For a detailed breakdown of how text generation systems interact with watermarking concepts, see our analysis on whether Claude watermarks text.

How to scan and clean text

To inspect and sanitise text containing hidden code points, follow these steps:

  1. Scan the input: Paste your text into the hidden character checker. The detector inspects every code point against the registry, flagging isolated zero-width characters, bidi overrides, and tag characters while preserving valid emoji sequences.
  2. Remove unwanted characters: Use the hidden character remover to strip non-functional invisible characters from the string.
  3. Re-scan the output: Always re-scan the cleaned text through the checker to confirm that the identified code points have been removed from the character stream.

All processing executes locally in your browser. No text is sent across the network, ensuring that sensitive documents and private notes remain entirely on your own machine.

Frequently asked questions

What is an invisible Unicode character detector?

An invisible Unicode character detector identifies unrendered, zero-width, or control code points within digital text. It highlights hidden characters such as zero-width spaces or bidirectional overrides without sending your text to an external server.

Do invisible Unicode characters prove text was generated by AI?

No. Invisible characters do not prove text came from any particular AI system. They frequently appear when copy-pasting from web pages, word processors, and PDF documents.

Will stripping invisible characters break emoji?

Yes, if done indiscriminately. Code points like zero-width joiners (U+200D) and variation selectors (U+FE0F) are essential for rendering multi-person emoji and colour glyphs correctly.

Can removing hidden Unicode bypass AI text detectors?

No tool guarantees bypassing AI detection. Hidden Unicode characters are ordinary character data rather than statistical watermarks, and removing them does not affect detector output in any way NoWatermark has measured. That is a statement about what we have tested, not a promise about any particular detector.

Related tools

Related guides