Invisible Unicode character detector and inspection guide
Detect hidden zero-width spaces, bidi overrides, and invisible Unicode characters in your text locally in your browser with verified re-scanning.
By NoWatermark · Published
An invisible Unicode character detector identifies non-printing code points, zero-width spaces, and hidden formatting controls that sit undetected within digital text. Finding these characters requires inspecting the raw string byte by byte rather than relying on visual inspection in a standard text editor.
You can inspect text directly in your browser using the hidden character checker. All scanning runs entirely client-side, meaning your pasted text is never transmitted to an external server or stored in a remote database.
Why invisible characters appear in normal text
Invisible characters are not inherently malicious, nor do they indicate that a document has been tampered with. Unicode contains thousands of defined code points designed for layout formatting, script directionality, and typographic control.
A detector is an inspection tool, not an accusation. Invisible code points have numerous legitimate uses across standard digital publishing:
- Typographic layout: Soft hyphens specify where a word processing engine may insert a line break, while word joiners prevent line breaks across specific character boundaries.
- Bidirectional scripts: Languages written from right to left, such as Arabic and Hebrew, rely on directional marks and isolates to interleave correctly with left-to-right numbers and Latin phrases.
- Copy-paste artefacts: Copying text from web pages, PDF documents, or desktop word processors often pulls along layout markers, non-breaking spaces, and formatting isolates that remain invisible on screen.
Invisible characters do not prove that text originated from any specific AI system. While discussion around hidden text markers often focuses on automated generation, everyday publishing workflows routinely introduce zero-width characters into human-written documents.
Code points detected by NoWatermark
NoWatermark scans text strings against an explicit registry of zero-width markers, bidirectional controls, layout formatters, and unusual whitespace characters. Each detected character is assigned a risk classification based on its potential to alter rendering or conceal arbitrary payload data.
Individual code points
The scanner checks for specific single code points across zero-width, bidirectional, and non-standard whitespace categories:
| Code point | Name | Category | Risk |
|---|---|---|---|
| U+00AD | Soft hyphen | zero-width | medium |
| U+061C | Arabic letter mark | bidi control | medium |
| U+180E | Mongolian vowel separator | zero-width | medium |
| U+200B | Zero-width space | zero-width | high |
| U+200C | Zero-width non-joiner | zero-width | high |
| U+200D | Zero-width joiner | zero-width | high |
| U+200E | Left-to-right mark | bidi control | medium |
| U+200F | Right-to-left mark | bidi control | medium |
| U+2028 | Line separator | control | low |
| U+2029 | Paragraph separator | control | low |
| U+2060 | Word joiner | zero-width | high |
| U+2061 | Function application | zero-width | medium |
| U+2062 | Invisible times | zero-width | medium |
| U+2063 | Invisible separator | zero-width | medium |
| U+2064 | Invisible plus | zero-width | medium |
| U+2800 | Braille pattern blank | unusual space | medium |
| U+3164 | Hangul filler | unusual space | high |
| U+FEFF | Zero-width no-break space (BOM) | zero-width | high |
| U+FFA0 | Halfwidth Hangul filler | unusual space | high |
Detected character ranges
In addition to individual code points, the detector monitors structured ranges containing legacy control codes, formatting overrides, and variation selectors:
| Range | Name | Category | Risk |
|---|---|---|---|
| U+0000–U+0008 | Control character | control | medium |
| U+000B–U+000C | Control character | control | low |
| U+000E–U+001F | Control character | control | medium |
| U+007F–U+009F | Control character | control | medium |
| U+202A–U+202E | Bidirectional override | bidi control | high |
| U+2066–U+2069 | Bidirectional isolate | bidi control | high |
| U+FE00–U+FE0F | Variation selector | variation selector | medium |
| U+E0000–U+E007F | Tag character | tag character | high |
| U+E0100–U+E01EF | Variation selector supplement | variation selector | medium |
Non-standard space characters
Standard body text uses the basic ASCII space (U+0020). However, formatting engines and typography software frequently insert alternative whitespace code points that alter spacing widths:
- U+00A0: Non-breaking space
- U+2000–U+200A: Fixed-width typographical spaces (en quad, em quad, en space, em space, three-per-em, four-per-em, six-per-em, figure space, punctuation space, thin space, and hair space)
- U+202F: Narrow no-break space
- U+205F: Medium mathematical space
- U+3000: Ideographic space
While these characters render visually, they differ from standard spaces and can cause unexpected behaviour in parsers, code interpreters, and strict text-matching pipelines.
The false-positive challenge with emoji
The central technical difficulty in detecting and removing invisible characters is avoiding false positives. A naive script that strips every zero-width code point will corrupt valid text containing modern emoji.
Two specific code points are structurally load-bearing inside emoji sequences:
- U+200D (Zero-width joiner): Used to combine multiple separate emoji glyphs into a single composite sequence. For example, a family emoji is composed of individual person glyphs fused together with U+200D characters. If a cleaner indiscriminately removes U+200D, the single family emoji disintegrates into three or four distinct individual people.
- U+FE0F (Variation selector-16): Dictates whether a character should render as a full-colour graphical emoji rather than a black-and-white typographic symbol. Stripping U+FE0F causes supported symbols to fall back to monochrome glyphs.
Because these code points serve essential rendering functions inside emoji sequences, NoWatermark inspects the surrounding context. It classifies contextual emoji joiners and variation selectors as legitimate and leaves them intact by default.
Cleaning text correctly requires identifying which invisible characters serve a functional rendering purpose and which are isolated or unattached. For more background on these specific code points, read our guide on hidden Unicode characters.
High-risk characters and Unicode tags
Certain invisible characters carry a higher risk profile because they can alter text direction or conceal encoded binary payloads.
Tag characters (U+E0000–U+E007F)
Tag characters represent the highest-concern category detected by the tool. Originally defined in Unicode for language tagging, these code points mirror standard ASCII characters but remain completely invisible when rendered by modern text engines.
Because the tag character block corresponds directly to the standard ASCII table, a sequence of tag characters can invisibly mirror an entire plain-text string or binary payload inside otherwise ordinary prose. The text appears completely normal on screen, while the underlying byte sequence carries hidden data.
Bidirectional overrides (U+202A–U+202E)
Bidirectional overrides force the rendering engine to display subsequent characters in a specific direction (left-to-right or right-to-left), overriding the inherent directionality of the characters themselves. When misplaced or deliberately injected, these code points cause visual text to render in an order that contradicts the physical sequence of bytes in the file.
Fillers and zero-width spaces
Code points such as the Hangul filler (U+3164), halfwidth Hangul filler (U+FFA0), and zero-width space (U+200B) occupy space in the byte stream without producing a visible glyph. In structural parsing contexts, these characters can disrupt search indexing, string matching, and automated form processing.
Character data versus statistical watermarks
When analysing digital text, it is critical to distinguish between three entirely different concepts:
- Character data and metadata: Raw code points, EXIF blocks, and XMP packets exist physically in the file or string. They can be detected directly, removed programmatically, and their removal can be confirmed by re-scanning the output.
- Statistical text watermarks: Some generative systems alter token selection probabilities during text generation (such as SynthID). No local tool can detect these or confirm their absence; NoWatermark reports statistical text watermarks as unable to verify in all cases.
- Server-side provenance: A platform or service may maintain internal database records associating generated text with a user account or query timestamp. Local editing cannot affect server-side records.
Removing hidden Unicode characters from a string strips specific, non-printing byte sequences. It does not alter statistical token distributions, does not certify human authorship, and does not guarantee that text will pass statistical AI detectors. NoWatermark has not measured any effect on detector output from removing them — which is a statement about the limits of what we have tested, not a claim about every detector. For a detailed breakdown of how text generation systems interact with watermarking concepts, see our analysis on whether Claude watermarks text.
How to scan and clean text
To inspect and sanitise text containing hidden code points, follow these steps:
- Scan the input: Paste your text into the hidden character checker. The detector inspects every code point against the registry, flagging isolated zero-width characters, bidi overrides, and tag characters while preserving valid emoji sequences.
- Remove unwanted characters: Use the hidden character remover to strip non-functional invisible characters from the string.
- Re-scan the output: Always re-scan the cleaned text through the checker to confirm that the identified code points have been removed from the character stream.
All processing executes locally in your browser. No text is sent across the network, ensuring that sensitive documents and private notes remain entirely on your own machine.
Frequently asked questions
What is an invisible Unicode character detector?
An invisible Unicode character detector identifies unrendered, zero-width, or control code points within digital text. It highlights hidden characters such as zero-width spaces or bidirectional overrides without sending your text to an external server.
Do invisible Unicode characters prove text was generated by AI?
No. Invisible characters do not prove text came from any particular AI system. They frequently appear when copy-pasting from web pages, word processors, and PDF documents.
Will stripping invisible characters break emoji?
Yes, if done indiscriminately. Code points like zero-width joiners (U+200D) and variation selectors (U+FE0F) are essential for rendering multi-person emoji and colour glyphs correctly.
Can removing hidden Unicode bypass AI text detectors?
No tool guarantees bypassing AI detection. Hidden Unicode characters are ordinary character data rather than statistical watermarks, and removing them does not affect detector output in any way NoWatermark has measured. That is a statement about what we have tested, not a promise about any particular detector.
Related tools
Related guides
- Hidden Unicode characters explainedZero-width spaces, bidi overrides and tag characters — what they do, why they turn up in your text, and how to strip them safely.
- Does Claude watermark text?Two very different things get called "AI text watermarks". Only one of them is something you can find in your browser.