Capability matrix
Every format this tool opens, every signal it can and cannot act on, and the date each claim was last checked.
Version 2.1 · Last updated . This page is generated from the same configuration the scanner and cleaners read, so it cannot describe a capability the code does not have.
Formats
“Supported” on its own is not a useful claim, so each row also states which parts of the container are opened and what is left uncovered.
JPEG
Inspect and cleanimage/jpeg · .jpg · image
Inspected
- APP1 EXIF
- APP1 XMP
- APP13 IPTC/Photoshop
- APP11 JUMBF (C2PA)
- COM comments
Not covered
- Pixel-domain watermarks cannot be measured
- Thumbnails inside EXIF are removed with it
PNG
Inspect and cleanimage/png · .png · image
Inspected
- tEXt, zTXt and iTXt chunks
- eXIf chunk
- XMP in iTXt
- caBX chunk (C2PA)
Not covered
- Pixel-domain watermarks cannot be measured
WebP
Inspect and cleanimage/webp · .webp · image
Inspected
- EXIF chunk
- XMP chunk
- C2PA chunk
- RIFF container structure
Not covered
- Pixel-domain watermarks cannot be measured
SVG
Inspect and cleanimage/svg+xml · .svg · text
Inspected
- XML comments
- metadata and RDF elements
- script elements and event handlers
- external references
- embedded data: URI images
Not covered
- An embedded raster image is scanned, but its own pixel content is not analysed
Markdown
Inspect and cleantext/markdown · .md · text
Inspected
- YAML frontmatter
- HTML comments
- hidden Unicode in body text
Not covered
- Linked and transcluded files are not followed
application/pdf · .pdf · document
Inspected
- the /Info dictionary of every revision, not only the newest
- XMP packets anywhere in the file
- document JavaScript and embedded files
- C2PA associated files and JUMBF containers
- the cross-reference chain and revision history
Not covered
- Encrypted PDFs are refused rather than modified
- A document whose structure cannot be rebuilt safely is refused, not guessed at
- Text content itself is not analysed
- Embedded file attachments are carried over and must be checked separately
Signals
Four outcomes, and the difference between the last two is the entire point of this product. Detected only means the signal is there and we will tell you, but we cannot take it out. Unable to verify means we cannot determine whether it is present at all — it is not a “no”, and it will never become one.
Provenance
A signed record of where a file came from and how it was edited.
Tags naming the AI tool or model that produced the image.
Google's imperceptible watermark, embedded in the pixels themselves.
Metadata
Camera, device, timestamp and settings data.
Adobe's XML metadata block, used by editors and AI tools.
Captions, keywords and rights information used by publishers.
PNG text chunks and JPEG comments, often holding prompts.
Colour interpretation data — preserved on purpose.
Privacy
Coordinates recorded when the photo was taken.
When the file was created or last modified.
Camera or phone make and model.
The application that created or last edited the file.
Creator name, copyright and ownership fields.
Script that runs when the file is opened.
Links to content fetched from another server on open.
Previous versions of the document, still inside the file.
Hidden text
Invisible characters that can carry a payload inside text.
Watermarks encoded in word choice rather than in characters.
What backs these claims
A capability table is only worth as much as the checking behind it. These are the validations the current version rests on.
733 real-world PDFs scanned locally: 99.3% parsed cleanly, no exceptions, median 0.2 ms. 3.7% carried metadata in a revision the current one had superseded.
727 of 733 real PDFs rebuilt successfully, 6 refused, none damaged. Every cleaned file parsed back, held exactly one revision, and contained no trace of its original author string. Page counts and content-stream bytes were identical before and after in all 727.
Tests compare the JPEG scan stream, PNG IDAT and WebP VP8/VP8L payloads before and after cleaning and require them to be byte-identical.
Every removal claim is produced by scanning the cleaned output a second time and diffing against the original scan, never by a cleaner reporting its own success.
Standing limits
These are not gaps waiting to be closed in a later release. They are consequences of where the processing happens.
- Pixel and statistical watermarks cannot be measured. SynthID and statistical text watermarks are detected by systems the model providers hold. A tool running on your device cannot confirm their presence or their absence, so it reports neither.
- Server-side provenance is out of reach. A service can re-associate a file with its records using a perceptual fingerprint held in its own database. Removing everything inside the file does not affect that, and nothing local can.
- C2PA signatures are not cryptographically verified. A manifest is reported as present, never as valid or invalid.
- Removal is only claimed after a re-scan. The cleaned output is scanned again by the same engine and diffed against the original. A cleaner reporting its own success is not evidence.
- Nothing is recompressed. Cleaners rewrite the container and copy the compressed image data byte for byte, so cleaning never costs image quality.
For how any of this works, see the methodology. For what leaves your device — which is nothing, unless you explicitly opt into text rewriting — see theprivacy page.