Methodology
What we inspect, what we can remove, and — the part most tools leave out — what we cannot determine at all.
Capability matrix
This table is generated from the same configuration the scanner and cleaners read. If a row says No, the product cannot make that claim anywhere in the interface.
For the full version — every format, which parts of each container are opened, what is not covered, and the dated validation behind each claim — see thecapability matrix.
| Signal | Detect | Remove | Verify |
|---|---|---|---|
| EXIF | Yes | Yes | Yes |
| XMP | Yes | Yes | Yes |
| IPTC | Yes | Yes | Yes |
| GPS | Yes | Yes | Yes |
| PNG text chunks | Yes | Yes | Yes |
| C2PA / Content CredentialsPresence only — signatures are not cryptographically verified, and cloud-side recovery is out of scope. | Yes | Yes | Yes |
| AI generator metadata | Yes | Yes | Yes |
| Hidden Unicode | Yes | Yes | Yes |
| PDF document metadataThe document is rebuilt from scratch, never appended to, and the output is checked against its raw bytes before it is offered. | Yes | Yes | Yes |
| PDF earlier revisionsPrevious versions left in the file by incremental saves. A rebuild drops them entirely — the cleaned file has exactly one revision. | Yes | Yes | Yes |
| SVG active contentScripts and event handlers in SVG files. | Yes | Yes | Yes |
| SVG remote referencesExternal URLs that the viewer's browser would fetch on open. | Yes | Yes | Yes |
| SynthIDEmbedded in pixels. We cannot confirm presence or absence. | No | No | No |
| Statistical text watermarksEncoded in word choice. Not detectable client-side. | No | No | No |
How processing works
When you select a file, the browser reads it into memory and hands the bytes to a Web Worker running on your own device. The worker identifies the real format from its magic bytes — never the extension or the reported MIME type — walks the container structure, and extracts the metadata blocks it recognises.
Cleaning rewrites the container without the metadata structures and copies the compressed image data through untouched. Nothing is decoded and nothing is re-encoded, so the pixels in the output are bit-identical to the input.
The cleaned bytes are then scanned a second time, and the before-and-after report you see is the difference between those two scans. A signal is reported as removed only because a fresh scan can no longer find it — never because a cleaner asserted it.
What we keep on purpose
Colour profiles. ICC profiles carry no personal or provenance information, and removing one visibly shifts an image's colours. We report them as detected and deliberately preserved.
Image rotation. Phones record orientation in an EXIF tag rather than rotating pixels. Stripping it makes portrait photos display sideways. When a file carries a rotation, we preserve exactly that one field, say so in the report, and offer a switch to remove it anyway.
Structural data. JFIF headers, Adobe colour-transform markers, and every critical chunk needed to decode the image.
The three statuses
Detected — the signal was found in the file.
Not detected — we inspected the places that signal lives and it was not there. We only use this where absence is genuinely establishable.
Unable to verify — we have no way to look. SynthID reads this on every image, permanently, because detecting it requires Google's own detector. This is not a weaker form of "not detected"; it is a statement about our instruments rather than about your file, and it never appears in a removed list.
Known limitations
- C2PA manifests are detected and partially read, but signatures are not verified against a trust list. We never report a manifest as valid or invalid.
- Durable Content Credentials can be re-associated with an image by a provider after the manifest is removed, using an invisible watermark or content fingerprint. Removing a manifest cleans your copy of the file, not anyone else's records.
- Pixel-embedded watermarks such as SynthID are neither detectable nor removable here.
- Statistical text watermarks are not detectable client-side by anyone, including us. We offer an optional rewrite for pasted text, which rewords it in an attempt to disturb one. It is the only feature on this site that sends anything anywhere: it asks every time, shows you the exact payload first, and sends only the text you pasted — never a file. It runs on Gemini's paid tier, where Google does not train on the text, though it is logged for up to 55 days for abuse detection — see privacy. Because we have no detector, its result is reported as"Rewritten — unverified" and never counted as a removal.
- We can scan JPG, PNG, WebP, SVG, Markdown and PDF, and we can clean all of them. A PDF is cleaned by rebuilding the document rather than appending an update, and one whose structure cannot be rebuilt safely is refused rather than guessed at. AVIF, HEIC, GIF, TIFF, audio and video are not supported at all.
- In a Markdown file we remove the frontmatter keys that describe how the document was made — generator, model, prompt, author, dates — and leave the keys that drive your site build, such as title, tags and layout. Where the frontmatter uses YAML anchors, one key can depend on another, so we report what is there and change nothing.
- PDF is inspect-only. A PDF is append-only: saving it again writes a new version on top of the old one rather than replacing it, so every earlier revision stays in the file and stays readable. Tools that "remove" PDF metadata by appending an update leave the original recoverable. We report what is in the file, including how many earlier revisions it carries, and we do not offer a removal we cannot yet do correctly. Encrypted PDFs are reported and left alone.
- In an SVG, images embedded as data URIs are unpacked and cleaned in place, so metadata inside a nested JPEG is removed rather than overlooked. An SVG that references an image on another server has that reference removed — the content was never in your file, and fetching it would reveal your IP address to that server.
- Very large files are limited by browser memory rather than by any server constraint, since there is no server.
Verifying our privacy claim yourself
Load any tool page, disconnect from the internet, then scan and clean a file. Everything works, because processing was always local. Alternatively, open your browser's network panel and watch while you use the tool: no request carries your file, its contents, its name, or anything derived from them.