File privacy scanner
Every file format keeps notes. Photos record where they were taken; documents record who wrote them and what was deleted; spreadsheets hide worksheets rather than removing them.
A file privacy scanner looks at all of that before the file leaves your hands, and tells you in plain language what it found and what it would mean.
Works today, in your browser. Photos open in the editor and Word documents are checked for hidden data; either way the file is read on your device and never uploaded.
What gets checked, by format
Images and Word documents are live today. The rest is listed honestly below.
- Photos — GPS, device, timestamps, thumbnails, faces, QR codes and barcodes
- Word documents — author, last editor, company, comments, tracked changes, hidden text
- PDFs — document properties, annotations, attachments and redactions that do not actually redact (not switched on yet)
- Excel — hidden worksheets, hidden rows and columns, cell comments, revealing column headings (not switched on yet)
- Any text the file contains — card numbers, IBANs, passport MRZ lines, national IDs, tokens
The file type is read from the bytes, not the name
A PDF renamed to .jpg and handed to an image parser comes back with nothing found — and the person reads that as "nothing to worry about". That is a dangerous way for a privacy tool to be wrong.
So the type is detected from the magic bytes at the start of the file, and when the name disagrees with the content, the scanner says so.
Fixable means fixable
Every finding is marked with whether cleaning will actually remove it, and that is computed from where the finding is rather than from what type it is. Metadata, comments, revision history and attachments are removable. Something visible in the content is not.
Marking a finding fixable when no sanitizer touches it would promise a one-click fix that does nothing — and someone who believes that promise shares the file. A test asserts that no dirty fixture has a fixable finding left after cleaning.
Cleaning never destroys something you might need
Metadata, comments, revision history and hidden text go without asking. Hidden worksheets do not: deleting one can turn a working spreadsheet into a page of reference errors, so it is reported for you to decide about instead.
Anything that cannot be cleaned losslessly comes back unchanged with a reason — never silently unmodified, which reads as "there was nothing to remove".
Questions
Images and Word documents in the browser. PDF, Excel and CSV are built and tested in the engine but not switched on in the website yet.
Signed out, 10 MB per file and five checks a day. A free account raises that to 25 MB and twenty checks. The engine itself handles up to 64 MB.
No. The browser never sends it. On the API the bytes are held in memory for one request and are never written to disk, a database or a log — there is no column in the schema that could hold one.
That none of our detectors matched. It does not mean the file is safe, and we will not use that word. Detection is best-effort and a human look at the content is still worth the ten seconds.
Related
- Photo privacy checkerCheck a photo for GPS coordinates, camera details, timestamps and faces before you post it. Runs in your brows…
- Remove the author from a Word documentStrip the author, last-edited-by, company and comments from a .docx file, and accept tracked changes so delete…
- Check a document before you upload it to ChatGPTSee what a file really carries before you paste it into an AI chatbot: author names, comments, edit history an…
We report no issues found, never “safe”. Absence of detections is not proof of absence, and detection is best-effort.