Remove PDF metadata

See the author and dates a PDF carries, then clear them out.

Your inputs stay on your deviceFREE · NO SIGN-UP

A small task starts here.

Drop your PDF here, paste it, or choose from your device.

Up to 100 MB total · Processed locally
Remove

Choose a PDF and the result appears here automatically.

THE LITTLE DETAILS

Remove PDF metadata, without the extra steps.

A PDF says who made it, what they called the file, which program wrote it and when - and it usually says all of it twice. The document properties are the pair of strings a reader shows you; the XMP packet is the same facts again as XML, in plain text inside the file, and a tool that clears only the first leaves the name sitting there for anyone who opens the document in a text editor. This lists everything it finds, in the words a person would use, and then clears both. The pages themselves are copied across untouched.

How to use this tool

  1. 1Choose a PDF. What the document says about itself is listed straight away.
  2. 2Remove everything, or keep the title and drop the rest.
  3. 3Check each line is marked Removed, then download the clean file.

When Remove PDF metadata is the right tool

  • A CV is going to a stranger and the file still says it was written by whoever the laptop's copy of Word was registered to.
  • A tender has to be anonymous for blind review, and the author, the file name and the creation date each give it away.
  • A report was exported from InDesign, which writes the whole team's names into an XMP packet as well as the properties panel.
  • You are publishing a document and would rather not advertise which version of which program made it.
  • A contract came back from the other side and you want to see what their copy records before you pass it on.
  • You want to check whether a PDF someone sent you carries anything you were not meant to see.

The name is in the file twice, and most tools clear it once

Document properties are what a reader shows in File, Properties: title, author, subject, the program, two dates. XMP is the same information again as an XML packet stored inside the PDF, usually uncompressed, which means the author's name is sitting in the file as readable text. Clear the properties alone and the properties panel goes blank while the name stays exactly where it was. Both go here, and the page lists what was in each before anything is downloaded.

Removed means it was checked, not that it was attempted

Every line in the report is marked from the file that was actually written, not from the code that wrote it: the finished PDF is read back and each thing found in the original is looked for again. That is also why an unreferenced object is not good enough. Deleting the pointer to an XMP packet stops a reader showing it and leaves the packet in the file; the object itself is deleted, so the words are gone rather than merely unlisted.

What is kept, and why it is kept

Files attached inside a PDF are contents rather than metadata. They are named on the page and left alone, because deleting a document's attachments is a different act from anonymising it - but they carry their own author fields, so an attached original is worth checking. A digital signature cannot survive: it signs the exact bytes, and removing anything writes new ones. Keeping the title is offered because it is often the document's real title, and it is what a screen reader announces.

The pages are not touched at all

The page contents, images and fonts are copied across as they are; nothing is rasterised, re-compressed or re-flowed, so a 40 MB scan comes out the same 40 MB scan. That also sets the limit of what this can do. Anything printed on a page is still printed on it - a letterhead, a name in a signature block, a watermark. This changes what the file says about itself, not what it shows.

GOOD TO KNOW

A few quick answers.

Choose a PDF. What the document says about itself is listed straight away. Remove everything, or keep the title and drop the rest. Check each line is marked Removed, then download the clean file.

Tool inputs are processed in your browser. Crate has no upload endpoint and no accounts. Downloaded site assets and a cookie-free page-view count still require a connection; your files and text are not included in those requests.

Both stores of metadata are cleared: the document properties and the XMP packet, which most tools miss. A password-protected PDF cannot be written again here. A digital signature does not survive the file being rewritten, and the page says so before you download. Files attached inside the PDF are left alone, with their own metadata intact. Anything visible on the page stays visible.

It should not. The library this is built on stamps its own name and a fresh modification date into every document it opens, which would mean a metadata remover adding metadata, so that behaviour is switched off on every read and write here - including the one that reads the finished file back to check it. If you see a date in the report marked Removed, the check found it gone from the written file.

No, and that is worth being clear about. This clears what the document records about itself. Redacting what is printed on a page is a different job with different risks - text covered by a black box is still text underneath - and it is not what this does.

Not here. An encrypted PDF can be read well enough to list what it records, but writing it again would need the password, so the page names the problem and locks the download rather than handing back something broken. Remove the password in your PDF reader, then come back.

An XML block Adobe introduced so that metadata could travel in the file rather than beside it. In a PDF it holds the title, the author, the program, timestamps and often a document ID that links every copy and revision of the same file. It is usually stored uncompressed, which is why it matters: no tools are needed to read it, only a text editor.

Back to all tools