Or drop it here. One file at a time, all in your browser.

Recovers text and tables into a clean docx. Complex layouts, images and macros do not come through. Pixel-faithful conversion needs Word or LibreOffice.

Text and tables come out of a binary .doc. Bold and italic are recovered only from RTF files and web pages.

The file never leaves this page: no upload, no server, nothing saved.

The same work, elsewhere

How it works

  1. Drop the document, or pick it with the selector.
  2. Read what was detected and check the extracted text.
  3. Download the .docx.

What it does

It recovers text and tables from a document in Word’s old binary format and writes them into a clean .docx. It also reads the two files that only resemble a .doc by name — RTF files and web pages saved by Word — and from those it recovers bold and italic too.

The promise, in full

It recovers text and tables into a clean docx. Complex layouts, images and macros do not come through. Pixel-faithful conversion needs Word or LibreOffice.

It is written above the button as well, because that is the thing to know before converting, not after opening the result.

The format is read from the bytes, not the extension

A file called .doc is often an RTF — that is what programs that “export to Word” without having Word produce — or an HTML page, if somebody used “save as web page”. They open in Word, which is what convinces people they are Word documents.

This tool looks at the first bytes and says what it found before converting. From RTF files and web pages it recovers the shape of the text as well; from a real binary .doc it recovers text and tables, and nothing more.

Why bold does not come out of the binary format

In a real .doc the text and its formatting live in two different places: the text in a piece table, the character properties in a separate structure, inside property pages indexed by position. It is a second format inside the first.

Promising it halfway would be worse than not promising it: a document where bold shows up in three places out of ten is harder to fix than one where it does not show up at all.

What it does not do

It is not a faithful conversion. Margins, columns, headers, footers, styles, numbered lists, section breaks and text boxes do not come through. The content comes out, not the layout. If the document matters because of how it is laid out, that needs Word or LibreOffice.

Images and macros do not come through. The first are embedded objects in formats that would each need decoding; the second are code, and this tool neither runs them nor carries them across.

Password-protected documents do not open, and neither do Word 6.0 or 95 files: they have a different header structure, outside what this page promises. In both cases it says which of the two it is, rather than handing back an empty file.

A document that yields not a single word is rejected. A valid, empty .docx is the worst defect in this category, because it looks like it worked.

About your data

Nothing leaves this page and nothing is saved.

From here you also go here

Convert files

Your files never leave your browser. Images, spreadsheets, documents and PDFs go in together, and each one says for itself what it can become.

Office

XLS to XLSX

An old .xls spreadsheet becomes an .xlsx that opens anywhere. It also reads CSVs and HTML tables disguised as Excel. All in the browser, the file never leaves.

Office

Images to PDF

Photographs and scans into one PDF, one per page. Page size, orientation and margins are yours to pick. All in the browser, the files never leave.

Office

All tools