ALL DEVELOPER TOOLS56
DOCX to Markdown
Convert DOCX to Markdown — Free, No Upload
Read a Word document in your browser and get Markdown, with a count of everything the trip cost.
About DOCX to Markdown
A .docx is a ZIP full of XML, and this reads it the way the format works: the package is unzipped in your browser, the body is mapped to a small, plain HTML, and that HTML is written out as Markdown by the same writer the HTML-to-Markdown page uses. Nothing is uploaded. Word can express far more than Markdown can, so the question is not whether something is lost — it is what, and how much. Rather than print a standing list of caveats, this page opens the package a second time and counts. Page headers and footers, comments, tracked deletions, underline, highlight, colour, embedded objects and macros are all real parts of a .docx and none survives; the panel above the preview names the ones your document has, with the numbers. What survives is the structure worth keeping. Headings come from Word's Heading styles rather than from font size, lists keep their nesting, footnotes become GFM footnotes, and tables become real pipe tables — except where a vertical merge means no pipe table could be honest, in which case the HTML is kept and the page says so.
Questions
My headings all came through as ordinary paragraphs.
In Word a heading is a style, not a size. The reader maps Word's built-in Heading 1 to Heading 6 styles onto # through ######, and a paragraph that is merely bold 18pt Calibri is, structurally, body text — there is nothing in the file to say otherwise, and guessing from font size would invent headings in documents that have none. The fix is in Word: put the cursor in the line, pick Heading 1 from the Styles gallery, and convert again. The HEADINGS count above the preview tells you how many the file really had before you download anything. If it reads 0 on a document full of titles, that is the reason.
Where did my pictures go?
Nowhere you did not send them. IMAGES has three settings. Link, the default, writes each picture into a media/ folder and points the Markdown at it, so the download becomes a ZIP — keep that folder beside the .md or the links break. Embed writes each picture into the .md itself as a base64 data URI: one self-contained file, at four characters of text for every three bytes of picture. Drop removes them and reports the count. A picture used only in a page header, a footer or a list bullet appears under none of the three, because only the document body is read; the panel says how many of those there are. One more kind goes missing under all three: a picture inserted with Link to File rather than Insert > Picture is not inside the .docx at all — it is a path to your disk, which a browser cannot follow — so it reaches neither the Markdown nor the download. The page counts those separately and says so rather than promising a media/ folder it could not write.
It says my file is an OLE compound document, and it is definitely a Word file.
It is a Word file — just not a .docx. Two different things arrive under that extension and both begin with the same eight bytes, D0 CF 11 E0 A1 B1 1A E1: a Word 97–2003 .doc, which is an entirely different binary format, and a .docx with a password on it, because encrypting an Office file wraps the whole ZIP inside an OLE container. Neither can be unzipped, so neither can be read here, and the page says which one it saw rather than reporting a broken ZIP. Open it in Word or LibreOffice, remove the password if there is one, and save it again as .docx.
A merged cell knocked my table out of line.
A Markdown pipe table has no merged cells at all, so the two kinds of merge are handled differently and both are counted. A cell spanning columns keeps its text in the first of them and leaves the rest empty, which keeps every later column under the right heading. A cell spanning rows cannot be faked that way — the rows beneath it are short, and every value in them would slide one column to the left — so a table containing one is written out as its original HTML instead, which most Markdown renderers display and a plain-text reader shows as tags. Set TABLES to Keep HTML to get that for every table.
Are the comments and tracked changes in the Markdown?
No comments, and tracked changes only in their accepted form. Comments live in word/comments.xml; this page opens that part purely to count them, and none of their text reaches the output — the commented sentence comes through, the comment on it does not. Tracked changes are the more dangerous half: inserted text is kept, deleted text is gone, and nothing marks which was which, so the Markdown reads as though every change had been accepted. The panel counts insertions and deletions separately, so a contract or a redline tells you before you rely on it. If that distinction matters, accept or reject the changes in Word first.
Is my file uploaded to a server?
No. Transmute processes everything locally in your browser using JavaScript and WebAssembly. Your files never leave your device — there is no server, no upload, no cloud processing.