将Word转成Markdown:word2markdown
jopen
10年前
这个工具能够将 Word 转成 Markdown,包含图片和Math。 它由9个连续的步骤:
- Exporting to HTML using Microsoft Word 2012. We automated this on OS X using Automator. Solutions for other platforms are welcome!
- Extracting image types that we want to use. Keeps the original quality, unless that's a proprietary .emz file. In this step we also fix some math.
- Converting HTML to XML using tagsoup.
- Covert OOML (proprietary Word format) into MathML equations, using Microsoft's own conversion XSLT, and a custom version of this XSLT. Uses Saxon 8.
- Some intermediate fixes for whitespace and math.
- Conversion back into HTML using Tidy. Also strips a lot of stuff.
- More intermediate fixes to deal with shortcomings of Tidy and Pandoc.
- Conversion into Markdown using Pandoc.
- Lots of cleanup and final fixes to the Markdown.