• POI转换word doc文件为(html,xml,txt)


    在POI中还存在有针对于word doc文件进行格式转换的功能。我们可以将word的内容转换为对应的Html文件,也可以把它转换为底层用来描述doc文档的xml文件,还可以把它转换为底层用来描述doc文档的xml格式的text文件。这些格式转换都是通过AbstractWordConverter特定的子类来完成的。 

    1 转换为Html文件

    将doc文档转换为对应的Html文档是通过WordToHtmlConverter类进行的。它会尽量的利用Html的方式来呈现原文档的样式。示例代码:

       /**
        * Word转换为Html
        * @throws Exception
        */
       @Test
       public void testWordToHtml() throws Exception {
          InputStream is = new FileInputStream("D:\test.doc");
          HWPFDocument wordDocument = new HWPFDocument(is);
          WordToHtmlConverter converter = new WordToHtmlConverter(DocumentBuilderFactory.newInstance().newDocumentBuilder().newDocument());
          //对HWPFDocument进行转换
          converter.processDocument(wordDocument);
            Writer writer = new FileWriter(new File("D:\converter.html"));
           Transformer transformer = TransformerFactory.newInstance().newTransformer();
           transformer.setOutputProperty( OutputKeys.ENCODING, "utf-8" );
           //是否添加空格
            transformer.setOutputProperty( OutputKeys.INDENT, "yes" );
           transformer.setOutputProperty( OutputKeys.METHOD, "html" );
           transformer.transform(
                       new DOMSource(converter.getDocument() ),
                       new StreamResult( writer ) );
       }

    2 转换为Xml文件

           将doc文档转换为对应的Xml文件是通过WordToFoConverter类进行的。它可以把doc文档转换为底层用来描述doc文档的Xml文档。示例代码:

       /**
        * Word转Fo
        * @throws Exception
        */
       @Test
       public void testWordToFo() throws Exception {
          InputStream is = new FileInputStream("D:\test.doc");
          HWPFDocument wordDocument = new HWPFDocument(is);
          WordToFoConverter converter = new WordToFoConverter(DocumentBuilderFactory.newInstance().newDocumentBuilder().newDocument());
          //对HWPFDocument进行转换
          converter.processDocument(wordDocument);
            Writer writer = new FileWriter(new File("D:\converter.xml"));
           Transformer transformer = TransformerFactory.newInstance().newTransformer();
           transformer.setOutputProperty( OutputKeys.ENCODING, "utf-8" );
           //是否添加空格
            transformer.setOutputProperty( OutputKeys.INDENT, "yes" );
    //     transformer.setOutputProperty( OutputKeys.METHOD, "html" );
           transformer.transform(
                       new DOMSource(converter.getDocument() ),
                       new StreamResult( writer ) );
       }
     

    3  转换为Text文件

           将doc文档转换为text文档是通过WordToTextConverter来进行的。它可以把doc文档转换为底层用于描述doc文档的Xml格式的text文档。示例代码:

       /**
        * Word转换为Text
        * @throws Exception
        */
       @Test
       public void testWordToText() throws Exception {
          InputStream is = new FileInputStream("D:\test.doc");
          HWPFDocument wordDocument = new HWPFDocument(is);
          WordToTextConverter converter = new WordToTextConverter(DocumentBuilderFactory.newInstance().newDocumentBuilder().newDocument());
          //对HWPFDocument进行转换
          converter.processDocument(wordDocument);
            Writer writer = new FileWriter(new File("D:\converter.txt"));
           Transformer transformer = TransformerFactory.newInstance().newTransformer();
           transformer.setOutputProperty( OutputKeys.ENCODING, "utf-8" );
           //是否添加空格
            transformer.setOutputProperty( OutputKeys.INDENT, "yes" );
           transformer.setOutputProperty( OutputKeys.METHOD, "text" );
           transformer.transform(
                       new DOMSource(converter.getDocument() ),
                       new StreamResult( writer ) );
       }
  • 相关阅读:
    6.VUE事件处理
    springmvc在使用@ModelAttribute注解获取Request和Response会产生线程并发不安全问题
    IDEAhttp://lookdiv.com/index/index/indexcodeindex.html
    不四舍五入保留...4(round(273.86015,4,1);)
    spring security中@PreAuthorize、@PostAuthorize、@PreFilter和@PostFilter四者的区别
    @RepeatSubmit spring boot 防止重复提交
    权限设计的杂谈
    vue设置全局样式变量 less
    坐标轴刻度取值算法-基于魔数数组-源于echarts的y轴刻度计算需求
    less使用
  • 原文地址:https://www.cnblogs.com/estellez/p/4091385.html
Copyright © 2020-2023  润新知