📜  Apache POI Word-文本提取

📅  最后修改于: 2020-11-18 08:57:07             🧑  作者: Mango


本章介绍如何使用Java从Word文档中提取简单文本数据。如果要从Word文档中提取元数据,请使用Apache Tika。

对于.docx文件,我们使用org.apache.poi.xwpf.extractor.XPFFWordExtractor类从Word文件中提取并返回简单数据。同样,我们有不同的方法来从Word文件中提取标题,脚注,表格数据等。

以下代码显示了如何从Word文件中提取简单文本-

import java.io.FileInputStream;
import org.apache.poi.xwpf.extractor.XWPFWordExtractor;
import org.apache.poi.xwpf.usermodel.XWPFDocument;

public class WordExtractor {

   public static void main(String[] args)throws Exception {

      XWPFDocument docx = new XWPFDocument(new FileInputStream("create_paragraph.docx"));
      
      //using XWPFWordExtractor Class
      XWPFWordExtractor we = new XWPFWordExtractor(docx);
      System.out.println(we.getText());
   }
}

将上面的代码另存为WordExtractor.java。从命令提示符下编译并执行它,如下所示:

$javac WordExtractor.java
$java WordExtractor

它将生成以下输出:

At tutorialspoint.com, we strive hard to provide quality tutorials for self-learning
purpose in the domains of Academics, Information Technology, Management and Computer
Programming Languages.