Document Classification and Information Extraction

doi:10.1007/978-1-4613-1295-6_4

Book Chapter•DOI•

Document Classification and Information Extraction

Qianhong Liu¹, Peter A. Ng¹•Institutions (1)

01 Jan 1996-pp 97-145

TL;DR: In TEXPROS, the task of document classification is to determine the types of the office documents and the document classification subsystem identifies the corresponding frame template of the document.

read less

Abstract: In Chapter 4 and 5, we turn our attention to the techniques used for document classification and information extraction [60, 61, 62, 174, 175]. In TEXPROS, the task of document classification is to determine the types of the office documents. That is, given an office document, the document classification subsystem identifies the corresponding frame template of the document. By identifying the defined type of the documents, it is possible to implement efficient storage and access methods to enhance the performance of retrieval. The task of information extraction is extracting from the contents of the document the most relevant information pertinent to the user. That is, given an office document, the information extraction subsystem forms its frame instance by instantiating its corresponding frame template. The document classification and information extraction can be achieved in aid of analyzing the document structures.

...read moreread less

Document Classification and Information Extraction

Citations

Related Papers (5)