Skip to content
PatentGenius

United States Patent

US Patent 6973423: Article and method of automatically determining…

US 6973423  ·  granted 2005-12-06

Abstract

A processor implemented method of identifying the text genre of a machine-readable, untagged text. The processor implemented method begins by generating a cue vector from the text, which represents occurrences in the text of a first set of nonstructural, surface cues, which are easily computable. Afterward, the processor determines whether the text is an instance of a first text genre using the cue vector and a weighting vector associated with the first text genre.

Patent Number 6973423
Title Article and method of automatically determining text genre using surface features of untagged texts
Filed 1998-06-18
Granted 2005-12-06
Inventor(s) Grefenstette; Gregory, Kessler; Brett L., Nunberg; Geoffrey D., Pedersen; Jan O., Schuetze; Hinrich
Assignee Xerox Corporation
Number of Claims 99

Abstract

A processor implemented method of identifying the text genre of a machine-readable, untagged text. The processor implemented method begins by generating a cue vector from the text, which represents occurrences in the text of a first set of nonstructural, surface cues, which are easily computable. Afterward, the processor determines whether the text is an instance of a first text genre using the cue vector and a weighting vector associated with the first text genre.

Claim 1

A processor implemented method of identifying a document type of a document in machine-readable form without structurally analyzing the document text, the processorimplemented method comprising: a) selecting a first set of nonstructural surface cues; b) generating a cue vector from the text, the cue vector having a valve for each of the selected cues and representing a frequency of occurrences in the text of thefirst set of nonstructural surface cues; c) associating a weighing vector with a first text genre; and d) determining whether the text is an instance of the first text genre using the cue vector and the weighting vector associated with the first textgenre, wherein the first set of cues includes a punctuational cue.

Claims

99 total

A processor implemented method of identifying a document type of a document in machine-readable form without structurally analyzing the document text, the processorimplemented method comprising: a) selecting a first set of nonstructural surface cues; b) generating a cue vector from the text, the cue vector having a valve for each of the selected cues and representing a frequency of occurrences in the text of thefirst set of nonstructural surface cues; c) associating a weighing vector with a first text genre; and d) determining whether the text is an instance of the first text genre using the cue vector and the weighting vector associated with the first textgenre, wherein the first set of cues includes a punctuational cue.