Institutions | About Us | Help | Gaeilge
rian logo

Go Back
A study on mutual information-based feature selection for text categorization
Xu, Yang; Jones, Gareth J.F.; Li, Jintao; Wang, Bin; Sun, ChunMing
Feature selection plays an important role in text categorization. Automatic feature selection methods such as document frequency thresholding (DF), information gain (IG), mutual information (MI), and so on are commonly applied in text categorization. Many existing experiments show IG is one of the most effective methods, by contrast, MI has been demonstrated to have relatively poor performance. According to one existing MI method, the mutual information of a category c and a term t can be negative, which is in conflict with the definition of MI derived from information theory where it is always non-negative. We show that the form of MI used in TC is not derived correctly from information theory. There are two different MI based feature selection criteria which are referred to as MI in the TC literature. Actually, one of them should correctly be termed "pointwise mutual information" (PMI). In this paper, we clarify the terminological confusion surrounding the notion of "mutual information" in TC, and detail an MI method derived correctly from information theory. Experiments with the Reuters-21578 collection and OHSUMED collection show that the corrected MI method’s performance is similar to that of IG, and it is considerably better than PMI.
Keyword(s): Information retrieval; feature selection; text categorization; text categorisation
Publication Date:
Type: Other
Peer-Reviewed: Unknown
Language(s): English
Institution: Dublin City University
Citation(s): Xu, Yang, Jones, Gareth J.F. ORCID: 0000-0003-2923-8365 <>, Li, Jintao, Wang, Bin and Sun, ChunMing (2007) A study on mutual information-based feature selection for text categorization. Journal of Computational Information Systems, 3 (3). pp. 1007-1012. ISSN 1553-9105
Publisher(s): Binary Information Press
File Format(s): application/pdf
Related Link(s):
First Indexed: 2011-06-15 05:13:47 Last Updated: 2019-05-22 06:18:19