|
An Assessment of Case-Based Reasoning for Spam Filtering |
Delany, Sarah Jane; Cunningham, Padraig; Coyle, Lorcan
|
|
|
|
TCD-CS-2004-44 Because of the changing nature of spam, a spam filtering
system that uses machine learning will need to be dynamic. This
suggests that a case-based (memory-based) approach may work well.
Case-Based Reasoning (CBR) is a lazy approach to machine learning
where induction is delayed to run time. This means that the case base
can be updated continuously and new training data is immediately
available to the induction process. In this paper we present a detailed
description of such a system called ECUE and evaluate design
decisions concerning the case representation. We compare its
performance with an alternative system that uses Naive Bayes (NB).
We find that there is little to choose between the two alternatives in
cross-validation tests on data sets. However, ECUE does appear to have
some advantages in tracking concept drift over time.
|
|
Keyword(s):
|
Computer Science |
Publication Date:
|
2004 |
|
Type:
|
Report |
|
Peer-Reviewed:
|
Unknown |
|
Language(s):
|
English |
|
Institution:
|
Trinity College Dublin |
|
Citation(s):
|
Delany, Sarah Jane; Cunningham, Padraig; Coyle, Lorcan. 'An Assessment of Case-Based Reasoning for Spam Filtering'. - Dublin, Trinity College Dublin, Department of Computer Science, TCD-CS-2004-44, 2004, pp15 |
|
Publisher(s):
|
Trinity College Dublin, Department of Computer Science |
|
File Format(s):
|
application/pdf |
|
First Indexed:
2014-05-13 05:31:17 Last Updated:
2015-04-10 05:13:57 |