IRMA-International.org: Creator of Knowledge
Information Resources Management Association
Advancing the Concepts & Practices of Information Resources Management in Modern Organizations

Spam Detection Approaches with Case Study Implementation on Spam Corpora

Author(s): Biju Issac (Swinburne University of Technology (Sarawak Campus), Malaysia)
Copyright: 2011
Pages: 19
EISBN13: 9781609606534

Purchase

View Spam Detection Approaches with Case Study Implementation on Spam Corpora on the publisher's website for pricing and purchasing information.

View Sample PDF


Abstract

Email has been considered as one of the most efficient and convenient ways of communication since the users of the Internet has increased rapidly. E-mail spam, known as junk e-mail, UBE (unsolicited bulk e-mail) or UCE (unsolicited commercial e-mail), is the act of sending unwanted e-mail messages to e-mail users. Spam is becoming a huge problem to most users since it clutter their mailboxes and waste their time to delete all the spam before reading the legitimate ones. They also cost the user money with dial up connections, waste network bandwidth and disk space and make available harmful and offensive materials. In this chapter, initially we would like to discuss on existing spam technologies and later focus on a case study. Though many anti-spam solutions have been implemented, the Bayesian spam detection approach looks quite promising. A case study for spam detection algorithm is presented and its implementation using Java is discussed, along with its performance test results on two independent spam corpuses – Ling-spam and Enron-spam. We use the Bayesian calculation for single keyword sets and multiple keywords sets, along with its keyword contexts to improve the spam detection and thus to get good accuracy. The use of porter stemmer algorithm is also discussed to stem keywords which can improve spam detection efficiency by reducing keyword searches.

Related Content

Stuart Palmer, Dale Holt. © 2010. 18 pages.
Moh’d Jarrar. © 2013. 24 pages.
Christina Badman, Matthew DeNote. © 2013. 25 pages.
Kevin Gosselin, Hansel Burley. © 2012. 20 pages.
Evan S. Smith, Terrie Nagel. © 2010. 18 pages.
Body Bottom