Data Mining Feature Subset Weighting and Selection Using Genetic Algorithms

Abstract

We present a simple genetic algorithm (sGA), which is developed under Genetic Rule and Classifier Construction Environment (GRaCCE) to solve feature subset selection and weighting problem to have better classification accuracy on k-nearest neighborhood (KNN) algorithm. Our hypotheses are that weighting the features will affect the performance of the KNN algorithm and will cause better classification accuracy rate than that of binary classification. The weighted-sGA algorithm uses real-value chromosomes to find the weights for features and binary-sGA uses integer-value chromosomes to select the subset of features from original feature set. A Repair algorithm is developed for weighted-sGA algorithm to guarantee the feasibility of chromosomes. By feasibility we mean that the sum of values of each gene in a chromosome must be equal to 1. To calculate the fitness values for each chromosome in the population, we use K Nearest Neighbor Algorithm (KNN) as our fitness function. The Euclidean distance from one individual to other individuals is calculated on the d-dimensional feature space to classify an unknown instance. GRaCCE searches for good feature subsets and their associated weights. These feature weights are then multiplied with normalized feature values and these new values are used to calculate the distance between features.

Open PDF

Document Details

Document Type
Technical Report
Publication Date
Mar 01, 2002
Accession Number
ADA401635

Entities

Organizations

  • Air Force Institute of Technology

Tags

Communities of Interest

  • Autonomy
  • Energy and Power Technologies

DTIC Thesaurus Topics

  • Artificial Intelligence
  • Birds
  • Computer Programming
  • Computers
  • Data Mining
  • Databases
  • Evolutionary Algorithms
  • Experimental Design
  • Fish
  • Fungi
  • Information Science
  • Machine Learning
  • Medical Personnel
  • Network Science
  • Neural Networks
  • Operating Systems
  • Pattern Recognition

Fields of Study

  • Computer science

Readers

  • Approximation Theory.
  • Molecular and genetic basis of cancer.
  • Neural Network Machine Learning.

Technology Areas

  • AI & ML
  • AI & ML - Machine Learning Algorithms
  • Biotechnology
  • Space
  • Space - Space Objects