Detecting disguised missing data

Download
2009
Belen, Rahime
In some applications, explicit codes are provided for missing data such as NA (not available) however many applications do not provide such explicit codes and valid or invalid data codes are recorded as legitimate data values. Such missing values are known as disguised missing data. Disguised missing data may affect the quality of data analysis negatively, for example the results of discovered association rules in KDD-Cup-98 data sets have clearly shown the need of applying data quality management prior to analysis. In this thesis, to tackle the problem of disguised missing data, we analyzed embedded unbiased sample heuristic (EUSH), demonstrated the methods drawbacks and proposed a new methodology based on Chi Square Two Sample Test. The proposed method does not require any domain background knowledge and compares favorably with EUSH.

Suggestions

Classification of remotely sensed data by using 2D local discriminant bases
Tekinay, Çağrı; Çetin, Yasemin; Department of Information Systems (2009)
In this thesis, 2D Local Discriminant Bases (LDB) algorithm is used to 2D search structure to classify remotely sensed data. 2D Linear Discriminant Analysis (LDA) method is converted into an M-ary classifier by combining majority voting principle and linear distance parameters. The feature extraction algorithm extracts the relevant features by removing the irrelevant ones and/or combining the ones which do not represent supplemental information on their own. The algorithm is implemented on a remotely sensed...
Constructing linear unequal error protection codes from algebraic curves
Özbudak, Ferruh (Institute of Electrical and Electronics Engineers (IEEE), 2003-06-01)
We show that the concept of "generalized algebraic geometry codes" which was recently introduced by Xing, Niederreiter, and Lam gives a natural framework for constructing linear unequal error protection codes.
Spatially Coupled Codes Optimized for Magnetic Recording Applications
Esfahanizadeh, Homa; Hareedy, Ahmed; Dolecek, Lara (2017-02-01)
© 1965-2012 IEEE.Spatially coupled (SC) codes are a class of sparse graph-based codes known to have capacity-approaching performance. SC codes are constructed based on an underlying low-density parity-check (LDPC) code, by first partitioning the underlying block code and then putting replicas of the components together. Significant recent research efforts have been devoted to the asymptotic, ensemble-averaged study of SC codes, as these coupled variants of the existing LDPC codes offer excellent properties....
A genetic-based intelligent intrusion detection system
Özbey, Halil; Şen, Tayyar; Department of Industrial Engineering (2005)
In this study we address the problem of detecting new types of intrusions to computer systems which cannot be handled by widely implemented knowledge-based mechanisms. The solutions offered by behavior-based prototypes either suffer low accuracy and low completeness or require use data eplaining abnormal behavior which actually is not available. Our aim is to develop an algorithm which can produce a satisfactory model of the target system̕s behavior in the absence of negative data. First, we design and deve...
Using Pad-Stripped Acausally Filtered Strong-Motion Data
Boore, David M.; Sisi, Aida Azari; Akkar, Dede Sinan (2012-04-01)
Most strong-motion data processing involves acausal low-cut filtering, which requires the addition of sometimes lengthy zero pads to the data. These padded sections are commonly removed by organizations supplying data, but this can lead to incompatibilities in measures of ground motion derived in the usual way from the padded and the pad-stripped data. One way around this is to use the correct initial conditions in the pad-stripped time series when computing displacements, velocities, and linear oscillator ...
Citation Formats
R. Belen, “Detecting disguised missing data,” M.S. - Master of Science, Middle East Technical University, 2009.