Using visual statistical inference to better understand random class separations in high dimension, low sample size data

Niladri Roy Chowdhury, Dianne Cook, Heike Hofmann, Mahbubul Majumder, Eun Kyung Lee, Amy L. Toth

Research output: Contribution to journalArticle

5 Citations (Scopus)

Abstract

Statistical graphics play an important role in exploratory data analysis, model checking and diagnosis. With high dimensional data, this often means plotting low-dimensional projections, for example, in classification tasks projection pursuit is used to find low-dimensional projections that reveal differences between labelled groups. In many contemporary data sets the number of observations is relatively small compared to the number of variables, which is known as a high dimension low sample size (HDLSS) problem. This paper explores the use of visual inference on understanding low-dimensional pictures of HDLSS data. Visual inference helps to quantify the significance of findings made from graphics. This approach may be helpful to broaden the understanding of issues related to HDLSS data in the data analysis community. Methods are illustrated using data from a published paper, which erroneously found real separation in microarray data, and with a simulation study conducted using Amazon’s Mechanical Turk.

Original languageEnglish (US)
Pages (from-to)293-316
Number of pages24
JournalComputational Statistics
Volume30
Issue number2
DOIs
StatePublished - Nov 2 2014

Fingerprint

Statistical Inference
Higher Dimensions
Sample Size
Model checking
Microarrays
Statistical Graphics
Projection
Projection Pursuit
Exploratory Data Analysis
High-dimensional Data
Microarray Data
Model Checking
Data analysis
Quantify
Simulation Study
Vision
Class
Sample size
Statistical inference
Inference

Keywords

  • Data mining
  • Lineup
  • Projection pursuit
  • Statistical graphics
  • Visualization

ASJC Scopus subject areas

  • Statistics and Probability
  • Computational Mathematics
  • Statistics, Probability and Uncertainty

Cite this

Using visual statistical inference to better understand random class separations in high dimension, low sample size data. / Roy Chowdhury, Niladri; Cook, Dianne; Hofmann, Heike; Majumder, Mahbubul; Lee, Eun Kyung; Toth, Amy L.

In: Computational Statistics, Vol. 30, No. 2, 02.11.2014, p. 293-316.

Research output: Contribution to journalArticle

Roy Chowdhury, Niladri ; Cook, Dianne ; Hofmann, Heike ; Majumder, Mahbubul ; Lee, Eun Kyung ; Toth, Amy L. / Using visual statistical inference to better understand random class separations in high dimension, low sample size data. In: Computational Statistics. 2014 ; Vol. 30, No. 2. pp. 293-316.
@article{2fc15ba3f2e04aff9a89ce00e9cc111a,
title = "Using visual statistical inference to better understand random class separations in high dimension, low sample size data",
abstract = "Statistical graphics play an important role in exploratory data analysis, model checking and diagnosis. With high dimensional data, this often means plotting low-dimensional projections, for example, in classification tasks projection pursuit is used to find low-dimensional projections that reveal differences between labelled groups. In many contemporary data sets the number of observations is relatively small compared to the number of variables, which is known as a high dimension low sample size (HDLSS) problem. This paper explores the use of visual inference on understanding low-dimensional pictures of HDLSS data. Visual inference helps to quantify the significance of findings made from graphics. This approach may be helpful to broaden the understanding of issues related to HDLSS data in the data analysis community. Methods are illustrated using data from a published paper, which erroneously found real separation in microarray data, and with a simulation study conducted using Amazon’s Mechanical Turk.",
keywords = "Data mining, Lineup, Projection pursuit, Statistical graphics, Visualization",
author = "{Roy Chowdhury}, Niladri and Dianne Cook and Heike Hofmann and Mahbubul Majumder and Lee, {Eun Kyung} and Toth, {Amy L.}",
year = "2014",
month = "11",
day = "2",
doi = "10.1007/s00180-014-0534-x",
language = "English (US)",
volume = "30",
pages = "293--316",
journal = "Computational Statistics",
issn = "0943-4062",
publisher = "Springer Verlag",
number = "2",

}

TY - JOUR

T1 - Using visual statistical inference to better understand random class separations in high dimension, low sample size data

AU - Roy Chowdhury, Niladri

AU - Cook, Dianne

AU - Hofmann, Heike

AU - Majumder, Mahbubul

AU - Lee, Eun Kyung

AU - Toth, Amy L.

PY - 2014/11/2

Y1 - 2014/11/2

N2 - Statistical graphics play an important role in exploratory data analysis, model checking and diagnosis. With high dimensional data, this often means plotting low-dimensional projections, for example, in classification tasks projection pursuit is used to find low-dimensional projections that reveal differences between labelled groups. In many contemporary data sets the number of observations is relatively small compared to the number of variables, which is known as a high dimension low sample size (HDLSS) problem. This paper explores the use of visual inference on understanding low-dimensional pictures of HDLSS data. Visual inference helps to quantify the significance of findings made from graphics. This approach may be helpful to broaden the understanding of issues related to HDLSS data in the data analysis community. Methods are illustrated using data from a published paper, which erroneously found real separation in microarray data, and with a simulation study conducted using Amazon’s Mechanical Turk.

AB - Statistical graphics play an important role in exploratory data analysis, model checking and diagnosis. With high dimensional data, this often means plotting low-dimensional projections, for example, in classification tasks projection pursuit is used to find low-dimensional projections that reveal differences between labelled groups. In many contemporary data sets the number of observations is relatively small compared to the number of variables, which is known as a high dimension low sample size (HDLSS) problem. This paper explores the use of visual inference on understanding low-dimensional pictures of HDLSS data. Visual inference helps to quantify the significance of findings made from graphics. This approach may be helpful to broaden the understanding of issues related to HDLSS data in the data analysis community. Methods are illustrated using data from a published paper, which erroneously found real separation in microarray data, and with a simulation study conducted using Amazon’s Mechanical Turk.

KW - Data mining

KW - Lineup

KW - Projection pursuit

KW - Statistical graphics

KW - Visualization

UR - http://www.scopus.com/inward/record.url?scp=84930822749&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84930822749&partnerID=8YFLogxK

U2 - 10.1007/s00180-014-0534-x

DO - 10.1007/s00180-014-0534-x

M3 - Article

VL - 30

SP - 293

EP - 316

JO - Computational Statistics

JF - Computational Statistics

SN - 0943-4062

IS - 2

ER -