Skip to main content
Social Sci LibreTexts

12: Quantitative Analysis Descriptive Statistics

  • Page ID
    124612
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)
    Learning Objectives
    • Describe data coding, data entry, and handling missing data.

    • Define a frequency distribution.

    • Define the three measures of central tendency.

    • Discuss dispersion.

    • Differentiate between univariate and bivariate data.

    Introduction

    Numeric data collected in a research project can be analyzed quantitatively using statistical tools in two different ways. Descriptive analysis refers to statistically describing, aggregating, and presenting the constructs of interest or associations between these constructs. Inferential analysis refers to the statistical testing of hypotheses (theory testing). In this chapter, we will examine statistical techniques used for descriptive analysis, and the next chapter will examine statistical techniques for inferential analysis. Much of today’s quantitative data analysis is conducted using robust software programs such as SPSS, SAS, or increasingly popular open-source programming languages like R and Python. Readers are advised to familiarize themselves with at least one of these tools for understanding and applying the concepts described in this chapter.

    Data Preparation

    In research projects, data may be collected from a variety of sources: mail-in surveys, digital questionnaires, interviews, pretest or posttest experimental data, observational data, and so forth. This data must be converted into a machine-readable, numeric format, such as in a spreadsheet or a text file, so that it can be analyzed by statistical software. Data preparation usually follows these steps:

    Data Coding. Coding is the process of converting data into numeric format. A codebook should be created to guide the coding process. A codebook is a comprehensive document containing a detailed description of each variable in a research study, items or measures for that variable, the format of each item (numeric, text, etc.), the response scale for each item (i.e., whether it is measured on a nominal, ordinal, interval, or ratio scale), and how to code each value into a numeric format. For instance, if we have a measurement item on a seven-point Likert scale with anchors ranging from “strongly disagree” to “strongly agree,” we may code that item as 1 for strongly disagree, 4 for neutral, and 7 for strongly agree. Nominal data such as industry type can be coded in numeric form (e.g., 1 for manufacturing, 2 for retailing), though nominal data cannot be analyzed as a continuous statistic. Coding is especially important for large complex studies involving many variables to help the coding team maintain consistency. Today, digital survey platforms (like  Google Forms) often automatically code data based on pre-set parameters, drastically reducing manual coding errors.

    Data Entry. Coded data can be entered into a spreadsheet, database, text file, or directly into a statistical program. Most statistical programs provide a data editor for entering data. However, these programs store data in their own native format (e.g., SPSS stores data as .sav files), which makes it difficult to share that data with other statistical programs. Hence, it is often better to enter data into a universally readable spreadsheet (like a .csv file) or database, where they can be reorganized as needed and shared across programs. The entered data should be frequently checked for accuracy via occasional spot checks. Furthermore, the coder should watch out for obvious evidence of bad data, such as a respondent selecting “strongly agree” for all items without reading them. Such "straight-lining" data should be excluded from subsequent analysis.

    Missing Values. Missing data is an inevitable part of any empirical data set. Respondents may not answer certain questions if they are ambiguously worded or too sensitive. During data analysis, the default mode of handling missing values in most software programs is to simply drop the entire observation containing even a single missing value, a technique called "listwise deletion." Because such deletion can significantly shrink the sample size and make it extremely difficult to detect small effects, researchers often use imputation to replace missing values with an estimated value. While simple mean substitution (averaging the remaining responses) was once popular, modern researchers typically rely on more robust, unbiased estimates like maximum likelihood procedures, k-Nearest Neighbors (kNN) imputation, or multiple imputation methods, which are supported in modern software programs like R, Python, and SPSS.

    Univariate Analysis

    Univariate analysis, or analysis of a single variable, refers to a set of statistical techniques that can describe the general properties of one variable. Univariate statistics include: (1) frequency distribution, (2) central tendency, and (3) dispersion.

    The frequency distribution of a variable is a summary of the frequency (or percentages) of individual values or ranges of values for that variable. For instance, we can measure how many times a sample of respondents attend religious services using a categorical scale. If we count the percentage of observations within each category and display it as a bar chart, we have a visual frequency distribution .

    With very large samples where observations are independent and random, the frequency distribution tends to follow a plot that looks like a bell-shaped curve, where most observations cluster toward the center of the range, and fewer observations fall toward the extreme ends. Such a curve is called a normal distribution. See the image below.

     Gauss distribution. Standard normal distribution. Gaussian bell graph curve. Business and marketing concept. Math probability theory. Editable stroke. Vector illustration isolated on white background. Source: Getty Images

    Central tendency is an estimate of the center of a distribution of values. There are three major estimates of central tendency: mean, median, and mode.

    • The arithmetic mean (often simply called the “mean”) is the simple average of all values in a given distribution.

    • The median is the middle value within a sorted range of values in a distribution.

    • The mode is the most frequently occurring value in a distribution of values. Note that any value that is estimated from a sample is called a statistic.

    Dispersion refers to the way values are spread around the central tendency, for example, how tightly or how widely the values cluster around the mean. Two common measures of dispersion are the range (the difference between the highest and lowest values) and the standard deviation. Because the range is particularly sensitive to extreme outliers, researchers rely heavily on standard deviation, which corrects for outliers by taking into account how close or far each value is from the distribution mean.

    Bivariate Analysis

    Bivariate analysis examines how two variables are related to each other. The most common bivariate statistic is the bivariate correlation (often simply called “correlation” or Pearson's r), which is a number between -1 and +1 denoting the strength of the relationship between two variables.

    Let’s say that we wish to study how age is related to self-esteem in a sample of 20 respondents that is, as age increases, does self-esteem increase, decrease, or remain unchanged?

    • If self-esteem increases as age increases, we have a positive correlation.

    • If self-esteem decreases as age increases, we have a negative correlation.

    • If age has no systematic bearing on self-esteem, we have a zero correlation.

    To calculate the value of this correlation, consider a hypothetical dataset where Age is a ratio-scale variable and Self-Esteem is an average score computed from a multi-item scale measured using a 7-point Likert scale (ranging from “strongly disagree” to “strongly agree”).

    When plotted on a scatter plot (with self-esteem on the vertical y-axis and age on the horizontal x-axis), the visual shape of the data points reveals the relationship. A plot that roughly resembles an upward-sloping line indicates a positive correlation. If the two variables were negatively correlated, the scatter plot would slope downward. If the two variables were uncorrelated, the scatter plot would approximate a horizontal line, implying no systematic relationship.

    Hypothesis Testing for Correlation

    After computing the bivariate correlation, researchers are often interested in knowing whether the correlation is statistically significant (i.e., a real relationship) or caused by mere chance. Answering such a question requires testing the following hypotheses:

    • H0: r = 0 (Null Hypothesis)

    • H1: r not equal 0 (Alternative Hypothesis)

    H0 is called the null hypothesis, and H1 is called the alternative hypothesis (sometimes represented as Ha). Although they may seem like two separate hypotheses, they actually represent a single framework since they are direct opposites. We are ultimately interested in testing H1. Note that this specific H1 is a non-directional hypothesis since it does not specify whether r is greater than or less than zero. A directional hypothesis would be specified as H0: r equal to or less than 0 and H1: r > 0 (if we were specifically testing for a positive correlation). Significance testing of a directional hypothesis is done using a one-tailed t-test, while a non-directional hypothesis utilizes a two-tailed t-test.

    In statistical testing, the alternative hypothesis cannot be tested directly. Rather, it is tested indirectly by rejecting the null hypothesis with a certain level of probability. Statistical testing is always probabilistic because we can never be absolutely sure if our inferences based on sample data apply perfectly to the entire population.

    • p-value: The probability that a statistical inference is caused by pure chance.

    • Significance level (alpha): The maximum level of risk we are willing to take that our inference is incorrect (typically set to 0.05).

    A p-value less than alpha = 0.05 indicates that we have enough statistical evidence to reject the null hypothesis, thereby indirectly accepting the alternative hypothesis. If p > 0.05, we do not have adequate statistical evidence to reject the null hypothesis.

    The easiest way to test this is to look up the critical value of r from statistical tables (though most software programs perform this significance testing automatically). The critical value depends on the desired significance level (alpha = 0.05), the degrees of freedom (df), and whether the test is one-tailed or two-tailed. The degrees of freedom equal n - 2. For our sample of 20, df = 20 - 2 = 18.

    In a standard two-tailed table, the critical value of r for alpha = 0.05 and df = 18 is 0.44. For a computed correlation to be significant, its absolute value must be larger than the critical value (greater than 0.44 or less than -0.44). If our computed correlation was 0.79, we would conclude that there is a significant positive correlation between age and self-esteem in our dataset. The odds are less than 5% that this correlation is a chance occurrence, allowing us to reject the null hypothesis.

    Correlation Matrices

    Most research studies involve more than two variables. If there are n variables, we will have a total of n(n-1)/2 possible correlations between them. Such correlations are easily computed using a software program like SPSS and represented using a correlation matrix.

    A correlation matrix lists the variable names along the first row and the first column, depicting bivariate correlations between pairs in the intersecting cells. The values along the principal diagonal (from the top left to the bottom right) are always 1.00 because a variable is always perfectly correlated with itself. Because correlations are non-directional, the correlation between V1 and V2 is identical to V2 and V1. Therefore, the lower triangular matrix is a mirror reflection of the upper triangular matrix, and researchers often only display the lower half for simplicity. When these correlations involve interval or ratio scale variables, they are specifically called Pearson product-moment correlations.

    Cross-Tabulation and Chi-Square

    Another useful way of presenting bivariate data is cross-tabulation (often abbreviated to cross-tab, and formally known as a contingency table). A cross-tab describes the frequency or percentage of all combinations of two or more nominal or categorical variables.

    As an example, let us assume we observed the Gender (nominal: Male/Female) and Letter Grade (categorical: A, B, C) of 20 students. A simple 2x3 cross-tabulation matrix displays the joint distribution to help us see if grades are equally distributed across gender:

    Table 13.3: Example of Cross-Tab Analysis

    Gender Grade A Grade B Grade C Marginal Row Total
    Male 1 6 3 10
    Female 4 5 1 10
    Marginal Col Total 5 11 4 Total = 20

    Note: The last row and the last column are called "marginal totals" because they indicate the totals across each category displayed along the margins of the table.

    Although we can see a distinct pattern in the distribution-female students lean toward A grades, while male students lean toward C grades—is this pattern statistically significant? Do these frequency counts differ from what might be expected by pure chance?

    To answer this, we compute the expected count for each cell by multiplying the marginal column total by the marginal row total and dividing by the grand total. For example, for the Male/A Grade cell:

    • Expected count = (5 \times 10) / 20 = 2.5

    We expected 2.5 male students to receive an A grade by pure chance, but in reality, only 1 male student received an A. We test whether the overall difference between expected and actual counts across all cells is significant using a chi-square (chi^2) test.

    The chi-square statistic computes the average difference between observed and expected counts across the entire matrix. We then compare this computed number to a critical value based on our desired probability level (p < 0.05) and the degrees of freedom. For a cross-tab, $f = (rows - 1) \times (columns - 1).

    • In this 2x3 table: df = (2 - 1) \times (3 - 1) = 2.

    According to standard chi-square tables, the critical value for p = 0.05 and df = 2 is 5.99. If our computed chi-square value based on the observed data is 1.00 (which is less than 5.99), we fail to reject the null hypothesis. We must conclude that the observed grade pattern is not statistically different from a pattern that could occur by pure chance.

    Key Takeaways
    • Data Handling: Understanding how to accurately code, enter, and handle missing data (via techniques like multiple imputation) is critical before statistical analysis begins.

    • Univariate Analysis: Methods like frequency distributions, measures of central tendency (mean, median, mode), and dispersion (standard deviation) are useful for describing the properties of a single variable.

    • Bivariate Analysis: Techniques like correlation and cross-tabulation compare two variables and are essential for exploring relationships, significance, and potential causation.

    Contributors and Attributions

    CC licensed content, Shared previously

    This page titled 12: Quantitative Analysis Descriptive Statistics was last modified on Thu, 11 Jun 2026 22:34:33 GMT and is shared under a CC BY license and was authored, remixed, and/or curated by William Pelz (Lumen Learning) .

    • Was this article helpful?