Stats Without Walls User Guide

Open the app

Stats Without Walls is a free statistical tool for all levels. It runs entirely in your browser: nothing you load is uploaded anywhere, there is no account, and it never costs anything. This guide explains every menu, what each option does, and how to read what comes back.

1. Start here

Three ways to run it

  1. Web. Open learnwithoutwalls.com/stats in any modern browser (Chrome, Edge, Safari, Firefox). This version always has the newest features.
  2. Install from the browser. In Chrome or Edge, click the install icon in the address bar (or Chrome menu, then "Install Stats Without Walls"). It opens in its own window and works offline after the first visit.
  3. Desktop app. For now the browser version is the better choice: the app is updated almost daily and the desktop app cannot update itself yet (a self-updating version is coming). If you need it, download the Mac or Windows installer from the download page. Mac: pick Apple Silicon (M1 to M4) or Intel. The first time you open it on a Mac, right-click the app and choose Open, because it is not signed with an Apple developer certificate. If your Mac says the app is damaged and offers no Open Anyway button, see Troubleshooting for the one-line fix.

Privacy

Your data stays on your computer. The app has no server. If you close the tab, the last dataset and your results are remembered by your browser (local storage) so you can pick up where you left off; clearing your browser data removes them. Use Results, Save session to keep a permanent copy.

A two-minute first run

  1. Data, Sample datasets, pick "Our class data (40 students)", click Open.
  2. Exploration, Descriptives, select GPA, click Run. A card appears in the Analyses tab.
  3. Graph, Histogram, choose Hours_Sleep, click Run. The graph appears in the Graphs tab.
  4. T-Tests, Independent Samples T-Test, dependent variable Hours_Sleep, grouping variable Employed, click Run. Read the blue interpretation box at the bottom of the card.

2. The screen

PartWhat it is
HeaderTitle, a link to install or download the app, a link back to learnwithoutwalls.com.
Menu barData, Exploration, Graph, T-Tests, ANOVA, Regression, Frequencies, distrACTION, Nonparametric, Learn, Advanced, Results. Click a menu, then an item, and a dialog opens.
Data table (left)Your dataset. The name, row count and column count are above it. Under each column name is its type; click it to change the type. Click a column name to sort; click again to reverse. Click any cell to edit it. Rows greyed out are excluded by the current filter. Drag the bar between the table and the results to resize the two panels; double-click it to reset.
Filter boxType a condition such as Year == "Freshman" and press Enter. Every analysis then uses only the kept rows. Blank removes it.
Results (right)Two tabs. Analyses holds numeric results; Graphs holds every graph made from the Graph menu (the StatCrunch habit). The number on each tab is how many cards it holds.
DialogEvery analysis opens the same kind of box: pick variables, tick options, click Run. Cancel closes it. Where a list allows several variables, hold Cmd (Mac) or Ctrl (Windows) while clicking.

Variable types

The app guesses a type for each column when data loads. The type decides which menus can use the column.

TypeShown asMeaningUsed by
Continuousruler, continuousNumbers you can average: age, GPA, minutes.Descriptives, histograms, t-tests, ANOVA, regression, and so on.
OrdinalordinalOrdered categories: Freshman to Senior, agreement scales, small numeric codes.Frequency tables, bar plots, chi-square, grouping variables. Bar plots keep the natural order.
NominalnominalNames with no order: major, ZIP code, yes or no.Frequency tables, bar and pie charts, chi-square, grouping variables.
IDIDA label for each row (student number). Never analysed.Point labels on scatter plots.

If a guess is wrong (a numeric code that is really a category, or a numeric column that was read as text because one cell contains a word), click the type under the column name and pick Continuous, Ordinal, Nominal or ID, the same job as a data setup panel in other tools. A type you set by hand stays set, even after you edit cells. Data, Set variable type does the same thing from the menu.

3. Getting data in

Open CSV file

Data, Open CSV file. Any comma-separated file with one row of column names at the top. Exports from Excel, Google Sheets, Canvas, Qualtrics and survey tools all work. Quoted text with commas inside is handled. Dates stay as text.

Paste data

Data, Paste data. Copy a block from a spreadsheet (including the header row), paste it in, give it a name. Tabs and commas both work.

Sample datasets

Data, Sample datasets. Real and course datasets that ship with the app, listed in section 13. They also work offline in the installed app.

New data table

Data, New data table. Type column names separated by commas and how many rows you want. Then click cells to type values. Good for a small in-class exercise.

Editing cells

Click any cell, type, then click away or press Tab. Types are re-guessed after each edit. Column names can be changed with Data, Rename or delete column.

Download data as CSV

Saves the current table (including computed, recoded and z-score columns) as a CSV file you can open in Excel or load again later.

4. Preparing data

Filters

Data, Filters, or the filter box above the table. Write a condition using column names exactly as written. Rows that fail the condition are greyed out and left out of every analysis and graph until the filter is cleared.

Year == "Freshman"
Commute_Minutes < 30
Year == "Senior" AND GPA >= 3.5
Major == "Biology" OR Major == "Nursing"
Employed != "Full-time"

Operators: == equal, != not equal, < <= > >=, AND, OR. Text values go in quotes. A single = is accepted as equal.

Compute a new variable

Data, Compute. Name the new column and write a formula using existing columns.

Hours_Study * 7                       weekly hours
Commute_Minutes / 60                  minutes to hours
(GPA - mean(GPA)) / sd(GPA)           z-score by hand
log(Income)                            natural log
round(Age / 10) * 10                  decade
Hours_Sleep >= 7                       gives true or false

Functions: log (natural), log10, sqrt, abs, exp, round, floor, ceil, pow(x, k). Whole-column summaries: mean(X), sd(X), median(X), min(X), max(X), sum(X), n(X). A row with a missing value in any column used gets a blank.

Transform (recode)

Data, Transform. Turn values into new values, one rule per line, as old = new. Numeric ranges are written 0 to 29 = short. "Everything else becomes" catches what the rules miss; leave it blank to keep the original value.

Full-time = Employed
Part-time = Employed
Not employed = Not employed
0 to 29 = short
30 to 59 = medium
60 to 300 = long

Add z-score column

Data, Add z-score column. Adds z_Variable with (x minus mean) divided by the sample SD, the same as the by-hand formula.

Bin into classes

Data, Bin into classes. Turns a numeric variable into classes of a width you choose ("60 to under 70"), starting where you say. Each class includes its left end. The new column is categorical, kept in numeric order, ready for a frequency table or bar plot.

Stack columns

Data, Stack columns. When each group sits in its own column, stacking makes one column of values and one column naming where each value came from, the layout two-sample tests, ANOVA and side-by-side graphs need. It makes a new data table, so download the old one first if you need it.

Simulate random data

Data, Simulate random data. Random values from a normal, uniform, discrete uniform, binomial, Poisson, exponential or yes or no distribution, any number of rows and columns (one column per sample). If the row count matches the open table the columns are added to it; otherwise a new table is made. A seed repeats the same numbers, so a whole class can work with identical data.

Random sample of rows

Data, Random sample of rows. A simple random sample from the open table, with or without replacement, with an optional seed. Uses only the rows kept by a filter.

Sorting

Click a column header. Sorting changes the row order for Time Series and for the display; analyses do not depend on it.

5. Reading a result card

Every analysis produces a card with the same layout, top to bottom:

  1. Title and a grey line naming the dataset, the variables and n. Three buttons: edit (reopens the same dialog with your choices filled in; change anything and Run, and the new result replaces this card in place, the StatCrunch habit), copy text (copies the whole card as plain text for a report) and remove.
  2. Tables, named the way the course instructions name them (Descriptives, Independent Samples T-Test, Model Fit Measures, Contingency Tables, and so on), so the instructions transfer as written.
  3. Formula box (grey, monospace): the formula that produced the statistic, with the symbols the course uses: mu and x bar for means, sigma and s for standard deviations, p and p hat for proportions.
  4. Blue box: the decision and a written conclusion in plain English, including the p-value compared with alpha.
  5. Green or pink box: a condition check. Green means the condition for the method is met (for example n at least 30, or at least 10 successes and 10 failures). Pink means it is not, and says what to do instead.
  6. Graphs attached to the analysis (descriptives plots, residual plots), below the text.
Rules that are built in and cannot be switched off: inference about a mean uses t with degrees of freedom, because sigma is almost never known; the Z-Test is there for the textbook problems that do give sigma, and its card says to use t otherwise. Proportions use z only when the success and failure counts are large enough, and the card says so. Every card reports the p-value; "p = 0" is never printed, the smallest shown is "p < 0.0001".

Decimals. Numbers are shown to the number of decimals set under Results, Decimal places shown (3 by default). Internally everything is full precision.

6. Which analysis do I need

QuestionData you haveUse
Describe one numeric variableone continuous columnExploration, Descriptives; Graph, Histogram or Boxplot or Dotplot
Describe one categorical variableone nominal or ordinal columnExploration, Descriptives with Frequency tables; Graph, Bar Plot
Is the mean different from a claimed value?one continuous column and a numberT-Tests, One Sample T-Test
Do two groups have different means?continuous column plus a two-level group columnT-Tests, Independent Samples T-Test (Welch)
Before and after, same peopletwo continuous columns, one per timeT-Tests, Paired Samples T-Test
Do three or more groups have different means?continuous column plus a group columnANOVA, One-Way ANOVA (with Tukey)
Is a proportion different from a claim?one categorical column and a numberFrequencies, 2 Outcomes: Binomial test
Do two groups have different proportions?categorical outcome plus two-level groupFrequencies, Two proportions z test
Do the category counts match expected shares?one categorical columnFrequencies, N Outcomes: chi-square goodness of fit
Are two categorical variables related?two categorical columnsFrequencies, Contingency Tables
Are two numeric variables related? Predict one from the othertwo continuous columnsRegression, Correlation Matrix; Regression, Linear Regression; Graph, Scatter Plot
Predict from several variablescontinuous outcome plus several predictorsAdvanced, Multiple Linear Regression
Predict a yes or no outcometwo-level outcome plus predictorsAdvanced, Logistic Regression
Data are skewed or ranks, small samplesas above but not normalNonparametric menu (Mann-Whitney, Wilcoxon, Kruskal-Wallis)
A probability from a distributionparameters onlydistrACTION menu
Is the mean different from a claim, and sigma is given?a numeric column, or the mean and n, plus sigmaT-Tests, One Sample Z-Test
Is a standard deviation different from a claim? Do two groups differ in spread?a numeric column, or SDs and nAdvanced, Variance tests
Predict y for a new x, with an intervaltwo numeric columnsRegression, Linear Regression, Predict at x
Mean and SD of a discrete random variablex values and their P(x), typed or as two columnsdistrACTION, Custom (discrete x and P(x))
How many people do I need?a margin of error or an effect sizedistrACTION, Sample size; Advanced, Power and Sample Size
Understand what a p-value or a confidence interval isnothing, no data neededLearn, What a p-value is; Learn, What a confidence interval means
Show why a coin does not even out, or why "two of the next ten" is wrongnothing, no data neededLearn, Coin flips and the law of large numbers
Show the Central Limit Theorem, or where sigma / sqrt(n) comes fromnothing, or a column as the populationLearn, The Central Limit Theorem; Learn, Sampling distribution simulator

Exploration

Descriptives

Pick one or more variables; optionally split by a group. Tick the statistics you want: N, missing, mean, median, mode (all modes are listed when there is a tie), sum, standard deviation (sample, n minus 1), variance, population SD, range, minimum, maximum, standard error, IQR, quartiles, fences for outliers, skewness, kurtosis, coefficient of variation, and any other percentiles you list (such as 10, 90). Tick plots: histogram, density, boxplot, dotplot, QQ plot, bar plot for categorical variables. Frequency tables appear for nominal and ordinal variables, with counts, relative frequency, percent, cumulative frequency, cumulative relative frequency and cumulative percent. Classes named by numbers ("60 to under 70") are listed in numeric order.

Quartile method. "Textbook halves" (default) finds Q1 and Q3 as the medians of the lower and upper halves with the overall median excluded, which is what the course teaches by hand. "Type 7" is what R and Excel use by default; the two can differ for small n. Boxplots use the same rule as the descriptives.

Scatterplot

Quick scatterplot with an optional group colour and regression line; the Graph menu version has more options.

Graph

All graphs land in the Graphs tab. Every graph follows the honest-graph rules, and says so under the picture. Every graph dialog has an Appearance section (title, axis labels, color), and the edit button on the card reopens it with your choices filled in.

ItemPickOptionsRead it as
Bar Plota categorical variable, optional groupcounts or percent; order (automatic keeps ordered scales in order, otherwise tallest first); counts above bars; frequency table; horizontalWhich categories are common. With a group, percent within group compares groups of different sizes fairly.
Pie Charta categorical variableshow percent and countParts of one whole only. A bar plot is usually clearer.
Histograma numeric variable, optional groupfrequency, relative frequency or density; bin width and starting point; frequency above each bin; frequency table (with cumulative frequency and cumulative relative frequency); mean and median lines; normal curve overlay with the mean and SD of the dataShape (symmetric, skewed left or right, one peak or more), centre, spread, gaps and outliers. Change the bin width twice before trusting a feature.
Dotplota numeric variable, optional groupmean and median linesOne dot per observation, nothing hidden.
Boxplotone or more numeric variables, optional groupshow all points; horizontalBox from Q1 to Q3, line at the median, whiskers to the last values inside the fences, dots beyond. Skew shows as a long whisker or an off-centre median.
Stem and Leafa numeric variableleaf unitA histogram that keeps every value.
Scatter Plotx and y numeric, optional colour, optional label columnleast-squares lineDirection, form, strength, outliers, in that order.
QQ Plota numeric variable, optional groupPoints near the line mean roughly normal. Curving means skew; tails peeling off mean outliers. Use it to justify t when n is under 30.

T-Tests

Every t-test dialog has a Data switch: "from the data table" or "summary statistics" (type mean, SD and n from a textbook problem). Options: mean difference with confidence interval, effect size (Cohen's d), descriptives, descriptives plots.

One Sample T-Test

Tests whether the population mean equals a test value. Reports t, df, p, mean difference, the confidence interval, and the condition check (n at least 30, or a normal-looking QQ plot when n is smaller). The formula box shows t = (x bar minus mu0) / (s / sqrt(n)).

Independent Samples T-Test

Compares two group means. Welch's version (no equal-variance assumption) is the default; untick it for Student's. Hypothesis can be two-sided, group 1 greater, or group 1 less. If the grouping variable has more than two levels, name the two you want. Summary-statistics mode takes mean, SD and n for each group.

One Sample Z-Test and Two Sample Z-Test

Only for problems that give the population standard deviation sigma. You type sigma (one, or one per group); the card gives z, the p-value, and the interval x bar plus or minus z* sigma / sqrt(n). If only the sample SD s is known, use the t-test: the card says so.

Paired Samples T-Test

Two columns measured on the same subjects (before and after). The test is a one-sample t on the differences; tick "Histogram of differences" to see them.

ANOVA

One-Way ANOVA

Three or more group means. Fisher's (equal variances) and Welch's F are both available; the homogeneity check compares the largest SD with the smallest (ratio under 2 passes), and Levene's test (deviations from the median, the Brown-Forsythe version) is available as a formal check. Tukey post-hoc tests tell which pairs differ, with adjusted p-values. Descriptives table and plot (means with confidence intervals) are optional. Two-way ANOVA and repeated measures are under Advanced.

Regression

Correlation Matrix

Two or more numeric variables. Pearson (linear) and Spearman (rank) correlations, with p-values if requested, and a grid of scatter plots. r near 0 means no linear relationship, not no relationship.

Linear Regression

One dependent variable and one explanatory variable (covariate). Reports the coefficients table (intercept, slope, SE, t, p, confidence interval), R and R squared, the residual plot and QQ plot of residuals, and predictions at one or several x values (comma separated). Each prediction comes with a confidence interval for the mean response and a prediction interval for one new value, and a warning when x is outside the data. Tick "Show confidence and prediction bands" to draw both bands around the line. The slope interpretation sentence is written for you.

Frequencies

2 Outcomes: Binomial test

One proportion against a test value. The exact binomial p-value is always given; tick "Also show the z test" for the course's by-hand method with the success and failure check (n p0 and n (1 minus p0) both at least 10). Summary mode: successes and n. Interval method: standard (Wald, the by-hand formula, default), plus four, Agresti-Coull, Wilson score, or exact Clopper-Pearson. When there are fewer than 10 successes or failures, the card points you to one of the small-sample methods instead of the standard interval.

N Outcomes: chi-square Goodness of fit

Do observed counts across categories match expected proportions (equal by default, or the ones you type)? Shows expected counts and the check that all are at least 5.

Contingency Tables: Independent Samples

Two categorical variables, or a typed table of counts. Chi-square test of independence with expected counts, row, column and total percentages; for a 2 by 2 table also Fisher's exact test, the odds ratio and the relative risk with confidence intervals. Each cell shows whatever you tick under Cells and Percentages, separated by vertical bars. If you untick all of them the table falls back to the observed counts and says so, rather than printing empty cells.

Two proportions: z test

Compares a success proportion between two groups using the pooled z test taught in the course, with the confidence interval for the difference and the success-failure check in each group. Summary mode takes successes and n per group.

distrACTION

Calculators, no data needed (Custom can also read two columns). Each one draws the distribution and shades the probability.

Nonparametric

Rank-based alternatives for skewed data, ordinal scores or small samples.

Learn

Teaching tools that show where the theory comes from. The first four run with nothing loaded, so they work on a projector in a room where no one has a file open; the rest can use a column of your data as the population.

Advanced

For second courses, research methods and upper-division work. Each card still ends with a plain-English reading.

ItemWhat it doesKey options and reading
Multiple Linear RegressionOne numeric outcome, several numeric and categorical predictors (categories are dummy coded; the most frequent level is the reference).Squared terms, one interaction, residual and QQ plots, VIF for collinearity (above 5 is a concern), influence diagnostics (leverage, standardized residuals, Cook's distance), nested F test against a reduced model, backward elimination demo. Read each slope as the change in the outcome per unit of that predictor, holding the others fixed.
Logistic RegressionA two-level outcome. You name the level modelled as 1.Coefficients, odds ratios with confidence intervals, model chi-square, McFadden R squared, classification accuracy, predicted-probability plot. An odds ratio of 1.5 means the odds are 50 percent higher per unit.
Two-Way ANOVATwo factors, with or without their interaction (Type II sums of squares).Main effects, interaction, cell means table and interaction plot. Non-parallel lines in the plot are the interaction.
Repeated Measures and Mixed ANOVAWide data: one column per condition, optional between-subjects factor.Mauchly's sphericity test, Greenhouse-Geisser and Huynh-Feldt corrections, Holm-adjusted pairwise paired t-tests, ICC, profile plot.
Count RegressionPoisson and negative binomial for counts (0, 1, 2, ...).Rate ratios with confidence intervals and a dispersion check; when the dispersion is well above 1, the negative binomial is recommended.
Multinomial Logistic RegressionThree or more unordered outcome levels against a baseline.One set of coefficients per non-baseline level, relative risk ratios.
Ordinal Logistic RegressionAn ordered outcome (proportional odds). You type the levels from lowest to highest.One set of odds ratios that apply across every cut point.
Variance testsOne SD against a claimed value (chi-square), or two SDs compared (F). Data or summary statistics.Test and confidence intervals for sigma and sigma squared, or for the ratio of variances. Both tests need normal populations at any sample size; to check equal spread before ANOVA, Levene's test is safer.
McNemar testPaired yes or no (before and after on the same people).Uses only the discordant pairs; summary mode takes the two discordant counts.
Cochran-Armitage trend testDoes a proportion rise or fall across ordered groups?Trend chi-square and a plot of the proportions.
Power and Sample Sizet-tests, proportions, ANOVA, correlation, chi-square.Find power for a given n, or the n for a target power; power curve and effect-size sensitivity plot. Effect sizes should come from prior studies, not the data you are about to collect.
Bootstrap confidence intervalMean, median, SD, proportion, difference in means or medians, correlation, slope.Resamples, seed (same seed, same answer), percentile interval; a basic interval is added when bias is noticeable.
Permutation testTwo groups, paired, or correlation.Exact logic of a p-value with no normality assumption.
Time SeriesA numeric series in row order.Time plot with moving average, ACF and PACF, additive or multiplicative decomposition, lag-1 plot, linear trend with Durbin-Watson, Holt forecast with the seasonal pattern added back.
Principal Component AnalysisThree or more numeric variables.Eigenvalues, scree plot, loadings (varimax optional), score plot coloured by a group.
Exploratory Factor AnalysisItems of a questionnaire.Principal axis factoring with varimax; loadings below a cutoff hidden; communalities.
Reliability (Cronbach's alpha)Items of one scale.Alpha, alpha if item dropped, item-rest correlations; reverse-keyed items flipped for you. Alpha at least 0.7 is the usual bar.
k-means ClusteringTwo or more numeric variables.Standardized by default, k-means++ starts, elbow plot, cluster plot on the first two principal components, optional cluster column added to the data.
Survival AnalysisTime to event with censoring.Kaplan-Meier curves with confidence bands and censoring ticks, median survival, risk table, log-rank test between groups, Cox proportional hazards model with hazard ratios and a forest plot.
Bayesian InferenceA proportion (beta prior) or a mean (normal prior), from a column or typed numbers.Prior, likelihood and posterior on one plot; posterior mean, SD and credible interval; probability the parameter is above a value you choose.

Results

8. Honest graphs

Statistics software will happily draw a misleading picture. Stats Without Walls refuses to, and prints the rules it followed under each graph so students learn them.

9. Writing up a result

Use the "copy text" button, then write one paragraph with these parts: what was compared, the statistic with its degrees of freedom, the p-value, the decision at alpha, and what it means in context. The blue box gives you most of this sentence.

Employed students slept less than students who were not employed (Welch t(35.2) = 2.41, p = 0.021, mean difference 0.62 hours, 95% CI 0.10 to 1.14). At alpha = 0.05 we reject H0; employment is associated with less sleep in this class.
There was no evidence of an association between year and commute method (chi-square(6) = 5.12, p = 0.53). Fail to reject H0.
GPA was moderately related to study hours (r = 0.42, p = 0.007). Each extra study hour was associated with 0.03 higher GPA (slope 0.031, 95% CI 0.009 to 0.053), and study hours explained 18 percent of the variation in GPA.

Always report the condition check. If it failed, say so and report the nonparametric or exact result instead.

10. Saving and sharing work

11. How it differs from other tools

Menus, option names and table names match the course instructions, so they work as written. Where common statistics software and the course textbook disagree, Stats Without Walls follows the textbook:

TopicOther toolsStats Without Walls
QuartilesType 7 interpolationTextbook halves by default (median excluded); Type 7 available
Mode with tiesShows one valueLists every mode, or "no mode" when all values are unique
Standard deviationSample (n minus 1) onlySample by default, population SD available and labelled
One proportionExact binomial onlyExact binomial plus the z test with the success-failure check
Two proportionsOften missing from the base menusPooled z test with confidence interval
Meanst, or z with a known sigmat by default; the Z-Test needs sigma typed in and its card says to use t when only s is known
One proportion intervalOne methodStandard (Wald) by default; plus four, Agresti-Coull, Wilson and exact on request
Histogram binsAutomatic, may differ by groupEqual nice bins shared across groups, editable width and start
Summary statistics inputOften not availableEvery t-test, proportion test and chi-square accepts typed summaries
GraphsMixed with outputOwn Graphs tab
PriceFree to paid subscriptionsFree, always

12. Troubleshooting

SymptomCause and fix
A numeric column is not offered in a t-test or histogramOne cell contains text (such as "NA" or "n/a"), so the column was typed as nominal. Clear the cell, or use Data, Set variable type, Continuous.
"Column not found"A filter or formula uses a name that does not match exactly (case and underscores matter). Copy the header text.
Only one variable gets selected in a multi-select listHold Cmd (Mac) or Ctrl (Windows) while clicking.
Every row is grey and analyses say "no rows"The filter excludes everything. Clear the filter box and press Enter.
"Levels of X: ..." errorYou typed a level that is not in the data. The message lists the levels; copy one exactly.
Pink "small group" warningA group has fewer than 5 rows. The result is fragile; say so in the write-up or combine groups.
Results look stale after an updateThe browser cached an old version. Reload with Shift held (or Cmd Shift R). The browser-installed version updates itself on the next launch with internet; the desktop app does not, so download the new version from the download page when one is announced.
Mac says the app "cannot be opened" or "is damaged"The app is not Apple-signed, so macOS quarantines it. Right-click the app in Applications and choose Open, once. If there is no Open Anyway button (newer macOS), open Terminal and run xattr -d com.apple.quarantine "/Applications/Stats Without Walls.app", then open the app normally. Also check you downloaded the right build: Apple Silicon for M1 to M4, Intel for older Macs.
Windows SmartScreen warningClick More info, then Run anyway. The installer is unsigned but the same file as on the download page.
Nothing printsUse Results, Print or save as PDF, not the browser's own print button, so the data table is hidden.

13. Sample datasets

DatasetWhat it isGood for
Our class data (40 students)Survey of a statistics class: major, year, age, commute, study and sleep hours, coffee, exercise, employment, satisfaction, GPA.Everything in a first course.
Class data, exam 1 versionThe same class with exam scores added.Paired and regression questions.
Sleep follow-upSleep before and after an intervention.Paired t-test.
Finch beaksBeak depth in Galapagos finches before and after a drought.Two-sample t, histograms.
Global healthCountries: life expectancy, income, health spending, region.Correlation, regression, ANOVA by region.
Police calls for service, academy fitness, community surveyLaw-enforcement course datasets.Proportions, chi-square, one-way ANOVA, time patterns.
CDC teen survey 2023 (1500 students)Youth Risk Behavior Survey extract: marijuana use, sleep, grades, mood.Contingency tables, two proportions, logistic regression.
General Social Survey 2018 and 2022Astrology and science beliefs, politics, wellbeing of US adults.Chi-square, ordinal outcomes, bar plots of ordered scales.
Big Five personality (1200) and Big Five 50 items (500)Personality scores and raw questionnaire items.Correlation, PCA, factor analysis, reliability.
Cadet mile timesTimes at weeks 1, 4, 8, 12 for two training groups.Repeated measures and mixed ANOVA.
Mauna Loa CO2 monthlyNOAA monthly CO2 since 2000.Time series.
Lung cancer survival (227 patients)Survival time, death or censoring, age, sex, performance score.Survival analysis.

A codebook for the real-world datasets is in the app's data folder (README_real_world_datasets).

14. How the numbers were checked

Every procedure was run on the same data in R (and the jmv package where it exists) and compared to four decimals: t-tests, proportion tests, ANOVA and Tukey, chi-square, regression coefficients and intervals, nonparametric tests, multiple and logistic regression with diagnostics, two-way and repeated measures ANOVA, count, multinomial and ordinal models, power calculations, time-series decomposition and autocorrelation, PCA, factor analysis and varimax rotation, Cronbach's alpha, k-means, Kaplan-Meier, log-rank, Cox regression and the Bayesian conjugate updates. Distributions come from jStat; graphs are drawn with Plotly.

The distribution calculators (binomial, custom discrete, Poisson, geometric, hypergeometric, discrete uniform, uniform, exponential) were checked against R's d, p and q functions; the z tests, the five proportion intervals (against prop.test and binom.test), the variance tests (against var.test), Levene's test (against car::leveneTest), kurtosis, percentiles and the regression confidence and prediction intervals (against predict) the same way. Simulated data were checked by their mean and SD against the distribution they came from.

The Learn demonstrations are checked the same way. The simulated sampling distributions are compared against sigma / sqrt(n); the confidence interval demo is checked by its capture rate against the stated level, which also falls below that level on purpose when the success-failure condition fails; and the p-value demo is checked against the exact binomial, computed independently. Where the z formula and the exact answer disagree, the app prints both and says why rather than hiding the gap.

Stats Without Walls was built by Safaa Dabagh for students at Santa Monica College, West Los Angeles College and Loyola Marymount University, and is free for anyone to use.