Stats Without Walls is a free statistical tool for all levels. It runs entirely in your browser: nothing you load is uploaded anywhere, there is no account, and it never costs anything. This guide explains every menu, what each option does, and how to read what comes back.
1. Start here
Three ways to run it
- Web. Open learnwithoutwalls.com/stats in any modern browser (Chrome, Edge, Safari, Firefox). This version always has the newest features.
- Install from the browser. In Chrome or Edge, click the install icon in the address bar (or Chrome menu, then "Install Stats Without Walls"). It opens in its own window and works offline after the first visit.
- Desktop app. For now the browser version is the better choice: the app is updated almost daily and the desktop app cannot update itself yet (a self-updating version is coming). If you need it, download the Mac or Windows installer from the download page. Mac: pick Apple Silicon (M1 to M4) or Intel. The first time you open it on a Mac, right-click the app and choose Open, because it is not signed with an Apple developer certificate. If your Mac says the app is damaged and offers no Open Anyway button, see Troubleshooting for the one-line fix.
Privacy
Your data stays on your computer. The app has no server. If you close the tab, the last dataset and your results are remembered by your browser (local storage) so you can pick up where you left off; clearing your browser data removes them. Use Results, Save session to keep a permanent copy.
A two-minute first run
- Data, Sample datasets, pick "Our class data (40 students)", click Open.
- Exploration, Descriptives, select GPA, click Run. A card appears in the Analyses tab.
- Graph, Histogram, choose Hours_Sleep, click Run. The graph appears in the Graphs tab.
- T-Tests, Independent Samples T-Test, dependent variable Hours_Sleep, grouping variable Employed, click Run. Read the blue interpretation box at the bottom of the card.
2. The screen
| Part | What it is |
|---|---|
| Header | Title, a link to install or download the app, a link back to learnwithoutwalls.com. |
| Menu bar | Data, Exploration, Graph, T-Tests, ANOVA, Regression, Frequencies, distrACTION, Nonparametric, Learn, Advanced, Results. Click a menu, then an item, and a dialog opens. |
| Data table (left) | Your dataset. The name, row count and column count are above it. Under each column name is its type; click it to change the type. Click a column name to sort; click again to reverse. Click any cell to edit it. Rows greyed out are excluded by the current filter. Drag the bar between the table and the results to resize the two panels; double-click it to reset. |
| Filter box | Type a condition such as Year == "Freshman" and press Enter. Every analysis then uses only the kept rows. Blank removes it. |
| Results (right) | Two tabs. Analyses holds numeric results; Graphs holds every graph made from the Graph menu (the StatCrunch habit). The number on each tab is how many cards it holds. |
| Dialog | Every analysis opens the same kind of box: pick variables, tick options, click Run. Cancel closes it. Where a list allows several variables, hold Cmd (Mac) or Ctrl (Windows) while clicking. |
Variable types
The app guesses a type for each column when data loads. The type decides which menus can use the column.
| Type | Shown as | Meaning | Used by |
|---|---|---|---|
| Continuous | ruler, continuous | Numbers you can average: age, GPA, minutes. | Descriptives, histograms, t-tests, ANOVA, regression, and so on. |
| Ordinal | ordinal | Ordered categories: Freshman to Senior, agreement scales, small numeric codes. | Frequency tables, bar plots, chi-square, grouping variables. Bar plots keep the natural order. |
| Nominal | nominal | Names with no order: major, ZIP code, yes or no. | Frequency tables, bar and pie charts, chi-square, grouping variables. |
| ID | ID | A label for each row (student number). Never analysed. | Point labels on scatter plots. |
If a guess is wrong (a numeric code that is really a category, or a numeric column that was read as text because one cell contains a word), click the type under the column name and pick Continuous, Ordinal, Nominal or ID, the same job as a data setup panel in other tools. A type you set by hand stays set, even after you edit cells. Data, Set variable type does the same thing from the menu.
3. Getting data in
Open CSV file
Data, Open CSV file. Any comma-separated file with one row of column names at the top. Exports from Excel, Google Sheets, Canvas, Qualtrics and survey tools all work. Quoted text with commas inside is handled. Dates stay as text.
Paste data
Data, Paste data. Copy a block from a spreadsheet (including the header row), paste it in, give it a name. Tabs and commas both work.
Sample datasets
Data, Sample datasets. Real and course datasets that ship with the app, listed in section 13. They also work offline in the installed app.
New data table
Data, New data table. Type column names separated by commas and how many rows you want. Then click cells to type values. Good for a small in-class exercise.
Editing cells
Click any cell, type, then click away or press Tab. Types are re-guessed after each edit. Column names can be changed with Data, Rename or delete column.
Download data as CSV
Saves the current table (including computed, recoded and z-score columns) as a CSV file you can open in Excel or load again later.
4. Preparing data
Filters
Data, Filters, or the filter box above the table. Write a condition using column names exactly as written. Rows that fail the condition are greyed out and left out of every analysis and graph until the filter is cleared.
Year == "Freshman" Commute_Minutes < 30 Year == "Senior" AND GPA >= 3.5 Major == "Biology" OR Major == "Nursing" Employed != "Full-time"
Operators: == equal, != not equal, < <= > >=, AND, OR. Text values go in quotes. A single = is accepted as equal.
Compute a new variable
Data, Compute. Name the new column and write a formula using existing columns.
Hours_Study * 7 weekly hours Commute_Minutes / 60 minutes to hours (GPA - mean(GPA)) / sd(GPA) z-score by hand log(Income) natural log round(Age / 10) * 10 decade Hours_Sleep >= 7 gives true or false
Functions: log (natural), log10, sqrt, abs, exp, round, floor, ceil, pow(x, k). Whole-column summaries: mean(X), sd(X), median(X), min(X), max(X), sum(X), n(X). A row with a missing value in any column used gets a blank.
Transform (recode)
Data, Transform. Turn values into new values, one rule per line, as old = new. Numeric ranges are written 0 to 29 = short. "Everything else becomes" catches what the rules miss; leave it blank to keep the original value.
Full-time = Employed Part-time = Employed Not employed = Not employed
0 to 29 = short 30 to 59 = medium 60 to 300 = long
Add z-score column
Data, Add z-score column. Adds z_Variable with (x minus mean) divided by the sample SD, the same as the by-hand formula.
Bin into classes
Data, Bin into classes. Turns a numeric variable into classes of a width you choose ("60 to under 70"), starting where you say. Each class includes its left end. The new column is categorical, kept in numeric order, ready for a frequency table or bar plot.
Stack columns
Data, Stack columns. When each group sits in its own column, stacking makes one column of values and one column naming where each value came from, the layout two-sample tests, ANOVA and side-by-side graphs need. It makes a new data table, so download the old one first if you need it.
Simulate random data
Data, Simulate random data. Random values from a normal, uniform, discrete uniform, binomial, Poisson, exponential or yes or no distribution, any number of rows and columns (one column per sample). If the row count matches the open table the columns are added to it; otherwise a new table is made. A seed repeats the same numbers, so a whole class can work with identical data.
Random sample of rows
Data, Random sample of rows. A simple random sample from the open table, with or without replacement, with an optional seed. Uses only the rows kept by a filter.
Sorting
Click a column header. Sorting changes the row order for Time Series and for the display; analyses do not depend on it.
5. Reading a result card
Every analysis produces a card with the same layout, top to bottom:
- Title and a grey line naming the dataset, the variables and n. Three buttons: edit (reopens the same dialog with your choices filled in; change anything and Run, and the new result replaces this card in place, the StatCrunch habit), copy text (copies the whole card as plain text for a report) and remove.
- Tables, named the way the course instructions name them (Descriptives, Independent Samples T-Test, Model Fit Measures, Contingency Tables, and so on), so the instructions transfer as written.
- Formula box (grey, monospace): the formula that produced the statistic, with the symbols the course uses: mu and x bar for means, sigma and s for standard deviations, p and p hat for proportions.
- Blue box: the decision and a written conclusion in plain English, including the p-value compared with alpha.
- Green or pink box: a condition check. Green means the condition for the method is met (for example n at least 30, or at least 10 successes and 10 failures). Pink means it is not, and says what to do instead.
- Graphs attached to the analysis (descriptives plots, residual plots), below the text.
Decimals. Numbers are shown to the number of decimals set under Results, Decimal places shown (3 by default). Internally everything is full precision.
6. Which analysis do I need
| Question | Data you have | Use |
|---|---|---|
| Describe one numeric variable | one continuous column | Exploration, Descriptives; Graph, Histogram or Boxplot or Dotplot |
| Describe one categorical variable | one nominal or ordinal column | Exploration, Descriptives with Frequency tables; Graph, Bar Plot |
| Is the mean different from a claimed value? | one continuous column and a number | T-Tests, One Sample T-Test |
| Do two groups have different means? | continuous column plus a two-level group column | T-Tests, Independent Samples T-Test (Welch) |
| Before and after, same people | two continuous columns, one per time | T-Tests, Paired Samples T-Test |
| Do three or more groups have different means? | continuous column plus a group column | ANOVA, One-Way ANOVA (with Tukey) |
| Is a proportion different from a claim? | one categorical column and a number | Frequencies, 2 Outcomes: Binomial test |
| Do two groups have different proportions? | categorical outcome plus two-level group | Frequencies, Two proportions z test |
| Do the category counts match expected shares? | one categorical column | Frequencies, N Outcomes: chi-square goodness of fit |
| Are two categorical variables related? | two categorical columns | Frequencies, Contingency Tables |
| Are two numeric variables related? Predict one from the other | two continuous columns | Regression, Correlation Matrix; Regression, Linear Regression; Graph, Scatter Plot |
| Predict from several variables | continuous outcome plus several predictors | Advanced, Multiple Linear Regression |
| Predict a yes or no outcome | two-level outcome plus predictors | Advanced, Logistic Regression |
| Data are skewed or ranks, small samples | as above but not normal | Nonparametric menu (Mann-Whitney, Wilcoxon, Kruskal-Wallis) |
| A probability from a distribution | parameters only | distrACTION menu |
| Is the mean different from a claim, and sigma is given? | a numeric column, or the mean and n, plus sigma | T-Tests, One Sample Z-Test |
| Is a standard deviation different from a claim? Do two groups differ in spread? | a numeric column, or SDs and n | Advanced, Variance tests |
| Predict y for a new x, with an interval | two numeric columns | Regression, Linear Regression, Predict at x |
| Mean and SD of a discrete random variable | x values and their P(x), typed or as two columns | distrACTION, Custom (discrete x and P(x)) |
| How many people do I need? | a margin of error or an effect size | distrACTION, Sample size; Advanced, Power and Sample Size |
| Understand what a p-value or a confidence interval is | nothing, no data needed | Learn, What a p-value is; Learn, What a confidence interval means |
| Show why a coin does not even out, or why "two of the next ten" is wrong | nothing, no data needed | Learn, Coin flips and the law of large numbers |
| Show the Central Limit Theorem, or where sigma / sqrt(n) comes from | nothing, or a column as the population | Learn, The Central Limit Theorem; Learn, Sampling distribution simulator |
7. Menu by menu
Exploration
Descriptives
Pick one or more variables; optionally split by a group. Tick the statistics you want: N, missing, mean, median, mode (all modes are listed when there is a tie), sum, standard deviation (sample, n minus 1), variance, population SD, range, minimum, maximum, standard error, IQR, quartiles, fences for outliers, skewness, kurtosis, coefficient of variation, and any other percentiles you list (such as 10, 90). Tick plots: histogram, density, boxplot, dotplot, QQ plot, bar plot for categorical variables. Frequency tables appear for nominal and ordinal variables, with counts, relative frequency, percent, cumulative frequency, cumulative relative frequency and cumulative percent. Classes named by numbers ("60 to under 70") are listed in numeric order.
Quartile method. "Textbook halves" (default) finds Q1 and Q3 as the medians of the lower and upper halves with the overall median excluded, which is what the course teaches by hand. "Type 7" is what R and Excel use by default; the two can differ for small n. Boxplots use the same rule as the descriptives.
Scatterplot
Quick scatterplot with an optional group colour and regression line; the Graph menu version has more options.
Graph
All graphs land in the Graphs tab. Every graph follows the honest-graph rules, and says so under the picture. Every graph dialog has an Appearance section (title, axis labels, color), and the edit button on the card reopens it with your choices filled in.
| Item | Pick | Options | Read it as |
|---|---|---|---|
| Bar Plot | a categorical variable, optional group | counts or percent; order (automatic keeps ordered scales in order, otherwise tallest first); counts above bars; frequency table; horizontal | Which categories are common. With a group, percent within group compares groups of different sizes fairly. |
| Pie Chart | a categorical variable | show percent and count | Parts of one whole only. A bar plot is usually clearer. |
| Histogram | a numeric variable, optional group | frequency, relative frequency or density; bin width and starting point; frequency above each bin; frequency table (with cumulative frequency and cumulative relative frequency); mean and median lines; normal curve overlay with the mean and SD of the data | Shape (symmetric, skewed left or right, one peak or more), centre, spread, gaps and outliers. Change the bin width twice before trusting a feature. |
| Dotplot | a numeric variable, optional group | mean and median lines | One dot per observation, nothing hidden. |
| Boxplot | one or more numeric variables, optional group | show all points; horizontal | Box from Q1 to Q3, line at the median, whiskers to the last values inside the fences, dots beyond. Skew shows as a long whisker or an off-centre median. |
| Stem and Leaf | a numeric variable | leaf unit | A histogram that keeps every value. |
| Scatter Plot | x and y numeric, optional colour, optional label column | least-squares line | Direction, form, strength, outliers, in that order. |
| QQ Plot | a numeric variable, optional group | Points near the line mean roughly normal. Curving means skew; tails peeling off mean outliers. Use it to justify t when n is under 30. |
T-Tests
Every t-test dialog has a Data switch: "from the data table" or "summary statistics" (type mean, SD and n from a textbook problem). Options: mean difference with confidence interval, effect size (Cohen's d), descriptives, descriptives plots.
One Sample T-Test
Tests whether the population mean equals a test value. Reports t, df, p, mean difference, the confidence interval, and the condition check (n at least 30, or a normal-looking QQ plot when n is smaller). The formula box shows t = (x bar minus mu0) / (s / sqrt(n)).
Independent Samples T-Test
Compares two group means. Welch's version (no equal-variance assumption) is the default; untick it for Student's. Hypothesis can be two-sided, group 1 greater, or group 1 less. If the grouping variable has more than two levels, name the two you want. Summary-statistics mode takes mean, SD and n for each group.
One Sample Z-Test and Two Sample Z-Test
Only for problems that give the population standard deviation sigma. You type sigma (one, or one per group); the card gives z, the p-value, and the interval x bar plus or minus z* sigma / sqrt(n). If only the sample SD s is known, use the t-test: the card says so.
Paired Samples T-Test
Two columns measured on the same subjects (before and after). The test is a one-sample t on the differences; tick "Histogram of differences" to see them.
ANOVA
One-Way ANOVA
Three or more group means. Fisher's (equal variances) and Welch's F are both available; the homogeneity check compares the largest SD with the smallest (ratio under 2 passes), and Levene's test (deviations from the median, the Brown-Forsythe version) is available as a formal check. Tukey post-hoc tests tell which pairs differ, with adjusted p-values. Descriptives table and plot (means with confidence intervals) are optional. Two-way ANOVA and repeated measures are under Advanced.
Regression
Correlation Matrix
Two or more numeric variables. Pearson (linear) and Spearman (rank) correlations, with p-values if requested, and a grid of scatter plots. r near 0 means no linear relationship, not no relationship.
Linear Regression
One dependent variable and one explanatory variable (covariate). Reports the coefficients table (intercept, slope, SE, t, p, confidence interval), R and R squared, the residual plot and QQ plot of residuals, and predictions at one or several x values (comma separated). Each prediction comes with a confidence interval for the mean response and a prediction interval for one new value, and a warning when x is outside the data. Tick "Show confidence and prediction bands" to draw both bands around the line. The slope interpretation sentence is written for you.
Frequencies
2 Outcomes: Binomial test
One proportion against a test value. The exact binomial p-value is always given; tick "Also show the z test" for the course's by-hand method with the success and failure check (n p0 and n (1 minus p0) both at least 10). Summary mode: successes and n. Interval method: standard (Wald, the by-hand formula, default), plus four, Agresti-Coull, Wilson score, or exact Clopper-Pearson. When there are fewer than 10 successes or failures, the card points you to one of the small-sample methods instead of the standard interval.
N Outcomes: chi-square Goodness of fit
Do observed counts across categories match expected proportions (equal by default, or the ones you type)? Shows expected counts and the check that all are at least 5.
Contingency Tables: Independent Samples
Two categorical variables, or a typed table of counts. Chi-square test of independence with expected counts, row, column and total percentages; for a 2 by 2 table also Fisher's exact test, the odds ratio and the relative risk with confidence intervals. Each cell shows whatever you tick under Cells and Percentages, separated by vertical bars. If you untick all of them the table falls back to the observed counts and says so, rather than printing empty cells.
Two proportions: z test
Compares a success proportion between two groups using the pooled z test taught in the course, with the confidence interval for the difference and the success-failure check in each group. Summary mode takes successes and n per group.
distrACTION
Calculators, no data needed (Custom can also read two columns). Each one draws the distribution and shades the probability.
- Binomial: P(X = x), P(X at most x), P(X at least x), P(x1 to x2), plus mean and SD of the distribution.
- Poisson, Geometric, Hypergeometric, Discrete Uniform: the same layout as Binomial, with P(X = x), less than, at most, more than, at least and between, the mean, variance and SD, a full table and a shaded bar graph. Geometric lets you choose whether X counts the trials up to the first success (1, 2, 3, ...) or the failures before it (0, 1, 2, ...), since textbooks use both.
- Custom (discrete x and P(x)): any discrete probability distribution. Either put x in one column and P(x) in another (columns named x and p(x) are picked automatically), or type both lists separated by commas. Gives the mean (expected value), variance and SD, a work table with x P(x) and (x minus mean)^2 P(x) to match the hand method, and the probability P(X = x), less than, at most, more than, at least, or between. The bar graph shades that probability and marks the mean. It stops with a message if the probabilities are not between 0 and 1 or do not add to 1.
- Normal: probabilities below, above or between values, or the value for a given percentile (quantile). Type z scores by using mean 0 and SD 1.
- Uniform (continuous) and Exponential: probabilities below, above or between values, and percentiles, with mean and SD and the shaded area. Exponential takes the mean; if you are given a rate lambda, the mean is 1/lambda.
- T-Distribution: tail probabilities for a t statistic, or the critical t for a confidence level and df.
- Chi-square and F: right-tail p-value for a statistic, or the critical value for alpha.
- Sample size for a margin of error: for a proportion (with a guess for p, 0.5 if unknown) or a mean (with a guess for s).
- Power and sample size for a test: opens the Advanced power calculator.
Nonparametric
Rank-based alternatives for skewed data, ordinal scores or small samples.
- Mann-Whitney U: two independent groups (alternative to the independent t-test). Reports U, the p-value and the group medians.
- Wilcoxon signed-rank and sign test: paired data, or one column against a constant. Both tests are shown; the sign test uses only the direction of each difference.
- Kruskal-Wallis: three or more groups (alternative to one-way ANOVA). Follow a significant result with Mann-Whitney tests on the pairs that matter.
Learn
Teaching tools that show where the theory comes from. The first four run with nothing loaded, so they work on a projector in a room where no one has a file open; the rest can use a column of your data as the population.
- Coin flips and the law of large numbers: one coin, two coins, three coins, a die, two dice, or any probability you choose. It shows the first fifty outcomes, a table of the proportion after 10, 25, 50, 100, 250 and more trials, and the running proportion settling toward p. A second plot shows the count drifting further from the expected count even while the proportion settles, which is the answer to "19 percent, so 2 of the next 10 will be priority 1" and to "tails is due."
- The Central Limit Theorem: pick a population shape (right-skewed, normal, uniform, two humps, yes or no, or a column of your data) and several sample sizes at once. The sampling distributions are drawn on the same axes, with a table comparing the simulated SD against sigma / sqrt(n) and the skewness falling as n grows. The point students miss: the center never moves, the spread shrinks by sqrt(n), and only the shape is the theorem.
- What a confidence interval means: draws many samples from a population only you can see, builds an interval from each, and plots them with the true value as a fixed vertical line. It reports the actual capture rate against the stated level. The truth never moves; the interval does. Uses t with n - 1 degrees of freedom for a mean, and warns when n p or n (1 - p) falls under 10 for a proportion.
- What a p-value is: forces H0 to be true, draws thousands of fresh samples from that world, and counts how many land at least as far out as your result. For a proportion it prints the simulated p, the exact binomial p and the z formula p side by side, and explains the gap when the normal approximation is off.
- Mean versus median: click the number line to add values, click a dot to remove one. A long tail or one far value drags the mean and barely moves the median.
- Regression: outliers and influence: click to add points to a scatterplot and watch the least-squares line, r and the slope refit against the starting line. A point far out in x and off the pattern is influential; an outlier in the middle of the x range is not.
- Guess the correlation: a fresh scatterplot each time; type your guess for r, check it, and keep a running average of how far off you are.
- Sampling distribution simulator: pick a population (a column of your data, or a yes or no population with a chosen p), a statistic, a sample size and a number of samples. It draws the population, then the distribution of the sample statistic, and compares its SD with the formula (sigma / sqrt(n) or sqrt(p(1 minus p)/n)). Watch it turn normal as n grows.
- Bootstrap: a confidence interval with no formula. Resample with replacement, recompute the statistic thousands of times, take the middle 95 percent.
- Permutation: what a p-value is. Shuffle the group labels, recompute the difference, count how often the shuffled world beats the real one.
- Bayesian: prior, data, posterior on one picture, for a proportion or a mean.
Advanced
For second courses, research methods and upper-division work. Each card still ends with a plain-English reading.
| Item | What it does | Key options and reading |
|---|---|---|
| Multiple Linear Regression | One numeric outcome, several numeric and categorical predictors (categories are dummy coded; the most frequent level is the reference). | Squared terms, one interaction, residual and QQ plots, VIF for collinearity (above 5 is a concern), influence diagnostics (leverage, standardized residuals, Cook's distance), nested F test against a reduced model, backward elimination demo. Read each slope as the change in the outcome per unit of that predictor, holding the others fixed. |
| Logistic Regression | A two-level outcome. You name the level modelled as 1. | Coefficients, odds ratios with confidence intervals, model chi-square, McFadden R squared, classification accuracy, predicted-probability plot. An odds ratio of 1.5 means the odds are 50 percent higher per unit. |
| Two-Way ANOVA | Two factors, with or without their interaction (Type II sums of squares). | Main effects, interaction, cell means table and interaction plot. Non-parallel lines in the plot are the interaction. |
| Repeated Measures and Mixed ANOVA | Wide data: one column per condition, optional between-subjects factor. | Mauchly's sphericity test, Greenhouse-Geisser and Huynh-Feldt corrections, Holm-adjusted pairwise paired t-tests, ICC, profile plot. |
| Count Regression | Poisson and negative binomial for counts (0, 1, 2, ...). | Rate ratios with confidence intervals and a dispersion check; when the dispersion is well above 1, the negative binomial is recommended. |
| Multinomial Logistic Regression | Three or more unordered outcome levels against a baseline. | One set of coefficients per non-baseline level, relative risk ratios. |
| Ordinal Logistic Regression | An ordered outcome (proportional odds). You type the levels from lowest to highest. | One set of odds ratios that apply across every cut point. |
| Variance tests | One SD against a claimed value (chi-square), or two SDs compared (F). Data or summary statistics. | Test and confidence intervals for sigma and sigma squared, or for the ratio of variances. Both tests need normal populations at any sample size; to check equal spread before ANOVA, Levene's test is safer. |
| McNemar test | Paired yes or no (before and after on the same people). | Uses only the discordant pairs; summary mode takes the two discordant counts. |
| Cochran-Armitage trend test | Does a proportion rise or fall across ordered groups? | Trend chi-square and a plot of the proportions. |
| Power and Sample Size | t-tests, proportions, ANOVA, correlation, chi-square. | Find power for a given n, or the n for a target power; power curve and effect-size sensitivity plot. Effect sizes should come from prior studies, not the data you are about to collect. |
| Bootstrap confidence interval | Mean, median, SD, proportion, difference in means or medians, correlation, slope. | Resamples, seed (same seed, same answer), percentile interval; a basic interval is added when bias is noticeable. |
| Permutation test | Two groups, paired, or correlation. | Exact logic of a p-value with no normality assumption. |
| Time Series | A numeric series in row order. | Time plot with moving average, ACF and PACF, additive or multiplicative decomposition, lag-1 plot, linear trend with Durbin-Watson, Holt forecast with the seasonal pattern added back. |
| Principal Component Analysis | Three or more numeric variables. | Eigenvalues, scree plot, loadings (varimax optional), score plot coloured by a group. |
| Exploratory Factor Analysis | Items of a questionnaire. | Principal axis factoring with varimax; loadings below a cutoff hidden; communalities. |
| Reliability (Cronbach's alpha) | Items of one scale. | Alpha, alpha if item dropped, item-rest correlations; reverse-keyed items flipped for you. Alpha at least 0.7 is the usual bar. |
| k-means Clustering | Two or more numeric variables. | Standardized by default, k-means++ starts, elbow plot, cluster plot on the first two principal components, optional cluster column added to the data. |
| Survival Analysis | Time to event with censoring. | Kaplan-Meier curves with confidence bands and censoring ticks, median survival, risk table, log-rank test between groups, Cox proportional hazards model with hazard ratios and a forest plot. |
| Bayesian Inference | A proportion (beta prior) or a mean (normal prior), from a column or typed numbers. | Prior, likelihood and posterior on one plot; posterior mean, SD and credible interval; probability the parameter is above a value you choose. |
Results
- Decimal places shown: 2, 3 (default), 4 or 6.
- Print or save as PDF: prints only the results (data table and menus are hidden). Use your browser's Save as PDF.
- Export results as HTML: one file with every card and graph, opens in any browser, good for submitting.
- Save session: a
.sww.jsonfile with the data, filter, analyses and graphs. Open a saved session brings it all back, on any computer. - Clear analyses and Clear graphs empty the tabs. Individual cards have their own remove button.
8. Honest graphs
Statistics software will happily draw a misleading picture. Stats Without Walls refuses to, and prints the rules it followed under each graph so students learn them.
- The count axis starts at zero and cannot be zoomed on bar plots and histograms. A difference of two students can never be stretched to look like a cliff.
- Equal bins with nice edges. Histogram bins are all the same width, chosen from 1, 2, 2.5, 5, 10 times a power of ten, with tick marks on the edges. Each class includes its left edge; the last class also includes the maximum, so every value is counted once. When you group by a variable, all groups share the same bins so their shapes can be compared.
- Ordered scales stay in order. Freshman to Senior, Never to Always, Strongly disagree to Strongly agree, 1 to 5 and similar scales are drawn in their natural order. Sorting them tallest first hides the shape of the distribution.
- Same width for every bar; only height carries information. No 3D, no pictograms.
- Pies are flat and carry their percent and count, so nobody judges angles by eye.
- Numbers are on the picture. The count or percent is printed above each bar and bin by default, and a frequency table can sit under the graph.
9. Writing up a result
Use the "copy text" button, then write one paragraph with these parts: what was compared, the statistic with its degrees of freedom, the p-value, the decision at alpha, and what it means in context. The blue box gives you most of this sentence.
Always report the condition check. If it failed, say so and report the nonparametric or exact result instead.
10. Saving and sharing work
- Your browser remembers the current dataset, filter and results between visits on the same computer. This is a convenience, not a backup.
- For homework, Results, Export results as HTML or Print or save as PDF and upload the file.
- To continue later or on another computer, Results, Save session and keep the
.sww.jsonfile. - To share a dataset with a classmate, Data, Download data as CSV.
11. How it differs from other tools
Menus, option names and table names match the course instructions, so they work as written. Where common statistics software and the course textbook disagree, Stats Without Walls follows the textbook:
| Topic | Other tools | Stats Without Walls |
|---|---|---|
| Quartiles | Type 7 interpolation | Textbook halves by default (median excluded); Type 7 available |
| Mode with ties | Shows one value | Lists every mode, or "no mode" when all values are unique |
| Standard deviation | Sample (n minus 1) only | Sample by default, population SD available and labelled |
| One proportion | Exact binomial only | Exact binomial plus the z test with the success-failure check |
| Two proportions | Often missing from the base menus | Pooled z test with confidence interval |
| Means | t, or z with a known sigma | t by default; the Z-Test needs sigma typed in and its card says to use t when only s is known |
| One proportion interval | One method | Standard (Wald) by default; plus four, Agresti-Coull, Wilson and exact on request |
| Histogram bins | Automatic, may differ by group | Equal nice bins shared across groups, editable width and start |
| Summary statistics input | Often not available | Every t-test, proportion test and chi-square accepts typed summaries |
| Graphs | Mixed with output | Own Graphs tab |
| Price | Free to paid subscriptions | Free, always |
12. Troubleshooting
| Symptom | Cause and fix |
|---|---|
| A numeric column is not offered in a t-test or histogram | One cell contains text (such as "NA" or "n/a"), so the column was typed as nominal. Clear the cell, or use Data, Set variable type, Continuous. |
| "Column not found" | A filter or formula uses a name that does not match exactly (case and underscores matter). Copy the header text. |
| Only one variable gets selected in a multi-select list | Hold Cmd (Mac) or Ctrl (Windows) while clicking. |
| Every row is grey and analyses say "no rows" | The filter excludes everything. Clear the filter box and press Enter. |
| "Levels of X: ..." error | You typed a level that is not in the data. The message lists the levels; copy one exactly. |
| Pink "small group" warning | A group has fewer than 5 rows. The result is fragile; say so in the write-up or combine groups. |
| Results look stale after an update | The browser cached an old version. Reload with Shift held (or Cmd Shift R). The browser-installed version updates itself on the next launch with internet; the desktop app does not, so download the new version from the download page when one is announced. |
| Mac says the app "cannot be opened" or "is damaged" | The app is not Apple-signed, so macOS quarantines it. Right-click the app in Applications and choose Open, once. If there is no Open Anyway button (newer macOS), open Terminal and run xattr -d com.apple.quarantine "/Applications/Stats Without Walls.app", then open the app normally. Also check you downloaded the right build: Apple Silicon for M1 to M4, Intel for older Macs. |
| Windows SmartScreen warning | Click More info, then Run anyway. The installer is unsigned but the same file as on the download page. |
| Nothing prints | Use Results, Print or save as PDF, not the browser's own print button, so the data table is hidden. |
13. Sample datasets
| Dataset | What it is | Good for |
|---|---|---|
| Our class data (40 students) | Survey of a statistics class: major, year, age, commute, study and sleep hours, coffee, exercise, employment, satisfaction, GPA. | Everything in a first course. |
| Class data, exam 1 version | The same class with exam scores added. | Paired and regression questions. |
| Sleep follow-up | Sleep before and after an intervention. | Paired t-test. |
| Finch beaks | Beak depth in Galapagos finches before and after a drought. | Two-sample t, histograms. |
| Global health | Countries: life expectancy, income, health spending, region. | Correlation, regression, ANOVA by region. |
| Police calls for service, academy fitness, community survey | Law-enforcement course datasets. | Proportions, chi-square, one-way ANOVA, time patterns. |
| CDC teen survey 2023 (1500 students) | Youth Risk Behavior Survey extract: marijuana use, sleep, grades, mood. | Contingency tables, two proportions, logistic regression. |
| General Social Survey 2018 and 2022 | Astrology and science beliefs, politics, wellbeing of US adults. | Chi-square, ordinal outcomes, bar plots of ordered scales. |
| Big Five personality (1200) and Big Five 50 items (500) | Personality scores and raw questionnaire items. | Correlation, PCA, factor analysis, reliability. |
| Cadet mile times | Times at weeks 1, 4, 8, 12 for two training groups. | Repeated measures and mixed ANOVA. |
| Mauna Loa CO2 monthly | NOAA monthly CO2 since 2000. | Time series. |
| Lung cancer survival (227 patients) | Survival time, death or censoring, age, sex, performance score. | Survival analysis. |
A codebook for the real-world datasets is in the app's data folder (README_real_world_datasets).
14. How the numbers were checked
Every procedure was run on the same data in R (and the jmv package where it exists) and compared to four decimals: t-tests, proportion tests, ANOVA and Tukey, chi-square, regression coefficients and intervals, nonparametric tests, multiple and logistic regression with diagnostics, two-way and repeated measures ANOVA, count, multinomial and ordinal models, power calculations, time-series decomposition and autocorrelation, PCA, factor analysis and varimax rotation, Cronbach's alpha, k-means, Kaplan-Meier, log-rank, Cox regression and the Bayesian conjugate updates. Distributions come from jStat; graphs are drawn with Plotly.
The distribution calculators (binomial, custom discrete, Poisson, geometric, hypergeometric, discrete uniform, uniform, exponential) were checked against R's d, p and q functions; the z tests, the five proportion intervals (against prop.test and binom.test), the variance tests (against var.test), Levene's test (against car::leveneTest), kurtosis, percentiles and the regression confidence and prediction intervals (against predict) the same way. Simulated data were checked by their mean and SD against the distribution they came from.
The Learn demonstrations are checked the same way. The simulated sampling distributions are compared against sigma / sqrt(n); the confidence interval demo is checked by its capture rate against the stated level, which also falls below that level on purpose when the success-failure condition fails; and the p-value demo is checked against the exact binomial, computed independently. Where the z formula and the exact answer disagree, the app prints both and says why rather than hiding the gap.
Stats Without Walls was built by Safaa Dabagh for students at Santa Monica College, West Los Angeles College and Loyola Marymount University, and is free for anyone to use.