Distributions Field Guide

Built-in AI tutor. Not sure which shape you are looking at, or what to report once you know? Ask the helper on this page. It coaches you to read the distribution, it does not do the analysis for you.

Before you summarize any column of data, ask it four questions: where is the center, how wide is the spread, what is the shape, and where does a value sit by position. The shape question is the one most people skip, and it is the one that decides whether the average you are about to quote describes anyone. This guide names the five shapes you will meet in accounting data, shows you how to spot each one, and tells you what each implies for the decision in front of you.

The habit this page is teaching: name the shape, state the spread, flag the outlier. A number reported without its shape is a half-truth. The mean of a right-skewed column can describe almost nobody in the data.
The four questions Box-and-whisker Normal Right-skewed Long-tailed Bimodal Pareto Tell them apart The empirical rule Build your own

0The four questions: center, spread, shape, position

Every descriptive statistic answers one of four questions. Reach for the right tool by knowing which question it answers.

QuestionThe toolsThe rule of thumb
CENTER
where is the middle
mean (average), median (the middle value), mode (most common value) Median for skewed data, mean for symmetric data, mode for categorical or most-common questions. When the mean and median disagree, the data is telling you it is skewed.
SPREAD
how wide is it
range (max minus min), IQR (the middle 50%), standard deviation, CV (coefficient of variation, sd / mean) A number without its spread is a half-truth. Use the IQR when outliers would distort the range or the standard deviation. You can compute a standard deviation on any shape; it is the empirical-rule percentages that need a bell. Use CV to compare spread across columns measured on different scales, and only when the mean is positive and well away from zero, because a mean near zero makes the ratio blow up. A ratio that straddles zero has no useful CV. The box-and-whisker plot below is the picture of spread.
SHAPE
what does the curve look like
skewness (lopsidedness), kurtosis (how long the tails run) Mean greater than median usually means a right skew (a long tail of large values). Long tails (high kurtosis) mean the tails run wider than a normal bell, so more extreme values turn up than a bell curve would predict. This is the bridge to the field guide below.
POSITION
where does one value sit
percentiles, quartiles, the 5-number summary (min, Q1, median, Q3, max), z-score (how many spreads from the center) "Top decile" and "90th percentile" are position statements. A z-score of 3 means three standard deviations above the mean. The empirical rule lives here, with the caveat below.
The bridge: the SHAPE question is what the rest of this page is about. Once you can name the shape, you know which CENTER to report, which SPREAD fits, and whether the empirical rule applies at all.

BThe box-and-whisker plot: the picture of spread

The box plot draws the 5-number summary with one refinement: the whiskers stop at the most extreme values still inside the fences, so a value far out shows as its own dot instead of stretching the whisker. It makes spread, skew, and outliers visible at a glance, and it stacks cleanly side by side so you can compare groups without doing any arithmetic in your head.

fence fence Q1 Q3 outliers whisker end median whisker end 1.5 IQR IQR 1.5 IQR box = IQR (middle 50%)
The boxRuns from Q1 to Q3, so it spans the IQR, the middle 50 percent of the values. The line inside is the median.
The fencesSit 1.5 IQRs past each quartile: Q1 minus 1.5 times the IQR, and Q3 plus 1.5 times the IQR. Most charts do not draw them, but they are the rule the chart uses.
The whiskersEnd at the most extreme values still inside the fences. They are not the minimum and the maximum. When there is an outlier, the whisker stops short of it.
The dotsOutliers, every value past a fence. The box plot is where they declare themselves, with no separate test needed.
Read it: a median pushed to one end of the box usually means the data is skewed. A median near the left of the box (as drawn) with a long right whisker is a right skew, the same story the mean-greater-than-median rule tells, drawn instead of computed.
Use it: draw one box per group side by side (commercial versus consumer, by region, by month) and the comparison is quick. Whose spread is wider, whose median is higher, and who has the outliers all show in one chart, with no averages to mislead you.

1Normal: the symmetric bell

mean = median
Spot itSymmetric, both sides mirror each other, and the mean sits right on top of the median.
In accountingMeasurement error, large-sample averages, standardized test-like scores. Rare in raw financial amounts.
It impliesThe data is typical and predictable. The mean is a fair summary, and the empirical rule applies, so about 68 percent sits within one standard deviation and 95 percent within two.
The trap: assuming financial data is normal when it almost never is. Revenue, balances, and losses are skewed or long-tailed far more often than they are bell-shaped. Do not default to the bell just because it is the one shape everyone learned in school.

2Right-skewed: the long right tail

median mean
Spot itA tail stretching to the right, and the mean noticeably greater than the median.
In accountingRevenue, firm size, transaction amounts, executive pay, account balances. The most common shape in financial data.
It impliesMost observations are small and a few are huge. Report the median, because it describes the typical case, while the mean is dragged up by the big accounts in the tail.
The trap: quoting the mean. The tail inflates it until it describes almost nobody, "the average customer balance is 18,000 dollars" when 80 percent of customers owe under 5,000 and a handful owe millions. Lead with the median, then mention the tail.

3Long-tailed (also called heavy-tailed): sharp peak, tails that run far out (leptokurtic)

tails run wider than the bell
Spot itA sharp central peak plus tails that run far out, so extreme values turn up more often than a bell would predict. The tails are longer and wider than the dashed bell behind it, which shows what "normal" would look like.
In accountingStock returns, audit errors, fraud losses, daily cash swings.
It impliesExtreme events are more common than they look, so plan for them. A loss the bell would call "once in a century" can show up much more often than that.
Long-tailed versus Pareto: long-tailed is about how far a single measure can swing (how big one daily return or one loss can get); Pareto (below) is about how a few categories dominate a total (how a few customers hold most of the overdue dollars). Both involve a "tail," but they answer different questions, so keep them apart.
The trap: sizing risk off the normal curve and being blindsided by the "that shouldn't happen" event that does. If the tails run wider than a bell, a two-standard-deviation cushion is not the 95 percent of safety the bell promised.

4Bimodal: two modes in one distribution

the textbook picture
Spot itBimodal means two modes in one distribution: two humps with a dip between them, so the data has two centers rather than one.
In accountingTwo customer segments (commercial versus consumer, the Week 5 receivables split), manual versus automated journal entries (the Week 2 lab), two product lines.
It impliesYou likely have two populations stacked on top of each other. Split them and summarize each on its own.
Real dataTwo populations do not always show as two clean peaks. The pooled Week 5 receivables do not: Commercial is 543 invoices whose middle 80 percent spreads from about 4,300 to 18,000, so pooled with the 1,957 Consumer invoices it reads as a long shoulder to the right of the Consumer hump rather than a second peak. The two centers, a Consumer median of 1,502 and a Commercial median of 10,502, show when each segment is drawn on its own. The test is whether splitting changes the answer, not whether you can count two peaks.
The trap: reporting one average that describes nobody in either group. In the Week 5 receivables, commercial invoices average about 10,900 and consumer invoices about 1,500, so the blended "average invoice" of about 3,600 lands in a sparse stretch holding 24 of the 2,500 invoices. It describes almost no real customer. The blended median, 1,818, does no better, because it sits inside the Consumer book and describes nothing about Commercial.

5Pareto: the vital few, and the 80/20 nickname

Spot itRanked categories where a few dominate the total, then a long flat run of small ones. Sort descending and the bars drop fast while the cumulative line climbs steeply.
In accountingRevenue by customer, sales by SKU, overdue dollars by account, late-arrival causes.
It impliesFocus effort on the vital few, because working the top of the list recovers most of the dollars. Say how few, since the count is what turns the chart into a plan.
80/20The rule of thumb says 20 percent of the customers hold 80 percent of the dollars. It is a nickname for concentration, not a measurement of it, so measure your own data before you quote it.

The real Week 5 past-due balance at March 31, 2026: 85 customers, $1,421,128 in total. The bars (left axis) rank each customer's past-due dollars; the line (right axis) is the running cumulative percent.

Read the chart: the cumulative line crosses 80 percent at the 27th customer, so 27 of the 85 customers, 32 percent of them, hold 80 percent of the past-due dollars. This book is 80/32, not 80/20. Measured the other way, the top 20 percent of customers (17 of 85) hold 66 percent. The concentration is real and worth acting on, and it is milder than the nickname.
Pareto-squared, and why you check it too: the nickname says the concentration recurses. If 20 percent hold 80 percent, the top 20 percent of that group, about 4 percent of all items, would hold about 64 percent of the whole (0.8 times 0.8). That is arithmetic on the nickname, not a finding. In this book the top 3 customers, about 4 percent of the 85, hold 21 percent, so it does not meet Pareto-squared. The largest accounts are still where triage starts, but the number you report is the one you measured.
The trap: treating all items as equally worth your time, or its mirror, writing "it follows the 80/20 rule" without drawing the line. Spreading collection effort evenly across every account usually wastes it. A slide that says "27 accounts carry 80 percent of what is past due" is a plan the reader can act on, while "it follows the 80/20 rule" is a claim nobody checked.

★Tell them apart at a glance

When you are staring at a fresh column and not sure which shape it is, run down this short list. The fastest tell for each is the first thing to check.

ShapeThe fastest tellWhat to report
NormalOne symmetric bell: mean and median about equal, and the tails thin out fast on both sides.The mean is fair. The empirical rule applies.
Right-skewedMean is clearly above the median, one long tail to the right.Report the median, then mention the tail.
Long-tailedCenter is fine, but extreme values show up far more often than a bell would allow.The center is not the story; the tail risk is. Plan for extremes.
BimodalTwo modes, two humps with a dip between them. Two groups can also hide as one hump and a shoulder.If you suspect two groups, split them, summarize each on its own, and check whether the answer changes.
ParetoRank the categories; a few hold most of the total and the cumulative line climbs steeply.Work the vital few first, and say how many it takes to reach 80 percent.
The two "tail" shapes are the easy ones to confuse. Long-tailed asks how far one measure can swing (one return, one loss). Pareto asks how a few categories split a total (which customers own the overdue balance). Skew and long tails both bend the empirical rule, while bimodal and Pareto are telling you the data may really be several groups rather than one.

6The empirical rule, and when it breaks

For a normal distribution, the spread follows a fixed pattern. This is the 68-95-99.7 rule, and it is genuinely useful for the rare normal column.

68%
within 1 standard deviation of the mean
95%
within 2 standard deviations
99.7%
within 3 standard deviations

Highlight the bands and watch the rule hold, then break. Pick a distribution and a band. The shaded area is the mean plus or minus that many standard deviations, and the readout compares what the rule promises against what the data actually does.

Distribution
Band

Where it fails: the empirical rule holds for normal data, and most accounting data is not normal. On right-skewed data (revenue, balances) the mean is not the center the rule assumes, so the bands are lopsided and misleading. On long-tailed data (returns, losses) far more than 0.3 percent of values land beyond three standard deviations, so the rule badly understates how often extremes occur. Name the shape first, then decide whether the rule applies at all.
The better move: when the shape is not normal, describe the spread with the 5-number summary and the IQR. They make no assumption about the shape, so they stay accurate on skewed and long-tailed data. The standard deviation is still a valid number on any shape; what needs the bell is reading it as 68, 95 and 99.7 percent. The CV needs more care again, because it divides by the mean, so it only works when the mean is positive and well away from zero.

7Build your own: prompts for an interactive distribution tool

You do not need this page to be interactive, because the AI tool you already use can build you one on your own data in a few minutes. Copy a prompt below, fill in the [bracketed] slots, and paste it into your AI tool. Each one asks for a single HTML file you open in your browser, which keeps the result small enough to read and check. These buttons only copy text; nothing on this page sends your prompt anywhere.

1. Project 1: your ratio by sector, raw against winsorized

Build me a single self-contained HTML file, using plain JavaScript and no outside libraries, that I can open in my browser. It should let me pick a CSV file from my computer with a file picker. The CSV has one row per firm, a sector column named [your sector column, such as gsector] and a ratio column named [your ratio column, such as gross_margin].

Draw side-by-side horizontal box plots of the ratio, one per sector, on one shared axis. Use Tukey's rule: the box runs from Q1 to Q3, the whiskers end at the most extreme values still inside 1.5 IQRs of the quartiles, and every value past that is drawn as a dot. Compute quartiles the way Excel's QUARTILE.INC does.

Add one toggle that switches between the raw ratio and the ratio winsorized at the [1st and 99th] percentile, [computed across all firms in the file before grouping by sector]. Under the chart, show a table with n, the median, Q1, Q3 and the number of outliers for each sector in the current view. Skip rows where the ratio is blank or not a number, and print how many you skipped.
Why it works: it is the Project 1 comparison drawn the way you will present it, and the toggle shows how much of each sector's picture was driven by a few extreme firms, which is something your slide has to explain either way.

2. A histogram with a bin-width slider

Build me a single self-contained HTML file, using plain JavaScript and no outside libraries, that I can open in my browser. It should have a text box where I paste one column of numbers copied from Excel, one number per line. The column is [what the numbers are, such as invoice amounts in dollars].

Draw a histogram of the numbers, with a slider that sets the bin width from [a small width, such as 100] to [a large width, such as 5,000] and redraws the chart as I move it. Draw the mean and the median as two labeled vertical lines in different colors. Under the chart, print the count, mean, median, Q1 and Q3, with quartiles computed the way Excel's QUARTILE.INC does. Ignore blank lines, and tell me how many lines could not be read as a number.
Why it works: the same data can look smooth, lumpy or two-humped depending on the bin width alone, so moving one slider shows you how much of a shape is the data and how much is the chart.

3. A box plot you can drag

Build me a single self-contained HTML file, using plain JavaScript and no outside libraries, that I can open in my browser. Draw one horizontal box plot of these numbers: [paste 15 to 30 values, such as days to pay for a set of customers]. Draw each value as a dot just below the box.

Let me drag any dot left or right with a mouse or a finger. As I drag, recompute and redraw Q1, the median, Q3, both fences at 1.5 IQRs past the quartiles, the whisker ends at the most extreme values still inside the fences, and which dots count as outliers. Show those numbers in a small table that updates while I drag, with quartiles computed the way Excel's QUARTILE.INC does.
Why it works: moving one value at a time shows you which statistics a single extreme value can move, the whisker and the outlier flags, and which ones it barely touches, the median and the quartiles.
Check one number before you trust the chart. Put the same column in Excel and compare one number the tool prints, such as =MEDIAN(B2:B500) or =QUARTILE.INC(B2:B500,1). If they match, the chart is probably drawing what you think it is. If they do not, ask the tool why before you read anything off the chart, because the usual causes are a header row read as data, blank cells read as zero, or a different quartile method.
Carry this off the page: the four questions (center, spread, shape, position) work on any dataset in any tool. Naming the shape is what turns "the average is 18,000" into "the balances are right-skewed, so the typical account is 4,000 and a few large accounts pull the mean up." The second sentence is the one a CFO can act on.

ACCTG 6155, Fall 2026, Week 6. The companion Distribution Explorer is the sandbox for these shapes, and the Analytical Prompt Library shows how to ask an AI to name them for you and turn the finding into a story.