Histogram Hub

What Is a Histogram?

What a histogram is

A histogram is a chart that shows how a set of numbers is distributed. It splits the range of your data into equal intervals called bins, then draws a bar over each bin whose height is the count of values that fall inside it. Put simply, it shows you where your numbers pile up and where they thin out.

One histogram describes one variable. If you measure the height of 200 people, a histogram of those heights tells you the typical height, how much people vary, and whether anyone is unusually short or tall. It does all of that in a shape you can read in about two seconds.

The parts of the chart

Four pieces, and once you can name them the chart stops being mysterious.

  • The horizontal axis carries the variable you measured, on its own scale. Scores, dollars, minutes, millimetres.
  • The bins are equal slices of that axis laid end to end, like 60 to 70 and then 70 to 80.
  • The vertical axis carries the count, sometimes called the frequency.
  • The bars rise over each bin to the height of its count.

The bars touch. That is not a style choice. A bar chart puts gaps between bars because each bar is a separate category and there is nothing in between them. In a histogram the axis is continuous, so a gap would mean an empty bin, which is real information. If you see a gap in a histogram, some part of the range genuinely had no values in it.

A worked example

Here are the ages of 24 people at a community class:

19, 21, 22, 24, 25, 27, 28, 28, 30, 31, 33, 34, 35, 36, 38, 41, 43, 45, 47, 52, 55, 58, 63, 71

Slice the range into bins ten years wide and count:

BinCountBar height
10 to 201▍
20 to 307███
30 to 407███
40 to 504██
50 to 603█
60 to 701▍
70 to 801▍

The counts add to 24, which is the check worth doing every time. Now read it. Most people are in their twenties and thirties, the numbers thin out steadily above 40, and there is a long thin tail stretching to 71 with nobody to balance it on the left. That is a right skewed distribution, and it is the shape you get from almost anything that has a floor but no ceiling, like age, income, or house prices.

Notice what the table gave up. You can no longer see that someone was exactly 47. You traded the individual values for the shape, and that trade is the entire point of the chart.

What you learn from the shape

  • Center. Where the tall bars sit, roughly the typical value. In the example above, the low thirties.
  • Spread. How far the bars stretch. Tight bars mean the values agree with each other, wide bars mean they do not.
  • Shape. Symmetric, skewed toward one side, one peak or two. Two clear peaks usually means two groups got mixed into one dataset. The histogram shapes page walks through each named shape.
  • Gaps and outliers. A lonely bar far from the rest is worth a second look. It is either an interesting case or a typing mistake, and you want to know which.

Which bin does a boundary value go in

If your bins are 60 to 70 and 70 to 80, where does a value of exactly 70 go? The usual convention is that a bin includes its lower edge and excludes its upper one, so 70 goes in the 70 to 80 bin. The tool on this site follows that rule.

The rule itself matters less than being consistent. Pick one and hold to it, or your counts will not add up to the number of values you started with.

Bin width changes what you see

The same numbers can look like different data depending on how wide you make the bins. Too few bins and everything collapses into two or three fat bars that hide the shape. Too many and every bar is a count of one, which is a list of your data pretty much, not a summary of it.

There is no single correct answer, which is why several rules exist for picking a starting point. The how to choose bins page covers the practical version, and the post on Sturges, Scott and Freedman-Diaconis covers the formulas if you need to justify a choice to someone. The honest working method is to start with a rule, then nudge the width up and down and keep the version that shows the structure without inventing any.

What a histogram cannot do

Three real limits, worth knowing before you rely on one.

It hides individual values. Once a number is in a bin it is just part of a count, so you cannot recover it from the chart.

It needs enough data. Below about 20 values the shape is mostly noise, and you are better off with a dot plot or just reading the numbers.

It handles one variable at a time. If you want to know whether two things move together, a histogram is the wrong chart and a scatter plot is the right one.

The other histogram you might mean

If you arrived here from a camera or a photo editor, that histogram is the same idea applied to brightness. The horizontal axis runs from black on the left to white on the right, and the bars count how many pixels in the image sit at each brightness level. Bars bunched at the left mean a dark image, bunched at the right mean a bright one, and pixels stacked flat against either end mean detail has been clipped away and cannot be recovered.

Same chart, same reading method, different variable. The rest of this site is about the statistics version.

Common mistakes

  • Putting gaps between the bars. That turns it into a bar chart and tells the reader the axis is categorical when it is not.
  • Using unequal bin widths without adjusting. A bin twice as wide collects roughly twice as many values, so its bar looks important when it is only wider. If you need unequal bins, plot density instead of counts.
  • Reading the tallest bar as the average. It is the most common bin, which is the mode. On a skewed distribution the mean can sit well away from it.
  • Comparing two histograms drawn with different bins or different axes. Match both before you conclude anything from the difference.

Try it on your own numbers

Paste a column of values into the histogram maker and it draws the chart, builds the frequency distribution table behind it, and lets you switch bin methods to see how much the picture depends on that choice. If you are still deciding whether a histogram is the chart you want, the histogram vs bar chart page is the shortest way to settle it.

Frequently asked questions

What is a histogram used for?
To see the distribution of a single set of numbers: where values cluster, how spread out they are, and whether the shape is symmetric, skewed, or has more than one peak.
What is a bin in a histogram?
A bin is one equal-width interval of the number line. The bar above it counts how many data values fall inside that interval. Bins touch end to end, so histogram bars have no gaps.
Is a histogram the same as a bar chart?
No. A histogram shows numeric data grouped into ranges with touching bars, while a bar chart compares separate categories with gaps between the bars.
Why do histogram bars touch each other?
Because the horizontal axis is continuous. The bins run end to end with nothing between them, so a gap would mean a range of values that genuinely had no data in it rather than a break between categories.
Which bin does a value on the boundary go into?
By the usual convention a bin includes its lower edge and excludes its upper one, so a value of exactly 70 goes into the 70 to 80 bin rather than 60 to 70. What matters most is being consistent, otherwise your counts will not add up to the number of values you started with.
How much data do you need for a histogram?
Roughly 20 values as a floor. Below that the shape is mostly noise and a dot plot or the raw list tells you more.
What does the tallest bar in a histogram mean?
It is the most common bin, which is the mode. It is not the average. On a skewed distribution the mean can sit well away from the tallest bar.
Is the histogram in a camera the same thing?
Same chart, different variable. A camera histogram puts brightness on the horizontal axis, from black on the left to white on the right, and counts pixels instead of data values. Pixels stacked against either end mean detail has been clipped.