ANES Intro: American Public Opinion · Module 2 of 6

Layer 2: What Does the Data Look Like?

Distributions, means, and proportions in survey data. Students use the Profile Explorer and Trends tool to visualize ANES variables and compare groups across election years.

← Previous module ANES Intro: American Public Opinion / Module 2 Next module →

Describing Survey Data: Distribution, Center, and Spread

Before you can test a hypothesis about what predicts political opinions, you need to know what those opinions actually look like in the data. A distribution shows how survey responses spread across the possible answer choices. Is most of the public clustered at one end of a scale, or spread evenly? Does the typical respondent land near the middle, or at an extreme? These are factual questions about the data that descriptive statistics can answer — and that any analysis must get right before moving to inference.

distribution frequency table weighted percentage mean median missing data thermometer variable ordinal scale binary variable

The Three Main ANES Variable Shapes

Categorical / Binary

A question with a small number of named choices, often with no natural order or a simple yes/no split.

Example: Did you vote in the presidential election? (1=Yes, 2=No, 3=Do not know)

Summarize with: frequencies and percentages for each category. The mean is not meaningful when categories have no numeric order.

Ordinal Scale (e.g., 1-7) and Thermometer (0-100)

Ordinal: choices are ordered low to high but intervals are unequal. Example: 7-point ideology scale (1=Extremely liberal to 7=Extremely conservative). Summarize with median.

Thermometer: runs 0 to 100. Respondents rate a group or candidate: 0=very cold, 50=neutral, 100=very warm. Most continuous of all ANES variable types -- use mean.

A thermometer variable runs from 0 to 100. Respondents are told to rate a group, candidate, or party: 0 means very cold (unfavorable), 50 means neutral, 100 means very warm (favorable). These are the most continuous-looking ANES variables and the most appropriate for computing means.

Why the Numbers Are Weighted

ANES samples are not perfectly proportional to the U.S. population — some groups are overrepresented or underrepresented in any given year. Survey weights correct for this: each respondent is assigned a weight that makes their contribution to aggregate statistics reflect what fraction of the real population they represent. When Pulse 2.0 shows you a weighted mean or a weighted percentage, it means the calculation applied those corrections — the number describes what Americans collectively believe, not just what showed up in the specific ANES sample that year.

Understanding Missing Data in Distributions

The percentage missing for a variable is the share of all respondents who did not give a usable answer. High missing rates matter because:

1. If the question was only asked in some years, the missing cases are simply respondents from years when the question was not on the survey — not evidence of confusion or sensitivity.

2. If missing responses are not random (e.g., lower-education respondents are more likely to say "don't know" on policy items), then statistics computed on valid cases may not represent the full population.

Pulse 2.0 always displays N_total and N_valid so you can check the missing rate before interpreting any statistic.

Guess-then-reveal exercise: Before you look at the actual distribution for any ANES variable, try to predict whether the average American rates the Republican Party higher or lower than 50 on a thermometer. Then check the Pulse Layer 2 tool to see how close your intuition was.

Try it now: The interactive Layer 2 tool lets you pick any ANES variable, make a prediction about its distribution, then reveal the real weighted data. Open it at Layer 2: Distribution Explorer →