How to Calculate Sample Size for a Cross-Sectional Study (With Free Calculator)
If you are planning a cross-sectional study ā for example, a survey on dengue awareness in your city ā the first question your supervisor or ethics committee will ask is: "What is your sample size, and how did you calculate it?" This article shows you exactly how to answer, step by step, with a worked example and a free calculator. No laptop or SPSS needed.
Why Sample Size Matters
Sample size is simply the number of people (or households, or patients) you need to include in your study. Getting it right matters for three reasons:
- Too small: your results will be unreliable. You might miss a real finding, or your percentages will swing wildly if just a few people answer differently. Reviewers will reject your paper.
- Too large: you waste time, money, and effort collecting data you did not need. It can also be unethical to trouble more participants than necessary.
- Just right: your estimates are precise enough to trust, and you can defend your number in your thesis or paper with a proper formula.
Every cross-sectional study should report how its sample size was calculated. It usually goes in the methodology section in one or two lines ā and this article gives you everything you need to write those lines.
The Formula, in Simple Words
The standard formula used for a cross-sectional study (the same one the well-known OpenEpi calculator uses) is:
n = [DEFF Ć N Ć p(1āp)] / [(d²/Z²)(Nā1) + p(1āp)]
It looks scary, but each letter has a simple meaning:
- N ā Population size. How many people are in the whole population you are studying? For a city, this could be 1,000,000. If you truly do not know, use a very large number like 1,000,000 ā the formula handles it.
- p ā Anticipated frequency (as a decimal). What percentage of people do you expect to have the condition or characteristic you are studying? If previous research says 30% of people know about dengue prevention, p = 0.30. If you have no idea, use 50% (p = 0.50). This gives the largest sample size, so it is the safe choice.
- d ā Confidence limits / margin of error (as a decimal). How precise do you want your estimate? ±5% (d = 0.05) is the standard for most student research. Smaller d means a bigger sample.
- DEFF ā Design effect. Use 1.0 if you are picking people randomly one by one (simple random sampling). Use 1.5 to 2.0 if you are sampling whole groups at once, like entire streets or villages (cluster sampling), because people in one cluster tend to be similar.
- Z ā the value linked to your confidence level. For 95% confidence, Z = 1.959964. (Note: many textbooks round this to 1.96, but the precise value matters ā with 1.96 the example below wrongly gives 385 instead of 384.)
Worked Example: n = 384
Let us calculate the sample size for a dengue awareness survey in a city:
- Population size N = 1,000,000
- Anticipated frequency p = 50% = 0.50 (we have no previous data, so we use the safe value)
- Confidence limits d = ±5% = 0.05
- Design effect DEFF = 1.0 (simple random sampling)
- Confidence level = 95%, so Z = 1.959964
Plugging into the formula:
n = [1.0 à 1,000,000 à 0.50 à 0.50] / [(0.05² / 1.959964²) à 999,999 + 0.50 à 0.50]
n = 250,000 / 651.02 = 384
So you need 384 participants. In your methodology section you would write: "The sample size of 384 was calculated using the single-population proportion formula with 95% confidence level, 5% margin of error, 50% anticipated frequency, and design effect of 1.0."
Sample Size at Different Confidence Levels
Using the same inputs, here is what the required sample size looks like if you change the confidence level:
| Confidence level | Required sample size (n) |
|---|---|
| 80% | 165 |
| 90% | 271 |
| 95% | 384 |
| 97% | 471 |
| 99% | 664 |
| 99.9% | 1082 |
| 99.99% | 1512 |
Higher confidence demands a bigger sample. For most student research, 95% is the standard.
What If Your Population Is Small?
The formula above already includes the finite population correction ā that is the (Nā1) part. Its job is simple: if your total population is small, you do not need as big a sample. Watch what happens with the same p = 50%, d = 5%, DEFF = 1:
| Population size (N) | Required sample size (n) |
|---|---|
| 1,000,000 | 384 |
| 10,000 | 370 |
| 1,000 | 278 |
| 500 | 218 |
For example, if you are studying all 500 households in a small town, you only need 218 of them ā not 384. Rule of thumb: if your population is under a few thousand, always enter the real N instead of leaving it at 1,000,000.
Adjusting for Non-Response
In real life, some people refuse to answer or leave questions blank. If you expect a 10% non-response rate, inflate your sample size like this:
Final n = calculated n Ć· (1 ā non-response rate)
Example: 384 Ć· (1 ā 0.10) = 384 Ć· 0.90 = 427. So you would approach 427 people, expecting about 384 complete responses. Always round up, never down.
Quick Checklist Before You Start Collecting Data
- Write down your N, p, d, DEFF, and confidence level ā and why you chose each value.
- Calculate n with the formula (or the free calculator below).
- Add your non-response adjustment.
- Write the one-line justification for your methodology section.
- Only then start collecting data ā never the other way round.
Try the Free Calculator
Doing this by hand is good practice once, but you do not need to do it every time. Try the free FormStat sample-size calculator ā enter your N, p, d, and DEFF, and get the sample size for every confidence level instantly, right on your phone. No laptop needed.
Run this analysis in seconds with the FormStat app. No signup, no laptop, works offline.