Aug 24, 2026Leave a message

What are the minimum sample sizes for NHP Preclinical studies?

There is no single minimum sample size for every non-human primate (NHP) preclinical study.

 

The right number depends on the question, the primary endpoint, the expected treatment effect, biological variability, the statistical design, and the ethics of the work. NHP pharmacology groups run smaller than rodent groups as a rule, but that doesn't mean one fixed number carries across disease models or study aims.

 

The better question isn't "what's the fewest monkeys I can get away with?" It's: What's the smallest sample that can answer the primary question without wasting animals? That framing matters in NHP work, where each animal is a real scientific, ethical, and financial commitment.

 

Why NHP Sample Sizes Usually Run Smaller Than Rodent Studies

 

NHP studies tend to use fewer animals than rodent work, for a few reasons.

 

NHPs are closer to humans in physiology and anatomy, so some translational endpoints carry more weight per animal. NHP studies also lean heavily on longitudinal measurement: repeated PK draws, imaging, behavioral scoring, biomarker tracking. One animal supplies many data points across the study, so a small group can still yield a lot.

 

On top of that, primate research carries heavier ethical weight. Apes and monkeys aren't easy to source, and the welfare bar is high. So the goal isn't to use as many as possible, it's to use the fewest that still give an interpretable result.

 

ARRIVE 2.0 reflects this. The guidelines require authors to report the exact number of experimental units and explain how they got it, including any a priori sample-size calculation. They don't hand you a universal animal count.

 

So How Many NHPs Are Typically Needed?

 

For most exploratory and pharmacology work, NHP groups land in the single digits, not the tens. But the range is wide.

 

Minimum Sample Size for NHP Preclinical Studies

 

An exploratory study with a narrow aim, feasibility, model setup, exposure, early PD signals, can sometimes run with just a few animals per group. A controlled efficacy study needs more statistical backing. Several per group may do if the endpoint is tight and the effect is clear; you'll need more when the endpoint wobbles or the expected effect is small.

 

Longitudinal and crossover designs can get away with fewer animals, since each one acts as its own control. The opposite holds when you've got multiple arms, noisy endpoints, expected dropouts, or several independent comparisons stacked up.

 

So treat "5 per group" or "10 per group" as anchors, not laws. They're starting points that depend on the design, not regulatory floors.

 

What Determines the Minimum Sample Size?

 

The sample size should follow the study's primary objective, not the species.

 

1. Primary endpoint

Usually the biggest lever. A stable quantitative readout needs fewer animals than a noisy behavioral, imaging, or disease-severity score.

 

A PK study, for instance, pulls repeated concentration measurements from each animal, so a small group can still nail the exposure question. An efficacy study built on a wandering clinical score needs more units to see a real treatment effect. Size the study around whichever endpoint is carrying the main evidence.

 

2. Expected effect size

Effect size is the difference you're trying to detect. A large, biologically meaningful effect shows up with fewer animals; a small one needs many more.

 

That's why "NHP studies need at least six per group" is incomplete on its face. Six might be plenty for one endpoint and nowhere near enough for another. Anchor your expected effect in prior NHP data, a pilot, historical controls, published work, or a defensible estimate of the clinically relevant gap.

 

3. Endpoint variability

Variability drives sample size directly. A scattered group needs more animals to separate signal from noise; a tight, reproducible endpoint can do the job with fewer.

 

That scatter comes from age, sex, weight, baseline disease severity, PK differences, how the disease was induced, individual biology. In NHP work, controlling that variability at the design stage can matter as much as adding animals.

 

4. Statistical power and significance level

For hypothesis-driven studies, power analysis sets the number. The inputs are the expected effect size, the variability or standard deviation, the significance level (usually α = 0.05), the target power (often 80% or 90%), and the statistical model and design.

 

The direction is simple: smaller effect to detect, stricter Type I error, or higher power all push the required n up.

 

One caveat. Run the power analysis on the real primary endpoint. Borrowing variance from an unrelated model or a different endpoint produces sample-size estimates that don't hold, and that bite is sharper in NHP work where comparable historical datasets are thin.

 

5. Study design

The design itself changes the math. Parallel groups need separate animals in each arm. Crossover or repeated-measures designs squeeze more information out of each animal and can shrink between-animal variance.

 

Longitudinal tracking of the same animals is more efficient for the same reason. But repeated measures don't erase the need for adequate n. The correlation between timepoints, expected dropout, missing data, and the right model all belong in the planning.

 

Exploratory Studies vs. Confirmatory Efficacy Studies

 

The line that matters most is exploratory versus confirmatory.

 

Exploratory work might aim to show a disease model induces reliably, sketch a PK profile, find a dose range, check target engagement, or produce early signs of activity. These often run small on purpose, to learn enough before committing to a bigger efficacy program.

 

A confirmatory efficacy study is different. If you're trying to demonstrate a predefined effect against a control, the n should rest on a formal statistical rationale whenever the data exist to support one.

 

This is also why published NHP sample sizes swing so widely. A small exploratory study and a powered efficacy study aren't comparable just because both used macaques.

 

Why "More Animals" Isn't Always the Answer

 

Bigger n raises power, but it isn't a cure-all.

 

If the design is sloppy, more animals won't fix it. Variability from inconsistent disease induction, the wrong endpoint, technical noise, or weak randomization survives a larger sample.

 

In NHP work the smarter move is usually to cut avoidable variability first. Standardize how you induce disease, baseline the animals properly, balance allocation, set inclusion and exclusion criteria up front, keep sampling consistent, and use objective quantitative endpoints. That serves the science and the ethics at once, using animals only when justified.

 

When a Small NHP Sample Can Still Be Informative

 

A small sample doesn't equal a weak study.

 

NHP studies can pull a lot from each animal when they build in repeated measures and complementary endpoints. One animal can contribute longitudinal PK, PD biomarkers, imaging, clinical observations, and behavior scores across the study. Advanced in vivo imaging lets you track disease progression or response over time without killing the animal at every timepoint; Prisys' translational platform, for instance, supports MRI, CT, PET-CT, and DSA for this kind of longitudinal NHP work.

 

Quantitative behavior works the same way. Markerless 3D tracking, often AI-driven, measures movement and behavior over time and adds another longitudinal endpoint for pharmacology.

 

None of this replaces a proper sample-size calculation. It just makes each experimental unit earn its keep.

 

A Practical Approach to NHP Sample-Size Planning

 

A workable planning sequence looks like this:

 

1. Pin down the primary question and the endpoint that answers it. Design around the endpoint, not a predetermined headcount. 2. Estimate the expected effect and variability from historical NHP data, a pilot, the literature, or another defensible source. 3. Pick the statistical model and set the significance level and power you want. 4. Run the a priori sample-size calculation where it applies. 5. Before locking the number, factor in attrition, repeated measures, number of arms, sex mix, randomization, and welfare.

 

When reliable variance estimates don't exist, which is common in NHP work, you'll lean on a mix of statistics, prior experience, pilot data, and judgment. The literature flags exactly this: directly relevant preliminary data are often thin, and power calculations built on unrelated endpoints don't hold up.

 

Does Every NHP Study Need a Formal Power Calculation?

 

Not always.

 

A formal power calculation earns its keep in hypothesis-driven studies meant to detect a predefined effect. For some exploratory studies, it's hard or wrong to run one when you don't yet have reliable effect-size or variability estimates.

 

What matters then is transparency. Say why you picked the number, and what question the study is built to answer. ARRIVE 2.0 asks for the same: explain how the sample size was set, and show the a priori calculation if you did one.

 

What's a Reasonable Starting Point for an NHP Study?

 

There's no universal number, but a reasonable rule of thumb: exploratory NHP studies often run in single digits per group, while a properly powered efficacy study may need more depending on the endpoint and its variability.

 

Set that number from the objective and the statistical assumptions, not from a generic NHP sample-size table. An experienced NHP team will weigh the disease model, expected pharmacological effect, historical controls, endpoint variability, design, and available longitudinal measures before landing on the final count.

 

Sample Size Is a Design Question, Not an Animal Count

 

The right NHP sample size is the smallest one that answers the predefined question.

 

That means there's no universal "minimum" that covers pharmacology, disease modeling, PK/PD, imaging, or safety. Three per group can be right for one exploratory aim and wrong for another. And piling on more animals won't save a bad endpoint or uncontrolled variability.

 

A justified NHP study balances statistical sensitivity, translational goals, feasibility, and welfare from the first line of the protocol. For sponsors planning an efficacy or translational pharmacology study, sample-size planning belongs with model selection, endpoint definition, randomization, dosing, PK/PD sampling, and analysis, not bolted on at the end.

 

Reference

 

Percie du Sert N, Hurst V, Ahluwalia A, et al. The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research. PLoS Biology. 2020;18(7):e3000410. doi:10.1371/journal.pbio.3000410.

 

Contact Prisys Biotech

 

FAQ

Q: What is the minimum number of NHPs required for a preclinical study?

A: No universal minimum exists. The right size depends on the objective, primary endpoint, expected effect size, variability, statistical design, and ethics.

Q: Is 3 animals per group enough for an NHP study?

A: It can work for some exploratory or feasibility studies, but it's not a general floor. Whether three is enough depends on the endpoint and the question.

Q: How many NHPs are typically used in efficacy studies?

A: Often single digits per group, but the exact count swings with the model and endpoint. A powered study should use an endpoint-specific calculation or another defensible rationale.

Q: Can repeated measurements reduce the required number of NHPs?

A: Sometimes. Longitudinal and repeated-measures designs get more from each animal and can trim between-animal variance. But you still have to account for the model and the correlation between timepoints when setting n.

Q: Should sample size be calculated before starting an NHP study?

A: For hypothesis-driven work, do the a priori calculation when you have reliable effect-size and variability estimates. For exploratory studies, a formal power calc may be out of reach if the preliminary data don't exist, but document the rationale for whatever you chose.

 

 
 

Send Inquiry

whatsapp

Phone

E-mail

Inquiry