On-Demand Webinar

Design, Power & Sample Sizes for Multi-Arm Trials

Design, Power & Sample Sizes for Multi-Arm Trials
9:01
Download and explore the data yourself. Files included:
  •  Multiple Arms Single Stage Non-Inferiority - Supersuperiority Tests for Means.nqt


Can You Test More Treatments While Recruiting Fewer Patients?

Multi-arm clinical trial design has come under increasing focus as rising costs, long timelines and the need to evaluate multiple treatments simultaneously have pushed sponsors and statisticians to seek more efficient approaches to drug development. Rather than running several separate two-arm trials, each with its own control group and recruitment burden, a well-designed multi-arm trial consolidates comparisons into a single unified protocol, saving time, costs, and patients.

In this webinar, Calvin O'Brien, Research Statistician at nQuery, provides an overview of multi-arm clinical trial design with a particular focus on comparison structures, family-wise error rate control, and sample size determination and how these considerations affect trial design in practice.

Learning objectives of this webinar:

Multi-arm clinical trial design offers a more resource-efficient alternative to running several separate two-arm studies. By evaluating multiple treatments simultaneously within a single protocol, sponsors can reduce patient recruitment, save time, and lower costs without sacrificing statistical rigour. However, testing multiple treatments at once introduces the challenge of multiple comparisons, which must be carefully controlled at the design stage.

Key regulatory guidance from both the FDA and EMA reinforces the need for pre-specified multiple testing procedures and strong control of the family-wise error rate, particularly in confirmatory Phase III trials.

Whether it's many-to-one comparisons, all-pairs testing, sample size allocation, or family-wise error rate control, there is a clear interest in re-evaluating every aspect of multi-arm clinical trial design to solve the practical challenges that prevent efficient development programmes.

Four key areas are covered:

1. Overview of Multi-Arm Clinical Trial Design

We begin by examining the fundamentals of multi-arm trials:

  • What defines a multi-arm clinical trial
  • Comparison structures: many-to-one vs. all-pairs
  • Efficiency gains over separate two-arm trials
  • Shared control arms and their role in reducing patient burden
  • When multi-arm designs are most appropriate

This context sets the stage for understanding why multi-arm clinical trial design is not just a methodological option, but an increasingly necessary strategic choice.

2. Design Considerations for Multi-Arm Trials

Effective multi-arm clinical trial design requires careful pre-specification of several key elements:

  • Number of arms and treatment comparisons
  • Control of the family-wise error rate (FWER)
  • Endpoint hierarchy and success criteria
  • Multiple testing procedures and regulatory alignment
  • Sample size allocation across arms

With clear regulatory expectations from both the FDA and EMA, sponsors must ensure their multiplicity control strategy is robust, transparent, and pre-specified before confirmatory trials begin.

3. Sample Size Determination and Power

Accurately determining sample size in a multi-arm setting involves unique considerations compared to standard two-arm trials:

  • Power calculations for multiple simultaneous comparisons
  • Disjunctive vs. conjunctive power definitions
  • Impact of the number of arms on overall sample size requirements
  • Balancing statistical power with recruitment feasibility
  • Practical trade-offs in design and allocation choices

Getting sample size right is critical to ensuring a multi-arm clinical trial is both adequately powered and operationally viable.

4. Practical Demonstrations in nQuery

The webinar includes hands-on worked examples in nQuery, covering:

  • Sample size and power calculations for multi-arm trials
  • Controlling the family-wise error rate in practice
  • Step-by-step walkthroughs of real design scenarios
  • How nQuery supports both frequentist and adaptive multi-arm designs

These demonstrations show how the theoretical considerations of multi-arm clinical trial design translate directly into practical planning decisions.

Why Is There Increased Pressure to Adopt Multi-Arm Clinical Trial Design?

The traditional model of running separate two-arm trials for each treatment or dose is increasingly seen as inefficient. These pressures have led to:

  • Re-allocation of development resources toward more operationally efficient trial structures.
  • Growing emphasis on evaluating multiple candidates in parallel without compromising scientific rigour.
  • Increased focus on reducing patient burden by sharing a single control arm across all comparisons.

Multi-arm clinical trial design is therefore not simply a methodological enhancement, it is a strategic response to the systemic inefficiencies of running sequential, siloed development programmes.


What is a Multi-Arm Clinical Trial Design and How Does It Work?

A multi-arm clinical trial evaluates several treatments or doses simultaneously within a single study. The standard two-arm randomised controlled trial - one treatment versus one control is statistically robust, but it becomes an inefficient use of resources when multiple candidates need evaluation.

Multi-arm clinical trial design addresses this by testing several treatments at the same time using one of two core comparison structures:

  • Many-to-one: Each treatment arm is compared against a shared control group. This is the most common structure in confirmatory Phase III trials.
  • All-pairs: Every treatment is compared against every other. More typical in exploratory settings where relative treatment performance is the primary question.

By sharing a single control arm across multiple comparisons, multi-arm designs require fewer total patients than running equivalent separate trials, reducing both cost and recruitment timelines without sacrificing statistical power.


When should you use a many-to-one vs. all-pairs comparison structure?

The choice between a many-to-one and all-pairs comparison structure should be driven by the scientific question the trial is designed to answer.

Many-to-one comparisons

In a many-to-one design, each treatment arm is compared against a single shared control group, typically a placebo or standard of care. No direct comparisons are made between treatment arms.

This structure is most appropriate when the primary objective is to demonstrate that one or more treatments are superior or non-inferior to a control, the trial is confirmatory and operating in a Phase III setting where regulatory evidence is the goal, a clear control condition exists that represents the current standard, and the sponsor does not need to establish relative superiority between active treatments.

Many-to-one is the dominant structure in confirmatory multi-arm trials because it maps directly to the regulatory question: does this treatment work better than the comparator?

All-pairs comparisons

In an all-pairs design, every treatment arm is compared against every other, including direct comparisons between active treatments. This generates a complete picture of relative performance across all candidates.

This structure is most appropriate when the trial is exploratory and operating in a Phase II or dose-finding context, the primary objective is to rank or select between multiple candidates rather than compare each to a fixed control, no established control condition exists, and the sponsor needs to understand relative efficacy between active arms to inform downstream development decisions.

The practical decision

Most confirmatory trials default to many-to-one because it requires fewer total comparisons for a given number of arms, allows for a more straightforward FWER control strategy, and aligns directly with regulatory expectations. All-pairs designs generate richer comparative data but at the cost of a substantially larger number of hypothesis tests and therefore a more stringent multiple testing correction as the number of arms grows.

If the goal is to identify which treatment to take forward into a confirmatory programme, a many-to-one design in Phase II using a shared control is often the most efficient path. If the goal is to understand how a set of candidates perform relative to each other without a natural comparator, all-pairs is the more informative choice.


How Can Multi-Arm Clinical Trial Design Accelerate Development?

Multi-arm designs address one of the most persistent inefficiencies in drug development: the resources lost by evaluating treatments one at a time.

Multi-arm clinical trial design supports:

  • Parallel evaluation of multiple treatments or doses within a single protocol.
  • Shared control arms that reduce the total patient population required.
  • Faster portfolio decisions by generating comparative data across candidates simultaneously.
  • Reduced operational overhead by consolidating trial infrastructure, monitoring, and data management.

As sponsors face rising development costs and increasing time pressure, multi-arm clinical trial design offers meaningful efficiency gains, provided the statistical and regulatory frameworks are implemented correctly from the outset.


Why is family-wise error rate control important in multi-arm trials?

When a trial tests multiple treatments simultaneously, each additional comparison increases the probability of obtaining at least one false positive result by chance alone. This cumulative risk is known as the family-wise error rate (FWER).

In a standard two-arm trial, the Type I error rate is set at 5%, meaning there is a 1-in-20 chance of incorrectly concluding a treatment works when it does not. In a multi-arm trial testing three treatments against a control, that risk compounds across comparisons. Without correction, the probability of at least one spurious positive finding can rise well above 5%, undermining the scientific integrity of the trial.

Strong control of the FWER ensures the overall probability of any false positive across all comparisons in the trial remains within the pre-specified threshold, regardless of how many arms are tested.

This matters for several reasons:

Regulatory requirement: Both the FDA and EMA require pre-specified multiple testing procedures for confirmatory Phase III trials. FWER control must be addressed in the statistical analysis plan before the trial begins.

Scientific credibility: Results from a trial without appropriate multiplicity control are difficult to defend under regulatory scrutiny, regardless of how positive the findings appear.

Trial integrity: Pre-specifying the FWER control strategy removes the risk of selective reporting or post-hoc adjustment of the analysis.

Common approaches to FWER control in multi-arm trials include the Bonferroni correction, Holm's step-down procedure, the Hochberg step-up procedure, and gatekeeping strategies for hierarchically ordered endpoints. The choice of procedure depends on the comparison structure, the number of arms, and the endpoint hierarchy defined in the protocol.


How do you calculate sample size for a multi-arm trial?

Sample size calculation in a multi-arm trial is more complex than in a standard two-arm design because it must account for multiple simultaneous comparisons, the chosen error rate control strategy, and the definition of trial success.

The key inputs and considerations are:

1. Number of arms and comparison structure
The number of treatment arms and whether the design uses many-to-one or all-pairs comparisons directly affects the number of hypothesis tests being conducted, which in turn affects the multiple testing correction applied and the per-comparison alpha level available.

2. Power definition
Multi-arm trials require a precise definition of what it means for the trial to succeed.

Disjunctive power is the probability that at least one treatment arm demonstrates a statistically significant effect. This is typically used when any successful treatment is a meaningful outcome.

Conjunctive power is the probability that all treatment arms demonstrate a statistically significant effect. This is more conservative and used when all comparisons must succeed for the trial to meet its objectives.

The choice between these definitions has a substantial impact on the required sample size.

3. Multiple testing procedure
The FWER control strategy determines how the overall alpha is allocated across comparisons. More stringent corrections such as Bonferroni reduce the per-comparison significance threshold, which reduces power and requires a larger sample size to compensate.

4. Allocation ratio
Sample size must be distributed across treatment and control arms. Unequal allocation, for example assigning more patients to the shared control arm, can improve overall efficiency, but the trade-offs between allocation ratio and power must be evaluated carefully.

5. Effect size and variability assumptions
As with any sample size calculation, assumptions about the expected treatment effect and outcome variability are required for each comparison. In multi-arm trials, these assumptions may differ across arms.

Getting sample size right at the design stage is critical. Underpowered multi-arm trials waste resources and may fail to detect true treatment effects. Overpowered trials expose more patients than necessary to experimental treatments.


Why are multi-arm trials more efficient than separate two-arm trials?

The core inefficiency of running separate two-arm trials is the duplication of the control group. Each independent trial requires its own control arm, meaning patients are repeatedly allocated to the same comparator across multiple studies. A multi-arm trial eliminates this redundancy by sharing a single control arm across all treatment comparisons within one protocol.

Consider a sponsor evaluating three treatments against a common control. Running three separate two-arm trials requires three independent control groups. A multi-arm trial with a shared control arm can achieve the same set of comparisons with substantially fewer total patients. The control group is recruited once and contributes to all three comparisons simultaneously.

Beyond patient numbers, multi-arm designs reduce inefficiency across the entire development programme:

Time: All treatment comparisons are conducted in parallel rather than sequentially. Data across all arms is generated simultaneously, accelerating portfolio decisions.

Cost: A single protocol, single site network, single data management infrastructure, and single regulatory submission covers all comparisons.

Scientific consistency: All arms are evaluated under identical conditions, covering the same time period, same sites, and same patient population, removing the variability introduced by running comparisons at different times or in different settings.

Patient burden: Fewer total patients are exposed to experimental treatments to generate the same volume of comparative evidence.

The efficiency gains are most pronounced when the number of treatment arms is larger and when the control arm allocation can be optimised across comparisons. However, realising these gains requires careful design. The comparison structure, FWER control strategy, and sample size allocation must all be pre-specified with the same rigour applied to a standard confirmatory trial.


How does the number of arms affect statistical power?

Adding arms to a multi-arm trial introduces a direct tension between the efficiency gains of testing more treatments simultaneously and the statistical cost of controlling for more comparisons.

Each additional arm means one more hypothesis test. To maintain family-wise error rate control across a larger number of comparisons, the significance threshold for each individual comparison must be adjusted downward. Each test must clear a higher bar to be declared significant. This reduction in per-comparison alpha directly reduces statistical power for each individual arm, unless sample size is increased to compensate.

Disjunctive vs. conjunctive power

The relationship between number of arms and overall power also depends on the power definition.

Under disjunctive power, where the trial succeeds if at least one arm is significant, adding more arms can actually increase overall power, since there are more chances for at least one comparison to succeed.

Under conjunctive power, where all arms must be significant, adding more arms generally decreases overall power, since each additional arm is another comparison that must clear the adjusted threshold.

The sample size trade-off

Despite the reduced per-comparison alpha, multi-arm trials often remain more efficient than separate two-arm trials because the shared control arm reduces the total patient burden. The net effect on sample size depends on the number of arms, the allocation ratio between treatment and control, the multiple testing procedure applied, and whether disjunctive or conjunctive power is targeted.

In practice, there is a point at which adding further arms yields diminishing returns. The statistical cost of additional comparisons outweighs the efficiency of the shared control. Evaluating this trade-off explicitly during the design stage, using sample size software capable of modelling multi-arm designs, is an essential step in determining the optimal number of arms for a given trial.


What Other Considerations Shape Multi-Arm Clinical Trial Design?

Beyond comparison structure and sample size, several additional factors influence multi-arm clinical trial design decisions:

  • Control of the family-wise error rate using pre-specified procedures such as Bonferroni, Holm, Hochberg, or gatekeeping strategies.
  • Endpoint hierarchy to establish which comparisons are primary and which are secondary.
  • Adaptive elements that allow pre-specified modifications based on interim analyses.
  • Regulatory alignment with FDA and EMA guidance on multiplicity before the trial begins.

Each of these reflects the broader complexity of designing a multi-arm trial that is both scientifically rigorous and regulatorily defensible.


What Are the Practical Implications for and Statisticians?

The multi-arm clinical trial design environment requires proactive planning and structured implementation from both sponsors and statisticians.

Key implications include:

  • Early integration of multi-arm design considerations into protocol development.
  • Clear statistical justification for the chosen comparison structure and FWER control strategy.
  • Transparent documentation of the multiple testing procedure aligned with regulatory guidance.
  • Strategic alignment between scientific objectives, sample size constraints, and operational feasibility.

Trial teams must ensure that the efficiency gains of multi-arm clinical trial design enhance, rather than complicate, regulatory acceptability and scientific credibility.

 

About nQuery
nQuery helps make your clinical trials faster, less costly and more successful.
It is an end-to-end platform covering Frequentist, Bayesian, and Adaptive designs with 1000+ sample size procedures. 

nQuery Solutions
Sample Size & Power Calculations
Calculate for a Variety of frequentist and Bayesian Design

Adaptive Design
Design and Analyze a Wide Range of Adaptive Designs

Milestone Prediction
Predict Interim Analysis Timing or Study Length

Randomization Lists
Generate and Save Lists for your Trial Design

Who is this for?

This will be highly beneficial if you're a biostatistician, scientist, or clinical trial professional that is involved in sample size calculation and the optimization of clinical trials in:

 

  • Pharma and Biotech
  • CROs
  • Med Device
  • Research Institutes
  • Regulatory Bodies
Share on Twitter
Share on LinkedIn

Get started with nQuery today

Try for free and upgrade as your team grows