Showing posts with label group sequential design. Show all posts
Showing posts with label group sequential design. Show all posts

Friday, March 03, 2023

Risk of Stopping Trial Early for Treatment Benefit - Story of Veru's Sabizabulin for Treatment of Covid

Today, we learned that FDA declined to approve the emergency use authorization (EUA) for Veru's sabizabulin for the treatment of Covid. The decision is not surprising given that in Nov 9, 2022, the FDA advisory committee voted 8-5 against the sabizabulin treatment aimed at hospitalized patients with moderate to severe infection who are at high risk for acute respiratory distress syndrome.

However, this case is a perfect example demonstrating the risk of stopping a clinical trial early for treatment benefit (or for overwhelming efficacy). 

The pivotal study was an international, multicenter, randomized, double-blind, placebo-controlled, parallel group study of 60 days duration that evaluated the efficacy and safety of VERU-111 (sabizabulin) in hospitalized adult subjects with COVID-19 infection described as being at “high risk for ARDS”. Subjects were included if they met criteria for World Health Organization (WHO) ordinal scale category 5 (non-invasive ventilation or high-flow oxygen) or category 6 (mechanical ventilation); or category 4 (oxygen by mask or nasal prongs) with the following comorbidities: asthma, chronic lung disease, diabetes, hypertension, severe obesity (BMI ≥40), 65 years of age or older, primarily residing in a nursing home or long-term care facility, immunocompromised. The primary efficacy endpoint is all-cause mortality at Day 60. 

According to the study protocol, approximately, 300 subjects were planned to be randomized at a 2:1 ratio into two treatment arms (200 subjects in the VERU-111 treated group and 100 subjects in the Placebo treated group). The protocol also pre-specified a formal interim analysis conducted by an independent data monitoring committee:

An efficacy interim analysis of the primary endpoint, all-cause mortality at Day 60, will be conducted when approximately 67% of the subjects (~200 subjects) have completed day 60, died or withdrawn for other reasons. The interim and final analyses will follow an O'Brien-Fleming group sequential design, with a plan of stopping the trial at the interim analysis if the results are statistically significant in favor of VERU-111. The criterion for efficacy at the interim analysis will be a two-sided 0.0121 p-value (a one-sided p-value <=0.0061 in favor of VERU-111). If the criterion is not met, the trial will continue, and the final analysis will have a criterion for efficacy of a two-sided p-value <=0.0463 (one-sided 0.0232 in favor of VERU111).

The formal interim analysis was conducted after 204 patients were randomized into the study and followed up to 21 days (134 patients in Sabizabulin group and 70 patients in Placebo group). Sabizabulin treatment resulted in a 24.9 percentage point absolute reduction and a 55.2% relative reduction in deaths compared with placebo (odds ratio, 3.23; 95% CI confidence interval, 1.45 to 7.22; P=0.0042). The mortality rate was 20.2% (19 of 94) for sabizabulin versus 45.1% (23 of 51) for placebo.

Given that the p-value 0.0042 is less than the criterion for stopping the trial for efficacy (0.0121), the sponsor (Veru) stopped the study and declared the overwhelming efficacy in the hope of FDA's approval for emergency use authorization. 

The interim analysis results were then published in the top medical journal - New England Journal of Medicine "Oral Sabizabulin for High-Risk Hospitalized Adults with Covid-19: Interim Analysis".

The original sample size of 300 patients was already pretty small in trials for the treatment of COVID-19. The sample size was further decreased due to the early stop for 'overwhelming' efficacy. During the FDA's review, the small sample size was one of the big issues. FDA expressed their concern about the evidence for supporting the benefit outweighing the risk in their briefing document provided for the advisory committee:  

The FDA Review team acknowledges that Study 902 met its prespecified primary endpoint of all-cause mortality at Day 60. We also note that the VERU-111 program is quite small in size compared to other therapeutic programs for patients hospitalized with COVID-19. As detailed in the briefing document, our review has identified a number of uncertainties with the data, which we raise in the context of this small trial in critically ill patients. These include:
  • High placebo group mortality rate
  • Potential for unblinding events with enteral tube administration
  • Baseline imbalances in standard of care therapies
  • Differences in hospitalization duration prior to trial enrollment
  • Uncertain effects of goals of care decisions on all-cause mortality
  • Negative studies with other microtubule disruptors in COVID-19 
  • Uncertainty in identification of a clinically relevant patient population 

Based on our review, none of these uncertainties or imbalances alone invalidate the mortality benefit observed in Study 902, but all of these issues together in a small trial which is more vulnerable to imbalances raise questions about the results. We conducted sensitivity analyses to investigate the potential impact of the noted imbalances. However, these analyses cannot eliminate the concern that certain baseline imbalances across treatment groups may have impacted study outcomes, due to the small study size. In addition, these issues raise concern that, even when using an objective endpoint such as mortality, observed results can be subject to biases in a small trial of short duration in critically ill patients. We ask the AC panel to consider these uncertainties together and how they affect the interpretation of the mortality data. 

When considering whether to authorize the emergency use of a product under EUA, the Agency must determine, among other requirements (see section 2.1.3), whether “the known and potential benefits of the product when used to diagnose, prevent, or treat the identified serious or life-threatening disease or condition, outweigh the known and potential risks of the product”. Evaluation of the potential risks in the VERU-111 development program is limited by the atypically small safety database, comprising a total of 149 subjects who received VERU-111 for the proposed use in COVID-19. While the small safety database limits our ability to identify clinically significant safety signals, potential safety signals identified in our review are urinary tract infections (including serious infections), ......

There is no guarantee that Veru's EUA application would be approved had the study not been stopped early. But I think that the chance of EUA approval would be much higher if the study was allowed to be completed according to the protocol and 100 additional patients were randomized into the study. To decide whether a study should be stopped early for efficacy, the stopping criterion is not the only factor to consider. I wished that the sponsor had consulted with the regulatory authorities before the decision to stop the study early for efficacy. 

On a separate note, FDA's "Good Review Practice: Clinical Review of Investigational New Drug Applications" spelled out their concerns about early stopping for efficacy:



Thursday, December 01, 2022

Sample size re-estimation or sample size increase?

Recently, a press release from a biotech company caught my eye. This seems to be an example of the adaptive design with sample size re-estimation, however, it is unusual that the sample size is decreased as usually the sample size re-estimation results in an increase in sample size. 
Bellerophon Therapeutics, Inc. (Nasdaq: BLPH) (“Bellerophon” or the “Company”), a clinical-stage biotherapeutics company focused on developing treatments for cardiopulmonary diseases, announced today that the U.S. Food and Drug Administration (FDA) has accepted the Company’s proposal to reduce the study size for its ongoing registrational REBUILD Phase 3 trial of INOpulse® for the treatment of fibrotic Interstitial Lung Disease (fILD). The new study size of 140 subjects does not impact the trial’s principal objective or endpoints and maintains power of >90% (p-value < 0.01) for the primary endpoint of Moderate to Vigorous Physical Activity (MVPA) based on the effect size observed in Phase 2.

Following the evaluation of baseline MVPA characteristics, as measured by actigraphy, compliance to treatment and review of safety data of the randomized subjects in the ongoing Phase 3 REBUILD study, the trial’s independent Data Monitoring Committee (DMC) supported reducing the target study size from 300 to 140 subjects.
Sample size re-estimation is one type of adaptive design where the sample size can be adjusted during the study based on a prespecified rule. Sample size re-estimation has its special features:  

Group Sequential Design (GSD) and Sample Size Re-estimation

Clinical trials with adaptive design can be in different forms depending on what the adaptations are. Two commonly utilized adaptive designs are group sequential design (GSD) and sample size re-estimation (SSR). Implementation of both GSD and SSR is through the interim analyses conducted by the independent data monitoring committee. In GSD studies, we set a large sample size and hope to stop the trial early due to the overwhelming efficacy, futility, or safety at the interim analyses. In adaptive design with SSR, we start with a small study and possibly increase the sample size post an interim analysis. Both GSD and SSR can achieve the same benefits of reduced sample size and potentially an earlier conclusion. 

Blinded Sample Size Re-estimation and Unblinded Sample Size Re-estimation

In FDA guidance "Adaptive Designs for Clinical Trials of Drugs and Biologics", sample size re-estimation was described in section B "adaptations to the sample size". Blinded sample size re-estimation is based on interim estimates of nuisance parameters such as the standard deviation for continuous outcome measure and overall event rate for discreet outcome measure. The unblinded sample size re-estimation is a type of adaptive design where adaptation is to prospectively plan modifications to the sample size based on comparative interim results. Blinded sample size re-estimation may be conducted by the sponsor statistician while unblinded sample size re-estimation must be through an independent data monitoring committee.  

Sample Size Re-estimation and Sample Size Increase

In clinical trials with prospectively planned sample size re-estimation, the sample size is usually increased. It is very rare that the sample size is decreased after the interim analysis. For adaptive clinical trials with adaptation on sample size (i.e., sample size re-estimation), the initial sample size estimation can be based on more aggressive assumptions that result in a smaller sample size. In the middle of the study, interim analyses are performed and the decision can be made (by independent DMC and through a prespecified rule) whether or not the sample size should be increased. 

In FDA's guidance discussing the Adaptations to Sample Sizes, while the terms 'sample size re-estimation' and 'sample size adaptation' are used, the sample size increase is really implied. 

Sample Size Adaptation and Sample Size Increase by a Fixed Number

The sample size re-estimation or sample size adaptation is really a binary decision. If the decision is to increase the sample size (after the interim analysis), the sample size will be increased by a pre-specified, fixed number, not increased by a number that is based on the observed treatment effect at the interim analysis. 

If the sample size is increased by a very exact level calculated from the observed treatment effect at the interim analysis to bring the conditional power up to a target level, there is a potential to reverse calculate the effect size or to at least make an educated guess about what the effect size is from the interim analysis

This potential for an educated guess about the effect size is a huge issue from the regulatory point of view. This specific concern is discussed in FDA's guidance ""Adaptive Designs for Clinical Trials of Drugs and Biologics".

Finally, there are additional challenges in maintaining trial integrity in the presence of sample size adaptations. For example, sample size modification rules are often based on maintaining the conditional probability of a statistically significant treatment effect at the end of the trial (often called the conditional power) at or near some desired level. In this scenario, knowledge of the adaptation rule and the adaptively chosen sample size allows a relatively straightforward back-calculation of the interim estimate of treatment effect. Therefore, additional steps should be taken to limit personnel with this detailed knowledge so that trial integrity can be maintained.