Showing posts with label futility analysis. Show all posts
Showing posts with label futility analysis. Show all posts

Saturday, February 15, 2025

Sponsor’s Response to DMC Recommendation: A Case Study from a Phase 2b/3 Trial

 The DMC or DSMB is now commonly used in the clinical trials, especially the late phase clinical trials. According to FDA guidance for industry "Use of Data Monitoring Committees in Clinical Trials", DMC Responsibilities include:

1.             1.    Monitoring of Trial Conduct

2.      Monitoring of Results of Interim Analysis of Trial Data

·         Safety – to determine if there is a credibly increased risk of a serious adverse outcome in subjects receiving the investigational product, indicating that enrollment should be stopped. To determine a safety risk, review of unblinded efficacy data should also be conducted by the DMC as they evaluate a benefit-risk assessment

·         Implementing a predefined adaptive feature

                                                              i.      Efficacy – to determine if there is statistically significant evidence of efficacy such that enrollment should be stopped

                                                            ii.      Futility – to determine if there is no longer a reasonable likelihood that the trial will reach a conclusion of effectiveness, so that enrollment should be stopped to protect subjects from further exposure to a potentially ineffective investigational product and to conserve resources

                                                          iii.      Other adaptations – a DMC or a separate adaptation committee should determine if a prespecified adaptive aspect of the trial design is to be implemented. This can include modifying the sample size, changing a randomization ratio, or restricting future enrollment to a prespecified subgroup (adaptive enrichment

3.    Consideration of External Data

4.    Recommendations and Documentation

When a DMC is established, a DMC charter will be established to describe DMC Obligations, Responsibilities, and Standard Operating Procedures. DMC charter may also specify if there is any stopping rules to implement, decision trees to be followed, and any adaptation rules to be implemented. 

DMC communicates with the sponsor through the DMC recommendations. The FDA guidance has the following about the DMC recommendations:

A fundamental responsibility of a DMC is to make recommendations to the sponsor concerning the continuation of the trial.  Most frequently, a DMC’s recommendation after an interim review is for the trial to continue as designed.  Other less frequent but possible recommendations, however, as discussed previously, include trial termination, trial continuation with major or minor modifications (such as implementation of prespecified adaptive elements), or temporary suspension of enrollment and/or trial intervention until an identified uncertainty is resolved. 

A DMC should express its recommendations clearly to the sponsor because a DMC’s actions potentially affect the safety of trial subjects.  Both a written recommendation and an oral communication, with opportunity for questions and discussion, can be valuable.  Recommendations for modifications are best accompanied by the minimum amount of data critical for the sponsor to make a reasonable decision about the recommendation, and the rationale for such recommendations should be as clear and precise as possible.  Sponsors may wish to develop internal procedures to limit the interim data released by a DMC after a recommendation and until a decision is made regarding acceptance or rejection of the recommendation in order to help maintain confidentiality of the interim results should the trial continue.  We recommend that a DMC document its recommendations and rationale in a manner that can be reviewed by the sponsor and then circulated, as appropriate, to IRBs, FDA, and/or other interested parties, when based on interim data.  Major trial changes—such as early trial termination, change in population or entry criteria, or change in trial endpoints—can have substantial impact on the validity of the trial and/or its ability to support the desired regulatory decision.  Sponsors should discuss with FDA any proposed protocol changes based on review of interim data that were not planned for, before implementation, and submit such changes to FDA in accordance with 21 CFR 312.30 and 812.35.  However, if the sponsor learns of information that presents an imminent safety hazard to trial participants, sponsors should implement the necessary changes as quickly as possible to ensure the safety and welfare of study subjects (see 21 CFR 312.30(b)(2)(ii) and 812.35(a)(2)). 

In most of situation, the study is as expected and it is easy for DMC to make a recommendation of no changes to the study. However, in complicated situation, the DMC needs to make tough decision and recommend the termination of the study. In a paper by Wittes et al "The Data Monitoring Committee: A Collective or a Collection?", the following suggestions of consensus operating were made:

In a typical DMC meeting, data emerge as expected. No worrisome safety concern arises; the efficacy data are not surprising; and the DMC deems that trial is progressing as planned with, perhaps, some lag in recruitment and a less than desirable rate of follow-up of participants and incomplete capture of important efficacy and safety data. These and other quality metrics affect decision-making and the DMC may discuss them with study leadership. When, however, evidence of an unexpected harm arises, or the study operations appear unacceptable, or efficacy appears much different from anticipated, the deliberations of the DMC may reveal initial, perhaps strong, differences of opinion. A requirement to vote may curtail discussion and may lead to the failure to produce a recommendation that all find acceptable. Instead, we agree with those who urge DMCs to operate by consensus . Operating by consensus means that a DMC can have an odd or even number of members. Prior to reaching consensus, the Chair may elicit the opinion of each member to gauge the general views of members of the DMC or even call an informal straw vote. Regardless of how the DMC reached consensus, all members should agree to the language summarizing its recommendations....

This week, we saw an interesting example how the DMC recommendation was received and handled by the sponsor. Apparently, the sponsor did not trust the DMC's recommendation of stopping the study. The sponsor is now assembling an expert panel to review the unblinded data to determine how the DMC's recommendation is made and what are the rationales for DMC's recommendation.

Pliant Therapeutics announcedthat their phase 2b/3 study in IPF was suspended per DMC’s recommendation.  

          Pliant Brings in Outside Experts to Review IPF Study Pause 

          Pliant Therapeutics  has initiated assembly of outside panel of world-renowned experts to review 

          BEACON-IPF trial dataAnnounces Next Steps Following DSMB

SOUTH SAN FRANCISCO, Calif., Feb. 13, 2025 (GLOBE NEWSWIRE) -- Pliant Therapeutics, Inc. (Nasdaq: PLRX) today announced that, per the charter of the trial’s independent Data Safety Monitoring Board (DSMB), the Company has initiated the assembly of an outside expert panel to review unblinded data from the ongoing BEACON-IPF Phase 2b trial of bexotegrast in patients with idiopathic pulmonary fibrosis (IPF). The panel, consisting of world-renowned experts in pulmonary diseases and biostatistics, will provide an independent recommendation to Pliant regarding the BEACON-IPF trial. Subsequently, the panel will serve as part of an expanded DSMB with the goal to reach a consensus recommendation regarding BEACON-IPF. The decision to assemble the outside panel was taken as the Company has not been able, through review of blinded data, to determine the rationale for the DSMB’s recommendation to pause enrollment and dosing in the trial. The Company expects this process to conclude in two to four weeks.

Following the DSMB’s previously announced recommendation, Pliant voluntarily paused enrollment and dosing in the BEACON-IPF clinical trial. Pliant is committed to remaining blinded ensuring the data integrity of the BEACON-IPF 2b clinical trial with the goal of maintaining its potential to serve as a registrational trial.

It is very likely that the DMC focuses on the safety review and follows the FDA guidance which says:

Safety – to determine if there is a credibly increased risk of a serious adverse outcome in subjects receiving the investigational product, indicating that enrollment should be stopped. To determine a safety risk, review of unblinded efficacy data should also be conducted by the DMC as they evaluate a benefit-risk assessment

There may be an imbalance in the number of deaths or serious adverse events and there may no clear indication of efficacy. In such cases, the DMC's recommendation to stop the trial is based on a careful benefit-risk assessment. 

Sunday, January 09, 2022

Overrunning issue at the interim analysis for group sequential design

It is pretty common these days that clinical trials (especially the late phase, adequate, and well-controlled studies) employ interim analyses to determine if the efficacy results are too good so that the study should be stopped early for overwhelming efficacy, or if the efficacy results are not good so that the study should be stopped early for futility, or both. A study with formal interim analyses to look at the comparative efficacy is called 'group sequential design' even though the 'group sequential design' may not be formally used in the study protocol. Group sequential design is the most common type of adaptive design as described in the FDA guidance "adaptive designs for clinical trials". 

As mentioned in an early post "overrunning issues in adaptive design clinical trials", one of the issues with interim analyses in group sequential design is the overrunning issue. Overrunning consists of extra data, collected by investigators while awaiting results of the interim analysis (IA). Overrunning is the 
phenomenon that data will continue to accumulate after it is decided to stop a trial (Whitehead, 1992). In many cases there will be patients who have already been admitted to the trial but whose responses are not yet known. Also some extra patients will enter the trial because of the delay between the moment the data for the final interim analysis were retrieved, and the moment participating clinical centers receive instruction to stop recruitment.

EMA guidances "reflection paper on methodological issues in confirmatory clinical trials planned with an adaptive design" had a full paragraph discussing the overrunning issue: 


When planning for an interim analysis, the decision needs to be made about the timing of the interim analysis and what data is to be included in the interim analysis. For an event-driven study where the number of events is the endpoint, the timing of the interim analysis can be based on the percentage of the events, for example, the interim analysis can be performed when 50% of the total number of events are accrued. In other words, once 50% of the total number of events is achieved. For a longitudinal design, the endpoint is measured at various intervals. At any time during the study, there will be patients at different stages of the study (reaching the end of the study, reaching a specific duration of the study, or just being randomized into the study). It is more difficult to determine a good timing for interim analysis. Suppose that an interim analysis is planned when 50% of subjects reach the end of the study, by that time, there will be plenty of subjects in the various stage of the study already, just having not reached the end of study yet. If the study enrollment is pretty fast and the study endpoint is pretty long (for example 52 weeks), by the time 50% of subjects reach the study endpoint, the majority of subjects (if not all) have already been randomized into the study. 

If the interim analysis results trigger the recommendation of discontinuing the study early (either for efficacy or futility), the debate is whether the interim analysis needs to be re-run by including the overrunning subjects before adopting the recommendation to discontinue the study early. 

This exact issue about the handling of the overrunning subjects was discussed in Biogen's aducanumab clinical trials in Alzheimer's disease. As I discussed in a previous post "Futility Analysis and Conditional Power When Two Phase 3 Studies are Simultaneously Conducted", Biogen made the wrong decision and discontinued its pivotal studies (EMERGE and ENGAGE) where one of the studies was later found to have statistically significant treatment effects. The wrong decision was driven by two issues: 

  • they calculated the conditional powers based on the pooled data from both studies (instead of calculating the conditional powers separately based on the data from individual studies) - this was discussed in the previous post
  • they stopped the study early without re-running the interim analysis by including the overrunning subjects. 

The overrunning issue was mentioned in a recent article in the Wall Street Journal (Jan 4, 2022) "How Biogen Fumbled Its Alzheimer's Drug ---Once-promising Aduhelm is pricey and without proven efficacy"

"By evaluating data midstream in approval-seeking trials, companies can try to predict whether a drug will succeed if the trial continues. Stopping trials early for "futility," in industry parlance, can save millions of dollars and prevent patients from investing hope on an ineffective drug.

In a March 2019 meeting, Biogen executives on a small "senior decision team," as the company called it, concluded that the trials were doomed. Biogen pulled the plug and asked researchers around the world to shut down trials. It told more than 3,000 Alzheimer's patients who had volunteered that they would no longer receive treatment.

Biogen stock fell by nearly 30% the day of the announcement.

Biogen executives made errors in shutting down the trials. The trial plan called for analyzing data after half of patients completed the study treatment in late December 2018. By the time Biogen completed the analysis in March 2019, three more months of additional data were available -- but the decision team didn't scrutinize the additional data before the company halted the trials, Biogen has said.

A Biogen consultant in the summer of 2018 recommended to senior Biogen statisticians that they consider all available trial data, according to a person involved in the process. The consultant cautioned them that a plan to leave out consideration of additional trial data after the cutoff date -- and to leave out certain data from patients in the trial before the cutoff -- would open up Biogen to criticism and scrutiny, the person said. The statisticians didn't heed the consultant's advice, and it isn't clear whether the decision team or management considered the recommendation, the person said.

Biogen declined to comment on past discussions with its consultants but said it followed its pre-established statistical-analysis plan.

The decision not to consider the three months of additional data was a misstep, said some clinical-trial experts and statisticians. "Additional data after a study stops is called 'overrunning.' We plan for it," said Scott Emerson, a professor emeritus of biostatistics at the University of Washington who served on an FDA advisory committee that recommended against approving Aduhelm in November 2020. "In this case, the overrunning data was large."

The Biogen spokeswoman said: "Our decision to stop the trials, though clearly incorrect in hindsight, was based on putting patients at the forefront -- as it always should be. Cost was not considered in determining futility."

Only in the weeks after the trials stopped did Biogen scientists complete a preliminary analysis of the overrunning data and recognize their mistake, the Biogen spokeswoman said. The data seemed to show that one of the trials would have produced positive results, despite the likelihood of a negative outcome in the second trial. Initially, Biogen had analyzed combined data from both trials."

Even though the aducanumab was finally approved by FDA through the accelerated approval pathway, the approval was very controversial - the first drug and the only drug so far that was approved based on studies that had been stopped early for futility. There are strong pushback from academic about the use of aducanumab because of its unapproved efficacy and perhaps also because of the drama of resurrecting a drug that was declared failed. 

We all wonder: had these two pivotal studies not been stopped for futility by the sponsor, what would be the situation now for aducanumab? 

Saturday, January 01, 2022

Futility Analysis and Conditional Power When Two Phase 3 Studies are Simultaneously Conducted

In late-phase clinical trials, an independent Data Monitoring Committee (DMC) is usually set up. If the clinical program includes multiple late-phase studies, the same DMC will be responsible for the entire program. With DMC, the interim analyses can be performed for different purposes:
  • The interim analysis for safety
    • with pre-specified stopping rule (for example stop the trial if the significant imbalance in # of Serious Adverse Events or in # of deaths)
    • without pre-specified stopping rule (rely on DMC members to review the overall safety)
  • The interim analysis for efficacy: To see if the new treatment is overwhelmingly better than the control group  - then stop the trial for efficacy
  • The interim analysis for futility (futility analysis): To see if the new treatment is unlikely to be better than the control group or the study will be unlikely to achieve its objective given the data at the interim – then stop the trial for futility.
There seem to be more studies with built-in futility analysis without interim analysis for overwhelming efficacy, mainly because of the concerns about the alpha-spending for efficacy. The futility analysis will have an impact on the beta-spending and the statistical power, but not on the alpha-spending. For the decision-making, regulatory agencies are usually more concerned about the alpha level (incorrectly approves a drug that does not work) or the alpha level inflation. The sponsors are more concerned about the statistical power (incorrectly concludes a drug not working while the drug is actually working).

Futility analysis usually requires calculating the Conditional Power (CP) that is defined as the probability that the final study result will be statistically significant, given the data observed thus far at the time of the interim data cut and a specific assumption about the pattern of the data to be observed in the remainder of the study, such as assuming the original design effect (alternative hypothesis) or the effect estimated from the interim data.  

If there is one single pivotal trial, the stopping rule and the CP are relatively straightforward. However, it is uncommon that the sponsor may need to conduct two pivotal (phase 3) studies (two adequate and well-controlled (A&WC) trials in FDA's term) to demonstrate substantial evidence of effectiveness as outlined in FDA guidance for industry "Demonstrating Substantial Evidence of Effectiveness for Human Drug and Biological Products Guidance for Industry".

For a clinical program with two independent A&WC trials (usually with identical design), the futility analysis and CP calculation are a little bit more complicated. Two independent A&WC trials may have an identical design but be executed differently (i.e., may not be started at the same time; may be conducted in different geographic regions/countries; and may have different enrollment speeds,...). 

When futility analysis is performed for two A&WC trials, should the conditional powers be calculated for individual studies separately or should the conditional powers be calculated for both studies together (i.e. pooled data from both studies)? 

When there are two identical A&WC trials, the interim analysis for safety should be based on the pooled data sets from both studies because it will give a more definitive answer to the safety issues, the interim analysis for efficacy should be based on the individual study data because the decision about the overwhelming efficacy should be based on the individual study, not the integrated data from two studies; the interim analysis for futility is a little bit more complicated and the decision to use the data from an individual study or to use the data from the pooled data seems to be dependent on how close the observed results from two A&WC trials are at the time of the interim analysis. 

For futility analysis using stochastic curtailment procedure, While CPs can be calculated for each individual study assuming that the treatment effect in the remaining subjects in the same study will follow the treatment effect estimated from the data of this same study at the time of the interim data cut, 

There is an alternative way to calculate the CP, i.e., to calculate the CP for each individual study, but use the observed treatment effect from the pooled data at the interim from both studies to project the trend and pattern for the remaining subjects. 

According to the paper by Lan and Wittes (1988) "The B-Value: A Tool for Monitoring Data", the CP calculation involves the decomposition of overall critical value (B-value or B1 for example) into the sum of two statistically independent interval B-values: 
  • Bt, the value of B that accumulated up through time t when interim analysis is conducted; and 
  • (B1 - Bt), the incremental value of B that accumulates from time t through the end of the study. The legitimacy of the decomposition follows from the independence of distributions of the outcomes for successive study subjects
At the time t when the interim analysis is conducted, Bt is known and is estimated from the observed data up to the time t. (B1 - Bt) is a random variable that needs to be estimated. The conditional power is derived by fixing Bt and calculating the probability that Bt + (B1 - Bt) will exceed Z1-a/2.

To calculate the CPs when there are two identical A&WC studies, t, as a measure of the information fraction, will be different for different studies. At the time t, maybe 60% of subjects have been enrolled in study #1 while 50% of subjects are enrolled in study #2. In CP calculations, the Bt part will be obtained from the individual study. The (B1-Bt) part is estimated assuming the remaining data following the observed effect up to the interim time t, should the observed effect up to the interim time t be based on the data from the individual study or from the pooled data?

It turns out both approaches can be used: 
  • estimate the treatment differences for each individual study and calculate the CP assuming that the reminding data follows the trend and pattern based on the observed data from individual study
  • estimate the treatment difference from both studies and calculate the CP assuming that the remaining data follow the trend and pattern based on the observed data from the pooled data of two studies.           
For both of these approaches, the CPs will be calculated for each individual study (therefore one CP for each study). The difference between these two approaches is in the calculation of the (B1-Bt) part - based on the individual study itself or based on the pooled data from both studies. 

We can take a look at the famous and controversial case in Biogen's aducanumab program in Alzheimer's disease. Aducanumab program in Alzheimer's diseases consisted of two pivotal, phase 3 studies (EMERGE (study 301) and ENGAGE (study 302)), and both studies were designed the same and conducted simultaneously globally. Each study had two active arms (low dose and high dose of aducanumab) versus placebo - therefore two hypothesis tests (low dose vs. placebo and high dose vs. placebo). There was a total of four hypothesis tests (two for each study).  The protocol and SAP specified the interim analysis for futility. 

An interim analysis was performed after approximately 50% of the subjects had the opportunity to complete the Week 78 visit for both EMERGE and ENGAGE studies. An interim analysis for the futility of the primary endpoint was performed to allow early termination of the studies if it was evident that the efficacy of aducanumab was unlikely to be achieved. The futility criteria were based on conditional power, which was the chance that the primary efficacy endpoint analysis would be statistically significant in favor of aducanumab at the planned final analysis, given the data at the interim analysis. The CP was calculated assuming that the future unobserved effect was equal to the maximum likelihood estimate of what is observed in the interim data. 

For each study, two CPs were calculated. The pre-specified CP calculation was to use the pooled interim data from both EMERGE and ENGAGE studies for the (B1-Bt) part and assume that the treatment effect for the remaining of the study would follow the observed treatment effect at the interim analysis. At the interim analysis, the CPs were calculated to be 13% for low dose vs placebo and 0% for high dose vs. placebo in EMERGE study, and 11% for low dose vs placebo and 12% for high dose vs. placebo in ENGAGE study. Given all four CPs were lower than the threshold of 20% (a criterion for futility), the DMC recommended stopping both studies for futility.  Biogen followed the DMC recommendation and stopped both EMERGE and ENGAGE studies for futility
.

Only after two terminated studies were wrapped up, the reanalyses of the final data indicated that there were statistically significant treatment differences in one of the studies (the ENGAGE study). With the help of the FDA, Biogen was able to submit the BLA and obtain approval for aducanumab for Alzheimer's disease. Leading to the FDA approval, there was an advisory committee meeting to review the aducanumab data. In FDA's presentation, the conditional powers were retrospectively re-calculated - this time, the conditional powers were calculated for each individual study and assumed future unobserved effect would be similar to the interim data for each individual study (not the pooled interim data). FDA claimed that CPs using this approach were more appropriate and would have one of the four CPs above the threshold of 20% (CP=59% for high-dose vs placebo in ENGAGE study) - the studies would not be recommended for stopping for futility. 


Retrospectively, CPs calculated for each study independently (not using the pooled interim data to project the trend and pattern for the remaining data) seemed to be better in Biogen aducanumab program consisting of two A&WC trials. 

However, in a paper by Deng et al "Superiority of combining two independent trials in interim futility analysis", CP calculation using the observed treatment effects from the pooled interim data from two studies was considered a better approach. It concluded, "it is demonstrated that by leveraging data from the other study, the probability of making correct interim decision is increased if the treatment effects are similar between the two studies, and such benefit remains even if there is small to moderate between-study difference."

It is probably true that CP calculation using the pooled data at the interim to project the trend and pattern for the remainder data is a better approach if two studies are conducted in the same way and the results at the time of the interim analysis are similar. However, the CP calculation and the statistical analysis plan for interim analysis are usually pre-specified before seeing the unblinded data. At the time of the interim analysis, it is usually unknown whether or not the results (treatment effects) observed from two identical studies will be similar. Even though two A&WC studies are designed the same, the operation and execution of the trial can still be different: two studies may be conducted in different countries, enrollment speed may be different,... As evidenced by Biogen's EMERGE and ENGAGE trials, two identical designed studies may have different results - therefore calculating the CP entirely independently for each study may be more appropriate when two identical A&WC trials are conducted.   

Monday, December 27, 2021

Futility Analysis and Conditional Power

Adaptive design has been used to drug development programs more efficient. According to FDA's guidance Adaptive Designs for Clinical Trials of Drugs and BiologicsGuidance for Industry, an adaptive design is defined as a clinical trial design that allows for prospectively planned modifications to one or more aspects of the design based on accumulating data from subjects in the trial. The modifications to the design based on the accumulating data from an ongoing study are through 'interim analysis'. An interim analysis is any examination of data obtained from subjects in a trial while that trial is ongoing and is not restricted to cases in which there are formal between-group comparisons. The observed data used in the interim analysis can include one or more types, such as baseline data, safety outcome data, pharmacokinetic, pharmacodynamic, biomarker data, or efficacy outcome data.

when an adaptive design is proposed, which aspect(s) of the trial to be adapted will need to be pre-specified and agreed upon by the regulatory agencies such as FDA. In the list of adaptations, the most common type of adaptive design is 'group sequential design'. 

    • Group sequential design
    • Adaptations to the sample size
    • Adaptations to the patient population (e.e., adaptive enrichment)
    • Adaptations to treatment arm selection
    • Adaptations to patient allocation 
    • Adaptations to endpoint selection
    • Adaptations to multiple design features

Group sequential design is probably the most commonly used adaptive design (even before the adaptive design concept came out). Group sequential design was once categorized as 'well-understood' adaptive design. Ironically, many studies with group sequential design may not be called 'adaptive design' and the term 'group sequential design' may not be used in the study protocol at all. 

According to FDA's Adaptive Designs for Clinical Trials of Drugs and Biologics Guidance for Industry

"Group sequential designs may include rules for stopping the trial when there is sufficient evidence of efficacy to support regulatory decision-making or when there is evidence that the trial is unlikely to demonstrate efficacy, which is often called stopping for futility."

"There are a number of additional considerations for ensuring the appropriate design, conduct, and analysis of a group sequential trial. First, for group sequential methods to be valid, it is important to adhere to the prospective analytic plan and terminate the trial for efficacy only if the stopping criteria are met. Second, guidelines for stopping the trial early for futility should be implemented appropriately. Trial designs often employ nonbinding futility rules, in that the futility stopping criteria are guidelines that may or may not be followed, depending on the totality of the available interim results. The addition of such nonbinding futility guidelines to a fixed sample trial, or to a trial with appropriate group sequential stopping rules for efficacy, does not increase the Type I error probability and is often appropriate. Alternatively, a group sequential design may include binding futility rules, in that the trial should always stop if the futility criteria are met. Binding futility rules can provide some advantages in efficacy analyses (e.g., a relaxed threshold for a determination of efficacy), but the Type I error probability is controlled only if the stopping rules are followed. Therefore, if a trial continues despite meeting prespecified binding futility rules, the Agency will likely consider that trial to have failed to provide evidence of efficacy, regardless of the outcome at the final analysis. Note also that some DMCs might prefer the flexibility of nonbinding futility guidelines."

With group sequential design, interim analyses will be performed during the study to evaluate early evidence of efficacy or early evidence of futility. To stop the study for efficacy, the most common approach is so-called 'repeat significance testing'. to stop the study for futility, the most common approach is through calculating the conditional power.

Group sequential design:
  • The interim analysis for efficacy: To see if the new treatment is overwhelmingly better than control - then stop the trial for efficacy
    • Repeat significance testing
      • Pocock
      • O'Brien-Fleming
      • Alpha-spending by Lan and DeMets 
  • The interim analysis for futility (futility analysis): To see if the new treatment is unlikely to be superior to the control – then stop the trial for futility - this is called ‘futility analysis’.
    • Repeat significance testing
    • Stochastic curtailment approach with three families of stochastic curtailment tests
      • Conditional power tests (frequentist approach)
      • Predictive power tests (mixed Bayesian-frequentist approach).
      • Predictive probability tests (Bayesian approach)

The most common futility analysis requires the calculation of the conditional power (CP) that is the probability that the study will demonstrate statistical significance at the end of the study (i.e. final analysis to claim superiority), conditioning on the data observed in the study thus far, and an assumption about the trend of the data to be observed in the remainder of the study. 

According to the paper by Lachin "A review of methods for futility stopping based on conditional power":
“Conditional power (CP) is the probability that the final study result will be statistically significant, given the data observed thus far and a specific assumption about the pattern of the data to be observed in the remainder of the study, such as assuming the original design effect, or the effect estimated from the current data, or under the null hypothesis.”

In conditional power calculation, assumptions about the trend of the data in the remainder of the study can be as the following and the assumption of the remainder data following the observed data is probably more reasonable. The assumption of the remainder data following the alternative hypothesis can overestimate the overall treatment effect (especially if the alternative hypothesis was based on the aggressive, over-optimistic assumptions) resulting in inflated conditional power. On the other hand, the assumption of the remainder data following the null hypothesis can underestimate the overall treatment effect resulting in deflated conditional power.  

  • Observed data - the effect estimated from the current data so far
  • The alternative hypothesis - assuming the original design effect
  • The null hypothesis - assuming no effect in the remainder of the study 

To summarize, the futility analysis is through interim analysis to determine if the trial data indicates the inability of a clinical trial to achieve its objectives. Futility analysis usually requires the calculation of the conditional power (CP) that is defined as the probability that the final study result will be statistically significant, given the data observed thus far at the time of the interim data cut and a specific assumption about the pattern of the data to be observed in the remainder of the study, such as assuming the original design effect (alternative hypothesis) or the effect estimated from the current data. It is pretty common that the threshold for futility is defined as CP less than 20% - suggesting that the probability of the final result to be statistically significant is less than 20% given the data observed at the time of interim analysis. If the CP is less than 20% at the time of the interim analysis, the Data Monitoring Committee may recommend the sponsor stop the trial (stop the trial for futility).  

in a book chapter by Tin, Ming T "Conditional Power in Clinical Trial Monitoring", the pros and cons of conditional power were discussed. 

To put things in perspective, the conditional power approach attempts to assess whether evidence for efficacy or the lack of it based on the interim data is consistent with that at the planned end of the trial by projecting forward or using conditional likelihood given the eventuality. Thus it substantially alleviates the major inconsistency in all other group sequential tests where different sequential procedures applied to the same data yield different answers. ...

The advantage of the conditional power approach for trial monitoring is its flexibility. It can be used for unplanned analysis and even analysis whose timing depends on previous data. For example, it allows inferences from overrunning or underrunning (namely, more data come in after the sequential boundary is crossed, or the trial is stopped before the stopping boundary is reached. Conditional power can be used to aid the decision for early termination of a clinical trial to complement the use of other methods or when other methods are not applicable. 

The caveat is that the conditional power can be calculated with different assumptions about the remaining data. Depending on the assumptions about the remaining data following the observed data, the alternative hypothesis (original design effect), or others, the conditional power can sometimes be quite different resulting in different conclusions about the futility assessment. 

Some examples: 

in the SAP for "Randomized, Open-Label Study of Abiraterone Acetate (JNJ-212082) plus Prednisone with or without Exemestane in Postmenopausal Women with ER+ Metastatic Breast Cancer Progressing after Letrozole or Anastrozole Therapy", conditional power was described to be calculated with both the assumption of the remaining data following the original hazard ratio (alternative hypothesis) and the assumption of the remaining data following the observed hazard ratio at the interim.
3.1.2 Conditional Power

Conditional power is the probability that the study will demonstrate statistical significance at the end of the study (i.e. final analysis to claim superiority on PFS), conditioning on the data observed in the study thus far, and an assumption about the trend of the data to be observed in the remainder of the study. Two assumptions about the trend of the data were presented below: The futility boundary corresponds to a conditional power of approximately 39% if the original hazard ratio assumption is true, while only 4% conditional power will be achieved if the observed hazard ratio at interim is true for the remainder of the study. The efficacy boundary corresponds to a conditional power of approximately 90% if the original hazard ratio assumption is true, and 92% conditional power will be achieved if the observed hazard ratio at interim is true for the remainder of the study. The conditional power of stopping boundaries was computed using method of Lan (2009).
In Gilead's trial "A Multicenter, Adaptive, Randomized Blinded Controlled Trial of the Safety and Efficacy of Investigational Therapeutics for the Treatment of COVID-19 in Hospitalized Adults", the repeat significant test procedure (the alpha spending function) was used to evaluate the potential stop for overwhelming efficacy and the stochastic curtailment approach (conditional power) was used to evaluate the potential stop for futility. 


In a trial by Incyte "GRAVITAS-301: A Randomized, Double-Blind, Placebo-Controlled Phase 3 Study of Itacitinib or Placebo in Combination With Corticosteroids for the Treatment of First-Line Acute Graft-Versus-Host Disease", interim data monitoring for the potential stop for efficacy or futility is assessed and conditional power of 20% is used as the threshold for declaring the futility: 


Further reading: 

Saturday, June 19, 2021

About Controversial Approval of Biogen's Alzheimer Drug

Two weeks ago, the US Food and Drug Administration (FDA) approved aducanumab (brand name Aduhelm) as a treatment for Alzheimer's disease -- a historic decision not because it addresses the longstanding unmet medical need for a safe and effective cure of a devastating disease that affects nearly 6 million Americans, but because of the unprecedented irregularities of the agency's actions, undermining its mission to protect public health and ensure the "safety, efficacy, and security" of treatments made available in the United States. 

The winner is obviously the drug developer, Biogen and its collaborator Eisai. They probably never thought that FDA would be so collaborative and more desired to approve aducanumab than the sponsors themselves. They rescued a drug that had been declared 'unlikely' to work (futility) just two years ago. They got an unlimited label for all Alzheimer patients (beyond the early Alzheimer patients that were studied in their clinical trials). They can decide on the drug price whatever they want because there is no price control in the US once the drug is approved by the US FDA. They have at least 9 years to complete the post-marketing confirmatory study. There are no incentives for them to complete this confirmatory study as early as possible. The longer the study takes, the more time they can make the money from a drug with unproven efficacy. 

The losers include a long list: 
  • FDA - loses its credibility
  • Alzheimer's patients - are given false hope and may end up taking 'snake oil' for many years down the road
  • Patient Advocacy Group - Alzheimer's Association was unhappy with Biogen's $56,000/year/patient price tag. 
  • Medicare/Medicaid/Insurance Companies - extremely high cost associated with Aduhelm ($56,000/year/patient) and the broad label for Aduhelm can cost them a lot of money
  • FDA Adcom Committee - insulted by FDA's decision to approve even though the Adcom voted overwhelmingly against the approval
  • Regulatory science - FDA has touted for years about the regulatory science and the strict rules to be followed for drug approval - these rules are not followed by the FDA - what can you do?
  • FDA statisticians - It is clear that the FDA statistical reviewers had their dissenting opinions and questioned the data / results from two pivotal studies that were prematurely discontinued due to futility. Statisticians' opinions were overruled. 
......

Usually, approval like this will be heralded as historical and celebrated by all parties - not this time for aducanumab approval. The reactions are overwhelmingly negative. Here is a list of articles discussing the controversial approval from different angles.  
In approving Biogen's aducanumab, the boundaries between the FDA (as a regulator) and the sponsor (as a drug developer) were crossed. In the drug development field, the sponsor will try everything to exaggerate the benefit and minimize the side effects while FDA will need to be on the conservative side, tamper down the expectations, prevent the manipulation of the data and biases in data analyses,...  In the aducanumab case, FDA is determined to approving the drug no matter what and no matter whether the data/ results from clinical trials have demonstrated "Substantial Evidence of Effectiveness". FDA retrospectively find a regulatory pathway (accelerated approval pathway) for approval. In doing so, FDA failed to stand by the standards it established and both regulators and sponsors had followed.

In the drug development field, pre-specification is critical. The regulatory pathway, the number of clinical trials for clinical development program, the clinical trial design, study endpoints, and statistical analysis plan have to be discussed and agreed upon with FDA. As Eli Lilly's CEO said that in drug development, "where the gold standard for approval is you call your shot, and then you hit your shot, like Babe Ruth pointing at the left-field and then hitting his home run there." The sub-group analyses and post-hoc analyses are for hypothesis-generating and can not be used to support the regulatory approval. In Biogen's case, it is obvious that an additional clinical trial is needed before the approval. By switching to the accelerated approval pathway, FDA essentially agreed that the pivotal studies with cognitive and function measures provided insufficient evidence for approval and they had to retrofit to find accelerated approval that is based on the biomarker (amyloid).

In a letter from FDA to AdCom about switching to the accelerated approval pathway, Dr. Billy Dunn said this: 
Following the advisory committee meeting, further discussion within FDA considered the uncertainty introduced by the conflicting results of Study 302 and Study 301 and the committee’s discussion of that uncertainty. Our discussions raised further consideration of the accelerated approval pathway; a topic discussed earlier in the development program but not directly discussed during the advisory committee meeting given the focus at that meeting on the evidence of clinical benefit. As you may be aware, the accelerated approval pathway is for drugs to treat serious diseases that are expected to provide a meaningful advantage over available therapy, but where there is residual uncertainty regarding the drug’s ultimate clinical benefit. To be approved under this pathway, there must be substantial evidence of the drug’s effectiveness on a surrogate endpoint—usually an endpoint that reflects the underlying disease pathology (accelerated approval can also use an intermediate clinical endpoint). An effect on this surrogate endpoint must be shown to be reasonably likely to predict clinical benefit. We concluded that these requirements were met for aducanumab, with substantial evidence that the drug reduces amyloid beta plaque, and that this reduction is reasonably likely to predict clinical benefit. For drugs approved using the accelerated approval pathway, further study is required to verify anticipated clinical benefits
FDA is preoccupied and determined to approve aducanumab no matter which pathway is used. The following conclusion is subjective and a lot of people will certainly not agree: "We concluded that these requirements were met for aducanumab, with substantial evidence that the drug reduces amyloid-beta plaque, and that this reduction is reasonably likely to predict clinical benefit." Had the FDA been so sure about the biomarker 'amyloid-beta plaque' reduction is 'reasonably likely to predict clinical benefit', they would advise the sponsors (Biogen and other Alzheimer drug developers) to design their phase III studies with the primary efficacy endpoint being the lowering the amyloid-beta plaque, not the measuring the benefit in improving the cognitive and function. 

Accelerated approval pathway is described in FDA guidance for industry "Expedited Programs for Serious Conditions – Drugs and Biologics", but is only used in a situation where the confirmatory studies with clinical endpoints have not been conducted. In Biogen's case, two confirmatory studies with clinical endpoints had already been completed (actually was stopped early for futility). It is a round peg in a square hole to retrospectively going back to the accelerated approval pathway based on the biomarker because of the conflicting and unconvincing results from confirmatory trials with clinical endpoints. Approval of aducanumab based on an accelerated approval pathway breaks agency precedent. "Accelerated approval is traditionally used for treatments that haven't yet proved themselves in large trials. In Biogen's case, Aduhelm went through two Phase 3 studies and came up with conflicting evidence."

FDA also loses its fairness - there are a lot of diseases with unmet medical needs. The drugs for other unmet medical conditional have been tested and generated stronger evidence than Biogen's pivotal studies, but the drugs were rejected by FDA. Here is an article about ALS (amyotrophic lateral sclerosis) - more deadly than Alzheimer's disease.

FDA's controversial Aduhelm decision leaves ALS patients feeling spurned

The FDA's controversial approval of Biogen's Aduhelm drug for Alzheimer's disease has been met with fierce resistance from all corners of the biopharma industry, but few seem to be as upset with the decision as ALS patients and advocacy groups.

For all that's already been written and discussed about the agency's announcement, from the drug's exorbitantly high price of $56,000 per year to criticism over lowered standards, ALS patients see something more. ALS patients and associations say they largely regarded Aduhelm's approval as a bittersweet double standard: happy that those with Alzheimer's have a new drug available, but questioning how the FDA evaluated Biogen's drug compared to the experimental programs being studies for their own disease. 

Nothing punctuated the feeling harder than the agency's announcement in April that a promising drug under development by the biotech Amylyx would need another study to confirm efficacy. This program, called AMX0035, hit the primary endpoint for improving function specifically laid out in the FDA's 2019 guidelines for new ALS treatments, whereas Biogen halted two pivotal Aduhelm studies early because of futility in its own function measurements. 

In general, to demonstrate substantial evidence of effectiveness of the drug, two adequate and well-controlled trials are needed. In Biogen's case, two adequate and well-controlled trials ENGAGE and EMERGE to evaluate the efficacy and safety of aducanumab in patients. When two studies gave contradicting results (one positive and one not positive), a third adequate and well-controlled study will be needed (before the drug approval, not after the drug approval). I remembered other examples: Pirfenidone was developed for treating the rare disease of IPF (idiopathic pulmonary fibrosis). The sponsor conducted two pivotal studies with one study positive (p=0.01) and one study negative (p=0.5). Initial NDA submission with these two studies was rejected by the FDA. FDA demanded the sponsor to conduct a third study. A third study gave a positive result (p<0.01) and NDA was resubmitted, and FDA approved the Perfenidone for IPF. Another example is ciprofloxacin dispersion in non-CF bronchiectasis (rare disease without approved treatment). The sponsor conducted two identical phase III studies ORIBIT-3 and ORBIT-4 - two studies gave contradicting results (one positive and one not positive). The NDA was rejected by FDA and additional studies were not conducted due to funding issues - ciprofloxacin dispersion remains not approved for non-CF bronchiectasis. 

In Biogen's case, two years have passed since they revealed the results of their pre-maturely discontinued studies: one with positive and one with negative. They could have started the third study and would be able to complete the third study not far from now. Instead, with FDA's help, they got their aducanumab approved without doing the third study and they were given a long 9-years to do a post-marketing phase IV study.

References:

Saturday, November 07, 2020

The Saga of Biogen’s Alzheimer Drug Aducanumab

Aducanumab is an investigational compound being studied for the treatment of early Alzheimer’s disease co-developed by Biogen and Eisai. Aducanumab is a human immunoglobulin gamma 1 (IgG1) anti‐amyloid beta monoclonal antibody (mAb) targeting aggregated forms of amyloid beta - a fundamental pathological hallmark of the disease.

After the ‘successful’ phase I study (PRIME trial) to demonstrated that Aducanumab had an acceptable safety and tolerability profile and reduced brain amyloid-beta accompanied by a slowing of clinical decline measured by Clinical Dementia Rating-Sum of Boxes (CDR-SB) and Mini-Mental State Examination (MMSE) scores, Biogen designed two pivotal studies (ENGAGE and EMERGE) to evaluate the efficacy and safety of aducanumab in patients with mild cognitive impairment due to Alzheimer’s disease and mild Alzheimer’s disease dementia.

Then the saga began,… a roller-coaster year in 2019.

In March 2019, Biogen and Eisai announced to discontinue Phase 3 ENGAGE and EMERGE trials of aducanumab in Alzheimer’s disease after the interim analyses found that the futility boundaries were crossed. The independent data monitoring committee advised that aducanumab would be unlikely to meet primary endpoints even the studies would be continued to the completion.

While everybody thought that aducanumab was dead in the water, Biogen unexpectedly announced that they would plan regulatory filing for aducanumab in Alzheimer’s disease based on a new analysis of larger data set from Phase 3 studies.

On July 08, 2020, Biogen announced that they had completed the submission of a Biologics License Application (BLA) to the U.S. Food and Drug Administration (FDA) for the approval of aducanumab, an investigational treatment for Alzheimer's disease.

Given that their BLA submission was based on the re-analyses from studies that had been discontinued due to futility, the common understanding would be that the FDA would reject their BLA and require them to do another Phase 3 study.

Then came the last week,…a roller-coaster week last week

For a drug application with controversies, FDA will usually organize an advisory committee meeting to seek the opinions from outside experts including representatives from patients’ organization, patient advocate group, and consumer citizen groups. There is no exception to BLA for aducanumab. Peripheral and Central Nervous System (PCNS) Drugs Advisory Committee Meeting was scheduled for November 6, 2020 and the meeting materials were posted online two days before the meeting on November 4, 2020.

The documents released by FDA came as shocking to outsides. With one negative trial and one very positive trial in one of the dose groups, one would expect that the FDA would give a negative tone and demand a third trial to confirm the efficacy. Usually, the sponsor would try everything to convince FDA that the drug was efficacious and safe while FDA would be conservative and poke and probe the data to identify the issues to discredit the sponsor’s claim about the efficacy and safety.

However, this time for aducanumab, FDA is on the sponsor’s side. The briefing book from the FDA (actually combined FDA and Biogen Briefing Information) depicted a very rosy picture for aducanumab’s efficacy and safety.

With FDA’s backing, one would think that aducanumab is on its way to get a positive opinion from advisory committee members and eventually to be the first FDA approved novel medication for the treatment of Alzheimer’s disease since 2004.

Then the shocking news continued,…

At Friday’s advisory committee meeting, committee members resoundingly concluded Friday that clinical data did not support the approval of Biogen’s much-watched Alzheimer’s drug, aducanumab, while providing a rebuke to the Food and Drug Administration, whose reviewers had given the medicine a glowing appraisal.


Here are the voting results:

FDA’s Questions to the Advisory Committee

Voting Results

 

 

 

Does Study 302, Viewed independently and without regard for Study 301, provide strong evidence that supports the effectiveness of aducanumab for the treatment of AD?

1

8

2

Does Study 103 provide supportive evidence of the effective of aducanumab for the treatment of AD?

0

7

4

Has the applicant presented strong evidence of a pharmacodynamic effect on AD pathophysiology?

5

0

6

In light of the understanding provided by the exploratory analyses of Study 301 and Study 302, along with the results of Study 103 and evidence of a pharmacodynamic effect on AD pathophysiology, is it reasonable to consider Study 302 as primary evidence of effectiveness of aducanumab for the treatment of AD?

0

10

1

Note: Study 301 was the pivotal study (ENGAGE trial) with the negative outcome; Study 302 was the pivotal study (EMERGE trial) with a positive outcome in high dose group; Study 103 was the phase I proof-of-concept study (PRIME trial).

 

Not sure where aducanumab will go from here. It will be another shocking if FDA approves aducanumab for the treatment of Alzheimer’s disease given the extremely negative view/voting outcome from the advisory committee panel even though everybody understands there is a huge, urgent, unmet need for a new AD drug. As one of the experts said, "with FDA's reputation already in a precarious position, it could be difficult -- maybe impossible -- to go against an expert panel at this time no matter how badly they want this". The best path forward would be to conduct a third pivotal study if Biogen has confidence in aducanumab. In Chinese idiom, true gold fears no fire.

The briefing book from the combined FDA and Biogen Briefing Information revealed the discordance in viewers about the aducanumab efficacy and the study issues among FDA reviewers. The conclusions from the clinical reviewer and the statistical reviewers were dramatically different. As one of the panel members commented on this, “It feels like the audio and video on TV are out of sync”. Unfortunately, in the overall conclusion and the FDA’s presentation, the clinical reviewer’s opinion trumped the statistical reviewer’s opinion. The statistical reviewer wasn’t even given an opportunity to present at the steering committee meeting.

Here is the conclusion from the clinical reviewer:

“……the applicant has provided substantial evidence of effectiveness to support approval. Study 302 provided the primary evidence of effectiveness as robust and exceptionally persuasive study demonstrating a treatment effect on a clinically meaningful endpoint and reinforced by effects on secondary endpoints, biomarkers, and in relevant sugroups. Study 103 was an adequate and well-controlled study that included design components consistent with Study 302 and demonstrated a persuative treatment effect on both clinical endpoints. The dose-response relationship for Aβ reduction provides support for the positive finding in the 10 mgkg treatment arm to the apparently dose-related effects observed on clinical outcomes in Studies 103 and 302. Study 301 does not contribute to the evidence of effectiveness. The results of exploratory analyses, however, contribute to the overall understanding of Study 301 and together do not meaningfully detract from the persuasiveness of Study 302."

 Here is the conclusion from the statistical reviewers:

“In summary, the totality of the data does not seem to support the efficacy of the high dose. There is only one positive study at best and a second study which directly conflicts with the positive study. Both studies were not fully completed as they were terminated early for futility and had sporadic unblinding for dose management of ARIA cases which was much higher in the drug group(s). The Amyloid PET sub-study data suggested a larger effect in APOE- (non-carriers) which is the opposite of what was observed for the clinical outcome data. Within the high dose group at the patient level, there is no correlation between the Week 78 change in the primary biomarker Ab in the cerebellum and the Week 78 Change from baseline in CDR-SB. In study 302, the on-face positive study, the raw correlation had the wrong +/- sign to support a realistic link between biomarker and long-term clinical change in cognition/function as measured by CDR-SB. For these reasons, the reviewer believes there is no compelling substantial evidence of treatment effect or disease slowing and that another study is needed to confirm or deny the positive study and the negative study. "