Sunday, June 20, 2021

Early Phase Trial to Find Maximal Tolerated Dose (MTD) - 3+3, CRM, and BOIN Designs

In early-phase clinical trials, determining the dose range and therapeutic window is critical. The purpose of the early-phase studies may just be to identify the maximum tolerated dose or maximum tolerable dose. 

Definition of maximum tolerated dose (MTD)
The highest dose of a drug or treatment that does not cause unacceptable side effects. The maximum tolerated dose is determined in clinical trials by testing increasing doses on different groups of people until the highest dose with acceptable side effects is found. Also called MTD.
The studies for identifying the MTD are usually designed as a dose-escalation study and the dose-escalation study is defined as:
A study that determines the best dose of a new drug or treatment. In a dose-escalation study, the dose of the test drug is increased a little at a time in different groups of people (also called cohort) until the highest dose that does not cause harmful side effects is found. A dose-escalation study may also measure ways that the drug is used by the body and is often done as part of a phase I clinical trial. These trials usually include a small number of patients and may include healthy volunteers.

In dose-escalation studies, within each dose cohort, a placebo group can be included even though the majority of the dose-escalation studies for MTD are designed without placebo controls.

Identifying MTD is based on the number of dose-limiting toxicities (DTLs)that are observed in each dose cohort. DTLs are defined as: 

side effects of a drug or other treatment that are serious enough to prevent an increase in dose or level of that treatment.

In practice, DTLs are often defined as grade 3 or above adverse events according to Common Terminology Criteria for Adverse Events (CTCAEs) especially in the oncology area even though other customer-defined criteria for DTLs may be used in non-oncology areas. 

Clinical trials to identify the MTD are generally needed for phase I studies directly conducted in patients, not healthy volunteers. Areas that the phase I studies are conducted in patients, not healthy volunteers, include oncology drugs, drugs in severe diseases such as AIDS, Sepsis, ARDS, etc., the gene and cell therapies, human-plasma derived products.

There are different types of clinical trial designs for identifying the MTD. The commonly used designs are 3+3 design, Continuous Reassessment Method (CRM), and Bayesian Optimal INterval design (BOIN). 

3+3 Design was discussed in an early post Phase I Dose Escalation Study Design: "3 + 3 Design". It is a straightforward rule-based method and requires no statistical calculations. 3+3 design is the most frequently used method for identifying the MTD. 

The CRM is a model-based design for phase I trials, which aims to find the maximum tolerated dose (MTD) of a new therapy. The CRM has been shown to be more accurate in targeting the MTD than traditional rule-based approaches such as the 3 + 3 design. With CRM design, statistical inferences on the model parameter(s) need to be made using likelihood-based or Bayesian approaches and DLT probability at each dose needs to be estimated. The patient is assigned to the next dose level based on the probability of patients with DLTs at the current dose level. The toxicity risk of other dose levels is based on accrued data, which improves trial efficiency. 

Following articles or videos provided a great introduction/reference about the CRM method: 

The BOIN design shares the simplicity of the 3+3 design, which makes the decision of dose escalation/de-escalation by comparing the observed DLT rate with 0/3, 1/3, 2/3, 0/6, 1/6, and 2/6. The BOIN design makes the decision by comparing with two fixed boundaries, λe and λd, which is arguably even simpler.



BOIN design are described and explained in the following article and video:
Software for Sample Size Calculation for Phase I MTD Finding Studies:
  • trialdesign.org is a website developed and maintained by a research team at MD Anderson Cancer and it contains the literature and software for phase I designs including CRM and BOIN. 
Additional Videos: 

Saturday, June 19, 2021

About Controversial Approval of Biogen's Alzheimer Drug

Two weeks ago, the US Food and Drug Administration (FDA) approved aducanumab (brand name Aduhelm) as a treatment for Alzheimer's disease -- a historic decision not because it addresses the longstanding unmet medical need for a safe and effective cure of a devastating disease that affects nearly 6 million Americans, but because of the unprecedented irregularities of the agency's actions, undermining its mission to protect public health and ensure the "safety, efficacy, and security" of treatments made available in the United States. 

The winner is obviously the drug developer, Biogen and its collaborator Eisai. They probably never thought that FDA would be so collaborative and more desired to approve aducanumab than the sponsors themselves. They rescued a drug that had been declared 'unlikely' to work (futility) just two years ago. They got an unlimited label for all Alzheimer patients (beyond the early Alzheimer patients that were studied in their clinical trials). They can decide on the drug price whatever they want because there is no price control in the US once the drug is approved by the US FDA. They have at least 9 years to complete the post-marketing confirmatory study. There are no incentives for them to complete this confirmatory study as early as possible. The longer the study takes, the more time they can make the money from a drug with unproven efficacy. 

The losers include a long list: 
  • FDA - loses its credibility
  • Alzheimer's patients - are given false hope and may end up taking 'snake oil' for many years down the road
  • Patient Advocacy Group - Alzheimer's Association was unhappy with Biogen's $56,000/year/patient price tag. 
  • Medicare/Medicaid/Insurance Companies - extremely high cost associated with Aduhelm ($56,000/year/patient) and the broad label for Aduhelm can cost them a lot of money
  • FDA Adcom Committee - insulted by FDA's decision to approve even though the Adcom voted overwhelmingly against the approval
  • Regulatory science - FDA has touted for years about the regulatory science and the strict rules to be followed for drug approval - these rules are not followed by the FDA - what can you do?
  • FDA statisticians - It is clear that the FDA statistical reviewers had their dissenting opinions and questioned the data / results from two pivotal studies that were prematurely discontinued due to futility. Statisticians' opinions were overruled. 
......

Usually, approval like this will be heralded as historical and celebrated by all parties - not this time for aducanumab approval. The reactions are overwhelmingly negative. Here is a list of articles discussing the controversial approval from different angles.  
In approving Biogen's aducanumab, the boundaries between the FDA (as a regulator) and the sponsor (as a drug developer) were crossed. In the drug development field, the sponsor will try everything to exaggerate the benefit and minimize the side effects while FDA will need to be on the conservative side, tamper down the expectations, prevent the manipulation of the data and biases in data analyses,...  In the aducanumab case, FDA is determined to approving the drug no matter what and no matter whether the data/ results from clinical trials have demonstrated "Substantial Evidence of Effectiveness". FDA retrospectively find a regulatory pathway (accelerated approval pathway) for approval. In doing so, FDA failed to stand by the standards it established and both regulators and sponsors had followed.

In the drug development field, pre-specification is critical. The regulatory pathway, the number of clinical trials for clinical development program, the clinical trial design, study endpoints, and statistical analysis plan have to be discussed and agreed upon with FDA. As Eli Lilly's CEO said that in drug development, "where the gold standard for approval is you call your shot, and then you hit your shot, like Babe Ruth pointing at the left-field and then hitting his home run there." The sub-group analyses and post-hoc analyses are for hypothesis-generating and can not be used to support the regulatory approval. In Biogen's case, it is obvious that an additional clinical trial is needed before the approval. By switching to the accelerated approval pathway, FDA essentially agreed that the pivotal studies with cognitive and function measures provided insufficient evidence for approval and they had to retrofit to find accelerated approval that is based on the biomarker (amyloid).

In a letter from FDA to AdCom about switching to the accelerated approval pathway, Dr. Billy Dunn said this: 
Following the advisory committee meeting, further discussion within FDA considered the uncertainty introduced by the conflicting results of Study 302 and Study 301 and the committee’s discussion of that uncertainty. Our discussions raised further consideration of the accelerated approval pathway; a topic discussed earlier in the development program but not directly discussed during the advisory committee meeting given the focus at that meeting on the evidence of clinical benefit. As you may be aware, the accelerated approval pathway is for drugs to treat serious diseases that are expected to provide a meaningful advantage over available therapy, but where there is residual uncertainty regarding the drug’s ultimate clinical benefit. To be approved under this pathway, there must be substantial evidence of the drug’s effectiveness on a surrogate endpoint—usually an endpoint that reflects the underlying disease pathology (accelerated approval can also use an intermediate clinical endpoint). An effect on this surrogate endpoint must be shown to be reasonably likely to predict clinical benefit. We concluded that these requirements were met for aducanumab, with substantial evidence that the drug reduces amyloid beta plaque, and that this reduction is reasonably likely to predict clinical benefit. For drugs approved using the accelerated approval pathway, further study is required to verify anticipated clinical benefits
FDA is preoccupied and determined to approve aducanumab no matter which pathway is used. The following conclusion is subjective and a lot of people will certainly not agree: "We concluded that these requirements were met for aducanumab, with substantial evidence that the drug reduces amyloid-beta plaque, and that this reduction is reasonably likely to predict clinical benefit." Had the FDA been so sure about the biomarker 'amyloid-beta plaque' reduction is 'reasonably likely to predict clinical benefit', they would advise the sponsors (Biogen and other Alzheimer drug developers) to design their phase III studies with the primary efficacy endpoint being the lowering the amyloid-beta plaque, not the measuring the benefit in improving the cognitive and function. 

Accelerated approval pathway is described in FDA guidance for industry "Expedited Programs for Serious Conditions – Drugs and Biologics", but is only used in a situation where the confirmatory studies with clinical endpoints have not been conducted. In Biogen's case, two confirmatory studies with clinical endpoints had already been completed (actually was stopped early for futility). It is a round peg in a square hole to retrospectively going back to the accelerated approval pathway based on the biomarker because of the conflicting and unconvincing results from confirmatory trials with clinical endpoints. Approval of aducanumab based on an accelerated approval pathway breaks agency precedent. "Accelerated approval is traditionally used for treatments that haven't yet proved themselves in large trials. In Biogen's case, Aduhelm went through two Phase 3 studies and came up with conflicting evidence."

FDA also loses its fairness - there are a lot of diseases with unmet medical needs. The drugs for other unmet medical conditional have been tested and generated stronger evidence than Biogen's pivotal studies, but the drugs were rejected by FDA. Here is an article about ALS (amyotrophic lateral sclerosis) - more deadly than Alzheimer's disease.

FDA's controversial Aduhelm decision leaves ALS patients feeling spurned

The FDA's controversial approval of Biogen's Aduhelm drug for Alzheimer's disease has been met with fierce resistance from all corners of the biopharma industry, but few seem to be as upset with the decision as ALS patients and advocacy groups.

For all that's already been written and discussed about the agency's announcement, from the drug's exorbitantly high price of $56,000 per year to criticism over lowered standards, ALS patients see something more. ALS patients and associations say they largely regarded Aduhelm's approval as a bittersweet double standard: happy that those with Alzheimer's have a new drug available, but questioning how the FDA evaluated Biogen's drug compared to the experimental programs being studies for their own disease. 

Nothing punctuated the feeling harder than the agency's announcement in April that a promising drug under development by the biotech Amylyx would need another study to confirm efficacy. This program, called AMX0035, hit the primary endpoint for improving function specifically laid out in the FDA's 2019 guidelines for new ALS treatments, whereas Biogen halted two pivotal Aduhelm studies early because of futility in its own function measurements. 

In general, to demonstrate substantial evidence of effectiveness of the drug, two adequate and well-controlled trials are needed. In Biogen's case, two adequate and well-controlled trials ENGAGE and EMERGE to evaluate the efficacy and safety of aducanumab in patients. When two studies gave contradicting results (one positive and one not positive), a third adequate and well-controlled study will be needed (before the drug approval, not after the drug approval). I remembered other examples: Pirfenidone was developed for treating the rare disease of IPF (idiopathic pulmonary fibrosis). The sponsor conducted two pivotal studies with one study positive (p=0.01) and one study negative (p=0.5). Initial NDA submission with these two studies was rejected by the FDA. FDA demanded the sponsor to conduct a third study. A third study gave a positive result (p<0.01) and NDA was resubmitted, and FDA approved the Perfenidone for IPF. Another example is ciprofloxacin dispersion in non-CF bronchiectasis (rare disease without approved treatment). The sponsor conducted two identical phase III studies ORIBIT-3 and ORBIT-4 - two studies gave contradicting results (one positive and one not positive). The NDA was rejected by FDA and additional studies were not conducted due to funding issues - ciprofloxacin dispersion remains not approved for non-CF bronchiectasis. 

In Biogen's case, two years have passed since they revealed the results of their pre-maturely discontinued studies: one with positive and one with negative. They could have started the third study and would be able to complete the third study not far from now. Instead, with FDA's help, they got their aducanumab approved without doing the third study and they were given a long 9-years to do a post-marketing phase IV study.

References:

Tuesday, June 01, 2021

Decentralized clinical trials and in silico clinical trials

These days, the buzzword in the clinical trial field is 'decentralized clinical trials' or DCTs. The Covid-19 pandemic seems to push the clinical trials toward 'decentralized' or 'hybrid' of decentralized and traditional clinical trials. 

Traditional clinical trials are 'centered' around the clinical trial sites and the investigators. The patients (clinical trial participants) are recruited by the investigators who are the medical doctors responsible for the conduct of the clinical trial at trial sites. The trial sites are the clinics, hospitals, and medical centers. The patients would need to visit the trial sites regularly to see the investigators for clinical trial activities (signing the informed consent, screening for eligibility, receiving study treatments, performing efficacy and safety measures,...). The clinical trial data will then be recorded and entered into the database (for example EDC) by the study coordinator or investigator at investigational sites.

Decentralized clinical trials are defined as the decentralization of clinical trial operations where technology is used to communicate with study participants and collect data and the data collection will not depend on the frequent patient's visits to the investigational sites. According to CTTI (clinical trial transformation initiative) Recommendations: Decentralized Clinical Trials
DCTs using telemedicine and other emerging and novel information technology (IT)
services offer the potential for local HCPs to participate in clinical trials. This may
provide several advantages compared to traditional clinical trials conducted at more
centralized clinical trial sites, including the following:
  • Faster trial participant recruitment, which can accelerate trial participant access to important medical interventions and reduce costs for sponsors.
  • Improved trial participant retention, which may reduce missing data, shorten clinical trial timelines, and improve data interpretability.
  • Greater control, convenience, and comfort for trial participants by offering at home or local patient care.
  • Increased diversity of the population enrolled in clinical trials.
  • An opportunity for home administration or home use of the IMP, which may be
  • more representative of real-world administration/use post-approval.
FDA's Advancing Oncology Decentralized Trials - Learning from COVID-19 Trial Datasets also listed the advantages of the DCTs: 
  • Decentralized Clinical Trials (DCT) may have several potential benefits including reduced patient and sponsor burden and increased accrual and retention of a more diverse trial population.
  • Use of full or hybrid DCT designs by commercial sponsors has been rare in oncology, in part due to uncertainty surrounding the effect of remote assessments on data quality and outcomes.
  • COVID-19 has necessitated DCT-type trial modifications such as remote assessments to reduce patient exposure to COVID-19 infection from travel to trial sites.
  • Many of these remote assessment modifications were deployed in the middle of large ongoing cancer trials.
  • There is an opportunity to evaluate the effect of remote assessments on trial data to advance Decentralized Trials in oncology.
  • Better understanding of the effect of DCT modifications can reduce uncertainty for sponsors and regulatory bodies, and identify mitigation strategies for future prospective DCT designs.
Historically, DCTs may be called differently: virtual trials, siteless trials, remote trials, digital trials, direct-to-patient trials. These different names may just describe one specific aspect of the DCTs and can cause confusion. For example, 'virtual trials' can be confused with the in silico trials which are based on computer models and do not use real participants (patients) but computer programs to model participants to assess drug efficacy and safety during the preclinical phase or before a traditional trial.

"Decentralized clinical trial" is a “Terrible Name for a Promising Innovation” and is not a best terminology for patients and study participants. Alternative names such as direct-to-patient trials, patient-centric trials, and home-based trials seem to be more straightforward and better terms.

In essence, in traditional clinical trials, we bring the trial to the patients; in DCTs, we bring the patients to the trial.

There are still a lot of challenges and obstacles to implementing decentralized clinical trials. Application of DCTs may be limited to some special situations (such as post-marketing studies with patient-reported outcomes and outcomes measured digitally). It is still rare for pivotal and registration studies to use full DCTs. It seems to be more appropriate to adopt hybrid trials - the combination of the traditional and the decentralized trials. For example, in clinical trials in the rare disease area, it is difficult for patients to travel to the investigational sites, the patients may visit the investigational sites for some important visits (in-clinic visits) and then the home health care nurses may be used for in-home visits to patient's home.  


In 2019, Janssen, PRA Launch a Fully Virtual Trial (it should be called the decentralized trial) "A Study on Impact of Canagliflozin on Health Status, Quality of Life, and Functional Status in Heart Failure (CHIEF-HF)". The study design was described in the Circulation: Heart Failure "Novel Trial Design: CHIEF-HF". CHIEF-HF seems to be the first phase III trial being fully decentralized.
 
Here are some references for DCTs:   
The decentralized clinical trials still collect the data from the patients and should not be called 'virtual clinical trials'. On the contrary, the 'In Silico clinical trials' is more appropriately called 'virtual clinical trials' or 'patientless clinical trials' and it simulates the virtual subjects for modeling and prediction. In Silico clinical trials may use the data from pre-clinical and historical clinical trials for simulation but involves no real patients in the study.  

According to the senator bill "AGRICULTURE, RURAL DEVELOPMENT, FOOD AND DRUG
ADMINISTRATION, AND RELATED AGENCIES APPROPRIATIONS BILL, 2016", In Silico clinical trials use computer models and simulations to develop and assess devices and drugs, including their potential risk to the public, before being tested in live clinical trials."
 

In Silico clinical trials are part of the model-informed drug development (MIDD). the FDA has a MIDD pilot program managed by the Division of Pharmacometrics

Dr. Yaning Wang has multiple presentations promoting the MIDD and In Silico clinical trials, for example, in his presentation at the 2021 FDA Science Forum "Regulatory Applications and Research of Model-Informed Drug Development (MIDD)" (Youtube video at 2:38:25) and in his presentation at PMDA "Application of MIDD in New Drug Development and Approval". 

Here are some additional references on In Silico clinical trials:

Saturday, May 29, 2021

Patient Advocacy Groups in Drug Development and Clinical Trials For Patients With Rare Diseases

I just saw an article on clinicalleader.com "Best Practices For Designing And Running Clinical Trials For Patients With Rare Diseases" by my previous colleague, Mary L Smith. 


She had some excellent points about the important role of the patient advocacy group in the drug development process in rare disease areas.

Now that the drug development has been moved to the patient-centric, the patient's voice (usually through the patient advocacy group) is critical. Over the years, I have seen or directly interacted with some of the patient advocacy groups in various activities.

We saw that the patient advocacy group (i.e Parentprojectmd.org) played a critical role in pushing FDA to approve the first drug for Duchenne Muscular Dystrophy even though there wasn't substantial evidence to support the efficacy. 

Cystic Fibrosis Foundation is probably the most successful patient advocate group and it is very well run and organized. In the US, it is extremely to do any clinical trials in CF patients without going through the Cystic Fibrosis Foundation. Cystic Fibrosis Foundation may also be the richest patient advocacy group and received a lot of money from the royalties from CF drug developers. 
In conducting the clinical trials in patients with Alpha-1 antitrypsin deficiency (a genetic form of severe COPD), we were very closely working with Alpha 1 Foundation - a patient advocacy group created by three Alpha-1 Antitrypsin Deficiency patients.  Alpha-1 Foundation's help pushed FDA/NIH to organize the workshops to discuss the efficacy endpoints that are realistic in clinical trials in Alpha-1 Antitrypsin patients. One of the endpoints was to measure the lung density through CT scan - so-called lung densitometry that was eventually accepted by the FDA to be the primary efficacy measure in Alpha-1 Antriypsin Deficiency trials.   

Other examples of patient advocacy groups are:

Sunday, May 23, 2021

Are we moving away from data listings now that the standard data sets (such as CDISC-compliant SDTM and ADaM data sets) are mandated by FDA?

After a clinical trial is concluded, the statistical group (statisticians and statistical programmers) will generate the tables, listings, and figures (TLFs in short) for the clinical study report (CSR). If the clinical trial results are good, the CSR including the TLFs and the data sets where the TLFs are generated from will be submitted to the regulatory authorities (such as FDA) for marketing authorization application (such as new drug application (NDA) and biological license application (BLA). 

From the statistical standpoint, the clinical trial is a process to collect the data (demographic, efficacy, and safety). The data sets we collected will be converted and mapped to the standardized format according to the CDISC standards - SDTM for tabulation data sets and ADaM for analysis data sets. The standardized data sets are then used for generating the tables, listings, and figures (may also be called the post-text TLFs). The post-text TLFs will be the building block for constructing the CSRs. 

The data listings are required as part of the CSRs and post-text listings are included in the CSRs as appendices. As specified in ICH E3 "STRUCTURE AND CONTENT OF CLINICAL STUDY REPORTS", data listings will be organized according to the sections/numbers below: 


A data listing will be looking like the below with a listing of adverse events as an example (the mock-up shell). Notice that the listing for adverse events is numbered as Listing 16.2.7 corresponding to the section number indicated in ICH E3 (above).

With the FDA's mandate for submission of the standardized data sets, the CDISC standard tabulation data set is SDTM (study data tabulation model), the natural question is if the data listings are still needed. The data listings are just the simple display of the SDTM data set (or ADaM data set) with format/layout beautified. Some sponsors have moved to the direction of not generating the formal data listings and they think that the CDISC-compliant data sets will be sufficient to replace the data listings. 

According to an FDA Final Guidance Webinar Q&A at Pinnacle21 website, the response indicated that the data listings would be replaced by the standard data sets. 
28: With the advent of a new “misc folder” instead of listings do you think that FDA is getting away from generating listings for each study? It seems that the datasets (SEND, SDTM, and ADaM) would stand alone to support any listings.

Once the standards requirements are in effect, the idea would be that listings would be replaced by the SDTM tabulations data, so yes.
The Final FDA Guidance on Standardized Study Data was initially published December 17, 2014 and has subsequently been revised several times. The FDA Binding Guidance requires Sponsors whose studies start after December 17, 2016 must submit data in FDA-supported formats listed in the FDA Data Standards Catalog. The FDA Data Standards Catalog specifies the use of CDISC standards such as SDTM, ADaM, and Define-XML as well as Controlled Terminology.

FDA has organized or participated in multiple webinars to emphasize the criticality of submitting the data sets in standardized formation, however, there is no subsequent mention about the standardized data sets replacing the data listings. 

I did an informal survey and asked the statisticians and statistical programmers working in other pharmaceutical companies and CROs, almost all responses were that they continue to generate data listings and have no plan to stop generating data listings even though they have been fully compliant with the CDISC data standards for their clinical trials.

Even though the CDISC compliant data sets make the data listings redundant and unnecessary, stopping generating the data listings entirely is a risky approach at this point. The data listings should still be generated until we see FDA's guidance indicating otherwise or until the ICH E3 is revised to indicate that the data listings can be replaced by the standardized data sets. 

Some links to this topic: 



Monday, April 26, 2021

Within Patient Benefit-Risk Evaluation? Using Outcomes to Analyze Patients versus Using Patients to Analyze Outcomes?

In our daily life, benefit-risk evaluation is something we always do whether we realize it or not. Benefit-risk evaluation is especially critical in drug development and in the regulator's decision process. We often hear that a drug is approved because the benefits outweigh the risks. In the recent decision of resuming the J&J Covid-19 vaccine, the CDC and the FDA cited that the benefits of rolling out the J&J Covid vaccine outweigh the risks of developing the rare blood clot (so-called CVST Cerebral Venous Sinus Thrombosis) in some young women who received the J&J Covid vaccine. 

In a recent New York Times article "Irrational Covid Fears", the benefit and risk of the Covid-19 vaccine are compared to a fable of our times and automobiles. 
A fable for our times
Guido Calabresi, a federal judge and Yale law professor, invented a little fable that he has been telling law students for more than three decades.
He tells the students to imagine a god coming forth to offer society a wondrous invention that would improve everyday life in almost every way. It would allow people to spend more time with friends and family, see new places and do jobs they otherwise could not do. But it would also come with a high cost. In exchange for bestowing this invention on society, the god would choose 1,000 young men and women and strike them dead.
Calabresi then asks: Would you take the deal? Almost invariably, the students say no. The professor then delivers the fable’s lesson: “What’s the difference between this and the automobile?”
In truth, automobiles kill many more than 1,000 young Americans each year; the total U.S. death toll hovers at about 40,000 annually. We accept this toll, almost unthinkingly, because vehicle crashes have always been part of our lives. We can’t fathom a world without them.
It’s a classic example of human irrationality about risk. We often underestimate large, chronic dangers, like car crashes or chemical pollution, and fixate on tiny but salient risks, like plane crashes or shark attacks.
One way for a risk to become salient is for it to be new. That’s a core idea behind Calabresi’s fable. He asks students to consider whether they would accept the cost of vehicle travel if it did not already exist. That they say no underscores the very different ways we treat new risks and enduring ones.
I have been thinking about the fable recently because of Covid-19. Covid certainly presents a salient risk: It’s a global pandemic that has upended daily life for more than a year. It has changed how we live, where we work, even what we wear on our faces. Covid feels ubiquitous.
Fortunately, it is also curable. The vaccines have nearly eliminated death, hospitalization and other serious Covid illness among people who have received shots. The vaccines have also radically reduced the chances that people contract even a mild version of Covid or can pass it on to others.
Yet many vaccinated people continue to obsess over the risks from Covid — because they are so new and salient.
This article reminds me of the seminars presented by Scott Evans. In his seminars, for example, the one posted on youtube, he started with a hypothetical question:
If you are given a choice to choose drug A or drug B, Drug A increases your intelligence, but decreases your good looks; Drug B increases your good looks, but decreases your intelligence; which drug will you choose? 
This is a typical question about the benefit-risk evaluation or benefit-risk tradeoff. With this question, he brought up a topic about an alternative (supposed to be optimal) way to perform the benefit-risk evaluation (i.e., the benefit-risk assessment on each individual patient level before aggregating the data on the group level).  
Currently, in clinical trials, the benefit (efficacy) evaluation and risk (safety) evaluation are performed independently. The study protocol was designed for showing the benefit (efficacy) - selecting the sensitive and clinically meaningful efficacy endpoint, ensuring sufficient large sample size for statistical power, sound statistical analysis methods are all for ensuring that the efficacy results can be used to demonstrate the benefit of the new drug. FDA has issued specific guidance only for efficacy "Demonstrating Substantial Evidence of Effectiveness for Human Drug and Biological Products".

Risk (safety) evaluation is usually assessed separately from the efficacy. While we collect the data for risk (safety) analysis (adverse events, serious adverse events, death, clinical laboratory results, ECG results, vital signs,...), the analyses of safety data are usually based on the summaries (no hypothesis testing) to assess the nature/pattern of the serious adverse events, related to the investigational new drug, if there is elevated levels in certain laboratory parameters,... Safety analyses contain a lot of subjective judgment. Different reviewers may come to different conclusions. 

There is no separate guidance from FDA specifically about the risk (safety assessment). Instead, the safety assessment is included in FDA's Good Review Practice: Clinical Review Template - a checklist for FDA reviewers in evaluating the safety. 

Only after the efficacy and safety are separately analyzed and evaluated, are a benefit-risk section written as a formal evaluation of the benefit-risk - this is usually in CTD module 1 and 2. 

This approach of assessing the efficacy and safety separately evaluates the average effect (efficacy or safety) in the entire study population. The benefit or risk can not be easily translated into the individual patient level. In clinical trials, it is almost impossible to decide if a drug is good (the benefit outweighs the risk) for a specific patient. We have to wait for the aggregate data to determine the benefit and risk on a group level. 

With advances in precision medicine and pharmacogenomics, we hope that in the future, within-patient benefit-risk evaluation can be performed. In the present days (perhaps the foreseeable future), the benefit-risk evaluation (or efficacy-safety evaluation) will still be primarily based on the population level to assess the average group effect. 
  • Average effect (Using Patients to Analyze Outcomes)
  • Subgroup analyses to identify the prognostic factors (phenotypes) to help identify the patients who will more likely to respond to the therapy with fewer side effects
  • Targeted therapies, Precision Medicine to identify the genetic biomarkers (genes) to help identify the subgroup of patients who will more likely to respond to the therapy with few side effects  
  • Individual effect - within patient benefit-risk evaluation 
Even with targeted therapy, it is still not possible to be certain if a therapy will be good (the benefit outweighs the risks) for a specific patient. 

For the J&J Covid-19 vaccine issue, it seems to be clear that the vaccine does appear to increase the risk of the rare blood clot - CVST. Since the CVST is so rare, the benefit of receiving the Covid-19 vaccine outweighs the risk of the rare blood clot - this assessment is on the population as a whole. When it comes to the individual person, it will be his/her own choice - the risk is small, but maybe there.  

Monday, April 19, 2021

Restricted Mean Survival Time (RMST) for Handling the Non-Proportional Hazards Time to Event Data

Time to event analysis (or traditionally survival analysis) is one of the most common analyses in clinical trials. In general, the time to event analysis relies on the assumption of the proportional hazards. However, quietly frequently, we may find that the proportional hazards assumption is violated, especially in many immuno-oncology trials. When the proportional hazards assumption is violated, alternative approaches may be needed to analyze the data to achieve statistical power. As discussed in the previous post "Non-proportional Hazards: how to analyze the time-to-event data?", one of the alternative approaches is the restricted mean survival time (RMST) method. 

RMST is one of the Kaplan-Meier-based methods and is essentially calculating and comparing AUCs under Kaplan-Meier Curves for different treatment groups or different comparative groups. It has been said that RMST analysis has the following advantages:
  • Model-free, robust, and easily interpretable treatment effect information
  • Produces radically powerful patterns of difference as has been observed in some recent Oncology clinical trials
  • Accepted approach by regulatory agencies and industry leaders
RMST has been mentioned in the latest FDA guidance for Industry (2020): Acute Myeloid Leukemia: Developing Drugs and Biological Products for Treatment as an alternative approach to analyzing the data when the non-proportionality hazards occur (e.g., plateauing effect). 

"Plateauing Effect

Trials designed to cure AML often result in survival contours characterized by an initial drop followed by a plateauing effect after some time point post randomization. This is an example of nonproportional hazards. While the log-rank test is somewhat robust to nonproportionality, it generally results in loss of power. Furthermore, nonproportionality can cause difficulty in describing the treatment effect. FDA is open to discussion about analyses based on other approaches, such as weighted Cox regression or other weighted methods, or summarizing the treatment effect using restricted mean survival time (RMST) or landmark survival analysis. Plans that use these alternative approaches should include:
    • justification for what constitutes clinically meaningful difference,
    • justification of design parameters, such as sample size and follow-up duration, based on this endpoint, and
    • justification for the value of the threshold that will be used to calculate the RMST.
RMST analysis has also been used as a primary analysis approach or for sensitivity analysis in FDA reviews: 

In NDA of Baloxavir marboxil in treatment of acute, uncomplicated influenza, both applicants and the FDA reviewer analyzed the data using RMST. It stated:
Restricted mean survival time (RMST) up to Day 10 was estimated for each treatment group along with the difference between RMST in the two treatment groups. RMST is a measurement of the average survival from time 0 to a specified time point (e.g., 10 days) which is equivalent to the area under the Kaplan-Meier curve from the beginning of the study through that time point.

At an FDA CDRH Medical Devices Advisory Committee Circulatory System Panel meeting in 2019, the independent statistical consultant addressed the analysis issue when the proportional hazards assumption is violated:

The proposal they made was the restricted mean survival time. The restricted mean survival time is area under curve. Please note the word restricted. Mean survival time is over a period of time, according to the rules that have been laid out, so that you're not looking, like with proportional hazards, over all the follow-up that could have possibly happened or in binary where you're only looking at the patients that survive. The restricted mean would say we're going to look between, let's say, 0 and 5 years because we have sufficient information to make that kind of assessment.

The paper showed that the restricted mean has just as much power as proportional hazards when the assumptions are there for proportional hazards, and then has more power when the assumptions are violated.

There's also some advantages in terms for clinicians, in terms of explaining this to the patient. It's hard to talk about hazards or number needed to treat. But if you could say to a patient over a 60-month period the average survival time is 55 months with Device A versus 52 months with Device B, now they can look at what their life is going to look like in the next 60 months and make a decision.

Unfortunately, it was not me who noticed this. This was actually from a presentation by FDA. Several very smart statisticians had talked about the restricted mean and have made recommendations on using it for both proportional violations and for its interpretation.

In FDA Briefing Document for Oncologic Drugs Advisory Committee Meeting (December 17, 2019) to review Olaparib for the maintenance treatment of adult patients with deleterious or suspected deleterious germline BRCA mutated (gBRCAm) metastatic adenocarcinoma of the pancreas

FDA performed a test to evaluate whether the proportional hazard assumption was met. This test failed to detect evidence of non-proportionality; however, such a test may lack power to detect non-proportionality due to the small sample size. The Kaplan-Meier curves of PFS appear to show some degree of nonproportionality. The curves did not show separation until approximately 4 months, after approximately 53% of patients either had events or were censored. FDA performed additional sensitivity analyses by applying the restricted mean survival time (RMST) method using different truncation points (15 months and 18 months). The truncated time was selected (15 or 18 months) such that approximately 8-12% patients remained at risk. Based on the truncation times, the estimated RMST difference in PFS between arms ranged from 2.6 months (95% CI: 0.9, 4.3) to 3.1 months (95% CI: 1.0, 5.2). The range of the RMST differences again demonstrated great variation in the difference in PFS and the lower ends did not suggest that there was a clinically meaningful difference.

Thanks to the software, RMST analyses can be easily implemented in SAS or R. In the latest version (version 15.1 or above) of SAS/Stat, RMST is included in SAS Proc LIFETEST with RMST option and Proc RMSTREG. See a nice paper by 
With R, the package for RMST analysis is survRM2 that is developed by Hajime Uno from Dana-Farber Cancer Institute

For RMST analysis, it is important to select the cut-off value (tau) for the truncated time. The different selection of taus will give different results. The selection of tau can sometimes be arbitrary. In an FDA briefing document above, the FDA statistician chose the truncated time such that approximately 8-12% of patients remained at risk.

There are different ways to calculate the RMST:

  • Non-parametric method
  • Regression Analysis Method
  • Pseudo-value Regression Method
  • IPCW Regression - Inverse Probability of Censoring Weighting (IPCW) regression
  • Conditional restricted mean survival time (CRMST)

According to the paper by Guo and Liang (2019) "Analyzing Restricted Mean Survival Time Using SAS/STAT®", non-parametric analysis can be implemented using Proc Lifetest; regression analysis, pseudo-value regression, and IPCW regression can be implemented using SAS Proc RMSTREG. 

FDA statisticians also proposed an approach 'conditional restricted mean survival time' or CRMST. This approach was described in the paper by Qiu et al (2019) "Estimation on conditional restricted mean survival time with counting process" and also in a presentation by Lawrence and Qiu (2020) Novel Survival Analysis When Hazards Are Nonproportional and/or There Are Multiple Types of Events. CRMST can allow the AUC under K-M curves to be calculated from an interval time (not necessarily to be started from the 0 time). They claim CRMST is better for event-driven studies where the time to the first event is the interest. They concluded the following: 
CRMST possesses all the desirable statistical properties of RMST. In particular, it does not rely on proportional hazard assumption. In addition, CRMST measures an average event-free time in the time range at issue and has straightforward interpretation. In case that two survival curves cross, CRMST can be estimated separately before and after crossing and the CRMST differences can be used to assess benefit versus harm.

Further Reading:

Monday, April 05, 2021

Non-proportional Hazards: how to analyze the time-to-event data?

Time to event data is one of the most common data types in clinical trials. Traditionally, the log-rank test is used to compares the survival curves of two treatment groups.; the Kaplan Meier survival plot is used to illustrate the totality of time-to-event kinetics, including the estimated median survival time;  the Cox-proportional hazards model is employed to provide the estimated relative effect (i.e., hazard ratio) between treatment arms. The performance of these analyses largely depends on the proportional hazards (PH) assumption – that the hazard ratio is constant over time. In other words, the hazard ratio provides an average relative treatment effect over time.

Before the time to event data is analyzed, it is typical for statisticians to check the proportional hazards assumption. Various methods can be used to check the proportional hazards assumptions - see a previous post "Visual Inspection and Statistical Tests for Proportional Hazard Assumption".

Recently we have seen more examples of the time to event data not following the proportional hazards assumption, even more examples in immuno-oncology clinical trials. 

It is not the end of the world if the proportional hazards assumption is violated, various approaches have been proposed to handle the time to event data with non-proportional hazards. 

In practice, it is pretty common that in the statistical analysis plan, we prespecify the log-rank test to calculate the p-values and then use Cox-proportional hazards regression model to calculate the hazard ratio, its 95% confidence interval, and p-value - I call this 'Splitting p-value and estimate of the treatment difference". Two different p-values will be calculated: one from the log-rank test and one from the Cox regression. If the proportional hazards assumption is met, it is better to use the p-value from the Cox regression since all estimates and p-value are coming from the model. However, When the proportional hazard assumption is violated, the Cox-proportional hazard model may no longer be the optimal approach to determine treatment effect and the Kaplan-Meier estimate of median survival may not be the most valid measure to summarize the results. 

In a website post "Testing equality of two survival distributions: log-rank/Cox versus RMST", it stated:
“One thing to note is that the log-rank test does not assume proportional hazards per se. It is a valid test of the null hypothesis of equality of the survival functions without any assumptions (save assumptions regarding censoring). It is however most powerful for detecting alternative hypotheses in which the hazards are proportional.”
It is true that the log-rank test does not depend on the proportional hazards assumption. The log-rank test is still a valid test of the null hypothesis of equality of the survival functions without any assumptions even though that the log-rank test may not be optimal under non-proportional hazards. 

In a public workshop "Oncology Clinical Trials in the Presence of Non-Proportional Hazards" organized by Duke in 2018, Dr. Rajeshwari Sridhara from Division of Biometrics V, CDER/FDA stated (@40:45 of the youtube video) that in the non-proportional hazards situation, FDA is ok with presenting the p-value from the log-rank test and hazard ratio to measure the treatment difference. 

At this same workshop "Oncology Clinical Trials in the Presence of Non-Proportional Hazards", ASA Biopharmaceutical Section Regulatory-Industry Statistics Workshop presented their work and proposed the 'max-combo' test as the alternative method to address the non-proportional hazards situation. The “max-combo” test is based on Fleming-Harrington (FH) weighted log-rank statistics. The max-combo test tackles some of the challenges due to non-proportional hazards as it is able to robustly handle a range of non-proportional hazard types, can be pre-specified at the design stage, and can choose the appropriate weight in an adaptive manner (i.e. is able to address the control of family-wise Type I error). In workshop summaries, Max-Combo Test Design was described as the following:

Knezevic & Patil has a paper describing a SAS macro to perform Max-Combo test (or Combination weighted log-rank tests) "Combination weighted log-rank tests for survival analysis with
non-proportional hazards" (2020 SAS Global Forum). 

The NPH workshop has presented or published their work on numerous occasions, here is a list: 
In addition to the Max-Combo test, there are several other methods for handling the non-proportional hazards situation. 
  • RMST (restricted mean survival time): according to a presentation by Lawrence et al from FDA, The idea of Restricted Mean Survival Time (RMST) goes back to Irwin (1949) and is further implemented in survival analysis by Uno et al. (2014). RMST is defined as the area under the survival curve up to t*, which should be pre-specified for a randomized trial. RMST may be loosely described as the event free expectancy over the restricted period between randomization and a defined, clinically relevant time horizon, called t*. RMST analyses are now built into the SAS procedures with Proc Lifetest and Proc RSMTREG. See a paper by Guo and Liang (2019) "Analyzing Restricted Mean Survival Time Using SAS/STAT®"
  • Piecewise exponential regression allows for an early and late effect of treatment comparison. it is especially useful when the non-proportional hazards pattern is cross-over. Piecewise exponential regression can be fitted with SAS Proc MCMC and R package pch
  • Estimation via the average hazard ratios (AHR) method of Schemper (2009) and the average regression effects (ARE) method of Xu and O’Quigley (2000) - the method can be implemented using the COXPHW package in R. COXPHW package is described as:
This package implements weighted estimation in Cox regression as proposed by Schemper, Wakounig and Heinze (Statistics in Medicine, 2009, doi: 10.1002/sim.3623). Weighted Cox regression provides unbiased average hazard ratio estimates also in case of non-proportional hazards. The package provides options to estimate time-dependent effects conveniently by including interactions of covariates with arbitrary functions of time, with or without making use of the weighting option. For more details we refer to Dunkler, Ploner, Schemper and Heinze (Journal of Statistical Software, 2018, doi: 10.18637/jss.v084.i02).

in a presentation by Kaur et al "Analytical Methods Under Non-Proportional Hazards: A Dilemma of Choice", the following methods were described: 

Earlier this year, Mehrotra and West published a paper to describe their proposed method (5-START) to handle the heterogeneity of the patient population and potential non-proportional hazards (Lin et al (2021) Survival Analysis Using a 5-Step Stratified Testing and Amalgamation Routine (5-STAR) in Randomized Clinical Trials or here ):

"The power of the ubiquitous logrank test for a between-treatment comparison of survival times in randomized clinical trials can be notably less than desired if the treatment hazard functions are non-proportional, and the accompanying hazard ratio estimate from a Cox proportional hazards model can be hard to interpret. Increasingly popular approaches to guard against the statistical adverse effects of non-proportional hazards include the MaxCombo test (based on a versatile combination of weighted logrank statistics) and a test based on a between-treatment comparison of restricted mean survival time (RMST). Unfortunately, neither the logrank test nor the latter two approaches are designed to leverage what we refer to as structured patient heterogeneity in clinical trial populations, and this can contribute to suboptimal power for detecting a between-treatment difference in the distribution of survival times. Stratified versions of the logrank test and the corresponding Cox proportional hazards model based on pre-specified stratification factors represent steps in the right direction. However, they carry unnecessary risks associated with both a potential suboptimal choice of stratification factors and with potentially implausible dual assumptions of proportional hazards within each stratum and a constant hazard ratio across strata.
We have developed and described a novel alternative to the aforementioned current approaches for survival analysis in randomized clinical trials. Our approach envisions the overall patient population as being a finite mixture of subpopulations (risk strata), with higher to lower ordered risk strata comprised of patients having shorter to longer expected survival regardless of treatment assignment. Patients within a given risk stratum are deemed prognostically homogeneous in that they have in common certain pre-treatment characteristics that jointly strongly associate with survival time. Given this conceptualization and motivated by a reasonable expectation that detection of a true treatment difference should get easier as the patient population gets prognostically more homogeneous, our proposed method follows naturally. Starting with a pre-specified set of baseline covariates (Step 1), elastic net Cox regression (Step 2) and a subsequent conditional inference tree algorithm (Step 3) are used to segment the trial patients into ordered risk strata; importantly, both steps are blinded to patient-level treatment assignment. After unblinding, a treatment comparison is done within each formed risk stratum (Step 4) and stratum-level results are combined for overall estimation and inference (Step 5)."
Non-proportional hazards and the NPH pattern are usually identified after the study unblinding, which poses the challenges for pre-specifying the best approach to analyze the time to event data with non-proportional hazards. The safest way is to prespecify both the Log-rank test and the Cox proportional hazards regression. If the non-proportional hazards assumption is violated, the p-values from the log-rank test will be used as a measure of the significance. One can also pre-specify the Max-Combo method as the primary method regardless of the NPH assumption