Friday, November 03, 2023

Walk Distance versus Timed Walk - endpoints for measuring patients' function in clinical trials

Drug development is now moved to the patient-focused era. Patient-focused drug development (PFDD) is a systematic approach to help ensure that patients’ experiences, perspectives, needs, and priorities are captured and meaningfully incorporated into drug development and evaluation. With PFDD, the endpoint or outcome measure needs to be meaningful, it should reflect or describe how the patient feels, functions, and survives.

Various tools can be used to measure functions. The most commonly used functional measure may be the measure of the walk distance or the walking speed.

Measuring the walk distance by fixing the time: six-minute walk test (6MWT) to measure the distance the patient can walk in six minutes (6MWD), two-minute walk test (2MWT) to measure the distance the patient can walk in two minutes (2MWD). 6MWD and 2MWD can provide a functional, therapeutic response and prognostic data that is valuable in the care of patients with respiratory, cardiac, and neurological diseases. It can also be used as the endpoint for clinical trials to evaluate the treatment differences.

There is also a 10-minute walk test to measure fatigability and walking economy, but it is not commonly used as primary efficacy endpoint in clinical trials. 

Measuring the time/speed by fixing the distance: The 10-meter Walk Test is a performance measure used to assess walking speed in meters per second over a short distance. It can be employed to determine functional mobility, gait, and vestibular function. Timed 25 Foot Walk (T25FW) is a quantitative mobility and leg function performance test based on a timed 25-walk. The time to complete 25-foot walk can be used to calculate the walk speed (ft/s). 


Examples/applications:

6MWT/6MWD

The 6MWT is a sub-maximal exercise test used to assess exercise capacity and endurance. The distance covered over a time of 6 minutes is used as the outcome by which to compare changes in performance capacity. There is a specific guideline developed by ATS (American Thoracic Society): Guidelines for the Six-Minute Walk Test.

6MWT/6MWD was the primary efficacy outcome measure in pivotal studies in pulmonary arterial hypertension (PAH) and pulmonary hypertension associated with interstitial lung disease (PH-ILD).

Bridgebio had a pivotal study to assess the efficacy and safety of the acoramidis in treatment of
ATTRibute-CM (a heart disease) with two parts: 6MWD was the primary efficacy endpoint for part 1 of the study and

Part 1 of the study failed to demonstrate the treatment difference in 6MWD

Part 2 of the study successfully demonstrate the treatment difference in win ratio in clinical events (deaths and cardiovascular related hospitalization)

Alnylam's pivotal study (APOLLO-B study) demonstrated the statistical significant difference in primary efficacy endpoint of 6MWD at week 52. However, the magnitude of the treatment difference was merely 14.7 meters. The study results were published in New England Journal of Medicine and had a positive vote in favor of the approval by the Advisory Committee, however, FDA declined the approval

6MWT/6MWD may also be used in neurology diseases, for example: The 6-minute walk test and other endpoints in Duchenne Muscular Dystrophy: longitudinal natural history observations over 48 weeks from a multicenter study


2MWT/2MWD:

Both 6MWT and 2MWT are clinical assessments to evaluate a patient's functional capacity and endurance, particularly in individuals with cardiopulmonary or musculoskeletal conditions. Compared to 6MWT/6MWD, 2MWT/2MWD was less commonly used in clinical trials. However, 2MWT is a shorter, more focused test designed to quickly assess walking capacity and is often used in situations where a shorter test is preferred. 2MWT can be conducted in a smaller space, making it more suitable for clinics or confined settings. 2MWT is particularly useful for assessing functional capacity in situations where time constraints or physical limitations may necessitate a shorter test, and offers a quicker assessment of walking capacity and can be used for patients who may have difficulty completing a longer test.

Two- and 6-minute walk tests assess walking capability equally in neuromuscular diseases

Grifols conducted a pivotal study to assess the efficacy of IGIV in the treatment of post-polio syndrome and used 2MWT as the primary efficacy endpoint. The study is still ongoing. 

MedDay Pharmaceuticals SA conducted a phase 3 study "MD1003-AMN MD1003 in Adrenomyeloneuropathy" with 2MWD as primary efficacy endpoint

Adamas Pharmaceuticals conducted a phase 3 study "Safety and Efficacy of ADS-5102 in Multiple Sclerosis Patients With Walking Impairment" where 2MWT used as the secondary endpoint (T25FW used as the primary endpoint)

The 10 Metre Walk Test


Sarepta Therapeutics recently released their confirmatory study results of gene therapy for the treatment of DMD (Duchenne Muscular Disease) and 10-meter walk test was one of the secondary efficacy endpoints. The 10-meter walk test results are shown here. The treatment differences are expressed in time (seconds). While all treatment differences are statistically significant, the clinical meaningfulness needs to be vetted by the experts and the regulators. 



Timed 25 Foot Walk (T25FW)


The T25FW is a quantitative mobility and leg function performance test based on a timed 25-walk. The patient is directed to one end of a clearly marked 25-foot course and is instructed to walk 25 feet as quickly as possible, but safely. The time is calculated from the initiation of the instruction to start and ends when the patient has reached the 25-foot mark. The task is immediately administered again by having the patient walk back the same distance. Patients may use assistive devices when doing this task.

The drug AMPYRA® (dalfampridine) was approved for improving walking in adult patients with multiple sclerosis (MS). The drug label stated that the T25FW was the primary efficacy endpoints:

The primary measure of efficacy in both trials was walking speed (in feet per second) as measured by the Timed 25-foot Walk (T25FW), using a responder analysis. A responder was defined as a patient who showed faster walking speed for at least three visits out of a possible four during the double-blind period than the maximum value achieved in the five non-double-blind no treatment visits (four before the double-blind period and one after). 

Acorda Therapeutics conducted phase 3 studies "Study of Fampridine-SR Tablets in Multiple Sclerosis Patients" and "Study of Oral Fampridine-SR in Multiple Sclerosis" where T25FW was used as the primary efficacy measure.

The paper by Cohen et al "A Phase 3, double-blind, placebo-controlled efficacy and safety study of ADS-5102 (Amantadine) extended-release capsules in people with multiple sclerosis and walking impairment" stated the following:
Walking speed ft/s was used for T25FW since walking speed is more normally distributed as compared to walking time, and is therefore a preferred approach. A 20% change in T25FW is considered a meaningful change in patients with MS


All four measures discussed above (6MWD, 2MWD, 10-meter walk test, T25FW) can be an acceptable endpoint for confirmatory trials. Which measure to use in a specific trial depends on the indication and the study population. The endpoint selection should be discussed with the review division of regulatory agencies such as FDA. 

Friday, October 20, 2023

Human Challenge Study Design in Action - a Dengue Fever vaccine trial

A human challenge study, also known as a controlled human infection model (CHIM), is a type of clinical research study in which healthy volunteers are intentionally exposed to a specific pathogen (such as a virus, bacterium, or parasite) under controlled conditions. The primary goal of these studies is to better understand the pathogen's behavior, the human immune response to it, and to test the effectiveness of potential treatments, vaccines, or preventive measures. Human challenge studies can provide valuable insights into disease progression, immunity, and treatment efficacy in a controlled and ethical manner.

These studies are typically conducted under strict ethical and safety guidelines to minimize the risk to participants. Participants are closely monitored, and their informed consent is obtained. Human challenge studies have been used to study a variety of diseases, including influenza, malaria, Dengue fever, and COVID-19, among others. They play a crucial role in advancing medical and scientific knowledge and can accelerate the development of treatments and vaccines.

A human challenge study was mentioned as an alternative clinical trial design at the beginning of the COVID-19 pandemic when the world was desperate to find an effective and safe vaccine. I wrote an article about this: "Human Challenge Study Design for Covid-19 Vaccine Clinical Trials?"

Just this morning, Janssen Announces Promising Antiviral Activity Against Dengue in a Phase 2a Human Challenge Model. The results were from a phase 2a study titled "A Phase 2a, Randomized, Double-blind, Placebo Controlled Trial to Evaluate the Antiviral Activity, Safety, and Pharmacokinetics of Repeated Oral Doses of JNJ-64281802 Against Dengue Serotype 3 Infection in a Dengue Human Challenge Model in Healthy Adult Participants" that was posted on clinicaltrials.gov. Unfortunately, the clinical trial registration did not contain any description of the 'Challenge' part (i.e., how the healthy volunteers are exposed to the infectious agents (in this case, the Dengue virus). We will just need to wait for the formal publication of the study to know the details. 

In a paper by Porter et al "A human Phase I/IIa malaria challenge trial of a polyprotein malaria vaccine", the whole details about the human challenge study including the 'challenge' part were discussed. The 'sporozoite challenge' to the healthy volunteers was described below: 

 

Friday, October 13, 2023

Drugs Approved by FDA Despite Failed Trials or Minimal/Insufficient Data

I have been trying to collect the cases of that drugs were approved by the FDA despite the failed trials or minimal/insufficient data. For drugs treating rare diseases or diseases with unmet medical needs, the FDA may apply flexibility in approving the drug with loosened criteria. 

For diseases with clearly unmet medical needs such as ALS (Amyotrophic Lateral Sclerosis) and Alzheimer's disease, FDA officials have recently emphasized the urgent need for new treatments and pledged to use maximum "regulatory flexibility" when reviewing the NDA/BLA packages. By applying the maximum "regulatory flexibility", FDA has approved some drugs which do not meet the agency's traditional approval standards. Some of the approvals are really controversial and make me wonder if there is any boundary for the maximum "regulatory flexibility". 

The following paper on BioSpace.com listed six drugs that earned FDA approval without substantial evidence of effectiveness.

6 Drugs Approved Despite Failed Trials or Minimal Data
  • Ipsen’s Sohonos (palovarotene) for the ultra-rare genetic disease fibrodysplasia ossificans progressive (FOP)
  • Sarepta’s Elevidys as the first gene therapy for Duchenne muscular dystrophy (DMD)
  • Biogen's Qalsody (tofersen) to treat patients with superoxide dismutase 1 (SOD1)-ALS, a rare subtype of the fatal neurodegenerative disease
  • Biogen and Eisai got the nod for Aduhelm (aducanumab) for Alzheimer's diease,
  • Jazz Pharmaceuticals and PharmaMar’s Zepzelca (lurbinectedin) for small cell lung cancer (SCLC) that had progressed on or after platinum-based chemotherapy
  • Acadia Pharmaceuticals’ Nuplazid (pimavanserin) to treat hallucinations and delusions associated with psychosis in Parkinson’s disease.
Some of the approvals gave the sponsors the false hope that an innovative drug could be approved by the FDA even if the study failed to demonstrate the effectiveness as long as the drug was for the treatment of diseases with urgent unmet medical needs. A recent story about BrainStorm's ALS drug is exactly the case about this. 

Friday, October 06, 2023

MCID (Minimum Clinical Important Difference) for 6MWD - how low can we go?

I went back to watch the FDA CRDAC (Cardiovascular and Renal Drugs Advisory Committee) meeting to discuss Alnylam's drug Patisiran for the treatment of ATTR-CM (Transthyretin Amyloidosis) - a rare form of heart disease. The meeting discussion was centered on the clinical meaningfulness of the efficacy measures in the primary efficacy endpoint of 6MWD (how many meters patients can walk in 6 minutes) and the secondary endpoint of KCCQ - a patient-reported quality of life measure. 

The sponsor, Alnylam, conducted a phase III study called "APOLLO-B: A Study to Evaluate Patisiran in Participants With Transthyretin Amyloidosis With Cardiomyopathy (ATTR Amyloidosis With Cardiomyopathy)". The study results showed statistically significant differences in 6MWD and in KCCQ total score. However, the magnitude of the treatment differences was very small: 14.7 meters in 6MWD and 3.7 points in KCCQ at month 12.

To judge if the treatment difference is clinically meaningful, people will compare the magnitude of the treatment differences from the study with the MCID (minimal clinically important difference). MCID. MCID is the smallest change in a treatment outcome that individual patients would identify as important and which would indicate a change in the patients' management.  The MCID is a patient-centered concept that captures both the magnitude of the improvement and the value patients place on the change.  In other words, the MCID is the smallest amount of change in the score of a scale recognized by the patient without considering the side effects and cost. 

In FDA's briefing book for CRADAC meeting, FDA casted doubts about the Patisiram's efficacy: 
The 6MWT, a performance outcome (PerfO), is a practical simple test that measures the distance that a patient can quickly walk on a flat, hard surface in a period of 6 minutes (the 6MWD). It evaluates the global and integrated responses of all the systems involved during exercise. The results of the APOLLO-B trial showed a statistically significant but small treatment effect for the primary efficacy endpoint. Subjects treated with patisiran experienced an average decrease in their 6MWD of 13 m at Month 12 from an average 6MWD of 361 m at baseline, while subjects in the placebo arm experienced an average decrease in their 6MWD of 31 m at Month 12 from an average 6MWD of 375 m at baseline. The change from baseline at Month 12 in 6MWT (Hodges-Lehmann [HL] estimate of median difference) for patisiran vs. placebo was 14.7 m (95% confidence interval [CI] 0.7, 28.7; p-value 0.04). Literature has reported a range of meaningful differences (22 to 90 m) reflective of the heterogeneity in cardiomyopathy patients (Mathai et al. 2012; Shoemaker et al. 2012).
 The KCCQ, a patient-reported outcome (PRO) and a disease-specific measure for HF, is a 23-item self-administered questionnaire developed to measure the patient’s perception of their health status, which includes heart failure symptoms, impact on physical and social function, and how heart failure impacts their quality of life (QOL) within a 2-week recall period. The KCCQ-OSS has a 0-100 transformed score range where higher scores reflect better health status (based on the Physical Limitation, Symptom Frequency, Symptom Burden, Quality of Life and Social Limitations Domain Scores). In the APOLLO-B trial,the treatment effect for the first secondary efficacy endpoint, change from baseline at Month 12 in KCCQ-OSS was small (3.7 points on a 0 to 100 transformed score range; 95% CI 0.2, 7.2; p-value 0.04). On average, subjects treated with patisiran had an increase in KCCQ-OSS of 0.3 points at Month 12 from the average baseline score of 69.8 points, while subjects in the placebo arm had a decrease in KCCQ-OSS of 3.4 points at Month 12 from the average baseline score of 70.3 points.  

Sponsor, Alnyam's briefing book and presentation spent a lot of effort to defend that the small, but statistically significant treatment differences are clinically meaningful. 


Sponsor attempted to derive an MCID using KCCQ category as an anchor based on the data from the study itself (APOLLO-B study). 


Not surprisingly, the MCID they generated were much smaller (MCID in the range of 7 - 8 meters) than the MCIDs reported in the literature. If the MCID is indeed in the range of 7 - 8 meters, the 14.7 meters (treatment difference observed in Apollo-B study) would be clinically meaningful. 



During the FDA advisory committee meeting, most of the members were not convinced by sponsor's presentation to defend the clinical meaningfulness of  small treatment difference in 6MWD (about 14 meter). However, majority of them (9-3) still voted in favor of the Patisiran's efficacy and the benefit-risk profile. 

Pfizer's tafamidis is the only approved drug for the treatment of ATTR-CM. According to the product label, the treatment difference in 6MWD was much larger - 76 meters with 95% confidence interval 58, 94 meters at month 30. 

For the same 6MWD, the MCID may be different depending on the treating diseases, different patient population, whether patients receiving the background therapies,... However, a treatment difference of 14 meters is still not a convincing number to be clinical meaningful. Putting on the relative scale, the 14 meters in patients with baseline 6MWD 361 meter is less than 5%. It is difficult to convince people a treatment difference less than 5% is clinically meaningful. 

I am particularlly interested in the MCID of 6MWD in lung diseases (especially the pulmonary arterial hypertension). 

Anne E. Holland (2014) "An official European Respiratory Society/American Thoracic Society technical standard: field walking tests in chronic respiratory disease" stated
“Available evidence suggests a minimal important difference (MID) of 30 m for the 6MWD in adults with chronic respiratory disease.”
Jude Moutchia (2023) "Minimal Clinically Important Difference in the 6-minute-walk Distance for Patients with Pulmonary Arterial Hypertension" found:
The minimal clinically important difference in the derivation sample was 33 meters (95% confidence interval, 27–38), which was almost identical to that in the validation sample (36 m [95% confidence interval, 29–43]). The minimal clinically important difference did not differ by age, sex, race, pulmonary hypertension etiology, body mass index, use of background therapy, or World Health Organization functional class.

Here is a table containing some literatures with estimated MCID. The MCID was found to be in the range of 20 - 54 meters depending on the indication/disease. 

Study/Article

Indication/Disease

MCID Range

MCID Midpoint

Chan (2015)

ARF

20 – 30

25

du Bois (2011)

IPF

24 – 45

35

Gilbert (2009)

PAH

41

41

Granger (2015)

Lung Cancer

22 – 42

32

Holland (2009)

DPLD/IPF

29 – 34

32

Holland (2010)

COPD

25

25

Mathai (2012)

PAH

33

33

Nathan (2015)

IPF

22 – 37

30

Polkey (2013)

COPD

30

30

Puhan (2008)

COPD

35

35

Puhan (2011)

COPD

24 – 28

26

Redelmeier (1997)

CLD

54

54

Swigris (2010)

IPF

28

28


Latest update: 

In the end, the FDA did not approve Patisiran for the treatment of ATTR-CM  because the treatment difference in 6MWD was too small (way below the MCID)  and not clinically meaningful even though the FDA advisory committee voted in favor of the Patisiran's benefit and there was no issue with the safety and the manufacturing. 

Alnylam Announces Receipt of Complete Response Letter from U.S. FDA for Supplemental New Drug Application for Patisiran for the Treatment of the Cardiomyopathy of ATTR Amyloidosis

 "In its Complete Response Letter (CRL), the regulator said that Alnylam had not provided enough evidence of the therapy’s benefit in the proposed indication. At the same time, the FDA did not flag any problems with patisiran’s clinical safety, drug quality, manufacturing processes or study conduct.

“The CRL indicated that the clinical meaningfulness of patisiran’s treatment effects for the cardiomyopathy of ATTR amyloidosis had not been established,” according to the company’s announcement. In light of the rejection, Alnylam will no long work toward an expanded label for Onpattro in the U.S."

Tuesday, August 15, 2023

Platform trials in action beyond oncology trials

One of the complex innovative designs is the study with 'Master protocol' that has been getting popularity in recent years. Master protocol simply refers to the study of more than one drug with a single protocol. In clinical trials with Master protocol, the protocol document stands alone, without reference to specific drugs. Protocol amendments (or study-specific protocol) are used to provide details of each drug. Master protocols establish a trial network with infrastructure in place to streamline trial logistics and improve data quality, and facilitate data sharing and new data collection. Master protocols develop a common protocol for the network that incorporates innovative statistical approaches to study design and data analysis.


Platform trial is one type of Master protocol and refers to establishing the trial infrastructure and master protocol as a perpetuating effort, with drugs entering and leaving the platform.

Two prominent references for master protocols and platform trials are the following: 

 "Master Protocols to Study Multiple Therapies, Multiple Diseases, or Both" in 2017 by Woodcock and LaVange defines Master protocols include three type of trials: Umbrella, basket, and platform trials. 



FDA procedural guidance in 2022 "Master Protocols: Efficient Clinical Trial Design Strategies to Expedite Development of Oncology Drugs and Biologics Guidance for Industry" provided the definition for Master protocols:
a master protocol is defined as a protocol designed with multiple substudies, which may have different objectives and involve coordinated efforts to evaluate one or more investigational drugs in one or more disease subtypes within the overall trial structure.
A master protocol may be used to conduct the trial or trials for exploratory purposes or to support a marketing application and can be structured to evaluate, in parallel, different drugs compared with their respective controls or to a single common control. The sponsor can design the master protocol with a fixed or adaptive design with the intent to modify the protocol to incorporate or terminate individual substudies within the master protocol. 

Master protocols can include basket trial, umbrella trial, and platform trial and these trials can be schematically displayed as the following: 

Basket trial

Umbrella Trial
Platform Trial:


In  a paper by Berry et al, "The Platform Trial An Efficient Strategy for Evaluating Multiple Treatments", the platform trial was compared to the traditional trial: 


While platform trial has its challenges, we see the implementation trials in action, beyond the oncology trials. Here are examples of platform trials in areas other than the oncology trials. 

HEALEY ALS Platform Trial
The HEALEY ALS Platform Trial is a perpetual multi-center, multi-regimen clinical trial evaluating the safety and efficacy of investigational products for the treatment of ALS. This trial is designed as a perpetual platform trial. This means that there is a single Master Protocol dictating the conduct of the trial.

In this trial, multiple investigational products for ALS will be tested simultaneously or sequentially. Each investigational product will be tested in a regimen. Each regimen consists of a placebo-controlled trial, meaning that the active investigational product and matching placebo will be tested in each regimen.

The study website is here: https://www.massgeneral.org/neurology/als/research/platform-trial

The study master protocol can be accessed here

Several testing drugs have been completed or graduated from the platform, please see the news releases

PrecISE (Precision Interventions for Severe and/or Exacerbation-Prone Asthma) Network Study
PrecISE is an adaptive platform trial under master protocol with common biomarker screening where drugs enter when available and discontinue based on futility analysis. According to clinicaltrials.gov, 5 novel interventions are currently listed as the testing drug for study: Medium Chain Triglycerides (MCT), Clazakizumab, Broncho-Vaxom, Imatinib Mesylate, Cavosonstat; each active will be tested against its own control.
the study design was described in the paper by Israel et al "PrecISE: Precision Medicine in Severe Asthma: An adaptive platform trial with biomarker ascertainment"
REMAP-CAP: A Randomised, Embedded, Multi-factorial, Adaptive Platform Trial for Community-Acquired Pneumonia

REMAP-CAP platform trial was originally designed for identifying the treatments for community-acquired pneumonia. After the Covid pandemic in 2020, REMAP-CAP has quickly implemented the Pandemic Appendix to the Core Protocol so that the platform can respond rapidly in the event of widespread disease resulting from the novel 2019 coronavirus (COVID-19).

ACTIV (NIH) Accelerating Covid-19 Therapeutic Interventions and Vaccines
Working in an unprecedented time frame, the Accelerating COVID-19 Therapeutic Interventions and Vaccines (ACTIV) public–private partnership developed and launched 9 master protocols between 14 April 2020 and 31 May 2021 to allow for the coordinated and efficient evaluation of multiple investigational therapeutic agents for COVID-19. The ACTIV master protocols were designed with a portfolio approach to serve the following patient populations with COVID-19: mild to moderately ill outpatients, moderately ill inpatients, and critically ill inpatients.

The study design was described in the paper by LaVange et al "Accelerating COVID-19 Therapeutic Interventions and Vaccines (ACTIV): Designing Master Protocols for Evaluation of Candidate COVID-19 Therapeutics"

RECOVER Clinical Trials - for Long Covid
Long COVID is defined as "a multifaceted disease that can affect nearly every organ system" and can manifest as new or worsening chronic health problems, including but not limited to heart disease, diabetes, kidney disease, hematologic issues, and mental and neurologic conditions. The signs, symptoms, and conditions continue or arise anew 4 weeks or more after the initial symptomatic or asymptomatic infection and may be relapsing and remitting.
NIH RECOVER trial is also using platform protocols/clinical trials to investigate treatments for long covid and master protocol is posted on the study website: https://trials.recovercovid.org/design. There will be multiple master protocols: RECOVER-VITAL, RECOVER-NEURO, RECOVER-AUTONOMIC (coming soon), RECOVER-SLEEP (coming soon) to investigate different
It is ironic that the RECOVER trial is started after NIH spent 2.5 years and $1 billion on long Covid research and failed to test meaningful long Covid treatments. See the article "Underwhelming’: NIH trials fail to test meaningful long Covid treatments — after 2.5 years and $1 billion "

Clinicaltrialsarena.com has an featured article "Platform trials: an opportunity for rare dystrophies or gene therapies?". The article discussed the possibility of platform trials in rare diseases and gene therapies. PaVeGT platform trial was listed: 

PaVe-GT: Paving the Way for Rare Disease Gene Therapies

PaVe-GT will develop and test AAV-9, with a different gene for each indication, using a single master protocol. The NCATS-led Platform Vector Gene Therapy (PaVe-GT) pilot project seeks to increase the efficiency of clinical trial startup by using the same gene delivery system and manufacturing methods for multiple rare disease gene therapies. We will make program results and regulatory documents publicly available, with the intention of benefiting future gene therapy clinical trials for very rare diseases.

 In a paper by Collignon (2022) "An Economic Perspective on Platform Trials—The Gift and the Curse", the following conclusion was made about the platform trial: 

"...sharing a common control group in a platform trial can be viewed as a gift and a curse, and choosing whether to implement a platform trial or a more standard development program is complex since it is contingent on numerous factors that are disease and context dependent. To make such a decision, a clear framework needs to be implemented. Decision-making in the pharmaceutical industry is becoming increasingly more quantitative, and in practice, both development approaches would be compared according to a series of standard metrics, including
(1) duration of the clinical program;
(2) cost and expected reward of the clinical program;
(3) flexibility of the clinical program (eg, reporting of the analyses corresponding to the different treatments while maintaining data and trial integrity, ability to add sites, and so forth); (4) probability of success for each treatment, for all treatments, or for at least 1 treatment, assuming all or some of the treatments have a certain efficacy;
(5) probability of (multiple) false-positive findings (eg, achieving statistical significance) for each treatment or for at least 1 treatment, assuming all or some of the treatments are comparator-like; and
(6) the extent to which a positive readout informs the success of potential subsequent trials and the ability to de-risk the next phase of development.
Using these metrics, governance boards would then make their decisions on whether to progress developing a treatment according to a range of diverse considerations, such as portfolio opportunities and investment priorities."

Tuesday, August 01, 2023

Mediation Analysis vs. Landmark Analysis for Clinical Trial Data

Mediation analysis and Landmark analysis are two valuable statistical tools in different contexts, providing insights into different aspects of data analysis. Both methods can be utilized to investigate if a surrogate endpoint or short-term measure can predict the long-term clinical endpoint or to investigate if short-term measures can be mediators for the long-term clinical endpoint. Both mediation analysis and landmark analysis are useful tools for post-hoc, exploratory analyses, but not the primary analysis method for the primary efficacy endpoint. 

Mediation analysis is a statistical method used to explore and understand the mechanism or process through which an independent variable influences a dependent variable. It helps to determine whether the effect of an independent variable on a dependent variable is mediated (i.e., transmitted through) one or more intermediate variables, often referred to as mediators.

In mediation analysis, the main focus is on understanding the causal chain of relationships between variables. It involves three key components:
  • Independent variable (X): The variable that is hypothesized to influence the dependent variable directly or indirectly through one or more mediators.
  • Mediator (M): The variable(s) that mediates the relationship between the independent variable and the dependent variable. It explains how and why the independent variable affects the dependent variable.
  • Dependent variable (Y): The variable that is influenced by the independent variable either directly or indirectly through the mediators.
The analysis aims to estimate the direct effect of the independent variable on the dependent variable and the indirect effect mediated through the mediator(s). Several statistical techniques can be used for mediation analysis, such as regression-based methods (e.g., ordinary least squares regression) or more advanced methods like structural equation modeling.

In our published paper "Contemporary risk scores predict clinicalworsening in pulmonary arterial hypertension - Ananalysis of FREEDOM-EV",  mediation analysis was used to test the hypothesis that improvements in risk score (a surrogate endpoint) contributed to reduced likelihood for clinical worsening (a long-term clinical endpoint). 
 
In a paper by Blette et al, "Is low-risk status a surrogate outcome in pulmonary arterialhypertension? An analysis of three randomised trials ", the mediation analysis was used to investigate surrogacy. The author stated:
"we performed a mediation analysis considering each candidate surrogate as an intermediate outcome. To avoid inconsistent mediation (ie, when direct and indirect effects cancel each other out, the direct effect is even larger than the total effect, or other situations that can result in a negative proportion mediated), we first empirically tested four criteria to justify the mediation analysis. Next, Cox proportional hazards models were fit for the clinical worsening and survival outcomes conditional on each candidate surrogate separately, as well as treatment, corresponding baseline risk score, and a fixed effect variable with three levels for trial membership (to allow for similarity within each trial). Results from these models were combined with parallel models that did not condition on the candidate surrogates, but otherwise conditioned on the same set of variables to perform the difference method for mediation, estimating the total effect and direct effects of treatment on clinical worsening and survival, as well as the indirect effects through each candidate surrogate. The proportion of the effect mediated through each surrogate risk score was estimated, along with 95% CIs via a bootstrap procedure."
In a previous post, the mediation analysis was discussed. "Mediation analysis and SAS CAUSALMED procedure"

Landmark analysis, also known as landmark survival analysis, is a statistical method commonly used in survival analysis to investigate the impact of time-dependent variables on the occurrence of an event of interest. It allows for the assessment of time-varying effects in longitudinal studies or clinical trials where the values of variables may change over time.

In landmark analysis, the follow-up time is divided into predefined intervals or "landmarks," and the analysis is performed separately for each landmark. The key steps in landmark analysis are as follows:

  • Define landmarks: Choose specific time points of interest during the follow-up period.
  • Create landmark cohorts: At each landmark, divide the study population into subgroups based on the status of the time-dependent variable(s) of interest.
  • Analyze survival outcomes: Estimate survival probabilities or hazard rates for each subgroup defined by the landmark cohorts.
  • Compare survival outcomes: Compare the survival outcomes between the different subgroups to assess the impact of the time-dependent variable(s) on the event occurrence.

Landmark analysis allows researchers to capture time-varying effects and observe how the relationship between variables changes over the course of a study. It is particularly useful when analyzing data with time-dependent covariates, treatment interventions, or changes in exposure levels over time.

In a paper by McLaughlin et al "Pulmonary Arterial Hypertension-Related Morbidity Is Prognostic for Mortality", the landmark analysis was used to assess the impact of morbidity events on the risk of subsequent mortality.

In a paper by Eisenstein et al "Clopidogrel use and long-term clinical outcomes after drug-eluting stent implantation" , landmark analyses were performed to explore the association of extended clopidogrel use and long-term clinical outcomes of patients receiving drug eluting stents (DES) and bare-metal stents (BMS) for treatment of coronary artery disease.

The tutorial paper "Landmark Analysis at the 25-Year Landmark Point" by Dr Dafni is a good reference for Landmark analysis. 

Mediation analysis and Landmark analysis differ in their goals, focus, and types of data they analyze:

Goals: Mediation analysis aims to understand the mechanism or process through which an independent variable influences a dependent variable, exploring direct and indirect effects. Landmark analysis, on the other hand, focuses on assessing time-varying effects and understanding how the occurrence of an event is influenced by time-dependent variables.

Focus: Mediation analysis emphasizes identifying mediators that explain the relationship between an independent variable and a dependent variable. Landmark analysis focuses on investigating the impact of time-dependent variables on survival outcomes or event occurrences.


Data: Mediation analysis typically requires cross-sectional or longitudinal data with variables measured at different time points. It is applicable to both continuous and categorical variables. Landmark analysis is commonly used in survival analysis, analyzing time-to-event data, and requires longitudinal data with time-dependent variables.

Both mediation analysis and landmark analysis are valuable statistical tools in different contexts, providing insights into different aspects of data analysis. Here is a table to compare the mediation analysis and the landmark analysis generated by ChatGPT, but I added the last row for implementing the mediation and landmark analyses. 

Aspect

Mediation Analysis

Landmark Analysis

Statistical Method

Investigates relationships and effects between variables

Analyzes time-varying effects and event occurrences

Time Dependency

Considers temporal aspect of data

Accounts for time-varying effects and changing values over time

Causal Inference

Aims to understand causal chain of relationships

Examines impact of time-dependent variables on event outcomes

Focus

Identifying mediators that explain relationships

Assessing time-dependent variables and event occurrences

Variables

Independent, mediator, and dependent variables

Time-dependent variables influencing event occurrences

Data Type

Cross-sectional or longitudinal with measured variables

Longitudinal data with time-dependent variables

Analytical Steps

Estimating direct and indirect effects using regression

Dividing follow-up into landmarks, comparing survival outcomes

Research Questions

Mechanisms and processes of variable influence

Impact of time-dependent variables on events or survival

Implementation




In SAS, Procedure CASUALMED allows you to estimate direct and indirect effects using different mediation models. It supports various regression-based mediation approaches, including Sobel, bootstrapping, and Bayesian estimation.

In R, to conduct mediation analysis, the most commonly used package is "mediation." This package provides a comprehensive set of functions to estimate direct and indirect effects in mediation models. It supports various mediation methods, including the causal steps approach, bootstrapping, and structural equation modeling (SEM).

In SAS, Procedures LIFETEST, PHREG, and LIFEREG can all be used. You can divide the follow-up time into landmark intervals and analyze survival outcomes for different subgroups defined by the landmarks

In R, you can utilize the "survival" package, which is widely used for survival analysis. The "survival" package provides functions to handle time-to-event data, perform survival analysis, and estimate survival probabilities. You can divide the follow-up time into landmarks and analyze survival outcomes for different landmark cohorts.


Personally, I prefer the mediation analysis to the landmark analysis. In the landmark analysis, the subjects who did not reach the landmark timepoint were excluded from the analysis and the analyses are performed on a subset of the overall population - sort of principle stratum. The endpoint measure at the landmark timepoint and whether or not the subjects reach the landmark timepoint itself is meaningful. Excluding it from the analysis is against the intention-to-treatment principle and may cause biases.