Showing posts with label adequate & well-controlled study. Show all posts
Showing posts with label adequate & well-controlled study. Show all posts

Sunday, February 22, 2026

From "Two-Trial Dogma" to the Single Pivotal Standard: The Evolution of FDA Evidence Requirements

The "two-trial" rule was born from the 1962 Kefauver-Harris Amendment to the Federal Food, Drug, and Cosmetic Act, which mandated that manufacturers prove a drug was not just safe, but also effective. This established the "substantial evidence" standard, which the FDA historically interpreted as requiring at least two adequate and well-controlled clinical trials. This "two-trial dogma" served as a statistical insurance policy: in a world where biologic understanding was more limited, requiring a developer to be "lucky twice" reduced the probability of a false-positive result from 250 in 10,000 to just 6 in 10,000.

However, as of 2026, the regulatory landscape has reached a historic turning point. Here is how the "substantial evidence" requirement evolved from a rigid duplication rule into a flexible, precision-based standard.

1. The 1998 Foundation: Establishing Statutory Flexibility

The first major shift toward modern flexibility arrived with the 1998 Guidance: Providing Clinical Evidence of Effectiveness for Human Drug and Biological Products. Following the FDA Modernization Act (FDAMA) of 1997, the agency gained formal statutory authority to grant marketing authorization based on a single adequate and well-controlled study combined with "confirmatory evidence".

While this allowed for disease-by-disease flexibility—particularly in oncology and rare diseases, where single trials began to support the majority of approvals—manufacturers remained confused about exactly when a single trial would be accepted. For most "Main Street" drugs or drugs for common diseases, the two-trial expectation remained the functional default.

2. The 2023 Expansion: Defining "Confirmatory Evidence"

In September 2023, the FDA released updated draft guidance "Demonstrating Substantial Evidence of Effectiveness With One Adequate and Well-Controlled Clinical Investigation and Confirmatory Evidenceto clarify what constitutes "confirmatory evidence" when only one pivotal trial is conducted. This guidance acknowledged that modern drug development relies on both statistical and biologic inferences. Under this framework, a single trial could be bolstered by:

  • Clinical Evidence from a Related Indication

  • Mechanistic or Pharmacodynamic Evidence

  • Evidence from Relevant Animal Model

  • Evidence from Other Members of the Same Pharmacological Class

  • Natural History Evidence

  • Real-World Data/Evidence (RWD/RWE)

  • Evidence from Expanded Access Use of an Investigational Drug

3. The 2026 Paradigm Shift: The New Default

Last week, in February 2026, the FDA officially ended the "two-trial dogma." In a landmark article "One Pivotal Trial, the New Default Option for FDA Approval - Ending the Two-Trial Dogma" in the New England Journal of Medicine, FDA officials announced that a one-trial requirement is now the agency's new default standard for drug approval. We are waiting for the formal FDA guidance to provide the details about this paradigm shift.

Why the shift?

  • Precision and Biology: Modern drug discovery is increasingly precise. The FDA now considers biochemical changes and biomarkers that tell a "complete biologic story," making overreliance on a second trial unnecessary when the "mechanistic science is sound".

  • Economic Relief: A single pivotal study can cost between $30 million and $150 million and take years to complete. By moving to a one-trial default, the FDA aims to lower capital costs and remove a primary justification for high drug prices.

  • Quality Over Quantity: Officials argue that two trials can provide "false assurance" if their designs are deficient (e.g., substandard control arms or dubious endpoints). The agency will now focus its energy on ensuring the single required trial is "robust and sound".

The Guardrails: When Two Trials Are Still Required

The FDA will not abandon the two-trial standard entirely. Additional studies may still be required if:

  • An intervention has a nebulous or nonspecific mechanism of action.

  • The trial affects only a labile or short-term surrogate outcome.

  • The primary trial has underlying limitations or deficiencies.

Conclusion

We have moved from an era of Replication (the 1962-1998 standard) to Precision (the 2026 default). By formally changing the "default option," the FDA expects to spur a surge in biomedical innovation and speed life-saving drugs to the patients.

Additional Reading:


Tuesday, November 29, 2022

Randomized withdrawal design in action - Accord trial in Alzheimer's agitation

The biotech company, Axsome Therapeutics, announced the positive results from one of their pivotal phase 3 studies (Accord study). 

Axsome's approved depression drug clears Alzheimer's agitation trial months after Lundbeck-Otsuka duo

The unique side of the Accord study is the use of a randomized withdrawal design. The study was registered on clinicaltrials.gov as "A Double-blind, Placebo-controlled, Randomized Withdrawal Trial to Assess the Efficacy and Safety of AXS-05 for the Treatment of Agitation in Subjects With Dementia of the Alzheimer's Type"

With the randomized withdrawal design, all participants were given the active drug (AXS-05) in a run-in phase in an open-label manner. Then, those patients who responded to treatment during the run-in phase were randomly assigned, in a double-blind manner, to either continue treatment with AXS-05 or switch to a placebo.

According to the sponsor, the basic idea behind this randomized withdrawal study design is to see whether those who initially experience a benefit stop doing so when moved to a placebo, indicating that the therapy itself is effective — as opposed to results being due to a placebo effect. The randomized-withdrawal design of this phase 3 trial simultaneously improved signal detection and mitigated placebo response. 

With the randomized withdrawal design, the sample size was reduced. Total 178 patients with Alzheimer's disease agitation were enrolled into the study run-in phase. 108 patients who achieved a sustained clinical response were then included in the randomized withdrawal period. 
"The ACCORD study was a double-blind, placebo-controlled, multi-center, randomized withdrawal, U.S. trial which treated 178 patients with Alzheimer’s disease agitation. Patients achieving a sustained clinical response after open-label treatment with AXS-05 were randomized (n=108) in a 1:1 ratio to continue treatment with AXS-05 or to discontinue AXS-05 and switch to placebo."
According to an article on evaluate.com:

"When Axsome decided to stop the Accord study of AXS-05 in Alzheimer’s disease agitation early, hopes for that trial took a nosedive. So it was clearly a pleasant surprise today when the company announced that the study had hit. Axsome’s stock opened up 33%, and some investors might be hoping for an earlier-than-expected filing, despite the fact that results from the pivotal Advance-2 study are not due until 2025. Accord had initially been intended as a second pivotal, alongside the previously completed Advance-1, but when the number of agitation events turned out lower than expected, management decided to switch focus to Advance-2. Despite this, Accord met its primary endpoint, time to relapse of agitation, and a key secondary, relapse prevention. One potential fly in the ointment could be Accord’s randomised withdrawal design; it comprised an open-label lead-in phase in which all 178 patients were given AXS-05, and those that had a sustained clinical response to the agent were randomised to either continue treatment or switch to placebo. AXS-05, a combination of dextromethorphan and bupropion, is approved in depression as Auvelity and moving into Alzheimer’s agitation would be an important expansion."

Given that Alzheimer's agitation is a common disease (70% of Alzheimer's disease patients may have agitation), more than one adequate and well-controlled study (or pivotal, confirmatory studies) are needed to demonstrate substantial evidence for effectiveness. Besides the Accord study (with randomized withdrawal design), two additional studies were conducted by the sponsor: ADVANCE-1 trial was a Phase 2/3 study with an active control arm. ADVANCE-2 trial is a phase 3 confirmatory study with the largest sample size (350 patients in a 1:1 randomization ratio). ADVANCE-1 study results had already been announced. ADVANCE-2 study has just started the enrollment. Both ADVANCE-1 and ADVANCE-2 studies were designed as traditional RCT design - randomized, double-blind, placebo-controlled, parallel groups. 

In a clinical program containing multiple pivotal clinical trials, it is appropriate to select different clinical trial designs. In Axsome's Alzheimer's agitation clinical program, a randomized withdrawal design was used in one of the three pivotal trials, and a traditional RCT design was used in the other two pivotal trials. If all these three trials are successful, the evidence for effectiveness will be more substantial and stronger than three studies with the same study design. 

Monday, April 11, 2022

Randomization and elements of randomization specifications

Randomization is the process of assigning trial subjects to treatment or control groups using an element of chance to determine the assignments to reduce bias. Randomization is the most critical feature of the RCT (randomized, controlled trials). In FDA's Good Review Practice: Clinical Review of Investigational New Drug Applications, the randomization is defined as the following: 
In the context of clinical trial design, randomization is defined as the allocation of patients to the investigational drug and control arms by chance. Randomization is intended to prevent any systematic difference between patients assigned to the treatments being compared and is a critical assumption for valid statistical comparisons. It is also intended to produce groups that are comparable (statistically balanced) with respect to both known and unknown factors. 
Randomization Schedule (also called a randomization scheme) is a list of randomization numbers and the corresponding treatment assignments in a data set (or in a printout in the early days). Randomization Schedule can be generated using SAS Proc Plan. A SUGI paper"Generating Randomization Schedules Using SAS Programming" I wrote 20 years ago is still applicable. 

Three steps for generating the randomization schedule for use in clinical trials: 

  • Create randomization specifications according to the study protocol requirements
  • Create and validate dummy randomization schedule for review and approval
  • Create and validate the final randomization schedule - the final randomization schedule for implementation

The dummy randomization schedule and the final randomization schedule have the same display but are generated with different random seeds (therefore different treatment assignments). The dummy randomization schedule can be reviewed by the study team and the final randomization schedule can only be distributed to the designated recipients who are unblinded to the treatment assignments. 

Here is an example randomization specification: 


Here are the elements for the randomization specifications: 

Study Design: clinical trial design dictates how the subjects are assigned to receive different study treatments. In clinical trials with parallel design, subjects are randomized to receive the treatments; in clinical trials with cross-over design, subjects are randomized to different treatment sequences. 

The study design will also include a randomization strategy: 

Fixed randomization: 
  • fixed-randomization scheme (rarely used)
  • block randomization
  • stratified randomization, 
Dynamic randomization: 
       Adaptive randomization. 

See FDA's Good Review Practice: Clinical Review of Investigational New Drug Applications for definitions of these different types of randomizations. 

Blind and blindness: concealing treatment assignments and treatment allocations. 

Block: Block randomization works by randomizing subjects within blocks such that within each block, the # of subjects is balanced between treatment groups or according to the randomization ratio. 

Block Size: The size of each block. Block sizes must be multiples of the number of treatments and take the allocation ratio into account. For 1:1 randomization of 2 groups, blocks can be sizes 2, 4, 6 etc. For 1:1:1 randomization of 3 groups or 2:1 randomization of 2 groups, blocks can be sizes 3, 6, 9 etc. 

If the randomization is by site, to prevent the potential unblinding/guessing, the block size can be set up as variable for different blocks or is not revealed to the investigators and study team. With central randomization, potential unblinding is less of a concern and the block size can be the smallest multiples (for 1:1 randomization of 2 groups, the block size can be 2). 

Number of Blocks

Total Number of Randomizations: Total number of randomization numbers to be generated. Total number of randomizations = Number of blocks x Block size. Usually, randomization numbers more than the protocol-specified sample size are generated to make sure that there is a sufficient number of randomizations in the situation that the sample size may be increased or randomization errors that results in some randomization numbers not being used. If the study protocol specifies 300 subjects to be randomized, it may be good to generate 600 randomizations. 

Strata and Stratification Factors: Stratification factors are those known factors that may have an impact on treatment responses. Stratification factors are the known confounders. The most common stratification factor is the baseline disease severity which usually has an impact on the treatment responses. When stratification factors are specified, stratified randomization is employed to prevent imbalance between treatment groups for known factors that influence prognosis or treatment responsiveness. The randomization schedule is essentially generated for each stratum. 

See previous posts "Restricted randomization, stratified randomization, and forced randomization"; "Minimization Algorithm to Achieve Treatment Balance across Strata in Stratified Randomization", and "Handling Randomization Errors in Clinical Trials with Stratified Randomization"

Randomization Ratio (or allocation ratio): The ratio for treatment groups. The typical randomization ratio is balanced: 1:1 ratio for two treatment groups (if the block size is 2, for every 2 subjects randomized, there will be one assigned to group A and one assigned to group B); 1;1:1 ratio for three treatment groups, ... The randomization ratio can also be unbalanced such as 2:1 (if the block size is 3 (minimal), for every three subjects randomized, there will be two assigned to group A and one assigned to group B) and 3:1,...  FDA's Good Review Practice: Clinical Review of Investigational New Drug Applications described the randomization ratio (allocation ratio) as the following: 

Allocation of patients to treatment and control arms can be uniform or nonuniform. Uniform allocation (i.e., equal numbers allocated to each arm) is the usual practice and provides the most statistical power for a given total sample size. Nonuniform allocation  may lower costs (if one arm is substantially more expensive) and improve recruitment (if one arm is generally preferred) and may increase the size of the exposed patient safety database. In general, the loss of statistical power in seeking to detect a difference between treatments going from uniform allocation to 2:1, or even 3:1, is fairly small; however, as more imbalanced allocation occurs, power drops off more rapidly. A special case is where a trial seeks both to show effectiveness versus placebo and to compare the test drug with an active control. In that case, it usually is necessary for the active treatment groups to be substantially larger to examine the smaller differences between the active treatments. 

Randomization Number: a series of sequential numbers corresponding to treatment assignments. 'randomization number' is not random, the associated treatment assignments are random. 

Randomization can be recorded in the database and serve as the subject identifier (same as the subject number). Seeing the randomization number will not unblind the subject's treatment assignment.  

Treatment Code: short description or abbreviation for long treatment descriptions. Treatment code can be just the letters (such as A = Active; P = Placebo). 

Treatment Description: the detailed description of the treatment groups. It can be just 'Active', 'Placebo' or more descriptive as "Inhaled drug X BID', 'Inhaled Placebo BID'.

Dummy Randomization Schedule: also called surrogate randomization schedule - the randomization schedule for review and approval purposes. The dummy randomization schedule should have exactly the same features as the final randomization schedule except that a different random seed is used (therefore, the treatment assignments are different). 

Random Seed: A number (integer) used to initiate a pseudorandom number generator. Random Seed is a number used in SAS Proc Plan to generate the randomization schedule. Random Seed needs to be specified in the program in order to reproduce the same randomization schedule. 

After dummy randomization schedule is reviewed and approved, a final randomization schedule can be generated for implementation by changing the random seed. 

Randomization Envelopes: Envelopes that contain the treatment assignment information. The outside of the envelope contains the randomization number, and the inside of the envelope contains the randomization number and the corresponding treatment assignments and treatment descriptions. Randomization envelopes were used in the randomization process in the early days. The randomization process using randomization envelopes is now replaced with the Interactive Response Technology (IRT) including the Interactive Web Response System (IWRS) or Interactive Voice Response System (IVRS). 

Central Randomization is the opposite of randomization by the site. When a subject is eligible to be randomized, the site will contact a centralized contact (usually the computer system, IRT) to obtain the next available randomization number in the corresponding stratum regardless of the individual sites.

Monday, April 04, 2022

Common Issues in Implementing Randomization and Blinding

Randomization and blinding are two techniques to help prevent (conscious or unconscious) bias in clinical trials and they are the cornerstone of the randomized, controlled clinical trials (RCTs) and in FDA's terms, the cornerstone of the adequate & well-controlled clinical trials (A&WCs). As stated in FDA's Good Review Practice: Clinical Review of Investigational New Drug Applications:
Randomization and blinding are the two principal means of reducing bias and ensuring validity of trial conclusions. Randomization helps protect against the possibility that differences between groups at baseline will lead to outcome differences that might mistakenly be attributed to drug effect. Blinding protects against the possibility that differences in the on-trial treatment or assessment of subjects will lead to spurious outcome differences that are mistakenly attributed to a drug effect. 

In the context of clinical trial design, randomization is defined as the allocation of patients to the investigational drug and control arms by chance. Randomization is intended to prevent any systematic difference between patients assigned to the treatments being compared and is a critical assumption for valid statistical comparisons. It is also intended to produce groups that are comparable (statistically balanced) with respect to both known and unknown factors. 

While every effort is made to prevent the mistakes in implementing the randomization, it is inevitable to have the randomization errors and mistakes here and there. Here are some of the randomization errors that may be seen in clinical trials. 

Ineligible subjects are randomized: Clinical trials contain a screening period for verifying the eligibility of the study participants. All inclusion and exclusion criteria are checked during the screening period. If a subject meets all inclusion and exclusion criteria, the subject is eligible to be randomized. Sometimes, subjects are thought to be eligible for randomization, but only, later on, are found to be ineligible for one or more entry criteria. When ineligible subjects are randomized into the study and receive the assigned study treatments, the subjects are considered to be in the study. According to the intention-to-treat principle, the subjects will be included in the analyses regardless of the violation of the inclusion or exclusion criteria. If the critical criteria that are violated have an impact on the efficacy evaluation, the subjects may be excluded from the per-protocol population and sensitivity analyses are conducted with the per-protocol population to assess the robustness of the results from the primary analysis. 

Choosing the wrong stratum for randomization: For clinical trials with stratified randomization, the randomization is executed within each stratum. When a new subject is eligible to be randomized, the next available randomization number in the corresponding stratum (for example, based on the subject's gender, baseline disease severity category,...) is allocated to the subject. 

It is not uncommon that the investigational sites to select an incorrect stratum for the randomization especially when the strata information requires additional derivation and calculation, for example, if a  subject with or without using one class of background medications is a stratification factor, the information about the use one class of background medication may need to be derived. 

If a randomization stratification factor is measured more than one time, which measure will be used for randomization needs to be clearly stated in the protocol. If a spirometry parameter (for example, % predicted FEV1 >= 50% versus <50%) is used as a randomization stratification factor and spirometry tests are performed at both screening and baseline visits, the protocol needs to be specific regarding whether the results from screening visit or the baseline visit will be used for randomization - typically, the measures at the baseline visit should be used for randomization. If a laboratory parameter is used for randomization and there are both local lab and central lab, the protocol needs to be specific regarding which lab results will be used for randomization - typically the central lab results at the baseline will be used for randomization unless the central lab results can not be obtained in time for the randomization. 

See previous post "Handling Randomization Errors in Clinical Trials with Stratified Randomization"

Randomize the patients too early before all eligibility criteria are met: The investigator rushed to go to the randomization system (IRT) and randomized the subject to trigger the downstream activities, then realized that one or more screening results were still pending. 

Once the subject is randomized, it can’t be undone in the randomization system (IRT system). However, the site can hold on to the randomization information obtained and wait for the last pieces of the screening results to confirm the eligibility. If the last piece of the screening results confirms that the subject is eligible to be randomized, the previously obtained randomization information will then be used. The subject can move on to initiate the assigned study treatment. If the last piece of the screening results indicates that the subject is ineligible to be randomized, we will then need to decide if the subject is allowed to be in the study. If so, it will become the situation mentioned in the previous section "Ineligible subjects are randomized". 

In either situation, a protocol deviation needs to be recorded to document this incident. 

The PI was practicing the randomization system to see how the randomization works but accidentally randomized the subject in a live system. In this situation, there was no actual and real subject to be randomized. The subject information entered into the randomization system was not real, but one of the randomization numbers was assigned and treatment assignment was used. 

While this subject may remain in the randomization system (IRT), the subject is fake and should be removed from the downstream clinical database. This is usually a rare event, therefore, has no big impact on the integrity of the original randomization. 

Randomization system (IRT system) is down at the time of randomization or the internet is down: sometimes, the randomization needs to be performed immediately after the last eligibility criterion is confirmed. It is critical to have immediate access to the randomization information in order to randomize the subject in time for initiating the randomized treatment. However, it could happen that the IRT system is down or the internet is down when the randomization number and treatment assignments are needed. 

If this is a situation, the advice is to have a backup manual randomization system (for example, calling an unblinded person or group). 

Dispense the incorrect drug kit: nowadays, the randomization system is embedded in the system for clinical trial supplies (IRT system). In addition to the treatment assignments, a separate drug kit list will be generated. When a subject is randomized and a treatment group is assigned, the drug kits that are corresponding to the assigned treatment will be allocated and dispensed to the subject. 

See a previous post "Monitoring the double-blind study: unblinded pharmacist, unblinded monitor, and drug kit"

Due to human error, it is possible to have the correct randomization information but dispense the incorrect kit numbers. When this happens, it is adverse to verify with the clinical trial supply manager (who are unblinded) if the incorrect drug kit is for the same treatment group as the drug kit that is supposed to be dispensed (Don't communicate about the actual treatment group). It is less an issue if the incorrect kit numbers are in the assigned treatment group.

If the incorrectly dispensed drug kits are not in the assigned treatment group, the subjects received the incorrect treatment. For the statistical analysis, the subject will be included in the intention-to-treatment analysis and will be included in the randomized treatment group  (so-called 'as randomized). The subject can be excluded from the per-protocol population for sensitivity analysis. 

Recording the randomization date/time (local time versus backend system time): When a subject is randomized in the IRT system, a randomization message or printout, or randomization report will indicate the subject number, randomization number, the stratification factors used for randomization, randomization date/time. The investigator can record the randomization information in the case report form. Only blinded information can be included in the randomization report. 

One issue for this report is the randomization date/time - is it based on the IRT system date/time or the local date/time?  The local date/time should be used as the randomization date/time. If the IRT system is located in the UK and the subject is randomized in the US, the local and system time can differ in 5-8 hours. Local date/time, not the system date/time should always be used as the randomization date/time.  

The same subject is randomized twice (unless it is the micro-randomized trial)

The majority of these randomization errors that occurred in the study were not included in the publications and regulatory submissions - the randomization issues appear to be less than what actually occurred. Some of the examples of the randomization errors can still be found in the literature: 

"To start, at the beginning of CENTAUR, a randomization implementation problem was identified and addressed by the unblinded statistician. Let’s walk through the details. In CENTAUR, kits were shipped one by one after successful screening visits. While preparing for the first Data Safety Monitoring Board meeting in November 2017, the unblinded statistician found that the initial 18 study kits shipped were all active. This was due to an error at the distribution center. They proceeded to instruct the distribution center to balance these 18 kits by shipping a block of 9 placebo kits to maintain randomization. After correction, the 2:1 active:placebo ratio was maintained. The unblinded statistician notified Amylyx of this issue in January 2020, two months after study unblinding in November 2019. Participants, investigators, and study staff were never unblinded due to this error. Upon notification, Amylyx initiated a thorough investigation of the root cause, in consultation with the unblinded statistician and the distribution center. Amylyx also consulted with external statisticians to determine the best approach to assess the impact. The statisticians recommended a sensitivity analysis to exclude the participants affected by the error."


Monday, February 28, 2022

Human Preclinical Studies and Phase 0 Clinical Trials

The drug development process includes various steps from discovery (discovery of a new compound or biological product) to preclinical research (measuring the safety, toxicity, and efficacy in animal models), and then to clinical trials (testing the safety and efficacy in humans - either healthy volunteers or patients). The clinical trials are phased from phase 1 & 2 (early phase trials) to phase 3 (pivotal, confirmatory, late-phase trials), and to phase 4 (post-marketing clinical trials). A diagram indicates various stages of drug development. With innovative clinical trial designs, the clinical trial phases are blurred, for example, seamless phase 1/2 trial and seamless 2/3 trial. In rare disease areas, the drug development process may not have all phases of clinical trials. The adequate and well-controlled study may be phase 3, phase 2, or phase 1 (for example, expansion cohort studies).



Preclinical development, also called preclinical study or nonclinical study, is a stage of research that begins before clinical trials (testing in humans) and during which important feasibility, iterative testing and drug safety data are collected, typically in laboratory animals. Traditionally, the phase 1 study may be the first-in-human trials with the purpose of studying the drug's:

  • pharmacokinetics (ADME: absorption, distribution, metabolism, and excretion) and pharmacodynamics (enzyme, protein,...) 
  • toxicity, safety and side effects associated with increasing doses
  • maximum tolerable dose
  • early evidence of effectiveness

There are now two new steps that may be utilized in drug development: Preclinical human studies and phase 0 clinical trials. A revised drug development diagram is as follows: 


Human pre-clinical studies are those pre-clinical studies utilizing the human subjects (either utilizing the specimens collected from human subjects or performing testing in human decedents). Human pre-clinical studies are 'pre-clinical' because they are not conducted under IND (investigational new drug) and they are conducted for collecting the data to support the IND-enabling studies. 

In a paper by Abdallah et al, "A novel prostate cancer immunotherapy using prostate-specific antigen peptides and Candida skin test reagent as an adjuvant", they described a human pre-clinical study where peptides based on the prostate-specific antigen amino acid sequences were evaluated in terms of their recognition by peripheral immune cells from prostate cancer patients using interferon-γ enzyme-linked immunospot assay. A sample size of 10 patients with prostate cancer was selected for the study. The authors concluded: 

"We described a human preclinical study of a novel prostate cancer immunotherapy consisting of PSA peptides and Candida skin test reagent as an adjuvant. As solubility and formulation have been developed, it would be feasible to further evaluate the utility of this new therapy particularly when a proportion of prostate cancer patients seem to have immune cells with the ability to recognize these PSA peptides already. Therefore, whether this immunotherapy may enhance immune responses to PSA leading to tumor regression should be examined."

Dr. Locke's team in UAB recently conducted pioneer xenotransplantation of a gene-modified pig kidney. The study was described in the paper by Porrett et al "First clinical-grade porcine kidney xenotransplant using a human decedent model". The pig kidney was transplanted to a human decedent (brain dead patient). The purposes of this human preclinical study were stated as the following:
"Xenotransplantation is arguably the most pragmatic solution to the organ shortage crisis, but safety and efficacy concerns have limited advancement into humans. In preparation for a phase I clinical trial of porcine renal xenotransplantation at the University of Alabama at Birmingham, we asked what gaps in knowledge must be filled before such a clinical trial could be ethically offered to research subjects. We thus aimed to develop a human preclinical model which would permit the in vivo evaluation of critical safety and feasibility tenets of the pig-to-NHP model without risk to a living human. Our study was designed to test five central questions: (1) Is the current suite of porcine genetic modifications sufficient to avoid hyperacute rejection in humans? (2) Would prospective flow-based crossmatching correlate with graft survival free of hyperacute rejection? (3) Would life-threatening intraoperative complications occur during a renal porcine xenotransplant? (4) Would porcine cells and/or pathogens be detected in the blood of a human recipient? (5) Could porcine renal xenotransplantation be safely performed under the conditions necessary for a clinical trial? To this end, we designed and performed this experiment under clinical-grade conditions which included the transplantation of 10-GE porcine kidneys designed specifically for human transplantation into the conventional anatomic position using processes and facilities in compliance with multiple regulatory agencies."
In human preclinical studies, while human subjects are involved, there is no IND needed. Consents by the patients (in the first example) or by the relatives (in the second example) are needed. 

Phase 0 Clinical Trial: the concept of phase 0 clinical trial came from the FDA's guidance for industry "Exploratory IND Studies". The term 'phase 0 clinical trial' was not used in the guidance, but was used for exploratory IND studies. Phase 0 clinical trial may now be called 'early phase 1' clinical trial, for example, in NIH.gov website and in clinicaltrials.gov:

According to the FDA guidance "Exploratory IND Studies", the exploratory IND studies (therefore phase 0 clinical trials or early phase 1 clinical trials) are defined as the following:
"Exploratory IND studies usually involve very limited human exposure and have no therapeutic or diagnostic intent. Such studies can serve a number of useful goals. For example, an exploratory IND study can help sponsors
  • Determine whether a mechanism of action defined in experimental systems can also be observed in humans (e.g., a binding property or inhibition of an enzyme)
  • Provide important information on pharmacokinetics (PK)
  • Select the most promising lead product from a group of candidates5 designed to interact with a particular therapeutic target in humans, based on PK or pharmacodynamic (PD) properties
  • Explore a product’s biodistribution characteristics using various imaging technologies

 Whatever the goal of the study, exploratory IND studies can help identify, early in the process, promising candidates for continued development and eliminate those lacking promise. As a result, exploratory IND studies may help reduce the number of human subjects and resources, including the amount of candidate product, needed to identify promising drugs. The studies discussed in this guidance involve dosing a limited number of subjects with a limited range of doses for a limited period of time.
Existing regulations provide more flexibility with regard to the preclinical testing requirements for exploratory IND studies than for traditional IND studies. However, sponsors submitting the kinds of studies described in this guidance have not always taken full advantage of that flexibility. Sponsors often provide more supporting information in their INDs than is required by the regulations. Because exploratory IND studies involve administering either subpharmacologic doses of a product, or doses expected to produce a pharmacologic, but not a toxic, effect, the potential risk to human subjects is less than for a traditional phase 1 study that, for example, seeks to establish a maximally tolerated dose. Because exploratory IND studies present fewer potential risks than do traditional phase 1 studies that look for dose-limiting toxicities, such limited exploratory IND investigations in humans can be initiated with less, or different, preclinical support than is required for traditional IND studies.  "

In cancer.org website,"Types and Phases of Clinical Trials", Phase 0 clinical trials were specifically mentioned: 
Phase 0 clinical trials: Exploring if and how a new drug may work

Even though phase 0 studies are done in humans, this type of study isn’t like the other phases of clinical trials. The purpose of this phase is to help speed up and streamline the drug approval process. Phase 0 studies may help researchers find out if the drugs do what they’re expected to do. This may help save time and money that would have been spent on later phase trials.

Phase 0 studies use only a few small doses of a new drug in a few people. They might test whether the drug reaches the tumor, how the drug acts in the human body, and how cancer cells in the human body respond to the drug. People in these studies might need extra tests such as biopsies, scans, and blood samples as part of the process.

Unlike other phases of clinical trials, there’s almost no chance the people in phase 0 trials will benefit. The benefit will be for other people in the future. And because drug doses are low, there’s also less risk to those in the trial.

Phase 0 studies aren’t widely used, and there are some drugs for which they wouldn’t be helpful. Phase 0 studies are very small, often with fewer than 15 people, and the drug is given only for a short time. They’re not a required part of testing a new drug.

Phase 0 clinical trials are mainly conducted in the oncology area and there are quite some phase 0 (or early phase 1) studies are listed in clinicaltrials.gov. An example of a phase 0 study was published in JCO (Kummar et al 2009 "Phase 0 Clinical Trial of the Poly (ADP-Ribose) Polymerase Inhibitor ABT-888 in Patients With Advanced Malignancies").

Saturday, January 01, 2022

Futility Analysis and Conditional Power When Two Phase 3 Studies are Simultaneously Conducted

In late-phase clinical trials, an independent Data Monitoring Committee (DMC) is usually set up. If the clinical program includes multiple late-phase studies, the same DMC will be responsible for the entire program. With DMC, the interim analyses can be performed for different purposes:
  • The interim analysis for safety
    • with pre-specified stopping rule (for example stop the trial if the significant imbalance in # of Serious Adverse Events or in # of deaths)
    • without pre-specified stopping rule (rely on DMC members to review the overall safety)
  • The interim analysis for efficacy: To see if the new treatment is overwhelmingly better than the control group  - then stop the trial for efficacy
  • The interim analysis for futility (futility analysis): To see if the new treatment is unlikely to be better than the control group or the study will be unlikely to achieve its objective given the data at the interim – then stop the trial for futility.
There seem to be more studies with built-in futility analysis without interim analysis for overwhelming efficacy, mainly because of the concerns about the alpha-spending for efficacy. The futility analysis will have an impact on the beta-spending and the statistical power, but not on the alpha-spending. For the decision-making, regulatory agencies are usually more concerned about the alpha level (incorrectly approves a drug that does not work) or the alpha level inflation. The sponsors are more concerned about the statistical power (incorrectly concludes a drug not working while the drug is actually working).

Futility analysis usually requires calculating the Conditional Power (CP) that is defined as the probability that the final study result will be statistically significant, given the data observed thus far at the time of the interim data cut and a specific assumption about the pattern of the data to be observed in the remainder of the study, such as assuming the original design effect (alternative hypothesis) or the effect estimated from the interim data.  

If there is one single pivotal trial, the stopping rule and the CP are relatively straightforward. However, it is uncommon that the sponsor may need to conduct two pivotal (phase 3) studies (two adequate and well-controlled (A&WC) trials in FDA's term) to demonstrate substantial evidence of effectiveness as outlined in FDA guidance for industry "Demonstrating Substantial Evidence of Effectiveness for Human Drug and Biological Products Guidance for Industry".

For a clinical program with two independent A&WC trials (usually with identical design), the futility analysis and CP calculation are a little bit more complicated. Two independent A&WC trials may have an identical design but be executed differently (i.e., may not be started at the same time; may be conducted in different geographic regions/countries; and may have different enrollment speeds,...). 

When futility analysis is performed for two A&WC trials, should the conditional powers be calculated for individual studies separately or should the conditional powers be calculated for both studies together (i.e. pooled data from both studies)? 

When there are two identical A&WC trials, the interim analysis for safety should be based on the pooled data sets from both studies because it will give a more definitive answer to the safety issues, the interim analysis for efficacy should be based on the individual study data because the decision about the overwhelming efficacy should be based on the individual study, not the integrated data from two studies; the interim analysis for futility is a little bit more complicated and the decision to use the data from an individual study or to use the data from the pooled data seems to be dependent on how close the observed results from two A&WC trials are at the time of the interim analysis. 

For futility analysis using stochastic curtailment procedure, While CPs can be calculated for each individual study assuming that the treatment effect in the remaining subjects in the same study will follow the treatment effect estimated from the data of this same study at the time of the interim data cut, 

There is an alternative way to calculate the CP, i.e., to calculate the CP for each individual study, but use the observed treatment effect from the pooled data at the interim from both studies to project the trend and pattern for the remaining subjects. 

According to the paper by Lan and Wittes (1988) "The B-Value: A Tool for Monitoring Data", the CP calculation involves the decomposition of overall critical value (B-value or B1 for example) into the sum of two statistically independent interval B-values: 
  • Bt, the value of B that accumulated up through time t when interim analysis is conducted; and 
  • (B1 - Bt), the incremental value of B that accumulates from time t through the end of the study. The legitimacy of the decomposition follows from the independence of distributions of the outcomes for successive study subjects
At the time t when the interim analysis is conducted, Bt is known and is estimated from the observed data up to the time t. (B1 - Bt) is a random variable that needs to be estimated. The conditional power is derived by fixing Bt and calculating the probability that Bt + (B1 - Bt) will exceed Z1-a/2.

To calculate the CPs when there are two identical A&WC studies, t, as a measure of the information fraction, will be different for different studies. At the time t, maybe 60% of subjects have been enrolled in study #1 while 50% of subjects are enrolled in study #2. In CP calculations, the Bt part will be obtained from the individual study. The (B1-Bt) part is estimated assuming the remaining data following the observed effect up to the interim time t, should the observed effect up to the interim time t be based on the data from the individual study or from the pooled data?

It turns out both approaches can be used: 
  • estimate the treatment differences for each individual study and calculate the CP assuming that the reminding data follows the trend and pattern based on the observed data from individual study
  • estimate the treatment difference from both studies and calculate the CP assuming that the remaining data follow the trend and pattern based on the observed data from the pooled data of two studies.           
For both of these approaches, the CPs will be calculated for each individual study (therefore one CP for each study). The difference between these two approaches is in the calculation of the (B1-Bt) part - based on the individual study itself or based on the pooled data from both studies. 

We can take a look at the famous and controversial case in Biogen's aducanumab program in Alzheimer's disease. Aducanumab program in Alzheimer's diseases consisted of two pivotal, phase 3 studies (EMERGE (study 301) and ENGAGE (study 302)), and both studies were designed the same and conducted simultaneously globally. Each study had two active arms (low dose and high dose of aducanumab) versus placebo - therefore two hypothesis tests (low dose vs. placebo and high dose vs. placebo). There was a total of four hypothesis tests (two for each study).  The protocol and SAP specified the interim analysis for futility. 

An interim analysis was performed after approximately 50% of the subjects had the opportunity to complete the Week 78 visit for both EMERGE and ENGAGE studies. An interim analysis for the futility of the primary endpoint was performed to allow early termination of the studies if it was evident that the efficacy of aducanumab was unlikely to be achieved. The futility criteria were based on conditional power, which was the chance that the primary efficacy endpoint analysis would be statistically significant in favor of aducanumab at the planned final analysis, given the data at the interim analysis. The CP was calculated assuming that the future unobserved effect was equal to the maximum likelihood estimate of what is observed in the interim data. 

For each study, two CPs were calculated. The pre-specified CP calculation was to use the pooled interim data from both EMERGE and ENGAGE studies for the (B1-Bt) part and assume that the treatment effect for the remaining of the study would follow the observed treatment effect at the interim analysis. At the interim analysis, the CPs were calculated to be 13% for low dose vs placebo and 0% for high dose vs. placebo in EMERGE study, and 11% for low dose vs placebo and 12% for high dose vs. placebo in ENGAGE study. Given all four CPs were lower than the threshold of 20% (a criterion for futility), the DMC recommended stopping both studies for futility.  Biogen followed the DMC recommendation and stopped both EMERGE and ENGAGE studies for futility
.

Only after two terminated studies were wrapped up, the reanalyses of the final data indicated that there were statistically significant treatment differences in one of the studies (the ENGAGE study). With the help of the FDA, Biogen was able to submit the BLA and obtain approval for aducanumab for Alzheimer's disease. Leading to the FDA approval, there was an advisory committee meeting to review the aducanumab data. In FDA's presentation, the conditional powers were retrospectively re-calculated - this time, the conditional powers were calculated for each individual study and assumed future unobserved effect would be similar to the interim data for each individual study (not the pooled interim data). FDA claimed that CPs using this approach were more appropriate and would have one of the four CPs above the threshold of 20% (CP=59% for high-dose vs placebo in ENGAGE study) - the studies would not be recommended for stopping for futility. 


Retrospectively, CPs calculated for each study independently (not using the pooled interim data to project the trend and pattern for the remaining data) seemed to be better in Biogen aducanumab program consisting of two A&WC trials. 

However, in a paper by Deng et al "Superiority of combining two independent trials in interim futility analysis", CP calculation using the observed treatment effects from the pooled interim data from two studies was considered a better approach. It concluded, "it is demonstrated that by leveraging data from the other study, the probability of making correct interim decision is increased if the treatment effects are similar between the two studies, and such benefit remains even if there is small to moderate between-study difference."

It is probably true that CP calculation using the pooled data at the interim to project the trend and pattern for the remainder data is a better approach if two studies are conducted in the same way and the results at the time of the interim analysis are similar. However, the CP calculation and the statistical analysis plan for interim analysis are usually pre-specified before seeing the unblinded data. At the time of the interim analysis, it is usually unknown whether or not the results (treatment effects) observed from two identical studies will be similar. Even though two A&WC studies are designed the same, the operation and execution of the trial can still be different: two studies may be conducted in different countries, enrollment speed may be different,... As evidenced by Biogen's EMERGE and ENGAGE trials, two identical designed studies may have different results - therefore calculating the CP entirely independently for each study may be more appropriate when two identical A&WC trials are conducted.