Monday, July 22, 2019

Upversion of medical coding dictionaries (MedDRA and WHO-DD)

In clinical trials, adverse events, medical histories, concomitant medications (drug names) are usually entered as the free text fields (verbatim terms) in the case report forms. Data in free text format can not be systematically summarized and analyzed, medical coding becomes necessary to convert the free text fields into categories.

Adverse events and medical histories are coded using MedDRA (Medical Dictionary for Regulatory Activities). Drug names (generic or brand names) are coded using WHO-DD (World Health Organization – Drug Dictionary). Early dates, there are different dictionaries available for performing the medical coding. Nowadays, the coding dictionaries are fixed on MedDRA and WHO-DD.

Both MedDRA and WHO-DD dictionaries are updated periodically - specifically, MedDRA is updated twice a year and WHO-DD is updated four times a year.

When we do a clinical trial, we will start with the latest version of the coding dictionaries. The clinical studies usually take several years to complete. By the time we reach the completion of the study, we will do the database lock (i.e., no data changes after the database lock). The coding dictionaries selected for use at the beginning of the study will be several years old and become obsolete.

One way to resolve this issue is to perform the upversion (or up-version, upversioning) of the medical dictionaries. For a clinical trial (especially a trial with extended study duration), it is good to update the medical coding dictionaries periodically. If upversion can not be performed periodically during the study, it is better to do an upversion at least at the end of the study (before the database lock).

Historically, there was no requirement/mandate for upversion, different companies had different practice in terms of the upversion. Here is the survey result about the upversoin in practice.



However, upversion will soon become a mandate for both MedDRA and WHO-DD dictionaries.

There are two Federal Registers pertinent to the upversion of the MedDRA and WHO-DD (or WHODG) dictionaries: one for MedDRA and one for WHO-DD. Notice that WHO-DD is now called WHODG (World Health Organization Drug Global) in the Federal Register.

Electronic Study Data Submission; Data Standards; Support for Version Update of the Medical Dictionary for Regulatory Activities
“Generally, the studies included in a submission are conducted over many years and may have used different MedDRA versions to code adverse events. The expectation is that sponsors or applicants will use the most current version of MedDRA at the time of study start. However, there is no requirement to recode earlier studies. The transition date for support and requirement to use the most current version of MedDRA is March 15, 2018. Although the use of the most current version is supported as of this Federal Register notice and sponsors or applicants are encouraged to begin using it, the use of the most current version will only be required in submissions for studies that start after March 15, 2019….”
Electronic Study Data Submission; Data Standards; Support for Version Update of World Health Organization Drug Global
“FDA currently supports the use of WHODG for the coding of concomitant medications in studies submitted to CBER or CDER in NDAs, ANDAs, BLAs, and certain INDs in the electronic common technical document format. Generally, the studies included in a submission are conducted over many years and may have used different WHODG versions to code concomitant medications. The expectation is that sponsors and applicants will use the most current B3-format annual version of WHODG at the time of study start. However, there is no requirement to recode earlier studies. The transition date for support of the most current B3- format annual version of WHODG is March 15, 2018. Although the use of the current B3-format annual version of WHODG is supported as of this Federal Register notice and sponsors or applicants are encouraged to begin using it, the use of the most current B3- format annual version will only be required in submissions for studies that start after March 15, 2019.”
Additional Readings:

Thursday, July 18, 2019

Retire Statistical Significance and p-value?


In the March issue of The American Statistician, there was a special issue with 43 papers about “Statistical Inference in the 21st Century: A World Beyond p < 0.05”. The discussion about using the p-values was picked up by the scientific communities and triggered a lot of discussions. Some of the articles were provocative: “Retire Statistical Significance”, “Abandon / Retire Statistical Significance”. American Statistician Association’s president, Karen Kafadar, has also discussed this issue in his ‘president’s corner’.

For a long time, statisticians have been cautioned about the misuse of the p-values.
  • Don’t become the slave of the p-values
  • Don’t base your conclusions solely on whether an association or effect was found to be “statistically significant” (i.e., the pvalue passed some arbitrary threshold such as p < 0.05).
  • Don’t believe that an association or effect exists just because it was statistically significant.
  • Don’t believe that an association or effect is absent just because it was not statistically significant.
  • Don’t believe that your p-value gives the probability that chance alone produced the observed association or effect or the probability that your test hypothesis is true.
  • Don’t conclude anything about scientific or practical importance based on statistical significance (or lack thereof).

The intention of the special issue is to trigger a healthy debate about the p-values and statistical significance, trigger the development of better methods, and provide the educations about the appropriate use and interpretation of the p-values. However, there is a danger of the unintended consequences: non-statisticians may be confused about what to do. Worse, “by breaking free from the bonds of statistical significance” as the editors suggest and several authors urge, researchers may read the call to “abandon statistical significance” as “abandon statistical methods altogether”.

The drug development relies on the clinical trials to demonstrate the substantial evidence about the efficacy and the substantial evidence comes from adequate and well-controlled investigations.
“evidence consisting of adequate and well-controlled investigations, including clinical investigations, by experts, qualified by scientific training and experience to evaluate the effectiveness of the drug involved, on the basis of which it could fairly and responsibly be concluded by such experts that the drug will have the effect it purports or is represented to have under the conditions of use prescribed, recommended, or suggested in the labeling or proposed labeling thereof”

For a common disease, two pivotal studies (with each showing a statistical significance at alpha = 0.05) have been the requirement for FDA (see FDA guidance "Providing Clinical Evidence of Effectiveness for Human Drug and Biological Products")

FDA has applied more flexibilities in evaluating the evidence for drugs/biological products for treating rare diseases especially those with unmet medical needs. Frank Sasinowski has two articles discussing this issue.
If we need to avoid the overuse and misuse of p-values, we will need to start with the changes in the statute of the laws and changes in regulatory science.

In addition, the scientific journals and editors may judge the value of a paper based on the significance of the results and favors the studies with statistical significance for publication. However, this may be changed now. On July 18. 2019 issue of New England Journal of Medicine (NEJM), an editorial paper was published "New Guidelines for Statistical Reporting in the Journal".
"The new guidelines discuss many aspects of the reporting of studies in the Journal, including a requirement to replace P values with estimates of effects or association and 95% confidence intervals when neither the protocol nor the statistical analysis plan has specified methods used to adjust for multiplicity."
With NEJM leading the way, other journals may follow. We will see more reporting of the confidence intervals and less reporting of the p-values. 

Generate Real-World Data (RWD) and Real-World Evidence (RWE) for Regulatory Purposes

Real-world data (RWD) and real-world evidence (RWE) have been hot topics in the drug development field and in the statistical field. With the new technologies, electronic health records, and big data, it is no surprise that RWD and RWE are much discussed to be a new way to revolutionize the clinical trial design and the regulations.

FDA has its dedicated webpage for 'Real-World Evidence' and issued several guidelines for using real-world evidence to support the regulatory approvals. 

What is the Definition of RWD and RWE?

Real-world data (RWD) are the data relating to patient health status and/or the delivery of health care routinely collected from a variety of sources. RWD can come from a number of sources, for example:
  • Electronic health records (EHRs)
  • Claims and billing activities
  • Product and disease registries
  • Patient-generated data including in home-use settings
  • Data gathered from other sources that can inform on health status, such as mobile devices
Real-world evidence (RWE) is the clinical evidence regarding the usage and potential benefits or risks of a medical product derived from analysis of RWD. RWE can be generated by different study designs or analyses, including but not limited to, randomized trials, including large simple trials, pragmatic trials, and observational studies (prospective and/or retrospective).

Last week, Duke Margolis Center for Health Policy organized a symposium titled "Leveraging Randomized Clinical Trials to Generate Real-World Evidence for Regulatory Purposes". All presentation slides and the video recordings are available for free.

Day 1 focused on the study design using RWD and RWE for efficacy measure (presentation slides, presentation video)


Day 2 focused on safety monitoring using RWD and RWE (presentation slides, presentation video)


This is not the first Duke Margolis Center for Health Policy organizes the event for this topic. Here are some previous events.

Second Annual Duke-Margolis Conference on Real-World Data and Evidence

Enhancing the Application of Real-World Evidence In Regulatory Decision-Making DAY 1

Sunday, July 14, 2019

Regulations for Clinical Trials, Drug Development, Drug Approvals in China


The article by Bill Wang and  Alistair Davidson "An overview of major reforms in China’s regulatory environment" summarized the background about the regulatory environment in China. 
It is widely recognized that China is currently the second largest pharmaceutical market in the world. Historically the regulatory environment in China has been considered a highly challenging one, with: (1) major issues in the areas of comparative quality between international standards and some local products and manufacturers; (2) a timeframe for review and approval of new drugs that is longer than most major countries; and (3) a lack of capacity in the regulatory bodies that has resulted in a backlog of applications. In August 2015, the China State Council issued “Opinions on Reforming the Review and Approval System for Drugs and Medical Devices.” This was partly a result of dialogue with the local and international pharma industry that, for many years, has been pressing for major regulatory reform.1 The overarching intention of this was to “promote the structural adjustment, transformation and upgrade of the pharmaceutical industry and bring marketed products up to international standards in terms of efficacy, safety and quality, so as to better meet the public needs for drugs.” The main practical aims are to: (1) eliminate the existing backlog of registration applications; (2) establish an environment for maximizing the quality of generic drugs; (3) create a framework in China that encourages research and development of new drugs in line with global development; and (4) improve the quality and increase transparency of the review and approval process.
 Additional discussions about the regulatory environment in China for clinical trials and drug development can be found here:

Chinese regulatory agencies have published a flurry of regulations about the clinical trials and drug approval processes. Unfortunately, all regulatory guidance and policies are in Chinese. 

In the US, NIH’s ClinRegs website contains the English version of the updates about the Chinese regulations in clinical trials and drug development:



Regulatory Authorities in China and the US 

China
US
国家药品监督管理局National Medical Product Administration, NMPA
It has just launched its English version website at
http://subsites.chinadaily.com.cn/nmpa/drugs.html

国家药品监督管理局药品审评中心 (Center for Drug Evaluation)

国家药品监督管理局医疗器械技术审评中心 (Center for Medical Device Evaluation)


Guidance, Guideline, and Policies Regarding Clinical Trials, Drug Development, Drug Approval (Links to the Chinese version and the translation of the titles in English)

Guidelines for Post-marketing individual case safety reporting (ICSRs) E2B (R3)
Communications for Drug Development and Technical Evaluation (Trial)
Guidance for Accepting Data from Foreign Clinical Trials
Data Protection in Clinical Trials for Drug

Priority Review & Approval Procedure


Guidelines for Drug Application / Registration Submission
Decisions on the Adjustment of Imported Drug Registration
Pediatric Extrapolation

General Considerations to Clinical Trials for Drug

Bioequivalence Evaluation for Generic Drugs
Data Management Planning and Reporting of Statistical Analysis
Biostatistics Principles for Clinical Trials

Data Management Procedure for Clinical Trials
Clinical Trials in Pediatric Population
Self-inspection of Clinical Trial Data
Electronic Data Capture for Clinical Trials
Multi-regional Clinical Trial
Clinical Trial Registry
Adverse Drug Reaction Reporting and Monitoring
Guidance for Quality Control in Clinical Trials

Saturday, June 22, 2019

Historical Control vs. External Control in Clinical Trials

Last week, I had an opportunity to attend the annual ICSA Applied Statistics Symposium in Raleigh, North Carolina. The symposium had a lot of good sessions to discuss contemporary statistical issues. Representing the DIA NEED group, we presented a session about “historical control in clinical trials”.

What is the historical control?

(v) Historical Control. The results of treatment with the test drug are compared with experience historically derived from the adequately documented natural history of the disease or condition, or from the results of active treatment, in comparable patients or population. Because historical control populations usually cannot be as well assessed with respect to pertinent variables as can concurrently control populations, historical control designs are usually reserved for special circumstances. Examples include studies of diseases with high and predictable mortality (for example, certain malignancies) and studies in which the effect of the drug is self-evident (general anesthetics, drug metabolism).

“The external control can be a group of patients treated at an earlier time (historical control),…”

In a recent FDA guidance (2019) “Rare Diseases: Common Issues in Drug Development”, the historical control and external control was used interchangeably.
1. Historical (external) controlsFor serious rare diseases with unmet medical need, interest is frequently expressed in using an external, historical, control in which all enrolled patients receive the investigational drug, and there is no randomization to a concurrent comparator group (e.g., placebo/standard of care). The inability to eliminate systematic differences between nonconcurrent treatment groups, however, is a major problem with that design. This situation generally restricts use of historical control designs to assessment of serious disease when (1) there is an unmet medical need; (2) there is a well-documented, highly predictable disease course that can be objectively measured and verified, such as high and temporally predictable mortality; and (3) there is an expected drug effect that is large, self-evident, and temporally closely associated with the intervention. However, even diseases with a highly predictable clinical course and an objectively verifiable outcome measure may have important prognostic covariates that are either unknown or unrecorded in the historical data.
What is the difference between historical control and external control?

The historical control was used to be one type of external controls that had a time (early time) component. In more recent guidelines, the historical control and external control are used interchangeably. The concept of historical control has a broader meaning now. In terms of the clinical trial design and statistical analyses, the same issues will apply no matter it is a study using historical control or external control.

Examples of Clinical Trials with Successful use of Historical or External Control

The randomized, controlled trials (RCTs) are still the golden standard, the study with historical control or external control can be used when concurrent controls are impractical or unethical. Many drugs, biological products or medical devices have successfully been approved or cleared by regulatory agencies for marketing authorization using the evidence generated from the clinical trials with historical or external control.

Here are some examples:

Brineura for Battten Disease
Brineura for Batten disease was approved by FDA based on a non-randomized, single-arm study in 22 subjects and a comparison with 42 subjects from a natural history cohort (a historical control group)

Venetoclax for Relapsed/Refractory Chronic Lymphocytic Leukemia

Venetoclas for R/R CLL was approved by FDA based on a single-arm study in 106 subjects with a comparison of the overall response rate to a 40% response rate that was considered as clinically meaningful.


Multiple IGIV products were approved based on FDA guidance. The guidance suggested to measure the rate of serious bacterial infections during regularly repeated administration of the investigational IGIV product in adult and pediatric subjects for 12 months (to avoid seasonal biases) and compare the observed infection rate to a relevant historical standard - a statistical demonstration of a serious infection rate per person-year less than 1.0.




FDA recently approved XVIVI XPS EVLP device to help increase access to more lungs for transplant. According to Summary of Safety and Effectiveness, the PMA approval was based on a single-arm study with a matched control to demonstrate the lung transplants with EVLP lungs were not inferior to the matched control group (all other lungs transplanted at that transplant center during the same time period). The one-year survival rate was compared to the matched control group and also the large database from UNOS (United Network for Organ Sharing). This is a good example of a study using ‘external control’.   








Monday, June 03, 2019

Six-Minute Walk Test (6MWT), 2-Minute Walk Test (2MWT), 12-Minute Walk Test (12MWT), and Timed Walk (T25FW, T10MW)

Six-Minute Walk Test (6MWT) is to measure the distance in a fixed duration (6 minutes). It has been used as a clinical trial endpoint to measure the functional capacity in many therapeutical areas especially in pulmonary diseases (such as COPD, Pulmonary Hypertension) and neurology diseases (such as Duchenne Muscular Dystrophy) and others (such as the treatment of Mucopolysaccharidosis type VII (MPS VII, Sly syndrome)). The distances measured through 6MWT is 'Six-Minute Walk Distance' (6MWD).

Guidelines for Performing Standardized 6MWT

There are several guidelines for performing standardized 6MWT. The guidelines by the  American Thoracic Society is the one we usually follow:


6MWT is one of the approaches to measure 'exercise capacity' and is considered as a simulated test for measuring the function. FDA has a long-standing position that the clinical trial endpoint needs to measure patients' feel, function, and survival. 6MWT is measuring patients' function.

In FDA's guidance "Chronic Obstructive Pulmonary Disease: Developing Drugs for Treatment", 6MWT along with other exercise capacity measures were described as the following:
"Exercise capacity. Reduced capacity for exercise is a typical consequence of airflow obstruction in COPD patients, particularly because of dynamic hyperinflation occurring during exercise. Assessment of exercise capacity by treadmill or cycle ergometry combined with lung volume assessment potentially can be a tool to assess efficacy of a drug. Alternate assessments of exercise capacity, such as the Six Minute Walk or Shuttle Walk, also can be used. However, all these assessments have limitations. For instance, the Six Minute Walk test reflects not only physiological capacity for exercise, but also psychological motivation. Some of these assessments are not rigorously precise and may prove difficult in standardizing and garnering consistent results over time. These factors may limit the sensitivity of these measures and, therefore, limit their utility as efficacy endpoints, since true, but small, clinical benefits may be obscured by measurement noise."
History of the Six-Minute Walk Test: 

On "ATS Statement: Guidelines for the Six-Minute Walk Test" contained the following descriptions about the history of 6MWT.
Assessment of functional capacity has traditionally been done by merely asking patients the following: “How many flights of stairs can you climb or how many blocks can you walk?” However, patients vary in their recollection and may report overestimations or underestimations of their true functional capacity. Objective measurements are usually better than self-reports. In the early 1960s, Balke developed a simple test to evaluate the functional capacity by measuring the distance walked during a defined period of time. A 12-minute field performance test was then developed to evaluate the level of physical fitness of healthy individuals. The walking test was also adapted to assess disability in patients with chronic bronchitis. In an attempt to accommodate patients with respiratory disease for whom walking 12 minutes was too exhausting, a 6-minute walk was found to perform as well as the 12-minute walk. A recent review of functional walkingtests concluded that “the 6MWT is easy to administer, better tolerated, and more reflective of activities of daily living than the other walk tests”.
History of the Six-Minute Walk Test in Pulmonary Arterial Hypertension:

6MWD has been accepted by the FDA as the primary efficacy endpoint in the drug development in pulmonary arterial hypertension (PAH). According to a presentation by Dr. Barbara LeVarge "Exercise physiology and noninvasive assessment in PAH', the use of 6MWT in PAH started with the clinical development program of Epoprostenol.


6MWT versus Timed Walk
I am curious why 6MWT is a popular measure, but not the timed walk. To assess the functional capacity, we can either fix the time, then measure the distance (such as 6MWT), or fix the distance, then measure the time (such as Timed 25 Foot Walk [T25FW] and Timed 10-Meter Walk [T10MW or 10-MWT]). In sports, for all events in track and field and swimming, we always fix the distance and then measure the time.

In terms of the measurement accuracy, timed walk (such as T25FW and T6MW or 10-MWT) seems to be more accurate than 6MWT. For the timed walk, we need to make sure the recording of the time is accurate because the distance is fixed. For 6MWT, we need to make sure the recordings of both time and distance are accurate - while time is fixed, it usually needs to be measured as well.

The timed walk is actually used in clinical trials in neurology area and is accepted by the FDA as a clinical trial endpoint, for example, the timed walk is used to measure the improvement of walking ability in multiple sclerosis patients

2MWT, 6MWT, and 12MWT 

While 6MWT is the most commonly used, the 12-minute Walk Test (12MWT) was initially used to measure the functional capacity by Balke and 2-Minute Walk Test (2MWT) has also used in some clinical trials.

Leung et al (2006) did a study to validate the 6MWT in severe COPD "Reliability, Validity, and Responsiveness of a 2-Min Walk Test To Assess Exercise Capacity of COPD Patients" and they concluded:
The 2MWT was shown to be a reliable and valid test for the assessment of exercise capacity and responsive following rehabilitation in patients with moderate-to-severe COPD. It is practical, simple, and well-tolerated by patients with severe COPD symptoms.
Grifols is currently conducting a pivotal FORCE study "Study of the Efficacy and Safety of Immune Globulin Intravenous (Human) Flebogamma 5% DIF in Patients with Post-Polio Syndrome" where 2MWD is the primary efficacy endpoint.


Monday, May 13, 2019

Pediatric Extrapolation for Pediatric Indication


In a previous post "Pediatric Study Plan (PSP) and Paediatric Investigation Plan (PIP)", we discussed the requirements for PSP and PIP and the importance of incorporating the pediatric investigation plan into the overall clinical development program.
Doing clinical trials in the pediatric population is always challenging. It is not feasible to have a pediatric investigation plan that is too big to implement. Ethically, it is also not a wise decision to expose too many children in clinical trials (especially the placebo-controlled trials).
Regulatory agencies (such as FDA and EMA) realized the challenges in the clinical development program in children and have issued guidelines that encourage the sponsors to use an approach called 'pediatric extrapolation'.

We have already seen that some sponsors use the pediatric extrapolation to obtain the pediatric indication successfully.

 


Thursday, May 02, 2019

FDA and EMA Guidance on Adjusting for Covariates in Randomized Clinical Trials

Last week, FDA issues its draft guidance for industry titled 'Adjusting for Covariates in Randomized Clinical Trials for Drugs and Biologicals with Continuous Outcomes'. The guidance is short and sweet and gives five recommendations:

  • Sponsors can use ANCOVA to adjust for differences between treatment groups in relevant baseline variables to improve the power of significance tests and the precision of estimates of treatment effect. 
  • Sponsors should not use ANCOVA to adjust for variables that might be affected by treatment. 
  • The sponsor should prospectively specify the covariates and the mathematical form of the model in the protocol or statistical analysis plan. 
  • Interaction of the treatment with covariates is important, but the presence of an interaction does not invalidate ANCOVA as a method of estimating and testing for an overall treatment effect, even if the interaction is not accounted for in the model. The prespecified primary model can include interaction terms if appropriate. 
  • Many clinical trials use a change from baseline as the primary outcome measure. Even when the outcome is measured as a change from baseline, the baseline value can still be used advantageously as a covariate. 
Not sure why the guidance is only for clinical trials 'with continuous outcomes' and ANCOVA. Adjusting for covariates is also applicable for studies with other types of outcomes: time to event outcomes analyzed using Cox regression, categorical outcomes analyzed using logistical regression,... It is also important to follow the same rules in these studies when dealing with the covariates.

The second recommendation essentially said that the post-randomization or post-treatment variables should not be used as covariates in analyses - which is consistent with what was laid out in EMA's guidelines (see below). If there is a time-dependent covariate in longitudinal studies, the best way is to come up with a pre-adjustment formula instead of using the time-dependent covariate as a covariate in the model. There is no discussion of whether or not the post-treatment covariates can be used in the imputation model when multiple imputation method is used to impute the missing data. There was a misperception that the post-treatment variables could be used in the imputation model, but not in the analysis model. 

EMA had a similar guidance "Guideline on adjustment for baseline covariates in clinical trials" - it was in effect more than three years earlier than FDA's guidance and it was much longer (11 pages versus 3 pages in FDA's guidance) with more details about its recommendations. The pre-specification of the covariates to be used and avoidance of using the post-randomization variables as covariates was also emphasized. Here are executive summaries:
  • Stratification may be used to ensure balance of treatments across covariates; it may also be used for administrative reasons (e.g. block in the case of block randomisation). The factors that are the basis of stratification should normally be included as covariates or stratification variables in the primary outcome model, except where stratification was done purely for an administrative reason.
  • Variables known a priori to be strongly, or at least moderately, associated with the primary outcome and/or variables for which there is a strong clinical rationale for such an association should also be considered as covariates in the primary analysis. The variables selected on this basis should be pre-specified in the protocol. 
  • Baseline imbalance observed post hoc should not be considered an appropriate reason for including a variable as a covariate in the primary analysis. However, conducting exploratory analyses including such variables when large baseline imbalances are observed might be helpful to assess the robustness of the primary analysis.
  • Variables measured after randomisation and so potentially affected by the treatment should not be included as covariates in the primary analysis. 
  • If a baseline value of a continuous primary outcome measure is available, then this should usually be included as a covariate. This applies whether the primary outcome variable is defined as the ‘raw outcome’ or as the ‘change from baseline’.
  • Covariates to be included in the primary analysis must be pre-specified in the protocol. 
  • Only a few covariates should be included in a primary analysis. Although larger data sets may support more covariates than smaller ones, justification for including each of the covariates should be provided. 
  • In the absence of prior knowledge, a simple functional form (usually either linearity or categorising a continuous scale) should be assumed for the relationship between a continuous covariate and the outcome variable. 
  • The validity of model assumptions must be checked when assessing the results. This is particularly important for generalised linear or non-linear models where mis-specification could lead to incorrect estimates of the treatment effect. Even under ordinary linear models, some attention should be paid to the possible influence of extreme outlying values. 
  • Whenever adjusted analyses are presented, results of the treatment effect in subgroups formed by the covariates (appropriately categorised, if relevant) should be presented to enable an assessment of the model assumptions.
  • Sensitivity analyses should be pre-planned and presented to investigate the robustness of the primary analysis. Discrepancies should be discussed and explained. In the presence of important differences that cannot be logically explained – for example, between the results of adjusted and unadjusted analyses – the interpretation of the trial could be seriously affected. 
  • The primary model should not include treatment by covariate interactions. If substantial interactions are expected a priori, the trial should be designed to allow separate estimates of the treatment effects in specific subgroups. • Exploratory analyses may be carried out to improve the understanding of covariates not included in the primary analysis, and to help the sponsor with the ongoing development of the drug. • In case of missing values in baseline covariates the principles for dealing with missing values as outlined e.g. in the Guideline on missing data in confirmatory clinical trials(EMA/CPMP/EWP/1776/99 Rev. 1) applies. 
  • A primary analysis, unambiguously pre-specified in the protocol, correctly carried out and interpreted, should support the conclusions which are drawn from the trial. Since there may be a number of alternative valid analyses, results based on pre-specified analyses will carry most credibility.

Saturday, April 20, 2019

Hodges-Lehmann estimator of location shift: Median of Differences versus Difference in Medians or Median Difference

Hodges-Lehmann estimator has been used to compare the treatment effect while the data is non-normal distributed. See my previous posts:
Many of the journal articles used Hodges-Lehmann estimator to the difference in two medians
In a study by Perkins et al "A Randomized Trial of Epinephrine in Out-of-Hospital Cardiac Arrest",
"The Hodges–Lehmann method was used to estimate median differences with 95% confidence intervals for length-of-stay outcomes"
In a study by Devinsky et al "Trial of Cannabidiol for Drug-Resistant Seizures in the Dravet Syndrome"
"Analysis of the primary end point was performed with the use of a Wilcoxon rank-sum test. An estimate of the median difference between cannabidiol and placebo, together with the 95% confidence interval, was calculated with the use of the Hodges–Lehmann approach. Sensitivity analyses of this primary end point were prespecified in the trial protocol and statistical analysis plan"
Similarly, Hodges-Lehmann estimator was used to estimating the treatment effect in licensure trials:

FDA Clinical/Statistical Review for Vascepa (icosapent ethyl) for reduction of triglycerides in patients with very high triglycerides
The median differences between the treatment groups and 95% CIs were estimated with the Hodges-Lehmann method. P-value is from the Wilcoxon rank-sum test.
FDA Statistical review for RLY5016 for Oral Suspension (Veltassa) for Hyperkalemia
To compare Veltassa with placebo, the difference between the mean ranks was tested using a two-sided t-test. The difference and 95% CI between the treatment groups in median change from baseline was estimated using a Hodges-Lehmann estimator.
FDA Medical Review of Oral Treprostinil for Pulmonary Arterial Hypertension
The magnitude of the treatment effects was defined by the Hodges-Lehmann method to estimate the median difference between treatment groups for the change from baseline in 6MWD.
It sounds like we have found a solution to estimate the difference in medians when the data is not normally distributed. However, if we look at how the Hodges-Lehmann is calculated, we will see that it is not accurate to say the Hodges-Lehmann estimator is to compare the difference in medians, it is actually the estimator of the location shift (the term originally used by the authors) or the estimator of the median of differences (further explained below).

Let's check how medians are calculated using a very simple example: 

Median and the difference in Medians:

Group A
Group B
Original Measures
4, 7, 5, 3, 6
3, 2, 5, 1, 4
Rank the original measures in order
3, 4, 5, 6, 7
1, 2, 3, 4, 5
Median
5
3
The difference in Medians (A-B)
2

Hodges-Lehmann Estimator of Location Shif (median of differences)

Group A
Group B
Original Measures
4, 7, 5, 3, 6
3, 2, 5, 1, 4
Rank the original measures in order
3, 4, 5, 6, 7
1, 2, 3, 4, 5
Each number in Group A is compared to each number in Group B
3 is compared to numbers in Group B:    2, 1, 0, -1, -2
4 is compared to numbers in Group B:    3, 2, 1, 0, -1
5 is compared to numbers in Group B:    4, 3, 2, 1, 0
6 is compared to numbers in Group B:    5, 4, 3, 2, 1
7 is compared to numbers in Group B:    6, 5, 4, 3, 2
Rank the differences from these pair comparisons in order
-2, -1, -1, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 5, 5, 6
Hodges-Lehmann estimator of location shift
Median of all these differences, in this case, the Hodges-Lehmann estimator is 2

The calculations of the medians can be implemented in the following SAS codes: 

data HodgesLehmann;
  input group $ number @@;
  datalines;
  A 3 A 4 A 5 A 6 A 7
  B 1 B 2 B 3 B 4 B 5
;
proc means data=hodgeslehmann median maxdec=0;
   class group;
   var number;
run;

proc npar1way data=hodgeslehmann hl;
   class group;
   var number;
run;

The Hodges-Lehmann estimation of the location shift is confirmed to be 2. In this example, the Hodges-Lehmann estimation of the location shift (2) is exactly the same as the differences in two medians (5-3 = 2). 

However, in many situations, the Hodges-Lehmann estimation of the location shift will be different from the differences between the two medians. the Hodges-Lehmann should really be called the median of differences between the two groups or the location shift (as the original authors used). 

The example below shows that the Hodges-Lehmann estimation of the location shift can be very different than the differences between the two medians. 



Group A
Group B
Original Measures
50.6, 39.2, 35.2, 17.0, 11.2, 14.2, 24.2, 37.4, 35.2
38.0, 18.6, 23.2, 19.0, 6.6, 16.4, 14.4, 37.6, 24.4
Rank the original measures in order
11.2
14.2
17.0
24.2
35.2
35.2
37.4
39.2
50.6
6.6
14.4
16.4
18.6
19.0
23.2
24.4
37.6
38.0
Median
35.2
19.0
The difference in Medians (A-B)
16.2


data HodgesLehmann2;                   
   input Group $ number@@;
   datalines;
A 50.6
A 39.2
A 35.2
A 17.0
A 11.2
A 14.2 
A 24.2 
A 37.4 
A 35.2 
B 38.0 
B 18.6 
B 23.2 
B 19.0 
B 6.6 
B 16.4 
B 14.4 
B 37.6 
B 24.4 

proc means data=hodgeslehmann2 median maxdec=1;
  class group;
  var number;
run;

proc npar1way data=hodgeslehmann2 hl;
  class group;
  var number;
run;

As illustrated above, the Hodges-Lehmann estimation of the location shift is 7.8, however, the difference between two medians is 35.2 - 19.0 = 16.2 (the median for groups A is 35.2 and the median for Group B is 19.0).

While the Hodges-Lehmann estimator is often used to measure the treatment difference when the data is not normally distributed, we need to understand how the Hodges-Lehmann is calculated and how Hodges-Lehmann estimator can be very different than the simple difference between two medians.