Sunday, October 04, 2026

When Liver Signals Stall Late-Stage Pipelines: The Critical Role of Pre-Market DILI (Drug-Induced Liver Injury) Assessment

A recent report by BioSpace “BMS reports liver injury events in lung disease program, with key readout near” detailing liver injury events in Bristol Myers Squibb’s lung disease program serves as a stark reminder of an enduring reality in drug development: Drug-Induced Liver Injury (DILI) remains one of the primary drivers of clinical holds, regulatory rejections, and post-marketing withdrawals.

For clinical development teams and biostatisticians alike, evaluating hepatic safety is not merely a routine check of routine laboratory panels. It is an intricate, multi-dimensional assessment where subtle shifts in baseline aminotransferases and bilirubin can mean the difference between FDA NDA/BLA approval and a Complete Response Letter (CRL).

The Regulatory Framework: FDA Premarketing Guidance on DILI

The cornerstone of hepatic safety review is the FDA’s guidance, Drug-Induced Liver Injury: Premarketing Clinical Evaluation. The agency emphasizes that the liver possesses immense functional reserve; hepatocellular damage can occur extensively before overt organ dysfunction appears. Consequently, pre-market evaluation centers on distinguishing benign, transient enzyme elevations (adaptation) from signals of impending acute liver failure.

The Gold Standard: Hy’s Zimmerman’s Rule (Hy’s Law)

The FDA considers Hy's Law the single most reliable predictor that a drug carries severe hepatotoxic liability. Historically, finding even a single true Hy’s Law case in a clinical trial database implies a mortality risk of roughly 10% for severe DILI, signaling that broader population exposure could result in fatal liver failure.

To classify a subject as a Potential Hy’s Law Case, three core criteria must be met concurrently or in close temporal proximity:

  1. Hepatocellular Injury: Aminotransferase elevation (ALT or AST)  Upper Limit of Normal (ULN).
  2. Impaired Clearance/Excretion: Total Bilirubin (TBL)  ULN, without initial signs of cholestasis.
  3. Absence of Cholestasis: Alkaline Phosphatase (ALP)  ULN (or an -ratio , pointing to hepatocellular rather than cholestatic injury).
  4. Alternative Etiology Ruled Out: Comprehensive workups negative for viral hepatitis (A, B, C, E), autoimmune hepatitis, biliary obstruction, alcoholic liver disease, or ischemic hepatopathy.

FDA Standard Safety Tables & Visualizations for DILI

Under the FDA CDER Standard Safety Tables and Figures (ST&F) integrated review guidelines, hepatic safety requires standardized tabulations and visual analytics to systematically parse DILI risk across treatment arms.

1. Categorical Shift Tables (Threshold Counts)

Sponsors are expected to summarize the incidence of peak post-baseline laboratory elevations relative to ULN:

Parameter

Standard Evaluation Thresholds

Clinical & Regulatory Interpretation

ALT / AST

, , ,  ULN

Detects hepatocellular necrosis; helps differentiate mild adaptation () from severe necrosis ().

Total Bilirubin

,  ULN

Evaluates clearance capacity;  ULN triggers Hy's Law screening.

Alkaline Phosphatase (ALP)

, ,  ULN

Differentiates cholestatic injury patterns from hepatocellular damage.

INR

Assesses synthetic function failure (coagulopathy).


2. Standardized Figures: eDISH and mDISH

Visual exploration of DILI datasets relies primarily on the eDISH (evaluation of Drug-Induced Serious Hepatotoxicity) plot, a standard log-log scatter plot of individual patient peak post-baseline lab values:

  • Horizontal Axis (X-axis): Peak ALT or AST ( fold of ULN), split by the  ULN threshold.
  • Vertical Axis (Y-axis): Peak Total Bilirubin ( fold of ULN), split by the  ULN threshold.

     

  • Upper-Right Quadrant (Hy’s Law Zone): Patients with ALT  ULN and TBL  ULN. Points falling here immediately trigger multidisciplinary adjudication.
  • Lower-Right Quadrant (Temple’s Zone / Hepatocellular Necrosis): Patients with ALT  ULN but TBL  ULN. Indicates enzyme leakage without functional clearance compromise (often adaptive, but flags liability).
  • Upper-Left Quadrant (Cholestatic / Gilbert’s Syndrome): Elevated bilirubin without transaminase surge.
  • mDISH (Modified eDISH): Incorporates the -ratio () or color-codes points by ALP
          elevation to weed out obstructive cholestasis from true hepatocellular Hy’s Law events.

Regulatory Casualties: The Solithromycin Precedent

The pharmaceutical graveyard contains numerous programs halted by hepatic signals, but few provide a clearer case study in pre-market DILI review than Cempra Pharmaceuticals’ Solithromycin.

In 2016, Cempra submitted an NDA for solithromycin, a fourth-generation macrolide/ketolide intended to treat Community-Acquired Bacterial Pneumonia (CABP). Ketolides carried a notorious precedent: Aventis’ Ketek (telithromycin) was approved in 2004 but subsequently linked to severe liver failure and death post-marketing, leading to strict boxed warnings and restricted indications.

When the FDA evaluated Solithromycin’s NDA:

  1. Premarket Signal Detection: In Phase 3 trials involving approximately 920 patients exposed to oral solithromycin, transaminase elevations (ALT  ULN) occurred at a rate significantly higher than in the active control arm (moxifloxacin). While no definitive Hy's Law mortality occurred in the ~1,000-patient registration cohort, the frequency of marked ALT elevations (including patients exceeding  and  ULN) was troubling.
  2. Statistical Modeling and Rare Event Risk: Given the telithromycin precedent, FDA clinical reviewers and biostatisticians used the Rule of Three and binomial probability models to estimate that the true incidence of severe liver injury could be as high as 1 in 300 to 1 in 500 patients.
  3. The Advisory Committee and CRL: The Antimicrobial Drugs Advisory Committee concluded that the pre-market safety database was too small to characterize the risk of catastrophic liver failure in a common indication like CABP. In December 2016, the FDA issued a Complete Response Letter (CRL), requiring a post-baseline safety study of roughly 9,000 to 10,000 patients to rule out severe DILI prior to approval—a requirement that effectively halted the drug's commercial path.

Other notable regulatory actions driven by DILI include premarket 'not approval'/complete response letter or post-marketing market withdrawal due to the hepatotoxicity:

  • Exanta (ximelagatran): An oral direct thrombin inhibitor denied FDA approval in 2004 after approximately 8% of treated patients developed ALT  ULN, accompanied by cases of fatal acute liver injury.
  • Lumiracoxib (Prexige): Withdrawn globally and denied US approval due to high rates of serious hepatotoxicity. 
  • Troglitazone (Rezulin): Fast-tracked for Type 2 diabetes in 1997, only to be withdrawn in 2000 after accumulating at least 90 cases of liver failure and 63 deaths.

Strategic Takeaways for Late-Stage Programs

When an investigational agent—such as an IPF or chronic lung disease candidate—shows late-stage hepatic signals, sponsors must treat it as an existential program milestone:

  1. Protocol-Defined Stopping Rules: Ensure trials implement the strict FDA-recommended cessation triggers: ALT/AST  ULN; ALT/AST  ULN persisting  weeks; or ALT/AST  ULN accompanied by TBL  ULN or INR .
  2. Independent Hepatic Adjudication Committees (EAC): Blinded, prospective evaluation by independent hepatologists is mandatory to differentiate background comorbidities (e.g., right heart failure, passive hepatic congestion in respiratory patients) from intrinsic or idiosyncratic DILI.
  3. Comprehensive Re-Challenge/De-Challenge Tracking: Monitor the kinetic resolution of ALT and bilirubin upon drug withholding. True DILI should show rapid recovery kinetics unless fulminant failure is underway.
  4. Indication Context: Regulators calibrate risk-benefit tolerance based on clinical unmet need. In fatal progressive disorders like idiopathic pulmonary fibrosis (IPF), some hepatic lab shifts may be acceptable if accompanied by robust efficacy and clear mitigation strategies (e.g., intensive liver function monitoring regimens similar to pirfenidone or nintedanib). However, any confirmed Hy's Law signal without a defined monitoring pathway remains an insurmountable hurdle.

Immediate Implementation Checklist for Biostatistics & Clinical Safety

  1. Audit SDTM / ADaM Datasets: Ensure ADLBC contains derived flags for Hy's Law criteria (R2ANRHI for peak transaminases and bilirubin) and time-to-onset flags within 30 days of initial elevation.
  2. Generate Specific Summary Tables or Shift Tables for assessing DILI according to FDA's Standard Safety Tables and Figures (ST&F) integrated review guideline
  3. Automate eDISH Generation: Run dynamic eDISH displays with -ratio stratification across unblinded interim cuts to detect emergent clustering in the upper-right quadrant.
  4. Standardize Workup Case Report Forms (CRFs): Verify clinical sites automatically execute acute hepatitis panels, abdominal ultrasounds, and concomitant medication reconciliation for any subject crossing the  ALT +  TBL threshold.

Sunday, July 26, 2026

Comparative Analysis of Statistical Reporting Guidelines: The New England Journal of Medicine vs. The Lancet

Recently, while working on manuscripts submitted to The New England Journal of Medicine and The Lancet, I had the opportunity to compare their statistical reporting guidelines, especially their approaches to p-value presentation.

The communication of medical evidence relies on the rigorous, transparent, and standardized presentation of statistical results. Over the past decade, academic publishing has undergone a profound shift regarding how inferential statistics—specifically -values and confidence intervals (CIs)—are reported and interpreted. This evolution has been driven by widespread concerns over -hacking, data dredging, and the misinterpretation of statistical significance as clinical reality.

Among high-impact general medical journals, The New England Journal of Medicine (NEJM, impact factor 84.5) and The Lancet (impact factor 109.0) represent two distinct editorial philosophies regarding statistical reporting. In 2019, NEJM instituted major revisions to its statistical guidelines, restricting the use of -values to prespecified analyses with strict Type I error control and emphasizing point estimates with 95% CIs for unadjusted secondary endpoints. Conversely, The Lancet maintains a framework focused on complete quantitative disclosure, exact numerical reporting, and estimation-first presentation without imposing structural restrictions on -value inclusions for secondary outcomes.

Comparative Editorial Paradigms in Modern Medical Publishing

The editorial policies of NEJM and The Lancet reflect foundational differences in statistical governance and risk management. NEJM’s statistical framework, detailed by Harrington and colleagues and formalized in NEJM's author guidelines, functions as an inferential gatekeeper designed to minimize false-positive discoveries across published clinical trials. The journal explicitly addresses the accumulation of family-wise Type I error caused by multiple hypothesis testing, mandatory interim looks, and subgroup exploratory analyses. Consequently, NEJM’s guidelines require authors to suppress -values when error control is absent, substituting inferential statistics with point estimates and unadjusted 95% CIs. Furthermore, NEJM mandates that unless one-sided tests are required by the study design (such as in non-inferiority clinical trials), all reported -values must be two-sided.

The Lancet, while equally rigorous regarding trial registration and Statistical Analysis Plan (SAP) compliance, prioritizes continuous quantitative transparency over structural -value suppression. Rather than withholding -values for exploratory or secondary endpoints without multiplicity adjustments, The Lancet encourages the complete reporting of exact -values accompanied by effect sizes and 95% CIs, placing the burden of nuanced interpretation on the clinician and meta-analyst.

Journal

Primary Statistical Paradigm

Core Philosophy on p-Values

Primary Error Control Focus

New England Journal of Medicine

Structural Inferential Gatekeeping

Restrict p-values strictly to analyses with controlled Type I error; require two-sided testing; replace unadjusted p-values with point estimates and 95% CIs.

Stringent control of false-positive rate (Type I error) across confirmatory endpoints.

The Lancet

Transparent Estimation and Full Disclosure

Provide exact p-values alongside absolute effect sizes and 95% CIs for all pre-planned outcomes.

Comprehensive quantitative disclosure and precision estimation.

Decimal Precision, Formatting, and Typographical Conventions

Precision rules for statistical reporting serve to prevent inappropriate over-precision while preserving sufficient analytical granularity. Both journals maintain detailed conventions regarding decimal digits, significant figures, floor thresholds, and typographical styling.

NEJM enforces a magnitude-based decimal place rule for -values rather than a fixed significant-figure standard. For -values greater than 0.01, values must be reported to two decimal places, such as  or . For -values falling between 0.01 and 0.001, values are reported to three decimal places, such as  or . Any -value smaller than 0.001 is not reported as an exact decimal but expressed using the inequality floor threshold of . Exceptional allowances are made for studies involving genome-wide association studies (GWAS) or tests associated with clinical trial stopping rules, where exponentially small -values may be presented using scientific exponential notation, such as . Typographically, NEJM mandates a leading zero before the decimal point for values under 1.0 ( instead of ) and renders the symbol as an unitalicized, uppercase "".

Conversely, The Lancet relies on a significant-figure framework. Authors must report -values to two significant figures across their entire numeric range (capped at four decimal places), such as , , or . The floor threshold for reporting -values in The Lancet is lower than in NEJM, capping values smaller than 0.0001 as  (or  in print format). Regarding descriptive metrics, The Lancet guidelines mandate reporting standard deviations (SD) for mean values and interquartile ranges (IQR) for medians. Typographically, The Lancet utilizes a lowercase, italicized symbol "", includes leading zeros, and presents confidence intervals using explicit en-dashes or the word "to" when handling negative ranges.

Reporting Dimension

New England Journal of Medicine

The Lancet

p-Value Precision Mechanics

Magnitude-based tiers: 2 decimal places for P > 0.01; 3 decimal places for 0.001 <= P <= 0.01.

Two significant figures across all reported exact values, capped at 4 decimal places (e.g., p=0.12, p=0.032, p=0.0043).

Floor Threshold for p-Values

P < 0.001.

p < 0.0001 (or p < 0·0001).

Leading Zero Formatting

Mandatory leading zero included (e.g., P=0.02).

Mandatory leading zero included (e.g., p=0.021).

Symbol Presentation

Uppercase "P", non-italicized.

Lowercase "p", italicized.

Descriptive Statistics Standard

Means with SDs; Medians with IQRs.

Report SDs for mean values and IQRs for medians.

Genomic / High-Dimensional Data

Scientific notation permitted for GWAS and stopping rules (e.g., P=1 x 10^-5).

Scientific notation permitted for massive multi-testing datasets.

Multiplicity, Secondary Endpoints, and Controlled Error Paradigms

The divergence between NEJM and The Lancet is most pronounced in their handling of multiple hypothesis testing, secondary outcomes, and exploratory analyses. NEJM enforces a strict policy to eliminate unadjusted Type I error inflation. Under NEJM guidelines, -values may only be reported for primary and secondary outcomes if the study protocol and Statistical Analysis Plan (SAP) explicitly defined a formal procedure to control the overall family-wise Type I error rate, such as prespecified hierarchical gatekeeping procedure (or fixed sequence method) or Bonferroni corrections. Various statistical methods for multiplicity adjustment were described in FDA guidance "Multiple Endpoints in Clinical Trials" and EMA's "Guideline on multiplicity issues in clinical trials".

When multiplicity adjustments are conducted, NEJM requires that adjusted -values be presented and explicitly labeled as such within the manuscript text and tables. In hierarchical testing protocols, -values may only be reported sequentially until the last comparison for which the -value was statistically significant. As soon as a comparison fails to reach statistical significance, -values for that specific endpoint and all subsequent secondary comparisons must be omitted entirely. For prespecified exploratory analyses, investigators should use methods for controlling the False Discovery Rate (FDR) described in the SAP, such as Benjamini-Hochberg procedures.

When no method to adjust for multiplicity or control FDR was specified in the protocol or SAP, secondary and exploratory endpoints in NEJM reports must be limited strictly to point estimates of treatment effects with 95% CIs. No -values should be reported for these unadjusted analyses. Furthermore, authors must include a mandatory disclaimer in the Methods section stating that the widths of the intervals have not been adjusted for multiplicity and that the inferences drawn may not be reproducible.

The Lancet approaches multiple comparisons through complete transparency rather than structural suppression of inferential statistics. The journal does not mandate the withholding of -values for secondary endpoints lacking formal multiplicity control, provided trial registration and SAP compliance are fully documented. However, The Lancet strictly prohibits isolated -values. Every reported -value must be co-reported with an absolute effect size and a 95% CI. For risk changes or effect sizes, The Lancet specifically mandates reporting absolute values (such as absolute risk differences) rather than relative changes alone, ensuring that clinical relevance is evaluated alongside statistical probability.

Multiplicity Parameter

New England Journal of Medicine

The Lancet

Prespecified Multiplicity Adjustments

Report adjusted p-values explicitly labeled as multiplicity-adjusted.

Report exact p-values co-reported with absolute effect sizes and 95% CIs.

Hierarchical Sequence Truncation

Stop reporting p-values immediately at the first non-significant outcome.

No mandatory truncation of subsequent exact p-value reporting.

Prespecified Exploratory Analyses

Control False Discovery Rate (FDR) via SAP methods (e.g., Benjamini-Hochberg).

Report exact p-values alongside effect sizes and 95% CIs.

Unadjusted Secondary Endpoints

p-Values strictly prohibited; report point estimates and 95% CIs only.

Permitted; report exact p-values co-reported with point estimates and 95% CIs.

Effect Measure Requirement

Point estimates with 95% CIs.

Report absolute values (e.g., absolute risk changes) rather than relative changes alone.

Mandatory Methods Disclaimer

Explicit statement on unadjusted CIs and potential non-reproducibility required.

Detailed SAP compliance and prospective trial registration statement required.

Safety Endpoints and Adverse Event Reporting Protocols

Safety evaluations in clinical trials present a distinct methodological challenge. While efficacy testing aims to minimize Type I errors to prevent ineffective treatments from entering clinical practice, safety testing must minimize Type II errors to avoid overlooking potential adverse drug reactions or toxicities.

NEJM explicitly decouples safety evaluations from the stringent Type I error constraints applied to efficacy endpoints. Because information contained within safety endpoints may signal problems within specific organ classes, NEJM guidelines explicitly state that Type I error rates larger than 0.05 are acceptable for safety endpoints. Enforcing rigid  thresholds or applying multiplicity penalties to safety tables could inadvertently obscure critical safety signals. Furthermore, NEJM editors reserve the discretion to request -values for comparisons of adverse event frequencies among treatment groups, regardless of whether such comparisons were prespecified in the SAP. Point estimates and 95% CIs remain the primary format for presenting risk differences and incidence rate ratios in safety tables. Recently, presenting only the risk difference and its 95% CI is also acceptable (see example, Inhaled Treprostinil for Idiopathic Pulmonary Fibrosis).

The Lancet prioritizes complete quantitative accounting and absolute risk presentation for safety data in its RCT guidelines. Authors must summarize adverse events using actual numbers (), denominators (), and exact percentages (%) in both intervention and control groups. Treatment-related deaths and serious adverse events (SAEs) must be explicitly itemized and tabulated. The Lancet requires that risk changes for adverse events be reported as absolute risk differences rather than relative changes alone, enabling direct assessment of clinical harm. Hypothesis testing and -value generation within safety tables are discouraged unless a specific safety hypothesis was prespecified and powered in the study protocol.

Safety Parameter

New England Journal of Medicine

The Lancet

Type I Error Threshold for Safety

Error rates > 0.05 explicitly acceptable to prevent Type II errors.

Focus on descriptive precision and absolute risk estimation.

Editorial Discretion on AE p-Values

Editors may request p-values for AE comparisons regardless of SAP pre-specification.

p-Values in safety tables discouraged unless prespecified in protocol.

Adverse Event Table Formatting

Point estimates, 95% CIs, and optional requested p-values.

Actual numbers (n), denominators (N), and percentages (%) across all groups.

Required Safety Disclosures

Key adverse events and organ-class toxicities.

Itemized treatment-related deaths and serious adverse events (SAEs).

Risk Metrics

Absolute risk differences or hazard ratios with 95% CIs.

Absolute risk changes rather than relative changes alone.

Baseline Characteristics and Subgroup Frameworks

Both NEJM and The Lancet strictly prohibit the inclusion of -values in baseline demographic and clinical tables (Table 1) for randomized controlled trials. This universal policy is rooted in statistical theory: in a properly randomized trial, any baseline imbalance between treatment groups occurs purely by chance. Performing significance tests on baseline characteristics tests the null hypothesis of proper randomization rather than true population differences, rendering baseline -values methodologically inappropriate. Instead, baseline tables in both journals present summary statistics, including sample sizes, proportions, means with standard deviations (SD), and medians with interquartile ranges (IQR).

For subgroup analyses, both journals require tests of interaction rather than relying on separate within-subgroup -values. Reporting isolated -values within individual subgroups can create false impressions of differential treatment efficacy driven by sample size variations across strata. NEJM requires subgroup interaction tests to be prespecified in the SAP to report -values, whereas The Lancet mandates that interaction terms be co-reported alongside Forest plots to demonstrate the consistency of treatment effects across clinical cohorts.

Practical Submission Checklist for Authors

When preparing a manuscript for submission, authors should apply the following journal-specific workflow for statistical reporting:

Workflow for NEJM Submissions

  1. Verify Alpha Control: Check the SAP to confirm whether secondary outcomes were included in a formal Type I error control scheme (e.g., Bonferroni or hierarchical testing).
  2. Apply Hierarchical Truncation: In hierarchical testing sequences, stop reporting  values immediately after the first comparison that fails to achieve statistical significance.
  3. Suppress Unadjusted  Values: If no multiplicity adjustment was prespecified in the SAP, remove all  values from secondary/exploratory outcome tables and text; present only point estimates and 95% CIs.
  4. Insert Mandatory Methods Disclaimer: Add the required disclaimer to the Methods section stating that CI widths have not been adjusted for multiplicity and that inferences may not be reproducible.
  5. Format -Value Decimals: Format two-sided  values to 2 decimal places if , 3 decimal places if , and  for smaller values (use exponential notation for GWAS or trial stopping rules). Use uppercase, non-italicized .

Workflow for The Lancet Submissions

  1. Format Exact -Values: Report all exact -values to two significant figures, capped at four decimal places, using  (or ) as the lower floor threshold. Use lowercase, italicized .
  2. Co-report Absolute Effect Estimates: Ensure every reported -value is accompanied by an absolute effect size (e.g., absolute risk difference) and a 95% CI.
  3. Format Descriptive Metrics: Report means with SDs and medians with IQRs throughout the text and tables.
  4. Detail Safety Tables: Summarize adverse events with complete participant counts (), denominators (), and percentages (%), explicitly detailing treatment-related deaths and SAEs.
  5. Omit Baseline -Values: Ensure Table 1 (baseline characteristics) contains no baseline -values.
REFERENCES: