Sunday, April 05, 2009

Least squares means (marginal means) vs. means


If you work with SAS, you probably heard and used the term 'least squares means' very often. Least squares means (LS Means) are actually a sort of SAS jargon. Least square means is actually referred to as marginal means (or sometimes EMM - estimated marginal means). In an analysis of covariance model, they are the group means after having controlled for a covariate (i.e. holding it constant at some typical value of the
covariate, such as its mean value).

I often find that it is neccessary to use a very simple example to illulatrate the difference between LS Means and Means to my non-statistician colleagues. I made up the data in Table 1 above. There are two treatment groups (treatment A and treatment B) that are measured at two centers (Center 1 and Center 2).

The mean value for Treatment A is simply the summation of all measures divided by the total number of observations (Mean for treatment A = 24/5 = 4.8); similarly the Mean for treatment B = 26/5 = 5.2. Mean for treatmeng A > Mean for treatment B.

Table 2 shows the calculation of least squares means. First step is to calculate the means for each cell of treatment and center combination. The mean 9/3=3 for treatment A and center 1 combination; 7.5 for treatment A and center 2 combination; 5.5 for treatment B and center 1 combination; and 5 for treatment B and center 2 combination.

After the mean for each cell is calculated, the least squares means are simply the average of these means. For treatment A, the LS mean is (3+7.5)/2 = 5.25; for treatment B, it is (5.5+5)/2=5.25. The LS Mean for both treatment groups are identical.

It is easy to show the simple calculation of means and LS means in the above table with two factors. In clinical trials, the statistical model often needs to be adjusted for multiple factors including both categorical (treatment, center, gender) and continuous covariates (baseline measures). The calculation of LS mean is not easy to demonstrate. However, the LS mean should be used when the inferential comparison needs to be made. Typically, the means and LS means should point to the same direction (while with different values) for treatment comparison. Occasionally, they could point to the different directions (treatment A better than treatment B according to mean values; treatment B better than treatment A according to LS Mean).

SAS procedure GLM has a nice discussion about the comparison of Least Square Means vs. Means. A small article "Means vs LS Means and Type I vs Type III Sum of Squares"by Dan may also help.

Sunday, March 29, 2009

Jadad Scale to assess the quality of clinical trials

In a cost-effectiveness assessment report, a detail descriptions were provided for the approaches in choosing the clinical trial data for meta analysis. After many clinical trials are selected, a 'Jadad Scale' was used to assess the quality of clinical trials.

Jadad Scale sometimes known as Jadad scoring or the Oxford quality scoring system, is a procedure to independently assess the methodological quality of a clincal trial. It is the most widely used such assessment in the world.

The Jadad score was used as the 'gold standard' to assess the methodological quality of studies. This validated score lies in the range 0-5. Studies are scored according to the presence of three key methodological features of randomization, blinding and accountability of all patients, including withdrawals.

According to NIH website Appendix E: The Jadad Score
A Method for assessing the quality of controlled clinical trials
Basic Jadad Score is assessed based on the answer to the following 5 questions.
The maximum score is 5.

Question Yes No
1. Was the study described as random? 1 0
2. Was the randomization scheme described and appropriate? 1 0
3. Was the study described as double-blind? 1 0
4. Was the method of double blinding appropriate? (Were both the patient and the assessor appropriately blinded?) 1 0
5. Was there a description of dropouts and withdrawals? 1 0

Quality Assessment Based on Jadad Score

Range of Score Quality
0–2 Low
3–5 High

Wikipedia has a pretty good summary of the use of Jadad Scale.

Jadad Scale has been frequently used as a study selection criteria when the literature review or meta analysis are performed.

References:
1. Jadad AR, Moore RA, Carroll D, et al. Assessing the quality of reports of randomized clinical trials: Is blinding necessary? Control Clin Trials 1996;17:1-12.

Stratified randomization to achieve the balance of treatment assignment within each strata

Stratified randomization refers to the situation in which strata are constructed based on values of prognostic variables or baseline covariates and a randomization scheme is performed separately within each stratum. One misconception is to think that the stratified randomization is going to require the equal number of subjects for each strata.

For example, suppose that in a two-arm, parallel design study, we would like to stratify the randomization for age group (<18 versus >=18 years old). But we don't know how many subjects in each age group we could enroll. The purpose is to make sure that within each age group, there are equal numbers of subjects assigned to treatment A or treatment B.

After the study, there may be quite different total number of subjects in each age group, but within each age group, there should be approximately equal number of subjects in treatment A or treatment B.

The strata size usually vary (maybe there are relatively fewer young males and young females with the disease of interest). The objective of stratified randomization is to ensure balance of the treatment groups with respect to the various combinations of the prognostic variables. Simple randomization will not ensure that these groups are balanced within these strata so permuted blocks are used within each stratum are used to achieve balance.

When the stratified randomization is utilized, the # of stratification factors is typically limited to 1 or 2. The number of strata is exponentially increased if too many randomization factors are included. For example, if we have 4 stratification factors and each factor has two levels, then the # of strata = 2^4 = 16 strata, which is not practical.

If there are too many strata in relation to the target sample size, then some of the strata will be empty or sparse. This can be taken to the extreme such that each stratum consists of only one patient each, which in effect would yield a similar result as simple randomization. Keep the number of strata used to a minimum for good effect.

I have also seen a trial to require the equal number of subjects for each strata and with each strata, then equal number of subjects assigned to two treatment groups. In a trial to study the IBS (irritable bowel Syndrome), the protocol required the equal number of subjects in two type of IBSs (IBS-C vs. IBS-M). Within IBS-C or IBS-M group, there should be equal number of subjects assigned to treatment A or treatment B. The things turned out not nice because there were a lot of more subjects with IBS-C than IBS-M. During the study, while enrollment target for IBS-C was achieved, there was still a lot of IBS-M subjects to be enrolled.

IBS-C=Irritable Bowel Syndrome (constipation dominant)
IBS-M=Irritable Bowel Syndrome (mixed - constipation and diarrria)

Saturday, March 28, 2009

Too good to be true?

Typically, the regulatory authority requires two pivotal studied to demonstrate the efficacy. If the results from two studies show the conflicting or inconsistent results, the evidence for efficacy may be considered as not convincing.

On the other side, if two studies show the results almost identical, it could raise the issue with regulatory reviewers for suspicious fraud. In the most recent ASA's biopharmaceutical report, two examples were discussed.

NDA 022145 Merck's Isentress

Nearly identical results were observed in the investigational treatment group in two pivotal phase III trials for the applicant’s primary efficacy endpoint.
As part of the data verification process, the statistical review team requested copies of original source documents (laboratory reports) for HIV RNA data from the four sites that were inspected, from the site with the largest number of patients and from an additional site that had highly statistically significant results in favor of the investigational drug.
Because the applicant used an IVRS, there were no fixed randomization lists available prior to enrollment of the patients in the trial and no treatment codes available in envelopes at the sites that DSI inspected. Therefore the statistical review team also requested that copies of original source documents for treatment randomization schedules be sent directly to the FDA from the external vendors. In addition, the statistical reviewer requested the applicant’s standard operating procedures for randomization schedule generation and certification from the external vendors that the randomization code documents were obtained from the original electronic file sent to the vendors from the applicant prior to study initiation. A sample of treatment codes and laboratory data were compared to corresponding values in the SAS data sets and appeared to match.

Of note, in this NDA, two pivotal studies were allowed to be combined. The final assessment is based on the integrated summary of efficacy (ISE). It appears that the dynamic randomization was used in these two studies even though there was no detail description about the randomizaton procedure (ie, dynamic allocation for baseline covariate or dynamic allocaton for response?)

GSK's Relenza (NDA021036)


Two phase III studies assessed post-exposure prophylaxis in household contacts of an index case of influenza. In the first household study, the index case was treated while the index case was untreated in the second study. The primary efficacy endpoint for the two phase III household prophylaxis studies was the proportion of households with at least one previously uninfected household member who contracted symptomatic, laboratory-confirmed influenza.

Nearly identical rates were observed for the primary efficacy endpoint in the two household studies. Such a high degree of coincidence is rare.

Of note, the biopharmaceutical report is a quarterly report by biopharmaceutical section under american statistical association. Unfortunately, the report was not updated on their website. I have to put the report under a temporary web location.

Wednesday, March 18, 2009

Should expected clinical outcomes of the disease under study, which are efficacy endpoints, be reported as AEs/SAEs?

The paragraphs below are from the following website:

http://firstclinical.com/journal/2008/0806_GCP35.pdf

Some protocols instruct investigators to record and report all untoward events that occur
during a study as AEs/SAEs, which could include common symptoms of the disease under
study and/or other expected clinical outcomes. This approach enables frequency
comparisons of all events between treatment groups, but can make event recording in the
CRF burdensome, result in more expedited reports from investigators to sponsors, and fill
safety databases with many untoward events that most likely have no relationship to study
treatment and that could obscure signal identification.
In some clinical trials, disease symptoms and/or other expected clinical outcomes
associated with the disease under study, which might technically meet the ICH definition of
an AE or SAE, are collected and assessed as efficacy parameters rather than safety
parameters. An example might be severity scoring of prospectively defined disease
symptoms at each clinic visit during a rheumatoid arthritis study. The hypothesis underlying
this approach is that the study treatment will have a positive impact on disease symptoms.
If prospectively defined clinical outcomes, such as symptoms of a studied chronic disease or
death due to disease progression in an oncology trial, are to be assessed as efficacy
endpoints and not as AEs/SAEs, the methods for recording and analyzing these data should
be clearly described in the protocol. In addition, sponsors are advised to consult with
applicable regulatory authorities to ensure that safety reporting instructions in protocols are
acceptable, especially if certain clinical outcomes are to be excluded from traditional AE/SAE
reporting.
In high morbidity/mortality trials, independent data monitoring committees (IDMC)
generally monitor all acquired AE/SAE and clinical outcomes data to assess benefit and risk
on an ongoing basis. A reviewing IDMC could halt a trial if there was significant
improvement in pre-specified clinical outcomes in the treatment group compared to the
control group. It is also possible that a study treatment might unexpectedly worsen prespecific
disease symptoms and/or other clinical outcomes that are being assessed as
efficacy parameters.1

Reference
1. “Good Clinical Practice: A Question & Answer Reference Guide”, Barnett International,
2007, #9.10 p. 215

Source
“Good Clinical Practice: A Question & Answer Reference Guide 2007,” is available for $39.95
at http://www.barnettinternational.com/

Adverse events (AE), treatment emergent adverse events (TEAE), and adverse drug reaction (ADR)

There is some debate and inconsistencies regarding the definition of Adverse Drug Reactions. If you call it an adverse event, you may not have a culprit drug in mind, whereas calling it an adverse drug reaction is already linking it to a suspected drug. Regardless of whether or not there is a suspected drug, an AE or an ADR is commonly defined as any adverse change in health or un-desired "side-effect" that occurs in a person while on a medical treatment (for example, drug or device) or within a pre-specified period after treatment is complete. Not every adverse event is causally related to the treatment or test being studied. However, regardless of causality, people who experienced adverse reactions, or their doctors, are encouraged to report these events to the FDA or the relevant regulatory authority in the country where the drug or device is registered.

Adverse event (AE) is any untoward medical occurrence including:
  • undesirable signs & symptoms
  • disease or accidents
  • abnormal lab finding (leading to dose reduction/discontinuation/intervention)
during treatment with a pharmaceutical product in a patient or a human volunteer that does not necessarily have a relationship with the treatment given.
Adverse events is typically collected after signing the informed consent form and could be related or unrelated to the study drug.
Adverse drug reaction (ADR) is defined as:
  • For approved pharmaceutical product: a noxious and unintended response at doses normally used or tested in humans;
  • for a new unregistered pharmaceutical product: a noxious and unintended response at any dose.
WHO defines "a response to a drug which is noxious & unintended and which occurs at doses normally used for prophylaxis diagnosis or therapy of a disease or for modification of a physiological function.
The difference between AE and ADR is that AE event does not imply causality, but for ADR, a causal rule is suspected.

Another confusion is about the term 'treatment-emergent adverse event (TEAE)'. A treatment-emergent adverse event is defined as any event not present prior to the initiation of the treatments or any event already present that worsens in either intensity or frequency following exposure to the treatments. Since the starting point for AE collection is the signing of the informed consent, not the start of the study treatment, there are some adverse events occurred prior to the initiation of the study treatment. These AEs may be called "baseline-emergent adverse event" which defined as any event which occurs or worsens during the staged screening process (after informed consent) including the randomization visit. It is common to have separate summaries for AEs occurred piror to the initiation of the treatment and AEs occurred after the initiation of the treatment (ie, summary of treatment emergent adverse events).

I was asked about a programming practice to define the TEAE used in some companies. For any AE with onset date/time after the first study drug administration date/time,they compare if there is a same AE with the same severity. If yes, AE is not counted as TEAE (even though the onset date/time is after the study drug administration). For example, a subject has a mild headache 30 days after using the study medication and subjects also has a mild headache event before using the study medication,the programming will identify this event as non treatment emergent. However I think this is wrong these are two distinct events and the second one should be counted as treatment emergent AE.

The TEAE is different from the drug-related adverse events. While the treatment emergent AEs refers to adverse events temporally related to the study treatment, the drug-related AEs refers to the causality assessment by the investigator.

Friday, March 06, 2009

What is the easiest way to start a meta analysis?

I used Revman program to do my meta analysis and generated the nice Forest plot two years ago. It worked for me very well at that time. Revman is a free program developed for Cochrane Review - the most reliable reviews for evidence-based medicine.

http://community.cochrane.org/tools/review-production-tools/revman-5

Read the instruction and tutorial about how to use this program. The algorithm used in this program is described in the attached file (following the weblink below).

http://community.cochrane.org/tools/review-production-tools/revman-5/resources
http://community.cochrane.org/sites/default/files/uploads/inline-files/RevMan_5.3_User_Guide.pdf


One of the authors Julian Higgins, is one of the speakers in last year’s Meta analysis workshop sponsored by SAMSI. His topic then is titled "Practical obstables in Meta Analysis".

Since I am a heavy SAS user, I also try to do meta analysis in SAS. The following references may be useful:

Biosimilar, Follow-up Biologics, Biogenerics, and Generic Biologics

Biosimilars or Follow-on biologics are terms used to describe officially approved new versions of innovator biopharmaceutical products, following patent expiry.

Unlike the more common "small-molecule" drugs, biologics generally exhibit high molecular complexity, and may be quite sensitive to manufacturing process changes. The follow-on manufacturer does not have access to the originator's molecular clone and original cell bank, nor to the exact fermentation and purification process. Finally, nearly undetectable differences in impurities and/or breakdown products are known to have serious health implications. This has created a concern that copies of biologics might perform differently than the original branded version of the drug. However, similar concerns also apply to any production changes by the maker of the original branded version. So new versions of biologics are not authorized in the US or the European Union through the simplified procedures allowed for small molecule generics.

While the term 'biosimilar' or 'follow-on biologics' are getting popular, other terms may also be used in one way or another. Other terms include 'biogenerics', 'generic biologics',...

The Obama administration supports the use and introduction of generic drugs into the market. In his new budget proposal, Obama calls for generic biotech drugs (see CNBC news or forbes news).

on November 21, 2008, FTC held a Roundtable on Follow-on Biologic Drugs: Framework For Competition and Continued Innovation. This workshop signals continuing interest in the issue. The trascript and the videos are available from the website.


Some other readings:

Sunday, March 01, 2009

iDMC, iSTAT, iDM, and more

I guess that the iDMC (stands for independent Data Monitoring Committee) has been used for a while, however, I first time heard the terms iSTAT and iDM in the recent data monitoring committee conference. iSTAT stands for independent Statistician and iDM stands for independent Data Management.

To support the iDMC who could review the interim data during the study, an independent statistical programming team is typically needed. Within the same organization (sponsor or CRO), there could be two teams: one is the study team and is always blinded to the study treatment (prior to the study unblinding) and one is the independent team that could have access to the randomization codes and prepare the unblinded interim information for iDMC.

Currently there are many different structures in arranging the iDMC operation with statistical support. The iSTAT could be with the sponsor, with CROs (contract research organization), or ARO (academic research organization). Each modol has its own pros and cons.

IS (independent statistician). See the talk about Pat O'meara
IDC (independent data center)

Saturday, February 28, 2009

Liability and Indemnification of data monitoring committee members

As stated in Demets' paper "Liability issues for data monitoring committee members" (Clinical
Trials 2004; 1: 525–531): In randomized clinical trials, a data monitoring committee (DMC) is often appointed to review interim data to determine whether there is early convincing evidence of intervention benefit, lack of benefit or harm to study participants. Because DMCs
bear serious responsibility for participant safety, their members may be legally liable
for their actions.

With increasing DMC monitoring in clinical trials, the liability and indemnification issues are the topic of the recent data monitoring committee conference. In the situation where a study was terminated based on DMC's suggestion, the study participants could file lawsuit on either the DMC members or the sponsor for not doing the diligent work to stop the trial or stop the trial sooner enough. For example, Pfizer was sued for its Torcetrapib trial even though Pfizer is cleared of any wrongdoing. Recent events (eg Cox-IIs, Vioxx) have raised the potential for litigation and DMC members have been gotten a subpoena. For protection, DMC charters for industry trials now often cover indemnification clauses.

However, there is no indemnification yet for government-sponsored trials. For example, in NCI's guidance, it is specified "The government is prohibited by statute from indemnifying any party without specific legislative authority and consultation with the United States Department of Justice. Government liability for its own actions is usually limited by the Federal Tort Claims Act."

So what is 'indemnification'?

According to Wikipedia, "An indemnity is a sum paid by A to B by way of compensation for a particular loss suffered by B. The indemnifying party (A) may or may not be responsible for the loss suffered by the indemnified party (B). Forms of indemnity include cash payments, repairs, replacement, and reinstatement."

In the United States, Indemnification is a legal document laying down the legal protection or exemption from liability for compensation or damages from a third party, investigator and/or hospital or institution from claims made by the study subject (or relatives) that harm
was caused to the subject as a result of participation in the clinical trial.

Friday, February 27, 2009

Confidence interval for correlation coefficient

When we perform the correlation analysis, we typically calculate the correlation coefficient and then test if this correlation coefficient is statistically significant or not. We will then judge the degree of the correlation based on the numerical value of the correlation coefficient. Sometimes, we may want to calculate the confidence interval for correlation coefficient to see if the correlation coefficient has reasonable precision.

The easy way to calculate the confidence interval for correlation coefficient is to use FISHER option in SAS procedure. FISHER option is available after SAS version 9. FISHER option specifies the Fisher's z transformation to estimate 95% confidence intervals for a correlation.



When we use the confidence interval to make a judgment about the procision, we need to be aware that this is largely related to the sample size used in the calculation of the correlation coefficient. The larger the sample size, the narrower the confidence interval.

Thursday, February 19, 2009

Most testing for US drug industry's late-stage human trials done outside the country, study indicates

The Wall Street Journal (2/19, Wang) reports, "Most testing for the US drug industry's late-stage human trials is now done at sites outside the country, where results often can be obtained cheaper and faster, according to a study" published in the New England Journal of Medicine. What "make overseas trials cheaper and faster, [is that] patients in developing countries are often more willing to enroll in studies because of lack of alternative treatment options, and often they aren't taking other medicines. Such 'drug-naïve' patients can be sought after because it is easier to show that experimental treatments are better than placebos, rather than trying to show an improvement over currently available drugs."
According to the New York Times (2/19, B7, Singer), the study "raises questions about the ethics and the science of increasingly conducting studies outside the United States -- when the studies are meant to gather evidence for new drugs to gain approval in this country." The study conducted "by several Duke University researchers, suggests an ethical quagmire when drugs intended for wealthy nations are tested on people in developing countries." The researchers "suggest that human volunteers in foreign countries may be unduly influenced with the promise of financial compensation or free medical care to participate in clinical trials. The report, 'Ethical and Scientific Implications of the Globalization of Clinical Research,' also asks whether drug research conducted in developing countries is relevant to the treatment of American patients." Individuals of East Asian origin, for example, have a genetic variance that may reduce the effects of nitroglycerin treatment.
The researchers' "review of a US government clinical trials registry and of 300 published reports in major medical journals revealed this: A third (157 of 509) of Phase III trials -- typically the largest and most significant trial in the development of a drug -- led by major US pharmaceutical companies were being conducted entirely outside the United States," HealthDay (2/18, Gardner) reported. "In addition, half of the study sites (13,521 of 24,206) used in these trials were located overseas, with many in Eastern Europe and Asia."
On its website, CNN (2/19, Watkins) adds that the researchers "reported one study that found only 56 percent of 670 researchers surveyed in developing countries said their work had been reviewed by a local institutional review board or a health ministry. Another study reported that 18 percent of published trials carried out in China in 2004 adequately discussed informed consent for subjects considering participating in research."

Saturday, February 14, 2009

Evidence-based medicine - the Evidence Gap

New York Times had a series of articles to explore medical treatments used despite scant proof they work and examining steps toward medicine based on evidence.

Evidence-based medicine (EBM) aims to apply evidence gained from the scientific method to certain parts of medical practice. It seeks to assess the quality of evidence relevant to the risks and benefits of treatment (including lack of treatment). According to the Centre for Evidence-Based Medicine, "Evidence-based medicine is the conscientious, explicit and judicious use of current best evidence in making decisions about the care of individual patients."

The key for evidence-based medicine is the quality of evidence. Obviously the regulatory such as FDA applied very strict efficacy standard. According to a slide on FDA's website, FDA does not permit Sponsors To Promote Off-Label Uses because such behaviour
  • Would diminish or eliminate incentive to study the use and obtain definitive data.
  • Could result in harm to patients from unstudied uses that actually lead to bad results, or that are merely ineffective.
  • Would diminish the use of evidence-based medicine.
  • Could ultimately erode the efficacy standard.

However, there are also different voices.

Cronbach's alpha - reliability coefficient

Cronbach's Alpha is a tool for assessing the reliability of scales (for example a quality of life instrument). Cronbach's alpha can be easily calculated from SAS Proc Corr.

To compute Cronbach's alpha for a set of variables, use the ALPHA option in PROC CORR as follows:
PROC CORR DATA=dataset ALPHA;
VAR item1-item10;
RUN;

SAS website provides an example about calculating the Cronbach's alpha.
Very often, 95% confidence interval may be required, the calculation is not straightforward, but there are SAS macros available from the SAS web site.

Some references about Cronbach's alpha can be found below:



To assess the reliability of an instrument, the good reliability features include:

  • Internal consistency = Cronbach's alpha >= 0.70 for new measures
  • Stability = reliability coefficient >= 0.70
  • Equivalence = Kappa statistic >= 0.61
Reference: Nunnally & Bernstein, 1994; Landis & Koch, 1977

In one of comments on FDA's guidance on PROM (patient reported outcome measures), Cronbach's alpha was cited to measure Internal Consistency and construct validity (with scale analysis) - Cronbach's alpha > 0.70
http://www.fda.gov/ohrms/dockets/dockets/06d0044/06d-0044-EC13-Attach-1.pdf

Thursday, February 12, 2009

Blood plasma and serum

Blood plasma, or plasma, is prepared by obtaining a sample of blood and removing the blood cells. The red blood cells and white blood cells are removed by spinning with a centrifuge. Chemicals are added to prevent the blood's natural tendency to clot. If these chemicals include sodium, than a false measurement of plasma sodium content will result. Serum is prepared by obtaining a blood sample, allowing formation of the blood clot, and removing the clot using a centrifuge. Both plasma and serum are light yellow in color.

Plasma is the liquid portion of the blood that is separated from the blood cells by centrifugation. One of the characteristics of plasma is that it clots easily which is important for hemophiliacs needing a transfusion but is a nuisance in most other applications. By agitating the plasma, one can precipitate the clotting factors as a large clot, and the leftover fluid is called serum. So, serum plus clotting factors is plasma, and clotted plasma yields serum (as an interesting aside, "serum" is Latin for whey, the liquid portion of clotted milk removed in making cheese).

The following course note describes the contents of the blood, plasma, and serum.

Thursday, February 05, 2009

DMC (Data Monitoring Committee) vs. DSMB (Data Safety Monitoring Board)

DMC (Data Monitoring Committee) vs. DSMB (Data Safety Monitoring Board) are the same thing. The term DMC is used more now because it is the term used in FDA's guidance "Establishment and Operation of Clinical Trial Data Monitoring Committees" and EMEA's Guidance on Data Monitoring Committees. However, the World Health Organization used DSMB in its guidance titled "Operational Guidelines for the Establishment and Functioning of Data Safety Monitoring Boards". When searching for articles, it is recommended to try both terms "DMC" and "DSMB".

Some further discussions prior to FDA's issurance of DMC guidance are worth to read. These include:
It is also useful to know how to write the DMC charter. There are some template/example from the public domain. For example,
Some articles/books related to DMC:
  • Slutsky et al (2004) Data Safety and Monitoring Board. NEJM 350:1143-1147
  • Freidlin, B., Korn, E. L. (2009). Monitoring for Lack of Benefit: A Critical Component of a Randomized Clinical Trial. JCO 27: 629-633
  • Miller and Wendler (2008). Is it ethical to keep interim findings of randomised controlled trials confidential?. J. Med. Ethics 34: 198-201
  • Borer et al (2008) When should data and safety monitoring committees share interim results in cardiovascular trials? JAMA Apr 9;299(14):1710-2
  • Mueller et al (2007) Ethical Issues in Stopping Randomized Trials Early Because of Apparent Benefit. ANN INTERN MED 146: 878-881
  • Goodman (2007) Stopping at Nothing? Some Dilemmas of Data Monitoring in Clinical Trials. ANN INTERN MED 146: 882-887
  • Silverman (2007) Ethical Issues during the Conduct of Clinical Trials. Proc Am Thorac Soc 4: 180-184
  • Chen-Mok et al (2006) Experiences and challenges in data monitoring for clinical trials within an international tropical disease research network. Clin Trials 3: 469-477
  1. Ellenberg, Fleming, Demets (2002) Data Monitoring Committees in clinical trials: a practical perspective
  2. Demets, Friedman, Furberg (2006) Data Monitoring in clinical trials: a case studies approach
  3. Moffett (2006) Statistical monitoring of clinical trials: a unified approach
Should DMC report and meeting minutes be part of clincal study report or inlcuded in regulatory submissoin?
EMEA guidance said "In case of a submission the working procedures of a DMC as well as all DMC reports (open and closed sessions) should form part of the submission."
The internal discussion notes said "A special circumstance is the case in which the sponsor wishes to use interim data in support of a regulatory submission, with the intent to continue the trial to its conclusion. Because of the risks to the trial’s credibility, analysis and use of interim data for this purpose is often ill advised. Exceptional circumstances may arise, however, in which such use could be appropriate. Before accessing and using interim data for this purpose, sponsors should confer with FDA and the DMC (or DMC chair) and consider all potential implications of such actions. "
According to FDA guidance "The agency recommends in the guidance that the DMC or the group preparing the interim reports to the DMC maintain all meeting records. This information should be submitted to FDA with the clinical study report (Sec. 314.50(d)(5)(ii) (21 CFR 314.50(d)(5)(ii)))."
Post-analysis DMC meeting: what are the pros and cons of having the DMC convene post-analyiss so they can make an assessment on complete and clean data?
The principal role of DMC is to ensure the safety of patients, which they do by analyzing adverse events and by performing interim analyses of the clinical outcome data. Due to the time constraints, the DMC analyses are typically based on the data that is incomplete or not totally cleaned. Analyses post DMC meeting are typically not needed unless there are serious issues with the data.
One interesting question is the role of the DMC after the study has been completed. My understanding is that the DMC plays the big role during the study. After the study has been completed, DMC would hand the responsibilities back to the sponsor and investigator since all subjects have been off the study. If there is any DMC meeting after the study completion, it is mainly for the courtesy or information purpose.
If DMC made the suggestion to stop the trial after reviewing the interim analysis data, after their suggestion, it is up to the sponsor and investigators (or steering committees or executive committees) to handle the rest (close out the study, disclose the study results, write manuscript,…). In this situation, no post-DMC meeting is needed. The final analyses will be performed by the sponsor or investigators. Investigators will publish the study results. Some examples are: Novartis ACCOMPLISH trial - stopped for efficacy; Pfizer’s ILLUMINATE trial - stopped for futility.

Tuesday, February 03, 2009

Standard Error of Mean vs. Standard Error of Measurement

Everybody with basic statistical knowledge should understand the differences between the standard deviation (SD) and the standard error of mean (SE or SEM). However, people may be confused with the terms of Standard Error of Mean (SEM) vs. Standard Error of Measurement (SEM). While both shares the same acronym, the meaning and the calculation are quite different. At least, this is the situation when I saw the term 'standard error of measurement'.

I first saw this term in a literature discussing various approaches to identify the minimal clinically important difference (MCID). In an article by Copay et al, SEM (standard error of measurement) was quoted as one of the many approaches in evaluating the MCID. This method was also discussed in a paper by Wyrwich et al. Initially, I mistakenly thought that SEM was for standard error of mean. After further exploration, I realized that this SEM is quite different from that SEM.

The standard error of the mean (SEM) is the standard deviation of the sample mean estimate of a population mean. (It can also be viewed as the standard deviation of the error in the sample mean relative to the true mean, since the sample mean is an unbiased estimator.) SEM is usually estimated by the sample estimate of the population standard deviation (sample standard deviation) divided by the square root of the sample size (assuming statistical independence of the values in the sample).

The standard error of measurement (SEM) estimates how repeated measures of a person on the same instrument tend to be distributed around his or her "true" score. The true score is always an unknown because no measure can be constructed that provides a perfect reflection of the true score. SEM is directly related to the reliability of a test; that is, the larger the SEm, the lower the reliability of the test and the less precision there is in the measures taken and scores obtained. Since all measurement contains some error, it is highly unlikely that any test will yield the same scores for a given person each time they are retested.

Ar article by Dr. James Brown at University of Hawai'i at Manoa gave an good comparison of these two concepts. Also, an free paper by Harvill LM from East Tennessee State University explained in detail how the standard error of measurement is calculated.

Tuesday, January 20, 2009

Multiple Comparisons

In statistics or biostatistics, the multiple comparisons problem occurs when one considers a set, or family, of statistical inferences simultaneously. Errors in inference, including confidence intervals that fail to include their corresponding population parameters, or hypothesis tests that incorrectly reject the null hypothesis, are more likely when one considers the family as a whole.

Multiple comparison issues were nicely summarized in EMEA's guidance titled "Points to consider on multiplicity issues in clinical trials". This guidance also discussed the situations where the adjustment for multiplicity is not needed.

Adjustment for multiplicity is also mentioned in many regulatory guidance, for example, FDA guidance on ISE and its importance has been recognized in may medical journal review process.

SAMSI held a workshop in 2005 to discuss teh multiplicity issues which included the issue in Multiple Testing, Reproducibility, and Subgroup analysis.

For an introduction about multiple comparisons, refer to Wikipedia "http://en.wikipedia.org/wiki/Multiple_comparisons"

SAS Proc Multitest can be an easy tool to compute the adjusted p-values (with different methods) if the raw p-values from multiple tests are provided. For example, with the following program, we would be able to obtain a set of adjusted p-values.

data integrated;
input Method$ Raw_P;
datalines;
method1 .331
method2 .090
method3 .105
method4 .xxx
;
proc multtest pdata=integrated holm hoc fdr bon;
run;


Monday, January 19, 2009

Trial Biomarker Analysis More Than Data Dredging

members of the FDA's oncology Drugs Advisory Committee cautioned sponsors against treating retrospective clinical trial biomarker analysis like an exercise in data dredging.

The committee met last month go consider the adequacy of retrospectivly mined data in determining whether a biomarker is truly predictive of patient response. The discussion stemmed from a retrospective data analysis conducted to show that the KRAS biomarker status of patient tumors helps predict responses to Amgen's Vectibix (panitumumab) and ImClone and Bristol-Myers Squibb's Erbitux (cebuximab) cancer drugs.

See meeting transcribts here or the slides.

In other news, the US FDA encourage the integration of biomarkers in drug development and their appropriate use in clinical practice.

Data dredging vs. Data mining; Post-hoc vs. Ad-hoc

Data dredging (data fishing, data snooping) is the inappropriate (sometimes deliberately so) search for 'statistically significant' relationships in large quantities of data. This activity was formerly known in the statistical community as data mining, but that term is now in widespread use with an essentially positive meaning, so the pejorative term data dredging is now used instead.

Data mining is the process of extracting hidden patterns from data. As more data is gathered, with the amount of data doubling every three years, data mining is becoming an increasingly important tool to transform this data into information. It is commonly used in a wide range of applications, such as marketing, fraud detection and scientific discovery. Data mining can be applied to data sets of any size. However, while it can be used to uncover hidden patterns in data that has been collected, obviously it can neither uncover patterns which are not already present in the data, nor can it uncover patterns in data that has not been collected.

Post-hoc:
In or of the form of an argument in which one event is asserted to be the cause of a later event simply by virtue of having happened earlier: coming to conclusions post hoc; post hoc reasoning.
[Latin, short for post hoc, ergō propter hoc, after this, therefore because of this : post, after + hoc, neuter of hic, this.]

AD-Hoc:
adv.
For the specific purpose, case, or situation at hand and for no other: a committee formed ad hoc to address the issue of salaries.adj.
Formed for or concerned with one specific purpose: an ad hoc compensation committee.
Improvised and often impromptu: “On an ad hoc basis, Congress has . . . placed . . . ceilings on military aid to specific countries” (New York Times).
[Latin : ad, to + hoc, neuter accusative of hic, this.]

While both post-hoc and ad-hoc analysis may be performed based on the data or results we have seen, the ad-hoc analysis typically occurred alongside the project while the post-hoc analysis occurred absolutely after the project or after the unblinding of the study or after the pre-specified analyses results have been reviewed. In this sense, the ad-hoc analysis is better than post-hoc analysis.

Sunday, January 11, 2009

EQ-5D

EQ-5D is a standardized instrument for use as a measure of health outcome. Applicable to a wide range of health conditions and treatments, it provides a simple descriptive profile and a single index value for health status. EQ-5D was originally designed to complement other instruments but is now increasingly used as a 'stand-alone' measure.

An EQ-5D health state (or profile) is a set of observations about a person defined by a descriptive system. An EQ-5D health state may be converted to a single summary index by applying a formula that essentially attaches weights to each of the levels in each dimension. This formula is based on the valuation of EQ-5D health states from general population samples.

EQ-5D was established and subsequently developed by the EuroQol Group, established in 1987. The aim of the group is to test the feasibility of jointly developing a standardized non-disease-specific instrument for describing and valuing health-related quality of life.

As a matter of fact, EQ-5D is becoming popular and one day may replace the SF-36 as the most popular generalized health-related quality-of-life instrument. The main advantage of EQ-5D may be:
  • Preference-based and suitable for cost-utility analysis
  • EQ-5D value sets can be easily converted to the QALY which is the denominator in cost-utility analysis.
  • Less questions and easy to implement within short time

In one of my studies, SF-36 was performed as a quality-of-life measure. However, in order to perform the cost-utility analysis, these SF-36 scores have to be converted into something similar to EQ-5D - SF-6D . There is also other discussions about the mapping of SF-36 to EQ-5D. QualityMetric, the company for developing SF-36, is also providing the mapping for SF-6D.

However, the analysis of EQ-5D is not as easy as the questions presented in the instrument. According to a book titled "EQ-5D value sets: inventory, comparative review and user guide" (see UNC catalog), two terms seem to be important, but I may need to do a complete study using EQ-5D to figure out how to use these value sets.

  • Time Trade-off (TTO) value sets
  • Visual Analog Scale value sets

9 ways to stay alive when the worst happens

The followings are copied from PARADE (Jan 11, 2009). I am not sure if these arguments (or suggestions) have any scientific merit, but I copy here just for fun.

  1. Escape a plane crash
    The safest seats on a plane are within five rows of any exit. The No. 1 safest seats are in an exit row or one row away.
  2. Get out of a hotel five
    Most fire departments use ladders that, at their maximum, can extend around 80 feets into the air. That means in order to be able to climb out of your building's window and onto a truck's ladder, you should be on or below the seventh floor.
  3. Leave the hospital alive
    If you need to go to the hospital, weekdays are much safer than weekends. Possible explanatoins are, during the weekends, there are lower staffing levels and the presence of workers who are less experienced and less familiar with procedures and patients.
  4. Don't bo back to hospital
    beware of checking out of the hospital on a Friday. Friday is the most common hospital discharge day, but the individuals released on Friday also have an increased readmissions rate to hospital.
  5. Get an initial boost
    In one intriguing study, California researchers analyzed death records to find out whether there was any correlation between people's initials and how long they lived. They divided their subjects' initials into positive and negative groups. The good-initial group included ACE, WIN, WOW, and VIP; the bad contained RAT, BUM, SAD, and DUD. They matched up initials with lifespans and looed for any correlation. The results were stunning (and also hotly debated): a person's initial actually may influence the time and cause of his or her death. "A symbol as simple as one's initials can add four years to life or subtract three years"
    In related news, last names that begin with letters occurring later in the alphabet can be associated with a phenomenon that Scotish researchers call "alphabetical prejudice." They found that when medical teams in a brain-injury rehabilitation center met to discuss patients, people with surnames that came early in the alphabet tended to receive three to four minutes' more attention than people with names later in the alphabet.
  6. Outlive a heart attack
    One of the best places to be is in a casino in Las Vegas. The heart-attack survival rate in Las Vegas is 53%. Compare that to rates of 16% in Seattle (which has some of the nation's best response systems) or 2% in Chicago.
  7. Walk away rom a car accideng
    The rear middle seat was 16% safer than any other place in the vehicle. Overall, riding in the back is 59% to 86% safer than riding in the front, and riding on the hump is 25% safer than riding in the rear window seats.
    Compared with white cars in daylight ours, black cars had a 12% higher crash risk; gray, 11%; silver, 10%; blue and red, 7%. At dawn or dusk, black cars had a 47% higher crash risk than white cars; gray, 25%; silver 15%.
  8. Cross the street safely
    The three deadliest days for pedestrians are Jan 1, Dec 23, and Oct 31.
  9. Beware of your birthday
    Women are more likely to die in the week after their birthdays than any other week of the year, while mean's deaths peak before their birthdays.

Tuesday, January 06, 2009

FDAAA and clinical trial data bases

The recently-enacted FDA Amendments Act (“FDAAA”) has lots of requirements that may have impact on statistical analysis and programming. It is not new that the study information needst o be registered in clinicaltrials.gov database. However, a major change to the clinical trial database is that not later than December 25, 2007, the database must include links to information on clinical trial results. The term “results” is used fairly broadly in FDAAA to include summary information FDA has posted from an advisory committee meeting that considered a particular study, FDA public health advisories, FDA’s application review documents, Medline citations to any publications focused on the results of the trial, and the drug entry in the National Library
of Medicine database of structured product labels (if available).



The results requirements include demographic and baseline characteristics of the study participants, results values for each of the primary and secondary outcomes for each arm of the study, point of contact for scientific queries, and information on sponsor agreements with investigators that could restrict their ability to discuss or publish trial results.



What makes the results requirement most complicated is the format in which the results must be submitted: rather than uploading study results that have already been compiled into a clinical study report for example, using Clinicaltrials.gov's online Protocol Registration System (PRS), sponsors must first create results tables and then enter the data and statistical analyses.


This requirement means that the statistician needs to step in when the study information needs to be entered in the correct way.

Some of the statements in the amendment act are worth attention. The act stated that only applicable drug clinical trials are required to have results published. "IN GENERAL.—The term ‘applicable drug clinical trial’ means a controlled clinical investigation, other than a phase I clinical investigation, of a drug subject..." This seems to imply that the phase I study can be exempted from this requirement. However, the assignment of the study phases sometimes is arbitrary especially when a phase I study is conducted in the patients rather than the healthy volunteers.

The requirement of presenting the results for all primary and secondary could provide the misleading information to the reader if the readers have no knowledge about the interpretation of the results. Not everybody can read the results of scientifically appropriate tests of the statistical significance. Statement says ‘‘(ii) PRIMARY AND SECONDARY OUTCOMES.—The
primary and secondary outcome measures as submitted under paragraph (2)(A)(ii)(I)(ll), and a table of values for each of the primary and secondary outcome measures for each arm of the clinical trial, including the results of scientifically appropriate tests of the statistical
significance of such outcome measures."

Regarding the AE and SAE reporting, the statement says ‘‘(I) SERIOUS ADVERSE EVENTS.—A table of anticipated and unanticipated serious adverse events grouped by organ system, with number and frequency of such event in each arm of the clinical trial.
‘‘(II) FREQUENT ADVERSE EVENTS.—A table of anticipated and unanticipated adverse events that are not included in the table described in subclause (I) that exceed a frequency of 5 percent within any arm of the clinical trial, grouped by organ system, with number and frequency of such event in each arm of the clinical trial." The confusion from this is that there is no clear definition for anticipated and unanticipated SAE and AE. Perhaps for the future study protocols, the anticipated SAE and AE need to be listed in the protocol. Subsequently, the summary table of SAEs and AEs need to be separated for anticipated events and unanticipated events.

Further readings on this topic can be found from:
http://www.fda.gov/oc/initiatives/advance/fdaaa.html
http://www.fda.gov/oc/initiatives/hr3580.pdf
http://prsinfo.clinicaltrials.gov/fdaaa.html
http://www.fdalawblog.net/fda_law_blog_hyman_phelps/2008/02/fdaaa-enforceme.html

Wednesday, December 31, 2008

Significance of the Correlation Coefficient

People can be confused about the interpretation of the correlation coefficient, especially when we observe a small, but statistically significant correlation coefficient. The following paragraphs are from "http://janda.org/c10/Lectures/topic06/L24-significanceR.htm", which explain nicely about the interpretation of the correlation coefficient. In addition, the Wikipedia provides a good introduction about correlation and it also contains a small table to categorize the size (or strength) of the correlation.

Test for the significance of relationships between two CONTINUOUS variables

  • We introduced Pearson correlation as a measure of the STRENGTH of a relationship between two variables
  • But any relationship should be assessed for its SIGNIFICANCE as well as its strength.

A general discussion of significance tests for relationships between two continuous variables.

  • Factors in relationships between two variables

The strength of the relationship: is indicated by the correlation coefficient: r
but is actually measured by the coefficient of determination: r^2

  • The significance of the relationship
    is expressed in probability levels: p (e.g., significant at p =.05)
    This tells how unlikely a given correlation coefficient, r, will occur given no relationship in the population
    NOTE! NOTE! NOTE! The smaller the p-level, the more significant the relationship
    BUT! BUT! BUT! The larger the correlation, the stronger the relationship

  • Consider the classical model for testing significance
    It assumes that you have a sample of cases from a population.
    The question is whether your observed statistic for the sample is likely to be observed given some assumption of the corresponding population parameter.
    If your observed statistic does not exactly match the population parameter, perhaps the difference is due to sampling error.
    The fundamental question: is the difference between what you observe and what you expect given the assumption of the population large enough to be significant -- to reject the assumption?
    The greater the difference -- the more the sample statistic deviates from the population parameter -- the more significant it is.
    That is, the lessl ikely (small probability values) that the population assumption is true.

  • The classical model makes some assumptions about the population parameter:
    Population parameters are expressed as Greek letters, while corresponding sample statistics are expressed in lower-case Roman letters:
    r = correlation between two variables in the sample
    (rho) = correlation between the same two variables in the population
    A common assumption is that there is NO relationship between X and Y in the population: r = 0.0
    Under this common null hypothesis in correlational analysis: r = 0.0
    Testing for the significance of the correlation coefficient, r
    When the test is against the null hypothesis: r_xy = 0.0
    What is the likelihood of drawing a sample with r_xy ­ 0.0?
    The sampling distribution of r is
    approximately normal (but bounded at -1.0 and +1.0) when N is large
    and distributes t when N is small.
    The simplest formula for computing the appropriate t value to test significance of a correlation coefficient employs the t distribution:

t=r*sqrt((n-2)/(1-r^2))

The degrees of freedom for entering the t-distribution is N - 2

  • Example: Suppose you obsserve that r= .50 between literacy rate and political stability in 10 nations
    Is this relationship "strong"?
    Coefficient of determination = r-squared = .25
    Means that 25% of variance in political stability is "explained" by literacy rate
    Is the relationship "significant"?
    That remains to be determined using the formula above
    r = .50 and N=10
    set level of significance (assume .05)
    determine one-or two-tailed test (aim for one-tailed)

t=r*sqrt((n-2)/(1-r^2))=0.5*sqrt((10-2)/(1-.25)) = 1.63
For 8 df and one-tailed test, critical value of t = 1.86
We observe only t = 1.63
It lies below the critical t of 1.86
So the null hypothesis of no relationship in the population (r = 0) cannot be rejected

  • Comments
    Note that a relationship can be strong and yet not significant
    Conversely, a relationship can be weak but significant
    The key factor is the size of the sample.
    For small samples, it is easy to produce a strong correlation by chance and one must pay attention to signficance to keep from jumping to conclusions: i.e.,
    rejecting a true null hypothesis,
    which meansmaking a Type I error.
    For large samples, it is easy to achieve significance, and one must pay attention to the strength of the correlation to determine if the relationship explains very much.


  • Alternative ways of testing significance of r against the null hypothesis
    Look up the values in a table
    Read them off the SPSS output:
    check to see whether SPSS is making a one-tailed test
    or a two-tailed test
  • Testing the significance of r when r is NOT assumed to be 0
    This is a more complex procedure, which is discussed briefly in the Kirk reading
    The test requires first transforming the sample r to a new value, Z'.
    This test is seldom used.
    You will not be responsible for it.

LogMAR in Ophthalmology trials

Vision is typically reported as xxx/yyy where the xxx value is usually 20 for US assessments. As vision gets worse, for the same numerator, the denominator increases.

logMAR is log10(denominator/numerator) or -log10(numerator/denominator)
"normal" vision is 20/20, or logMAR = 0
20/100 is worse than 20/20 and logMAR = 0.69897
So the logMAR increases as vision gets worse and decreases as vision gets better
if you are doing change = visit - baseline, a negative change would be improvement in vision
a positive change would be worsening in vision.

Another interpretation of change in logMar values is to take the antilog of the change in logMAR values - this would be the "number of lines" in which vision changed. FDA often applies a 3 lines of change (ETDRS chart) criteria as this change is a doubling of the visual angle.

A useful reference on calculating average visual acuity and the whole logMAR concept is the article by Jack Holladay "Proper Method for Calculating Average Visual Acuity".

It is also useful to refer to FDA Guidelines for Multifocal Intraocular Lens IDE Studies and PMAs and Guidance for Industry Guidance for Premarket Submissions of Orthokeratology Rigid Gas Permeable Contact Lenses.

AstraZeneca considers pursuing "biosimilars."

From "http://www.delawareonline.com/article/20081230/BUSINESS/812300331"

AstraZeneca is considering joining several of its peers in pursuing "biosimilars" -- generic versions of high-priced biotechnology drugs.

The London-based drug maker has made a push into the $94 billion market for biologics in recent years with the acquisition of Cambridge Antibody Technologies in 2006 and last year's $15.6 billion purchase of Maryland-based MedImmune.
Generic versions of biologics -- drugs made from living cells rather than chemicals -- are not yet approved for sale in the United States.
The complexity of dealing with the larger biological molecules makes it impossible to create an exact copy of a biologic drug, prompting concerns that the biosimilar medicine may end up working differently than the original drug.
But amid the growing popularity and high price tags of many biologics, Congress is expected to consider a regulatory pathway next year to bring biosimilars to market. President-elect Barack Obama has said he supports biosimilars.
Several large drug makers, threatened by patent expirations on top-selling products, are looking at biosimilars as a potential source of revenue. Merck said earlier this month it would start a new unit to copy biologics, and Eli Lilly has also expressed interest in the market.
In an interview published last week by the Financial Times, AstraZeneca CEO David Brennan said the company was studying the launch of biosimilar products, although he said such a move would depend on the legislation being considered by Congress.
AstraZeneca, whose U.S. headquarters is in Fairfax, said in a statement that MedImmune has facilities well-equipped to produce biosimilars, "should we choose to do so and if the legal and regulatory framework allowed.
"However, at the current time, we see the strongest opportunities for the business in flexing its track record of innovation, developing its pipeline of potential biologic candidates to treat or prevent a number of debilitating or life-threatening diseases," the company said.
U.S. and European regulators have a streamlined approval process for generic versions of conventional small-molecule drugs, which are easier to copy than biologics. The European Union has an approval procedure for certain biologics.
Novartis AG's generic-drug unit two years ago became the first company to have a biosimilar product approved: the growth hormone Omnitrope.
The European Commission last August cleared Novartis' anemia drug that is similar to Johnson & Johnson's Eprex and Amgen's Epogen.

Thursday, November 20, 2008

Biosimilar

Biosimilars or Follow-on biologics are terms used to describe officially approved new versions of innovator biopharmaceutical products, following patent expiry.
Unlike the more common "small-molecule" drugs, biologics generally exhibit high molecular complexity, and may be quite sensitive to manufacturing process changes. The follow-on manufacturer does not have access to the originator's molecular clone and original cell bank, nor to the exact fermentation and purification process. Finally, nearly undetectable differences in impurities and/or breakdown products are known to have serious health implications. This has created a concern that copies of biologics might perform differently than the original branded version of the drug. However, similar concerns also apply to any production changes by the maker of the original branded version. So new versions of biologics are not authorized in the US or the European Union through the simplified procedures allowed for small molecule generics. In the EU a specially-adapted approval procedure has been authorized for certain protein drugs, termed "similar biological medicinal products". This procedure is based on a thorough demonstration of "comparability" of the "similar" product to an existing approved product. In the US the FDA has taken the position that new legislation will be required to address these concerns. Additional Congressional hearings have been held, but no legislation had been approved as of December 2007.

Recent FDA Actions Fuel Debate Over Copycat Biotech Drugs11-19-08 3:47 PM EST

NEW YORK -(Dow Jones)- The Food and Drug Administration's scrutiny of production changes by Genzyme Corp. (GENZ) and Amylin Pharmaceuticals Inc. ( AMLN) may signal a tougher stance in eventually evaluating generic versions of biologic drugs - should they ever become legal.

The market for so-called biosimilars could grow to as much as $200 billion a year by the middle of the next decade, as recently estimated by an industry executive, but their regulation will likely be more rigorous than that enjoyed by chemical counterparts. That scrutiny could make their development more difficult and expensive for generic drug makers, possibly hurting sales and forming a barrier to entry that allows only the largest companies to participate.
"This could be a concerted effort on the part of the FDA to draw a line in the sand in advance of a biosimilar pathway," said analyst Chris Raymond with Robert Baird & Co.
The FDA denied that it has changed its policies, saying that it "has had clear and consistent guidance about comparability since 1996." The agency wouldn't comment further.
No pathway for generic biologics exists in the U.S., but legislation to provide a pathway for generic versions is widely expected to be among President- elect Barack Obama's agenda. An official with Obama's transition team declined to comment on the issue.
Currently, generic drug makers can receive approval of copycat small-molecule drugs, like cholesterol-fighting statins, by showing they have the same active ingredient and the same action as the brand-name version, which allows the generics to depend on the original clinical trials and avoid having to pay for new ones.
Biotech drugs, made by culturing specially engineered organisms, are large proteins that are sometimes thousands of times bigger than small-molecule drugs. Their manufacturing makes them sensitive to minor changes in the process, potentially altering their complicated structures and even how they work in the body.
The biotech industry has long argued the complicated nature of the drugs makes it hard for a generic company to copy the drug, and expensive clinical trials should be used to prove similarity.
Earlier this year, the FDA decided that a version of Genzyme's Myozyme, to treat a rare enzyme disorder, produced on a larger scale had slight differences and had to be reviewed as a separate product with clinical data.
"I think what they have done with Myozyme is a pretty big departure," said Raymond, who notes that treating the larger-scale production as a separate brand was "unimaginable" until recently.
The FDA also recently requested more information on the comparability of Amylin Pharmaceuticals' Byetta LAR, an experimental once-weekly version of already approved twice-daily Byetta for diabetes. The issue is between batches of the drug made by partner Alkermes Inc. (ALKS) in its facility, used in previous clinical studies, and batches made on a commercial scale in Amylin's Ohio facility.
Barrier To Entry
The size of the generic biologics markets is unclear. In a 2007 report, Cowen & Co. estimated that U.S. sales of major biologics totaled $25 billion in 2006. Assuming lower prices, and limited penetration of generics, the firm estimates that the total generic revenue from those sales at $2 billion to $7 billion.
That differs greatly with the more recent projection of worldwide biosimilars sales of $200 billion by 2015 from Teva Pharmaceutical Industries Ltd.'s (TEVA) North American chief executive, Bill Marth.
But Raymond points to his own research that shows biosimilars of Amgen Inc.'s (AMGN) anemia treatments aren't being widely adopted in Europe yet.
Understandably, the biotech industry is hoping that the U.S. policies are tougher than in Europe, and it has long pushed for heavy scrutiny, citing the complexity of the products and processes.
The industry, led by the Biotechnology Industry Organization, advocates for clinical data requirements and fighting interchangeability, which allows the generic to be substituted for the branded drug, citing potential safety issues from imperfect drug copies.
While all parties involved are concerned about safety, those policies erect a number of hurdles for the generic companies.
Many observers expect biosimilars to require clinical data to some degree and be distinct products that must be marketed and specifically prescribed by physicians. That may make the drugs more expensive to develop and possibly less lucrative.
Furthermore, the scientific, manufacturing and marketing investment needed to enter such a market will likely allow only the biggest of the generic drug makers to take part, including Teva, Mylan Inc. (MYL) and Novartis AG (NVS).
"This is going to be a big thing. This is going to be very expensive, very intensive," Marth said. "I can't imagine somebody investing less than $1 billion and getting involved in this."
Teva has positioned itself to benefit from any regulatory pathway for biosimilars in the U.S., including its pending $7.46 billion acquisition of Barr Pharmaceuticals Inc. (BRL).
Evan McCulloch, a mutual fund manager with Franklin Templeton, believes that generic companies will have a tougher time selling generic biologics than small- molecule drugs.
He expects clinical trial requirements and companies having to sell biosimilars like a branded product using an expensive sales force, which is a new strategy for most generic companies. All of that could bode well for the biotechnology companies that would face sales pressure from generic competition.
"It is one thing when that drug goes generic and essentially disappears within three months," said McCulloch, referring to the situation seen with small- molecule drugs when generics enter the market, "and another thing entirely when you can bet that that drug is going to hold onto some of its revenues into perpetuity."
-By Thomas Gryta, Dow Jones Newswires; 201-938-2053; thomas.gryta@dowjones.com

Sunday, November 16, 2008

Herbal medicine

I have been thought that the chinese traditional medicine from herbal is typically safe. However, recent discussions with my friends make me extremely nervous about the safety of the herbal medicine. The recent report (see below) is just one of the examples. The reason could be in multifold: 1) the safety is rarely tested in human trials; 2) counterfeit or shoddily made medications ; 3) the original herbal was now grown and harvested in total different climate/environment - the ingredient might be different from the original intended ingredient, some could be toxical. 4) contamination of the herbal raw materials.

China recalls hemorrhoid medicine
The Associated Press
Published: November 12, 2008
BEIJING: China's drug regulator ordered a nationwide recall of a hemorrhoid medicine Wednesday because of concerns it may cause liver problems.
The State Food and Drug Administration said in a statement on its Web site that it had ordered Vital Pharmaceutical Holdings Ltd., based in Sichuan Province, to stop producing Zhixue capsules and begin a nationwide recall. Twenty-one people around the country developed liver problems after taking the medicine in recent months.
"An obvious connection can be found between the hemorrhoid medicine and the liver damage after case analysis, but the cause of the adverse reactions remains unknown," the statement said.

China's pharmaceutical industry is highly lucrative but poorly regulated, resulting in some companies using fake or substandard ingredients. In recent years, a string of fatalities blamed on counterfeit or shoddily made medications has been reported.
Several herbal medicines have been recalled in recent months because of suspicions they have caused deaths, according to the official Xinhua News Agency.

The recalls come as China tries to reassure consumers over a scandal involving the spread of the industrial chemical melamine into the food chain, the latest incident to mar its already troubled product safety record.

Wednesday, November 12, 2008

Analysis Problems with Subgroup Analyses

Sub-grouping damages the balance obtained by randomization

  • If the randomization is stratified for one factor (for example, disease severity), it will ensure the balance of the treatments inside the subgroups defined by that factor but not necessarily the balance of other prognostic factors (unless the subgroups are very large)
  • When minimization is used, the balance for other stratification factors (eg., age category) inside the subgroups is not guaranteed.

Treatment comparisons within subgroups lack power

  • the planned sample size N is large enough for detecting a specified difference in the WHOLE group
  • Sub-grouping -> smaller sample size for each comparison -> lower power
  • The statistical power to detect a treatment by subgroup interaction (ie. different treatment effects between subgroups) is usually very low

It is always possible to find subgroups in which the treatment effect is more extreme than the overall effect (data dredging)

  • It is always possible to find a grouping of the sample such that the treatment effect is more pronounced in one subgroup and less pronounced in the other
  • Indeed, the overall treatment effect is a sort of average of the subgroup treatment effects
  • It is always possible to find a subgroup with a significant difference just by chance!

Subgroup anlaysis induce multiple testing problems

  • Suppose you perform K tests, each of them at the alpha=0.05 significant level, the overall type I error rate (the risk of finding at least one spurious statistically significant result among the K tests) is alpha(overall) = 1-(1-alpha)^k
  • The Bonferoni adjustment must be used to maintain the overall alpha close to 0.05: use alpha/K for each test

Improper subgroups

  • Improper sugroups: subgroups of patients classified by an event measured after randomization and potentially affected by treatment - Response, means or survival comparisons to therapy, by compliance, by severity of side effects, or any factor not stratified for
  • Inherent prognostic features inflence both the endpoint and the event
  • Lead time bias: those who have the event early necessarily fall in the "poor" classification
  • No causality relationship can be demonstrated

Thursday, November 06, 2008

FDA Revises Process for Responding to Drug Applications

The following annoucement really makes sense. Previously, FDA could issue an "approvable" letter that could be very confusing. A product is 'approvable' based on efficacy, but can not be approved due to other safety concern.

http://www.fda.gov/bbs/topics/NEWS/2008/NEW01859.html

The U.S. Food and Drug Administration today announced that it is revising the way it communicates to drug companies when a marketing application cannot be approved as submitted.

Under new regulations that govern the drug approval process, FDA's Center for Drug Evaluation and Research (CDER) will no longer issue "approvable" or "not approvable" letters when a drug application is not approved. Instead, CDER will issue a "complete response" letter at the end of the review period to let a drug company know of the agency's decision on the application.
"These new regulations will help the FDA adopt a more consistent and neutral way of conveying information to a company when we cannot approve a drug application in its present form," said Janet Woodcock, M.D., director of the agency's Center for Drug Evaluation and Research (CDER). "Thorough and timely review of drug applications is a priority of the FDA, and these new processes will make our communications with sponsors of applications more consistent."
Taking the place of "approvable" and "not approvable" letters, a "complete response" letter will be issued to let a company know that the review period for a drug is complete and that the application is not yet ready for approval. The letter will describe specific deficiencies and, when possible, will outline recommended actions the applicant might take to get the application ready for approval.

Currently, when assessing new drug applications, the FDA can respond to a sponsor in one of three types of letters: an "approval" letter, meaning the drug has met agency standards for safety and efficacy and the drug can be marketed for sale in the United States; an "approvable" letter, which generally indicates that the drug can probably be approved at a later date provided that the applicant provides certain additional information or makes specified changes (such as to labeling); or a "not approvable" letter, meaning the application has deficiencies generally requiring the submission of substantial additional data before the application can be approved.
"Complete response" letters are already used to respond to companies that submit biologic license applications. The process for drugs and biologics will be consistent under the new regulations.

The revision should not affect the overall time it takes the FDA to review new or generic drug applications or biologic license applications. These changes, which will become effective on Aug. 11, 2008, are not expected to directly affect consumers.
In July 2004, the FDA issued a proposed rule on these topics. At that time the agency asked for comments on the proposal. Today's final rule addresses comments submitted to the agency.
For more information, see:

Link to the Complete Response Final Rulehttp://www.fda.gov/cder/regulatory/complete_response_FR/default.htm
Link to the drug approval process pagehttp://www.fda.gov/fdac/special/testtubetopatient/default.htm

Wednesday, November 05, 2008

PRO, CRO, and Laboratory tests / device measurements

In an article by Willke et al (Controlled Clinical Trials, 25, 2004), the study endpoints were classified as three major categories. Endpoints were classified into the following three major categories, and the presence or absence of each of these categories was noted for each product reviewed. Each product may have employed one, two, or all three types of endpoints:

  • Laboratory tests and device measurements,
  • Clinician-reported outcomes (CROs)
  • Patient-reported outcomes (PROs).

Laboratory and device measurements included highly objective typically numerical measures often performed by machine.

Clinician-reported outcomes included those that might be considered traditional endpoints, either observed by the physician (e.g., cure of infection and absence of lesions) or requiring interpretation by the physician (e.g., radiologic results and tumor response). In addition, CROs included both formal and informal scales completed by the physician using information about the patient. CROs requiring patient input are distinguished from clinicianadministered PROs in that the former requires clinician judgment or interpretation when recording answers, while the latter involves recording precise, unmodified patient responses to prespecified questions.

Finally, endpoints classified as patient-reported outcomes included formal health-related
quality of life measures and any other endpoint that was primarily based on a direct patient report. PROs categorized as "formal" scales are those multiitem questionnaires that have a well-defined standardized format, well-documented procedures for administration and scoring, demonstrated reliability and validity, and some guidelines for interpretation of scores. Other PROs included informal symptom scales, patient global assessments, or visual analog scales, as well as patientreported endpoints recorded in event logs (e.g., specific events). In some cases, nonclinician proxies reported the outcome from the perspective of the patient (e.g., when vaccines were tested in infants); these endpoints were considered patient-reported.

Tuesday, November 04, 2008

Declaration of Helsinki and FDA

The newly released Declaration of Helsinki was issued by the 59th World Medical Association General Assembly in October 2008. This document details ethical principles for medical research involving human subjects.

Section 19 requiring every clinical trial to be registered before recruitment of the first subject. Also note Section 30 on the obligation to make public the results of research on human subjects and requirements for publications. The additional contents are in line with the recent push for registry of the clinical studies and publication of the clinical trial results.

http://www.wma.net/e/policy/pdf/17c.pdf
http://www.wma.net/e/index.htm
http://en.wikipedia.org/wiki/Declaration_of_Helsinki

However, the FDA is moving away form the Helsinki accords because of what it says about placebo. The following two links discussed this issue.
http://www.socialmedicine.org/2008/06/01/ethics/fda-abandons-declaration-of-helsinki-for-international-clinical-trials/
In 21 CFR 312, "Human Subject Protection; Foreign Clinical Studies Not Conducted Under an Investigational New Drug Application- Notice of Final Rule", FDA states
" The final rule replaces the requirement that these studies be conducted in accordance with ethical principles stated in the Declaration of Helsinki (Declaration) issued by the World Medical
Association (WMA), specifically the 1989 version (1989 Declaration), with a requirement that the studies be conducted in accordance with good clinical practice (GCP), including review and approval by an independent ethics committee (IEC)."

A article on EMBO report (7(7), 2006) titled "The Battle of Helsinki" is worth to read.

North Carolina's Triangle Business Journal (2/19, Gallagher) reports that "the study also questions the decision by the US Food and Drug Administration in 2008 to abandon the Declaration of Helsinki, a set of standards adopted by the World Medical Association in 1984 that required trials to compare new drugs with the most effective alternative." The Food and Drug Administration "dropped the Helsinki standards in favor of the policy of Good Clinical Practice adopted by the International Conference on Harmonisation of Technical Requirements for Registration of Pharmaceuticals for Human Use. That policy, which allows drug manufacturers to compare the results of the new drug with those of a placebo, is considered by some to be less stringent than the Declaration of Helsinki."