Friday, June 03, 2011

Restructure FDA's Drug Review Process?


Last week, I had a chance to listen to a speech by Dr Scott Gottlieb. While he touched several topics in related to the health care reforms, I was specifically interested in his discussion on restructuring FDA’s drug approval process.

Dr Gottlieb gave a lot of insights within FDA and analyzed the root cause of the very long and inefficient FDA drug review process.

Following his speech, I located his paper “SHOULD FDA RESTRUCTURE ITS DRUG REVIEW PROCESS?” from FDLI’s website. A lot of his analyses are so true and to the point.

For example, he elaborated why FDA adopt a matrix management structure for its review program.

“Prior to FDA’s adoption of a matrix management structure for its drug review program, agency scientists were organized largely around the clinical areas in which they worked (oncology, cardio-renal, antiviral, etc). This therapeutically focused structure had some advantages, but also led to some of its own challenges. FDA’s adoption of a matrix structure was aimed at solving some of these problems.

For one thing, sponsors complained that the advice they received about disciplines like biostatistics or clinical pharmacology varied (sometimes significantly) across different therapeutic divisions. Statisticians in one clinical division would be interpreting certain principles of statistics or evaluating a particular protocol design in a manner different than statisticians inside another therapeutic division.

These discrepancies still occur. But it is believed that the matrix organizational structure cuts down on this sort of conflict.

Grouping all of the statisticians or pharmacologists inside the same office increases opportunities for comparable training and cross-calibration on key principles. CDER management pulled the first group of the review divisions—the chemists—in 1995. The impetus was differences in pharmaceutical quality requirements being maintained among the different clinical divisions. Ultimately, having the chemists organized as a single group fostered the development of consistent standards. It also enabled FDA to negotiate the standards established by the International Conference on Harmonization (ICH).

Another reason for establishing the matrix was to improve morale. FDA remains a very “physician centered” culture, but was much more so prior to adoption of the matrix. Staff who lacked medical degrees or who weren’t the clinical reviewers on an application sometimes complained that they felt marginalized in the review process. As one statistician told me, “we were treated like second-class citizens.” Specialists from non-clinical disciplines like statistics also complained that remaining immersed in a single therapeutic area didn’t give them the breadth of experience that they needed for their own professional development.

Similarly, the organization of scientific personnel by therapeutic area was also seen as an impediment to their continued training in their chosen disciplines. For example, statisticians came together for the equivalent of grand rounds or other kinds of shared learning experiences. But these kinds of cross-training opportunities were challenging because staff were ultimately accountable to their divisions. The shared training opportunities weren’t prioritized. Efforts were also made to rotate non-clinical experts across different therapeutic areas. But the challenges endured.”

He then went on discussing the issues with FDA’s weak matrix management structure.

“But in practical terms, the weak matrix means that FDA project managers have limited dominion over key aspects of the review. Key disciplines involved in the review aren’t accountable to the project manager, or the division director. It is a system where there are few management carrots and no sticks. This weak structure also makes it harder to organize collaborative projects or even team meetings. The project manager doesn’t have strong authority when it comes to managing the collaboration between the different scientists involved in a drug’s review.”

He also mentioned the quality of the FDA review scientist and issues with FDA’s policy to allow very flexible working hours and work-from-home schedule. The current FDA drug review team is loosely organized and inefficient in many aspects. However, it is not easy to make big changes to the current process.

Monday, May 30, 2011

Differences between SOPs and WPs


Working in industry setting (versus academic), we all understand that it is critical for us to be familiar with the Standard Operating Procedures (SOPs) / Working Procedures (WPs) and follow these procedures. This is even more obvious in pharmaceutical and drug development industry. An employee could be fired not because of incompetence, but because of violating these procedures. A couple of days ago, I attended ChineseAssociation for Science and Technology Annual Convention, one session was to discuss about the requirements for starting up a new company. Several local entrepreneurs elaborated the intellectual property (IP), business plan, funding, a loyalty team, and so on. Nobody mentioned these procedures. But I think that it is equally important to have a set of necessary SOPs and WPs in order to start a new company, especially a service-type company in highly regulated drug development field such as contract research organization. Scientists working in the academic setting often find the difficulties when they try to land a position in industry mainly because they lack the ‘industry experience’ – one critical part is that they lack the experience in following up the strict industry SOPs. If you interview a candidate who does not even understand the term ‘SOP’, you know that your candidate will have a long way to adapt to the industry setting. 

While there are definitions for Standard Operating Procedures (SOPs), it is hard to define Working Procedures (WPs), not to mention the differences between SOP and WP.

Standard Operating Procedure is established procedure to be followed in carrying out a given operation or in a given situation. In clinical trial setting, the SOP is defined in ICH E6 as “detailed, written instructions to achieve uniformity of the performance of a specific function. “  SOP is one of the components of a Quality Management System (QMS) implemented in many large organizations - improve quality by providing a standard format, approach, and application of common processes.. 

The issue may be with the wording ‘detailed’ in this definition. The common understanding within the industry is that SOP needs to be written in a little bit high level and the working procedure provides the detailed steps according to SOP. If the strict definition of SOP from ICH E6 is followed, there perhaps should not be any working procedures. All working procedures should be called SOPs. There is really no defined boundary for separate the SOPs and WPs. For SOPs, say what you do... but do what you say. The documentation is always important to the SOP compliance. 

My personal opinion is that SOPs are the procedures we must follow, typically on high level, with GCP compliance while WPs are procedures we should follow, typically in more detail, for internal standardization and efficiency. For example, in biostatistics, a SOP must be in place to require the independent validation for statistical outputs. A WP may be created to detail how independent validation should be implemented.

Here are some observations about SOPs and WPs:

1) Within FDA, the term SOPP (standard operating procedure and policies) is used for SOPs and WPs.

 2) There are two papers containing good discussions about SOP:
 3) There seems to be no formal definition ‘working procedure’. Different companies may call this differently and different people may have different understanding for ‘working procedure’.
There are some discussions on the web: 

Having SOPs and WPs in place is one thing; having everybody trained on SOPs and WPs is another thing. If SOPs and WPs are written in great detail like step-by-step instructions, the procedures may become a burden to be followed. This could in turn cause issues in SOP compliance when there is an audit. 

Last week, we had a group meeting during the lunch time and we ordered some food from Jason’s Deli. It was good to see that the delivery guy brought a Bagset Checklist which listed the items and quantities for each item. Unfortunately, the guy clearly did not follow their procedure and did not check the check list. When he delivered, he realized that he forgot to bring ‘bulk chips’. He had to go back to pick up this missed item. Had he follow the procedure and check his Bagset Checklist, he would have not missed any item. In this case, the consequence of the non-compliance is minor, but in clinical trials, the consequence of the non-compliance could sometimes be very significant.

Saturday, May 14, 2011

I-Spy 2 Trial - a study with fancy name and fancy design

If you see the term “I-spy”, you would typically think something related to movie, kids game or TV series. Recently, I-spy has been linked to a novice clinical trial design in breast cancer.

For a short description about this study, clinicaltrials.gov is a good place. In clinicaltrials.gov, I-SPY 2 TRIAL is listed as “Neoadjuvant and Personalized Adaptive Novel Agents to Treat Breast Cancer”. For more detail descriptions about this study, you should check out the following websites:
This study got a lot of publicities. For example, in FDA’s ‘advancing Regulatory Science for Public Health’, it specifically mentioned the I-Spy 2 trial,

Personalized treatment for cancer

The “I-SPY 2 TRIAL,” launched in March 2010, represents a groundbreaking new clinical trial model that will help scientists quickly and efficiently test the most promising drugs in development for women with higher risk, rapidly growing breast cancers. During the trial, drugs in development are individually targeted to the biology of each woman’s tumor using specific genetic or biologi­cal markers, known as “biomarkers.” By applying an innovative trial design, researchers will use data from one set of patients’ treatments to treat other patients – more quickly eliminating inef­fective treatments and drugs. The I-SPY 2 trial was developed under the Biomarkers Consortium, a unique public-private partnership that includes the FDA, the National Institutes of Health, and major pharmaceutical companies, led by the Foundation for the National Institutes of Health
. “

The claimed advantages for this trial design are:
  • Utilized Personalized Medicine
  • Uses genetic or biological marker (“biomarkers”) from individual patients’ tumors to screen promising new treatments, identifying which treatments are most effective in specific types of patients
  • Improved Efficiency (fewer patients, less time, and fewer resources)
Enable researchers to use early data from one set of patients to guide decisions about which treatments might be more useful for patients later in the trial, and eliminate ineffective treatments more quickly Enable the development of more informed, smaller phase III trials

So what is this trial design? Can this trial really achieve its purpose? Here is what I would say:

  • I-Spy 2 is a phase II study, anything generated from this study will need to be confirmed in Phase III confirmatory studies.
  • I-Spy 2 is a government-sponsored study and supported by Foundation of NIH, Biomarkers Consortium, NCI, and others. It may never happen if it is an industry-sponsored study.
  • I-Spy 2 is academic study and not a study for product licensure. If it is study for product licensure, there may be difficulties for study design to be accepted by regulatory agencies.

  • I-Spy 2 design is complicated and it employs Bayesian Adaptive, Response Adaptive Randomization (play-the-winner approach), Biomarker Adaptive, Stopping and adding treatment arms. The algorithm and stopping rules are complicated and the assumptions / basis for generating these algorithms / stopping rules may eventually be demonstrated incorrect. 
  • I-Spy 2 is a randomized, open label study. It would be more difficult to implement if this is a double-blinded study
  • I-Spy 2 got it fame due to its use of ‘personalized medicine’, ‘adaptive design’ – all fit into the hot areas
  • I-Spy 2 is a multi-center, NOT multi-national study. It would be much more challenging if this is a multi-national study
Although I-Spy 2 got a lot of attentions, not everyone is convinced with this design. The study is closed watched – let’s see the fate once it is completed.

Friday, April 29, 2011

Mobile phone text message used in clinical trial?

Recent discussions with my friends in China surprised me a little bit. They are way ahead in applying the advanced technology in public health area. They used the mobile phone text message to promote the smoking cessation. They recently applied research grant to study the effect of mobile phone text message in improving the maternal and child care. For pregnant women or new mothers, the mobile phone text message is used to send reminders, maternal care tips, child care tips,...

There have been some publications demonstrating the effectiveness of text message in these public health promotion areas. For example, a study in New Zealand demonstrated that smoking cessation using mobile phone text messaging is as effective in Maori as non-Maor. The newscientist.com reported that text messages double young smokers' quit rates.There are

It is natural to think that such text message approach can be used in clinical trials. When I search the clinicaltrials.gov, I can find many clinical trials using text message mostly for improving the adherence of clinical visits or adherence of drug taking.

In my opinion, effectiveness of using text message in clinical trials depends on the study population and the country the study is conducted. In China, if the study is conducted in urban areas, mobile phone text message could be very effective because 1) almost everyone has mobile phone; 2) people use text message more often than actual calling. In the United States, for general population, text message may not be a good approach because some family may still rely on residence-line phone instead of cell phone. Even though they have cell phone, they rarely use text message on daily basis. If a study is conducted in high school students or college students, text message could be an effective tool since all students like to use text messages.

Saturday, April 16, 2011

Emerging Statistical Issues in the Conduct and Monitoring of Clinical Trials

This Wednesday, I had a chance to attend “University of Pennsylvania Annual Conference on Statistical Issues in Clinical Trials”. The topic for this year is “Emerging Statistical Issues in the Conduct and Monitoring of Clinical Trials”. The number of participants was just right in size and the conference was organized pretty well.

In terms of the topics, there are some of them I like and some of them I don’t like. The presentation slides will eventually be posted on conference’s website, however, I would like to give one or two sentences commenting on each topic.

“Sample size estimation incorporating disease progression” – the key issue is the adequacy of the study endpoint. A good endpoint will incorporate the impact of the disease progression.

“Hurdles and future work in adaptive designs” – it is good to hear the discussion about the hurdles, caveats of the adaptive designs. Still very often, a lot of people only talk about the advantages of adaptive designs – too good to be true.  Similarly, a recent article "a once-rare type of clinical trial that violates one of the sacred tenets of trial design is taking off, but is it worth the risk? " from The-Scientist magazine gave some objective assessments on implementing the adaptive designs.

“Predicting accrual in ongoing trials” – utilizing the complicated statistical model to predict the accrual is a waste of time. Accrual in ongoing clinical trials is 95% clinical operations issues, 5% related to statistics. Is it worth to modeling the accrual?

“New incentive approaches for adherence” – money incentives including lottery is a sensitive topic and ethic issue could follow no matter it is incentive for adherence or for study visit compliance. Money incentives are different depending on participants’ social economic status (family income). $100 lottery may be very incentive to some, but not to others.  

“Efficient source dada verifications in cancer trials” – I always thought that all data fields had to be 100% source data verified. It is not entirely true in large scale trials in oncology or in studies with cardiovascular endpoint. In industry, we are rather conservative.

“Estimation of effect size in trials stopped early” – trials stopped early due to efficacy is not very common and should not be encouraged. Difficulty in estimating the effect size still exists for trials stopped early.

“Accounting in analysis for data errors discovered through sampling” – Unreliable data or large % of missing data is always a concern, even for observational studies. Statistical approach may not be a good option. When data is garbage, the results we draw from the data will also be garbage – so called ‘garbage in, garbage out’ no matter which statistical model is utilized to address the data issues.

“Some practical issues in the evaluation of multiple endpoints” – It is so correct that we should play down the importance of differentiating ‘primary endpoint’, ‘secondary endpoint’, ‘tertiary endpoint’,… Multiple comparison has been expanded so much and is everywhere now (co-primary, primary and secondary, co-secondary, secondary superiority test after non-inferiority test, interim analysis, meta analysis, ISE,…). Are we overdoing this?

Saturday, April 09, 2011

Sparse sample and population pharmacokinetics

In drug development, it is necessary to understand the pharmacokinetics profiles (or time concentration profiles) of the experimental drug and calculate the pharmacokinetic (PK) parameters (Area Under the Curve – AUC, Clearance – CL, or Volume of distribution –Vd). These PK parameters can provide the estimate of the dose exposure and assist in the decision on dose timing and dose interval. In order to calculate the PK parameters, we typically need a serial of blood samples at multiple time points (usually more than 6) after the drug administration. In some situations, it is not feasible or not practical to obtain these many blood samples. The obvious example is in pediatric studies where it is not feasible to obtain multiple blood samples due to the blood volume restriction. The specimen may not just be blood samples. If the PK is conducted using other specimens, it is usually difficult to obtain multiple PK samples. For example, we could obtain middle ear fluid (MEF) sample to determine the antibiotic drug concentration in the ear and bronchoalveolar lavage (BAL) to determine the drug exposure in the lung. It is not practical to obtain multiple samples for these special specimens due to the safety concern.

When very few samples are available for each patient, we call it ‘sparse sampling’. With sparse data, we would need to employ a
Population PK
approach to estimate the PK parameters, describe the PK profile, or do PK/PD modeling. The use of population PK during the drug development has been steadily increasing. Regulatory agencies have issued several guidance on the use of population pharmacokinetics.




There are different sparse sample designs. Below are some of the sparse sample designs I have seen.  

An example of sparse sampling at fixed time points is described in a paper by Vogelmeier et al. They used BAL fluid sample to study the intrapulmonary half-life of aerosolized product in Normal Volunteers”.

For BAL fluid samples, it is not feasible to obtain serial samples at all six time points (at screening, 0.5, 6, 12, 24, and 36 h). Therefore, in this study, “each volunteer underwent two BALs. The first lavage was done in the screening phase with an interval of between 3 and 7 d before inhalation of the drug. The volunteers were randomly assigned to one of five groups with the second lavage following 0.5, 6, 12, 24, or 36 h after aerosol administration. Each of the groups consisted of six individuals…”

Subjects in group 1 contributed two BAL samples at Screening and at 0.5 hours after inhalation.
Subjects in group 2 contributed two BAL samples at Screening and at 6 hours after inhalation.
Subjects in group 3 contributed two BAL samples at Screening and at 12 hours after inhalation.
Subjects in group 4 contributed two BAL samples at Screening and at 24 hours after inhalation.
Subjects in group 5 contributed two BAL samples at Screening and at 36 hours after inhalation.

With subjects from all five groups combined, a overall picture of the PK profiles over 24 hours after inhalation could be described. Original paper provided only the summary analysis. Nowadays, the data could be further analyzed using nonlinear mixed model from population PK model with software such as NONMEM.

In FDA guidance on Population Pharmacokinetics, an example was provided for estimating the AUC using sparse data (1-2 middle ear fluid samples per subject) in pediatric subjects.

“The penetration of drug X into middle ear fluid (MEF) was investigated using population PK analysis with sparse data (1-2 samples per subject) obtained from 36 pediatric patients (2 months to 2.0 years of age) who underwent clinical therapy with drug X. The estimated area under the concentration-time curve (AUC) that was above the minimum inhibitory concentration (MIC) (AUCMIC) and the half-life of drug X are 12.5 ug.hr/ml and 6.1 hours in MEF, respectively, vs. 23.7 ug.hr/ml and 3.2 hours in plasma, respectively….”

With this short description, we don’t know if MEF samples are taken from subjects at various times or fixed times. However, non-linear mixed model must have been used for analyzing the data.  

FDA’s guidance on population pharmacokinetics states, “the full population PK sampling design is sometimes called experimental population pharmacokinetic design or full pharmacokinetic screen. When using this design, blood samples should be drawn from subjects at various times (typically 1 to 6 time points) following drug administration. The objective is to obtain, where feasible, multiple drug levels per patient at different times to describe the population PK profile. This approach permits an estimation of pharmacokinetic parameters of the drug in the study population and an explanation of variability using the nonlinear mixed-effects modeling approach. “

If a full population PK sampling design is used, the sampling scheme will be something like below. The different subject could contribute different number of samples at various times.  


Subject number
Blood sampling time (t)
concentration at time t
 Ct
001
Predose
xxx
001
24 hours post dose
xxx
002
Predose
xxx
002
8 hours post dose
xxx
002
12 hour post dose
xxx
003
Immediately postdose
xxx
003
5 hour post dose
xxx
004
4 hour post dose
xxx




Then when non-linear mixed model such as NONMEM is used to fit the data to characterize the PK profile with PK parameter (such as AUC) = function of concentration (Ct) at time t.

In multiple dose studies, if the purpose is to characterize the PK profile at steady state, one could implement a strategy of splitting the number of samples into different dose intervals.
Suppose we need 8 serial blood samples (t1 to t8) to calculate AUC and the dose interval is weekly, we can have these 8 samples split into 4 dose cycles. For each subject, we would only take two samples for each dose cycle. At steady state, for each subject, we expect PK profile after each repeat dose is not much different; the concentration at day 1 after repeat dose #1 would be similar to the concentration at day 1 after repeat dose #4, and so on. In this case, we would be able to calculate AUC for each subject with 8 samples from four dose intervals (instead of 8 samples from one dose interval over 7 days). The drawback is that the study period would be longer.

Friday, March 11, 2011

Use of SF-36 in Clinical Trials

The SF-36 is a multi-purpose, short-form health survey with 36 questions. SF-36 is one of the most popular instruments for generic health surveys and it can be used across age, disease, and treatment group, and are appropriate for a wide variety of applications. Conversely to generic health surveys, disease specific health surveys are focused on a particular condition or disease. In clinical trials, SF-36 remains as one of the most common instruments for assessing the Health Related Quality of Life (HQOL), especially in diseases where there is no valid disease-specific tool.
 
SF-36 yields an 8-scale profile of functional health and well-being scores (so called domain scores) as well as psychometrically-based physical and mental health summary measures [physical component summary (PCS) and mental component summary (MCS)] and a preference-based health utility index (question #2).
 
The mapping from the original questions -> 8 domains -> PCS or MCS is sketched in the diagram below. Notice that only 35 out of 36 questions are used in this diagram. The question 2 asks about the general health status and does not contribute to the calculation of domain scores and component summaries. A good use of question 2 is to use its responses as anchor in identifying the minimal clinically important difference (MCID). In one of our publications in J Neurol Neurosurg Psychiatry, we indeed used this approach to identify the MCID.
 
For these 36 questions, the response categories vary depending on the question. The response categories range from 2 (yes, no) to 6 (all of the time, most of the time, a good bit of the time, some of the time, a little of the time, none of the time). Therefore, in order to calculate the domain score, a scoring method or algorithm has to be employed. For PCS and MCS, the calculation will be based on equations with coefficients from the regression models generated from the General Healthy Populatoin. In US, it is the Healthy General US Population. If different healthy population is used, the factor score coefficients for the Z_scores will be different and PCS and MCS values will be different.
 
The details about scoring method can be found at QualityMetric’s website. The scoring and calculation of component summaries require the programming. Some of the example programs (but not validated) can be found from the web:
Some questions and answers on using SF-36 in clinical trials:
 
Q: Is SF-36 free for using in clinical trials?
A: It is not free. License has to be obtained for using in industry-sponsored clinical trials. See qualitymetric website for detail.
 
Q: Why do we have question #2 that is used in calculation of any domain score and component summary?
A: It can be used as an assessment of general health status and also as an anchor for identifying MCID.
 
Q: Which general health population should be used for norm-based scoring?
A: The advantage of norm-based scoring is to facilitate the comparisons. If a study is a US domestic study, General Healthy US Population should be used. If it is an international study, the country-specific General Healthy Populations are preferred. SF-36 has been validated in many languages.
 
Q: What will be language to describe the statistical analysis plan for SF-36
A: For study protocol or for journal article statistical method section, analysis plan for SF-36 should be kept simple. In one of our publications on SF-36, we simply said:
"The corresponding physical component summary and mental component summary values for the randomized participants were calculated using the reported means, SDs, and factor score coefficients that came from the healthy general US population in 1990. A linear T-score transformation method was used so that both the physical component summary and the mental component summary scores were standardized with a range of 0 (lowest) to 100 (highest)"
 
Q: Could SF-36 be used in cost utility analysis?
A: No. SF-36 is not a utility score. However, Sf-36 can be converted to utility score (such as EQ-5D). See my previous blog
 
Q: Could we have one overall score for SF-36?
A: No. PCS and MCS have to be analyzed separately. You can not add PCS and MCS to have a single overall score.
 
Q: How to analyze the domain scores and component summaries?
A: Typically, 8 domain scores and 2 component summaries can be analyzed separately using analysis of variance or analysis of covariance or other methods such as repeat measurement depending on the study design.
A good approach in analyzing the SF-36 is to compare the each domain score with the General Healthy Population to show how much difference between the patients in the study and the General Healthy Population for pre-treatment and for end treatment visits. This approach was utilized in our SF-36 publication in Neurology.

Thursday, March 03, 2011

Incidence Rate (IR) – How could this be wrongly calculated?

I am very surprised to see how a simple concept of ‘incidence rate’ can be wrongly calculated in documents  submitted to regulatory agencies (such as FDA). In a briefing document titled “Tiotropium (SPIRIVA): Pulmonary Allergy Drug Advisory Meeting – November 2009” submitted by a sponsor, there were wrong statements every where about the calculation of the incidence rate for safety variables.

For example, on page 50, it says “Incidence rates of adverse events were computed as the number of patients experiencing an event divided by the person-years at risk”; In Section 8.1.5 (Statistical methods), it says “For each event, an incidence rate (IR) was calculated from the number of patients with an event divided by the cumulative time at risk within a treatment group and expressed as patient-years.”  In their summary tables, they footnoted “the number of patients with an event” (instead of the number of total events) was used in calculating the incidence rate. They never listed the total number of patient year (the denominator) for their Incidence rate calculation. In ‘Statistical method’ section, they even tried to justify the use of “the difference in incidence rate” because “most Tiotropium trials have significantly greater number of patients in the placebo group discontinuing the trial early compared to tiotropium treated patients.”
“Incidence Rate” is a basic concept from epidemiology studies and is calculated as the number of events divided by the number of patient years. According to free medical dictionary, “incidence rate is the probability of developing a particular disease during a given period of time; the numerator is the number of new cases during the specified time period and the denominator is the population at risk during the period. “   According to Wikipedia, “The incidence rate is the number of new cases per population in a given time period. When the denominator is the sum of the person-time of the at risk population, it is also known as the incidence density rate or person-time incidence rate. In the same example as above, the incidence rate is 14 cases per 1000 person-years, because the incidence proportion (28 per 1,000) is divided by the number of years (two). Using person-time rather than just time handles situations where the amount of observation time differs between people, or when the population at risk varies with time. Use of this measure implicitly implies the assumption that the incidence rate is constant over different periods of time, such that for an incidence rate of 14 per 1000 persons-years, 14 cases would be expected for 1000 persons observed for 1 year or 50 persons observed for 20 years.”
In an article by Marco et al “Incidence of Chronic Obstructive Pulmonary Disease in a Cohort of Young Adults According to the Presence of Chronic Cough and Phlegm”, the incidence rate is correctly defined for calculation.
“Incidence rates of COPD were estimated as the ratio of the number of new cases and the number of person-years at risk (per 1,000), which were considered equal to the length of the follow-up for each member of the cohort.”

The key is that if you calculate the ‘incidence rate’, your numerator must be ‘number of events’, not ‘number of patients with an event’. For events that can only occur once in a lifetime for a specific patient (such as cancer), there may not be much difference between ‘number of events” and “number of patients with an event”. However, for events occurr more than one time for a specific patient, “number of events” and “number of patients with an event” are very different concepts.

In Tiotropium briefing document, the correct calculation for incidence rate should be ‘number of events (AEs or COPDs)’ divided by ‘the patient year’. It was simply wrong when they used ‘number of patients with an event’ as the numerator in their calculation of incidence rate. Their justification for using the difference in incidence rate is just the opposite of their statement. If placebo group has more dropouts, their way of calculating the incidence rate will overestimate the rate for placebo group and underestimate the rate for Tiotropium group. This can be easily illustrated using an example below:


Assuming 10 patients in Tiotropium and 10 subjects in Placebo group, 5 patients in Tiotropium group and 5 patients in Placebo group had at least one COPD during the study. The incidence of COPD will be 5/10 = 50% in both groups. Suppose it is a one-year trial, all patients in Tiotropium group completed the one-year and all patients in Placebo group completed only 6 months. The patient year will be 10X1 = 10 for Tiotropium group and 10x0.5 = 5 for Placebo group. The incidence rates now become 5/10 = 50% in Tiotropium group and 5/5 = 100% in Placebo group – this is just simply wrong. In this case, when the patient year (or person year) is used as denominator, the numerator used in the calculation should be the number of events, not the number of patients with an event.    

It is unfortunate this simple concept of ‘incidence rate’ has been wrongly calculated in Tiotropium studies. This wrong calculation may have been embedded in their paper published in prestigious New England Journal of Medicine.

If ‘number of patients with an event’ is used in the numerator, the denominator has to be the total number of patients (not the number of patient year). ‘Number of patients with an event’ divided by ‘number of total patients’ is called ‘incidence of events’ – this is a typical way when we summarize the adverse events in clinical trials.  

Friday, February 25, 2011

Study Center Pooling Strategy in Multicenter Clinical Trials

Pooling the study center for statistical analysis purpose is rather an old issue. However, we can still see the discussion o f study center pooling strategy or algorithm in the study protocol or the statistical analysis for multi-center clinical trials. When a clinical trial has multiple centers, study center or investigator site is usually included in the statistical analysis either by including as an exploratory variable in the model (for example ANOVA or ANCOVA) or by conducting the categorical analysis adjusted by study center (for example, Mantel-Haenszel test, Elteren's test, Wilcoxon rank sum test stratified by pooled center). However, there could be situation that some study centers have very few subjects and can not be directly included as a stand alone center for the analysis. In this situation, a pooling strategy is often employed to combine the small centers together. The reason for pooling the small centers instead of using center as random effect may be due to the factor that centers in the clinical trial are rarely a random sample of all possible centers. It is not uncommon to find the statistical analysis including pooled center in regulatory submission or in publications, for example, in NDA for Refludan (the analysis was stratified by pooled center) and in FDA advisory committee documents (… were analyzed using Wilcoxon rank sum test stratified by pooled center (centers that entered fewer subjects than a complete block were pooled by country)). Here are some of the example languages describing such pooling strategies:

“Statistical tests will be performed as two-sided tests and will be adjusted to the multi-centric design of the study. A center must have enrolled at least 8 subjects to be a standalone center in the analysis (centers enrolling less than 8 subjects will be pooled – will be done before the study unblinding”

“Study centers were pooled from largest to smallest until the pooled center had more than 5 subjects with post baseline data in each treatment group. No pooled center had more than 15% of the total number of subjects”

“The majority of study centers were small. A small center was defined as any center with <5 patients with postbaseline data in any treatment group, resulting in 5 large and 25 small centers. To avoid loss of information, small centers were pooled from largest to smallest until the pooled center had 5 patients in each treatment group. These centers were grouped into 11 pooled centers for the purpose of analysis."

In one of hypertension clinical trials, the pooling strategy is described as “To avoid loss of information, small centers (<5 per protocol patients) were pooled from largest to smallest until the pooled center had 5 per protocol patients in each treatment group. These centers were grouped into 19 pooled centers for the purpose of analysis. The pooling algorithm was predetermined before unblinding the data, and the pooling algorithm was described in the statistical analysis plan for the study. Considering the subjective nature of the pooling algorithm, albeit prespecified before completion of the study, an exploratory analysis was also performed with actual center as a fixed effect in contrast to pooled centers. This analysis did not change the inference.”

In a type 2 diabetes trial, a different pooling strategy was used “For all center stratified analyses, centers with <24 randomized and treated subjects were pooled on a geographical basis, independently of treatment identification.”

In a recent brief book for PDAC, the sponsor provided the detail pooling strategy for centers “Pooling algorithm for centers: For non-US sites, all investigative sites within a country with fewer than 10 randomized subjects will be combined into a single pooled site for analysis purposes. If a resulting pooled site still has fewer than 10 randomized subjects, then this pooled site will be further combined with the smallest unpooled site within that country. If there is not another unpooled site within that country, then the pooled site will be combined with the smallest pooled site from another country. This pooling process will continue until there are at least 10 randomized subjects in each pooled site. For US sites, all investigative sites within a geographic region with fewer than 10 randomized subjects will be combined into a single pooled site for analysis purposes. If a resulting pooled site still has fewer than 10 randomized subjects, then this pooled site will be further combined with the smallest unpooled site within that region. If there is not another unpooled site within that region, then the pooled site will be combined with the smallest pooled site from another region within the US. This pooling process will continue until there are at least 10 randomized subjects in each pooled site.”

As we can see from the examples above, the cut point for center pooling (5, 8, 10, or 24) is really arbitrary and there is no scientific basis for choosing one or another. The decision on the cut point may be based on the distribution of the number of subjects across centers.

Center pooling strategy could sometimes be questioned by the regulatory reviewers. For example, in BLA review of Rebif, FDA reviewer had concerns about the pooling strategy “The sponsor’s study center pooling strategy: Per the pre-specified strategy in the sponsor’s statistical analysis plan (SAP), pooling of study centers for inclusion of center as a main effect in analyses was to have been based on geographic considerations for small centers. In fact, the pooling strategy actually used was data driven which is problematic. NOTE: There were 56 participating centers from 9 countries. The smallest recruiting center had 3 subjects, 2 centers contributed 4 subjects, and 5 centers contributed 6 subjects each. The remaining centers contributed between 6 – 24 subjects each (CSR, Table 3, pp. 65-66). This reviewer performed analyses of major efficacy endpoints based on strict geographic pooling of centers into 3 groups (US, Canada, and Europe) as well as un-pooled analyses (not including the center effect). In addition, descriptive analyses for individual centers were also performed for the primary and major secondary efficacy endpoints. The sponsor’s positive statistical findings were found to be robust based on these analyses.”

In Biopharmaceutical Report (Summer 1998), Paul Gallo wrote an article titled “Practical Issues in Linear Models Analyses in Multicenter Clinical Trials” which contained a section discussing “construction of composite centers”. The caveats of using the composite centers are also discussed in the paper.

“In performing unweighted analyses, a practice of defining artificial “pooled” or “composite” centers is often employed; that is, data from different centers are treated in the analysis as if they came from the same center. A number of small centers may be combined, or one or more small centers may be combined with a larger center. This practice attempts to minimize the large variance inflation and data instability of unweighted analyses when there are very small centers. Composites may be constructed to the extent of eliminating empty cells to ensure that treatment effects are estimable in models containing interaction terms. More commonly, this is done to achieve some minimum cell size felt to appropriately limit the influence of individual observations; values around 5 are often chosen. ”

Arbitrarily pooling the centers sometimes does not make sense at all. This is exactly true when the centers with small number of enrolled subjects are pooled even though these centers are scattered in totally unrelated geographic regions or countries. When pooled center is used and the statistically significant center effect is detected, the interpretation of the results is difficult. Instead of the center pooling purely based on the number of enrollees, the geographic distribution of centers should be considered. In many cases, instead of pooling centers by the number of enrollees, we could use country and geographic region in the analysis. In one of our multi-national clinical trials, we grouped centers by geographic region as North American, South American, Eastern Europe, Western Europe, and Eastern Asia. The strategy worked very well.

If possible, we could use the random effect model to include the study site / center as random effect to avoid the center pooling. We could also use a center weighting strategy that is similar to the Meta analysis where centers with more subjects are given more weights.

Tuesday, February 08, 2011

Guidelines for Blood Volumes in Clinical Trials (Especially in Pediatric Clinical Trials)

Nowadays, the clinical study protocols are becoming more and more complicated and require more and more blood sample draws for various purposes. The blood samples are needed for testing the hematology, chemistry, immunogenicity (for biological products), biomarkers (for diagnostic or other purposes), pharmacogenomics,… In some clinical trials, additional blood samples (sample retains) may be drawn for future studies (even though we may not know what the future study will be). If the study has the component of pharmacokinetics, the many more samples (series blood samples) will be drawn within a short period to characterize the pharmacokinetic profile, estimate the total drug exposure (AUC), and calculate other pharmacokinetic parameters.

With increasing in the number of blood draws or the blood volumes, the ethic issue often arises, especially in clinical trials with children.

US FDA and EMA do not really regulate the maximum blood volume that can be drawn from a subject during the clinical trials. The requirements for limiting the blood sample volume may come from the National Institute of Health (NIH), American Academy of Pediatrics, World Health Organization (WHO), and European Union (EU) and are typically enforced by the ethic bodies such as Institute Review Board (IRB) and Ethics Committee (EC). The requirements on blood volume during the clinical trials may be different depending on the country and local IRB.

The blood volume drawn for pharmacokinetic studies in the pediatric population is specifically a concern and has been discussed extensively. Stephen RC Howie (2010) reviewed blood sample volumes in child health research: a review of safe limits in the Bulletin of the World Health Organization (BLT). WHO also has its guidelines on drawing blood: best practices in phlebotomy. The guidelines are not specifically for clinical trials, rather for general blood donations. The guidelines contain specific technical requirements for the blood drawn in pediatric and neonatal subjects.

In US, Code of Federal Regulations has a specific chapter (Part 46) to discuss protection of human subjects and the chapter contains a subpart D to address additional Protections for Children Involved as Subjects in Research. While there is no specific requirement on the limit of blood volume, the CFR indicated that the research involves no more than minimal risk to the subjects and IRB should take into account the purposes of the research and the setting in which the research will be conducted and should be particularly cognizant of the special problems of research involving vulnerable populations, such as children, prisoners, pregnant women, mentally disabled persons, or economically or educationally disadvantaged persons. Similarly, the American Academy of Pediatrics has its policy on Guidelines on Ethical Conduct of Studies to Evaluate Drugs in Pediatric Populations. The policy requires “…with the growing number of pediatric drug studies, IRBs need to be familiar with the various research-design methods that minimize risk to the child. Examples include limiting research under some circumstances to pharmacokinetic and safety data, combining this approach with pharmacodynamic data, and minimizing the volume of blood withdrawn through the use of sensitive assays, pediatric enabled laboratories, and population pharmacokinetic approaches"

National Institute of Health Clinical Center has a guideline M95-9: Guidelines for Blood Drawn for Research Purposes in the Clinical Center.

Two articles from the web actually reflect the limit of blood volume in the US.
In EU, there are specific guidelines on "ETHICAL CONSIDERATIONS FOR CLINICAL TRIALS ON MEDICINAL PRODUCTS CONDUCTED WITH THE PAEDIATRIC POPULATION"



The guidelines on blood volume are usually based on the amount of blood in the percentage of total blood volume (BLV). BLV varies depending on age and body weight. A good reference for BLV for pediatrics can be found in pediatricareonline.com.

Friday, January 28, 2011

Edit check - a critical step to ensure the data quality during clincial trials

In clinical trial, one critical task is to ensure that the data collected or data entered into the system / database is valid, correct, and logically sound. This task requires a data quality plan starting from designing a good study protocol -> developing efficient case report forms -> providing clear instructions for completing case report forms -> implementing electronic edit checks -> monitoring the study data / source data verification -> data clarification process -> data review process. One of the steps is to implement the electronic edit checks.
Edit check is a program instruction or subroutine that tests the validity of input in a data entry program. According to the CDISC clinical research glossary from Applied Clinical Trials, the edit check is defined as:

An auditable process, usually automated, of assessing the content of a data field against its expected logical, format, range, or other properties that is intended to reduce error. NOTE: Time-of-entry edit checks are a type of edit check that is run (executed) at the time data are first captured or transcribed to an electronic device at the time entry is completed of each field or group of fields on a form. Back-end edit checks are a type that is run against data that has been entered or captured electronically and has also been received by a centralized data store.

Electronic edit checks allow us to use the power of the computer to check for illogical, incomplete or inconsistent data. In clinical trial, one of the most important tasks facing clinical data management personnel is to produce the electronic Edit Checks specifications for a study. Developing the electronic edit check specification -- and processing the queries that result from them -- is arguably the most vital and time-consuming data cleaning activity data management personnel undertakes. The study statistician should always participate in the process of developing the electronic edit checks to ensure that the critical edit checks are included. Effectively implementing the edit check can prevent the illogical, incomplete, or inconsistent data from entering into the data capture system or data set, which will make the downstream data analyses much easier.

There are two types of edit checks:

Univariate edit checks (include range checks): these are the edit checks only applicable to a single field or single variable. For example, for subject weight, we can set up an edit check to ensure that the extreme or unlikely value not to be entered. Let’s say we set up a range check if a data entry is smaller than 90 lb or greater than 300 lb. For lung function test, we may set up an edit check for predicted FEV1 to be no less than 20% because it is unlikely to have someone with predicted FEV1 <20%. The univariate edit checks are usually run instantly during the time of data entry.

Multivariate edit checks (also called aggregate edit checks): these are the edit checks with more than one fields or variables involved. These edit checks cross check the entries across multiple fields / variables to ensure the data is logical and consistency. For example, if the entry on Gender field is ‘Male’, there should not be data for pregnancy test result field. If the reason for subject dropping out the study is entered as ‘adverse events’, there should be a corresponding entry in AE data set. Statistician can provide great inputs in identifying the multivariable edit checks. Some multivariate edit checks could involve the complicated algorithm and take considerable time to run. In this situation, the multivariable edit checks can be run at back-end at a specified interval (for example, 2 am at night).

One misunderstanding is to think that all data issues can be resolved by implementing the edit checks. Edit check is only one of the steps in the data cleaning process. Also, there should be balance in terms of the number of edit checks. Too many edit checks for non-critical fields could be very annoying for people who enter the data. This is especially true for clinical trials using electronic data capture (EDC) where the data entry responsibility is delegated to the investigator and study coordinators who may lose patient if there are too many pop-up messages during the data entry. For example, if the telephone number needs to be entered, an edit check to enforce the data entry to follow xxx-xxx-xxxx would be unnecessary (xxxxxxxxxx and 1xxxxxxxxxx should also be accepted) – this is an example I see in some of the web forms – very annoying).

Sunday, January 23, 2011

Regulatory Guidance on Source Data in EDC Trials

When we move toward the clinical studies using electronic data capture, the ‘source data’ or ‘source document’ has been an issue. Unlike the paper-CRF (case report form) based study, the source data in EDC study can be confusing and sometimes vague. If the data was directly entered into EDC system, the EDC system is the direct source and there is no another source to be verified against. This could be worrisome to some people. In a 2008 article, I talked about this issue.
Recently, both FDA and EMEA published the guidance on this issue. FDA’s guidance "Electronic Source Documentation in Clinical Investigations" was issued in December, 2010. EMEA issued its guidance last June and the guidance titled “Reflection paper on expectations for electronic source data and data transcribed to electronic data collection tools in clinical trials”.

The guidance titles seem to suggest that they are written for the data management functions, however, the discussions in these two guidelines are more relevant to the clinical sites and study monitors. Switching the clinical study from paper CRF to EDC is not just about the shift of the data entry from data management group to the clinical sites, it actually has impact on how the entire study is operated.

Tuesday, January 11, 2011

FDA's New Website for Industry

Have you noticed the changes in the design of FDA website (http://www.fda.gov/) recently? Last August, I mentioned the FDA's initiatives on transparency. As part of FDA's continued push to increase transparency in an agency once notorious for making decisions behind closed doors, the FDA has launched a new Web-based resource that industry can use to keep abreast of the regulatory status for drugs, devices, food, and cosmetics. The new website is under http://www.fda.gov/ForIndustry/ and is supposed to provide a repository for industries to understand FDA's detail processes in submission, reviewing, approval, and surveillance of the regulated products, and even the processes for complaints (dispute resolution). The website includes the sections that are very pertinent to us working in the pharmaceutical industry:
  • Developing products for rare disease and conditions
  • Dispute resolution
  • Guidance documents
  • FDA eSubmitter
  • Data standards
  • FDA basics for industry
FDA basics for industry includes the kind of basic information about the regulatory process that is often requested by drug, device, and biologic companies and is aimed at improving communication between FDA and industry by making basic information about the regulatory process more accessible to industry in a user-friendly format.

The new website reflects the great improvement towards the transparency and is a great resource for professionals working in the drug development industry.

Also see:

Sunday, January 02, 2011

Agreement Statistics and Kappa

In clinical trial and medical research, we often have a situation where two different measures/assessments are performed on the same sample, same patient, same image,… the agreement needs to be calculated as a summary statistics. Depending on whether or not the measurement is continuous or categorical, the agreement statistics could be different. Lin L had a very nice overview for agreement statistics.

Specifically for categorical assessment, there are many examples where the agreement statistics is needed. In a clinical trial with imaging assessment, the same image (for example, CT Scan, arteriogram,…) can be read by different readers. For disease diagnosis, a new diagnostic tool (with advantage of less invasive or easier to implement) could be compared to an established diagnostic tool… Typically, the outcome measure is dichotomous (e.g., disease vs no disease, positive vs. negative…).

The choice of the methods of comparison is influenced by the existence and/or practical applicability of a reference standard (golden standard). If a reference standard (golden standard) is available, we can estimate sensitivity and specificity – ROC (receiver operation characteristics) analysis. If a reference standard is not available or there is no golden standard for comparison, we can not perform ROC analysis. Instead, we can assess the agreement and calculate the Kappa. This has been discussed in detail in FDA’s Guidance for Industry and FDA Staff “Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests”. For example, for comparing the assessment from two different readers, we would calculate Kappa, overall percent agreement, positive percent agreement, and negative percent agreement. We would not use ROC statistics and would not calculate the sensitivity and specificity.
If we would like to assess the agreement between the urine pregnancy test and the serum pregnancy test, we could use the ROC and calculate the sensitivity, specificity, positive predictive value, and negative predictive value since the serum pregnancy test could be considered as a reference standard or golden standard for pregnancy test.

Kappa Statistic(K) is a measure of agreement between two sources, which is measured on a binary scale (i.e. condition present/absent). K statistic can take values between 0 and 1.
  • Poor agreement : K < 0.20
  • Fair agreement : K = 0.20 to 0.39
  • Moderate agreement : K = 0.40 to 0.59
  • Good agreement : K = 0.60 to 0.79
  • Very good agreement : K =0.80 to 1.00
A good review article about Kappa Statistics is the one written by Karemer et al “Kappa Statistics in Medical Research”.

SAS procedures can calculate Kappa Statistics easily. Here is a list of papers: