Saturday, June 18, 2011

Is blinded study really blinded? - assessment of blinding / unblinding in clinical trials


Randomization and blinding are critical components of the clinical trial from the start (design) to the end. Randomized, controlled, and double-blinded trial (RCT) has been the ideal clinical trial design. Inappropriate randomization and blinding (or potential unblinding) affect the integrity for the clinical trial. If the patient or investigator is aware of the treatment assignment, there will be conscious or unconscious biases in assessing efficacy, safety, or patient-reported outcome. With available software and computer programs, generating a randomization schedule is relatively easy. Ten years ago, I wrote a paper on “Generating randomization schedule using SAS programming” to show how easily randomization can be generated. With the interactive response technologies (IRT) including interactive voice response system (IVRS) and interactive web response system (IWRS), implementation of the randomization can also be easily managed. However, maintaining the blinding during the study may not be as easy as we thought.

I still remember the time when I was one of the randomization team members in PPD. After we generated the randomization schedule, we had to put the randomization schedule into an envelope and sealed with signatures. Then we had to put the envelope into a locked security box in a secured randomization room. In order to get the randomization schedule, at least two statisticians had to be present in order to open the security box.

While the actual randomization schedule is locked and secured, the randomization information or treatment assignment concealment can still be compromised by what happened at the site, how the patient and investigator guess the treatment assignment, and how the unblinded personnel communicate with the blinded team members.

There are many factors that can cause the potential unblinding. Here are some examples:

  • Guess treatment assignment by the experience of adverse events and side effects. Suppose an intravenously administered drug can cause more headaches than Placebo, a patient with headache may guess he/she is on treatment group and not on Placebo. While this guess may not be 100% accurate, majority of patients may guess their treatment assignment correctly. In a book by Chow et al, ‘Design and analysis of clinical trials: concepts and methodologies’, an example about challenge in maintaining blinding was described “beta-blocker (e.g., pro-pranolol) have specific pharmacologic effects such as lowering blood pressure and the heart rate and distinct adverse effects such as fatigue, nightmares, and depression. Since blood pressure and heart rate are vital signs routinely evaluated at every visit in clinical trials, if a drug such as propranolol is known to lower blood pressure and the heart rate, then preservation of blindness is a huge challenge and seems almost impossible” In a large scale study (BHAT study), at the conclusion of the trial, patients, investigators, and clinical coordinators were asked to guess the patient’s treatment assignment, 79.9%, 69.6%, and 67% of patients, investigators, and clinical coordinators respectively guessed correctly the patient was on Propranolol and 42.8%, 58.6, and 70.6% of patients, investigators, and clinic coordinators respectively guessed correctly that the patient was on Placebo.
  • Guess treatment assignment by improvement or no improvement in efficacy. If there is a prior knowledge that an treatment is effective (lack of equipoise), the investigator or patient can guess which treatment the patient is on based on the lack of effect.
  • Guess treatment assignment by knowing the blood concentration of the drug or analytes. If a treatment is for augmentation purpose, a patient could have his/her blood sample tested to know whether or not the concentration for augmented drug is increased or not, then guess which treatment group he/she is on.
  • In double-blinded studies, there are always some unblinded groups. These groups could include global drug safety for safety monitoring, laboratories that measure drug concentration or biomarkers, study drug supplies, site unblinded pharmacist… all of these groups could potentially reveal the treatment assignment to other study team unintentionally.
  • For a study with DMC that involved a third party to prepare the unblinded data for DMC, treatment concealment could potentially be compromised during the information exchange with the blinded study team. This is critical for studies with adaptive designs where the patient data needs to be constantly reviewed and analyzed. An interesting example was discussed by Janet Witts regarding an awkward situation in an adaptive design where the DMC knew the event rate by treatment assignment and the sponsor didn’t. 

There is a dilemma when we develop the informed consent form. On the one hand, we are required to put into the informed consent form as much information as we can. On the other hand, the more information we put into the informed consent form, the more likely we enable the patients to guess their treatment assignment (based on their experience of side effects or perceived efficacy).

Ideally, in a double-blind trial, it is a good practice to evaluate for both the subjects and investigators whether or not blinding / masking has been preserved. However, in the real world, it is rare in double-blinded clinical trials to include a formal assessment of how well the blinding has been preserved. If the assessment of blinding becomes a routine, I think that many studies will show that subjects/investigators guessed correctly more frequently than they should have done by chance alone. Part of the reason this assessment has not been done often is perhaps the difficulty to explain the study results if the blinding is found to be compromised. It will be extremely difficult to assess the magnitude of the impact on the safety and efficacy evaluation if the blinding/treatment assignment concealment is compromised.

Further readings:

Wednesday, June 15, 2011

Bland-Altman Plot for Assessing Agreement


Bland-Altman plot is a scatter plot of variable means plotted on the horizontal axis and the differences plotted on the vertical axis which shows the amount of disagreement between the two measures (via the differences) and lets you see how this disagreement relates to the magnitude of the measurements.

When I was in graduate school, the statistical analysis of microarray data just started to be a hot topic. In collaboration with Dr Rick Song, we looked at the microarray data and wrote a manuscript titled “On Graphical Presentation and Quantitative Analysis of cDNA Microarray Data” and we presented in JSM. In this manuscript, we proposed to use Bland-Altman plot. In clinical trials, I have not got a chance to apply this approach, but I do often see articles using the Bland-Altman plot. For example, an article titled “Using the Bland–Altman method to measure agreement with repeated measures” from British Journal of Anaesthesia.

When data is appropriate, Bland-Altman plot can be a handy tool to use. It is worth relaying the paragraphs from our original paper on graphical presentation of micro-array data using Bland-Altman plot.

“Graphical presentation is usually the first step for data analysis of microarray data. In the case without duplication (this is typical in microarray experiment), scatter plots will be drawn and then a regression line drawn through the data. This helps the eye in gauging the degree of agreement between two measurements and also may help us to identify the "outliers" that represent the differentially expressed genes in microarray experiment.

In clinical medicine, to assess agreement between two methods of clinical measurement, Bland and Altman proposed to plot the difference between the methods (A-B) against the mean (A+B)/2[12,13,14,15]. This approach has been extensively used in medical research for assessing measurement error and comparing different measurements for the same quantity. Bland and Altman’s method can be also applied to the microarray data. We can plot (Rm-Gm) against (Rm+Gm)/2 (figure2 above).

Calculating or plotting a regression line is not our focus as we are not concerned with the estimated prediction of one color intensity by another but with the theoretical relationship of equality and deviations from it.

There are several advantages for presenting the microarray data using Brand and Altman’s approach:

The plot of difference against mean allows us to investigate any possible relationship between the discrepancies and the true value. The plot will also show clearly any extreme or outlying observations. If two different samples are used in the experiment, these extreme or outlying observations could indicate the differentially expressed genes. It is often helpful to use the same scale for both axes when plotting differences against mean values. This feature helps to show the discrepancies in relation to the size of the measurement.

Brand and Altman's method makes it easier for us to estimate the precision of the estimated limits of agreement between two color intensities. We want a measure of the agreement that is easy to estimate and to interpret for a measurement on the color intensity of an individual gene. An obvious starting point is the difference between measurements by the two channels on the same gene. There may be a consistent tendency for one channel to exceed the other. This is called calibration factor and can be estimated by the mean difference. There will also be variation about this mean, which we can estimate by the standard deviation of the differences. These estimates are meaningful only if we can assume that calibration factor and variability are uniform throughout all genes.”

More references on Bland-Altman Plot:

Friday, June 03, 2011

Restructure FDA's Drug Review Process?


Last week, I had a chance to listen to a speech by Dr Scott Gottlieb. While he touched several topics in related to the health care reforms, I was specifically interested in his discussion on restructuring FDA’s drug approval process.

Dr Gottlieb gave a lot of insights within FDA and analyzed the root cause of the very long and inefficient FDA drug review process.

Following his speech, I located his paper “SHOULD FDA RESTRUCTURE ITS DRUG REVIEW PROCESS?” from FDLI’s website. A lot of his analyses are so true and to the point.

For example, he elaborated why FDA adopt a matrix management structure for its review program.

“Prior to FDA’s adoption of a matrix management structure for its drug review program, agency scientists were organized largely around the clinical areas in which they worked (oncology, cardio-renal, antiviral, etc). This therapeutically focused structure had some advantages, but also led to some of its own challenges. FDA’s adoption of a matrix structure was aimed at solving some of these problems.

For one thing, sponsors complained that the advice they received about disciplines like biostatistics or clinical pharmacology varied (sometimes significantly) across different therapeutic divisions. Statisticians in one clinical division would be interpreting certain principles of statistics or evaluating a particular protocol design in a manner different than statisticians inside another therapeutic division.

These discrepancies still occur. But it is believed that the matrix organizational structure cuts down on this sort of conflict.

Grouping all of the statisticians or pharmacologists inside the same office increases opportunities for comparable training and cross-calibration on key principles. CDER management pulled the first group of the review divisions—the chemists—in 1995. The impetus was differences in pharmaceutical quality requirements being maintained among the different clinical divisions. Ultimately, having the chemists organized as a single group fostered the development of consistent standards. It also enabled FDA to negotiate the standards established by the International Conference on Harmonization (ICH).

Another reason for establishing the matrix was to improve morale. FDA remains a very “physician centered” culture, but was much more so prior to adoption of the matrix. Staff who lacked medical degrees or who weren’t the clinical reviewers on an application sometimes complained that they felt marginalized in the review process. As one statistician told me, “we were treated like second-class citizens.” Specialists from non-clinical disciplines like statistics also complained that remaining immersed in a single therapeutic area didn’t give them the breadth of experience that they needed for their own professional development.

Similarly, the organization of scientific personnel by therapeutic area was also seen as an impediment to their continued training in their chosen disciplines. For example, statisticians came together for the equivalent of grand rounds or other kinds of shared learning experiences. But these kinds of cross-training opportunities were challenging because staff were ultimately accountable to their divisions. The shared training opportunities weren’t prioritized. Efforts were also made to rotate non-clinical experts across different therapeutic areas. But the challenges endured.”

He then went on discussing the issues with FDA’s weak matrix management structure.

“But in practical terms, the weak matrix means that FDA project managers have limited dominion over key aspects of the review. Key disciplines involved in the review aren’t accountable to the project manager, or the division director. It is a system where there are few management carrots and no sticks. This weak structure also makes it harder to organize collaborative projects or even team meetings. The project manager doesn’t have strong authority when it comes to managing the collaboration between the different scientists involved in a drug’s review.”

He also mentioned the quality of the FDA review scientist and issues with FDA’s policy to allow very flexible working hours and work-from-home schedule. The current FDA drug review team is loosely organized and inefficient in many aspects. However, it is not easy to make big changes to the current process.

Monday, May 30, 2011

Differences between SOPs and WPs


Working in industry setting (versus academic), we all understand that it is critical for us to be familiar with the Standard Operating Procedures (SOPs) / Working Procedures (WPs) and follow these procedures. This is even more obvious in pharmaceutical and drug development industry. An employee could be fired not because of incompetence, but because of violating these procedures. A couple of days ago, I attended ChineseAssociation for Science and Technology Annual Convention, one session was to discuss about the requirements for starting up a new company. Several local entrepreneurs elaborated the intellectual property (IP), business plan, funding, a loyalty team, and so on. Nobody mentioned these procedures. But I think that it is equally important to have a set of necessary SOPs and WPs in order to start a new company, especially a service-type company in highly regulated drug development field such as contract research organization. Scientists working in the academic setting often find the difficulties when they try to land a position in industry mainly because they lack the ‘industry experience’ – one critical part is that they lack the experience in following up the strict industry SOPs. If you interview a candidate who does not even understand the term ‘SOP’, you know that your candidate will have a long way to adapt to the industry setting. 

While there are definitions for Standard Operating Procedures (SOPs), it is hard to define Working Procedures (WPs), not to mention the differences between SOP and WP.

Standard Operating Procedure is established procedure to be followed in carrying out a given operation or in a given situation. In clinical trial setting, the SOP is defined in ICH E6 as “detailed, written instructions to achieve uniformity of the performance of a specific function. “  SOP is one of the components of a Quality Management System (QMS) implemented in many large organizations - improve quality by providing a standard format, approach, and application of common processes.. 

The issue may be with the wording ‘detailed’ in this definition. The common understanding within the industry is that SOP needs to be written in a little bit high level and the working procedure provides the detailed steps according to SOP. If the strict definition of SOP from ICH E6 is followed, there perhaps should not be any working procedures. All working procedures should be called SOPs. There is really no defined boundary for separate the SOPs and WPs. For SOPs, say what you do... but do what you say. The documentation is always important to the SOP compliance. 

My personal opinion is that SOPs are the procedures we must follow, typically on high level, with GCP compliance while WPs are procedures we should follow, typically in more detail, for internal standardization and efficiency. For example, in biostatistics, a SOP must be in place to require the independent validation for statistical outputs. A WP may be created to detail how independent validation should be implemented.

Here are some observations about SOPs and WPs:

1) Within FDA, the term SOPP (standard operating procedure and policies) is used for SOPs and WPs.

 2) There are two papers containing good discussions about SOP:
 3) There seems to be no formal definition ‘working procedure’. Different companies may call this differently and different people may have different understanding for ‘working procedure’.
There are some discussions on the web: 

Having SOPs and WPs in place is one thing; having everybody trained on SOPs and WPs is another thing. If SOPs and WPs are written in great detail like step-by-step instructions, the procedures may become a burden to be followed. This could in turn cause issues in SOP compliance when there is an audit. 

Last week, we had a group meeting during the lunch time and we ordered some food from Jason’s Deli. It was good to see that the delivery guy brought a Bagset Checklist which listed the items and quantities for each item. Unfortunately, the guy clearly did not follow their procedure and did not check the check list. When he delivered, he realized that he forgot to bring ‘bulk chips’. He had to go back to pick up this missed item. Had he follow the procedure and check his Bagset Checklist, he would have not missed any item. In this case, the consequence of the non-compliance is minor, but in clinical trials, the consequence of the non-compliance could sometimes be very significant.

Saturday, May 14, 2011

I-Spy 2 Trial - a study with fancy name and fancy design

If you see the term “I-spy”, you would typically think something related to movie, kids game or TV series. Recently, I-spy has been linked to a novice clinical trial design in breast cancer.

For a short description about this study, clinicaltrials.gov is a good place. In clinicaltrials.gov, I-SPY 2 TRIAL is listed as “Neoadjuvant and Personalized Adaptive Novel Agents to Treat Breast Cancer”. For more detail descriptions about this study, you should check out the following websites:
This study got a lot of publicities. For example, in FDA’s ‘advancing Regulatory Science for Public Health’, it specifically mentioned the I-Spy 2 trial,

Personalized treatment for cancer

The “I-SPY 2 TRIAL,” launched in March 2010, represents a groundbreaking new clinical trial model that will help scientists quickly and efficiently test the most promising drugs in development for women with higher risk, rapidly growing breast cancers. During the trial, drugs in development are individually targeted to the biology of each woman’s tumor using specific genetic or biologi­cal markers, known as “biomarkers.” By applying an innovative trial design, researchers will use data from one set of patients’ treatments to treat other patients – more quickly eliminating inef­fective treatments and drugs. The I-SPY 2 trial was developed under the Biomarkers Consortium, a unique public-private partnership that includes the FDA, the National Institutes of Health, and major pharmaceutical companies, led by the Foundation for the National Institutes of Health
. “

The claimed advantages for this trial design are:
  • Utilized Personalized Medicine
  • Uses genetic or biological marker (“biomarkers”) from individual patients’ tumors to screen promising new treatments, identifying which treatments are most effective in specific types of patients
  • Improved Efficiency (fewer patients, less time, and fewer resources)
Enable researchers to use early data from one set of patients to guide decisions about which treatments might be more useful for patients later in the trial, and eliminate ineffective treatments more quickly Enable the development of more informed, smaller phase III trials

So what is this trial design? Can this trial really achieve its purpose? Here is what I would say:

  • I-Spy 2 is a phase II study, anything generated from this study will need to be confirmed in Phase III confirmatory studies.
  • I-Spy 2 is a government-sponsored study and supported by Foundation of NIH, Biomarkers Consortium, NCI, and others. It may never happen if it is an industry-sponsored study.
  • I-Spy 2 is academic study and not a study for product licensure. If it is study for product licensure, there may be difficulties for study design to be accepted by regulatory agencies.

  • I-Spy 2 design is complicated and it employs Bayesian Adaptive, Response Adaptive Randomization (play-the-winner approach), Biomarker Adaptive, Stopping and adding treatment arms. The algorithm and stopping rules are complicated and the assumptions / basis for generating these algorithms / stopping rules may eventually be demonstrated incorrect. 
  • I-Spy 2 is a randomized, open label study. It would be more difficult to implement if this is a double-blinded study
  • I-Spy 2 got it fame due to its use of ‘personalized medicine’, ‘adaptive design’ – all fit into the hot areas
  • I-Spy 2 is a multi-center, NOT multi-national study. It would be much more challenging if this is a multi-national study
Although I-Spy 2 got a lot of attentions, not everyone is convinced with this design. The study is closed watched – let’s see the fate once it is completed.

Friday, April 29, 2011

Mobile phone text message used in clinical trial?

Recent discussions with my friends in China surprised me a little bit. They are way ahead in applying the advanced technology in public health area. They used the mobile phone text message to promote the smoking cessation. They recently applied research grant to study the effect of mobile phone text message in improving the maternal and child care. For pregnant women or new mothers, the mobile phone text message is used to send reminders, maternal care tips, child care tips,...

There have been some publications demonstrating the effectiveness of text message in these public health promotion areas. For example, a study in New Zealand demonstrated that smoking cessation using mobile phone text messaging is as effective in Maori as non-Maor. The newscientist.com reported that text messages double young smokers' quit rates.There are

It is natural to think that such text message approach can be used in clinical trials. When I search the clinicaltrials.gov, I can find many clinical trials using text message mostly for improving the adherence of clinical visits or adherence of drug taking.

In my opinion, effectiveness of using text message in clinical trials depends on the study population and the country the study is conducted. In China, if the study is conducted in urban areas, mobile phone text message could be very effective because 1) almost everyone has mobile phone; 2) people use text message more often than actual calling. In the United States, for general population, text message may not be a good approach because some family may still rely on residence-line phone instead of cell phone. Even though they have cell phone, they rarely use text message on daily basis. If a study is conducted in high school students or college students, text message could be an effective tool since all students like to use text messages.

Saturday, April 16, 2011

Emerging Statistical Issues in the Conduct and Monitoring of Clinical Trials

This Wednesday, I had a chance to attend “University of Pennsylvania Annual Conference on Statistical Issues in Clinical Trials”. The topic for this year is “Emerging Statistical Issues in the Conduct and Monitoring of Clinical Trials”. The number of participants was just right in size and the conference was organized pretty well.

In terms of the topics, there are some of them I like and some of them I don’t like. The presentation slides will eventually be posted on conference’s website, however, I would like to give one or two sentences commenting on each topic.

“Sample size estimation incorporating disease progression” – the key issue is the adequacy of the study endpoint. A good endpoint will incorporate the impact of the disease progression.

“Hurdles and future work in adaptive designs” – it is good to hear the discussion about the hurdles, caveats of the adaptive designs. Still very often, a lot of people only talk about the advantages of adaptive designs – too good to be true.  Similarly, a recent article "a once-rare type of clinical trial that violates one of the sacred tenets of trial design is taking off, but is it worth the risk? " from The-Scientist magazine gave some objective assessments on implementing the adaptive designs.

“Predicting accrual in ongoing trials” – utilizing the complicated statistical model to predict the accrual is a waste of time. Accrual in ongoing clinical trials is 95% clinical operations issues, 5% related to statistics. Is it worth to modeling the accrual?

“New incentive approaches for adherence” – money incentives including lottery is a sensitive topic and ethic issue could follow no matter it is incentive for adherence or for study visit compliance. Money incentives are different depending on participants’ social economic status (family income). $100 lottery may be very incentive to some, but not to others.  

“Efficient source dada verifications in cancer trials” – I always thought that all data fields had to be 100% source data verified. It is not entirely true in large scale trials in oncology or in studies with cardiovascular endpoint. In industry, we are rather conservative.

“Estimation of effect size in trials stopped early” – trials stopped early due to efficacy is not very common and should not be encouraged. Difficulty in estimating the effect size still exists for trials stopped early.

“Accounting in analysis for data errors discovered through sampling” – Unreliable data or large % of missing data is always a concern, even for observational studies. Statistical approach may not be a good option. When data is garbage, the results we draw from the data will also be garbage – so called ‘garbage in, garbage out’ no matter which statistical model is utilized to address the data issues.

“Some practical issues in the evaluation of multiple endpoints” – It is so correct that we should play down the importance of differentiating ‘primary endpoint’, ‘secondary endpoint’, ‘tertiary endpoint’,… Multiple comparison has been expanded so much and is everywhere now (co-primary, primary and secondary, co-secondary, secondary superiority test after non-inferiority test, interim analysis, meta analysis, ISE,…). Are we overdoing this?

Saturday, April 09, 2011

Sparse sample and population pharmacokinetics

In drug development, it is necessary to understand the pharmacokinetics profiles (or time concentration profiles) of the experimental drug and calculate the pharmacokinetic (PK) parameters (Area Under the Curve – AUC, Clearance – CL, or Volume of distribution –Vd). These PK parameters can provide the estimate of the dose exposure and assist in the decision on dose timing and dose interval. In order to calculate the PK parameters, we typically need a serial of blood samples at multiple time points (usually more than 6) after the drug administration. In some situations, it is not feasible or not practical to obtain these many blood samples. The obvious example is in pediatric studies where it is not feasible to obtain multiple blood samples due to the blood volume restriction. The specimen may not just be blood samples. If the PK is conducted using other specimens, it is usually difficult to obtain multiple PK samples. For example, we could obtain middle ear fluid (MEF) sample to determine the antibiotic drug concentration in the ear and bronchoalveolar lavage (BAL) to determine the drug exposure in the lung. It is not practical to obtain multiple samples for these special specimens due to the safety concern.

When very few samples are available for each patient, we call it ‘sparse sampling’. With sparse data, we would need to employ a
Population PK
approach to estimate the PK parameters, describe the PK profile, or do PK/PD modeling. The use of population PK during the drug development has been steadily increasing. Regulatory agencies have issued several guidance on the use of population pharmacokinetics.




There are different sparse sample designs. Below are some of the sparse sample designs I have seen.  

An example of sparse sampling at fixed time points is described in a paper by Vogelmeier et al. They used BAL fluid sample to study the intrapulmonary half-life of aerosolized product in Normal Volunteers”.

For BAL fluid samples, it is not feasible to obtain serial samples at all six time points (at screening, 0.5, 6, 12, 24, and 36 h). Therefore, in this study, “each volunteer underwent two BALs. The first lavage was done in the screening phase with an interval of between 3 and 7 d before inhalation of the drug. The volunteers were randomly assigned to one of five groups with the second lavage following 0.5, 6, 12, 24, or 36 h after aerosol administration. Each of the groups consisted of six individuals…”

Subjects in group 1 contributed two BAL samples at Screening and at 0.5 hours after inhalation.
Subjects in group 2 contributed two BAL samples at Screening and at 6 hours after inhalation.
Subjects in group 3 contributed two BAL samples at Screening and at 12 hours after inhalation.
Subjects in group 4 contributed two BAL samples at Screening and at 24 hours after inhalation.
Subjects in group 5 contributed two BAL samples at Screening and at 36 hours after inhalation.

With subjects from all five groups combined, a overall picture of the PK profiles over 24 hours after inhalation could be described. Original paper provided only the summary analysis. Nowadays, the data could be further analyzed using nonlinear mixed model from population PK model with software such as NONMEM.

In FDA guidance on Population Pharmacokinetics, an example was provided for estimating the AUC using sparse data (1-2 middle ear fluid samples per subject) in pediatric subjects.

“The penetration of drug X into middle ear fluid (MEF) was investigated using population PK analysis with sparse data (1-2 samples per subject) obtained from 36 pediatric patients (2 months to 2.0 years of age) who underwent clinical therapy with drug X. The estimated area under the concentration-time curve (AUC) that was above the minimum inhibitory concentration (MIC) (AUCMIC) and the half-life of drug X are 12.5 ug.hr/ml and 6.1 hours in MEF, respectively, vs. 23.7 ug.hr/ml and 3.2 hours in plasma, respectively….”

With this short description, we don’t know if MEF samples are taken from subjects at various times or fixed times. However, non-linear mixed model must have been used for analyzing the data.  

FDA’s guidance on population pharmacokinetics states, “the full population PK sampling design is sometimes called experimental population pharmacokinetic design or full pharmacokinetic screen. When using this design, blood samples should be drawn from subjects at various times (typically 1 to 6 time points) following drug administration. The objective is to obtain, where feasible, multiple drug levels per patient at different times to describe the population PK profile. This approach permits an estimation of pharmacokinetic parameters of the drug in the study population and an explanation of variability using the nonlinear mixed-effects modeling approach. “

If a full population PK sampling design is used, the sampling scheme will be something like below. The different subject could contribute different number of samples at various times.  


Subject number
Blood sampling time (t)
concentration at time t
 Ct
001
Predose
xxx
001
24 hours post dose
xxx
002
Predose
xxx
002
8 hours post dose
xxx
002
12 hour post dose
xxx
003
Immediately postdose
xxx
003
5 hour post dose
xxx
004
4 hour post dose
xxx




Then when non-linear mixed model such as NONMEM is used to fit the data to characterize the PK profile with PK parameter (such as AUC) = function of concentration (Ct) at time t.

In multiple dose studies, if the purpose is to characterize the PK profile at steady state, one could implement a strategy of splitting the number of samples into different dose intervals.
Suppose we need 8 serial blood samples (t1 to t8) to calculate AUC and the dose interval is weekly, we can have these 8 samples split into 4 dose cycles. For each subject, we would only take two samples for each dose cycle. At steady state, for each subject, we expect PK profile after each repeat dose is not much different; the concentration at day 1 after repeat dose #1 would be similar to the concentration at day 1 after repeat dose #4, and so on. In this case, we would be able to calculate AUC for each subject with 8 samples from four dose intervals (instead of 8 samples from one dose interval over 7 days). The drawback is that the study period would be longer.

Friday, March 11, 2011

Use of SF-36 in Clinical Trials

The SF-36 is a multi-purpose, short-form health survey with 36 questions. SF-36 is one of the most popular instruments for generic health surveys and it can be used across age, disease, and treatment group, and are appropriate for a wide variety of applications. Conversely to generic health surveys, disease specific health surveys are focused on a particular condition or disease. In clinical trials, SF-36 remains as one of the most common instruments for assessing the Health Related Quality of Life (HQOL), especially in diseases where there is no valid disease-specific tool.
 
SF-36 yields an 8-scale profile of functional health and well-being scores (so called domain scores) as well as psychometrically-based physical and mental health summary measures [physical component summary (PCS) and mental component summary (MCS)] and a preference-based health utility index (question #2).
 
The mapping from the original questions -> 8 domains -> PCS or MCS is sketched in the diagram below. Notice that only 35 out of 36 questions are used in this diagram. The question 2 asks about the general health status and does not contribute to the calculation of domain scores and component summaries. A good use of question 2 is to use its responses as anchor in identifying the minimal clinically important difference (MCID). In one of our publications in J Neurol Neurosurg Psychiatry, we indeed used this approach to identify the MCID.
 
For these 36 questions, the response categories vary depending on the question. The response categories range from 2 (yes, no) to 6 (all of the time, most of the time, a good bit of the time, some of the time, a little of the time, none of the time). Therefore, in order to calculate the domain score, a scoring method or algorithm has to be employed. For PCS and MCS, the calculation will be based on equations with coefficients from the regression models generated from the General Healthy Populatoin. In US, it is the Healthy General US Population. If different healthy population is used, the factor score coefficients for the Z_scores will be different and PCS and MCS values will be different.
 
The details about scoring method can be found at QualityMetric’s website. The scoring and calculation of component summaries require the programming. Some of the example programs (but not validated) can be found from the web:
Some questions and answers on using SF-36 in clinical trials:
 
Q: Is SF-36 free for using in clinical trials?
A: It is not free. License has to be obtained for using in industry-sponsored clinical trials. See qualitymetric website for detail.
 
Q: Why do we have question #2 that is used in calculation of any domain score and component summary?
A: It can be used as an assessment of general health status and also as an anchor for identifying MCID.
 
Q: Which general health population should be used for norm-based scoring?
A: The advantage of norm-based scoring is to facilitate the comparisons. If a study is a US domestic study, General Healthy US Population should be used. If it is an international study, the country-specific General Healthy Populations are preferred. SF-36 has been validated in many languages.
 
Q: What will be language to describe the statistical analysis plan for SF-36
A: For study protocol or for journal article statistical method section, analysis plan for SF-36 should be kept simple. In one of our publications on SF-36, we simply said:
"The corresponding physical component summary and mental component summary values for the randomized participants were calculated using the reported means, SDs, and factor score coefficients that came from the healthy general US population in 1990. A linear T-score transformation method was used so that both the physical component summary and the mental component summary scores were standardized with a range of 0 (lowest) to 100 (highest)"
 
Q: Could SF-36 be used in cost utility analysis?
A: No. SF-36 is not a utility score. However, Sf-36 can be converted to utility score (such as EQ-5D). See my previous blog
 
Q: Could we have one overall score for SF-36?
A: No. PCS and MCS have to be analyzed separately. You can not add PCS and MCS to have a single overall score.
 
Q: How to analyze the domain scores and component summaries?
A: Typically, 8 domain scores and 2 component summaries can be analyzed separately using analysis of variance or analysis of covariance or other methods such as repeat measurement depending on the study design.
A good approach in analyzing the SF-36 is to compare the each domain score with the General Healthy Population to show how much difference between the patients in the study and the General Healthy Population for pre-treatment and for end treatment visits. This approach was utilized in our SF-36 publication in Neurology.

Thursday, March 03, 2011

Incidence Rate (IR) – How could this be wrongly calculated?

I am very surprised to see how a simple concept of ‘incidence rate’ can be wrongly calculated in documents  submitted to regulatory agencies (such as FDA). In a briefing document titled “Tiotropium (SPIRIVA): Pulmonary Allergy Drug Advisory Meeting – November 2009” submitted by a sponsor, there were wrong statements every where about the calculation of the incidence rate for safety variables.

For example, on page 50, it says “Incidence rates of adverse events were computed as the number of patients experiencing an event divided by the person-years at risk”; In Section 8.1.5 (Statistical methods), it says “For each event, an incidence rate (IR) was calculated from the number of patients with an event divided by the cumulative time at risk within a treatment group and expressed as patient-years.”  In their summary tables, they footnoted “the number of patients with an event” (instead of the number of total events) was used in calculating the incidence rate. They never listed the total number of patient year (the denominator) for their Incidence rate calculation. In ‘Statistical method’ section, they even tried to justify the use of “the difference in incidence rate” because “most Tiotropium trials have significantly greater number of patients in the placebo group discontinuing the trial early compared to tiotropium treated patients.”
“Incidence Rate” is a basic concept from epidemiology studies and is calculated as the number of events divided by the number of patient years. According to free medical dictionary, “incidence rate is the probability of developing a particular disease during a given period of time; the numerator is the number of new cases during the specified time period and the denominator is the population at risk during the period. “   According to Wikipedia, “The incidence rate is the number of new cases per population in a given time period. When the denominator is the sum of the person-time of the at risk population, it is also known as the incidence density rate or person-time incidence rate. In the same example as above, the incidence rate is 14 cases per 1000 person-years, because the incidence proportion (28 per 1,000) is divided by the number of years (two). Using person-time rather than just time handles situations where the amount of observation time differs between people, or when the population at risk varies with time. Use of this measure implicitly implies the assumption that the incidence rate is constant over different periods of time, such that for an incidence rate of 14 per 1000 persons-years, 14 cases would be expected for 1000 persons observed for 1 year or 50 persons observed for 20 years.”
In an article by Marco et al “Incidence of Chronic Obstructive Pulmonary Disease in a Cohort of Young Adults According to the Presence of Chronic Cough and Phlegm”, the incidence rate is correctly defined for calculation.
“Incidence rates of COPD were estimated as the ratio of the number of new cases and the number of person-years at risk (per 1,000), which were considered equal to the length of the follow-up for each member of the cohort.”

The key is that if you calculate the ‘incidence rate’, your numerator must be ‘number of events’, not ‘number of patients with an event’. For events that can only occur once in a lifetime for a specific patient (such as cancer), there may not be much difference between ‘number of events” and “number of patients with an event”. However, for events occurr more than one time for a specific patient, “number of events” and “number of patients with an event” are very different concepts.

In Tiotropium briefing document, the correct calculation for incidence rate should be ‘number of events (AEs or COPDs)’ divided by ‘the patient year’. It was simply wrong when they used ‘number of patients with an event’ as the numerator in their calculation of incidence rate. Their justification for using the difference in incidence rate is just the opposite of their statement. If placebo group has more dropouts, their way of calculating the incidence rate will overestimate the rate for placebo group and underestimate the rate for Tiotropium group. This can be easily illustrated using an example below:


Assuming 10 patients in Tiotropium and 10 subjects in Placebo group, 5 patients in Tiotropium group and 5 patients in Placebo group had at least one COPD during the study. The incidence of COPD will be 5/10 = 50% in both groups. Suppose it is a one-year trial, all patients in Tiotropium group completed the one-year and all patients in Placebo group completed only 6 months. The patient year will be 10X1 = 10 for Tiotropium group and 10x0.5 = 5 for Placebo group. The incidence rates now become 5/10 = 50% in Tiotropium group and 5/5 = 100% in Placebo group – this is just simply wrong. In this case, when the patient year (or person year) is used as denominator, the numerator used in the calculation should be the number of events, not the number of patients with an event.    

It is unfortunate this simple concept of ‘incidence rate’ has been wrongly calculated in Tiotropium studies. This wrong calculation may have been embedded in their paper published in prestigious New England Journal of Medicine.

If ‘number of patients with an event’ is used in the numerator, the denominator has to be the total number of patients (not the number of patient year). ‘Number of patients with an event’ divided by ‘number of total patients’ is called ‘incidence of events’ – this is a typical way when we summarize the adverse events in clinical trials.  

Friday, February 25, 2011

Study Center Pooling Strategy in Multicenter Clinical Trials

Pooling the study center for statistical analysis purpose is rather an old issue. However, we can still see the discussion o f study center pooling strategy or algorithm in the study protocol or the statistical analysis for multi-center clinical trials. When a clinical trial has multiple centers, study center or investigator site is usually included in the statistical analysis either by including as an exploratory variable in the model (for example ANOVA or ANCOVA) or by conducting the categorical analysis adjusted by study center (for example, Mantel-Haenszel test, Elteren's test, Wilcoxon rank sum test stratified by pooled center). However, there could be situation that some study centers have very few subjects and can not be directly included as a stand alone center for the analysis. In this situation, a pooling strategy is often employed to combine the small centers together. The reason for pooling the small centers instead of using center as random effect may be due to the factor that centers in the clinical trial are rarely a random sample of all possible centers. It is not uncommon to find the statistical analysis including pooled center in regulatory submission or in publications, for example, in NDA for Refludan (the analysis was stratified by pooled center) and in FDA advisory committee documents (… were analyzed using Wilcoxon rank sum test stratified by pooled center (centers that entered fewer subjects than a complete block were pooled by country)). Here are some of the example languages describing such pooling strategies:

“Statistical tests will be performed as two-sided tests and will be adjusted to the multi-centric design of the study. A center must have enrolled at least 8 subjects to be a standalone center in the analysis (centers enrolling less than 8 subjects will be pooled – will be done before the study unblinding”

“Study centers were pooled from largest to smallest until the pooled center had more than 5 subjects with post baseline data in each treatment group. No pooled center had more than 15% of the total number of subjects”

“The majority of study centers were small. A small center was defined as any center with <5 patients with postbaseline data in any treatment group, resulting in 5 large and 25 small centers. To avoid loss of information, small centers were pooled from largest to smallest until the pooled center had 5 patients in each treatment group. These centers were grouped into 11 pooled centers for the purpose of analysis."

In one of hypertension clinical trials, the pooling strategy is described as “To avoid loss of information, small centers (<5 per protocol patients) were pooled from largest to smallest until the pooled center had 5 per protocol patients in each treatment group. These centers were grouped into 19 pooled centers for the purpose of analysis. The pooling algorithm was predetermined before unblinding the data, and the pooling algorithm was described in the statistical analysis plan for the study. Considering the subjective nature of the pooling algorithm, albeit prespecified before completion of the study, an exploratory analysis was also performed with actual center as a fixed effect in contrast to pooled centers. This analysis did not change the inference.”

In a type 2 diabetes trial, a different pooling strategy was used “For all center stratified analyses, centers with <24 randomized and treated subjects were pooled on a geographical basis, independently of treatment identification.”

In a recent brief book for PDAC, the sponsor provided the detail pooling strategy for centers “Pooling algorithm for centers: For non-US sites, all investigative sites within a country with fewer than 10 randomized subjects will be combined into a single pooled site for analysis purposes. If a resulting pooled site still has fewer than 10 randomized subjects, then this pooled site will be further combined with the smallest unpooled site within that country. If there is not another unpooled site within that country, then the pooled site will be combined with the smallest pooled site from another country. This pooling process will continue until there are at least 10 randomized subjects in each pooled site. For US sites, all investigative sites within a geographic region with fewer than 10 randomized subjects will be combined into a single pooled site for analysis purposes. If a resulting pooled site still has fewer than 10 randomized subjects, then this pooled site will be further combined with the smallest unpooled site within that region. If there is not another unpooled site within that region, then the pooled site will be combined with the smallest pooled site from another region within the US. This pooling process will continue until there are at least 10 randomized subjects in each pooled site.”

As we can see from the examples above, the cut point for center pooling (5, 8, 10, or 24) is really arbitrary and there is no scientific basis for choosing one or another. The decision on the cut point may be based on the distribution of the number of subjects across centers.

Center pooling strategy could sometimes be questioned by the regulatory reviewers. For example, in BLA review of Rebif, FDA reviewer had concerns about the pooling strategy “The sponsor’s study center pooling strategy: Per the pre-specified strategy in the sponsor’s statistical analysis plan (SAP), pooling of study centers for inclusion of center as a main effect in analyses was to have been based on geographic considerations for small centers. In fact, the pooling strategy actually used was data driven which is problematic. NOTE: There were 56 participating centers from 9 countries. The smallest recruiting center had 3 subjects, 2 centers contributed 4 subjects, and 5 centers contributed 6 subjects each. The remaining centers contributed between 6 – 24 subjects each (CSR, Table 3, pp. 65-66). This reviewer performed analyses of major efficacy endpoints based on strict geographic pooling of centers into 3 groups (US, Canada, and Europe) as well as un-pooled analyses (not including the center effect). In addition, descriptive analyses for individual centers were also performed for the primary and major secondary efficacy endpoints. The sponsor’s positive statistical findings were found to be robust based on these analyses.”

In Biopharmaceutical Report (Summer 1998), Paul Gallo wrote an article titled “Practical Issues in Linear Models Analyses in Multicenter Clinical Trials” which contained a section discussing “construction of composite centers”. The caveats of using the composite centers are also discussed in the paper.

“In performing unweighted analyses, a practice of defining artificial “pooled” or “composite” centers is often employed; that is, data from different centers are treated in the analysis as if they came from the same center. A number of small centers may be combined, or one or more small centers may be combined with a larger center. This practice attempts to minimize the large variance inflation and data instability of unweighted analyses when there are very small centers. Composites may be constructed to the extent of eliminating empty cells to ensure that treatment effects are estimable in models containing interaction terms. More commonly, this is done to achieve some minimum cell size felt to appropriately limit the influence of individual observations; values around 5 are often chosen. ”

Arbitrarily pooling the centers sometimes does not make sense at all. This is exactly true when the centers with small number of enrolled subjects are pooled even though these centers are scattered in totally unrelated geographic regions or countries. When pooled center is used and the statistically significant center effect is detected, the interpretation of the results is difficult. Instead of the center pooling purely based on the number of enrollees, the geographic distribution of centers should be considered. In many cases, instead of pooling centers by the number of enrollees, we could use country and geographic region in the analysis. In one of our multi-national clinical trials, we grouped centers by geographic region as North American, South American, Eastern Europe, Western Europe, and Eastern Asia. The strategy worked very well.

If possible, we could use the random effect model to include the study site / center as random effect to avoid the center pooling. We could also use a center weighting strategy that is similar to the Meta analysis where centers with more subjects are given more weights.

Tuesday, February 08, 2011

Guidelines for Blood Volumes in Clinical Trials (Especially in Pediatric Clinical Trials)

Nowadays, the clinical study protocols are becoming more and more complicated and require more and more blood sample draws for various purposes. The blood samples are needed for testing the hematology, chemistry, immunogenicity (for biological products), biomarkers (for diagnostic or other purposes), pharmacogenomics,… In some clinical trials, additional blood samples (sample retains) may be drawn for future studies (even though we may not know what the future study will be). If the study has the component of pharmacokinetics, the many more samples (series blood samples) will be drawn within a short period to characterize the pharmacokinetic profile, estimate the total drug exposure (AUC), and calculate other pharmacokinetic parameters.

With increasing in the number of blood draws or the blood volumes, the ethic issue often arises, especially in clinical trials with children.

US FDA and EMA do not really regulate the maximum blood volume that can be drawn from a subject during the clinical trials. The requirements for limiting the blood sample volume may come from the National Institute of Health (NIH), American Academy of Pediatrics, World Health Organization (WHO), and European Union (EU) and are typically enforced by the ethic bodies such as Institute Review Board (IRB) and Ethics Committee (EC). The requirements on blood volume during the clinical trials may be different depending on the country and local IRB.

The blood volume drawn for pharmacokinetic studies in the pediatric population is specifically a concern and has been discussed extensively. Stephen RC Howie (2010) reviewed blood sample volumes in child health research: a review of safe limits in the Bulletin of the World Health Organization (BLT). WHO also has its guidelines on drawing blood: best practices in phlebotomy. The guidelines are not specifically for clinical trials, rather for general blood donations. The guidelines contain specific technical requirements for the blood drawn in pediatric and neonatal subjects.

In US, Code of Federal Regulations has a specific chapter (Part 46) to discuss protection of human subjects and the chapter contains a subpart D to address additional Protections for Children Involved as Subjects in Research. While there is no specific requirement on the limit of blood volume, the CFR indicated that the research involves no more than minimal risk to the subjects and IRB should take into account the purposes of the research and the setting in which the research will be conducted and should be particularly cognizant of the special problems of research involving vulnerable populations, such as children, prisoners, pregnant women, mentally disabled persons, or economically or educationally disadvantaged persons. Similarly, the American Academy of Pediatrics has its policy on Guidelines on Ethical Conduct of Studies to Evaluate Drugs in Pediatric Populations. The policy requires “…with the growing number of pediatric drug studies, IRBs need to be familiar with the various research-design methods that minimize risk to the child. Examples include limiting research under some circumstances to pharmacokinetic and safety data, combining this approach with pharmacodynamic data, and minimizing the volume of blood withdrawn through the use of sensitive assays, pediatric enabled laboratories, and population pharmacokinetic approaches"

National Institute of Health Clinical Center has a guideline M95-9: Guidelines for Blood Drawn for Research Purposes in the Clinical Center.

Two articles from the web actually reflect the limit of blood volume in the US.
In EU, there are specific guidelines on "ETHICAL CONSIDERATIONS FOR CLINICAL TRIALS ON MEDICINAL PRODUCTS CONDUCTED WITH THE PAEDIATRIC POPULATION"



The guidelines on blood volume are usually based on the amount of blood in the percentage of total blood volume (BLV). BLV varies depending on age and body weight. A good reference for BLV for pediatrics can be found in pediatricareonline.com.