Saturday, April 09, 2011

Sparse sample and population pharmacokinetics

In drug development, it is necessary to understand the pharmacokinetics profiles (or time concentration profiles) of the experimental drug and calculate the pharmacokinetic (PK) parameters (Area Under the Curve – AUC, Clearance – CL, or Volume of distribution –Vd). These PK parameters can provide the estimate of the dose exposure and assist in the decision on dose timing and dose interval. In order to calculate the PK parameters, we typically need a serial of blood samples at multiple time points (usually more than 6) after the drug administration. In some situations, it is not feasible or not practical to obtain these many blood samples. The obvious example is in pediatric studies where it is not feasible to obtain multiple blood samples due to the blood volume restriction. The specimen may not just be blood samples. If the PK is conducted using other specimens, it is usually difficult to obtain multiple PK samples. For example, we could obtain middle ear fluid (MEF) sample to determine the antibiotic drug concentration in the ear and bronchoalveolar lavage (BAL) to determine the drug exposure in the lung. It is not practical to obtain multiple samples for these special specimens due to the safety concern.

When very few samples are available for each patient, we call it ‘sparse sampling’. With sparse data, we would need to employ a
Population PK
approach to estimate the PK parameters, describe the PK profile, or do PK/PD modeling. The use of population PK during the drug development has been steadily increasing. Regulatory agencies have issued several guidance on the use of population pharmacokinetics.




There are different sparse sample designs. Below are some of the sparse sample designs I have seen.  

An example of sparse sampling at fixed time points is described in a paper by Vogelmeier et al. They used BAL fluid sample to study the intrapulmonary half-life of aerosolized product in Normal Volunteers”.

For BAL fluid samples, it is not feasible to obtain serial samples at all six time points (at screening, 0.5, 6, 12, 24, and 36 h). Therefore, in this study, “each volunteer underwent two BALs. The first lavage was done in the screening phase with an interval of between 3 and 7 d before inhalation of the drug. The volunteers were randomly assigned to one of five groups with the second lavage following 0.5, 6, 12, 24, or 36 h after aerosol administration. Each of the groups consisted of six individuals…”

Subjects in group 1 contributed two BAL samples at Screening and at 0.5 hours after inhalation.
Subjects in group 2 contributed two BAL samples at Screening and at 6 hours after inhalation.
Subjects in group 3 contributed two BAL samples at Screening and at 12 hours after inhalation.
Subjects in group 4 contributed two BAL samples at Screening and at 24 hours after inhalation.
Subjects in group 5 contributed two BAL samples at Screening and at 36 hours after inhalation.

With subjects from all five groups combined, a overall picture of the PK profiles over 24 hours after inhalation could be described. Original paper provided only the summary analysis. Nowadays, the data could be further analyzed using nonlinear mixed model from population PK model with software such as NONMEM.

In FDA guidance on Population Pharmacokinetics, an example was provided for estimating the AUC using sparse data (1-2 middle ear fluid samples per subject) in pediatric subjects.

“The penetration of drug X into middle ear fluid (MEF) was investigated using population PK analysis with sparse data (1-2 samples per subject) obtained from 36 pediatric patients (2 months to 2.0 years of age) who underwent clinical therapy with drug X. The estimated area under the concentration-time curve (AUC) that was above the minimum inhibitory concentration (MIC) (AUCMIC) and the half-life of drug X are 12.5 ug.hr/ml and 6.1 hours in MEF, respectively, vs. 23.7 ug.hr/ml and 3.2 hours in plasma, respectively….”

With this short description, we don’t know if MEF samples are taken from subjects at various times or fixed times. However, non-linear mixed model must have been used for analyzing the data.  

FDA’s guidance on population pharmacokinetics states, “the full population PK sampling design is sometimes called experimental population pharmacokinetic design or full pharmacokinetic screen. When using this design, blood samples should be drawn from subjects at various times (typically 1 to 6 time points) following drug administration. The objective is to obtain, where feasible, multiple drug levels per patient at different times to describe the population PK profile. This approach permits an estimation of pharmacokinetic parameters of the drug in the study population and an explanation of variability using the nonlinear mixed-effects modeling approach. “

If a full population PK sampling design is used, the sampling scheme will be something like below. The different subject could contribute different number of samples at various times.  


Subject number
Blood sampling time (t)
concentration at time t
 Ct
001
Predose
xxx
001
24 hours post dose
xxx
002
Predose
xxx
002
8 hours post dose
xxx
002
12 hour post dose
xxx
003
Immediately postdose
xxx
003
5 hour post dose
xxx
004
4 hour post dose
xxx




Then when non-linear mixed model such as NONMEM is used to fit the data to characterize the PK profile with PK parameter (such as AUC) = function of concentration (Ct) at time t.

In multiple dose studies, if the purpose is to characterize the PK profile at steady state, one could implement a strategy of splitting the number of samples into different dose intervals.
Suppose we need 8 serial blood samples (t1 to t8) to calculate AUC and the dose interval is weekly, we can have these 8 samples split into 4 dose cycles. For each subject, we would only take two samples for each dose cycle. At steady state, for each subject, we expect PK profile after each repeat dose is not much different; the concentration at day 1 after repeat dose #1 would be similar to the concentration at day 1 after repeat dose #4, and so on. In this case, we would be able to calculate AUC for each subject with 8 samples from four dose intervals (instead of 8 samples from one dose interval over 7 days). The drawback is that the study period would be longer.

Friday, March 11, 2011

Use of SF-36 in Clinical Trials

The SF-36 is a multi-purpose, short-form health survey with 36 questions. SF-36 is one of the most popular instruments for generic health surveys and it can be used across age, disease, and treatment group, and are appropriate for a wide variety of applications. Conversely to generic health surveys, disease specific health surveys are focused on a particular condition or disease. In clinical trials, SF-36 remains as one of the most common instruments for assessing the Health Related Quality of Life (HQOL), especially in diseases where there is no valid disease-specific tool.
 
SF-36 yields an 8-scale profile of functional health and well-being scores (so called domain scores) as well as psychometrically-based physical and mental health summary measures [physical component summary (PCS) and mental component summary (MCS)] and a preference-based health utility index (question #2).
 
The mapping from the original questions -> 8 domains -> PCS or MCS is sketched in the diagram below. Notice that only 35 out of 36 questions are used in this diagram. The question 2 asks about the general health status and does not contribute to the calculation of domain scores and component summaries. A good use of question 2 is to use its responses as anchor in identifying the minimal clinically important difference (MCID). In one of our publications in J Neurol Neurosurg Psychiatry, we indeed used this approach to identify the MCID.
 
For these 36 questions, the response categories vary depending on the question. The response categories range from 2 (yes, no) to 6 (all of the time, most of the time, a good bit of the time, some of the time, a little of the time, none of the time). Therefore, in order to calculate the domain score, a scoring method or algorithm has to be employed. For PCS and MCS, the calculation will be based on equations with coefficients from the regression models generated from the General Healthy Populatoin. In US, it is the Healthy General US Population. If different healthy population is used, the factor score coefficients for the Z_scores will be different and PCS and MCS values will be different.
 
The details about scoring method can be found at QualityMetric’s website. The scoring and calculation of component summaries require the programming. Some of the example programs (but not validated) can be found from the web:
Some questions and answers on using SF-36 in clinical trials:
 
Q: Is SF-36 free for using in clinical trials?
A: It is not free. License has to be obtained for using in industry-sponsored clinical trials. See qualitymetric website for detail.
 
Q: Why do we have question #2 that is used in calculation of any domain score and component summary?
A: It can be used as an assessment of general health status and also as an anchor for identifying MCID.
 
Q: Which general health population should be used for norm-based scoring?
A: The advantage of norm-based scoring is to facilitate the comparisons. If a study is a US domestic study, General Healthy US Population should be used. If it is an international study, the country-specific General Healthy Populations are preferred. SF-36 has been validated in many languages.
 
Q: What will be language to describe the statistical analysis plan for SF-36
A: For study protocol or for journal article statistical method section, analysis plan for SF-36 should be kept simple. In one of our publications on SF-36, we simply said:
"The corresponding physical component summary and mental component summary values for the randomized participants were calculated using the reported means, SDs, and factor score coefficients that came from the healthy general US population in 1990. A linear T-score transformation method was used so that both the physical component summary and the mental component summary scores were standardized with a range of 0 (lowest) to 100 (highest)"
 
Q: Could SF-36 be used in cost utility analysis?
A: No. SF-36 is not a utility score. However, Sf-36 can be converted to utility score (such as EQ-5D). See my previous blog
 
Q: Could we have one overall score for SF-36?
A: No. PCS and MCS have to be analyzed separately. You can not add PCS and MCS to have a single overall score.
 
Q: How to analyze the domain scores and component summaries?
A: Typically, 8 domain scores and 2 component summaries can be analyzed separately using analysis of variance or analysis of covariance or other methods such as repeat measurement depending on the study design.
A good approach in analyzing the SF-36 is to compare the each domain score with the General Healthy Population to show how much difference between the patients in the study and the General Healthy Population for pre-treatment and for end treatment visits. This approach was utilized in our SF-36 publication in Neurology.

Thursday, March 03, 2011

Incidence Rate (IR) – How could this be wrongly calculated?

I am very surprised to see how a simple concept of ‘incidence rate’ can be wrongly calculated in documents  submitted to regulatory agencies (such as FDA). In a briefing document titled “Tiotropium (SPIRIVA): Pulmonary Allergy Drug Advisory Meeting – November 2009” submitted by a sponsor, there were wrong statements every where about the calculation of the incidence rate for safety variables.

For example, on page 50, it says “Incidence rates of adverse events were computed as the number of patients experiencing an event divided by the person-years at risk”; In Section 8.1.5 (Statistical methods), it says “For each event, an incidence rate (IR) was calculated from the number of patients with an event divided by the cumulative time at risk within a treatment group and expressed as patient-years.”  In their summary tables, they footnoted “the number of patients with an event” (instead of the number of total events) was used in calculating the incidence rate. They never listed the total number of patient year (the denominator) for their Incidence rate calculation. In ‘Statistical method’ section, they even tried to justify the use of “the difference in incidence rate” because “most Tiotropium trials have significantly greater number of patients in the placebo group discontinuing the trial early compared to tiotropium treated patients.”
“Incidence Rate” is a basic concept from epidemiology studies and is calculated as the number of events divided by the number of patient years. According to free medical dictionary, “incidence rate is the probability of developing a particular disease during a given period of time; the numerator is the number of new cases during the specified time period and the denominator is the population at risk during the period. “   According to Wikipedia, “The incidence rate is the number of new cases per population in a given time period. When the denominator is the sum of the person-time of the at risk population, it is also known as the incidence density rate or person-time incidence rate. In the same example as above, the incidence rate is 14 cases per 1000 person-years, because the incidence proportion (28 per 1,000) is divided by the number of years (two). Using person-time rather than just time handles situations where the amount of observation time differs between people, or when the population at risk varies with time. Use of this measure implicitly implies the assumption that the incidence rate is constant over different periods of time, such that for an incidence rate of 14 per 1000 persons-years, 14 cases would be expected for 1000 persons observed for 1 year or 50 persons observed for 20 years.”
In an article by Marco et al “Incidence of Chronic Obstructive Pulmonary Disease in a Cohort of Young Adults According to the Presence of Chronic Cough and Phlegm”, the incidence rate is correctly defined for calculation.
“Incidence rates of COPD were estimated as the ratio of the number of new cases and the number of person-years at risk (per 1,000), which were considered equal to the length of the follow-up for each member of the cohort.”

The key is that if you calculate the ‘incidence rate’, your numerator must be ‘number of events’, not ‘number of patients with an event’. For events that can only occur once in a lifetime for a specific patient (such as cancer), there may not be much difference between ‘number of events” and “number of patients with an event”. However, for events occurr more than one time for a specific patient, “number of events” and “number of patients with an event” are very different concepts.

In Tiotropium briefing document, the correct calculation for incidence rate should be ‘number of events (AEs or COPDs)’ divided by ‘the patient year’. It was simply wrong when they used ‘number of patients with an event’ as the numerator in their calculation of incidence rate. Their justification for using the difference in incidence rate is just the opposite of their statement. If placebo group has more dropouts, their way of calculating the incidence rate will overestimate the rate for placebo group and underestimate the rate for Tiotropium group. This can be easily illustrated using an example below:


Assuming 10 patients in Tiotropium and 10 subjects in Placebo group, 5 patients in Tiotropium group and 5 patients in Placebo group had at least one COPD during the study. The incidence of COPD will be 5/10 = 50% in both groups. Suppose it is a one-year trial, all patients in Tiotropium group completed the one-year and all patients in Placebo group completed only 6 months. The patient year will be 10X1 = 10 for Tiotropium group and 10x0.5 = 5 for Placebo group. The incidence rates now become 5/10 = 50% in Tiotropium group and 5/5 = 100% in Placebo group – this is just simply wrong. In this case, when the patient year (or person year) is used as denominator, the numerator used in the calculation should be the number of events, not the number of patients with an event.    

It is unfortunate this simple concept of ‘incidence rate’ has been wrongly calculated in Tiotropium studies. This wrong calculation may have been embedded in their paper published in prestigious New England Journal of Medicine.

If ‘number of patients with an event’ is used in the numerator, the denominator has to be the total number of patients (not the number of patient year). ‘Number of patients with an event’ divided by ‘number of total patients’ is called ‘incidence of events’ – this is a typical way when we summarize the adverse events in clinical trials.  

Friday, February 25, 2011

Study Center Pooling Strategy in Multicenter Clinical Trials

Pooling the study center for statistical analysis purpose is rather an old issue. However, we can still see the discussion o f study center pooling strategy or algorithm in the study protocol or the statistical analysis for multi-center clinical trials. When a clinical trial has multiple centers, study center or investigator site is usually included in the statistical analysis either by including as an exploratory variable in the model (for example ANOVA or ANCOVA) or by conducting the categorical analysis adjusted by study center (for example, Mantel-Haenszel test, Elteren's test, Wilcoxon rank sum test stratified by pooled center). However, there could be situation that some study centers have very few subjects and can not be directly included as a stand alone center for the analysis. In this situation, a pooling strategy is often employed to combine the small centers together. The reason for pooling the small centers instead of using center as random effect may be due to the factor that centers in the clinical trial are rarely a random sample of all possible centers. It is not uncommon to find the statistical analysis including pooled center in regulatory submission or in publications, for example, in NDA for Refludan (the analysis was stratified by pooled center) and in FDA advisory committee documents (… were analyzed using Wilcoxon rank sum test stratified by pooled center (centers that entered fewer subjects than a complete block were pooled by country)). Here are some of the example languages describing such pooling strategies:

“Statistical tests will be performed as two-sided tests and will be adjusted to the multi-centric design of the study. A center must have enrolled at least 8 subjects to be a standalone center in the analysis (centers enrolling less than 8 subjects will be pooled – will be done before the study unblinding”

“Study centers were pooled from largest to smallest until the pooled center had more than 5 subjects with post baseline data in each treatment group. No pooled center had more than 15% of the total number of subjects”

“The majority of study centers were small. A small center was defined as any center with <5 patients with postbaseline data in any treatment group, resulting in 5 large and 25 small centers. To avoid loss of information, small centers were pooled from largest to smallest until the pooled center had 5 patients in each treatment group. These centers were grouped into 11 pooled centers for the purpose of analysis."

In one of hypertension clinical trials, the pooling strategy is described as “To avoid loss of information, small centers (<5 per protocol patients) were pooled from largest to smallest until the pooled center had 5 per protocol patients in each treatment group. These centers were grouped into 19 pooled centers for the purpose of analysis. The pooling algorithm was predetermined before unblinding the data, and the pooling algorithm was described in the statistical analysis plan for the study. Considering the subjective nature of the pooling algorithm, albeit prespecified before completion of the study, an exploratory analysis was also performed with actual center as a fixed effect in contrast to pooled centers. This analysis did not change the inference.”

In a type 2 diabetes trial, a different pooling strategy was used “For all center stratified analyses, centers with <24 randomized and treated subjects were pooled on a geographical basis, independently of treatment identification.”

In a recent brief book for PDAC, the sponsor provided the detail pooling strategy for centers “Pooling algorithm for centers: For non-US sites, all investigative sites within a country with fewer than 10 randomized subjects will be combined into a single pooled site for analysis purposes. If a resulting pooled site still has fewer than 10 randomized subjects, then this pooled site will be further combined with the smallest unpooled site within that country. If there is not another unpooled site within that country, then the pooled site will be combined with the smallest pooled site from another country. This pooling process will continue until there are at least 10 randomized subjects in each pooled site. For US sites, all investigative sites within a geographic region with fewer than 10 randomized subjects will be combined into a single pooled site for analysis purposes. If a resulting pooled site still has fewer than 10 randomized subjects, then this pooled site will be further combined with the smallest unpooled site within that region. If there is not another unpooled site within that region, then the pooled site will be combined with the smallest pooled site from another region within the US. This pooling process will continue until there are at least 10 randomized subjects in each pooled site.”

As we can see from the examples above, the cut point for center pooling (5, 8, 10, or 24) is really arbitrary and there is no scientific basis for choosing one or another. The decision on the cut point may be based on the distribution of the number of subjects across centers.

Center pooling strategy could sometimes be questioned by the regulatory reviewers. For example, in BLA review of Rebif, FDA reviewer had concerns about the pooling strategy “The sponsor’s study center pooling strategy: Per the pre-specified strategy in the sponsor’s statistical analysis plan (SAP), pooling of study centers for inclusion of center as a main effect in analyses was to have been based on geographic considerations for small centers. In fact, the pooling strategy actually used was data driven which is problematic. NOTE: There were 56 participating centers from 9 countries. The smallest recruiting center had 3 subjects, 2 centers contributed 4 subjects, and 5 centers contributed 6 subjects each. The remaining centers contributed between 6 – 24 subjects each (CSR, Table 3, pp. 65-66). This reviewer performed analyses of major efficacy endpoints based on strict geographic pooling of centers into 3 groups (US, Canada, and Europe) as well as un-pooled analyses (not including the center effect). In addition, descriptive analyses for individual centers were also performed for the primary and major secondary efficacy endpoints. The sponsor’s positive statistical findings were found to be robust based on these analyses.”

In Biopharmaceutical Report (Summer 1998), Paul Gallo wrote an article titled “Practical Issues in Linear Models Analyses in Multicenter Clinical Trials” which contained a section discussing “construction of composite centers”. The caveats of using the composite centers are also discussed in the paper.

“In performing unweighted analyses, a practice of defining artificial “pooled” or “composite” centers is often employed; that is, data from different centers are treated in the analysis as if they came from the same center. A number of small centers may be combined, or one or more small centers may be combined with a larger center. This practice attempts to minimize the large variance inflation and data instability of unweighted analyses when there are very small centers. Composites may be constructed to the extent of eliminating empty cells to ensure that treatment effects are estimable in models containing interaction terms. More commonly, this is done to achieve some minimum cell size felt to appropriately limit the influence of individual observations; values around 5 are often chosen. ”

Arbitrarily pooling the centers sometimes does not make sense at all. This is exactly true when the centers with small number of enrolled subjects are pooled even though these centers are scattered in totally unrelated geographic regions or countries. When pooled center is used and the statistically significant center effect is detected, the interpretation of the results is difficult. Instead of the center pooling purely based on the number of enrollees, the geographic distribution of centers should be considered. In many cases, instead of pooling centers by the number of enrollees, we could use country and geographic region in the analysis. In one of our multi-national clinical trials, we grouped centers by geographic region as North American, South American, Eastern Europe, Western Europe, and Eastern Asia. The strategy worked very well.

If possible, we could use the random effect model to include the study site / center as random effect to avoid the center pooling. We could also use a center weighting strategy that is similar to the Meta analysis where centers with more subjects are given more weights.

Tuesday, February 08, 2011

Guidelines for Blood Volumes in Clinical Trials (Especially in Pediatric Clinical Trials)

Nowadays, the clinical study protocols are becoming more and more complicated and require more and more blood sample draws for various purposes. The blood samples are needed for testing the hematology, chemistry, immunogenicity (for biological products), biomarkers (for diagnostic or other purposes), pharmacogenomics,… In some clinical trials, additional blood samples (sample retains) may be drawn for future studies (even though we may not know what the future study will be). If the study has the component of pharmacokinetics, the many more samples (series blood samples) will be drawn within a short period to characterize the pharmacokinetic profile, estimate the total drug exposure (AUC), and calculate other pharmacokinetic parameters.

With increasing in the number of blood draws or the blood volumes, the ethic issue often arises, especially in clinical trials with children.

US FDA and EMA do not really regulate the maximum blood volume that can be drawn from a subject during the clinical trials. The requirements for limiting the blood sample volume may come from the National Institute of Health (NIH), American Academy of Pediatrics, World Health Organization (WHO), and European Union (EU) and are typically enforced by the ethic bodies such as Institute Review Board (IRB) and Ethics Committee (EC). The requirements on blood volume during the clinical trials may be different depending on the country and local IRB.

The blood volume drawn for pharmacokinetic studies in the pediatric population is specifically a concern and has been discussed extensively. Stephen RC Howie (2010) reviewed blood sample volumes in child health research: a review of safe limits in the Bulletin of the World Health Organization (BLT). WHO also has its guidelines on drawing blood: best practices in phlebotomy. The guidelines are not specifically for clinical trials, rather for general blood donations. The guidelines contain specific technical requirements for the blood drawn in pediatric and neonatal subjects.

In US, Code of Federal Regulations has a specific chapter (Part 46) to discuss protection of human subjects and the chapter contains a subpart D to address additional Protections for Children Involved as Subjects in Research. While there is no specific requirement on the limit of blood volume, the CFR indicated that the research involves no more than minimal risk to the subjects and IRB should take into account the purposes of the research and the setting in which the research will be conducted and should be particularly cognizant of the special problems of research involving vulnerable populations, such as children, prisoners, pregnant women, mentally disabled persons, or economically or educationally disadvantaged persons. Similarly, the American Academy of Pediatrics has its policy on Guidelines on Ethical Conduct of Studies to Evaluate Drugs in Pediatric Populations. The policy requires “…with the growing number of pediatric drug studies, IRBs need to be familiar with the various research-design methods that minimize risk to the child. Examples include limiting research under some circumstances to pharmacokinetic and safety data, combining this approach with pharmacodynamic data, and minimizing the volume of blood withdrawn through the use of sensitive assays, pediatric enabled laboratories, and population pharmacokinetic approaches"

National Institute of Health Clinical Center has a guideline M95-9: Guidelines for Blood Drawn for Research Purposes in the Clinical Center.

Two articles from the web actually reflect the limit of blood volume in the US.
In EU, there are specific guidelines on "ETHICAL CONSIDERATIONS FOR CLINICAL TRIALS ON MEDICINAL PRODUCTS CONDUCTED WITH THE PAEDIATRIC POPULATION"



The guidelines on blood volume are usually based on the amount of blood in the percentage of total blood volume (BLV). BLV varies depending on age and body weight. A good reference for BLV for pediatrics can be found in pediatricareonline.com.

Friday, January 28, 2011

Edit check - a critical step to ensure the data quality during clincial trials

In clinical trial, one critical task is to ensure that the data collected or data entered into the system / database is valid, correct, and logically sound. This task requires a data quality plan starting from designing a good study protocol -> developing efficient case report forms -> providing clear instructions for completing case report forms -> implementing electronic edit checks -> monitoring the study data / source data verification -> data clarification process -> data review process. One of the steps is to implement the electronic edit checks.
Edit check is a program instruction or subroutine that tests the validity of input in a data entry program. According to the CDISC clinical research glossary from Applied Clinical Trials, the edit check is defined as:

An auditable process, usually automated, of assessing the content of a data field against its expected logical, format, range, or other properties that is intended to reduce error. NOTE: Time-of-entry edit checks are a type of edit check that is run (executed) at the time data are first captured or transcribed to an electronic device at the time entry is completed of each field or group of fields on a form. Back-end edit checks are a type that is run against data that has been entered or captured electronically and has also been received by a centralized data store.

Electronic edit checks allow us to use the power of the computer to check for illogical, incomplete or inconsistent data. In clinical trial, one of the most important tasks facing clinical data management personnel is to produce the electronic Edit Checks specifications for a study. Developing the electronic edit check specification -- and processing the queries that result from them -- is arguably the most vital and time-consuming data cleaning activity data management personnel undertakes. The study statistician should always participate in the process of developing the electronic edit checks to ensure that the critical edit checks are included. Effectively implementing the edit check can prevent the illogical, incomplete, or inconsistent data from entering into the data capture system or data set, which will make the downstream data analyses much easier.

There are two types of edit checks:

Univariate edit checks (include range checks): these are the edit checks only applicable to a single field or single variable. For example, for subject weight, we can set up an edit check to ensure that the extreme or unlikely value not to be entered. Let’s say we set up a range check if a data entry is smaller than 90 lb or greater than 300 lb. For lung function test, we may set up an edit check for predicted FEV1 to be no less than 20% because it is unlikely to have someone with predicted FEV1 <20%. The univariate edit checks are usually run instantly during the time of data entry.

Multivariate edit checks (also called aggregate edit checks): these are the edit checks with more than one fields or variables involved. These edit checks cross check the entries across multiple fields / variables to ensure the data is logical and consistency. For example, if the entry on Gender field is ‘Male’, there should not be data for pregnancy test result field. If the reason for subject dropping out the study is entered as ‘adverse events’, there should be a corresponding entry in AE data set. Statistician can provide great inputs in identifying the multivariable edit checks. Some multivariate edit checks could involve the complicated algorithm and take considerable time to run. In this situation, the multivariable edit checks can be run at back-end at a specified interval (for example, 2 am at night).

One misunderstanding is to think that all data issues can be resolved by implementing the edit checks. Edit check is only one of the steps in the data cleaning process. Also, there should be balance in terms of the number of edit checks. Too many edit checks for non-critical fields could be very annoying for people who enter the data. This is especially true for clinical trials using electronic data capture (EDC) where the data entry responsibility is delegated to the investigator and study coordinators who may lose patient if there are too many pop-up messages during the data entry. For example, if the telephone number needs to be entered, an edit check to enforce the data entry to follow xxx-xxx-xxxx would be unnecessary (xxxxxxxxxx and 1xxxxxxxxxx should also be accepted) – this is an example I see in some of the web forms – very annoying).

Sunday, January 23, 2011

Regulatory Guidance on Source Data in EDC Trials

When we move toward the clinical studies using electronic data capture, the ‘source data’ or ‘source document’ has been an issue. Unlike the paper-CRF (case report form) based study, the source data in EDC study can be confusing and sometimes vague. If the data was directly entered into EDC system, the EDC system is the direct source and there is no another source to be verified against. This could be worrisome to some people. In a 2008 article, I talked about this issue.
Recently, both FDA and EMEA published the guidance on this issue. FDA’s guidance "Electronic Source Documentation in Clinical Investigations" was issued in December, 2010. EMEA issued its guidance last June and the guidance titled “Reflection paper on expectations for electronic source data and data transcribed to electronic data collection tools in clinical trials”.

The guidance titles seem to suggest that they are written for the data management functions, however, the discussions in these two guidelines are more relevant to the clinical sites and study monitors. Switching the clinical study from paper CRF to EDC is not just about the shift of the data entry from data management group to the clinical sites, it actually has impact on how the entire study is operated.

Tuesday, January 11, 2011

FDA's New Website for Industry

Have you noticed the changes in the design of FDA website (http://www.fda.gov/) recently? Last August, I mentioned the FDA's initiatives on transparency. As part of FDA's continued push to increase transparency in an agency once notorious for making decisions behind closed doors, the FDA has launched a new Web-based resource that industry can use to keep abreast of the regulatory status for drugs, devices, food, and cosmetics. The new website is under http://www.fda.gov/ForIndustry/ and is supposed to provide a repository for industries to understand FDA's detail processes in submission, reviewing, approval, and surveillance of the regulated products, and even the processes for complaints (dispute resolution). The website includes the sections that are very pertinent to us working in the pharmaceutical industry:
  • Developing products for rare disease and conditions
  • Dispute resolution
  • Guidance documents
  • FDA eSubmitter
  • Data standards
  • FDA basics for industry
FDA basics for industry includes the kind of basic information about the regulatory process that is often requested by drug, device, and biologic companies and is aimed at improving communication between FDA and industry by making basic information about the regulatory process more accessible to industry in a user-friendly format.

The new website reflects the great improvement towards the transparency and is a great resource for professionals working in the drug development industry.

Also see:

Sunday, January 02, 2011

Agreement Statistics and Kappa

In clinical trial and medical research, we often have a situation where two different measures/assessments are performed on the same sample, same patient, same image,… the agreement needs to be calculated as a summary statistics. Depending on whether or not the measurement is continuous or categorical, the agreement statistics could be different. Lin L had a very nice overview for agreement statistics.

Specifically for categorical assessment, there are many examples where the agreement statistics is needed. In a clinical trial with imaging assessment, the same image (for example, CT Scan, arteriogram,…) can be read by different readers. For disease diagnosis, a new diagnostic tool (with advantage of less invasive or easier to implement) could be compared to an established diagnostic tool… Typically, the outcome measure is dichotomous (e.g., disease vs no disease, positive vs. negative…).

The choice of the methods of comparison is influenced by the existence and/or practical applicability of a reference standard (golden standard). If a reference standard (golden standard) is available, we can estimate sensitivity and specificity – ROC (receiver operation characteristics) analysis. If a reference standard is not available or there is no golden standard for comparison, we can not perform ROC analysis. Instead, we can assess the agreement and calculate the Kappa. This has been discussed in detail in FDA’s Guidance for Industry and FDA Staff “Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests”. For example, for comparing the assessment from two different readers, we would calculate Kappa, overall percent agreement, positive percent agreement, and negative percent agreement. We would not use ROC statistics and would not calculate the sensitivity and specificity.
If we would like to assess the agreement between the urine pregnancy test and the serum pregnancy test, we could use the ROC and calculate the sensitivity, specificity, positive predictive value, and negative predictive value since the serum pregnancy test could be considered as a reference standard or golden standard for pregnancy test.

Kappa Statistic(K) is a measure of agreement between two sources, which is measured on a binary scale (i.e. condition present/absent). K statistic can take values between 0 and 1.
  • Poor agreement : K < 0.20
  • Fair agreement : K = 0.20 to 0.39
  • Moderate agreement : K = 0.40 to 0.59
  • Good agreement : K = 0.60 to 0.79
  • Very good agreement : K =0.80 to 1.00
A good review article about Kappa Statistics is the one written by Karemer et al “Kappa Statistics in Medical Research”.

SAS procedures can calculate Kappa Statistics easily. Here is a list of papers:

Monday, December 27, 2010

Bootstrap and SAS

In statistics, bootstrapping is a resampling technique used to obtain estimates of summary statistics. In clinical trials, bootstrapping technique could be a useful approach in obtaining the precision of an estimator. Most common application of the bootstrapping technique may be in obtaining the confidence interval for an estimator while the typical way of obtaining the confidence interval through the standard error approach is impossible or difficult.

Here are two examples that the bootstrapping technique needs to be implemented. The first example is for a manuscript. When we submitted our paper to European Respiratory Journal, one of the reviewer comments was a request for evaluating the internal consistency. The comment says “The statistical method is sample-based as it consists in a regression performed on this sample. Such a method needs at least evaluation for internal consistency (by measuring the regression correlation on a subsample then validating on another subsample or better by using bootstrap and jackknife methods).”

The second example is a request from the regulatory agency for calculating the 95% CI for % relative dif
ference. When there are two treatment means: A and B; % relative difference is defined as %RD= (A-B)/A. There may be other approaches in this case, but bootstrapping technique could come handy in calculating the 95% CI for %RD.

Bootstrap can be easily implemented in SAS and it contains three main steps: 1) resample the data from the observed data set (observed data is only one sample) – SAS Proc Surveyselect can serve this purpose 2) obtain the statistics (or estimator) by performing the analysis for each sample / resample 3) perform the summary statistics from the collection of the statistics or estimator.




Bootstrap is a suggested statistical approach for obtaining the confidence interval for individual and population bioequivalence criteria.

Some good references about how to do bootstrapping using SAS are included here:

Ten years ago, I had to use a SAS macro to do the bootstrap for my PhD dissertation. The macro is still there on SAS website.

Bootstrap technique has also been built into several SAS procedures (such as Proc Multtest, Proc MI).

When bootstrap is used in regression situation, 'Bootstrap Pairs' technique may be employed. Freedman (1981) proposed to resample directly from the original data: that is, to resample the couple dependent variable and regressor, this is called bootstrapping pairs.  Bootstrap pairs is described in a paper by Flachaire. The SAS macro for bootstrapping discussed two main ways to do bootstrap resampling for regression models, depending on whether the predictor variables are random or fixed.If the predictors are random, you resample observations just as you would for any simple random sample. This method is usually called "bootstrapping pairs". If the predictors are fixed, the resampling process should keep the same values of the predictors in every resample and change only the values of the response variable by resampling the residuals. 

Sunday, December 12, 2010

Counting the study day


For every clinical trial, we need to count the study day for calculating the follow-up visits and for assessing the temporal relationship between events. The study day starts with the day that the subject is randomized and receives the first dose of the study medication. Usually, the randomization date and the first dose of the study medication date are the same. In clinical study protocol, there should always be a ‘schedule of events’ or ‘schedule of evaluations’ table which defines the study procedures and the study visits. This table should include the study day.

There is one critical difference in counting the study days. The protocol could count the day of subject receiving the first dose of the study medication as “day 0” or “day 1”.

If the first dose date is counted as day 0, the day immediately after the first dose date will be counted as day 1 and the date immediately before will be counted as day -1. Therefore, the study day is counted continuously as … day -7, day -6, day -5, day -4, day -3, day -2, day -1, day 0, day 1, day 2,… In this case, for programming, the study day variable can be created using the formula:

          The event/visit date – first dose date  

The problem with this counting is in 'day 0'. People are used to calling the first day of the study medication as the 'day 1'.

If the first dose date is counted as day 1, the day immediately after the first dose date will be counted as day 2 and the date immediately before will be counted as day 0 – which is confusing. In practice, if the first dose date is counted as day 1, the day 0 will not be used in the study day counting. The date immediately before will be counted as day -1 (skipped day 0). Therefore, the study day is counted as: day -7, day -6, day -5, day -4, day -3, day -2, day -1, day 1, day 2,… For programming, the study day variable would be created using two separate formulas for predose and postdose visits.
For pre-dose:
           the event/visit date – first dose date
For post-dose:
           the event/visit date – first dose date + 1

Both of these approaches (counting including study day 0 or not including study day 0) are not wrong, but sometimes confusions can arise when we calculate the study day variable. Even for CDISC, there are disagreements in handling this between Submission Data Set Tabulation Model (SDTM) (not allowing study day 0) and Analysis Data Set  Model (ADaM) (allowing study day 0).

The following clinical trial protocol templates indicate that the study day counting starts with day 0:
The following clinical trials indicate that the study day counting starts with day 1. There are more industry trials like this.

The unit used in counting the study day depends on the length of the clinical trials. For a trial with months and years in duration, instead of counting by day, it is more practical to count by week, month, or year. For example, for a clinical trial with three years treatment duration, the last treatment date would be three years away. If we count by day, it will be something like day 1095. Even worse, some people may apply the time window to this date to have the last treatment date 1095 +/- 7 days. Sound stupid, isn’t it?

Counting the study day correctly is important for study investigators/coordinators to avoid the protocol deviation. The Barnettinternational actually developed a tool to facilitate the study day /visit scheduling.  

Monday, November 29, 2010

A conditional probability issue?

There is a question and answer from 'AskMarilyn' at Parade.com. I copy the question and answer here since it is a probability issue.

Question: Four identical sealed envelopes are on a table. One contains a $100 bill. You select an envelope at random and hold it in your hand without opening it. Two of the three remaining envelopes are then removed and set aside, still sealed. You are told that they are empty. You are not given the choice of keeping the envelope you selected or exchanging it for the one on the table. What should you do? A) Keep your envelope; B) switch it; or C) it doesn't matter.

Marilyn said you should switch envelopes. Here's her reason: Imagine playing this game repeatedly. You start with a 25% chance of choosing the envelope with the cash. Then two empty ones are taken away on purpose. (Only someone with knowledge of the contents can inform you that sealed envelopes are empty.) so if the $100 bill is in any of the three unchosen envelopes - which it is 75% of the time - you'll get it by switching.

However, I would choose the answer C) it doesn't matter. This is a conditional probability issue. In the beginning, with all four envelopes sealed, the probability of choosing one envelope with $100 bill is 25%. When two envelopes are revealed not to contain the $100 bill, for the remaining two envelopes, each now has 50% probability with $100 bill in it. It doesn't matter if you keep the envelope on hand or switch it for the one on the table.

Saturday, November 20, 2010

Using RevMan to Conduct the Meta Analysis

RevMan (or Review Manager) is designed as a review tool to facilitate the literature review and the meta analyses by the Cochrane Collaboration Group. RevMan can be downloaded from website for free. It can be installed into your system without requiring the system administer privilege. Thousands of systematic reviews and meta analyses published on the Cochrane Library are performed using RevMan. These systemic reviews and meta analyses have been one of the leading resources in evidence-based medicine.

RevMan can be easily used by the medical researchers who are non-statisticians. For statisticians who work in the medical research area, RevMan is an easy tool to perform the meta analyses and generate the graphs (forest plot, funnel plot) in publication standard.

The statistical method and statistical model are described in the document Standard statistical algorithms in Cochrane reviews by Jon Deeks and Julian Higgins and Cochrane Handbook for Systemic Review of Interventions. For statistical models, both fixed model and random model are included in the RevMan. For random models, DerSimonian and Laird random-effects models are used. This is most common random effects model used in Meta Analysis.  

RevMan 5 is extremely easy to use. Various tutorials, tips, webinars are provided in RevMan documentation website and The Cochrane Collaboration Open Learning Material. I find it is extremely useful to watch two webinars (especially the part 2 regarding the data and analyses. For

To perform a Meta analysis, RevMan is just a tool. There are a lot of works to be done prior to enter the information including data into the RevMan. Considerable time needs to be spent on the literature search. Since the data used in Meta analyses relies on the publications, some data needs to be converted first. For example, for outcomes measured in continuous variable, the published article may only provide the Standard Error or just the 95% confidence interval. The SE can be easily converted to the Standard Deviation by multiplying the square root of the sample size. If only the 95% confidence interval is available, the standard deviation can be approximated by normal approximation using upper bound = mean +/- 1.96 * SE.

Sunday, November 07, 2010

Good Review Practice

In previous article, 'regulatory science' is discussed. 'Good Review Practice' can be considered one aspect of the regulatory science. Here Good Review Practice is specifically refer to a “documented best practice” within CDER that discusses any aspect related to the process, format, content and/or management of a product review.

On the industry side, the sponsor needs to establish the standard operating procedures (SOP) and the working procedure documents (WPDs) to ensure the compliance of the regulatory guidance and GCP and to improve the efficiency. On the regulatory side, it is important to establish the good review practice to ensure that the same standard procedures are following during the review process for drug approval.

These good review practices could cover the review process in different areas: efficacy, safety, pregnancy, CMC,... They are supposed to be written for FDA reviewers, however, understanding the good review practice is also very helpful for sponsor to prepare the regulatory submission documents in a way that is amenable to the reviewers. The mis-communication between the sponsor and the regulatory could be minimized. All necessary information/analyses required per good review practice are included in the submission documents.

Below are some links related to good review practice:

Sunday, October 31, 2010

Regulatory Science


Regulatory Science is the science of developing new tools, standards, and approaches to assess the safety, efficacy, quality, and performance of all FDA-regulated products including drug, biological products, medical device and more. On February 24, 2010, FDA along with NIH launched its Advancing Regulatory Science Initiative (ARS) aim to accelerate the process from scientific breakthrough to the availability of new, innovative medical therapies for patients.

On October 6, 2010, the U.S. Food and Drug Administration unveiled an overview of initiatives to advance regulatory science and help the agency assess the "safety, efficacy, quality and performance of FDA-regulated products." And published its white paper  Advancing Regulatory Science for Public Health - A Framework for FDA's Regulatory Science Initiative” The white paper outlines the agency's effort to modernize its tools and processes for evaluating everything from nanotechnology to medical devices to tobacco products.

In companion to the release of the white paper, FDA commissioner, Dr Hamburg gave a speech to the National Press Club in Washington, DC

In the white paper, the section I “Accelerating the Delivery of New Medical Treatments to Patients” has specific meaning to statisticians. “Adaptive design” was not specifically mentioned in the white paper, however, any approach or methodology in clinical trial design that can expedite the drug development process should be encouraged. The personalized medicine should also be encouraged.

Even though the regulatory science or regulatory affairs is critical in drug development field, the professionals working in the field are very diversified and come from variety of different backgrounds. Perhaps, you can only learn the regulatory science through the experience and on-job training. However, I do notice that USC has a graduate program in regulatory science. Considering that FDA is increasing its investment in regulatory science and the regulatory laws are getting more and more complicated, the graduates from this program should not have any difficulty in finding a job.


Friday, October 08, 2010

Missing data in clinical trials - the new guideline from EMEA and National Academies

Missing data issues have been discussed and debated for many years. Handling of missing data in clinical trials has been recognized as an important issue not only for statisticians who analyze the data, but also for the clinical study team who conduct the study.  While we are still waiting for FDA to issue its guidance on missing data in clinical trials, there are several guidelines published recently.

EMEA just issued its final rule of "Guideline on missing data in confirmatory clinical trials". This guideline provided the guidance on handling the missing data from the perspective of European regulatory authorities. Comparing to the FDA's guidance on non-inferiority and adaptive design, EMEA's missing data guidance is written in plain language and can be easily understood by the non-statisticians.

The recent trend is to discourage the use of LOCF and other single imputation methods (ie, replace the missing value with the last measured value, with averaged value, or with baseline value,...). It is noted that LOCF is mentioned as one of the single imputation methods in EMEA's guideline. The guideline acknowledged that "Only under certain restrictive assumptions does LOCF produce an unbiased estimate of the treatment effect. Moreover, in some situations, LOCF does not produce conservative estimates. However, this approach can still provide a conservative estimate of the treatment effect in some circumstances.". The guideline further elaborated that LOCF may be a good technique for studies (e.g. depression, chronic pain) where the condition is expected to improve spontaneously over time, but may not be conservative for studies (e.g. Alzeimer's disease) where the condition is expected to worsen over time.

In the United States, the Division of Behavioral and Social Sciences and Education under National Research Council of the National Academies have been working on a project "Handling missing data in clinical trials". The working group recently makes its draft report available. The draft report is titled "The prevention and treatment of missing data in clinical trials". I like the word 'prevention' in the title since it is critical to prevent or minimize the occurrence of missing data. Once the missing data has happened, there is no universal method to handle the missing data perfectly. The assumptions of MACR, MAR, and MNAR can never been fully verified.

Academies' report on missing data has a stronger language in discouraging the use of LOCF and other simple imputation approaches. The recommendation #10 stated "Single imputation methods like last observation carried forward and baseline observation carried forward should not be used as the primary approach to the treatment of missing data unless the assumptions that underlie them are scientifically justified."

So far, there is no official guideline from FDA regarding the missing data handling (even though the topic has been the perennial topic in almost all statistics conferences and workshops). Nevertheless, a presentation by Dr. O'Neill to the International Society of Clinical Biostatistics may give some insides.