Sunday, October 14, 2012

Using Area Under the Curve (AUC) as Clinical Endpoints


Area Under Curve (AUC) has been frequently used as the endpoint measure in clinical trials. We use AUC commonly in clinical pharmacology - Area under the time concentration curve or in diagnostic research – Area Under the ROC curve. The use of AUC is much more broader than what we think. Many clinical endpoints can utilize the AUC as a measure for the aggregate effect over a period of time. Below are some of the examples that I have experienced where AUC is used in clinical trials not for the purpose of pharmacokinetics measure or ROC measure.  

AUC Used in Pain Assessment

In the study of pain medications (usually acute pain medications), the pain intensity or pain relief in scales are measured at pre and serial time points post analgesic drug administration. The Summed Pain Intensity Difference (SPID) and total pain relief (TOTPAR) are usually calculated and used as the efficacy endpoints. TOTPAR is a time-weighted measure of AUC or total area under the pain relief curve and is a summary measure that integrates serial assessments of a subject’s pain over the duration of the study. The area under the pain relief vs. time curve can be used to derive the proportion of patients experiencing typically 50% pain relief over a specified time frame. This can be calculated as the ratio of two AUCs: TOTPAR vs. maxTOTPAR (maximum potential value for TOTPAR) as illustrated in the “Analysis of scale results – summary measures” of pain.

FEV1 AUC

In Asthma and COPD studies, FEV1 can be measured at pre-dose and at several serial time points post the treatment. The AUC will then be calculated from the time-FEV1 curve.

In  DULERA drug label, FEV1 AUC(0-12 hr) was mentioned as the efficacy measure:

FEV1 AUC (0-12hr) was assessed as a co-primary efficacy endpoint to evaluate the contribution of the formoterol component to DULERA. Patients receiving DULERA 100 mcg/5 mcg had significantly higher increases from baseline at Week 12 in mean FEV1 AUC (0-12 hr) compared to mometasone furoate 100 mcg (the primary treatment comparison) and vs. placebo ......"

In a recent news release “Results of Phase II Study of Boehringer Ingelheim's Investigational Bronchodilator for COPD Presented at 2012 ATS International Conference”, FEV1 AUC was used to measure the treatment effect in COPD.


“Results of the study found olodaterol 5 microgram QD provided significant improvement in lung function as measured by FEV1 AUC(0-12) versus twice-daily olodaterol 2 microgram, while twice-daily dosing of olodaterol 5 microgram had a better FEV1 AUC(0-12) profile versus once-daily olodaterol 10 microgram”

AUC in Type 1 Diabetes

In type-1 diabetes research, the main purpose of the treatment is to preserve the beta-cell function. The assessment of beta-cell function is through the measurement of the C-peptide concentration after simulated Mixed Meal Tolerance Test (MMTT) - the gold standard measure of endogenous insulin secretion

In the mixed-meal tolerance test (MMTT), commonly used in the U.S., a liquid meal (Sustacal/Boost) is ingested in the fasting state with timed measurements of C-peptide over the subsequent 2–4 h. The AUC is then calculated for the area under time-C-Peptide curve over 2 hour (AUC0-2hr) or 4 hours (AUC0-4hr) (see Greenbaum at al “Mixed-Meal Tolerance Test Versus Glucagon Stimulation Test for the Assessment of β-Cell Function in Therapeutic Trials in Type 1 Diabetes”).

In type 1 diabetes research, a concept of mean AUC is also used. Mean AUC is calculated by the AUC divided by the time duration (i.e., AUC0-2 hr / 120 minutes or AUC0-4 hr / 240 minutes)

AUCs to Assess the Responsiveness

We recently published a paper “Vigorimeter grip strength in CIDP: a responsive tool that rapidly measures the effect of IVIG – the ICE study” where we used AUCs to compare the responsiveness of two difference measures. Since two different measures used different scales, we had to calculate the SRM (standardized response mean) before we calculated the AUCs for INCAT scale and for Grip Strength . The larger the AUC, the higher the responsiveness to the treatment. The results indicated that the Vigorimeter grip strength could be more sensitive measure comparing to INCAT scale to evaluate the treatment effect of IVIG in CIDP patients.

AUC for Visual Analog Scale (VAS) for Dyspnea in Acute Heart Failure

In FDA's Cardiovascular and Renal Drugs Advisory Committee Meeting in March 27, 2014 for Serelaxin for Acute Heart Failure, one of the statistical issues discussed was the use of VAS AUC to assess the dyspnea in Acute Heart Failure. The FDA presentation included the detail calculation of the VAS-AUC and the results from this endpoint.

Friday, October 05, 2012

Missingness Mechanism (MCAR, MAR, and MNAR) - A Great Explanation of These Terms

For statisticians working in clinical trial field, the best challenge may not be in the statistical methodologies. The best challenge may be in communication with non-statisticians (such as physicians, clinical team members, corporate executives) about the statistical concepts and the statistical terminologies in plain languages.  

The missing data is very common in clinical trials and the concept of the missing data is very easy to understand. However, the categories for missing data mechanisms (or taxonomy of missingness) are not so easy to understand. A formal taxonomy exists for classifying missing data mechanisms, including for longitudinal and event history data. The mechanisms can be classified as MCAR (missing completely at random), MAR (missing at random), and MNAR (missing not at random). Take a look at the definition of MCAR, MAR, and MNAR below, you will see that these definitions are not easy to be understood by non-statisticians.


EMA
MCAR
For the dependent variable (conditional on the covariates in the model), if the probability of an observation being missing does not depend on observed or unobserved measurements then the observation is Missing Completely At Random (MCAR).
In the case of MCAR, the missing data are unrelated to the study variables: thus, the participants with completely observed data are in effect a random sample of all the participants assigned a particular intervention. With MCAR, the random assignment of treatments is assumed to be preserved, but that is usually an unrealistically strong assumption in practice.
*        
MAR
Conditional on the covariates in the model, if the probability of an observation being missing depends only on observed measurements then the observation is Missing At Random (MAR).
In the case of MAR, whether or not data are missing may depend on the values of the observed study variables. However, after conditioning on this information, whether or not data are missing does not depend on the values of the missing data.
MNAR
When observations are neither MCAR nor MAR, they are classified as Missing Not At Random (MNAR), i.e. the probability of an observation being missing depends on unobserved measurements. In this scenario, the value of the unobserved responses depends on information not available for the analysis (i.e. not the values observed previously on the analysis variable or the covariates being used), and thus, future observations cannot be predicted without bias by the model.
In the case of MNAR, whether or not data are missing depends on the values of the missing data.


Thanks to Ziad Taib, the following example for three different missingness mechanisms were explained very well and were easy to be understood by the non-statisticians.  

Suppose you are modelling weight (Y) as a function of sex (X). Some respondents wouldn't disclose their weight, so you are missing some values for Y. There are three possible mechanisms for the nondisclosure:
  • There may be no particular reason why some respondents told you their weights and others didn't. That is, the probability that Y is missing may has no relationship to X or Y. In this case our data is missing completely at random (MCAR)
  • One sex may be less likely to disclose its weight. That is, the probability that Y is missing depends only on the value of X. Such data are missing at random (MAR)
  • Heavy (or light) people may be less likely to disclose their weight. That is, the probability that Y is missing depends on the unobserved value of Y itself. Such data are not missing at random or missing not at random (MNAR)
 
Understanding the concept of missing mechanism is one thing, fully understanding missing mechanism in practice is another story. The reason for missing data is often not collected or incompletely collected in the clinical trials. Patients may not tell the real reason for them to withdraw from the study (discontinue from the study earlier). Academy’s suggestions below are reasonable, however, ‘full and detailed documentation for each individual of the reasons for missing records or missing observations’ is not the reality in the current clinical trial practice.
"Reasons for missing data must be documented as much as possible. This includes full and detailed documentation for each individual of the reasons for missing records or missing observations. Knowing the reason for missingness permits formulation of sensible assumptions about observations that are missing, including whether those observations are well defined.
Missing data in clinical trials can seriously undermine the benefits provided by randomization into control and treatment groups. Two approaches to the problem are to reduce the frequency of missing data in the first place and to use appropriate statistical techniques that account for the missing data. The former approach is preferred, since the choice of statistical method requires unverifiable assumptions concerning the mechanism that causes the missing data, and so always involves some degree of subjectivity.”

Tuesday, October 02, 2012

Subgroup Analyses for Clinical Trial Data

Subgroup analyses have been used in the clinical trials for many years and when the sample size is adequate, subgroup analyses can assess the qualitative consistency of treatment effect across different subgroups and can provide the information for identifying a sub-population that may have greater benefit from the treatment. Recently, the purpose of subgroup analyses has been expanded. As mentioned in the “Concept paper on the need for a Guideline on the use of Subgroup Analyses in Randomised Controlled Trials”, the subgroup analyses can be used to:

-          Assess internal consistency,
-          Try to rescue trials that ‘fail’ based on the full analysis set
-          try to identify patient groups with the most favourable benefit-risk profile

The subgroup analyses may:
-          be pre-specified in the trial protocol, based on demographic, genomic or disease characteristics (e.g. sub-entities of a disease that are widely recognised within the medical community)
-          materialise based on a need or desire to further explore study results.

Sub-group analyses may be especially useful in personalized medicine. Through sub-group analyses, certain biomarkers/subgroups may be identified so that we can develop the tailored therapeutics.

If the sub-group analyses are not pre-specified and are post-hoc after the data dredging / mining, the interpretation of the findings from the sub-group analyese needs to be cautioned. If we introduce the ‘learning/confirming’ concept, the post-hoc sub-group analyses is a learning process and the results of interest need to be confirmed in further prospectively designed trials.

There are many examples of clinical trials where the statistically significant treatment effect in a specific sub-group identified from a study can not be subsequently verified / confirmed in prospectively designed trials.

In PRASE (THE PROSPECTIVE  RANDOMIZED AMLODIPINE SURVIVAL EVALUATION) study, the study was powered to detect the treatment difference in death from any cause and hospitalization for major cardiovascular events between Amlodipine and Placebo in overall population. The study result was not statistically significant. Sub-group analyses were then performed to compare the treatment difference in the patients with ischemic heart disease and in the patients with nonischemic cardiomyopathy.The statistically significant difference between the Amlodipine and Placebo Groups was obtained among Patients with Nonischemic Dilated Cardiomyopathy. The author concluded “Amlodipine did not increase cardiovascular morbidity or mortality in patients with severe heart failure. The possibility that amlodipine prolongs survival in patients with nonischemic dilated cardiomyopathy requires further study”

A subsequent PRAISE-2 trial was conducted to confirm the finding from the sub-group analyses of the PRAISE study. Unfortunately, the results are negative. The study results were not published in peer-reviewed paper (since it is negative), but was presented in the scientific meeting by American Heart Association

“The Trial: PRAISE-2

Presenter: Milton Packer, Columbia University College of Physicians and Surgeons, New York, NY.
The study: A randomized, double-blind, placebo-controlled trial of amlodipine in patients with nonischemic cardiomyopathy on maximal medical therapy. A total of 1652 patients were randomized to receive either amlodipine (initially 5 mg/d, then increased to 10 mg/d after 2 weeks) or placebo. The primary end point of the study was all-cause mortality. The study was powered at 90% to detect a 25% difference in mortality between the treatment arms.
The results: No significant differences existed in all-cause mortality between the 2 arms (placebo, 31.7%; amlodipine, 33.7%; hazard ratio, 1.09; log-rank P=0.32). A pooled analysis of the PRAISE-1 and PRAISE-2 trials showed no significant affect of amlodipine on mortality (placebo, 34%; amlodipine, 33.4%; hazard ratio, 0.98; log-rank P=0.81).
Summary: Despite the fact that in PRAISE-1 a survival benefit was noted with amlodipine in patients with nonischemic cardiomyopathy, no such difference was noted in PRAISE-2 or when PRAISE-1 and PRAISE-2 were combined. Long-term treatment with amlodipine does not seem to be of benefit in patients with severe, chronic heart failure. “
 

In Sepsis indication, the Kybersept study compared the 28-day all cause mortality between High-Dose Antithrombin III and Placebo treatment groups. The results indicated “High-dose antithrombin III therapy had no effect on 28-day all-cause mortality in adult patients with severe sepsis and septic shock when administered within 6 hours after the onset. High-dose antithrombin III was associated with an increased risk of hemorrhage when administered with heparin. There was some evidence to suggest a treatment benefit of antithrombin III in the subgroup of patients not receiving concomitant heparin.”

Subsequently, a paper based on the post-hoc sub-group analyses were published and concluded “High-dose AT without concomitant heparin in septic patients with DIC may result in a significant mortality reduction. The adapted ISTHDIC score may identify patients with severe sepsis who potentially benefit from high dose AT treatment.”

Unfortunately, there was no formal randomized clinical trial to study this specific sub-group. Had a prospectively designed study been conducted to study the subjects without concomitant heparin, the results might not be significant as anticipated.

If the results from the sub-group analyses are used for supporting the drug approval or claim in the product label, it is obvious that the multiplicity issue will arise. A good paper about this is “A flexible strategy for testing subgroups and overall population” by Alosh and Huque.

For statistical issues arising from the clinical trial practice, European Medicines Agency (EMA) seems to be always ahead of the US. For the sub-group analysis, EMA organized an expert workshop on subgroup analysis in November, 2011. Various topics related to sub-group analyses were discussed during the workshop. The presentation materials and the workshop summary can all be found at EMA’s website.




Sunday, September 09, 2012

Adverse Event Collections for Screening Failures

In the last issue, I stated that a more accurate definition of Screening Failures may be as following “Potential subjects who were screened for the study participation, but were not enrolled (randomized or dosed) for the study”

Once we know the definition of the screening failures, the next question is about the data collection for screening subjects in clinical trials. How much data should we collect for screening failures? Should AE be collected and entered into the clinical database? for screening failures, should SAE be reported to regulatory agency and should SAE narratives be written?


Question:

I have a question related to the collecting and recording of screening failure adverse events during clinical trials. I work for a data management group of a medium sized pharma company. Currently we collect and database all AEs that occur to screening failures in our clinical trials.
On further research I have not been able to find any regulation or guidance document that requires this. Can you tell me what the FDA position is on this?
Should we
1) Record and database all AEs for screening failures?
2) Not record and database AEs for screening failures?
Clinical Site Quality Control
3) Not record or database and put systems in place that can detect an unusually large number of AEs in a specific site for screening failures.

Answer:

I'm not sure what you mean by "screening failure adverse events." Are you referring to an intercurrent illness or condition that occurs between the time that the subject was enrolled but before randomization that leads to the subject being considered a screening failure? If so, we do not consider these to be adverse events; because the subject has not yet received any study drug, it would not be "associated" with the use of the test article. Nevertheless, because the intercurrent illness or condition affects the subject's eligibility for the study, it should still be recorded by the study site and reported to the sponsor. The clinical investigator should also ensure that the subject receives appropriate medical care, either by providing it directly or by referring the subject back to the subject's primary care physician.
 

The answer above was not all accurate. During the screening period, subjects were exposed to the screening procedures which could cause adverse events; subjects could have emotional changes / nervousness / anxiety just because of the study participation;  subjects could be asked to change their regular treatment/medication to meet the inclusion /  exclusion criteria that in turn could cause side effects (such as withdrawal effect). Therefore, it is very possible for subjects to develop adverse events during the screening period. In other words, adverse events can be reported for screening failure subjects even though the subject has not yet received any study drug.

Different situations for adverse event collections can be listed in the table blow:

All Subjects Screened
Eligible for study participation
Screening failures
Adverse events during the screening period starting from ICF signing
Non treatment-emergent adverse events
Non treatment emergent adverse events
Adverse events reported after the first dose of the study drug
Treatment-emergent adverse events
-
Statistical analysis
Non-treatment emergent and treatment emergent adverse events are summarized separately
Not included in statistical analysis since screening failures are not included in safety population


From the table above, we can see the followings:

1. For subjects who are eventually randomized and receive the study medication, all adverse events need to be captured and recorded in the clinical database starting from the informed consent signing. By comparing the AE onset date/time with the first dose date/time, the AEs can be separated as non-treatment emergent AEs and treatment emergent AEs.
For this group of subjects, it is accurate to say that any AE occurred after the informed consent signing should be recorded.

2. For screening failures, whether or not the AEs reported during the screening period should be recorded in the clinical database is up for debating.

  • Some companies do not record any data into the clinical database for screening failures. All information about the screening failures are maintained in a screening log.
  • Some companies record only the key information into the clinical database for screening failures. The key information may be demographics, reason for screening failures
  • Some companies choose extreme conservative way and record all available information for screening failures in the clinical database.

Unfortunately, there is no clear regulatory guidance on what information (especially adverse events) should be recorded into the clinical database for screening failures. The languages from In ICH E3 (STRUCTURE AND CONTENT OF CLINICAL STUDY REPORTS), seems to suggest that for screening failures, only information needed may be the reason for screening failures (and adverse event could be one of the reasons for screening failure).

The most extreme (or conservative) situation could be that in a study, the SAE narratives would be written for all subjects including the screening failures. In order to have sufficient information for SAE narrative writing for screening failures, all details about SAE and the ancillary information (physical example, medical history, vital signs, laboratory, …) would need to be collected. A lot of time and efforts would be spent on the data collection, but the collected data would not be very useful or at least not relevant to the purpose of the study since in the end, the screening failures would be excluded from the safety population for the safety analysis. This practice of collecting almost every detail about the screening failures is not wrong, but is not an efficient way for conducting the clinical trials.

Nowadays, the industry trend is moving toward to being compliant with CDISC standards (SDTM, aDaM). There seems to be a lot of confusions about whether or not data for screening failures should be included in the database and if so, where to include. The following weblinks from CDISC Public Discussion Forum show the confusions.


 To summarize, for screening failures, the best way for data collection may be to collect only the demographic information and the reason for screening failures in clinical database. The reason for screening failure should include a category of “AE” since subject can be screening failure due to AE (or precisely non-treatment emergent AE) during the screening period. The details about the AE / SAEs for screen failure subjects are not necessary to be entered into the clinical database.

Thursday, September 06, 2012

Definition of Screening Failures in Clinical Trials

In all clinical trials, the typical process starts with a screening period. The screening period starts with the signing of the informed consent. During the screening period, inclusion/exclusion criteria for the study participation will be checked / tested. Subjects who meet all inclusion criteria and do not meet any exclusion criterion will be eligible to be randomized (in randomized trial) or to be dosed (in non-randomized trial). Those who are not eligible for randomization or dosing will be considered as ‘screening failures”.

It seems to be a straightforward concept. However, there could be confusions if there are subjects who are not randomized or dosed due to other reasons (for example consent withdrawal, family relocation, death during the screening period,…). These situations may not be part of the inclusion/exclusion criteria, but still cause the subjects not to be randomized or dosed.

What is the definition of “screening failures”? Will screening failures only refer to subjects who do not meet the inclusion/exclusion criteria?

In the most recent version of CDISC Clinical Research Glossary, the term Screening (of Subjects) is defined as “A process of active consideration of potential subjects for enrollment in a trial’ and the term Screen Failure is defined as “Potential subject who did not meet one or more criteria required for participation in a trial.” This definition of Screening Failure is accurate only if all other situations (such as consent withdrawal, lost to follow up,…) are part of the inclusion/exclusion criteria.

In ICH E3 (STRUCTURE AND CONTENT OF CLINICAL STUDY REPORTS), while no definition of Screening Failures are provided, it has the following statement and the example flow chart.

“…It may also be relevant to provide the number of patients screened for inclusion and a breakdown of the reasons for excluding patients during screening, if this could help clarify the appropriate patient population for eventual drug use.”
 

The annex IV b above implied that there could be multiple reasons for screening failures and inclusion/exclusion criteria would just be one of these reasons.

For example, in a clinical trial, we could have a case report form to ask the reasons for screening failures and we could have the following list of reasons:

Primary reason for screening failure:
  • Adverse Event
  • Patient Non-compliance
  • Consent Withdrawn
  • Inclusion/exclusion criteria not met
  • Lost to follow-up
  • Death
  • Other

Therefore, a more accurate definition of Screening Failures may be as following “Potential subjects who were screened for the study participation, but were not enrolled (randomized or dosed) for the study”

INSET statement in SAS Procedures

 
Recently, I find out how convenient to include some summaries statistics in a statistical graph with a statement called INSET. An INSET statement places a box or table of summary statistics, called an inset, directly in a graph created with a CDFPLOT, HISTOGRAM, PPPLOT, PROBPLOT, or QQPLOT statement. INSET statement is available in many SAS procedures (Proc Univeriate, Proc Boxplot, Proc Lifereg,...).
 
If we run the following program in SAS, the INSET statement used in Proc Univariate will place a box on the left corner of the CDF graph to indicate the mean and standard deviation.
 
data Cord;
label Strength="Breaking Strength (psi)";
input Strength @@;
datalines;
6.94 6.97 7.11 6.95 7.12 6.70 7.13 7.34 6.90 6.83
7.06 6.89 7.28 6.93 7.05 7.00 7.04 7.21 7.08 7.01
7.05 7.11 7.03 6.98 7.04 7.08 6.87 6.81 7.11 6.74
6.95 7.05 6.98 6.94 7.06 7.12 7.19 7.12 7.01 6.84
6.91 6.89 7.23 6.98 6.93 6.83 6.99 7.00 6.97 7.01
;
run;
 
title 'Cumulative Distribution Function of Breaking Strength';
 
proc univariate data=Cord noprint;
histogram strength /normal;
cdf Strength / normal;
inset normal(mu sigma);
run;
 
 The following are more examples of using INSET statements:
 
 

Monday, September 03, 2012

Free Lectures on Statistics and Medical Research

Now that we are in the internet era, the learning is not limited to be in the school. There are great resources on the web. The elite universities now post their video lectures for the public.

One great resource for mathematics/statistics and varriety of other topics is academicearth.org which features the lectures from Universities such as Harvard, MIT, Yale, Stanford,... For statistics,  there are six classes listed. Unfortunately, there is no topic specifically to the biostatistics. Another resource is open course which currently listed 500 free online courses from top universities.

For biostatistics, while there is no video lectures, there are recorded lectures in mp3 format. For example, there are five biostatistics classes listed at education-portal.com.
For topics in medical research (not necessarily clinical trials), there are more resources available.

Friday, August 17, 2012

Confidence Intervals for difference between two proportions and for the ratio of two proportions


For clinical trials with binary outcomes, the results can usually be presented as a 2x2 contingency table as below:


Responder
Non-responder
Total
Treatment 1
n11
n12
n1
Treatment 2
n21
n22
n2

We can then calculate the proportion of responders for two treatment groups:

       p1=n11/n1

       p2=n21/n2

We have two ways to compare two treatment groups:
  • The difference between two proportions: p1-p2
  • The ratio of two proportions: p1/p2

p1-p2 may be called the absolute risk difference and p1/p2 is called relative risk (RR) or risk ratio. 

The confidence interval can be constructed for the difference between two proportions and for the relative risk.

For the difference between two proportions, the asymptotic confidence interval is ca1culated using the following formula:

                                 (p1-p2) +/- Z(alpha/2)*sqrt((p1 *(1-p1)/n1)+(p2*(1-p2)/n2))

Reference: Stokes, Davis, and Kock (2000) Categorical Data Analysis using the SAS System, 2nd edition

The notations may be different in the reference book and in SAS manual, but the results should be the same.

I had a posting a while ago about “Confidence Interval for Difference in Two Proportions” where I mentioned the corrections and the SAS codes.

For relative risk, the asymptotic confidence interval is calculated using the following formula:

Exp(log(RR) +/- Z(alpha/2) * sqrt((1-p1)/(n1*p1) + (1-p2)/(n2*p2)))

Reference: Agresti A (2007) An Introduction to Categorical Data Analysis, 2nd edition, JohnWiley & Sons, Inc.,

The notations may be different in the reference book and in SAS manual, but the results should be the same.

The confidence interval for relative risk can be obtained from SAS Proc Freq and can also be manually calculated using the formula above and the formula from SAS manual.

Suppose we have study results as below:


Success
Non-success
Total
Trt1
63
3
66
Trt2
56
13
69


data example;
  length trt $8;
  input trt $ success $ count;
  datalines;
  trt1   yes 63
  trt1   no  3
  trt2   yes 56
  trt2  no  13
;
proc freq data=example;
  weight count;
  tables trt*success/measures nopercent nocol;
  title 'outputs from SAS Proc Freq';
run;


data agresti;
  n11=63;
  n21=56;
  n1=66;
  n2=69;
  p1=n11/n1;
  p2=n21/n2;
  rr = p1/p2;
  v = (1-p1)/(n1*p1) + (1-p2)/(n2*p2);
  upper = exp(log(rr) - probit(0.025)*sqrt(v));
  lower = exp(log(rr) + probit(0.025)*sqrt(v));
run;
proc print data=agresti;
  title "using the formula from Agresti's book"
run;


data sasmanual;
  n11=63;
  n21=56;
  n1=66;
  n2=69;
  p1=n11/n1;
  p2=n21/n2;
  rr = p1/p2;
  v = (1-p1)/n11 + (1-p2)/n21;
  upper = rr * exp(-probit(0.025)*sqrt(v));
  lower = rr * exp(probit(0.025)*sqrt(v));
run;

proc print data=sasmanual;
  title "using the formula from SAS manual";
run;

I recently read a paper by Fischer et al. The confidence interval for relative risk was constructed using a method by Koopman. In Koopman’s paper “Confidence Intervals for the Ratio of Two Binomial Proportions”, a Chi-square method was proposed and the method required using numerical procedure and the iterative computations. There is no SAS program available for the calculation using Koopman's method.

There are other approaches proposed for computing confidence intervals for the ratio of two proportions. However, the method for calculating the asymptotic confidence interval adopted in SAS Proc Freq is commonly used.  

Further reading: