For years, online prediction markets were largely associated with sports and political forecasting
Kalshi (a CFTC-regulated prediction-market exchange) has launched a pilot program for event contracts on late-stage drug trials and FDA outcomes. Each contract is a binary “Yes/No” bet tied to a specific public event (e.g. “Will Drug X meet its Phase 3 endpoint by date Y?”), with settlement based on published data (ClinicalTrials.gov records, FDA letters, advisory votes, etc.). Kalshi’s rules require that no market opens until after a trial’s enrollment closes, and that all traders verify employment and have no material nonpublic information (MNPI). The partners (Kalshi and AppliedXL) define clear resolution criteria before trading begins and review all evidence from primary sources to determine outcomes.
These markets aggregate dispersed
expectations about trial outcomes into a public probability. Advocates argue
they can reveal hidden insights in drug development and “democratize
forecasting” by giving stakeholders (investors, smaller companies, patients) a
shared “odds” number. Critics warn of manipulation risks, ethical issues, and
unintended effects on trials. For example, bioethicist Jonathan Kimmelman notes that market signals might corrupt
trials themselves by influencing patient enrollment, investigator behavior, or
dropout patterns. Both academic commentary and industry voices highlight pros
(efficient aggregation, speed) and cons (insider trading, feedback loops,
patient recruitment impacts) of these markets.
We compare these markets to (a)
statistical trial simulations (power analyses and Monte Carlo), (b) Wall
Street/analyst success-probability models, and (c) informal FDA-approval
forecasts. These approaches differ in inputs (e.g. design parameters vs expert
judgment vs historical rates), outputs (market price vs model probability vs
policy statement), transparency (public vs private), users (traders vs sponsors
vs regulators), and demonstrated accuracy. A summary table contrasts these
attributes.
The STAT Readout Loud podcast
(July 16 2026) featuring Kimmelman underscored
that prediction markets, by making expectations public, create a paradox: “the
very market that aggregates expert knowledge… can also nudge the patients,
doctors, and researchers inside a clinical trial toward the outcome that the
market predicts”. Kimmelman argued the most socially valuable markets would
bet on broad scientific paradigms (e.g. the amyloid vs tau hypotheses) rather
than individual trials.
Prediction markets offer novel, real-time signals about trials, but
carry risks. Regulators and sponsors should proceed cautiously. Stakeholders might
benefit from markets that aggregate evidence, but must guard trial integrity
and privacy. We recommend strict safeguards (as Kalshi is starting) and focus
on post-recruitment or paradigm-level markets to avoid distorting
ongoing trials. Broadening transparent data sharing (while respecting
confidentiality) and encouraging regulatory guidance will help realize
potential benefits (better forecasting, informed decisions) without undermining
ethical trial conduct.
1. How Kalshi’s Clinical-Trial Markets Work
Kalshi’s new markets are structured as CFTC-regulated event contracts.
Each contract asks a binary question about a named trial or FDA event (e.g. “Will
Company A’s Phase 3 trial meet its registered primary endpoint by date D?”
or “Will the FDA approve Drug X by date D?”). The contract has a price
from $0 to $1, interpretable as the market-implied probability of the outcome.
For example, a $0.72 price implies ~72% probability of “Yes”. Contracts are
settled in cash: if the event occurs (“Yes”), contracts pay $1; if not (“No”),
they pay $0.
Market Creation and Eligibility: Kalshi (not
AppliedXL) decides which markets to list. To protect trial integrity, Kalshi
and AppliedXL only list late-stage (Phase 3) trials by large companies, and
only after enrollment completes. This limits insider influence and
enrollment effects. Traders must pass employment verification (to ensure U.S.
residency) and Kalshi bars anyone with nonpublic insider information. These KYC
and MNPI rules mirror those in finance.
Contract Terms & Resolution: Each
contract’s terms specify exactly how the outcome will be determined. The
resolution source is a specific public document or record: e.g. the posted
results on ClinicalTrials.gov, the FDA’s published approval letter, or the
transcript of an advisory committee vote. AppliedXL predefines a logic for
interpreting that source (e.g. what constitutes “meeting” the endpoint) before
trading starts. If multiple data sources conflict or are delayed, the
contract’s resolution rules (and Kalshi’s exchange rules) dictate how to
proceed.
In practice, once results emerge, AppliedXL gathers the evidence
(registries, filings, publications) and maps it to the contract criteria. A
human reviewer checks for ambiguities. Kalshi then independently reviews the
compiled record and makes the final settlement call. In short, Kalshi is the
ultimate arbiter: its published exchange rules govern any dispute. AppliedXL
does not bind the outcome; it only provides an “auditable evidence
package” to support resolution.
Regulatory Status: Kalshi operates as a
CFTC-designated contract market (DCM). This means U.S. federal rules for
derivative trading (similar to futures) apply. Kalshi’s exchange complies with
transparency and surveillance rules (including MNPI prohibitions) common to
financial markets. In sum, the markets are legally structured as regulated
binary futures on factual, objectively verifiable drug-development events.
flowchart LR
TrialStart[Trial Initiated] --> EnrollEnd[Enrollment Closes]
EnrollEnd
--> MarketOpen[Contract Listed on Kalshi]
MarketOpen
--> Trading{Trading Period}
Trading
--> DataPub[Results Published (e.g. CT.gov, FDA letter)]
DataPub
--> ApplyGather[AppliedXL Gathers Evidence]
ApplyGather
--> KalshiReview[Kalshi Reviews & Adjudicates]
KalshiReview
--> Settled[Contract Settles $0 or $1]
Settled
--> MarketClose[Market Closes]
This flowchart shows the resolution process: only after enrollment ends
does Kalshi list the market. Traders then bet until results are published.
AppliedXL compiles relevant evidence, and Kalshi officially determines the
outcome and pays out.
2. Pros and Cons of Prediction Markets for Clinical Outcomes
Pros – Information Aggregation and Efficiency:
Prediction markets are praised for efficiently pooling dispersed information.
In theory, many participants (scientists, investors, even well-informed
patients) each contribute “private signals” by trading, and the market price
continuously reflects the aggregate wisdom. As one commentator noted, markets
draw on people “touching different parts of the elephant” of knowledge to reach
the truth. By trading, insiders and outsiders alike reveal their beliefs. This
can produce a real-time probability that updates with every new data point,
unlike sporadic polls or reports.
FierceBiotech quotes Endpoint Arena’s CEO arguing that markets “democratize”
trials and “incentivize” people to forecast scientific outcomes, potentially
speeding up discovery. Patients and doctors, Fischer claimed, could use markets
as one more tool to gauge trial success alongside medical advice. Indeed,
markets make explicit what is often implicit: they turn qualitative analyses
and rumors into a single number. And unlike a stock price (which lumps all
assets), a prediction market isolates the specific event (one trial or drug).
Cons – Manipulation, Bias, and Ethics: Many
experts warn that betting on trials risks distorting them. A constant concern
is insider trading: participants with nonpublic knowledge (e.g. biotech execs,
CRO employees, trial investigators) could trade on early data. While Kalshi
bans MNPI and job-related insiders, enforcement is challenging. A small, thinly
traded contract is vulnerable: even a modest bet could swing the odds,
misleading outsiders. For example, a drug’s sponsor might (illegally) bet on
success to move prices, or its competitor might bet on failure. The
Science/AAAS policy forum warns modern markets can be “designed to leverage
legal ambiguities” and can be “gamed by a single firm” in thin markets.
Critics also highlight feedback loops: public odds might influence
patient behavior and investigator judgments. As Windisch observes, clinical
trial recruitment is sensitive to perceived risks. If a market price dips low,
potential volunteers or their doctors might decide not to enroll, making
failure more likely. Conversely, if odds are high, control-arm patients might
drop out or report fewer side effects (to get the active drug), biasing results.
Kimmelman’s analysis formalizes this: he lists channels by which market signals
could bias trials (e.g. enrollment reluctance, investigator drift, differential
dropout).
The ethical stakes run deep: if a prediction market becomes
sufficiently accurate, Kimmelman argues, it implies trials are collecting data
we largely already know—raising a “Moneyball paradox” for pharma. In his view, “we
ought not to know the answer… if the trial is going to be ethical”. In
other words, profoundly predictable trials might violate the principle of
clinical equipoise. There are also patient confidentiality concerns: open
markets could tempt leaks of trial details (or patient-level experiences) into
public view.
Finally, social and legal downsides echo broader gambling risks. The Science/AAAS warning notes that large commercial prediction platforms prioritize engagement
and profit, potentially normalizing gambling-like behavior. Prediction markets
could encourage addictive speculation or distract from evidence-based
decision-making.
In summary, the benefits (aggregating knowledge, transparent signals)
must be weighed against serious risks: insider abuse, changes in trial conduct,
and ethical objections. Many recommend strict limits (as Kalshi is doing) or
even pausing markets during active recruitment.
3. Impacts on Clinical-Trial Operations
Introducing betting on trials could alter many aspects of how trials
are run:
- Trial Design: Sponsors may rethink
blinding and endpoints. Knowing a market will be watching, companies might
favor harder endpoints (e.g. overall survival) over subjective
measures, or avoid early-phase exploratory designs. Conversely, some worry
sponsors could game designs to influence markets (e.g. choosing broad
endpoints that are harder to “bet down”). Kalshi’s pilot avoids this by
focusing on agreed-upon primary endpoints and late trials.
- Recruitment & Enrollment: The
clearest impact is on patient enrollment. If public odds suggest a low
chance of success, fewer patients may volunteer. Melinda Chu (oncologist)
warned that a small biotech could fail to meet enrolment targets if a
market is pessimistic. The Stat/SensibleMed article models this: a patient
Googling a trial’s odds and seeing them low “may decide not to
participate”. This could slow enrollment or force sponsors to extend
trials. Even after closing enrollment, participants already enrolled might
drop out if they perceive an arm as failing. Kalshi’s requirement to list after
enrollment closes is meant to blunt this effect, but as Kimmelman notes,
markets can still influence behavior during follow-up.
- Investigator Behavior: Knowledge of
market odds could bias clinical assessments. Investigators might
(consciously or unconsciously) grade patient outcomes more favorably for
the arm predicted to win, or be extra vigilant for side effects in the
underdog arm. Windisch notes that even supposedly “blinded” trials are
often effectively unblinded by side effects. Thus outcome adjudication
could drift to align with expectations, compromising data integrity.
- Reporting and Sponsor Actions: Sponsors
watch markets too. If negative odds rise, a sponsor might hasten a press
release or request an earlier interim analysis. Mike Abrams (Numerof)
speculates that poor market pricing “invites those with access to
non-public information to compromise their integrity”. In extreme
cases, a sponsor might consider altering trial conduct (e.g. unblinding
early or stopping a trial) in response to market pressure. Transparency
moves by regulators (like releasing complete response letters) have
already strained companies; adding prediction markets could amplify
investor and media pressure around each trial result.
- Data Integrity and Privacy: Public odds
effectively publish a piece of interim information. The FDA and trialists
guard interim data carefully (Data Monitoring Committees are kept blind).
A real-time market price is a novel public signal. As Windisch warns, this
breaks the traditional data firewall. Moreover, if individuals feed inside
info (e.g. one site’s data) into trades, patient confidentiality could be
indirectly breached.
- Investigator Incentives: While most
discussion focuses on patients/sponsors, even trial investigators have
stakes. They cannot legally trade on MNPI, but they will see market odds.
A clinician who built their career on a drug might feel pressure if the
odds are low, or pride if they are high. It’s unclear how that social
factor might subtly affect trial conduct.
In sum, markets introduce new dynamics into trials. Kalshi’s safeguards
(no early-phase bets, post-enrollment launch, MNPI bans) are aimed at
minimizing distortion. But evidence (from social science and medical ethics
literature) suggests even later-phase, after-enrollment bets could influence
participation and conduct. Regulators and trial sponsors will need to monitor
for these effects carefully if such markets grow.
4. Comparison: Forecasting Methods for Clinical Success
We compare four approaches to forecasting trial outcomes:
- Prediction Markets (Kalshi) – Inputs:
collective trader info (public news, science publications, unofficial
reports). Outputs: live market price (probability). Transparency:
prices and contract terms are public (but trader identities are
anonymized). Incentives: financial profit; participants have skin
in the game. Users: investors, analysts, smaller companies, patient
advocates, curious public. Accuracy Evidence: Empirical data is
sparse for biotech markets specifically, but in other domains (politics,
sports) markets often match or beat polls/experts. (For example, markets
famously predicted elections better than most polls.) Accuracy depends on
liquidity and broad participation.
- Statistical Simulations (Power Analysis / Monte Carlo) – Inputs: known trial design parameters (sample size,
event rates, effect size assumptions, dropout rates), historical data. Outputs:
probability of detecting an effect (power), expected distribution of
outcomes under hypotheses. Transparency: usually private or
technical; sponsors and regulators see them, but models are rarely public.
Incentives: none of profit – used for trial planning and regulatory
justification. Users: trial statisticians, CROs, regulatory
reviewers. Accuracy Evidence: These models are “accurate” only to
the extent the assumptions (e.g. effect size) are correct. They help
design the trial (e.g. set sample size) but are not predictive of actual
real-world outcome beyond those assumptions.
- Analyst/Model Predictions (Wall Street) –
Inputs: company disclosures, clinical data, published literature,
historical success rates, competitive landscape, sometimes proprietary
databases. Outputs: often a “probability of success (POS)” or
recommendation (“buy/hold/sell”) but usually qualitative. Transparency:
low. Each firm’s model is proprietary, and analysts typically do not
publish their probability models; investors only see summaries or final
forecasts. Incentives: institutional profit, reputation. Analysts
may overestimate success to maintain stock coverage. Users: biotech
equity investors, pharma strategists. Accuracy Evidence: Historically
mixed; analysts incorporate many factors, but can be swayed by hype or
ignore failures. AppliedXL’s research shows traditional models often miss
trial-specific execution issues. (For example, analysts grossly
overestimated some high-profile trials that ultimately failed.)
- FDA/Historical Forecasts – Inputs:
broad historical success rates, class-effect knowledge, advisory committee
input. Outputs: not formally published, but the FDA effectively
uses internal probabilities to guide decisions (e.g. pre-specified
approval benchmarks). Some outside groups infer likely outcomes from trial
context. Transparency: limited. FDA guidelines and public data give
clues (e.g. FDA reports average success rates: historically ~48–50% for
Phase 3 to approval. Incentives: regulatory (public safety and
efficacy). Users: regulators, large pharma (for portfolio
planning). Accuracy Evidence: Reflects aggregate historical truth:
if “FDA forecast” means just using the baseline approval rate, it’s about
50% for a Phase 3 trial. But FDA also evaluates each trial's data
rigorously, so final decisions are tailored, not fixed by formula.
Below is a comparative summary table:
|
Attribute |
Prediction Markets (Kalshi) |
Trial Simulations (Statistical) |
Analyst Probability Models |
FDA/Historical Forecasts |
|
Inputs |
Public signals, news, expert tips (and sometimes private info) from
many traders |
Assumed effect size, variance, enrollment rates, epidemiology;
mathematical models |
Published data (trials, disclosures), historical success rates,
expert judgment |
Historical approval rates; broad clinical and drug-class knowledge
(e.g. “Amyloid vs Tau”) |
|
Outputs |
Market price (real-time probability) |
Statistical power curves, simulated outcome distributions |
Percentage likelihood or ordinal scores (often unpublished) |
None formal; general heuristic (“~50% if phase3”) often inferred |
|
Transparency |
High: contract terms and prices visible to all |
Low/Medium: typically internal to sponsor/regulator; not public |
Low: models proprietary; only sometimes summary data or odds are
shared |
Medium: FDA shares some data (success rates reports); but no
real-time “forecast” given |
|
Incentives |
Monetary gain from correct prediction; crowdsourced wisdom |
None (academic/regulatory goal of sound trial design) |
Career/reputation of analysts; investment profits |
Regulatory mandate; patient safety and public trust |
|
Typical Users |
Investors/traders, biopharma analysts, journalists, patient advocates |
Sponsors’ statisticians, clinical trial designers, FDA reviewers |
Institutional investors, biotech equity funds |
FDA reviewers, big pharma R&D planners |
|
Accuracy Evidence |
Empirically good in other domains (e.g. politics); untested in
pharma. Dependent on liquidity and diverse participation. |
Relies on quality of assumptions; can accurately predict power given
assumptions, but no real “success rate” metric. |
Mixed: many documented misses. (E.g. analysts missed NKTR-214 failure.) |
Tied to history: ~50% of Phase 3 trials succeed. Provides baseline
odds but not trial-specific factors. |
This table highlights that prediction markets provide a real-time
public signal tied narrowly to a defined event, whereas simulations are
forward-looking design tools, analysts rely on patchwork models, and FDA
forecasting is mostly implicit and aggregate. Markets score high on
transparency but raise unique incentive issues.
5. Insights from STAT’s “Predicting Biotech Clinical Trials” Podcast
STAT’s Readout Loud podcast (July 16, 2026) discussed these
markets with bioethicist Jonathan Kimmelman. Key points included:
- The Moneyball Metaphor: Kimmelman
described drug development as a “Moneyball problem” for pharma.
Traditional decision-making pools opinions of a few insiders, whereas
markets could synthesize broader expertise. He noted that markets can “do
a pretty good job getting us as close as possible to the truth” by
aggregating dispersed knowledge.
- Risk of “Infecting” Trials: He warned
that trading on trials “effectively bet[s] on the behavior of human
beings” – patients, doctors, researchers. He outlined four bias
channels (enrollment reluctance, assessment drift, patient-reported
distortions, asymmetric dropouts) through which public odds can influence
trial data. Even blinding doesn’t fully protect, since patients infer
their assignment over time.
- Ethical Paradox: Kimmelman’s standout
argument: “If… we can predict the outcomes of… trials, it suggests we
know too much at the point where we’re running clinical trials”. In
other words, if the market price is consistently accurate, one must ask
why randomization is ethical at all. He concluded “if the trial is
going to be ethical, we ought not to know the answer to that trial in
advance”.
- Stock Market vs. Prediction Market: When
asked why not just use stock prices, Kimmelman replied stocks are “clumsy
predictors” (they mix all company factors). A pure contract “focuses
the lens on one particular question”, isolating the trial from
corporate noise. Still, he cautioned that prediction markets “will not
eliminate the biases” inside companies – the market is only one input
among many.
- Better Targets – Scientific Paradigms:
Kimmelman argued the most valuable bets are “paradigm-level”: e.g. “Will
anti-tau therapy meaningfully alter Alzheimer’s progression by year X?”.
Such questions aggregate across trials and cannot be gamed by one
company’s data. They would help allocate R&D resources more broadly,
unlike narrow Phase 3 bets whose “development decision has already been
made”. He noted, however, that designing and resolving such long-horizon
bets is complex and not part of Kalshi’s initial launch.
These perspectives reinforce caution: even STAT’s hosts and guests
recognized the dual nature of the tool. Fischer (Endpoint Arena CEO) remained
optimistic about motivation and speed, but Kimmelman and others stressed
potential trial impacts. As the podcast summarized, one must guard “whether
traders, patients, and trial investigators can resist the impulse to let the
odds shape the outcome.”
6.
Conclusion & Recommendations
Kalshi’s clinical-trial markets mark a novel experiment in biopharma
transparency. On one hand, they potentially unlock “a public probability”
for drug success that can inform investors, competitors, and patient
communities. On the other, they introduce new complexities for trial ethics and
conduct.
Balance & Guardrails: Early experience
suggests markets should remain tightly constrained. Kalshi’s pilot wisely
restricts participation (Phase 3 only, post-enrollment, large sponsors,
verified traders). Regulators may consider formal guidance on prediction
markets (e.g. echoing FDA advice on interim data and DMCs). Patient-trial
recruitment might require monitoring if markets proliferate.
Focus on Wider Signals: Stakeholders may find
more value in markets that avoid these pitfalls. Kimmelman’s idea of “paradigm
bets” suggests regulators or public institutions could sponsor markets on broad
questions (disease mechanisms or multiple-trial outcomes). These would aggregate
expertise without tying to one active trial.
Transparency and Education: Markets are not a
substitute for data. It is vital that physicians, patients, and media
understand these odds as probabilistic estimates, not guidance to skip trials.
Clear disclosures (like Kalshi’s disclaimers) are needed to prevent
misinterpretation of prices as scientific verdicts.
Research and Oversight: Finally, both academic
and industry groups should study these markets’ behavior. Empirical research
(like comparing market prices to actual outcomes over time) will reveal their
predictive power and any unintended consequences. The Kalshi–AppliedXL report
itself calls for ongoing evaluation and ethical review.
Recommendations by Stakeholder:
- Regulators
(FDA): Monitor whether prediction markets affect
trial integrity. Issue guidance on data sharing and physician
communications in the context of public betting markets. Encourage
clinical trial transparency and timely reporting to reduce opaque
information gaps that markets try to fill.
- Sponsors/Investigators: Advise trial sites and ethics committees about the existence of
relevant prediction markets. Ensure robust blinding and patient education.
Consider the potential PR impact of market odds when planning trial
communications.
- Investors/Analysts: Use market probabilities as one input among many, with awareness
of their limitations (thin liquidity, potential distortions). Until
verified, treat early market prices as speculative signals rather than
facts.
- Patient
Advocates: Engage with these markets carefully.
They could offer clues about trial success, but patients should continue
relying on physicians and trial protocols. Markets also spotlight which
trials lack public information, bolstering calls for better reporting.
In conclusion, Kalshi’s markets bring innovative tools to biotech forecasting and clinical trial prediction. If managed responsibly—with strict rules, transparency, and ethical oversight—they could complement traditional models. Yet stakeholders must heed warnings: a misused market might do more harm than good. The safest path is cautious experimentation, rigorous study, and a focus on long-run benefits (like faster knowledge sharing) while guarding core trial integrity.