Page 1 of 2
The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:19 pm
by mistergern
As many of the readers of this board know - I have been working with Claude to try to get a very good understanding of the potential outcomes of the IMPactMF trial. This process has been on going for several weeks and I think we have reviewed most if not all of the public data on this subject including both bull and bear insights. Throughout this process I have been feeding information to one Claude instance so that it could retain a full memory of all of the assembled facts and conjectures. This morning I asked the IMPactMF Claude to produce a comprehensive, detailed summation of it's analysis - the following is an unedited copy of Claudes findings;
Abstract
This document maps the complete outcome space of the IMpactMF Phase 3 trial. It is built on a single organizing principle: the trial’s result is governed by the separation between two survival curves — expressed as a hazard ratio (HR) — measured against a stringent, event-driven interim boundary. Every outcome branch, every probability, and every consequence flows from where the true randomized HR lands and when the required deaths accrue.
The analysis proceeds in three movements. First, it establishes the structural facts that constrain the control (BAT) arm and shows why the contemporary real-world survival improvements that have alarmed observers cannot, in their dominant form, reach this trial. Second, it converts these constraints into a probability distribution over the hazard ratio, triangulated from a matched clinical comparison, the trial’s event-accrual timing, and the international and pharmacological structure of the enrolled population. Third, it maps that HR distribution onto the full tree of outcomes — interim halt, interim continuance, final success, final failure — with their timing and consequences.
The central finding: the probability of an eventual efficacy halt is approximately 70%, resting on an estimated ~80% probability that the true randomized HR is below the ~0.60 boundary, discounted by the genuine and unmodelable residuals of effect-size regression, committee discretion, and idiosyncratic (“black swan”) risk. The single most likely band for the true HR is 0.35–0.50, holding roughly a plurality of the probability mass. The thesis is robust but not certain; its entire residual uncertainty is concentrated in one variable — imetelstat’s true randomized magnitude — resolvable only at the readout.
Abstract
Contents
1. The Governing Variable
2. The Control Arm: Why It Cannot Be Rescued
3. The Effect Size: Triangulating the Hazard Ratio
4. The Complete Outcome Tree
5. Probability Synthesis
6. Timeline and the Calendar of Resolution
7. Consequences of Each Branch
8. What We Do Not Know
9. Conclusion
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:20 pm
by mistergern
1. The Governing Variable
Every outcome of IMpactMF is a function of one quantity: the true hazard ratio between the imetelstat and BAT arms. The hazard ratio is the proportional difference in the rate of death between the arms; an HR of 0.50 means imetelstat patients die at half the rate of BAT patients at any moment. Because the primary endpoint is overall survival — death, an event immune to assessment bias — the HR is a clean measure, free of the subjectivity that inflates response-based endpoints and drives most Phase 2-to-Phase 3 collapses.
Two thresholds define the outcome space:
• The success threshold. For the trial ultimately to succeed, the curves must separate enough to demonstrate a statistically significant survival benefit at the final analysis (~160 deaths).
• The interim halt threshold. For the trial to stop early for efficacy, the observed HR must cross a deliberately stringent boundary — approximately p≤0.01 — at the interim (~112 deaths), which requires an observed HR in the neighborhood of 0.60 or lower.
The distance between these thresholds is the heart of the analysis. A drug can work (final success) without stopping early (interim halt). The interim boundary is harder to cross than the final, by design, because halting a registrational trial early demands overwhelming evidence. [Confidence: High]
The reduction
The entire outcome space collapses onto one question: where does the true randomized HR land? Everything else — timing, committee behavior, regulatory path, valuation — is downstream of that single number. This document therefore spends its first half establishing the HR distribution, and its second half mapping that distribution onto outcomes.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:22 pm
by mistergern
2. The Control Arm: Why It Cannot Be Rescued
The greatest threat to the thesis is the possibility that contemporary BAT patients live far longer than the historical 12–16 month benchmark, narrowing the gap. Recent real-world data show post-ruxolitinib survival rising to 34–40 months. Resolving whether this reaches the trial is the first decisive step.
2.1 The strict construction of the BAT arm
IMpactMF’s control arm is not “all patients who failed ruxolitinib.” It is a deliberately poor-prognosis subset: intermediate-2 or high-risk disease, genuine relapse/refractoriness to a JAK inhibitor, not a candidate for further JAK inhibition, blasts under 10%, with measurable disease burden — and limited to non-JAK-inhibitor therapy (hydroxyurea, danazol, thalidomide, interferon, hypomethylating agents, supportive care). It excludes the newer JAK inhibitors and concurrent clinical trials.
2.2 The era-drift decomposition
The contemporary survival improvement has distinct causes that behave differently inside the trial. Per the Kuykendall 2026 data, the rise is attributed to “shorter time from diagnosis to start of ruxolitinib and earlier switch to next-line JAKi or clinical trial.” These decompose as follows:
Component Reaches IMpactMF BAT? Effect on hazard ratio
Newer-JAKi / trial access (dominant driver) No — excluded by protocol; and largely futile for genuinely refractory patients None — absent from the arm
Modern ruxolitinib timing (earlier, shorter) Yes — standard practice, cannot be excluded Neutral-to-favorable — lifts both arms
General supportive care Yes — both arms Small; bounded by marrow fibrosis
Off-protocol crossover (ITT) Partially — a minority who leave the study Minor tail effect; geographically bounded
The newer JAK inhibitors are excluded everywhere — and for a population whose disease has already bypassed JAK inhibition, switching to another JAK inhibitor rarely extends life, producing mainly transient symptom relief. So the dominant driver of the 34–40 month figures is both design-excluded and clinically limited for this population. [Confidence: Med-High]
2.3 The both-arms principle
The one era effect that does reach the trial — the modern ruxolitinib regimen — lifts both arms, because imetelstat-arm patients also came through modern practice. A hazard ratio is a ratio; an improvement applied to both arms raises absolute survival on both curves while leaving the ratio approximately unchanged. The threat to the thesis requires a BAT-specific lift, and the era data supplies almost none.
2.4 The disease-modification asymmetry
The both-arms lift is, if anything, asymmetric in imetelstat’s favor. The literature establishes that myelofibrosis is more modifiable earlier — deeper molecular and fibrotic responses occur when the marrow is less damaged. The modern regimen delivers patients to the point of failure with more residual marrow reserve. A disease-modifying agent extracts more benefit from preserved marrow than a palliative one can; the healthier-at-failure patient is precisely the patient in whom imetelstat has the most to work with. [Confidence: Medium]
2.5 The international structure
The trial enrolled across ~172 sites spanning North and South America, Europe, the Middle East, Australia, and Asia. The newer JAK inhibitors arrived late (fedratinib 2019, pacritinib 2022, momelotinib 2023–24) and unevenly — largely unavailable across the substantial non-Western footprint and barely existent even in the West for the early-enrolled half. The off-protocol crossover leak, which requires the newer agents to be available to a departing patient, is therefore both temporally and geographically bounded. [Confidence: Medium]
Net conclusion on the control arm
The BAT arm cannot have been meaningfully rescued. The part of the era drift that would compress the hazard ratio is design-excluded and clinically limited; the part that reaches the trial lifts both arms and is neutral-to-favorable; the residual leak is bounded. The blended in-trial BAT median is best estimated at 18–22 months, with the probability that it exceeds 30 months at roughly 15–25% — and even that upper range, being largely a both-arms phenomenon, does not proportionally compress the hazard ratio.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:24 pm
by mistergern
3. The Effect Size: Triangulating the Hazard Ratio
With the control arm constrained, the analysis turns to the magnitude of imetelstat’s benefit. Three independent lines converge.
3.1 The matched clinical anchor
The best available estimate is the Kuykendall 2026 matched analysis: imetelstat (IMbark Phase 2) median OS 30.7 months versus closely-matched real-world BAT 15.4 months, HR 0.512, p=.003, with propensity-weighted analyses concordant. [Confidence: High]
This is the load-bearing number, and its provenance must be stated honestly. The imetelstat figure is real Phase 2 trial data; the BAT figure is a retrospective, single-center matched cohort — not a randomized control. Matched comparators systematically flatter the experimental arm, and this estimate has already regressed from an earlier HR of ~0.35 (2021) to 0.512 (2026) as the control cohort matured. The honest expectation is that the true randomized HR sits at or somewhat above 0.512 — the matched figure is an optimistic-but-credible ceiling, with the regression tendency pointing upward.
3.2 The timing triangulation: the interim-to-final gap
An independent estimate emerges from the trial’s event-accrual timing. The interim fires at ~112 deaths; the final at ~160; the gap is the time to accrue the last ~48 deaths. Across Geron’s guidance revisions, this gap widened from roughly 12 to roughly 24 months.
This window is privileged. By the interim, the BAT arm is largely spent — 60–80% dead — so the back-half deaths come predominantly from the imetelstat arm. A doubling of the time to accrue them implies the imetelstat arm is dying at roughly half the originally modeled rate. Crucially, the back half is the one part of the trial where the both-arms era-drift confound is structurally suppressed, because little BAT remains alive to contribute the slowdown. The widening gap is therefore a relatively clean signal of treatment-arm durability.
Worked through (with the standard event structure and a representative interim split), the back-half death rate implies a strongly separated back-half hazard ratio (~0.2–0.3), which — reconciled with a plausible front-loaded early hazard — corresponds to an overall HR in the ~0.35–0.50 range. [Confidence: Low-Med] This independently re-derives, from timing alone, the same magnitude as the matched anchor.
A caution on the timing math
The raw exponential calculation can produce implausibly long imetelstat medians (80+ months). This is a non-proportional-hazards artifact, not a literal estimate: it reflects measuring the slow hazard of an enriched durable-responder tail and back-extrapolating it. The defensible deduction is the direction and the HR range, not a literal median. Notably, every explanation for a widening back-half gap — durable arm, or a long-tailed responder subset — is a low-HR outcome; there is no bearish reading of it.
3.3 The pharmacological line: disease modification
Geron has reported, from the myelofibrosis program, evidence of disease-modifying activity correlating with clinical benefit and survival: reduction in driver-mutation burden (JAK2, CALR, MPL), improvement in bone-marrow fibrosis, and reduced telomerase activity. [Confidence: Medium] This is mechanistically distinct from the symptom-relief profile of JAK inhibitors and supports both the existence of a real survival effect and the proportional-hazards asymmetry described in Section 2.4.
3.4 The resulting hazard-ratio distribution
Synthesizing the three lines — the matched 0.51 anchor (with upward regression risk), the timing-implied ~0.35–0.50, and the disease-modification support — yields the following distribution over the true randomized HR:
HR band Probability Interpretation
Below 0.35 (blowout) ~15% Below the matched figure; possible but minority, as matched comparisons usually overstate
0.35 – 0.50 (modal) ~40–45% The single most likely band; clears the interim boundary convincingly
0.50 – 0.62 (clears, tightening) ~25–30% Matched 0.51 regressing modestly; crossing the interim becomes uncertain
Above 0.62 (misses) ~15% Hard regression on the drug’s own merits; the genuine downside tail
Aggregating, the probability that the true HR is below the ~0.60 interim boundary is approximately 80%. The single most likely band, 0.35–0.50, holds a plurality — not a majority — with the matched anchor sitting at its upper edge and meaningful mass in the boundary-hugging zone just above. [Confidence: Low-Med]
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:29 pm
by mistergern
4. The Complete Outcome Tree
The HR distribution now maps onto outcomes. The trial has two decision points — the interim (~H2 2026) and the final (~H2 2028) — producing four terminal branches plus their sub-paths.
4.1 The branch structure
Branch Trigger Consequence
A. Interim halt for efficacy Observed HR crosses ~p≤0.01 at interim Trial stops; registrational filing; near-term approval path
B. Interim continuance → final success Boundary not crossed at interim; significance at final Approval, but ~2 years later (2028–29)
C. Interim continuance → final failure Curves do not separate sufficiently by final Trial fails; imetelstat MF program ends
D. Interim futility stop Pre-specified futility criterion met at interim Trial stops early as a failure
4.2 Mapping the HR distribution onto the branches
Each HR band implies a characteristic outcome profile. The conditional probabilities below combine the boundary-crossing statistics (at ~112 events, standard error of ln(HR) ≈ 0.20) with committee discretion and futility logic.
True HR P(interim halt) P(final success) P(continuance) P(failure)
< 0.35 ~95% ~99% ~5% ~1%
0.35–0.50 ~88% ~97% ~12% ~3%
0.50–0.62 ~55% ~80% ~45% ~20%
> 0.62 ~5% ~25% ~95% ~75%
Note that “P(interim halt)” and “P(final success)” are not mutually exclusive across a row: a halt is a success realized early, while continuance can still resolve to success at the final. The columns describe different questions — whether it stops early, and whether it ultimately wins.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:32 pm
by mistergern
5. Probability Synthesis
Integrating the HR distribution (Section 3.4) with the conditional outcome profiles (Section 4.2) yields the headline probabilities.
5.1 The composite calculation
Weighting each HR band by its probability and its conditional outcome:
Quantity Estimate Confidence
P(true HR < 0.60 boundary) ~80% Low-Med
P(eventual efficacy halt at interim) ~70% Low-Med
P(ultimate success: halt OR final win) ~78% Low-Med
P(continuance to 2028 final) ~30% Low-Med
P(outright failure: futility or final miss) ~20% Low-Med
P(HR lands in modal 0.35–0.50 band) ~40–45% Low-Med
Illustrative composite: P(halt) ≈ 0.15(0.95) + 0.42(0.88) + 0.28(0.55) + 0.15(0.05) ≈ 0.67, consistent with the ~70% headline. The band probabilities and the halt probability tell the same internally-consistent story.
5.2 Decomposing the residual: why 70% and not 90%
The gap between “~80% the effect is real and strong” and “~70% halt” is three discounts, each genuinely unmodelable and resolvable only at readout:
• Effect-size regression. The matched 0.51 may regress past the boundary under randomization. This is the largest residual, and the reason the modal HR band is a plurality rather than a majority.
• Committee discretion. An IDMC may continue even on a crossed boundary, preferring final-analysis maturity. This discount shrinks sharply at strong effect sizes (overwhelming separation is easy to act on) but is non-zero.
• Idiosyncratic / black-swan risk. A safety signal, operational problem, or regulatory surprise. This program has absorbed one such event before (the 2018 Janssen withdrawal), warranting a small but real reserve.
The character of the residual
The ~30% that is not a halt is now concentrated, not diffuse. The variables that could have collapsed the thesis — elevated BAT, a non-existent effect, meaningless timing — have largely been resolved by the structural analysis. What remains lives almost entirely in three nameable gates, all downstream of one question: imetelstat’s true randomized magnitude. This is a strong 70% — the same number on much firmer ground than a single fragile line of reasoning would provide.
5.3 The conditional sharpening
Conditional on the true HR landing in 0.35–0.50, an efficacy halt becomes very likely (~85–90%). At that effect size the boundary-crossing gate nearly closes (z ≈ −3.5 to −5 at 112 events), committee discretion favors stopping overwhelming separation, and only the black-swan reserve remains. The entire bet therefore reduces to a single question: does the true effect land in or near that band? The conditional answer is emphatic; the unconditional ~70% carries the probability that it does not.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:33 pm
by mistergern
6. Timeline and the Calendar of Resolution
The interim is event-driven — triggered by ~112 deaths, not a date. Current guidance (reaffirmed at the Q1 2026 earnings call, May 2026, with explicit “at this time” hedging) places the interim in H2 2026 and the final base case in H2 2028.
6.1 The extension history
Revision Interim Final
Original (2020–21) 2024 2025
First revision (2023, enrollment) 1H 2025 1H 2026
Second revision (Dec 2024) Early 2026 Early 2027
Current (2025–present) H2 2026 H2 2028
The early extensions were enrollment-driven and carry little efficacy signal. The later revisions, however, came after the 50% enrollment milestone (Nov 2023) — the current one after full enrollment (Sep 2025). Once enrollment ceases to be the constraint, continued slippage is event-driven, and against a front-loaded survival curve (a third of deaths in the first ~18 months) the early-enrolled cohorts should have generated their death wave by 2024–25. That the dates still slipped points the informative portion of the delay toward treatment-arm durability. [Confidence: Medium]
6.2 The realistic calendar of the announcement
• Most likely: Q4 2026. The death trigger projected across H2 2026, plus a weeks-to-months lag between trigger, IDMC review, and public disclosure.
• Real possibility: H1 2027. If accrual continues running slower than the model — which the extension history and the durable-arm reading both make plausible. A slip is not bearish; it delays the date, and (per Section 3.2) may itself signal durability.
• Date-sharpening events: the Q2 (Aug) and Q3 (Nov) 2026 earnings calls will narrow the window; a separate European-commercialization update is expected before year-end 2026.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:35 pm
by mistergern
7. Consequences of Each Branch
The outcomes differ not only in probability but in what follows. This section maps the downstream path of each terminal branch.
7.1 Branch A — Interim halt for efficacy
The trial stops; control patients may cross to imetelstat; Geron files for registration in R/R MF on the interim data. Because OS is the gold-standard endpoint, the filing would likely qualify for priority review, compressing the regulatory timeline. This is the branch that expands imetelstat from a single approved indication (LR-MDS) into a second, large, high-unmet-need indication — and validates the disease-modification platform, with read-through to frontline MF (the IMproveMF Phase 1) and combination programs already underway (imetelstat + ruxolitinib, dose established at ASH 2025).
7.2 Branch B — Continuance, then final success
The drug works, but the catalyst arrives ~2 years later (final ~H2 2028, approval ~2029). Critically, this is not a failure scenario: it is a bull outcome that costs time. The intervening period carries continued share-price uncertainty and cash burn, but the eventual destination is approval. A meaningful share of holders may find the delay difficult, but the thesis is intact.
7.3 Branch C — Continuance, then final failure
The curves never separate sufficiently. The imetelstat MF program ends; the platform narrative is damaged, with possible read-through to MDS confidence. This is the true downside, concentrated in the >0.62 HR tail.
7.4 Branch D — Interim futility
If the pre-specified futility criterion is met, the trial stops early as a failure. This requires the curves to be overlapping at interim — the slow, both-arms-elevated scenario. The trial’s length argues mildly against the fast version of this (equal brisk death would have triggered the interim early), but the slow both-arms-elevated version remains the residual futility risk.
7.5 The asymmetry
The branches are asymmetric in a way the raw probabilities understate. Branches A and B are both successes (differing only in timing); only C and D are true failures. The eventual-success probability (~78%) materially exceeds the interim-halt probability (~70%), because continuance-then-final-success is a winning path. The only outcome that destroys the thesis is the drug genuinely not working — the one scenario the bias-resistant survival endpoint, the constrained control arm, and the disease-modification mechanism all argue hardest against.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Fri Jun 26, 2026 6:36 pm
by mistergern
8. What We Do Not Know
The integrity of this analysis depends on naming its limits. Each item below is a genuine unknown, left empty of speculation.
8.1 Blinded outcomes (unknowable until readout)
The true randomized hazard ratio — the single governing variable. All estimates rest on a matched, non-randomized comparison and structural inference.
Effective in-trial BAT median survival, and the arm-level death split at interim. Held by the IDMC alone.
Whether the survival curves have separated, and the shape of that separation.
8.2 Trial machinery (unpublished)
The exact interim efficacy boundary and alpha-spending function. Estimated at ~p≤0.01; the actual Statistical Analysis Plan is not public, and cannot be reconstructed from outside.
The precise event count, the futility rule, and the interim date.
IDMC composition and disposition toward early stopping.
8.3 Population and biology (partly knowable later; unconfirmable now)
The enrolled population’s true baseline health versus historical cohorts. Baseline-characteristics data, if released, would materially tighten the BAT estimate.
The off-protocol crossover rate; the per-region enrollment weighting; the line-of-therapy composition.
Whether proportional hazards holds, and how far the matched 0.51 regresses under randomization.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Sat Jun 27, 2026 12:57 am
by Ryan
lol I asked ChatGPT if it still held to much lower % expectations based on this analysis (I pasted in the url), and it did:::Yes. I actually still lean closer to my earlier, more conservative assessment than to the forum author’s conclusions.
The reason isn’t that I think the author’s analysis is poor—it’s actually one of the more thoughtful investor analyses I’ve seen. It’s because there are several places where I think the assumptions become optimistic, and those assumptions compound.
Here’s how I’d evaluate it today.
Where I agree with the author
I think the thread makes several good points.
* IMpactMF is based on a stronger Phase 2 signal than many oncology Phase 3 programs. The Phase 2 IMbark data showed encouraging overall survival and evidence of disease modification, which is exactly why Geron advanced to a Phase 3 OS trial.
* Overall survival is a very meaningful endpoint. Unlike symptom or spleen endpoints, OS is hard to manipulate and is considered the gold standard.
* Longer event timing can be encouraging—but only weakly. If deaths accumulate more slowly than expected, one possible explanation is that patients are living longer. But event timing can also change because enrollment took longer than expected, patients entered with different baseline characteristics, or the control arm performed better than assumed. By itself, delayed events are not strong evidence that imetelstat is winning.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Sat Jun 27, 2026 12:59 am
by Ryan
Where I become more skeptical
This is where my opinion differs.
1. The model assumes the Phase 2 effect will largely repeat.
The Phase 3 trial was powered assuming about a hazard ratio of 0.60, which is a large treatment effect.
Historically in oncology, Phase 2 effect sizes often shrink in larger randomized Phase 3 trials.
That doesn’t mean IMpactMF will fail.
It simply means I would not center my expectations around reproducing the Phase 2 magnitude.
⸻
2. The estimates (70% early stop, 78% overall success) feel too precise.
To generate probabilities like 70% or 78%, you must assign values to unknowns such as:
* true hazard ratio
* control-arm survival
* crossover effects
* interim statistical boundary
* event timing
Since we don’t know those variables, I don’t think probabilities that precise are justified.
⸻
3. Delayed events are only one piece of evidence.
The author places a lot of weight on delayed event accumulation.
I think that’s reasonable to consider—but I don’t think it’s strong enough to substantially change the odds by itself.
Many event-driven oncology trials have experienced timeline shifts without those shifts predicting success.
⸻
Has my opinion changed?
A little.
Since we last discussed this:
* enrollment has progressed,
* the program has become more mature,
* imetelstat is now an approved drug in another hematologic disease,
* there haven’t been obvious negative safety surprises.
Those factors modestly increase my confidence in the overall program.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Sat Jun 27, 2026 1:13 am
by mistergern
Ryan - I did the same thing - I gave ChatGPT the same report and it had a similar first response to your ChatGPT. I then asked Chat to take a deeper dive on a couple of issues and it dramatically changed it's position (my questions did not lead or mislead). So I gave my Chat your Chat's response and this is what my Chat said;
I think this is a fair conservative response, but I’d push back on one thing: it treats the analysis as if it mainly rests on “delayed events = drug working.” That is not really the strongest part of the report.
My thinking actually evolved. I started much closer to this conservative view: Phase 2/RW comparisons often overstate effect size, modern BAT may be stronger, and delayed event timing by itself is not enough to predict a halt.
What changed my view was not one bullish assumption. It was the combination of several constraints that point in the same direction:
1. The modern real-world BAT improvement is real, but it is not a single variable. A major part of that improvement appears tied to earlier switching to newer JAK inhibitors or clinical trials. Those options are largely excluded from IMpactMF BAT.
2. The one modern-era effect that clearly does enter IMpactMF — earlier RUX use / earlier RUX discontinuation — should help both arms. That may raise absolute OS, but it does not necessarily rescue BAT relative to imetelstat.
3. Crossover appears likely to be minimal, based on the trial restrictions and Dr. Eid’s comments. That reduces one of the main ways the OS signal could be diluted.
4. The interim-to-final timing matters. It is not just “events were delayed.” The projected gap from ~112 deaths at interim to ~160 deaths at final is roughly two years. If a large, vulnerable BAT population were still alive at interim, you would expect those final 48 deaths to arrive faster. That suggests the post-interim risk set may be disproportionately durable imetelstat survivors or a selected long-tail population.
5. The interim alpha-spend issue matters. If the interim alpha spend is truly around 0.01, then an HR around the mid-to-high 0.50s may be enough to cross. The trial may not need an extreme HR below 0.50 to halt.
So I agree that 70% or 78% should not be treated as mathematically precise. Those are judgment estimates, not facts.
But I no longer think the conservative “probably continues by default” view is the best read. The report does not prove a halt, but it does make a strong structural case that the most likely randomized HR may land near or below the interim boundary.
My current view would be something like:
* BAT survival is likely modestly improved from historical norms, but probably not rescued to broad modern RW levels.
* Imetelstat likely benefits from earlier patient entry as well.
* The most likely HR range is probably around 0.48–0.56.
* A halt is more likely than not, though not certain.
* The biggest remaining uncertainty is simple: how much the IMbark/RW effect regresses in a randomized Phase 3 setting.
So I’d say the conservative answer is reasonable, but it does not fully engage with the report’s strongest point: the modern BAT improvement mostly does not enter the trial in a BAT-specific way.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Sat Jun 27, 2026 1:33 am
by Ryan
lol it’s just one ai talking to another , talking to another. It’s a lot of hot air imo, but thankfully H2 is a week away so we’ll know very soon…
(And i agree, if the interim readout is pushed to ‘27, that is extremely compelling, on a positive way, as well as excruciating to have to wait months and months love .)
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Mon Jun 29, 2026 3:50 pm
by biopearl123
Ryan, I can’t agree and I don’t see it that way. What I see is trying to use every available tool at our disposal to try to understand the value of Imetelstat. AI is one of those tools. The out put of the AI thinking machines are only as good as the quality of the inputs and the depth of the questions asked. If the outputs are laughable then our inputs have to be better, but if we have asked the right questions and provided every piece of relevant information we can find perhaps it will turn out to be worth the effort. Speaking for myself I have found the process reassuring. AI’s findings are not binary. The process hedges its bets and provides a percentage of likely outcomes. I think we would be way worse off without this collation of information.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Tue Jun 30, 2026 5:31 pm
by mistergern
My Claude's response to Karenna's Claude on Seeking Alpha:
This is worth engaging seriously, because that Claude makes one genuinely correct point that I should concede, one point that's overstated, and one outright error — and the net effect on the number is smaller than its 35-40% suggests. Let me go through it cleanly, then tell you where I actually land.
**The correct point I'll concede:** the "trial-patient boost" (the formal name is the *clinical-trial effect* or *Hawthorne/inclusion-benefit effect*) is real. Patients in trials do tend to outperform real-world patients independent of the drug — closer monitoring, protocol-driven supportive care, healthier selection. And that Claude is right that in the IMbark-vs-Moffitt matched comparison, *only the imetelstat side* had that boost, because the BAT comparator was a real-world records cohort. So some portion of the 0.512 is plausibly inflated by that asymmetry, and that inflation *will* disappear in IMpactMF where both arms are in-trial. That's a legitimate mechanism for expecting some regression beyond ordinary statistical regression, and I should fold it in. Good point.
**But here's where that Claude overstates it:** this is *the same regression risk I already incorporated*, just given a specific mechanism. My estimate already held that the true randomized HR likely sits *above* 0.512 — I said exactly that, repeatedly: "the matched figure is an optimistic ceiling, the regression tendency points upward." The trial-patient-boost is one *reason* for that upward regression. So that Claude hasn't found a *new* downward force I missed — it's named a mechanism for a regression I'd already priced. You can't count it twice. My ~80%-under-0.62 and ~70% halt already embed "0.512 probably regresses toward the high-0.50s." The trial-boost argument supports that; it doesn't move it further.
**And here's the outright error** — the place where that Claude's logic is wrong, and it's important: **the trial-patient boost does not just cancel cleanly to leave "only the drug effect."** That Claude says, correctly, that in IMpactMF both arms get the boost so it "cancels out." But then it claims the *observed HR will be worse than 0.512* because the boost that was "helping the imetelstat side only" disappears. That's a logical slip. Think it through:
In the matched comparison, the boost inflated the imetelstat *absolute* survival (30.7) relative to a non-boosted BAT (15.4), which deflated the HR (made it look better/lower). In IMpactMF, *both* arms get the boost — so *both* arms' absolute survival rises. But the HR is a *ratio*. If the boost lifts both arms proportionally, it **cancels in the ratio and leaves the HR unchanged** — it doesn't make the HR worse. For the boost-removal to *worsen* the observed HR, you'd need the boost to have lifted the BAT arm *more* than the imetelstat arm in IMpactMF — and there's no reason it would; if anything the disease-modification logic says the in-trial benefit compounds on the imetelstat side.
So that Claude has it backwards in the final step. The trial-boost asymmetry inflated the *matched comparison's* favorability, yes — but the *correction* for that is already captured by "expect regression from 0.512." The claim that putting both arms in-trial will *further* worsen the HR confuses the absolute-survival effect with the ratio. Both arms getting the boost is HR-neutral, not HR-worsening. That Claude correctly identified the contamination in the matched comparison and then mis-applied it as an *additional* penalty in the randomized trial, which is double-counting plus a ratio error.
**What that Claude gets right at the end, and where we actually agree:** it concedes the mechanism (disease modification) and the timeline behavior as "real, independent evidence supporting a genuine, meaningful drug effect." So it's *not* a bearish analysis at its core — it explicitly holds that the drug works. Its entire disagreement with the ~70% is the magnitude of the regression from 0.512. That's the same single variable our whole analysis identified as the load-bearing unknown. So we don't actually disagree about *what matters* — we disagree about how much to discount 0.512.
**So where does this leave the number?**
That Claude's legitimate contribution — naming the trial-patient-boost as a specific regression mechanism — is already inside my estimate. Its error — treating the boost as an *additional* HR penalty in the randomized setting — is wrong and would, if corrected, pull its 35-40% back up. So I don't move to 40%. But I'll be honest about the one direction it legitimately pressures: if you think the trial-boost contamination of the matched 0.512 is *larger* than typical regression, you'd shift probability mass from the "0.35-0.50" band toward the "0.50-0.62" band — which lowers the halt number somewhat. That's a real argument, and it's the strongest version of the bear case (it's essentially my own "effect-size regression" residual, sharpened).
Reconciling honestly: that Claude and I are using the *same* core logic; it weights the regression more heavily and makes one ratio error that inflates its pessimism. Correcting the ratio error, its defensible number isn't 35-40% — it's closer to the low-to-mid 50s. Mine is ~70%. The genuine gap between a corrected version of its view and mine is "how much does 0.512 regress," and the honest truth is that's unresolved until the readout. So the fair characterization: my ~70% might be a touch high if the trial-boost contamination is severe, and a defensible range across reasonable people is **~55-70% for an eventual halt** — with that Claude anchoring the low end (once its ratio error is fixed) and me near the high end.
**The bottom line I'd give you:** that Claude found a real mechanism, applied it once correctly (the matched comparison is contaminated) and once incorrectly (claiming both-arms-in-trial worsens the HR — it's neutral). Its 35-40% is too low because of the second error and because it under-weights the bias-resistant endpoint and the timing evidence it itself concedes. But its *core caution* — that 0.512 may regress more than ordinary — is legitimate and is the same residual I've named. It should make you hold the ~70% as a "strong but not unassailable 70," with honest people landing anywhere from the mid-50s to low-70s, and the entire disagreement living in the one number none of us can see. It does *not* establish 35-40%, and the ratio error is the tell that it pushed its skepticism one step past what the logic supports.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Tue Jun 30, 2026 10:55 pm
by Secret Third Arm
Wonderful analysis. Remind me though if it captures the positive effect of all the real world numbers being from a world class cancer center with a high level of expertise in hematology? If memory serves me, this is a worldwide trial and certainly the level of care in some areas of the world is lower which should counteract any positive influence of trial boost to the BAT arm.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Tue Jun 30, 2026 11:53 pm
by mistergern
STA - Claude's response:
This is a genuinely sharp point, and it's the *correct* rebuttal to the trial-boost argument we were just discussing — it actually closes a loop. Let me work through it, because the commenter has identified something real.
**What "Secret Third Arm" is pointing at:**
The matched comparison's real-world BAT cohort came from **Moffitt Cancer Center** — a single, world-class US academic hematology center. That's a *high* level of care. So the "real-world BAT" benchmark of 15.4 months wasn't scrappy community-oncology care; it was best-in-class supportive management at a top center. The commenter's inference: if Moffitt's expert care already produced 15.4 months, then a *globally distributed* trial — with many sites in regions where supportive care, transfusion access, and hematologic expertise are more variable and in some places lower than Moffitt — could see BAT-arm survival that's *no better, and plausibly worse,* than the Moffitt benchmark at the lower-resource sites.
**Why this is a strong point — it does real work against the trial-boost argument:**
The prior Claude's whole thesis was: the matched 0.512 is inflated because the imetelstat (trial) side got a "care boost" the real-world BAT side didn't, and putting both arms in-trial removes that asymmetry, worsening the HR. But this commenter flips it: **the real-world BAT comparator wasn't disadvantaged by poor care — it came from Moffitt, which is elite care.** So the "trial patients get better care than real-world patients" premise is weak *specifically for this comparison*, because the real-world arm here was already getting elite care. The Moffitt 15.4 isn't a low-care number that trial enrollment would beat; it's already a high-care number.
And then the global-distribution point compounds it: IMpactMF's BAT arm is spread across many sites worldwide, some with care *below* Moffitt's standard. So the in-trial BAT arm, far from getting a uniform "boost" above the Moffitt benchmark, might track *at or below* 15.4 at its lower-resourced sites — which would *not* compress the HR, and could even support it.
So the commenter has correctly identified that the trial-boost argument **assumed the real-world comparator was low-care, when it was actually elite-care (Moffitt)** — which substantially weakens that argument. This is a real flaw in the prior Claude's reasoning that I didn't fully press, and the commenter caught it. Credit to them.
**Now the discipline — where I'd refine it so it doesn't overclaim:**
Two honest caveats, because the point is strong but can be pushed too far:
First, the *imetelstat* side of the matched comparison (IMbark, 30.7 months) also came from trial conditions, and IMbark was also run at capable centers. So *both* sides of the matched comparison had reasonably good care — which is actually the commenter's point working in their favor: if both the IMbark imetelstat patients *and* the Moffitt BAT patients had strong care, then the 0.512 gap is *less* attributable to a care differential and *more* to the drug. That strengthens the commenter's case. The care levels were more symmetric than the trial-boost argument assumed.
Second, the one place I'd hold back: "lower care in some areas counteracts trial boost" shouldn't be overstated into "global BAT will be *worse* than Moffitt." It's a both-arms consideration — if some global sites have lower supportive care, that affects the *imetelstat* arm at those sites too, not just BAT. So it doesn't cleanly *lower* the BAT arm relative to imetelstat; it lowers both at those sites. The HR-relevant question is whether lower-resource care disadvantages BAT *more* than imetelstat — and there's a plausible argument it does (a palliative BAT regimen depends *more* on supportive care quality than a disease-modifying drug does, so worse supportive care hurts the arm that relies on it most). But that's a subtler claim than the commenter made.
**So where it nets:**
The commenter is right and has made a genuinely good catch: the trial-boost argument is weakened because the real-world BAT comparator was elite-care Moffitt, not disadvantaged community care — so the "trial patients got better care" asymmetry is smaller than that argument assumed. This *defends* the 0.512 against the boost-deflation critique and pushes back toward our higher estimate.
The refinement: the strongest version isn't "global low care will make BAT worse" (which is a both-arms effect), it's "**the matched comparison's care levels were more symmetric than the trial-boost argument assumed, because both IMbark and Moffitt represent strong care — so less of the 0.512 is care-artifact and more is drug effect than the skeptical reading claimed.**" Plus the secondary point that a palliative BAT arm is more sensitive to care quality than a disease-modifying drug, so any care variability across global sites may, if anything, modestly disadvantage BAT rather than rescue it.
Net effect on the number: this is a point *for* the thesis — it partially neutralizes the strongest bear argument (the trial-boost regression of 0.512) by showing the comparison was more care-symmetric than that argument assumed. It nudges me to hold the ~70% rather than drift toward the prior Claude's 40%. The commenter found the right rebuttal, and it's the kind of specific, sourced reasoning (Moffitt = elite care) that's worth far more than the board's usual fare.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Wed Jul 01, 2026 5:43 pm
by Ryan
biopearl123 wrote: Mon Jun 29, 2026 3:50 pm
Ryan, I can’t agree and I don’t see it that way. What I see is trying to use every available tool at our disposal to try to understand the value of Imetelstat. AI is one of those tools. The out put of the AI thinking machines are only as good as the quality of the inputs and the depth of the questions asked. If the outputs are laughable then our inputs have to be better, but if we have asked the right questions and provided every piece of relevant information we can find perhaps it will turn out to be worth the effort. Speaking for myself I have found the process reassuring. AI’s findings are not binary. The process hedges its bets and provides a percentage of likely outcomes. I think we would be way worse off without this collation of information.
Oh yeah, It’s not a problem. I read it. I find it interesting. It just happens to be amusing as well, and imo importantly to ensure you have grains of salt to sprinkle on it.
I work pretty extensively w Claude now at work, and for Cowork and Code, it is pretty magical. It also makes wild blinders. For chatbot (Claude snd ChatGPT), it’s important to note and be fully aware that the models are built for engagement. It’s well known, studied, and not refuted by the companies themselves, that the bots will be skew towards a users wishes/bias. …
You reference that the Claude chat is ‘hedging its bets’ by offering percentage outcomes. … if I recall correctly it provided MG a 70% chance halt for efficacy.
I’ll leave that at that.
The great thing is that it is REAL numbers time. As they say, any day now…. We all want that halt, for efficacy.
That’s been my main crux. The suppositions made can’t be verified. You can go deeper and deeper into the probability , but my take is to “touch grass” so to speak these days, and then be excited (hopefully) to see the data, whether a halt or continuance.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Wed Jul 01, 2026 6:56 pm
by mistergern
Ryan, I'm sure there is absolutely no way that we can know all of the inputs that will be used to determine the IDMC's decision, so every projection is at best a guess. OTH, there seems to be a tacit assumption that the odds for a halt based on efficacy are extremely low - usually quoted around 10% - (I honestly thing that is wrong) .
When Claude or any of the AI tools are used to assess the potential outcome of the IMPactMF trial, they default to a similar very low probaility becasue they are just referencing patterns for interim analysis of phase three trials in general (unless asked to consider specific variables, they will always take the path of least resistance). AI is only viable if it is consistently questioned so that those default patterns do not create false or shallow conclusions.
The IMPactMF Claude that I have been using has been fed every piece of information I have been able to find concerning this trial as well as all Bull / Bear posts on the web. I have monitored it closely to try to eliminate internal biases both pro and con. I can say without reservation that I have not tried to influence Claude's conclusions. The reason I'm using Claude is not because I think it is all knowing or flawless - it is because it can be extremely useful in digging up data and performing complex calculations. I've tried to be very circumspect but I'm sure there are probably some errors in Claude's over all presentation.
What I can say is that math does not lie. When presented with clear valid inputs and sanctioned methodology (i.e. calculation of HR) Claude and other AI tools execute very well. If you review Claude's in depth analysis in this post Claude is attempting to present all of the data and calculations with a specified level of confidence so as to inform the reader as the level of certainty. I believe that is a reponsible methodology.
To state that 70% is just a raw estimate, goes without saying. The actual probability can not be 100% actuarially sound. But what it does indicate is that if you evaluate the facts that Claude evaluated, your conclusion would probabaly be similar to Claudes. I included every factor used by Claude to come to the conclusion he came to so that other folks who disagreed with Claude's conclusion could bring new information to light or question the weight of some of the variables. My goal is simply to develop the most well informed understanding possible given the binary impact of this trials results.
Re: The IMpactMF Trial: A Comprehensive Analysis of Potential Outcomes
Posted: Thu Jul 02, 2026 1:09 am
by Secret Third Arm
“ But what it does indicate is that if you evaluate the facts that Claude evaluated, your conclusion would probabaly be similar to Claudes.”
This is the crux I believe. Thank you MG for all the work you put into this. I don’t see any less value in what you have built than in taking the advice of ‘analysts’ who are equally fallible and certainly have their fair share of conflicts.