Category: Position Papers

Reltronic position papers on diagnostic latency, data sovereignty and clinical evidence.

  • Evidence Standards for External Control Arms in Small Populations

    Sanjay Ahuja, PhD
    Reltronic, Inc.

    Where a condition is rare enough that randomization is infeasible or unethical, sponsors increasingly propose an external control arm. Regulatory receptivity has grown, but it remains conditional, and the conditions are frequently misunderstood.

    The regulatory frame is guidance, not rule

    Two documents govern. ICH E10, on the choice of control group, establishes that the control must be considered in the context of available standard therapies, the adequacy of the evidence supporting the chosen design, and the applicable ethical considerations. In the United States, the Food and Drug Administration issued Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products on 1 February 2023, with the comment period closing on 2 May 2023.

    One feature of that document is routinely overlooked in commercial discussion. It was issued in draft, states on its face that it is not for implementation, and contains non-binding recommendations. Sponsors designing against it are therefore designing against a stated direction of travel rather than a settled requirement. The practical consequence is that early engagement with the review division carries more weight in this design than in a conventional randomized study, because the acceptability of the comparison group is a matter for discussion rather than a matter already determined by rule.

    The three conditions, which must hold together

    External control is generally considered acceptable where randomization is not feasible or ethical and three circumstances hold together. The conjunction matters. Sponsors sometimes present a strong case on one condition and treat the remaining two as satisfied by implication, which is the most common structural weakness in these submissions.

    ConditionWhat must be demonstrated
    Disease severity and unmet needThe condition is life-threatening or seriously debilitating and no satisfactory alternative therapy exists.
    Predictable natural historyThe untreated course is well characterized and highly predictable, so the counterfactual outcome can be established with confidence.
    Effect size and temporalityThe anticipated effect is substantial and self-evident, and closely related in time to administration, such that it is implausibly attributable to confounding.
    Circumstances in which external control is generally considered acceptable.

    The second condition carries most of the weight and is the one most frequently overstated. A natural history described in the literature is not the same as a natural history characterized well enough to serve as a counterfactual for an individual patient. Where the condition is heterogeneous, or where the rate of progression varies substantially between patients, the requirement is not met by aggregate description however extensive.

    Where external controls come from in practice

    Published surveys of external controls used in regulatory decision making show a distribution that is informative in itself.

    SourceSharePrincipal limitation
    Retrospective natural history, including record review44%Ascertainment differs from trial conditions
    Baseline control, patient serving as own comparator33%Confounded by regression to the mean and secular trend
    Published data11%Patient-level covariates unavailable
    Data from a previous clinical study11%Eligibility criteria and endpoints rarely align
    Observed sources of external control arms.

    The largest category, retrospective review of medical records, is also the category in which the comparability problem is most acute and most tractable. It is acute because records are generated for clinical rather than research purposes, so the timing, frequency and definition of measurement differ systematically from the trial. It is tractable because the underlying source data exist and can, in principle, be interrogated to establish how far they differ.

    Six failure modes

    Most regulatory objections to an external control arm reduce to one of the following. None is primarily a question of statistical technique.

    1. Differential ascertainment. Outcomes in the trial arm are measured on protocol at fixed intervals; outcomes in the external arm are measured when a patient happened to present. Apparent differences in event timing may reflect observation schedule rather than biology.
    2. Outcome definition drift. The endpoint as defined in the protocol does not correspond to what was recorded in the source data, particularly where the external data span a period in which diagnostic criteria or coding practice changed.
    3. Unmeasured confounding at baseline. Covariates known to affect prognosis are absent from the external source, so no adjustment can be made and none can be shown to be unnecessary.
    4. Secular trend in standard of care. Supportive management improved between the period from which the external control is drawn and the trial period, so part of any observed benefit is attributable to era rather than to treatment.
    5. Selection into the external cohort. Patients appearing in a registry or specialist center record differ systematically from the trial population, frequently by severity, and the direction of that difference is not always predictable.
    6. Immortal time. The external control accrues follow-up from a point that is not equivalent to randomization, creating survival time in one arm with no counterpart in the other.

    What to establish before construction, not after

    The recurring error in these programs is sequential rather than analytical. A comparison group is constructed, differences are discovered, and statistical adjustment is applied to close them. Adjustment can correct for measured imbalance. It cannot correct for imbalance in what was measured, or when, or by whom, and it is that class of difference which generates most objections.

    The following eight questions are answerable before a comparison group is built, and considerably more expensive to answer afterward.

    1. Is the untreated course of this condition characterized at the level of the individual patient, or only in aggregate?
    2. Over what period were the external data generated, and what changed in supportive care across that period?
    3. How was the endpoint ascertained in the external source, at what frequency, and by whom?
    4. Which baseline covariates known to affect prognosis are present in the external source, and which are absent?
    5. What proportion of the external cohort would have met the trial’s eligibility criteria had they been applied?
    6. From what index point does follow-up accrue in each arm, and are those points equivalent?
    7. Can the derivation of every value in the comparison group be traced to a source record and reproduced independently?
    8. Has the review division been consulted on the proposed comparison group, and at what stage?

    The question is provenance, not statistics

    External control is a legitimate design in the circumstances the guidance describes, and in a sufficiently rare condition it is frequently the only design available. Its acceptability turns on a narrower question than is generally supposed. Not whether the comparison group can be adjusted to resemble the trial arm, but whether what was measured in each, and when, and under what definition, is comparable in the first place.

    That question is one of provenance rather than of statistics. It is settled by documentation of how each value came to exist, which is available at the point the comparison group is designed and progressively less available thereafter.

    Sources

    1. United States Food and Drug Administration. Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products. Draft guidance for industry, issued 1 February 2023; comment period closed 2 May 2023. Draft status, non-binding recommendations.
    2. International Council for Harmonisation. ICH E10, Choice of Control Group and Related Issues in Clinical Trials.
    3. The Use of External Controls in FDA Regulatory Decision Making. Therapeutic Innovation and Regulatory Science, 2021.

    Disclosure: the author is affiliated with Reltronic, Inc. This paper makes no claim regarding any Reltronic product and describes no proprietary method. It is not regulatory advice; sponsors should engage the relevant review division directly on the acceptability of any proposed comparison group. The principal United States guidance cited was in draft at the time of writing and its status should be verified before reliance.

  • Data Sovereignty Is an Architecture Decision, Not a Contract Decision

    Reltronic, Inc.
    For information technology, security, privacy and legal review

    The legal basis for moving European health data to the United States has been rebuilt twice in ten years and is under appeal a third time. Clinical infrastructure procured today will still be running when that appeal is decided. An institution choosing where analysis happens is therefore making a decision with a longer life than the legal instrument that currently permits the alternative.

    What the European Health Data Space actually requires, and when

    Regulation (EU) 2025/327, establishing the European Health Data Space, entered into force on 26 March 2025. A good deal of confusion follows from that date, because entry into force and application are not the same thing. Most of the substantive obligations are not yet live, and they arrive in stages over the following decade.

    • 26 March 2025. The Regulation enters into force.
    • 26 March 2027. General application. The primary use provisions and the requirements on electronic health record systems begin to bite.
    • 26 March 2029. Chapter III applies. This is the secondary use regime, covering most categories of electronic health record data.
    • 26 March 2031. Sharing obligations extend to medical imaging studies and imaging reports, medical test results and discharge reports, and to previously excluded categories including genetic, molecular and clinical trials data.

    The practical consequence is straightforward. A clinical system selected in 2026 and deployed over the following year will be operating inside the general application regime almost immediately, and inside the secondary use regime well within its service life. Institutions assessing infrastructure against today’s obligations alone are assessing against the wrong set.

    The transfer question has been reopened twice already

    For any architecture that moves patient data outside the jurisdiction where it was collected, the governing question is whether a lawful transfer mechanism exists. That question has an unusual history.

    The Safe Harbor arrangement was invalidated by the Court of Justice of the European Union in 2015. Its replacement, Privacy Shield, was invalidated by the same court in 2020. The current instrument is the EU-US Data Privacy Framework, adopted by adequacy decision on 10 July 2023. It was challenged, and the General Court dismissed that challenge on 3 September 2025, upholding the Framework. The applicant appealed to the Court of Justice on 31 October 2025 in Case C-703/25 P, and that appeal is pending.

    The Framework is valid law today. It remains in force unless the Commission repeals it or the Court annuls it, and nothing here should be read as advice that transfers under it are presently unlawful, because they are not. The point is narrower and more useful than that. Two of the last three instruments governing this question were struck down, the third is under appeal, and separate representations to the Commission during 2026 have questioned whether identified deficiencies can be remedied at all. An institution whose clinical architecture depends on the outcome is carrying a risk it did not choose and cannot price.

    Contracts allocate liability. They do not move data.

    The standard institutional response to this exposure is contractual. Standard contractual clauses, data processing agreements, business associate agreements, transfer impact assessments. Each is necessary, and none of them changes the underlying fact pattern.

    A contract determines who is answerable when something goes wrong. It does not determine where the data physically sits, which jurisdictions can compel its production, or what happens to the arrangement if the adequacy decision underneath it is annulled. When Privacy Shield fell in 2020, the institutions affected did not have a contract problem that better drafting would have prevented. They had an architecture that assumed a legal instrument, and the instrument was withdrawn.

    This is the distinction that matters during procurement. Contractual controls are remedies. Architectural choices are preventions, and only one of the two survives a change in the law.

    What changes when the analysis moves to the data

    Where computation happens inside the institution that already holds the record, several questions that dominate a cloud assessment do not arise at all. There is no onward transfer to assess, because nothing is transferred. There is no additional processor to enter on the register, because no third party processes the data. There is no cross-border question, because no border is crossed. There is no dependency on an adequacy decision, because no adequacy decision is engaged.

    Institutions frequently expect an installed system to be reviewed more heavily than a hosted one, on the reasonable intuition that anything on the network is the institution’s problem. In practice the review is usually narrower, because the questions that consume the most time in a cloud assessment are the transfer, subprocessor and jurisdiction questions, and those questions have no subject matter when the data does not move.

    This is not an argument that installed architecture is superior for every purpose. It is an argument that the two are assessed against different question sets, and that institutions which apply a cloud checklist to installed infrastructure will overstate the work involved, while institutions which apply an installed checklist to a cloud service will understate it.

    Questions worth asking during procurement

    The following separate architectures that hold data in place from those that move it and manage the consequences contractually.

    • Where does computation occur, physically, and in which legal jurisdiction does that location sit? Not where the data is stored at rest, but where it is processed.
    • Which entities other than the institution can technically access identifiable data, irrespective of whether they are contractually permitted to?
    • If the applicable adequacy decision were annulled next year, what would have to change in the deployment, and how long would that take?
    • Does the vendor require any standing access to the environment, and if access is granted for support, is it scoped, time-limited and logged in a record the institution controls?
    • Who holds the encryption keys, and can the vendor decrypt without the institution’s participation?
    • Under the EHDS secondary use regime arriving in 2029, which party is the data holder, and what obligations does the architecture place on the institution that a different architecture would not?

    The decision has a longer life than the instrument

    Clinical infrastructure is not replaced often. A system selected now will plausibly still be running in 2031, when the final tranche of EHDS sharing obligations reaches imaging, test results and genetic data. Over that period the transfer mechanism underpinning cross-border processing may be upheld, replaced, or struck down for a third time. None of those outcomes can be predicted, and none needs to be if the architecture does not depend on the answer.

    The question for an institution is not which arrangement is compliant today. It is which arrangement remains compliant without renegotiation if the legal position changes, because the legal position has changed twice already within the service life of systems still in use.

    Sources

    1. Regulation (EU) 2025/327 establishing the European Health Data Space. Entry into force 26 March 2025.
    2. European Commission implementing decision on the adequacy of the EU-US Data Privacy Framework, 10 July 2023.
    3. General Court of the European Union, Case T-553/23, judgment of 3 September 2025.
    4. Court of Justice of the European Union, Case C-703/25 P, appeal lodged 31 October 2025, pending.
    5. Court of Justice of the European Union, Case C-362/14 (2015) and Case C-311/18 (2020), invalidating Safe Harbor and Privacy Shield respectively.

    Disclosure: this paper is published by Reltronic, Inc. It makes no claim about any Reltronic product and describes no proprietary method. It is not legal advice, and institutions should take their own counsel on the application of these instruments to their circumstances.

  • Alert Fatigue Is a Patient Safety Problem, Not a Usability Complaint

    Manuel Rios, MD
    Reltronic, Inc.

    Clinical decision support alerts are overridden between 46 and 96 percent of the time. That figure is usually presented as evidence of a design problem, and it is one. But it is first a safety problem, because the alert that would have mattered arrives through the same channel as the ninety that did not, and it is dismissed by the same reflex.

    The number is not a satisfaction score

    A systematic review of alert fatigue measurement published in 2026 found average override rates across studies ranging from 46.2 to 96.2 percent. Severity does not protect against this. One study within that literature found that 88.2 percent of alerts classified as very severe drug-drug interactions were overridden. The alerts a system flags as most urgent are dismissed at close to the same rate as the rest.

    It is tempting to read those numbers as a workforce problem, as evidence that clinicians are not paying sufficient attention. The literature does not support that reading. When overridden alerts are independently reviewed for appropriateness, a large share of the overrides turn out to be clinically correct. Reported appropriateness ranges from 63.4 to 100 percent for drug-allergy alerts and from 27 to 87.5 percent for renal alerts. In many categories, the clinician who dismissed the alert was right to dismiss it.

    That reframes the problem entirely. Override behavior is not inattention. It is a rational response to a channel with a low base rate of useful signal. A clinician who has learned that most of what arrives in a given channel does not require action has learned something true about that channel, and has adapted correctly.

    Desensitization is dose-dependent

    The mechanism has been characterized. Ancker and colleagues studied 112 ambulatory primary care clinicians over a three and a half year period and found that clinicians became measurably less likely to accept an alert as the number of alerts they received rose. The effect was strongest for repeated alerts on the same patient.

    The implication is uncomfortable for anyone building these systems. Alert volume is not a neutral quantity that a well-designed interface can absorb. Volume is itself the mechanism of harm. Every low-value alert added to a system degrades the response to every alert already in it, including the ones that were working. A system that adds coverage without regard to specificity is not neutral. It is actively spending down a finite resource, and the resource is clinical attention.

    This is why the framing matters. Filed as a usability complaint, alert fatigue is a matter of preference, to be traded off against feature completeness and addressed at some later point in the roadmap. Filed correctly, as a patient safety problem, it becomes a design constraint that binds from the beginning. Overrides have been associated with medication errors and serious adverse events, including death, when clinically important information was inadvertently passed over.

    What follows for the design of these systems

    If clinical attention is treated as a fixed and exhaustible resource, several design positions follow that are otherwise easy to argue against.

    • Specificity is the binding constraint, not coverage. A system that surfaces fewer findings at higher positive predictive value is not a less capable system. It is a system that has made the correct trade.
    • Volume warrants an explicit budget. If every additional alert degrades response to the existing ones, then adding a new rule should require displacing an existing one, or demonstrating that the new rule clears a threshold the current set does not.
    • Within-patient repetition deserves particular scrutiny, since it is the pattern most strongly associated with desensitization in the published work.
    • Interruption should be the exception rather than the default. Information that a clinician can consult when it is relevant occupies a different cognitive category from information that stops work to be dismissed.
    • The reasoning should travel with the finding. A statement that can be checked against its supporting observations in a few seconds is a different object from a conclusion that must be either trusted or dismissed. Only the first kind can be evaluated rather than pattern-matched.

    The measurement problem underneath

    Most institutional reporting on decision support counts alerts fired and rules deployed. Neither quantity describes whether the system is helping. The systematic review noted that alert fatigue measurement is not consistently defined across the literature, which means institutions are frequently comparing figures that are not comparable.

    Two measures are more informative than volume. The first is the acceptance rate over time, which shows whether the system is being adapted to or attended to. The second is positive predictive value by alert category, which shows where a system is earning attention and where it is spending it. A vendor that can report the second figure has confronted the problem. A vendor that reports coverage instead has not.

    The question worth asking

    Institutions evaluating any clinical decision support system, including systems that describe themselves as using artificial intelligence, are entitled to ask one question before any other. Not what the system can detect, but what proportion of what it surfaces a clinician acts on, and how that proportion has moved over the period it has been running.

    A system whose acceptance rate is falling is not being adopted. It is being tuned out, and the interval between those two states is where the risk sits.

    References

    1. Alert fatigue measurement in clinical decision support: a systematic review. PubMed 42148822.
    2. Appropriateness of Alerts and Physicians’ Responses With a Medication-Related Clinical Decision Support System: Retrospective Observational Study. JMIR Medical Informatics, 2022.
    3. Ancker JS, Edwards A, Nosal S, Hauser D, Mauer E, Kaushal R. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Medical Informatics and Decision Making, 2017; 17:36.

    Disclosure: the author is affiliated with Reltronic, Inc. This paper makes no claim about any Reltronic product and describes no proprietary method. Every figure cited is drawn from the published literature listed above.