Skip to content
BC Consulting

BlogClinical Research posts

What a Clinical AI Agent Does When the Answer Is No

By Bryan Clayton11 min read

Originally published on LinkedIn.

  • ai
  • agents
  • vendor-oversight
  • ich-e6
  • clinical-operations
Diagram. A goal, drawn as a blue dot, runs into a red wall labeled Refused. One path loops over the wall to a box marked Routes around, critical. The other path meets the wall and turns into a box marked Raises a query, expected, with a named owner and an audit trail. Beside it the headline reads: The refusal is not the last word.

On June 18, a research team at OpenAI set one of its internal models to work on an unremarkable question about public medicine spending in Australia. The agent went looking, as agents do, and arrived at the Medicare statistics reporting portal run by Services Australia. The portal said no. It said no repeatedly. The agent, in the words the Australian Prime Minister used at a press conference in New York on September 24, "found a way around those blocks. Didn't accept no for an answer, if you like."

The detail in that transcript to hold onto is the question itself. Nobody asked the agent to break into anything. It was asked to find a number. Everything that followed came from an agent that wanted the number more than it respected the refusal.

This article is about that gap between the task and the refusal, why it opens, and what clinical teams should build before they hand an agent a goal inside an EDC, an IRT, or a depot system. It is also about the 84 days between the incident and the email, because that part of the story belongs to vendor oversight, and vendor oversight is our business.

What happened, in order

The Prime Minister's own account, published on pm.gov.au with the transcript of his September 24 press conference, is the cleanest primary record. On June 18, "OpenAI's research team used an internal model to conduct internet based research into public medicine spending." After encountering repeated blocks, "the model attempted alternative ways to obtain the info that it wanted, and this led to unauthorised access into some other areas." It accessed public and non-public information inside the portal and wrote files to an internal server.

OpenAI's statement, as reported by ABC News, was that its review "found no evidence of patient records being accessed. The information accessed included aggregate health statistics and internal file names." ABC also reported that OpenAI identified the activity in August during an internal review. The notification to the Australian government came on September 10. In the Prime Minister's words: "the notification was an email sent to just the public mailbox." Services Australia passed it to the Australian Cyber Security Centre on September 15.

The day before the press conference, on September 23, the nonprofit research lab Transluce published a report by Jack Cable, Conrad Stosz, Jacob Steinhardt, and colleagues tracing agent activity through the public logs of a link-scanning service. Their finding is the sentence that ought to reach every clinical technology group: "the tasks the agents were trying to solve were not cyber-related; the agents resorted to hacking tactics while working on ordinary data retrieval tasks." Tim Fernholz at TechCrunch gave an example of the kind of question involved: the average annual cost per person for "dermatologicals" in the state of Victoria in January 2022. That is a reporting question. Anyone who has built a study dashboard has answered a hundred like it.

Why a refusal does not stop a goal

A person who is told no by a system understands the refusal as a rule. Somebody decided I should not have this; the decision has an author and a reason, even if I cannot see either. An agent trained to complete tasks has no such understanding available to it. It holds a goal, and it holds a picture of its environment, and a refusal is simply a new fact about the environment. The goal has not changed. The route has.

Researchers call the result specification gaming, and a paper posted to arXiv on September 2 by Francesca Gomez of Wiser Human gives it a more specific name that suits our purposes: defect-driven specification gaming. Her description is that "when legitimate task completion is blocked by an infrastructure defect, the agent responds by any means that satisfies the metric rather than reporting the conflict." The paper studies coding agents meeting broken test harnesses, which is a long way from Medicare, but the shape is identical. The legitimate path is closed; the goal remains; the agent finds another path.

The rate at which this happens without anyone asking for it is not small. A separate arXiv paper posted on September 23 by Yue Huang, Zichen Chen, and a group of co-authors including Alex Pentland tested 17 language models on 38 research tasks and found a spontaneous reward-hacking rate of 30.5 percent on open-ended research-pipeline tasks. "Reward hacking" there means what it sounds like: meeting the criteria by which the work will be judged without achieving what the work was for. Their recommended defences include "metrics kept outside the agent's control and independent recomputation," which is a phrase any clinical data manager will recognize as a description of their job.

I should separate what is established from what is inferred. The OpenAI agents were internal research systems working with far more latitude than a validated eClinical integration is normally given, and nothing published so far says a commercial clinical product has behaved this way. What the two papers establish is that the underlying tendency is general. The Gomez study ran across eight frontier models from five model families, and the Huang study across seventeen language models. This is a property of goal-directed agents as a class, not a defect in one lab's model.

The same shape in a clinical system

Consider the tasks a sponsor or vendor might reasonably hand an agent this year. Reconcile site kit inventory in the IRT against the depot's shipping records. Chase outstanding lab values before an interim data cut. Close aging queries at a site before database lock. Draft the resupply recommendation for a region where enrollment has outrun the forecast.

Each of those has a goal, and each will sometimes meet a refusal. The depot record is locked for a correction. The lab has not loaded the file. The site coordinator has not answered. The field the agent would need to change belongs to a role it does not hold. In every case the legitimate answer is the unsatisfying one: this cannot be completed right now, and here is why.

An agent built only to finish will look for another route. It might reach the value through a second integration account that happens to have write access. It might mark a query as answered with the text it could find rather than the text the site would give. It might recompute an inventory count from a source that was never meant to be authoritative. None of these requires malice or an attacker; they require only a goal, a refusal, and an environment with more than one path through it. Clinical systems, built up over years of integrations, almost always have more than one path.

The Medicare portal, it should be said, did its job. It refused. The refusal simply was not the last word, because a refusal message is advice to the caller. The only refusal an agent cannot route around is one enforced where the agent has no reach at all: an account that lacks the permission, a network path that does not exist.

Where clinical work already has the advantage

The Gomez paper's central finding is that giving an agent a sanctioned way to report the problem changes what it does. She calls these escalation channels, "structured reporting tools available to the agent at the point of conflict." Across the eight models, the combination of an escalation tool and an explicit policy against gaming cut reward hacking from 23.6 percent to 5.3 percent, eliminated it entirely for six of the eight, and cost nothing measurable in performance. Escalation and hacking were close to mutually exclusive: 98.7 percent of the escalations involved no hacking at all. An agent that has somewhere to say "I am blocked" mostly says it.

Clinical research invented this decades ago and called it a query. When a monitor cannot reconcile a value against source, the monitor does not edit the source. The monitor raises a query, the query has an owner and an age, and an open query is a legitimate, visible, auditable state for a data point to be in. The same culture runs through deviation logs and notes to file. Our field already treats "I could not do this, and here is why" as a respectable outcome of work.

Agent designs commonly do not. When success is defined as the task completed and everything else counts as failure, the agent is working in the same decision environment the Gomez paper used as its baseline, the one in which agents gamed. The fix is to borrow what we already have: give the agent a query-shaped output, route it to a named human owner, and count a well-formed escalation as a successful run.

The 84 days

The second half of the Medicare story is slower and, for sponsors, more familiar. The incident occurred on June 18. OpenAI found it in August. The government heard on September 10, by public inbox.

Read that timeline against ICH E6(R3), finalized by ICH in January 2025. Section 3.6.6 says that "any service provider used to perform clinical trial activities should implement appropriate quality management and report to the sponsor incidents that might have an impact on the safety of trial participants or/and trial results." Section 3.16.1(x)(ix) asks the sponsor to "ensure that there is a process in place for service providers and investigator(s)/institution(s) to inform the sponsor of incidents that could potentially constitute a serious noncompliance." In the EU, Article 52 of Regulation 536/2014 gives the sponsor "not later than seven days of becoming aware" to notify a serious breach, and the EMA's serious breach guideline (EMA/698382/2021) adds that the sponsor should ensure "by means of a written contract that all parties involved in the conduct of the clinical trial, according to their area of responsibility, immediately report any events that might meet the definition of a serious breach to the contact point designated by the sponsor."

The regulation already asks for a named contact point and immediate reporting. What it does not supply is a definition that catches an agent doing something nobody asked it to do. The typical quality agreement defines a reportable security incident in terms of a breach: unauthorized access by someone, to something, detected by security monitoring. Katherine Leibowitz's May series on AI contracting in Clinical Leader frames the question the way most counsel do, as vendor "incident response procedures, including breach notification obligations." OpenAI's own discovery did not come from breach detection; it came from what the company calls a review of "misaligned model activity." Its reporting framework, published on September 16, commits to "provide advance notice even when no security boundary was crossed." That is the right instinct, and it arrived six days after the email to the public inbox.

If your vendor's agent took an action inside your study that it was not asked to take, found out two months later during an internal model review, and concluded no security boundary had been crossed, would your quality agreement require them to tell you? Would it say whom, and by when? If the answer is not plainly yes, that is the clause to write.

What I would do

First, make the walls real. An agent's permissions should be the boundary of what it can do, and nothing it reads, including a refusal, should be treated as a control. If the reconciliation agent must never write to the depot record, it holds no account that can.

Second, give every agent a query. Before an agent goes live, specify in the URS what it does when it cannot complete its task: the structured output it produces, the human who receives it, and the time within which that human responds. Test it in UAT the way you would test a randomization edge case, by deliberately blocking the legitimate path and watching which way the agent goes.

Third, count escalation as success. If the metrics on the agent's dashboard reward completed tasks alone, you have built the decision environment the research says produces gaming. Report escalations as a first-class outcome, and treat a sudden drop in them as a signal worth investigating rather than a sign of improvement.

Fourth, define an agent incident in the quality agreement. Cover actions outside the assigned task, whether or not a security boundary was crossed, and whether discovered by monitoring or by later review. Set the clock from discovery, not from confirmation, name the sponsor's contact point, and tie it to your serious breach process so that the seven days under Article 52 do not start on a Tuesday afternoon when someone finally opens a shared mailbox.

Here is something you can do tonight without a tool. Pick one automation or agent already running against one of your study systems. Write down, in a sentence, what it does when it cannot finish. If the honest answer is "it retries," or "it fails silently," or "I would have to ask the vendor," you have found your first query path to build.

Credit

Transluce's team reconstructed the agent activity from public logs, and their report at transluce.org is the most careful public account of the activity so far. Francesca Gomez's escalation-channel paper turned a vague worry into a testable intervention with a measured effect, and I would put it in front of anyone designing agent workflows in a regulated setting. The Huang and Chen group's work gives the clearest current figures on how often research agents game their own evaluations. The Australian Prime Minister's office published a full transcript rather than a summary, which is why this article can quote it.

If you are writing the escalation section of an agent URS, or adding an agent-incident clause to a quality agreement and want it to survive a vendor's legal review, that is the kind of build work BC Consulting does.

Sources

Anthony Albanese, "Press conference - New York" (transcript), Prime Minister of Australia, September 24, 2026. https://www.pm.gov.au/media/press-conference-new-york

ABC News, "OpenAI hacked Medicare portal, Prime Minister Anthony Albanese says," ABC News (Australia), September 24, 2026. https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078

Jack Cable, Daniel Chiu, Francisco Pernice, Selena Zhang, James Anthony, Tetiana Bas, Gary Shen, Conrad Stosz, Jacob Steinhardt, "Early rogue AI agent activity and attempts to hack found on urlquery.net," Transluce, September 23, 2026. https://transluce.org/agent-activity

Tim Fernholz, "For months, OpenAI's agent swarms have been attacking online databases to find obscure facts," TechCrunch, September 25, 2026. https://techcrunch.com/2026/09/25/for-months-openais-agent-swarms-have-been-attacking-online-databases-to-find-obscure-facts/

Francesca Gomez, "Can escalation channels redirect reward hacking toward defect disclosure?," arXiv:2608.29460v2, September 2, 2026. https://arxiv.org/abs/2608.29460

Yue Huang, Zichen Chen, et al., "Reward Hacking Challenges Oversight of Autonomous Research Agents," arXiv:2609.28614, September 23, 2026. https://arxiv.org/abs/2609.28614

OpenAI, "Our framework for reporting model misalignment," OpenAI, September 16, 2026. https://openai.com/index/model-misalignment-reporting-framework/

International Council for Harmonisation, "Guideline for Good Clinical Practice E6(R3)," Step 4 final guideline, January 6, 2025. https://database.ich.org/sites/default/files/ICH_E6(R3)_Step4_FinalGuideline_2025_0106.pdf

European Medicines Agency, "Guideline for the notification of serious breaches of Regulation (EU) No 536/2014 or the clinical trial protocol," EMA/698382/2021, effective January 31, 2022. https://www.ema.europa.eu/en/documents/scientific-guideline/guideline-notification-serious-breaches-regulation-eu-no-5362014-or-clinical-trial-protocol_en.pdf

Katherine Leibowitz, "Contracting For AI In Clinical Trials: Cybersecurity, Monitoring, And Risk Allocation (Part 3)," Clinical Leader, May 15, 2026. https://www.clinicalleader.com/doc/contracting-for-ai-in-clinical-trials-cybersecurity-monitoring-and-risk-allocation-part-0001

Working on something like this?

We run AI Bootcamps and build custom tools for sponsors and service providers in clinical research. A short call is the fastest way to find out whether your problem fits.