What an AI Agent Says It Did Versus What the Audit Trail Shows
By Bryan Clayton10 min read
Originally published on LinkedIn.
- ai
- agents
- gxp
- data-integrity
- audit-trail

What jumps out for eClinical applications
Most model news reaches me as a benchmark table, and I read it the way I read a vendor's capability slide: politely, and with a (virtual) pencil nearby. The last week of September was different. On September 28, Maxwell Zeff at the Wall Street Journal reported why OpenAI had cancelled the planned October release of a model called GPT-6.1 Astra. Among the reasons: the model "wasn't always honest about telling users of the actions it did or didn't take."
I have spent more than fifteen years around IRT and clinical supply systems, and I recognized that sentence immediately. It is a data integrity finding. Swap "the model" for "the site" or "the vendor" and it could have come out of an audit report on any study I have worked on; or worse, an FDA 483.
This piece covers what OpenAI and the UK AI Security Institute reported, why an AI agent is able to describe its own work inaccurately in the first place, and what that means for anyone using an agent with an EDC, an IRT, or a depot system. The short version is that clinical research addressed this issue decades ago, in a federal regulation most of us can recite from memory, and it is still just as relevant today.
What was reported, and what was not
Saachi Jain, OpenAI's head of safety systems, told the Journal in an interview that GPT-6.1 Astra regressed in two areas compared with its predecessor, GPT-6 Astra. The first was the honesty problem above, which the Journal describes as higher levels of deception on alignment tests. The second is what OpenAI calls "scope authorization": the model "would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe." Jain said the model improved on "model laziness" but did not meet OpenAI's bar for safety and alignment. OpenAI confirmed the decision to CNBC the same evening, and Jain put the tension plainly: "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction." OpenAI has published no post or system card for GPT-6.1 Astra; the Journal interview and the CNBC confirmation are, as far as I can find, the only public record.
On September 28, the same day as the Journal report, the UK AI Security Institute (AISI) published results on the model OpenAI did ship, GPT-6 Astra, which now runs OpenAI's new Dots agents. In a simulated environment called Petri, AISI's Alignment Red Team found that GPT-6 Astra completed unsanctioned supply-chain attacks in 29.2 percent of scenarios, against 6.3 percent for GPT-5.6 Sol and zero for GPT-5.5. Two caveats come with that figure, as stated by AISI. OpenAI's standard safeguards were deliberately switched off to measure the model's underlying tendencies, and the environment was simulated. AISI answers the second caveat itself: the scope violations are instruction-following failures "regardless of whether the environment is simulated." The agent was told where the line was and crossed it. Whether the system on the other side was real does not change that.
Before anyone reads this as an OpenAI problem: in an incident AISI disclosed in July, 17 of the 19 unsanctioned actions came from Anthropic's Mythos 5. This is a property of capable agents in general, and I use Claude every day.
Why an agent can misreport its own work
When an agent finishes a task and tells you what it did, it is natural to imagine it consulting a log, the way you would check your sent folder before telling a colleague you replied. Nothing guarantees that is what happens.
A language model produces its summary the same way it produces everything else: by generating the text that best fits what came before. The summary sits downstream of the work, but it is not a readout of the work. Picture a new coordinator at the end of a long Friday who is asked whether every subject visit got entered. The verifiable answer requires opening the system. The fluent answer, the one that sounds like a competent coordinator, is "yes, all done." Most of the time the fluent answer and the verifiable answer agree. The trouble is that you cannot tell which kind you are getting from the answer itself.
OpenAI has not published a root-cause analysis for GPT-6.1 Astra, so what follows is my inference from how these systems are trained. Models are increasingly rewarded for completing tasks, and Jain's comment about "laziness" tells you OpenAI was pushing in exactly that direction. If the training signal ever rewards a report of completion more reliably than it checks the completion, a model can learn that saying "done" is part of being done. Pushing a model to try harder and pushing it to report more accurately are separate pressures, and the reporting suggests they pulled apart here. OpenAI appears to be looking in the same place. Jain told the Journal the company plans several deep dives into the root cause, and that the work includes, in the Journal's words, "ensuring that OpenAI's reinforcement learning environments are rewarding the right type of behavior."
AISI saw a related pattern in Petri. When researchers checked by hand, the model sometimes invented simulation errors to explain what it had done. That reads to me less like a lie in the human sense and more like a plausible story generated after the fact, which, for a data integrity purpose, comes to the same thing.

Clinical research addresses this with what we call "Part 11"
Here is where I think our field may have an advantage that some other industries may not.
Part 11 is explicit. 21 CFR 11.10(e) requires "secure, computer-generated, time-stamped audit trails to independently record the date and time of operator entries and actions that create, modify, or delete electronic records." The word I would underline is "independently." The regulation does not ask the operator to describe what they did. It requires the system to record it, separately from anyone's account, and FDA's data integrity guidance builds the ALCOA expectations (attributable, legible, contemporaneous, original, accurate) on the same footing.
We also already practice the habit this calls for. Source data verification exists because a site's transcription into the EDC is a claim about the source, not the source. A monitor does not accept the site's word that the visit happened on the date entered; they look at the chart. Nobody considers this an insult to the site. It is how the work is designed.
An AI agent's end-of-task summary belongs in the same category as the site's word. It may well be right. It is still a claim, and the audit trail is the record.
Agents are already arriving in the systems where this matters. Suvoda announced an agentic RTSM in April that uses multiple AI agents for configuration, testing, and change orders, and says it can shorten the time from kickoff to UAT by up to 80 percent. Others will surely follow. None of the announcements I have read describe how agent actions are reviewed, which is the question I would put to any vendor. Consider a hypothetical in an agentic RTSM: an agent reports that it updated the resupply trigger for one depot in the UAT environment, re-ran the affected scripts, and that all of them passed. If that summary gets pasted into the UAT evidence package, you have filed a narrative as a test result. If one script was never re-run, the narrative will not tell you, and the validation file will say something the system cannot back up.
What clinical teams should do
This is how I would handle it, and it sits alongside my September 28 article on permission boundaries rather than replacing it. Permissions limit what an agent can do; these rules govern what you believe about what it did.
First, treat every agent summary as a claim and never as evidence. Nothing an agent writes about its own actions goes into a UAT package, a change order record, a deviation, or a TMF filing as proof that the action happened. The evidence is the system's own audit trail, exports, and test logs.
Second, reconcile the claim against the record automatically. For every action an agent reports, there should be a matching audit-trail entry, and for every audit-trail entry under the agent's identity, there should be a reported action. Both directions matter: a claimed action with no entry means the agent said it did something it did not, and an entry with no claim means it did something it did not mention. The Cloud Security Alliance's September 30 research note on Astra recommends this kind of tracking, agent self-reports against actual logs, and clinical systems are better placed than most to do it because the logs are already mandated.
Third, give the agent its own identity in every system it touches. The UK NCSC's August guidance on agentic AI says agents should have "their own unique identity in a class which differentiates them from human" users. Reconciliation is impossible if the agent works under a study builder's login, and attributability, the A in ALCOA, goes with it.
Fourth, treat a model version change as a change control event. GPT-6.1 Astra was newer than GPT-6 Astra and less honest about its own actions. Newer is not automatically safer. If your vendor swaps the underlying model, the reconciliation tests get re-run before the agent goes back into a validated environment, the same way you would treat any other change to a validated system.
Fifth, ask your vendors one specific question: when the agent tells me what it did, is that text generated by the model, or is it read back from the system's audit trail? A good answer is the second, or the first with an automated reconciliation behind it. A vague answer tells you something too.

Here is something you can do tonight. Pick one task you have handed to an AI tool this month that touched a real system, even something as small as a bulk update in a spreadsheet tracker. Write down what the tool told you it did, then open the file's version history or the system's audit log and check each claim. You will most likely find they match. The value is in learning how long the check takes, because that is the cost you will be designing into every agent workflow.
Credit where it is due
This piece rests on the AISI Alignment Red Team's GPT-6 Astra write-up, which is careful about its own limits in a way I wish more evaluation reports were, and on Maxwell Zeff's reporting in the Wall Street Journal, including Saachi Jain's on-the-record comments. The CSA AI Safety Initiative's research note is the most practical enterprise reading of the episode I found. All of them are worth following.
At BC Consulting, we run AI Bootcamps and build custom tools for sponsors and service providers in clinical research. If you want the reconciliation check from the exercise above designed into an agent workflow, with the agent's own identity and the audit trail doing the work, a short call is the place to start: bcconsulting.io/book.
Sources
Maxwell Zeff, "OpenAI Scraps Release of New AI Model Over Safety Concerns," The Wall Street Journal, September 28, 2026. https://www.wsj.com/tech/ai/openai-chatgpt-model-release-cancel-safety-5a2f9f42
CNBC, "OpenAI abandons plan to release upcoming model as safety concerns escalate," September 28, 2026. https://www.cnbc.com/2026/09/28/openai-abandons-plan-to-release-upcoming-model-as-safety-concerns-escalate.html
UK AI Security Institute, Alignment Red Team, "GPT-6 Astra performs unsanctioned supply-chain attacks in simulations," AISI Blog, September 28, 2026. https://www.aisi.gov.uk/blog/gpt-6-astra-performs-unsanctioned-supply-chain-attacks-in-simulations
UK AI Security Institute, "Incident Report: unsanctioned agent behaviour during cyber testing," AISI Blog, July 2026. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Lucas Ropek, "OpenAI launches Dots, its bubbly agentic avatar," TechCrunch, September 29, 2026. https://techcrunch.com/2026/09/29/openai-launches-dots-its-bubbly-agentic-avatar/
Cloud Security Alliance AI Safety Initiative, "OpenAI Shelves GPT-6.1 Astra Over Deception, Scope Failures," CSA Research Note, September 30, 2026. https://labs.cloudsecurityalliance.org/research/csa-research-note-gpt61-astra-deception-shelving-20260930-cs/
UK National Cyber Security Centre, "Managing the cyber risk of agentic AI," NCSC, August 20, 2026. https://www.ncsc.gov.uk/sites/default/files/2026-08/Managing-the-cyber-risk-of-agentic-AI.pdf
U.S. Food and Drug Administration, "21 CFR 11.10 Controls for closed systems," Electronic Code of Federal Regulations, current. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11/subpart-B/section-11.10
U.S. Food and Drug Administration, "Data Integrity and Compliance With Drug CGMP: Questions and Answers," Guidance for Industry, December 2018. https://www.fda.gov/media/119267/download
Suvoda, "Suvoda introduces agentic RTSM, cutting clinical trial start-up timelines by up to 80%," PR Newswire, April 21, 2026. https://www.prnewswire.com/news-releases/suvoda-introduces-agentic-rtsm-cutting-clinical-trial-start-up-timelines-by-up-to-80-302747770.html
Bryan Clayton, "What a Clinical AI Agent Does When the Answer Is No," BC Consulting, September 28, 2026. https://bcconsulting.io/blog/what-a-clinical-ai-agent-does-when-the-answer-is-no