When AI Becomes an Actor: The Accountability Architecture Nobody Built
The question everyone will ask is whether AI can replace doctors. That is the wrong question. The right question is this. When the AI says one thing and the doctor says another, and the AI turns out to be right, who carries the liability for the decision the human made? And when the AI turns out to be wrong and the doctor followed it without sufficient scrutiny, who carries that liability?
We do not have answers to either question. No court has ruled on them. No hospital protocol has been written for them. No medical board has defined the standard of care for a physician operating alongside an AI diagnostic tool that has just been shown to outperform them at triage. The accountability architecture in medicine was built for a world where AI was a tool that humans operated. It was not built for a world where AI is an actor in the room making independent assessments that clinicians must either follow, override, or explain their way around in a documented patient record.
That world arrived this week. The architecture is still in the post.
What the Harvard trial actually measured and why it matters
The study involved 76 patients arriving at the emergency room of a Boston hospital. OpenAI's o1 and pairs of emergency medicine doctors each received a standard electronic health record, vital signs, demographics, and a brief triage nurse summary. No physical examination. No direct patient interaction. The same information each party would have at the earliest and most critical moment of an ER presentation.
The 67% to 50-55% accuracy gap at intake is the number that will dominate the coverage. The more significant finding sits one level deeper. When given more clinical information, o1 climbed to 82% accuracy on complex diagnostic questions. Human doctors reached 70 to 79%. The authors wrote that large language models have eclipsed most benchmarks of clinical reasoning. Independent reviewers called it a genuine step forward.
McKinsey's 2024 healthcare AI report found that 74% of health system leaders plan to deploy AI diagnostic support within two years. The Harvard trial just provided the evidence base that will accelerate those deployments. The deployment decision is essentially made. What is not made is the governance framework that should precede it.
The liability map was drawn for human actors
In medical malpractice law the standard of care asks what a reasonably competent physician with similar training in similar circumstances would have done. That standard is built entirely around human professional judgment. It has no provision for the scenario where a reasonably competent physician had access to an AI diagnostic tool demonstrably more accurate than unassisted human judgment at triage, chose not to follow its recommendation, and the patient was harmed.
Did that physician fall below the standard of care by ignoring the AI? Or did they exercise appropriate independent professional judgment? A plaintiff's attorney will argue one position. The hospital's legal team will argue the other. The medical board has not defined a standard for this scenario. No court has ruled on it.
Across advisory engagements in healthcare IT strategy this liability vacuum is the single most consistent barrier to AI deployment in clinical environments, more significant than model accuracy, more significant than integration complexity, more significant than staff training overhead. Health systems know the technology works. They do not know who absorbs the legal exposure when it is wrong, and they do not know what their obligation is when they choose to proceed without it and it would have been right.
The NIST AI RMF 1.0, published in January 2023 and the current US federal standard for AI risk management, explicitly frames human oversight as a core govern function requirement. But the oversight it describes is primarily concerned with monitoring AI output for errors, not with designing AI systems and workflows that account for the scenario where the AI is right and the human is wrong. That gap in the framework is not an oversight. It is an honest reflection of the fact that the governance community had not yet confronted a situation where the model demonstrably wins the comparison. It has now.
The courtroom story that connects directly
At almost exactly the same moment the Harvard trial results were circulating, a completely separate but structurally identical story broke from the legal domain. CNN reported that prosecutors and plaintiffs' attorneys are subpoenaing ChatGPT logs and using them as evidence in criminal and civil cases. The Florida Attorney General opened a criminal investigation of OpenAI after alleging ChatGPT gave significant advice to the Florida State University shooter. Families of victims in a Canadian school shooting filed suit against OpenAI alleging the model was complicit in the attack. Courts are treating chatbot conversation logs the way they treat Google search history. Subpoena-able. Discoverable. Admissible.
The connection between the ER trial and the courtroom story is not immediately obvious. It becomes clear when we ask the same governance question of both. In medicine the accountability architecture was built for human actors and AI has just become a consequential actor in clinical decisions. In law the privacy architecture was built for human confidential relationships and AI has become a repository of disclosures that millions of people treat as confidential without any of the legal protections that confidentiality actually requires.
A doctor cannot be subpoenaed for what a patient told them in consultation. A therapist cannot be compelled to disclose session content. A lawyer cannot be forced to reveal privileged communications. No equivalent protection exists for AI conversations. Every prompt typed into a chatbot is a written record on a server owned by a company that responds to court orders. Millions of users are sending ChatGPT medical questions, legal questions, and emotional disclosures under an implicit assumption of something resembling privacy. That assumption has no legal foundation anywhere in the world right now.
The pattern across both stories
The Harvard trial and the ChatGPT evidence story are superficially unrelated. One is about clinical accuracy. One is about legal discovery. What they share is more important than what separates them.
In both cases AI has crossed a threshold from experimental to consequential. In the ER it is making assessments that materially affect patient outcomes and the accuracy data now makes its involvement in clinical decisions professionally arguable. In the courtroom its conversation logs are material to criminal investigations and civil litigation. In both cases the institutional frameworks governing accountability, liability, professional standards, privacy, and privilege were designed before AI crossed that threshold. They have not been updated to reflect a world where it has.
Microsoft's launch of Agent 365 this week is instructive as a contrast. The company deliberately built its AI agent governance framework to mirror the identity, permissions, and audit controls it already uses for human employees. Every AI agent in the Microsoft 365 environment is governed by the same accountability infrastructure as a person with equivalent access. IBM's simultaneous launch of Bob, its AI coding platform with multi-model routing and human checkpoints baked into the architecture, makes the same point from a different angle. Human oversight is not a feature that gets added after deployment when something goes wrong. It is a design requirement that gets built in before the first consequential decision is made.
The ER and the courtroom are learning this lesson from the exposed side. The enterprise software vendors are demonstrating what it looks like to learn it from the designed side. The distance between those two approaches is the accountability architecture gap that the next generation of AI governance frameworks needs to close.
What this means for enterprise AI strategy now
For every enterprise team in healthcare, legal services, financial services, or any other domain where AI is moving from advisory to consequential, three things need to be true before the next deployment decision is made.
The liability allocation needs to be defined before deployment, not litigated after the first disputed outcome. Who owns the decision when the AI and the human disagree? Who owns the outcome when the AI-assisted decision produces harm? These are not rhetorical questions. They are procurement requirements.
The privacy and confidentiality architecture needs to match the actual use behaviour of the people interacting with the system. If users are treating AI conversations as confidential disclosures the governance framework needs to address that expectation honestly, either by extending genuine protections or by making the absence of protections unmistakably clear before the disclosure is made.
The human oversight design needs to be built into the architecture before deployment rather than added as an afterthought when an incident creates regulatory pressure to demonstrate accountability. The Microsoft and IBM examples from this week show that this is achievable. It requires prioritising it.
The Harvard trial marks the moment when AI diagnostic accuracy became an argument that health systems cannot responsibly ignore. The ChatGPT evidence story marks the moment when AI conversation data became legally material in ways that users were not warned about. Both mark the same underlying shift. AI has stopped being a tool and started being an actor. The accountability architecture needs to catch up before the next consequential decision arrives without a governance framework to hold it.
Data source: Romero-Brufau S. et al. Large language models versus clinicians in emergency medicine. Science. 2026. Harvard Medical School. 76 patients. Boston ER. Additional references: McKinsey State of AI 2024. NIST AI RMF 1.0. CNN Legal Analysis 2026. Microsoft Agent 365 GA 2026. IBM Bob 2026.
The strategic observations in this piece draw from advisory engagements across healthcare IT strategy, BFSI, and enterprise AI governance, and from co-authored research on responsible AI governance presented at BIGS 2025, AIS eLibrary.


Comments