Posts

The Provenance Problem

Image
  A leading private bank and a major wealth management firm asked us to build AI systems that could improve investment decision making and surface portfolio risk earlier than existing processes allowed. The solution we designed was a RAG architecture, retrieval augmented generation, that pulled dynamically from multiple enterprise data sources, applied business rules and policy parameters, and generated recommendations a relationship manager could act on in near real time. The system was validated carefully. The data sources were approved. The model was tested against historical portfolios. The workflow was documented end to end. The systems performed well in production. Recommendations were sharper. Risk flags arrived earlier. Portfolio teams reported the tools were genuinely useful. Several months into production, a specific recommendation came under internal review. The outcome itself was not the issue. The recommendation was reasonable. The process had followed policy. What nob...

The Accountability GAP

Image
A global financial services organization once asked us to review an incident that had triggered far more internal discussion than the incident itself. The issue had already been resolved. The system was functioning normally again. The business impact had been contained. Yet weeks later, senior leaders were still debating what had happened. Not because nobody cared. Because everybody cared. Technology had logs. Operations had reports. Risk had assessments. Compliance had approvals. Every team could produce evidence supporting its own perspective. Every team could explain what it owned and what it had done. What nobody could produce was a single coherent explanation of the decision path that led to the outcome. The logs existed. The reports existed. The evidence existed in fragments. The decision pathway did not exist at all. The organization did not have an accountability problem. It had an evidence problem. We had seen a sharper version of this earlier in a fintech engagement running A...

When Nobody Owns the Outcome

Image
A leading private bank asked us to build an AI platform that could sharpen investment decisions for its wealth management clients. We implemented a RAG architecture with vector embeddings and LLM integration, the kind of system that retrieves relevant market intelligence, reasons across portfolio context, and generates recommendations a relationship manager can act on quickly. The results during the pilot phase were strong. Recommendations were more precise. Risk flags arrived earlier. Client portfolio performance metrics moved in the right direction. Six months after deployment, a question nobody had asked at the start became unavoidable. When the model flags a risk and the portfolio manager does not act on it, who owns that outcome? The technology team built the model and maintained the architecture. The investment team received the recommendations and made the final calls. The risk function had signed off on the controls. Compliance had approved the process. A major wealth managemen...

The Hidden Cost of AI Pilots

Image
Every Successful AI Pilot Creates a Hidden Liability A large US insurer once asked us to help modernize how claims got processed. RPA bots went in, two legacy systems got replaced with one platform, and the pilot numbers looked excellent, faster intake, fewer manual touchpoints, happier claims adjusters in the room where it was tested. Everyone in that room had reason to be pleased. The metrics were real, the adjusters were not exaggerating their relief, and on paper this looked like exactly the kind of automation story organizations like to tell about themselves. Eighteen months later, that same pilot was still running in exactly one regional office. The model had not gotten worse. Nothing about its performance had changed. What changed is that scaling it meant pulling in people who were never part of the pilot, a budget owner who had to fund the rollout, a compliance team that suddenly had to sign off on something that used to be a contained experiment, an operations lead who had to ...

The Governance Line Nobody Draws: Why Enterprises Keep Regulating the Wrong Layer

Image
Most enterprise AI governance conversations I sit in on still treat "the model" and "the system around the model" as the same thing. A risk committee asks whether the AI is safe, someone answers with a benchmark score, and the conversation moves on. It is a comfortable shortcut, and it is also the reason so many governance frameworks fail the moment an agent gets real access to a real environment. Anthropic gave the industry an unusually clean way to see why that shortcut breaks down. In late May 2026, the company published a long engineering account of how it contains Claude across its three agentic products, claude.ai, Claude Code, and Claude Cowork. It is candid in a way corporate security writing rarely is, naming specific incidents, specific failure rates, and specific architectural choices that did not work the first time. Read against the backdrop of April's decision to hold back Claude Mythos Preview after it engineered its own way out of a sandbox dur...

AI Doesn't Fail in Isolation. Organizations Do.

Image
  Most AI Failures Are Not AI Failures A regional bank in the northeastern US came to us with a familiar ambition. Move from a hierarchical structure to a project based operating model, faster decisions, less layered approval, technology and operations working as one team instead of two. The technology side was never the hard part. The hard part was that nobody had touched decision rights. People kept reporting the way they always had, escalating the way they always had, getting evaluated the way they always had. The bank wanted agility without redesigning who owned what, and that gap is where the actual work began. We ended up redesigning the operating model for the technology and operations group, building a new talent platform and reward framework around it, because the structure had to change before any process inside it could. This is the pattern we keep seeing, and it rarely gets named correctly. Pilots succeed on a narrow, well defined task with a small group of engaged user...

From Agency Costs to Delegation Costs: Revisiting Agency Theory in the Age of Agentic AI

Image
1989, Kathleen M. Eisenhardt published a landmark review of Agency Theory that would influence decades of thinking across economics, organizational behavior, finance, governance, and management. While the theory is often associated with incentives and monitoring mechanisms, its deeper contribution was to illuminate a more fundamental organizational challenge: how do we govern delegated authoritywhen information is imperfect and uncertainty is unavoidable? At its core, Agency Theory examines the relationship between a principal and an agent. Shareholders delegate authority to executives, boards delegate authority to management, and clients delegate authority to advisors. The challenge arises because the principal cannot perfectly observe what the agent knows, what actions are being taken, or whether those actions remain aligned with the principal's interests. Agency Theory provided a framework for understanding these tensions through concepts such as information asymmetry, incentive...