WEKID™
EXECUTIVE SUMMARY

Governing Machine Learning Decisions

Why predictive models and AI agents require different controls but the same discipline of evidence and authority
SEPTEMBER 2026
The executive issue: Enterprises are building new controls for generative and agentic AI while leaving conventional machine-learning systems inside older model-risk practices. That separation is dangerous. The systems work differently, but both convert uncertain evidence into decisions. The financial, operational, legal and human risk begins when an organization decides what those outputs are allowed to do.

Machine learning and agentic AI are different, but the governance failure is the same

A machine-learning model normally produces a score: the probability of fraud, default, illness or equipment failure. An AI agent may interpret information, choose a tool and execute a sequence of actions. The controls cannot be identical. A predictive model must be governed for training data, labels, thresholds, calibration, drift and repeated decisions at scale. An agent must also be governed for context, tool selection, sequencing, permission boundaries and recovery from an unsafe action.

But both systems cross the same boundary. Inputs are accepted as evidence, evidence becomes an inference, and the inference is granted authority to affect the world. Traditional ML research calls its failures underspecification, shortcut learning, leakage, distribution shift and data cascades. Agentic-AI discussions use terms such as hallucination, context failure, tool misuse and excessive autonomy. Different vocabulary has encouraged organizations to build separate governance programs around what is fundamentally the same executive question: what does the evidence justify allowing this system to do?

Governance dimensionMachine learningGenerative and agentic AI
Typical outputA score, classification or forecastGenerated content, a recommendation or a proposed action
Where judgment hidesThe threshold that turns a score into an actionThe prompt, policy, tool permission and action plan
Primary control cadencePeriodic population testing plus a gate on each predictionEvaluation of each output, tool call or proposed action
Characteristic riskOne flawed inference repeated consistently at scaleA novel failure emerging from context, reasoning or action sequence
Shared requirementEvidence must establish the limit of authority, and consequence may require more human control.

What the evidence shows

These are examples of publicly discosed failure records the reject the notion that better model accuracy is enough. In one widely deployed sepsis model, vendor performance claims of 0.76 to 0.83 AUC fell to about 0.63 under independent evaluation; local validation had not preceded deployment.  AUC is a common measure of how well a model separates higher-risk from lower-risk cases, with 1.0 representing perfect separation and 0.5 little better than chance. Measured independently, the model came out at about 0.63.  A state benefits system produced 40,195 automated fraud determinations, with later reviews finding error rates of roughly 85% to 93%; verification and appeal came after penalties, garnishments and damaged lives. An automated home-buying operation used models that performed acceptably in stable conditions, then deteriorated as the market changed, contributing to roughly $881 million in losses.

These cases failed in different places. Evidence supplied by the builder was accepted without sufficient challenge. Historical patterns or measurable substitutes were treated as the real outcome. Operating conditions changed while authority remained in place. Probabilities were treated as judgments, and decisions were automated without a proportionate route for review or appeal.

The academic literature reaches the same conclusion from another direction. Models with nearly identical test performance can behave differently after deployment. They can succeed by learning convenient shortcuts rather than durable relationships. Leakage can make development results look stronger than real use permits. Small changes anywhere in an ML system can invalidate evidence gathered against the previous version. The problem is not that probabilistic systems are unusable. The problem is that a clean, repeatable score can make uncertainty disappear from view precisely when a business rule converts that score into action.

Authority belongs to the decision, not the model.

What WEKID changes

WEKID separates the quality of the evidence from the authority granted to the decision. The Epistemic Maturity Model evaluates Data, Information, Knowledge, Experience and Wisdom. It asks whether inputs are authentic and traceable, whether they mean what the model assumes, whether the conclusion is sound and bounded, whether it has survived operation, and whether delegation is justified. The layers gate rather than average: strong performance at one layer cannot compensate for a foundational failure at another.

The AI Decision Authority Model then evaluates the decision. Evidence establishes a maturity ceiling, the most authority the evidence can support. Consequence establishes a human-authority floor, the minimum human control required by the stakes. The more conservative requirement governs.

This matters because a single model can support several decisions with different consequences. A predictive-maintenance model might be trusted to order an inexpensive part, require human approval to reschedule an outage and remain advisory when personnel safety could require shutting down a plant. The model, score and evidence are identical. What changes is the consequence of being wrong.

Between the two models sits the Trust Bridge. It is WEKID’s governed admission point from Epistemic Maturity, or knowledge maturity, into Decision Authority. Only an output that satisfies the required maturity thresholds and hard gates may cross the bridge and be considered for delegated authority. Crossing it does not grant authority. It admits the maturity result, supporting evidence, assumptions and uncertainty into Model Two, where consequence, reversibility and accountability determine what authority, if any, may be assigned.

The control model executives should require

For agentic systems, the Trust Bridge can evaluate each proposed output or action because every response may be a new artifact. Machine learning needs two clocks. The model and the population it serves must be evaluated periodically, while each prediction receives an efficient gate confirming that required inputs are present and current, confidence clears the approved threshold and the case resembles the conditions represented in training.

Experience is the pivot. Test results are not operating experience. Models should run in shadow mode before receiving consequential authority, be observed across difficult periods rather than only convenient windows and be compared with a credible alternative. A digital twin can expose a new release to complete business cycles, stressed conditions and rare events without placing customers, operations or assets at risk.

That evidence is still point-in-time evidence. Data, populations, processes and incentives change. A drift alert should therefore do more than turn a dashboard red. A failed gate should reduce authority automatically and invoke a predefined rule or human process. Authority should expire unless the evidence is re-established, and retraining should begin a new authority assessment because it produces a new release.

What executives should do now

  1. Build a decision inventory. Start with every approval, denial, alert, prioritization and operational action triggered by AI. A model register alone cannot show where authority has accumulated.
  2. Separate evidence from the builder. Treat builder claims as declared evidence. Require independent validation before an output can support delegated action.
  3. Govern ML and agents through distinct control cadences. Evaluate agent outputs and actions as they occur. Evaluate ML populations over time and gate each prediction against established bounds.
  4. Assign authority to each decision. Record the evidence ceiling, consequence floor and resolved Authority Level for every use of a model, not once for the model as a whole.
  5. Make drift and failed gates operational controls. Define the condition that revokes authority, the fallback that takes over and the evidence required for restoration.
  6. Require operating experience. Use shadow operation, full-cycle testing, credible alternatives and digital twins where appropriate before granting consequential authority.
  7. Publish accountability. Record who approved the assignment, who monitors it, how affected people can contest it and when the authority expires.

Conclusion

Machine learning is not a lesser governance problem because it predates generative AI. It is the quieter one: deeply embedded, operating at scale and already connected to consequential decisions. Agentic AI makes the authority problem more visible, but it did not create it.

Executives do not need one identical control system for every form of AI. They need one consistent governance principle applied through controls suited to how each system works: evidence must earn authority, consequence must bound it, and accountability must remain visible when either changes.

Read the full analysis Explore the WEKID framework