Organizations deploying agents into regulated workflows generally encode governing rules in the agent's system context, such as a rule against proceeding without a second approval. The rule is included in the system context once and is expected to apply throughout the conversation. PACT, a benchmark published on September 16 by Trace AI Labs researchers Mika Okamoto and Ansel Kaplan Erol, tests whether models continue to follow such rules when users pressure them to make exceptions.

PACT includes 48 scenarios across 12 regulated industries, including hiring, healthcare, and finance. Each scenario follows an employee asking an agent to complete a task that would require breaking a rule in its system context. The employee then applies common workplace pressure, such as asking repeatedly, citing an urgent deadline, or arguing that an exception would make the task easier. The researchers tested different versions of these requests and system prompts. They also reviewed each scenario to ensure the correct response was clear and the conversation resembled a normal workplace interaction.

Across 22 models from multiple providers and model families, user pressure increased the average violation rate by 65%. Failures were also present before the pressure began. Even the strongest assistants misapplied the relevant rule in 6-10% of the initial cases.

The benchmark evaluates compliance using six separate metrics. These cover robustness under pressure, consistency as the conversation develops, transparency about the agent’s actions, and the ability to determine whether a rule applies to the current situation. That final capability is especially important in decision workflows because an agent cannot reliably enforce a rule it has failed to identify as relevant. The results show substantial variation between models and across individual dimensions, meaning a model that performs well on one aspect of compliance may still fail on another.

The results also show why policy enforcement should not depend on the agent alone. System-context rules compete with repeated user instructions during inference and can lose influence over the course of a long conversation. An external policy layer can inspect each proposed tool call or action before execution and block anything that violates the rule, regardless of repetition, urgency, or other conversational pressure.

This failure mode is unlikely to appear in standard single-turn evaluations. An agent may identify the relevant rule and follow it on the first request, but violate the same rule after the user reframes the task, introduces a deadline, invokes managerial authority, or repeatedly asks for an exception. Teams deploying agents in regulated workflows should therefore test policy adherence across full multi-turn conversations. They should also verify that blocked actions cannot reach external tools or production systems if the model eventually concedes.

The DecideWise Insider

Angela Carducci, Founding Member of the DecideWise community, argues that human judgment becomes more important as AI agents gain autonomy and become harder to monitor. She points to the investigation of OpenAI agents that hacked Hugging Face, where access to chain-of-thought traces helped investigators reconstruct their actions, and GPT-6 Astra, whose system card reports lower monitorability than GPT-5.6 Sol. Carducci uses the AI canvas from Prediction Machines to divide enterprise decisions into seven elements: prediction, input, judgment, training, action, outcome, and feedback. The framework helps teams decide which tasks AI should perform, where humans should retain authority, and how outcomes should inform future decisions.

Market Pulse

  • Adthena has introduced Decision Intelligence, which combines a company’s Google Ads or ChatGPT Ads data with more than 14 years of labeled paid-search data. The system identifies issues such as declining quality scores, competitor-driven market shifts, keyword gaps, and spending on non-converting searches. It currently produces editable reports and upload-ready files, with in-app recommendations and one-click implementation planned.

  • RWS has released public previews of Tridion Agent and Tridion Connect, two components of its agentic platform for enterprise content operations. Tridion Connect monitors external sources for changes to regulations, policies, products, and technical information, then identifies which internal content may need updating. Tridion Agent uses that information to draft revisions and coordinate tasks across content workflows, while keeping source references, approvals, version histories, and audit records. The tools are intended for organizations managing large volumes of regulated or technical content across multiple formats, systems, and languages. General availability is planned for 2027.

  • Sixfold has launched an AI system that reviews medical, prescription, driving, and financial records against an insurer’s underwriting manual. It identifies impairments, flags missing evidence, cites the relevant sources and rules, and recommends whether to rate, refer, decline, or postpone a case. Customers have reported a 55% reduction in case-evaluation time and 30% more premiums written per underwriter.

  • IBM was one of several vendors named a Leader in the 2026 IDC MarketScape for Decision Intelligence Platforms. FICO, Quantexa, Experian, Palantir, SAS, and Aera Technology also placed in the Leaders category, highlighting the breadth of established players competing in the decision intelligence market.

Resources and Events

📅 Enterprise AI Summit (Charlotte, NC - October 7-8, 2026)

Hosted by Gene Kim at the Carolina Theatre, the two-day summit brings together enterprise technology leaders using AI and AI-assisted coding. Sessions will examine agent-written code review, application security, legacy modernization, developer productivity, staffing, budgeting, and the changing structure of engineering teams. Speakers include executives and engineers from OpenAI, Mastercard, Morgan Stanley, GitLab, Cisco Security, Red Hat, Vanguard, adidas, and Northrop Grumman.

📅 2027 SDP Annual Conference (Golden, CO - March 22-26, 2027)

The Society of Decision Professionals will hold its annual conference and workshops at the Colorado School of Mines, with the main conference running March 23-25. Under the theme "Clarity Amid Complexity: Better Decisions for a Dynamic World," the program will cover decision-making methods, case studies, research, applications, and lessons from unsuccessful projects. The call for presentations and posters is open to all industries, with particular interest in speakers from energy and pharma/life sciences. Proposals are due October 1, 2026.

📊 Report Spotlight: Worldwide Decision Intelligence Platforms 2026 Vendor Assessment (IDC)

IDC evaluated platforms that design, automate, monitor, and refine decisions through closed feedback loops. It found that vendors are combining deterministic rules with machine learning, using generative AI primarily to author, modify, and explain decision logic in natural language. Quantexa was named a Leader for its dynamic entity resolution and knowledge graphs, which connect structured and unstructured data before rules, graph analytics, models, workflows, or agents act on it. Its agentic stack includes an MCP-based gateway, embedded copilots, and task-specific agents, but deployment depends on data readiness, graph-model configuration, and mature AI governance.

Join the conversation: How can enterprises govern millions of AI agents?

Join the DecideWise discussion on agent discovery, shared context, inherited permissions, unexpected behavior, human oversight, and control as agent autonomy increases.

For the Commute

Technology, Change, and Management (Decision Intelligence Lab)

Technology CFO and author Karthik Gada joins Vijay Mehrotra and Northwestern University’s Michael Watson to explain why technological progress accelerates as one generation’s outputs become the next generation’s inputs. Gada estimates that industries improving by more than 10% annually now account for 4% of the global economy, up from 2% eight years ago, and could reach 8% within another eight years. The discussion examines how institutional rigidity delays change in healthcare, education, government, and large companies, while cheaper AI allows individuals to automate work, develop specialized businesses, and manage increasing complexity.