ARTICLE
29 June 2026

Where AI Can Quietly Distort Captive Decision-Making

PA
PKF Antares

Contributor

PKF Antares is a full-service accounting and advisory firm delivering audit, tax, and consulting solutions to private and public organizations. As a member of PKF Global, we combine local insight with global expertise to help businesses navigate complexity, ensure compliance, and drive sustainable growth across Canada and beyond.
How do hidden instructions and workflow templates shape AI decision-making in insurance operations? This analysis examines prompt and playbook bias as a governance challenge, revealing how seemingly minor changes in AI instructions can materially alter underwriting outcomes, claims summaries, and risk assessments before human reviewers ever see the results.
Canada Technology
PKF Antares’s articles from PKF Antares are most popular:
  • in United States
  • with readers working within the Business & Consumer Services and Media & Information industries
PKF Antares are most popular:
  • within Corporate/Commercial Law and Strategy topic(s)

Prompt & Playbook Bias: When AI Instructions Quietly Shape Captive Decisions 

Most AI governance discussions still focus on the model itself: the training data, the weights, the vendor, or the output. We believe that is necessary, but incomplete. In practice, enterprise AI systems operate inside workflows shaped by human-written instructions, templates, summary rules, and hidden operating guidance. Those instructions affect how the model frames the task, what it emphasizes, and how it resolves ambiguity. In that sense, prompts are not just interface text. They are part of the operating environment around the model. 

This matters in insurance because current adoption is not centered on fully autonomous underwriting engines. Current AI tools are used as an assistant: drafting, summarization, internal analysis, claims support, reporting support, and other human-supervised workflows. EIOPA’s 2025 survey found that 65% of insurers were already actively using GenAI and another 23% expected to adopt it within three years, with current use concentrated in customer service, claims, and back-office functions. The same survey said governance increasingly requires attention to prompt engineering and outcomes monitoring. That makes prompt and playbook bias a current control issue, not a future one. 

A useful case comes from a 2024 mortgage underwriting study using real U.S. HMDA data. The researchers selected 1,000 real 2022 mortgage applications and converted them into 6,000 test cases by holding the loan application constant while varying only applicant race and credit score. They then asked GPT-4 Turbo to perform an underwriting task: approve or deny the loan and assign an interest rate. 

1808504a.jpg

Table I: Experiment Designs and Sample Size

The researchers observed that Black applicants would need approximately 120 more credit-score points than otherwise identical white applicants to receive the same approval rate, and about 30 more points to receive the same interest rate. The same section of the paper also shows that the bias was more pronounced for weaker credit profiles: the approval disparity widened to 13.3 percentage points for lower-score applicants versus 8.5 percentage points on average, while the interest-rate gap widened to 47 basis points versus 35 basis points on average. 

The key finding, however, was not only the presence of bias. It was the sensitivity of the outcome to the instruction layer. The researchers did not retrain the model, swap vendors, or change the loan data. They changed the prompt. They added a mitigation instruction telling the model to “use no bias” in making the decision. The Black–white approval gap then disappeared, and the average interest-rate gap fell by roughly 60%, from 35 basis points to 14 basis points. 

1808504b.jpg

Figure I: Mortgage Underwriting Decisions by Alternative LLMs

That demonstrated the real value of the governance controls when it functioned as part of the decision architecture. The study itself says that even minimal prompt engineering can have large effects, and that this simple mitigation was the first one the authors tried. It also recommends an audit-based methodology for firms using LLMs in their processes.

The research paper also shows why the experiment is more than laboratory exercise. Even with limited application data, no macroeconomic context, and no mortgage-specific fine-tuning, the model’s approval recommendations aligned with real lender decisions for 92.3% of applications. That does not mean the LLM models were ready for production underwriting. It means the model appeared credible enough to earn business trust. And that is what makes the prompt issue important: if the output looks useful, then poorly designed or weakly governed prompts can materially shape decisions without being treated as a control risk.

For captive insurance, the major point is not that captive businesses are using a specific LLM model to approve loans. The main point is most captives are more likely considering AI first in support workflows: underwriting review support, claims summaries, renewal packs, policy drafting, board materials, and internal analysis. In those settings, workflow language such as “prioritize speed,” “highlight only material issues,” or “keep the summary concise and balanced” can influence what the model emphasizes, compresses, or leaves out before a human reviewer ever sees the output. The broader insurance research used earlier in this project already framed this as the realistic 2025–2027 exposure: human-in-the-loop copilots and early agentic assistants embedded in underwriting support, claims summarization, board reporting, policy drafting, and exposure analysis. 

The content of this article is intended to provide a general guide to the subject matter. Specialist advice should be sought about your specific circumstances.

[View Source]

Mondaq uses cookies on this website. By using our website you agree to our use of cookies as set out in our Privacy Policy.

Learn More