A customer asks an organization's chatbot whether a fee will be waived, a payment will be timely, or a disputed transaction will be reversed. The chatbot answers quickly and confidently.
And the answer is wrong.
The organization may respond that the chatbot lacked authority to amend the governing terms, waive a requirement, or make a binding commitment. That may answer one question. It does not answer whether the customer reasonably relied on a representation delivered through a channel the organization selected and made available to users. It does not answer whether the response interfered with the customer’s exercise of a legal right. And it does not answer whether the circumstances require the organization to investigate the error, determine whether other users were affected, or remediate resulting harm.
“Obligation” in this context does not mean that every incorrect answer amends a contract or binds an organization. The interaction may instead trigger an existing statutory process, support a reliance-based claim, require investigation or correction, become subject to an existing preservation duty, or expose a control failure the organization must address.
A chatbot can also receive a communication that may invoke a legal right or trigger an existing process without sending it where the law expects it to go. It can describe a policy correctly and still apply it incorrectly to a particular customer. Internally, it can tell an employee that leave is unavailable, that a contract requires no legal review, or that an incident requires no escalation. The central deployment question is therefore not simply whether the chatbot's error rate is low enough. It is what the organization has made the chatbot available to say, receive, promise, and do.
Lack of Authority Does Not End the Analysis
Formal authority matters when a chatbot appears to waive a fee, modify terms, extend a deadline, or promise some other organizational action. Its absence may be highly relevant to whether the statement changed an existing obligation. Other consequences, however, do not necessarily turn on that question.
Depending on the circumstances, a chatbot interaction may implicate:
- Misrepresentation or unfair or deceptive practices;
- Apparent authority or related attribution arguments;
- Waiver, estoppel, or contractual arguments;
- Statutory duties governing disputes, notices, revocations, terminations, or appeals; and
- Requirements to investigate, correct, preserve, or remediate.
These theories do not collapse into one another. A chatbot’s statement may fail to modify terms and still mislead a customer. It may fail to create an enforceable commitment and still cause an organization to miss a dispute deadline. It may accurately describe a general rule while applying that rule incorrectly to the user’s specific circumstances.
“Did the chatbot have authority?” may therefore be the beginning of the analysis, but the analysis should not end there. The more practical inquiry is what kind of reliance environment the organization chose to create and what consequences it should reasonably expect to follow from that choice.
Making the Chatbot Available Creates the Reliance Environment
Making a chatbot available does not make reliance on every answer reasonable. An organization has nevertheless deployed an environment in which reasonableness will later be judged. It decides where the chatbot appears, whether it sits behind authentication, how prominently organizational branding surrounds it, what subjects it invites users to ask about, and what account information it may access. It also decides whether the chatbot cites governing documents, communicates uncertainty, applies general information to individual circumstances, takes actions, or sends particular questions to a person.
A chatbot on a public webpage that retrieves office hours presents one reliance profile. A chatbot inside an authenticated account that addresses the customer by name and provides a specific payment deadline presents a very different one. An organization may call the second system an assistant or informational tool, but the user may reasonably understand something simpler: they answered my question.
Disclaimers remain relevant, but they are not incantations. A general warning that automated answers may contain errors may do less work if an organization places the chatbot inside an official application, gives it account access, invites consequential questions, answers in unqualified language, and makes human assistance difficult to reach. A disclaimer is one fact in the reliance analysis, not a substitute for designing the interaction around the consequences of an incorrect answer.
The Correct Answer Somewhere Else May Not Cure the Wrong Answer Here
The Canadian decision in Moffatt v. Air Canada remains the most useful illustration. In Moffatt, a customer asked Air Canada’s website chatbot about bereavement fares. The chatbot said the customer could buy an ordinary ticket and apply for the reduced fare afterward. The airline’s actual policy, however, required that the request be made before travel. The customer relied on the chatbot’s advice when purchasing the ticket and was later denied the requested retroactive adjustment. The British Columbia Civil Resolution Tribunal ultimately found Air Canada liable for negligent misrepresentation.
Air Canada argued that it could not be held liable for information provided by the chatbot. The tribunal described Air Canada as suggesting that “the chatbot is a separate legal entity that is responsible for its own actions,” calling that “a remarkable submission.” The tribunal rejected the argument and held that Air Canada remained responsible for information supplied through the chatbot on its website.
The airline also argued that the correct policy was available elsewhere on its website. The tribunal was not persuaded that the customer should have regarded the separate policy page as inherently more trustworthy than the chatbot the airline had also placed there.
The lesson is fairly ordinary:
An organization should not assume it can publish the correct answer through one official channel and the wrong answer through another, then decide after reliance which channel the customer should have trusted.
Moffatt should not be made to carry more authority than it has. It is a decision of a Canadian small claims tribunal and is not controlling precedent for US courts. Its importance lies partly in how little AI-specific doctrine the tribunal needed. The chatbot was not treated as an independent legal actor that happened to wander onto the airline’s website; it was part of the airline’s chosen customer service environment. Traditional principles of duty, reasonable care, inaccurate representation, reliance, causation, and loss ultimately did the work.
The Output Problem and the Input Problem
Most discussions of chatbot risk focus on outputs: a wrong answer about a price, deadline, account status, eligibility, refund, or what an organization will do next. And while those errors matter, they are not the only failure mode.
A chatbot can create exposure without saying anything false. It can fail by treating a potentially operative communication as a conversation.
A customer might say something like:
- “I dispute this transaction.”
- “That transfer was not mine.”
- “I revoke authorization.”
- “I need an accommodation.”
The chatbot may respond politely, accurately, and uselessly. If the statement never reaches the function responsible for disputes, fraud, payments, or accommodations, an organization may have a problem that no hallucination metric captures.
Whether a particular message is legally operative depends on the governing law, its content, the channel through which it was sent, and an organization's disclosures and procedures.
Regulation E, for example, requires a financial institution to follow prescribed error resolution procedures after receiving a qualifying oral or written notice of an electronic fund transfer error. The notice must satisfy the rule’s timing and content requirements, including identifying the consumer and account and indicating why the consumer believes an error occurred. A financial institution may require written confirmation within 10 business days after oral notice, subject to the rule’s conditions, but the qualifying oral notice begins the process.
Whether a message entered through a particular chatbot constitutes qualifying notice received by the institution will depend on the rule, the content of the message, the channel, and the institution’s disclosures and procedures. That is precisely why routing should be resolved before deployment rather than during the later argument over whether the customer used the right interface.
Regulation Z, on the other hand, uses a different structure for formal billing error notices. The notice generally must be written, received within 60 days after the creditor transmitted the first periodic statement reflecting the alleged error, and sent to the address the creditor designated for billing error notices.
A chatbot message may therefore fail to satisfy Regulation Z’s formal notice procedure even though the creditor operates the channel. That does not resolve what happens if the chatbot says a dispute has been opened, directs the consumer away from the required procedure, or allows the deadline to pass without meaningful escalation.
In its June 2023 issue spotlight, Chatbots in consumer finance, the Consumer Financial Protection Bureau (“CFPB”) identified inaccurate answers, failures to recognize consumer disputes, and inadequate access to timely human assistance as recurring chatbot risks. The spotlight expressly stated that it was not intended to impose obligations, define rights, or interpret any statute or regulation. Its identified failure modes remain useful testing scenarios because the statutes and regulations governing disputes, notices, customer information, and error resolution apply according to their own terms.
Separately, the CFPB addressed automated handling of account information requests in a 2023 advisory opinion interpreting Section 1034(c) of the Consumer Financial Protection Act. That advisory opinion was withdrawn on May 12, 2025, as part of a broader withdrawal of agency guidance. Section 1034(c) remains in force, but the withdrawn opinion should be treated as a historical agency interpretation rather than current guidance.
The point is not that every chatbot message triggers every regime. Organizations should determine before deployment which communications the channel may receive, which legal rules attach to them, and what the system must recognize, preserve, and route.
The distinction is important:
- Output risk asks whether the organization said the wrong thing.
- Input risk asks whether the organization received something legally significant and failed to recognize or act on it.
The second problem can exist even when every sentence the chatbot generates is accurate.
Internal Chatbots Can Become Policy by Conversation
The same issue arises inside organizations. An internal chatbot may be offered as a convenient way to search policies, obtain benefits information, review a contract, or classify an incident. The practical effect may be to make the chatbot the place employees obtain organizational answers.
An internal assistant rarely needs authority to bind an organization directly. It supplies an answer to an employee who may already have authority to communicate, approve, deny, or escalate on an organization’s behalf. The governance question is not limited to whether the chatbot’s answer was correct. It also includes whether deploying the chatbot changed who exercised judgment, authority, or discretion in the first place.
An answer that leave is unavailable may implicate statutory rights. An answer that an incident is not reportable may become evidence in a later regulatory investigation. An answer that a contract needs no legal review may explain why an employee bypassed a required control.
An organization's presentation of the tool will matter. A drafting aid differs from a system positioned as the recognized channel for answers from HR, Legal, Compliance, IT, or Finance. If an organization routes employees toward the chatbot and away from those functions, a claimant may argue that reliance on the chatbot was not an unauthorized detour but the path the organization built.
For internal systems, the employee may become the delivery mechanism for the chatbot’s error.
Define the Reliance Boundary
Before deployment, organizations should define the chatbot’s reliance boundary: the line between information the system may safely provide and representations, notices, promises, or actions requiring stronger controls.
One useful framework has five levels. They are not legal categories but a governance continuum for deciding when stronger controls should attach.
- Inform. The chatbot retrieves low consequence, published information such as business hours or general procedures.
- Explain. The chatbot summarizes controlled sources, ideally with citations that let the user inspect the governing material.
- Apply. The chatbot interprets a policy, contract, or rule against a particular customer’s account or an employee’s circumstances.
- Promise. The chatbot represents that the organization will waive, refund, approve, extend, investigate, or refrain from something.
- Act. The chatbot executes a transaction, changes an account, records a dispute, or otherwise modifies the organization's legal or operational position.
A system suitable for informing may be unsuitable for applying. A system that can describe the dispute process may not be equipped to recognize that the user has just initiated one.
Controls should strengthen along the continuum. At the lower levels, the priorities include approved and version-controlled sources, visible citations, and rules for expressing uncertainty. At the higher levels, they include detection of legally significant statements, mandatory escalation, prohibited answer categories, authority limits, and human approval before action.
Throughout the continuum, organizations should preserve appropriate transcripts and, where available, the source references, versions, and retrieval context supporting consequential answers. Organizations also need severity-weighted monitoring and the ability to correct errors at scale.
A general instruction to “review important information” does not define the reliance boundary. It transfers the boundary question to the user without telling the user where the boundary is.
Organizations need to decide what counts as important before the user supplies the test case.
Measure the Errors That Matter
Aggregate accuracy is a useful technical measure; it is not a legal risk assessment. A chatbot can answer nearly every question correctly and remain unsuitable for the narrow questions it gets wrong. One wrong answer about office hours and one wrong answer about a statutory deadline may count equally in a comfort statistic. They do not present the same legal risk.
A more meaningful assessment measures errors by subject and function, with particular attention to money, deadlines, rights, eligibility, and account status. It tracks unsupported promises, contradictions with governing policies and contracts, failures to escalate or recognize legally significant communications, inconsistencies across languages and user populations, and errors repeated across conversations. It also asks whether affected users can be identified and whether the error can be corrected before its consequences expand. Testing should therefore be weighted toward high-consequence questions rather than treating every question as interchangeable.
Asking how wrong the chatbot can be, on which subjects, with what consequences, and for how long before anyone notices provides a clearer risk assessment.
A low hallucination rate can coexist comfortably with high legal risk if the remaining errors concern the subjects that users are most likely to rely on.
Build the Correction Workflow Before the First Wrong Answer
Every customer service and internal support channel is sometimes wrong. That does not make governance futile. It makes detection, containment, and remediation part of the deployment design. Before launch, organizations should decide:
- What triggers legal, compliance, or operational review;
- Who can suspend an answer category or the chatbot itself;
- How conversations, sources, and system versions will be retained and searched;
- How affected users will be identified and corrections delivered;
- Who may authorize refunds, waivers, deadline extensions, or other remediation;
- Whether and when regulators, counterparties, insurers, or other stakeholders must be notified; and
- How the underlying sources and workflows will be corrected, tested, and documented.
Much of that capability may depend on the chatbot vendor. The vendor agreement should address transcript retention and export, notice of changes to models and retrieval sources, cooperation in investigating and remediating errors, and allocation of responsibility when the system misstates organizational policy. A correction workflow the vendor contract does not support is a plan, not a control.
The scale that makes chatbot errors dangerous can also support more systematic remediation. If conversations are retained, searchable, and tied to identifiable users, organizations may be able to locate and correct repeated errors more consistently than they could after isolated mistakes by individual representatives. But that advantage exists only if the governance and records architecture preserves it. Organizations that cannot search their chatbot histories may possess an impressively scalable way to make mistakes and no comparably scalable way to find them.
Availability Is an Institutional Act
Making a chatbot available is an institutional act. An organization decides where the chatbot appears, what it can access, which questions it may answer, which statements it must escalate, whether it may act, what other guardrails are in place, and whether users have any practical reason to look elsewhere.
The chatbot may lack authority to amend a contract, waive a fee, or resolve a dispute. Its responses can still induce reliance, interfere with a legal right, influence an authorized employee, or leave the organization with consequences it must investigate, correct, preserve, or remediate.
Authority limits what the chatbot was permitted to do. It does not determine everything the organization must do afterward.
The content of this article is intended to provide a general guide to the subject matter. Specialist advice should be sought about your specific circumstances.
[View Source]