Applied AI

Doing AI to Say We Have AI

When organizations adopt AI to prove they have AI, they do not create strategy. They create cost, risk, complexity, and a public reminder that hype is not architecture.

A formal ribbon-cutting ceremony on open ground: a freestanding doorframe is planted in the dirt with a teal ribbon stretched across it, a dignitary holds oversize ceremonial scissors at the ribbon, a small group of polite attendees applauds on the left, photographers and a lectern with a microphone stand on the right, and the horizon behind the doorframe is empty.

Some enterprise AI initiatives begin with a need to be seen using AI rather than with a defined problem. The chatbot, copilot, or “intelligent” search box becomes a visible signal of momentum before anyone has named the workflow, owner, limits, or measure of value.

The problem is not using AI. Deliberate, scoped AI connected to a real workflow can create value. The problem is using AI as a status signal. That choice turns a demonstration into an operating commitment before the organization has decided what the system is for.

A useful test is whether the initiative still makes sense without the launch announcement. If its value depends mainly on being visible, the organization is optimizing the signal before it has designed the capability that must remain after attention moves elsewhere.

Why AI theater starts

Competitors announce AI. Boards ask about strategy. Vendors bring polished demos. Leaders want evidence of momentum. Under that pressure, the organization asks “Where can we put AI?” before it asks “What problem deserves AI?”

The first question predetermines the answer: a feature, surface, or demo. The second allows AI to compete with rules, traditional software, automation, search, process improvement, and the possibility that the real problem is poor data or unclear ownership.

Descriptive evidence should not be mistaken for causal proof. Gallup workplace surveys document growing use of AI at work. They do not establish why a given organization starts a project or why it fails. RAND’s 2024 report draws from 65 interviews with experienced data scientists and ML engineers and identifies leadership-related anti-patterns, including unclear problems and pressure to demonstrate AI activity. That sample supports a credible pattern, not a universal failure rate for all enterprise initiatives.

AI theater persists because it looks like progress. A demo is easy to present and difficult to criticize without sounding resistant to innovation. Once launched, political and operational costs make the system harder to withdraw. The organization starts defending a capability it never fully decided to build.

The public surface reveals the missing architecture

Public chatbots are not the whole problem. They are the clearest scene because customers can test the gap directly. DPD’s assistant produced inappropriate output and criticized the company. Air Canada’s chatbot gave a customer inaccurate guidance, and the British Columbia Civil Resolution Tribunal held the airline responsible for information presented through its channel.

These incidents do not prove that every chatbot is reckless or that generated text has unlimited authority. They show a narrower point: a public AI surface inherits the responsibilities of a public production system. Scope, refusal behavior, approved sources, escalation, monitoring, and accountable ownership must exist before launch.

The same gap appears inside the enterprise: a copilot nobody owns, generation layered over a stale knowledge base, or an agent connected to workflows without contracted interfaces. The surface changes. The architectural questions do not.

The strongest warning from the public incidents is not that a chatbot used surprising language. It is that the organizations had created official channels whose boundaries were unclear to users and, in some cases, to the organizations themselves. A production surface needs an approved purpose, a defined source of truth, a path to a person, and a response plan for outputs that escape the intended scope.

That does not make customer-facing AI categorically wrong. A narrow assistant over maintained information, with visible limits and reliable escalation, can reduce friction. The risk rises when the interface implies broad competence, uses open-ended generation for policy or transactional guidance, or lets fluent answers outrun the system’s actual authority.

Internal copilots deserve the same discipline even when failures are less public. An inaccurate summary can enter a case record. Generated code can carry an insecure assumption. A search answer can omit the document that changes the decision. The effect may surface as quiet rework rather than a viral screenshot, which makes measurement and ownership more important, not less.

The direct model charge may be material, but it is only one part of total cost. Other costs include retrieval and tool calls, rate abuse, support, monitoring, content maintenance, security testing, legal review, vendor change, incident response, and engineering attention diverted from higher-value work.

The relative size of those costs depends on the system. The important discipline is to estimate them before launch rather than treating the API invoice as the full business case.

Strategy begins with a problem and an operating model

Valuable enterprise AI often starts in less visible work: ticket classification, case summarization, approved-source search, document review, data extraction, anomaly explanation, or controlled drafting behind a qualified reviewer.

The common shape is narrow. The workflow pain, delay, cost, or error rate is documented before the model is selected. The job has explicit non-goals. Data sources and refresh cadence are approved. Refusal and escalation behavior are tested. Cost and rate limits exist. An accountable owner can explain success, failure, and shutdown conditions.

Sometimes the right decision is not to use AI yet. A stale knowledge base needs maintenance before generation. Inconsistent processes need standardization before copilots. Agents need stable interfaces rather than improvised screen scraping. Predictive systems need measured outcomes. Automation needs an owner.

Poor inputs affect systems in different ways: training data can shape a model, retrieval data can shape an answer, and evaluation data can hide failure. The remedy is not the vague instruction to “clean the data.” It is to identify which data enters which stage, who maintains it, and how quality is tested.

A production plan must also distinguish a model from the system around it. Retrieval, prompts, tools, permissions, user interface, review, telemetry, and support all affect the outcome. Model selection matters, but a benchmark score does not answer who may use the capability, which records it may reach, or what happens when an external dependency changes.

That distinction is where many demos become misleading. A controlled example uses selected inputs, a cooperative user, and a prepared path. Production brings stale records, ambiguous requests, inaccessible dependencies, bursts of demand, adversarial input, exceptions, and people who interpret fluent output as authoritative. The gap is not evidence that AI cannot work. It is evidence that the demo did not test the operating system around it.

Cost needs the same system view. Inference, retrieval, and tool calls are variable costs. Evaluation, observability, security testing, content maintenance, access reviews, incident response, and support are operating costs. Integration work, migration, and vendor changes are lifecycle costs. A useful estimate separates those categories and ties them to expected volume, service levels, and error handling.

Public surfaces add brand and legal exposure because generated text appears in the organization’s voice. Internal systems add a different risk: employees may quietly route decisions through an assistant whose authority was never defined. In both cases, labeling the output as AI does not resolve accountability. The interface should communicate uncertainty, approved use, escalation, and the point at which a person must decide.

Human review is not a universal safeguard. Reviewers need context, competence, time, and authority to reject the result. If they are expected to approve high volumes or cannot inspect the evidence behind an answer, the review step may transfer responsibility without adding control. The design should specify what the reviewer sees, what must be checked, and what happens after rejection.

Agents raise the stakes because they can act through tools. A stable API does not make an action safe by itself, but it creates a contract that can be authorized, logged, tested, rate-limited, and revoked. Screen automation against an interface built for people is more fragile and harder to constrain. The right sequence is to define the action boundary first, then decide whether a model should operate inside it.

Measurement should start before the pilot. If the current workflow has no baseline for time, quality, cost, backlog, or error, the team cannot distinguish improvement from novelty. Usage and satisfaction can help, but they do not replace operating outcomes. A frequently used assistant can still create rework; a narrowly used one can be valuable if it improves a high-cost decision.

Finally, the system needs a lifecycle. Owners should know when to change a prompt, replace a model, refresh a source, suspend a tool, notify users, or retire the capability. That work is less visible than launch. It is also the difference between an experiment the organization can learn from and a permanent surface no one feels authorized to remove.

The architecture test

This test separates a useful experiment from a status project. The more questions that lack a concrete answer, the less ready the initiative is to become an operating capability.

  • If we removed the word “AI,” would the problem still matter?
  • Do we have evidence of pain, delay, cost, risk, or error?
  • Why is AI better here than a simpler alternative?
  • What must the system never do?
  • What happens when the output is wrong or abused?
  • What data may it use, and who maintains that data?
  • Can we measure value within a defined window?
  • Who owns operation, incident response, and downstream effects?
  • Are cost and usage limits set before launch?
  • What condition triggers shutdown?

Leadership should ask for these answers, not for a larger count of AI surfaces. Surface count is a visibility metric. Operating discipline is the capability.

Key takeaway. If the owner cannot name the shutdown condition before launch, the initiative is still a demonstration, not an operating capability.

Source notes

  • RAND, “The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed” (2024). Based on 65 semi-structured interviews with experienced data scientists and ML engineers. The report identifies recurring leadership anti-patterns, including unclear problem definitions, use of AI where simpler tools may fit, and pressure to demonstrate activity. The article treats those findings as qualitative evidence, not as a population-wide causal estimate. Source: RAND RR-A2680-1.
  • Gallup, “AI Use at Work Has Nearly Doubled in Two Years” (2025). Workplace survey evidence describing adoption and usage patterns. It supports the claim that workplace AI use is growing; it does not establish why individual projects begin or fail. Source: Gallup Workplace.
  • Gartner, “Hype Cycle for Generative AI” (2025). Industry framing for the difficulty of moving from inflated expectations and proofs of concept into production. It is context for the pattern, not causal evidence for specific failures. Source: Gartner.
  • DPD chatbot incident. Evidence of brand and operational exposure when a public assistant produces inappropriate output. Source: The Guardian (Jan 2024).
  • Air Canada chatbot decision. Evidence that an organization can remain accountable for information provided through its official chatbot. The article does not generalize the tribunal’s decision beyond that context. Sources: Moffatt v. Air Canada, 2024 BCCRT 149; BBC (Feb 2024).
  • OWASP Top 10 for LLM Applications. Prompt injection and unbounded consumption illustrate two risks that public AI systems should address through design and operation rather than through a system prompt alone. Sources: LLM01: Prompt Injection; LLM10: Unbounded Consumption.

Discuss this article

Thoughtful comments, corrections, and notes from real-world practice are welcome. Discussion is managed through GitHub.

This space is intended for meaningful technical discussion, useful corrections, and field experience. Spam, personal attacks, low-effort comments, and vendor pitches may be removed.

Continue

Place AI deliberately, not decoratively.

Explore more structured writing on architecture, applied AI, and operating discipline, or use the speaking page for topic and audience fit.