Home / Insights / Onboard AI like a colleague

Onboard AI like a colleague, do not just buy it

AI is not a new CRM. It is a new way of working, with an additional pair of brains next to your colleagues, which is why deployment fails more often than technology does. Treating an AI system as a hire rather than a purchase forces the questions procurement leaves vague: what is the role, who is the accountable owner, what does probation look like, which permissions apply, when is performance reviewed, how is it decommissioned. The metaphor is operationally useful, and it has one hard limit. The AI Strategy Session

Four states of AI in an organisation: access, activity, adoption and value, with the evidence required at each stage

Why is AI adoption not the same as AI value?

Most organisations stop measuring at access or activity and call it adoption.

There are four distinct states, and conflating them is how AI programmes report success while nothing changes in the work.

  1. 1. Access. Employees can use a model or tool. Evidence: a licence or technical access record.
  2. 2. Activity. Employees use it in work. Evidence: usage logs and workflow observation.
  3. 3. Adoption. It is embedded in a repeatable process. Evidence: sustained use by a defined group.
  4. 4. Value. It improves a business outcome without unacceptable risk. Evidence: a baseline, a counterfactual, or pre and post measurement.

A serious programme reaches state four, with the outcome defined in advance rather than discovered afterwards. Licence counts and training attendance are states one and two dressed as progress. The gap shows up in the aggregate numbers: nearly 90% of organisations already use AI in at least one function, 60% see minimal returns, and only 5% reach outsized results (BCG).

What does the published evidence actually support?

Healthcare is the most instructive sector here because it publishes its failures. A 2025 BMJ Health and Care Informatics report describes most institutional generative AI efforts as pilots or early deployments with few full-scale institutional adoptions, and emphasises secure data access, interdisciplinary teams, bias analysis, regulatory alignment and structured evaluation. A 2023 global implementation study surveyed 2,525 decision-makers with AI experience across China, Germany, India, the United Kingdom and the United States, and interviewed sixteen implementation experts - useful evidence that implementation problems are organisational and cross-national, though it should not be quoted as a 2025 or 2026 statistic. The Stanford AI Index Report 2025, now in its eighth edition, expands coverage of corporate adoption and responsible AI practice.

Widely circulated pilot-to-production and abandonment percentages are a different matter. They vary dramatically depending on whether pilot means a technical experiment, a funded business trial or a production deployment, and whether success means usage, savings or profit. Publishing each study's sample, date, definition and outcome measure separately is more defensible than merging them into a single AI failure rate. I have stopped quoting the merged versions, and I would treat any adviser who quotes them without a definition as guessing.

Why does software procurement misfit AI?

Procurement assumes stable behaviour and deterministic acceptance testing, and AI offers neither.

Traditional software procurement assumes a defined feature set, predictable behaviour, testable acceptance criteria and a vendor responsible for maintenance. Generative and agentic systems introduce variable outputs, model updates, context sensitivity, prompt dependence, data leakage risk, uncertain failure modes and a continuing need for human review.

The NEJM AI ecosystem article makes this concrete. System performance depends on interacting components including the deep-learning model, training and update procedures, guardrails, data, prompts, interaction history, human behaviour, workflows and organisational resources, and it recommends local measurement and post-market-style surveillance for mission-critical uses ("AI as an Ecosystem", NEJM AI). Procurement does not end when the licence is signed. That single sentence is the entire argument for the onboarding frame.

There is a second mismatch that procurement cannot see: the process itself. Organisations fail with AI because they attempt to scale broken processes instead of stepping back to redesign them first, an observation Deloitte's implementation work and my own deployments agree on. Your teams have usually already fixed the official process with undocumented workarounds, which makes those workarounds the most valuable data in the organisation and the first thing to map.

One deployment I reviewed had a signed licence, a completed training day, an enthusiastic internal announcement and no named owner. Three months later nobody could say who decided a new use case, who stopped a failing one, or what the tool was supposed to improve. The technology worked. The hire had no manager.

How does the hiring metaphor map to governance?

Each employment step maps onto a governance question that procurement usually skips.
  • Employment metaphor - job description
  • Hiring manager
  • Probation
  • Access rights
  • Training
  • Performance review
  • Manager feedback
  • Offboarding
  • Operational equivalent - intended use, out-of-scope use, success criteria
  • A named accountable business owner
  • A time-limited pilot with predefined evaluation
  • Least-privilege permissions and data boundaries
  • Context, examples, escalation paths, user education
  • Regular quality, safety, cost and outcome review
  • Human correction, error logging, prompt and workflow updates
  • Revoked access, archived records, decommissioning

Wiley's human-centric AI guidance reaches the same practice from a different direction, recommending that AI systems require ongoing coaching, regular feedback, performance monitoring and continuous adjustment, and that you begin with the user's job to be done rather than the technology's capabilities (Dr Lisa Palmer, Show AI, Don't Tell It, Wiley).

Why does the frame work on people, not just systems?

It makes the human questions unavoidable. You cannot write a job description for an AI system without deciding what your people will stop doing, which is the conversation every organisation postpones.

The cost of postponing it is documented. 42% of C-suite executives say AI adoption is tearing their company apart, and 76% of executives believe their teams are excited about AI while 31% of employees actually are (BCG and Columbia Business School, 2025). That 45-point gap is not a communication problem. It is what happens when experienced professionals are asked to accept that the way they have worked through a successful career is no longer sufficient - and are then asked to scale a broken process they had already fixed themselves.

Which interventions actually address the failure modes?

Eight interventions, each matched to the specific failure it prevents and how you would know it worked.
  1. 1. Use-case selection. Prevents automating low-value or high-risk tasks first. Measure: baseline time, quality, risk and genuine user need.
  2. 2. Human-in-the-loop gate. Prevents silent errors and automation bias. Measure: override rate, error severity, escalation time.
  3. 3. Phased sequencing. Prevents premature autonomy. Simple prompts, then workflows with a human at every decision point, then agents. No one runs an ultramarathon on day one. Measure: compare the assisted, workflow and agentic stages separately.
  4. 4. Team co-design. Prevents workarounds and rejection. Measure: adoption by role, qualitative friction, observed process change.
  5. 5. Executive personal use. Prevents symbolic sponsorship without understanding. Measure: leadership usage tied to visible workflow change.
  6. 6. Literacy programme. Prevents misuse and over-trust. Measure: scenario-based competence rather than attendance.
  7. 7. Error-tolerance policy. Prevents hidden failures and non-reporting. Measure: incident rate, near misses, reporting volume.
  8. 8. Kill or continue review. Prevents zombie pilots. Measure: a quarterly value, risk and strategic-fit decision.

What does governance built in rather than bolted on look like?

The Barcelona GenAI Health Hackathon at Hospital Clínic offers a concrete model. Multidisciplinary teams worked with anonymised clinical data in a secure environment, with access policies, bias analysis, performance evaluation and regulatory alignment built into the process rather than added afterwards (BMJ Health and Care Informatics, 2025). The transferable lesson is sequencing: governance and evaluation were part of the build, not a review at the end.

The BMJ's clinical documentation work adds the practice that generalises furthest beyond medicine - maintain approved-tool registries, training, performance monitoring and incident reporting, and document explicitly how AI contributions were accepted, modified or rejected. Any organisation can adopt that last one this quarter at almost no cost, and it is the single best defence against the question "who decided this".

Where does the metaphor mislead?

An AI system is not an employee, and pretending otherwise moves accountability to something that cannot hold it.

It has no moral agency, no employment relationship, no interests and no independent accountability. Anthropomorphic language can lead users to overestimate understanding, underweight uncertainty, or treat a confident explanation as evidence of intention. That is not a pedantic objection. It is the mechanism behind automation bias, where fluent output gets approved rather than investigated, and it is documented most clearly in clinical settings where the stakes forced someone to study it.

The stronger governance principle: manage the system with the discipline you would apply to a high-impact operational capability, while keeping responsibility explicitly human. BMJ guidance states that clinicians retain responsibility for final decisions, and that organisations must supply safe tools, procurement standards, training, monitoring and incident reporting. That principle transfers to every sector without modification.

So use the frame for what it does well - forcing role definition, ownership, permissions, evaluation, feedback and an exit plan - and drop it the moment it starts to sound as though the system shares the blame. It does not. You do.

What should you measure after deployment?

The outcome you defined before deployment, plus the risk you agreed to accept.

Time per task. Share of outputs passing the quality gate first time. Number of people actually using the workflow. Override rate and error severity. Cost per unit of output. Incidents and near misses reported. And one line nobody writes down: what result would make you stop.

There is one predictor worth more than any dashboard. Personal, hands-on governance of AI by the chief executive shows the strongest link to AI-driven EBIT of any factor measured (Thomson Reuters). Not the stack, not the infrastructure budget - the calendar. A sandbox budget for experiments, a protected slot to actually try things, and senior people visibly stepping into learning shoes do more for adoption than another licence tranche. Alongside that, keep some capability deliberately manual: Carnegie Mellon and Microsoft Research work with 319 knowledge workers found higher confidence in the model associated with less critical thinking, and Wharton reports 43% of organisations already observing declining employee qualification.

Frequently asked questions about onboarding AI

Is hiring AI just a metaphor, or does it change anything operationally?
It is an operating metaphor with real governance consequences. It usefully forces role definition, an accountable owner, permissions, review cadence and an exit plan - all things software procurement leaves vague because it assumes stable behaviour. It becomes misleading when it implies shared accountability, because responsibility stays human regardless of how the system is described. Use it to structure the decisions, not to describe the technology. See the AI Strategy Session.
Who should be the accountable owner for an AI system?
A business owner senior enough that findings cannot be shelved, not IT by default. IT owns implementation and security; the business owns intended use, success criteria and the decision to stop. Without a named owner the deployment drifts into the state I see most often - a working tool nobody is responsible for, which means no one can authorise a new use case or retire a failing one. Name the owner before the licence, not after.
What does probation look like for an AI system?
A time-limited pilot with predefined evaluation: which metric, on which process, verified by whom, decided on which date. No open-ended trials, because they become unpaid integration projects that never produce a verdict. The evaluation should include a kill criterion written before you start - the result that would make you stop - since that is the number people avoid defining and the one that prevents zombie pilots.
Should we start with customer-facing use cases?
Rarely. Start internal, low-stakes and high-frequency, where errors are cheap and trust can be earned before anyone outside sees the output. Then move outward once the quality gates have been tested against real failures. The selection criteria matter more than the ambition: painful enough, frequent enough, structured enough, with a bearable cost of error. If a mistake in that output is unacceptable, it is not a first candidate. See choosing the first process.
How do we avoid automation bias in practice?
Design gates where a human must actively decide rather than passively approve, log every override, and make error reporting safe enough that people actually do it. Fluent output is easier to approve than to investigate, especially under time pressure, which is why the countermeasure has to be structural rather than a reminder in a training deck. Tracking override rate tells you whether the gate is real or ceremonial.
Why do you not quote AI failure rate statistics?
Because studies define pilot, success and agentic AI differently, so merging them produces a number that cannot survive scrutiny in a board meeting. I report the definition alongside the figure or leave the figure out. If you want a defensible internal number, measure your own: how many use cases entered evaluation, how many reached sustained use by a defined group, and what each cost. See the AI Strategy Session.
What is the fastest governance improvement we can make this quarter?
Document how AI contributions were accepted, modified or rejected in each workflow where AI is used. It is the practice BMJ recommends for clinical documentation and it costs almost nothing to adopt elsewhere. It creates an audit trail, surfaces which outputs never get corrected - usually a sign the gate is ceremonial - and answers the question that follows any incident: who decided this, and on what basis.
Does this apply to agentic AI or only to assistants?
It applies more strictly to agents, because autonomy raises the cost of every gap in the frame. The sequencing rule exists for that reason: prompts, then workflows with a human at each decision point, then agents, comparing the stages separately rather than assuming the last one inherits the trust earned by the first. Definitions of agentic AI vary between vendors, so specify what yours does before agreeing to any benchmark about it.

Diagnosis first, ideation last. The AI Strategy Session runs a readiness audit, an employee and manager survey, facilitated strategy days and a twelve-month follow-through loop. Owners and metrics before anyone leaves the room.

See the AI Strategy Session · Book a 30-minute call

Read nextWhy borrowed prompts write like everyone else · Why strong technology loses the pitch · The AI Strategy Session

What changed in this article (7 September 2026)
  • First published, built from the Ukrainian keynote on making AI a full member of the team
  • Added the four-state access to value model and the NEJM AI ecosystem evidence
  • Added the section on where the employment metaphor misleads
  • Replaced merged AI failure rate figures with definition-specific sources