Home / Articles / AI adoption
AI adoption

A practical 90-day AI adoption plan for a real business

Move from a promising idea to a controlled, measurable pilot in 90 days without buying a platform first, skipping governance or asking a proof of concept to carry the whole transformation.

AnswersHow do we begin without creating avoidable risk?
AI adoption15 min read6th Meridian

How do we begin without creating avoidable risk?

A pilot earns the right to scale through evidence: a clear baseline, a bounded test and a decision gate.

Ninety days is enough time to prove whether an AI use case deserves serious investment. It is not enough time to transform an entire enterprise, solve every data problem, or automate every function. That is precisely why a 90-day plan is useful: it forces leaders to choose one meaningful problem, establish controls, test the difficult parts, and make an evidence-based decision.

The goal at day 90 should not be “we launched AI.” It should be a decision package that answers five questions:

  • Did the pilot improve a defined business outcome?
  • Was the quality sufficient for the intended use?
  • Can the organisation operate it safely and affordably?
  • Do the people who perform the work actually want to use it?
  • Should the company stop, refine, or scale?

Microsoft’s Cloud Adoption Framework for AI follows the same logic. It recommends beginning with business value, evaluating AI maturity and data readiness, prioritising use cases by impact and feasibility, starting with focused proofs of concept, and embedding responsible AI, governance, security, and operations from the beginning.

This plan translates that guidance into thirteen practical weeks for a nontechnical leadership team.

Ninety daysOne use case, carried all the way to a decision
  1. Days 1–14

    Mission and inventory

    Name the accountable team, then gather candidate use cases without judging them yet.

  2. Days 15–30

    Prioritise and map

    Pick one. Map its process, data, risks and the baseline you will be measured against.

  3. Days 31–45

    Prepare the pilot

    Access, controls, evaluation method and support model — not just the model.

  4. Days 46–60

    Build the narrow version

    The smallest thing that could be useful, tested against the baseline.

  5. Days 61–90

    Operate and decide

    Run it with real users, evaluate honestly, then continue, change or stop.

The programme is designed to produce a decision, not a demonstration. Stopping at day 90 with good evidence is a successful outcome.

Before day one: define what this programme is not

Set three expectations before work begins.

First, the team is not choosing an enterprise-wide winner among every AI platform. It is validating one use case and a reusable way of working.

Second, the pilot is not permitted to bypass normal privacy, security, legal, procurement, or operational controls because it is “only an experiment.” A small experiment can still expose sensitive data or produce a harmful decision.

Third, the team will not measure success through demonstrations, prompt counts, licences assigned, or the volume of generated content. It will compare a business process before and after the pilot.

These expectations protect the programme from two common failures: endless strategy work with no usable result, and uncontrolled experimentation that produces activity without trust.

Days 1-7: establish the mission and accountable team

Begin with a short, written mission. It should describe the business outcome, not a technology.

Weak mission: “Deploy generative AI across operations.”

Stronger mission: “Reduce the time required for service representatives to find an approved answer while maintaining or improving answer accuracy.”

Name an executive sponsor who can resolve priorities and remove barriers. Name a business owner who remains accountable for the process and outcome. Add a small cross-functional team with process expertise, technology, data, security, privacy or legal, risk, and change support. Include two or three people who perform the work every day.

Microsoft recommends centralised standards with cross-functional oversight, whether through an AI Center of Excellence or another governance team. A 90-day pilot does not require a large permanent department, but it does require clear authority and representation beyond IT.

Create a one-page charter containing:

  • the problem and target users;
  • the current business owner;
  • the intended outcome and baseline measures;
  • the data likely to be used;
  • known legal, security, privacy, and regulatory considerations;
  • the decisions the AI may support;
  • the actions it may never take during the pilot;
  • who can approve, pause, or end the pilot;
  • the day-90 decision criteria.

At the end of week one, every participant should be able to explain why the pilot exists in the same sentence.

Days 8-14: build a use-case inventory

Run structured interviews or workshops across a few business areas. Ask where people repeatedly search, retype, reconcile, summarise, predict, classify, route, or wait for an approval. Look for measurable friction rather than fashionable ideas.

For each opportunity, record:

  • frequency and volume;
  • current time and cost;
  • error or rework rate;
  • impact on customers or employees;
  • data required and data owner;
  • degree of process variation;
  • consequence of an incorrect result;
  • current systems involved;
  • user willingness to change;
  • whether ordinary process improvement could solve it.

Separate opportunities into practical types: prediction, detection, perception, generation, knowledge retrieval, or multistep agent action. This prevents the team from using a generative model where a rule, search index, or traditional predictive model would be more reliable.

Microsoft’s current guidance explicitly advises against using agents for structured, predictable tasks and recommends a classic RAG approach for static knowledge retrieval that does not require tools or multistep reasoning. The rule for week two is simple: prefer the least complex approach that can create the outcome.

Days 15-21: prioritise one use case

Score the candidate use cases across four dimensions.

Business impact: Does it support a funded priority? Can it improve cost, speed, quality, risk, revenue, or experience in a way the organisation can measure?

Feasibility: Is the necessary data available, accessible, and of sufficient quality? Does the team have the skills and systems to run a pilot?

User desirability: Does the use case solve a problem users recognise? Will they participate and provide honest feedback?

Risk: What could go wrong? Does the system affect people, money, legal rights, safety, regulated advice, critical infrastructure, or sensitive data?

Do not automatically select the highest-impact idea. A narrow, moderately valuable, low-risk use case can build capability and confidence more effectively than an ambitious customer-facing agent.

Microsoft’s AI adoption plan recommends internal, non-customer-facing proofs of concept where possible. Good first candidates include internal knowledge retrieval, document extraction with human review, meeting or case summaries, draft assistance, and forecasting that informs rather than automatically executes a decision.

Select one primary use case and one reserve. Publish the scoring and rationale so the choice is auditable.

Days 22-30: map the process, data, risks, and baseline

The team now studies the selected process in detail.

Map the current workflow from trigger to outcome. Include handoffs, systems, decisions, exceptions, queues, workarounds, and approval points. Observe actual work rather than relying only on a procedure document. The gap between the written process and the lived process often determines whether the pilot succeeds.

Establish a baseline using several weeks or months of representative information where possible. Relevant measures might include cycle time, search time, backlog, conversion, forecast error, document accuracy, first-contact resolution, customer satisfaction, rework, escalations, or cost per completed task.

Inventory the data. Record its owner, location, sensitivity, quality, access rules, retention requirements, geographic restrictions, and refresh frequency. Identify conflicting sources and nominate which one is authoritative.

Complete an initial responsible AI and security assessment. Microsoft’s responsible AI framework uses six principles: fairness, reliability and safety, privacy and security, inclusiveness, transparency, and accountability. Convert each into a practical question:

  • Could the use case disadvantage a group or create unequal quality?
  • What happens when the system is wrong or unavailable?
  • Does it expose data beyond a user’s entitlement?
  • Can all intended users access and understand it?
  • Will people know AI is involved and understand its limits?
  • Who remains answerable for the outcome?

Define the human control points. Any action that is hard to reverse or affects people, money, compliance, or safety should require approval. Decide how ambiguous cases are escalated and what information the reviewer needs.

At day 30, the team should have a signed pilot brief, data map, risk register, baseline, success measures, and control design.

Days 31-45: prepare the pilot, not just the model

The middle of the programme is where strategy becomes an operating system.

Prepare a representative test set. Do not use only clean examples. Include ordinary work, difficult edge cases, incomplete input, conflicting information, sensitive content, stale material, unusual users, adversarial instructions, and cases where the correct response is to refuse or escalate.

For a knowledge assistant, curate a limited collection of approved sources with owners, dates, permissions, and a refresh process. For document processing, collect several formats and scan qualities. For forecasting, check missing values, changing business conditions, and whether historical patterns contain bias. For generative assistance, define approved source material, tone, prohibited claims, and review responsibility.

Set up access using least privilege. Users and applications should receive only the data and actions needed for the pilot. Separate development, test, and production-like environments where appropriate. Avoid shared credentials. Establish logs, cost budgets, alerts, and a way to revoke access quickly.

Define the user experience. Good controls should be visible and usable: source citations, confidence or uncertainty cues, clear AI disclosure, edit and feedback controls, and an obvious path to a person. Do not hide the escalation route inside documentation.

Create a simple operating guide:

  • intended and prohibited uses;
  • required review steps;
  • how to report a bad result;
  • who provides support;
  • incident and shutdown procedure;
  • content or data refresh ownership;
  • pilot hours and participant group.

Days 46-60: build and test the narrowest useful version

Build enough capability to test the business hypothesis, not every requested feature.

Suppose the use case is an internal policy assistant. The narrow version might answer questions from one approved policy set, cite the source, respect user access, and abstain when no sufficient evidence exists. It does not need voice, multiple languages, autonomous actions, or integration with every department on day one.

Test in layers.

Functional testing: Does the system complete the intended task?

Quality testing: Is the prediction, extraction, retrieval, or generated answer accurate enough for the use?

Safety testing: Can it handle harmful, biased, manipulative, or sensitive input appropriately?

Security testing: Can a user retrieve unauthorised information, inject malicious instructions, abuse a tool, or exceed permitted actions?

Operational testing: What happens under load, service failure, missing data, slow response, or budget limits?

Human testing: Do intended users understand the result, identify mistakes, and follow the review process?

For RAG systems, Microsoft Foundry describes evaluators for retrieval quality, groundedness, relevance, and response completeness. For generative AI more broadly, Microsoft Foundry observability covers quality, safety, agent behaviour, latency, errors, token use, and monitoring. The pilot does not need every available metric, but it does need measures tied to its actual failure modes.

Run a formal readiness review before real work enters the pilot. Confirm that critical defects are closed, logging works, escalation is staffed, permissions are correct, and the business owner accepts the residual risk.

Days 61-75: operate a controlled pilot with real users

Start with a small, motivated user group and a limited volume of work. Keep the existing process available as a fallback. In higher-risk scenarios, use shadow mode: let the AI produce a recommendation while the current process remains authoritative.

Collect structured feedback, not only satisfaction. Ask:

  • Was the output correct?
  • What did the user change?
  • What evidence was missing?
  • Did the system save or add time?
  • Was the escalation appropriate?
  • Would the user trust it for this task again?
  • Did it change the final business outcome?

Review failures at least twice a week. Group them by root cause: poor source data, missing permissions, retrieval failure, ambiguous prompt, model limitation, user misunderstanding, integration defect, or process problem. Fixing the root cause is more valuable than adding a longer prompt around it.

Monitor cost per completed task. A pilot can appear efficient while consuming expensive capacity on low-value interactions. Include support, review, integration, and governance effort in the economic picture, not only model usage.

Do not silently broaden scope. New users, data, languages, actions, or business decisions change the risk profile. Put enhancements into a backlog for the day-90 decision.

Days 76-83: evaluate outcomes and operational readiness

Compare the pilot with the baseline. Separate activity from effect.

Activity measures include prompts submitted, documents processed, or summaries generated. They show use, not value.

Outcome measures include shorter cycle time, fewer errors, improved resolution, reduced backlog, more accurate forecasts, better conversion, or lower cost per completed task. These determine whether the pilot changed the business.

Review quality by segment, not only as an average. A document model may perform well overall but poorly for one supplier. A knowledge assistant may answer common questions accurately but fail on one country’s policy. A prediction may perform differently across customer groups.

Review risk evidence: unauthorised access attempts, unsafe output, ungrounded answers, incorrect consequential recommendations, user complaints, and failed escalations. Review operational evidence: availability, latency, cost, logs, incident handling, support load, and the time required to update content or configuration.

Assess adoption honestly. If users avoid the tool, determine whether the problem is trust, workflow disruption, insufficient quality, weak training, or lack of value. Making use mandatory can hide the reason without solving it.

Days 84-90: make the decision and publish the evidence

Prepare a concise decision paper with four sections.

1. Business result: baseline, pilot result, sample size, limitations, and financial interpretation.

2. Quality and risk: passed and failed criteria, material incidents, residual risks, and required controls.

3. Operating model: owner, support, monitoring, data refresh, security, governance, cost, and change responsibilities.

4. Recommendation: stop, refine and retest, scale within the same process, or expand to a new process.

Stopping is a valid outcome. A disciplined pilot that prevents a weak investment has created value. Refinement is appropriate when the business case remains sound but quality, integration, data, or adoption is not ready. Scaling is appropriate only when the outcome, controls, economics, and ownership are repeatable.

If scaling, define the next boundary. Add one department, region, document type, or integration at a time. Repeat risk and readiness review whenever the system gains new data, users, tools, or authority.

An illustrative scenario: a proposal knowledge assistant

Consider this illustrative scenario. A professional-services firm has teams spending hours searching old proposals, service descriptions, staff biographies, and approved contract language. The firm selects an internal RAG assistant as its first pilot.

During the first three weeks, it measures current search and drafting time, identifies authoritative sources, and discovers that many “approved” files are outdated. The data work becomes a benefit before the AI is even built.

By day 30, the firm narrows the pilot to one service line and two offices. The assistant may retrieve approved content and draft a response, but it may not create pricing, legal commitments, client references, or claims about staff availability. Every answer cites its source and date.

During build and testing, the team includes questions with no source, conflicting service descriptions, restricted client material, and requests to ignore policy. The assistant is required to abstain or escalate in those cases.

A small group of bid professionals enters the controlled pilot. They record whether each answer was accepted, edited, rejected, or escalated. The team measures time to an approved first draft rather than raw generation speed.

Suppose the day-83 review finds shorter search time alongside weak metadata and inconsistent content ownership. It also finds outdated citations that the review controls caught before use, with no restricted client content exposed. Users value the source links more than the drafting feature. These are hypothetical findings that illustrate the decision, not claimed client results.

The day-90 decision is to refine and expand slowly. The company funds a content-owner role, adds automated freshness checks, and extends the pilot to one additional service line. It does not yet enable autonomous proposal assembly.

This is a successful AI adoption outcome because the company learned what must be true for the capability to scale safely.

The minimum governance pack

Every pilot should leave behind reusable artefacts:

  • AI use-case charter;
  • business baseline and scorecard;
  • data and permission map;
  • responsible AI and security assessment;
  • human approval and escalation matrix;
  • test set and evaluation results;
  • model, prompt, source, and configuration change log;
  • user guidance and training record;
  • incident, shutdown, and rollback procedure;
  • cost and capacity report;
  • day-90 decision paper.

These artefacts form the beginning of an AI operating model. Microsoft’s AI Center of Excellence guidance describes an internal cross-functional capability that prevents fragmented or ungoverned adoption. A company does not need to create a large central team immediately, but it should reuse standards, test methods, controls, and lessons across pilots.

Guardrails throughout the 90 days

  • Start with a business outcome and a named owner.
  • Use the least complex technology that solves the problem.
  • Keep the pilot narrow and changes reversible.
  • Apply least privilege to users, applications, data, tools, and agents.
  • Do not use sensitive data without a documented purpose and approval.
  • Disclose AI use and explain limitations to participants.
  • Require human approval for consequential or hard-to-reverse actions.
  • Test groundedness, fairness, privacy, safety, security, reliability, and accessibility in proportion to risk.
  • Maintain a manual fallback and clear escalation path.
  • Monitor quality, cost, usage, failures, and policy compliance after launch.
  • Define who can pause the system and practice the incident procedure.
  • Reassess whenever scope, data, users, models, integrations, or authority change.

A balanced day-90 scorecard

Business: target outcome versus baseline, cost per completed task, cycle time, backlog, revenue or risk effect.

Quality: accuracy, rework, groundedness, completeness, false positives, false negatives, correct abstention, and exception rate.

People: active use, task completion, user confidence, override reasons, training completion, and accessibility feedback.

Risk: unauthorised access, harmful or biased output, consequential errors, complaints, privacy events, and policy violations.

Operations: availability, latency, unit cost, support effort, data freshness, monitoring coverage, incident recovery, and change lead time.

The value of a 90-day plan is discipline. It gives ambition a boundary, governance a practical role, and leadership a date by which enthusiasm must become evidence. By day 90, the organisation should know more than whether the AI can produce an impressive answer. It should know whether the capability improves work, earns trust, fits the operating environment, and deserves the next investment.

  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/strategy
  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/plan
  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/govern
  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/responsible-ai-policies
  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/center-of-excellence
  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai-agents/business-strategy-plan
  • https://learn.microsoft.com/en-us/agents/center-of-excellence/responsible-ai
  • https://learn.microsoft.com/en-us/azure/foundry/concepts/observability
  • https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/rag-evaluators
  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/platform/security
  • https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/govern/monitor-cloud-governance

This article is original editorial work for a business audience. Product facts were checked against the official sources above. Any performance figures are targets to validate against your own baseline, not vendor guarantees.

Bring us the operating need, risk, or opportunity. We will connect the people, security, AI and enterprise systems required to act.

Start a conversation

Come with a problem. Leave with a decision.

One working session. You leave knowing what to build first, what it needs, who owns it, and how you will know it worked.