Point AI at Your Operation, Not Your Software

Point AI at Your Operation, Not Your Software

Part Two of MissionViewpoint's series: AI doesn't create organizational intelligence. It amplifies it.

In the first article in this series, I drew a line between using AI and building the organizational capability AI can amplify.

There's a reasonable objection to that distinction: building sounds like engineers, integrations, infrastructure—a year of work and a budget I don't have. Using the AI features in the tools I already own is the realistic version. Building my own is for organizations further along than mine.

It doesn't have to.

The smallest thing that counts as building rather than using is a loop you can run by hand this month.

Export a few reports. Reason across them with AI. Put the result in front of an operator who knows the operation. Capture what they know that the reports didn't. Change one thing, and see what happens.

That's not a feature you bought. It's capability you built, and it runs on exported spreadsheets and one operator's judgment, not on infrastructure you don't have.

The point isn't that the manual version is where you stay. It's that building was never gated behind a sophisticated stack. It's gated behind whether you run the loop at all - and you can run the first turn of it today, whatever your architecture looks like.

But I'd make one thing clear from the beginning: this is an operations initiative, not an IT initiative.

Technology and data teams may need to establish the environment, provide access to information, and make sure the appropriate security and governance are in place. But operations should own the question being asked, the people involved, the change that follows, and the outcome being measured.

And someone needs to guide the loop.

It may not be a new job. It could be an operational leader, an analyst, a transformation lead, or a technically fluent operator. Think of that person as the sherpa: the one responsible for moving the problem from question to data to AI to operator judgment to implementation, and then back around again.

You don't need the organizational structure figured out before you begin. But you do need someone responsible for making sure the experiment becomes an operating change rather than another interesting AI demo.

One other thing has to be settled before the first report leaves a system.

Before You Paste Anything In

If the reports contain protected health information, the AI environment has to be approved for that use.

That means treating AI like any other technology that will receive sensitive healthcare information: going through the appropriate privacy and security review, understanding how the vendor handles the data, and, where the vendor is acting as a business associate, having the appropriate Business Associate Agreement in place.

A consumer AI account shouldn't be treated as an acceptable place to experiment with PHI simply because it's convenient.

The specific products, contractual terms, and eligible tiers change quickly enough that providers should confirm them directly with the vendor and their own privacy and security teams before any protected information moves.

And even in an approved environment, keep pulling only what the question needs. The fact that a tool is permissible for PHI doesn't make it wise to dump every identifiable field into it. The discipline should still be to assemble the minimum information necessary to investigate a specific operating problem.

This is the part of the experiment you shouldn't shortcut.

Once the environment is appropriate, the rest can be surprisingly simple.

Reason Across the Reports, Don't Summarize Each One

Say the operating question is: why are some clients taking much longer to start care than others?

You export the four or five reports that touch it—referrals with dates and geography, open capacity, the recruiting pipeline, credentialing status, historical service-start timing—and trim them down to what the question needs.

The naive thing to do now is feed them in one at a time and ask the AI to summarize each. It will, competently. You'll get a tidy paragraph on your recruiting backlog, another on open capacity, another on where referrals cluster—and you may have learned very little you couldn't have read off the reports yourself.

The useful move is to put the relevant information in front of the tool together and ask the question none of the individual reports can answer on its own.

Given these extracts, which clients currently in intake share characteristics with clients who historically took longest to start? What combinations of geography, availability, staffing, and credentialing appear repeatedly among delayed starts? Where do these reports seem to tell different stories about the same client?

This is where a general-purpose model becomes useful: not summarizing each report independently, but helping you interrogate the relationships across them.

Time to care was never a fact sitting in one report. It's an outcome produced by multiple parts of the operation at once.

You'll get something rougher than a real production capability would produce. Some conclusions may be wrong, some relationships coincidental, some records won't line up cleanly across systems. That's fine. At this stage you aren't looking for an autonomous decision-maker or a polished prediction engine. You're looking for a specific, checkable set of hypotheses and cases that you can now hand to someone who knows better.

Which is exactly where it goes next.

Put It in Front of an Operator. Keep the Disagreement.

Whatever the AI hands back isn't the output. It's the setup.

Suppose the analysis flags fifteen clients currently in intake as unusually likely to experience a delayed start. Take that list - not a dashboard of the whole pipeline, just the fifteen the model is worried about - and put it in front of the person who actually runs that region. Not to check whether the AI is right, but to hear where she disagrees, and why.

She reads it and agrees that eleven really are at risk. But she immediately recognizes why the other four are different.

Two aren't staffing problems at all. Those families have very narrow availability the reports never captured. A third is waiting for a school schedule to settle before the family will commit to session times. The fourth looks difficult on paper, given geography and current staffing, but she knows an experienced technician is transferring into that area next week - a fact that lived in her head and nowhere else.

The AI found a pattern.

She supplied the context the pattern couldn't see.

And those four disagreements may be worth more than the eleven cases the tool got right.

The eleven confirm the operation can already recognize what it can see. The four expose exactly what it can't: a data element that isn't captured, a judgment that lives in one person, a signal that arrives too late to matter, an exception nobody thought to represent.

That's the material.

So capture it—not with a three-day knowledge-mapping offsite, but in a few seconds attached to a real decision.

Why doesn't this one worry you? What did the analysis miss?

A sentence or two, written down. Do it across a handful of cases and you have something no dashboard produces: the operation's own judgment, made visible enough to examine.

Don't assume every disagreement means the operator should win, either. Maybe she's identifying a legitimate exception, or relying on information the system simply doesn't have. But maybe she's following an old habit, applying a rule inconsistently, or missing a pattern that only becomes visible across hundreds of cases. That ambiguity is part of the value.

The objective isn't to make AI agree with experienced operators, or to make operators defer to the AI. It's to make the reasoning on both sides visible enough to test.

Don't optimize for the AI always being right. Make sure the organization learns something useful when it isn't.

That's the step that separates this from the AI experiment many organizations have already run - the one where an analyst asks a sharp question, the tool gives a sharp answer, someone builds a deck, and everyone goes back to work unchanged.

The deck version stops at the answer. This version treats the answer as bait for the disagreement, and the disagreement as the thing worth keeping. Same tools, completely different outcome. The difference is whether you stop at the output or mine the override.

And you can do all of this by hand—exported reports, one appropriately governed AI tool, one operator, a notes file. No pipeline, no integration, no data spine. The loop runs on manual labor, and it still runs.

Change One Thing. Then Outgrow This.

The loop runs. Now close it—because a loop that stops at insight is just a smarter version of the deck.

Change one thing in the operation.

Not ten. One.

Maybe intake starts capturing family availability as a real field because you just learned it materially affects time to care. Maybe recruiting gets visibility into anticipated staffing gaps before a client clears intake instead of after. Maybe the regional lead gets a short list of at-risk starts each week instead of another dashboard showing the whole pipeline.

Pick the change the operator's disagreements pointed toward most clearly, make it, and watch what happens to the outcome you were trying to move.

That's the full loop, run once, by hand:

Operating question → Relevant information → AI analysis → Human judgment → Captured knowledge → Operational change → Outcome

You didn't need connected data to do it. That was the point.

Now be equally honest about what you just built. This version doesn't scale. It doesn't persist. The operator's judgment lives in a notes file, not in anything the next analysis can automatically reach. It doesn't run itself; someone has to re-export and re-trim those reports every time. The insights don't appear where people actually work. And very little compounds: run it again next month and you start much of the way back at zero, because what you learned hasn't been wired into the next analysis.

The manual loop proves the loop works. It does almost nothing to make it keep working.

And that's the point of running it first.

Everything tedious about the manual version - the re-exporting, the re-trimming, the judgment disappearing into a notes file, the standing start every month - is a precise description of what a more durable version needs to fix. You don't discover that by reading an argument about infrastructure. You discover it the second time you rebuild the same five reports by hand.

The manual loop doesn't just build a little capability. It tells you what the lasting version needs to do - which problem was worth the effort, which information needs to be connected, which judgment is worth keeping, which missing data keeps tripping you up, and where the resulting intelligence needs to enter the workflow.

That's when investments in integrations, a data spine, persistent context, and workflow applications start having a specific purpose. You're no longer building AI infrastructure because AI seems important.

You're removing friction from a learning loop you've already proven matters.

Who Actually Does This?

None of this requires an AI department.

But it does require being clear about who owns which part, because the loop crosses boundaries that don't always talk to each other.

Operations owns the substance. The problem worth solving, the priorities, the change that follows, and the outcome being measured. That's the reason this is an operations initiative rather than an IT one.

Data and technology own the enabling layer. Access to the information, a sound technical environment, and the security and governance required to use it appropriately.

Frontline experts supply and test the operating judgment. But don't turn them into the project.

Your best regional operator shouldn't disappear into a six-week documentation exercise. Her highest-value contribution happens in small moments attached to real work.

Why did you reject this one? What was the analysis missing? Which exception applies?

A few seconds of structured judgment attached to an actual decision may be worth more than another forty-page SOP nobody opens. The experts supply and test the operating judgment. They don't run the initiative.

And someone has to keep those groups connected. That's where the sherpa comes back in.

As the manual experiment becomes repeatable capability, I would think of that person more formally as the capability owner.

Someone has to prioritize which problem comes next. Someone has to coordinate operations, frontline experts, data, and technology. Someone has to determine whether what the organization learned deserves to become a new data element, an operating rule, a workflow change, or additional context for the AI. And someone has to manage the change when that intelligence starts showing up in people's work.

The capability owner also has to watch two things at once.

Did the operating outcome improve? Are difficult starts caught earlier? Are interventions happening sooner?

And did the decision process improve? Are recommendations becoming more useful? Where are people still overriding them? Are the same exceptions appearing repeatedly? What information is still missing?

An override isn't necessarily evidence that the AI failed. It may be evidence that an experienced operator knows something the organization still hasn't captured.

Over time, the question that matters isn't whether the AI gets better on its own. It's whether the combination of AI, operating context, and human judgment produces better decisions.

That's the capability you're actually building.

Run the First Turn

So don't start by asking what your enterprise AI architecture should look like.

Start with an operating problem worth solving. Make sure you have an appropriate environment. Pull the minimum information the question requires. Reason across it. Put the result in front of someone who knows the operation. Keep the disagreement. Capture what it teaches you. Change one thing. Measure what happened.

Then run it again.

Capability isn't a feature you switch on.

It's a loop you run.

Run the first turn by hand, and you'll understand what the durable version is for better than any architecture diagram could tell you. You'll have built the argument for it yourself.

And once you do make that loop durable, something else begins to happen. The corrections, exceptions, operating judgments, and outcomes stop disappearing after each analysis.

They accumulate.

Which raises a different question:

Whose intelligence is actually compounding?

That's where I'll turn in the final article in this series.