Somewhere in most mid-size companies, someone spends a real chunk of their week pulling information out of a few different business systems and a handful of supplier emails or portals, and turning it into a spreadsheet or report that someone else then acts on. A weekly summary of which suppliers are running behind. A daily list of orders at risk of shipping late. A side-by-side comparison of price quotes from different suppliers. None of it involves actually placing an order or approving a payment. It's the research and the pulling-together of information that happens before someone makes that call, done by a person because there was no credible alternative.

There is now a credible alternative for a meaningful slice of that work. AI agents that can look things up across internal systems, reconcile information that doesn't quite match between them, and draft the report a planner used to write by hand are a real, buyable, buildable thing in 2026. Not the fully autonomous kind that places orders or approves invoices on its own, but the research-and-report kind. The catch is that most attempts to build this kind of automation don't make it to production. McKinsey's 2025 global AI survey found that roughly a quarter of organizations are meaningfully scaling an AI agent in at least one part of the business. That's a real number, but still a minority, and most of those are scaled in only one or two areas. Set against that, MIT's widely cited research on enterprise generative AI concluded that the large majority of pilots never produce a measurable financial return, and pinned the cause less on the quality of the AI itself than on projects that were never narrow or specific enough, and were never built to actually learn from how the work really gets done.

So this isn't a piece arguing that you should do this. It's a piece about how to do it so it doesn't become one of the projects that stalls.

The shape of the work worth automating first

The unit of work that succeeds here is not "automate supply chain planning." That's the framing that gets projects canceled. Gartner has pointed to unclear business value and open-ended, escalating costs as the leading reasons it expects a large share of these projects to be shelved by 2027, and has been blunt that a lot of what gets marketed as intelligent, autonomous AI is really just existing chatbot or automation tooling repackaged with a new name.

The unit that works is a single, named, recurring process with a human at the end of it who currently produces a document or a spreadsheet by hand. In supply chain and purchasing, that looks like:

A good candidate has three properties: it happens on a schedule, it draws from two or three systems you can name in plain terms (the system that tracks orders, the one that tracks warehouse stock, a supplier's portal), and it currently costs a specific, measurable number of hours. If you can't name where the information comes from or how many hours it currently takes, it isn't ready to scope yet. That's usually a sign you're describing a whole department, not a process.

Does the math actually work?

The calculation here is simple enough to write out as a comparison, and it's worth doing explicitly rather than trusting a vendor's promise:

what it costs to build and connect the tool, plus what it costs to keep running every month, against the hours it saves each week multiplied by what that person's time is actually worth to the company.

A couple of things matter before you run that comparison. First, for this kind of work, a narrow, recurring reporting task, the ongoing running cost is almost never the number that matters. Producing a weekly report costs very little to actually run; the hours of a person's time it replaces cost considerably more. The exception is a poorly designed tool that gets stuck repeating itself, which is a documented real risk and the reason sensible usage limits belong in the design from day one, not as an afterthought.

Second, the real financial risk isn't the running cost at all. It's spending the setup effort on the wrong workflow. That's the mechanism behind most of the "unclear business value" cancellations analysts describe: the project was approved before anyone measured a real, current baseline against it.

We're deliberately not putting specific figures on any of this. Published cost ranges for projects like this vary enormously depending on who's quoting them, and this particular corner of the market is dominated by vendors marketing their own services. The only number that should actually decide whether to proceed is your own measured baseline, compared against a quote scoped to your actual workflow, not a number from a blog post. What analysts do consistently report is that well-scoped projects like this tend to pay for themselves within the first year of use.

Bottom line: for a real, recurring, well-defined reporting task, the economics almost always work out. The risk was never the cost of running it. It's spending the setup effort on a process that turns out to be too rare, too undefined, or too dependent on human judgment to have a clean baseline in the first place.

How long this actually takes

Two different questions hide inside "how long," and they have different answers.

How long to observe the process before designing the workflow: for a single process you've already identified, the consistent range across people who actually do this work is one to four weeks, done by shadowing the person who performs the task today, not their manager. The most commonly cited cause of expensive rework in projects like this is exactly that substitution: discovery conversations held with a manager or a stand-in instead of the person who actually knows where the information gets messy, which surfaces the missing detail much later and forces the project to reset. If you haven't yet decided which process to target (say, you're weighing a few candidates across supply chain, purchasing, and finance), add a few more weeks for that decision.

How long from kickoff to a workflow you'd actually trust: for one narrowly scoped project, a realistic path is observation (two to four weeks), building and connecting it to your systems (four to eight weeks), then running it alongside your existing manual process for a minimum of thirty days before anyone relies on its output unsupervised. All in, that's roughly two to four months to a workflow you'd actually trust.

To be clear, that is a different number from the six-to-eighteen-month timelines you'll see quoted in broader "AI transformation" coverage. Those describe programs that touch many systems across an entire department. A single, well-scoped reporting workflow is a much smaller project, and mixing the two up is part of what makes this sound like a bigger undertaking than it needs to be for a mid-size company doing it narrowly. The two most common reasons real projects run long: the actual person who does the work wasn't part of the early conversations, and the scope quietly grew midway through. That question, "can it also handle this other thing?", is usually the most expensive sentence in a project like this.

Decide what the tool can touch before you decide what it does

This is a decision to make on day one, not something to add once the project is already built. The pattern most enterprises now use splits the work into two tiers: a look-and-report tier that can run with light supervision, and an act-and-change tier that stays behind human approval until it's earned trust.

For the workflows described above, that split is concrete: the tool can look up information in the systems that track orders, stock, and supplier details, and it can draft the report. It cannot place an order, approve an invoice, or send a payment. Work that only reads and summarizes information is the kind that can reasonably run with minimal oversight; anything that changes a record in a business system, or communicates with someone outside the company, should stay supervised until it's proven itself.

Setting this boundary explicitly, before you scope the build, does two things: it keeps the project inside the level of risk most mid-size companies actually want to take on for a first project like this, and it gives you a clean, later decision point. Expanding what the tool is allowed to do becomes a deliberate choice you make with evidence behind it, rather than something it drifts into.

The unglamorous part that determines success: getting the underlying data right

Across a wide range of surveys and after-the-fact reviews from this year (the exact figures vary depending on how each one was measured, but the direction is consistent), the way information is connected and organized is cited more often than the AI itself as the reason these projects stall before going live. That tracks with what the MIT research found: the projects that survived were the ones built directly into a specific, real process against real information, not general-purpose tools thrown at a broad problem.

For a mid-size company, this is where the real project risk actually lives, and this is the point where you want real specifics before committing to a timeline. What does the tool actually have access to, and across how many separate systems? Are supplier names, product codes, and dates recorded the same way in every system, or do they need to be reconciled by hand first? Is there a safe test version of these systems to build against, or will the project need access to real, live company data from day one? None of this is glamorous, and none of it shows up in a demo, but it's the actual determinant of whether the build estimate above holds or slips.

This is also, practically, where getting outside help pays for itself: someone who has connected AI tools to messy, real-world business data before will find the inconsistencies in week two instead of week eight.

Expand only on evidence, not on schedule

Once the supervised trial has run its thirty days against your existing manual process, the honest question is whether it earned more independence, not whether the calendar says it's time to grant it. The usual progression is from a person approving everything before it's used, to a person reviewing the results afterward and stepping in only when something looks off.

For some reports, the ones that feed a decision with real financial weight, it's entirely reasonable to keep a person reviewing everything, indefinitely. That's not a stalled project; that's a correctly calibrated one. The mistake worth avoiding in the other direction is expanding what the tool is allowed to do just because the trial technically worked, rather than because it was measured against the baseline you set out earlier and clearly beat it. Narrow and proven beats broad and assumed, every time this has been studied.

What this actually gets a mid-size company

Done this way, the realistic outcome is specific: hours of recurring, repetitive work returned to the person each week, more consistent coverage of a report that used to slip when someone was out sick or simply swamped, and a faster turnaround on something that's naturally time-boxed: a Monday-morning report that's actually ready Monday morning. What it isn't, at least not yet and not in this scope, is a replacement for the judgment that happens after the report lands, or a system you'd hand ordering or payment authority to on day one.

That's a deliberately smaller claim than most AI pitches make, and it's exactly why it's achievable in a two-to-four-month project rather than a multi-year transformation program.