Back to the Archive

LibraryFoundational AI: Do's and Don'ts12 min read

Which part of your monthly report is actually an AI job?

A recurring management report does three separate jobs: it pulls numbers, it computes them, and it explains them. Only one of those is an AI job, and Microsoft retiring a spreadsheet function on the fourteenth is a useful reminder of which one.

A manager reviews a monthly report while an AI assistant explains results apart from data collection and calculations.

On the fourteenth of this month Microsoft turned off the COPILOT function in Excel. If you never met it: it let you write =COPILOT("Summarize this feedback", A2:A20) in a cell and get a sentence back, the same way you write =SUM(). It was a preview feature, it was well liked by the people who found it, and Microsoft's documentation for the function now opens with a notice that it is no longer available. Results already calculated stay in the workbook as cached values. The next time one of those cells recalculates, it returns #NAME?.

That is a quiet way to break. Nothing fails on the fourteenth. The workbook fails whenever somebody next opens it, changes a filter, and sets the sheet recalculating, which for a monthly report means the first working day of October, roughly twenty minutes before it was due.

Most of you did not have that function in a report, so this is not a repair notice. It is worth raising anyway, because of what it exposes. A recurring management report does three separate jobs that feel like one job, and the retired function was doing one of them in the place where a different one belongs. Sorting the three out is the actual work, and it is worth doing whether or not any AI ends up in the finished thing.

Three jobs wearing one document

Here is an illustrative case, not a client: a services business of about thirty people runs a weekly operations report. It carries new opportunities and their value, utilization by team, invoices more than thirty days overdue, support tickets open longer than five days, and a short written summary at the top for the two people who read only the summary.

Producing that involves three jobs:

The pull. Getting figures out of the systems that hold them. The CRM, the time tracker, the accounting package, the helpdesk.

The compute. The arithmetic performed on what came out. Totals, week over week comparisons, percentages, an average that excludes two people on leave.

The explain. The paragraph at the top. What moved, against what, what it probably means, what somebody should do about it.

In most reports these three are smeared together. Somebody reads a figure off a dashboard and types it into a cell, which is a pull performed by a human being. Somebody works out a comparison in their head while writing the summary, which is a compute performed inside the explain. Both of those feel like efficiency and both are why the report takes a person most of a morning and cannot be handed to anybody else.

The short answer to the question in the title: AI belongs in the explain job. It is conditionally useful around the pull. It should not be doing the compute, and the vendors selling it to you now say so in their own documentation.

The rest of this is a worksheet. Take the last copy of a report you actually sent and work through it. It takes about forty minutes the first time.

Step one: write the decision beside every number

Make four columns and one row per number in the report.

| The number | Who reads it | What they decide with it | What breaks if it disappears | | ---------- | ------------ | ------------------------ | ---------------------------- |

Fill the third column in as a sentence containing a verb. "Utilization by team" is not a decision. "Whether to move two people off the delivery backlog next week" is a decision.

The rule for this step: a number that cannot name both a decision and a person leaves the report. Not into an appendix, not into a hidden tab. It leaves. A report that has run for four years is mostly sediment, and every one of those rows is costing somebody a pull, a compute and a line of explanation every single cycle.

Expect to cut somewhere between a quarter and a half of the rows. If you cut nothing, you have written aspirations in the third column rather than decisions, and it is worth a second pass with somebody who actually receives the report.

Step two: label each surviving row P, C or E

For every row still standing, mark which of the three jobs produces it. Use these tests, because the intuitive answer is often wrong.

It is a pull if you can name the system, the saved view or filter, and the date range, and there is an export button at the end of that sentence. If the answer is that somebody opens a screen and reads a figure off it, this is not a pull yet. It is a person, and it stays a person until you have a filter saved and an export.

It is a compute if you can write it as a formula and show the formula to somebody. If the arithmetic currently happens in a head, write the formula down now, in this step. This is where you discover that two people have been computing "utilization" differently for a year.

It is an explain if it is sentences. Cause, context, comparison, consequence.

Most reports come out of this heavily weighted towards pull. That is the finding, and it is the opposite of what people assume before they do it. The explain job is usually one paragraph and ten minutes. The pull is usually four systems and two hours.

Step three: keep AI out of the compute layer

This is the step the retired Excel function makes easy to argue, so use it while it is fresh.

Microsoft's own FAQ for Copilot in Excel lists, under limitations, that Copilot "can sometimes make mistakes, misinterpret information, or produce inaccurate results" and advises avoiding it "for decisions in sensitive areas such as finance, legal, or medical topics." On trusting its output: review, edit and verify anything it creates before you rely on it. That is the vendor writing about its own paid feature, which makes it a reasonable line to quote to anybody on your team who wants to skip this step.

Google is equally specific about its equivalent. The AI function in Google Sheets, written as =AI() or =Gemini(), carries limitations that matter enormously for something you run every month:

  • Responses are text only.
  • The function cannot see the rest of your spreadsheet or your Drive. It only sees the range you hand it.
  • You cannot undo or redo it. You can only regenerate.
  • You cannot nest it inside another function, so it cannot sit inside an IF that checks its answer.
  • It does not run on its own. Somebody selects the cells and clicks Generate and Insert.
  • Only the first 350 selected cells generate at a time, and there are longer term limits that can stop you generating for 24 hours.

Read that list as a description of a recurring report and it disqualifies itself: a thing that cannot be scheduled, cannot be checked by a formula, cannot be undone, and can refuse to run for a day is not where your invoice ageing total should live.

There is a subtler reproducibility problem too. In Excel's Copilot pane you can switch between Claude and GPT models, and per the same FAQ that choice applies only to the active session; close Excel and it reverts to the default. So a report generated on Tuesday and regenerated on Thursday was not necessarily produced by the same model. For a sentence, that is tolerable. For a number that somebody compares to last month's number, it is not.

So the rule for this step: every figure in the report comes from an export or a formula. No exceptions, including the convenient ones. A figure you cannot produce that way stays a manual number with a person's name written next to it until you can.

Step four: fix the pull before you touch the AI part

Work through each P row in order:

  1. Name the system and the exact saved view or filter. Save the view inside the tool so it is not retyped.
  2. Fix the date range as a rule rather than a pair of dates. "The seven days ending last Sunday", not "15th to 21st".
  3. Export to the same place with the same filename pattern, with the period in the name.
  4. Write down who has access to do this, and one other person who also does.

Then run the two run test: perform the whole pull twice in the same hour for the same period. If any number differs between the two runs, the filter is not pinned yet, and you have just found the reason two people disagreed about a figure in March.

This step is unglamorous and it is where the time goes. It is also where most of the benefit is, and a lot of teams should stop here for a cycle or two. A report whose numbers arrive the same way every week, from a filter nobody retypes, is more valuable than the same report with a generated paragraph on top of numbers that still move depending on who fetched them.

Step five: write the standing brief once, as a file

Now the explain job. The mistake here is treating it as a prompt somebody types fresh each cycle, which is why the summary reads differently every month and why nobody else can produce it.

Write a standing brief, once, containing:

  1. Who reads this and what they decide.
  2. The baseline every comparison is made against: previous week, same week last year, or plan. Pick one and name it.
  3. The threshold for mentioning a movement at all. Below it, say nothing.
  4. What it may not do. No recommendations it cannot support from the numbers in front of it. No causes it cannot see in the data. No arithmetic of its own: every figure it uses is quoted from the sheet.
  5. The shape: length, order, whether it names people.
  6. What to write when nothing material moved, because "nothing material moved" is a valid and useful summary and the default behavior is to invent significance rather than report a quiet week.

Microsoft now supports exactly this as a file. Custom skills for Copilot in Excel are folders in your OneDrive, each containing a SKILL.md with a name, a description telling Copilot when to use it, and the instructions themselves. Note the constraint in that documentation: skills are supported in English only, and your Office display language has to be set to English to reach them. On the Google side there is no equivalent file, so the brief lives in the spreadsheet itself, in a hidden tab, and gets pasted.

An illustrative brief for the weekly operations report above, written as it would actually be stored:

You write the summary at the top of the weekly operations report for the managing director and the operations lead. They decide staffing moves for the coming week and which overdue invoices get chased personally. Compare every figure to the previous week, which is in the column marked Prior. Mention a movement only if it is more than 10 percent or more than 5,000 in value. Quote figures exactly as they appear in the Summary tab and never calculate your own. Do not suggest causes that are not visible in this data. Four sentences maximum, in this order: pipeline, utilization, cash, support. If nothing crosses the threshold, write one sentence saying the week was flat and name the largest movement anyway.

That brief is transferable, reviewable, and arguable at a meeting. A prompt in somebody's chat history is none of those things.

Step six: the three checks that make it safe to send

Each of these takes under a minute, and together they are what turns this from a demo into a workflow.

The spot check. Pick three figures from the summary and trace them back to the export. Not the same three every week.

The sentence test. Read the generated paragraph and strike any sentence you would not have written yourself. If you strike the same kind of sentence twice in a row, that is not a bad generation, it is a missing line in the brief. Go and add it.

The second run test. Regenerate the summary without changing any data. The wording will differ. If a number or a direction differs, something is being computed in the explain layer, and it needs to go back to step three.

And one thing that is not a check but matters more than all three: a person's name goes at the bottom of the report. The workflow drafts. A person sends. That is not ceremony, it is the thing that makes the checks happen.

What this looks like when it is running, and when to revisit it

A report that has been through this has a pull that is mechanical, a compute that is formulas, an explain that takes a couple of minutes to generate and about ten to review, and a brief that somebody else could pick up on a Monday when the usual person is away.

Set a review point: once a quarter, redo step one only. Any number that nobody has asked a question about in two cycles comes out. Reports gain weight the way inboxes do, and the quarterly cut is cheaper than the annual rebuild.

Where this stops working

If a number lives in a system with no export and no API, a supplier portal or a bank interface, the pull is a person and it stays a person. Write that into the brief so the summary never implies a freshness the figure does not have.

If the report is a regulated financial statement rather than a management report, keep generated text out of it entirely. Microsoft's own advice about finance is the relevant standard, and an auditor's view of "the model wrote it" is not a conversation worth having.

If the real problem is that people keep asking questions the report does not answer, you do not have a reporting workflow problem, you have a dashboard problem, and adding a generated paragraph will not touch it.

And if the report is five numbers going to one person who already knows what they mean, this whole exercise costs more than the report does. Write the five numbers in an email and get on with your week.

Open the last report you sent, and mark each number P, C or E. Most people find the part they wanted to automate, the writing, is the smallest job in the document, and the part nobody wanted to look at, the pull, is where the morning actually goes.

Sources

Every claim above traces back to one of these. Go read them yourself.

  1. 01
  2. 02
  3. 03
  4. 04