Start with the person waiting
Imagine a programme coordinator at 4.40 on a Friday. A report is nearly finished. One figure does not match the source file. A colleague is waiting for approval. Someone the programme exists to support is waiting for an answer.
This is an illustrative scene, not a client story. But it puts the AI question where I think it belongs: close to the work, and closer to the people it serves.
If a tool helps that coordinator finish sooner, that matters. The next question matters more: who benefits from the time that has been released? Another form can always fill the gap. So can a conversation that should have happened days ago.
The time we save needs a destination.
The promise is real. So are its boundaries.
There is evidence worth being optimistic about. In Generative AI at Work, researchers studied the staggered introduction of an AI assistant among 5,172 customer-support agents. Productivity, measured as issues resolved per hour, rose by 15% on average. Less experienced workers benefited particularly.
That is an encouraging finding about support work in that setting. It is not evidence that every charity can cut costs by 15%, or serve 15% more people. Throughput, expenditure and impact are different things.
For a non-profit leader, I would turn the result into a testable question: could carefully designed assistance help colleagues find reliable information and complete a particular task? Then I would measure the whole task, including checking and corrections. That is a more useful starting point than importing somebody else's percentage into our business case.
The time we save needs a destination.
A real gain, in a particular setting
Customer-support issues resolved per hour. Comparison indexed to 100.
Afficher les données du graphique
| Category | Value | Group |
|---|---|---|
| Comparison baseline | 100 | |
| AI-assisted productivity | 115 |
Read past the headline
The uncomfortable findings deserve the same attention. In a 2025 METR experiment, 16 experienced developers working in familiar repositories took 19% longer with the AI tools studied, despite believing they had been faster.
But stopping there would also mislead. METR's February 2026 update reported signs of improvement while warning that selection effects and time measurement made its newer estimates unreliable. We should neither dismiss useful tools nor turn an older result into a timeless verdict.
One distinction deserves particular care: feeling productive is not the same as producing better work. Ask colleagues how the experience feels, certainly. Also examine completion time, corrections, exceptions and the outcome the task was meant to support. The gap between those views may be the most useful thing your pilot discovers.
Fluency can hide the difficult bit
A study involving 758 consultants, published in Organization Science in March 2026, sharpens the point. AI helped on tasks within its capabilities. On one harder task outside them, the combined AI groups were around 19 percentage points less likely to reach the correct answer.
The less familiar finding is in Table 9: AI-assisted recommendations could be more persuasive and coherent even when they were wrong. A beautifully framed answer can therefore make the reviewer's job harder, not easier. This was a particular experiment with older GPT-4, not a failure rate for today's systems.
My practical reading is simple. Put the source beside the claim. Ask what would disprove the recommendation. Keep uncertainty visible. Do not make the person approving the work reconstruct the evidence from a confident paragraph. Good design should make careful judgement easier.
A convincing answer is not always a correct one
Correct answers to one complex consulting task. Approximate group percentages.
Afficher les données du graphique
| Category | Value | Group |
|---|---|---|
| No AI | 84.5 | |
| GPT-4 only | 70.6 | |
| GPT-4 + overview | 60 |
Give the hours a job
Here is a deliberately hypothetical example. An update takes 60 minutes: 25 reading source material, 20 drafting and 15 checking. AI reduces reading to 10 and drafting to five, but checking rises to 25. The complete task now takes 40 minutes, not 15.
Across 120 updates, that would release 40 hours a month. It would not automatically remove 40 hours from payroll. Implementation, support and software have costs too. Nor would it prove a better service.
What could we choose to do with that capacity? Resolve overdue questions. Improve accessibility. Spend longer understanding why somebody could not use the service. Agree the destination before the pilot begins, alongside the people doing the work. Then check whether those hours actually reached it. Otherwise, a genuine efficiency gain can disappear into an unchanged system.
A bigger ambition, honestly measured
I want AI to make organisations doing extraordinary good more capable. That ambition needs more than impressive demonstrations. It needs the patience to ask what improved, for whom, and at what cost.
We do not have to choose between being brave and being careful. A bounded experiment can be both: small enough to understand, meaningful enough to matter, and open to stopping when the evidence does not support it.
Start with one piece of work and one person accountable for learning from it. Protect the checks that keep people safe. Decide where any released capacity should go.
The goal is not simply to get through more work. It is to make more room for the work only your organisation is here to do.
We do not have to choose between being brave and being careful.
Count the checking, too
Illustrative minutes per update. Lower is faster. No observed client data.
Afficher les données du graphique
| Category | Value | Group |
|---|---|---|
| Before: complete task | 60 | |
| After: complete task | 40 |
Take one useful step.
- With the workflow owner, record a baseline for completion time, corrections and unresolved questions before changing the process.
- Run a bounded comparison on suitable work; record tool versions and count verification, rework and exceptions, not just drafting time.
- Choose one service improvement to receive any released capacity, name its owner and review both benefit and harm before expanding.
Original AIHI commentary. Source authorship remains with the credited publication.

001