Start with the person waiting

Imagine a programme coordinator at 4.40 on a Friday. A report is nearly finished. One figure does not match the source file. A colleague is waiting for approval. Someone the programme exists to support is waiting for an answer.

This is an illustrative scene, not a client story. But it puts the AI question where I think it belongs: close to the work, and closer to the people it serves.

If a tool helps that coordinator finish sooner, that matters. The next question matters more: who benefits from the time that has been released? Another form can always fill the gap. So can a conversation that should have happened days ago.

The time we save needs a destination.

The promise is real. So are its boundaries.

There is evidence worth being optimistic about. In Generative AI at Work, researchers studied the staggered introduction of an AI assistant among 5,172 customer-support agents. Productivity, measured as issues resolved per hour, rose by 15% on average. Less experienced workers benefited particularly.

That is an encouraging finding about support work in that setting. It is not evidence that every charity can cut costs by 15%, or serve 15% more people. Throughput, expenditure and impact are different things.

For a non-profit leader, I would turn the result into a testable question: could carefully designed assistance help colleagues find reliable information and complete a particular task? Then I would measure the whole task, including checking and corrections. That is a more useful starting point than importing somebody else's percentage into our business case.

The time we save needs a destination.
AIHI
lollipop

A real gain, in a particular setting

Customer-support issues resolved per hour. Comparison indexed to 100.

A real gain, in a particular settingCustomer-support issues resolved per hour. Comparison indexed to 100. Staggered rollout among 5,172 support agents: average productivity increased by 15%. This is throughput, not a 15% cash saving or a non-profit impact estimate. Index: 100 x 1.15 = 115. 0255075100125indexComparison baseline100AI-assisted productivity115
Staggered rollout among 5,172 support agents: average productivity increased by 15%. This is throughput, not a 15% cash saving or a non-profit impact estimate. Index: 100 x 1.15 = 115.Source: Brynjolfsson, Li and Raymond, Generative AI at Work, revised November 2024.
Afficher les données du graphique
CategoryValueGroup
Comparison baseline100
AI-assisted productivity115

Read past the headline

The uncomfortable findings deserve the same attention. In a 2025 METR experiment, 16 experienced developers working in familiar repositories took 19% longer with the AI tools studied, despite believing they had been faster.

But stopping there would also mislead. METR's February 2026 update reported signs of improvement while warning that selection effects and time measurement made its newer estimates unreliable. We should neither dismiss useful tools nor turn an older result into a timeless verdict.

One distinction deserves particular care: feeling productive is not the same as producing better work. Ask colleagues how the experience feels, certainly. Also examine completion time, corrections, exceptions and the outcome the task was meant to support. The gap between those views may be the most useful thing your pilot discovers.

Fluency can hide the difficult bit

A study involving 758 consultants, published in Organization Science in March 2026, sharpens the point. AI helped on tasks within its capabilities. On one harder task outside them, the combined AI groups were around 19 percentage points less likely to reach the correct answer.

The less familiar finding is in Table 9: AI-assisted recommendations could be more persuasive and coherent even when they were wrong. A beautifully framed answer can therefore make the reviewer's job harder, not easier. This was a particular experiment with older GPT-4, not a failure rate for today's systems.

My practical reading is simple. Put the source beside the claim. Ask what would disprove the recommendation. Keep uncertainty visible. Do not make the person approving the work reconstruct the evidence from a confident paragraph. Good design should make careful judgement easier.

lollipop

A convincing answer is not always a correct one

Correct answers to one complex consulting task. Approximate group percentages.

A convincing answer is not always a correct oneCorrect answers to one complex consulting task. Approximate group percentages. Figure 5 / section 5.2: no AI about 84.5%; GPT-4 only 70.6%; GPT-4 with an overview 60%. The combined AI groups were about 19 percentage points below control. One task and older GPT-4: not a general failure rate for current tools. 020406080100%No AI84.5%GPT-4 only70.6%GPT-4 + overview60%
Figure 5 / section 5.2: no AI about 84.5%; GPT-4 only 70.6%; GPT-4 with an overview 60%. The combined AI groups were about 19 percentage points below control. One task and older GPT-4: not a general failure rate for current tools.Source: Dell’Acqua et al., Organization Science, 11 March 2026, Figure 5 and Table 9.
Afficher les données du graphique
CategoryValueGroup
No AI84.5
GPT-4 only70.6
GPT-4 + overview60

Give the hours a job

Here is a deliberately hypothetical example. An update takes 60 minutes: 25 reading source material, 20 drafting and 15 checking. AI reduces reading to 10 and drafting to five, but checking rises to 25. The complete task now takes 40 minutes, not 15.

Across 120 updates, that would release 40 hours a month. It would not automatically remove 40 hours from payroll. Implementation, support and software have costs too. Nor would it prove a better service.

What could we choose to do with that capacity? Resolve overdue questions. Improve accessibility. Spend longer understanding why somebody could not use the service. Agree the destination before the pilot begins, alongside the people doing the work. Then check whether those hours actually reached it. Otherwise, a genuine efficiency gain can disappear into an unchanged system.

A bigger ambition, honestly measured

I want AI to make organisations doing extraordinary good more capable. That ambition needs more than impressive demonstrations. It needs the patience to ask what improved, for whom, and at what cost.

We do not have to choose between being brave and being careful. A bounded experiment can be both: small enough to understand, meaningful enough to matter, and open to stopping when the evidence does not support it.

Start with one piece of work and one person accountable for learning from it. Protect the checks that keep people safe. Decide where any released capacity should go.

The goal is not simply to get through more work. It is to make more room for the work only your organisation is here to do.

We do not have to choose between being brave and being careful.
AIHI
lollipop

Count the checking, too

Illustrative minutes per update. Lower is faster. No observed client data.

Count the checking, tooIllustrative minutes per update. Lower is faster. No observed client data. Before: reading 25 + drafting 20 + checking 15 = 60 minutes. After: reading 10 + drafting 5 + checking 25 = 40. Net saving: 20 minutes. At 120 updates a month, 40 hours of capacity could be redirected. Implementation costs and cash savings are not estimated. 01224364860minutesBefore: complete task60After: complete task40
Before: reading 25 + drafting 20 + checking 15 = 60 minutes. After: reading 10 + drafting 5 + checking 25 = 40. Net saving: 20 minutes. At 120 updates a month, 40 hours of capacity could be redirected. Implementation costs and cash savings are not estimated.Source: AIHI illustrative arithmetic, 14 September 2026. Assumptions shown above; not a measured result.
Afficher les données du graphique
CategoryValueGroup
Before: complete task60
After: complete task40
VOTRE PROCHAINE CONVERSATION

Take one useful step.

  • With the workflow owner, record a baseline for completion time, corrections and unresolved questions before changing the process.
  • Run a bounded comparison on suitable work; record tool versions and count verification, rework and exceptions, not just drafting time.
  • Choose one service improvement to receive any released capacity, name its owner and review both benefit and harm before expanding.

Original AIHI commentary. Source authorship remains with the credited publication.

PASSEZ À LA PRATIQUE

Impactly

A connected journey from programme information and authorised field updates to a clearer donor view.

Découvrir cette solution
TROUVEZ VOTRE POINT DE DÉPART

Évaluer votre préparation à l’IA

Découvrez ce qui est prêt, où un soutien serait utile et quelle première étape choisir.

Faire l’évaluation gratuite