What the usage data shows AI is good and bad at, at work
Short answer
Across the 332 work activities behind this site, the pattern is consistent: the highest scores go to explaining and advising, and the lowest to collecting things in person and directing other people. Every figure here is classified by a language model rather than observed, and that classifier agrees with human annotators at a level the authors themselves call generally low, so read the ordering rather than any single value. Creative work splits instead of moving as one block: making images and objects sits at the bottom of the measured range, below every other activity that registers at all, while creative writing sits near the top. Everything here describes 2024 usage of one assistant, and the fastest-moving evidence in this field is about how quickly such figures expire.
The pattern in the usage data
Every work activity in the source dataset carries a value combining how often AI was used for it, whether the assistant finished, and how much of the activity it covered.1 All the figures below come from the user-goal series, meaning what people asked the assistant to help with. The dataset also publishes a separate series for what the assistant itself did; the site's headline score averages both.2
Grouping those 332 activities by their leading verb, which is this site's own editorial grouping and not part of the source data, gives an ordering. Every family below has at least five activities in it, which is the cut-off used here so that no average rests on one or two rows. The nineteen families cover 163 of the 332 activities; the rest sit in families too small to average.
| Kind of activity | Activities | Mean value | Scoring zero |
|---|---|---|---|
| Providing information | 5 | 0.56 | 1 |
| Advising | 10 | 0.56 | 2 |
| Researching | 8 | 0.48 | 1 |
| Investigating | 5 | 0.38 | 1 |
| Assisting | 5 | 0.37 | 2 |
| Analysing | 9 | 0.30 | 3 |
| Maintaining | 10 | 0.24 | 6 |
| Assessing | 5 | 0.21 | 3 |
| Preparing | 13 | 0.20 | 9 |
| Evaluating | 11 | 0.20 | 7 |
| Developing | 22 | 0.19 | 14 |
| Designing | 6 | 0.14 | 4 |
| Performing | 7 | 0.14 | 5 |
| Monitoring | 12 | 0.12 | 9 |
| Operating equipment | 14 | 0.10 | 11 |
| Inspecting | 5 | 0.10 | 4 |
| Coordinating | 5 | 0.09 | 4 |
| Collecting | 5 | 0.00 | 5 |
| Directing | 6 | 0.00 | 6 |
One family sits above everything in that table and is too small to appear in it: the four activities beginning "explain" average 0.80.2
The two families at the bottom need care, because the obvious reading of them is wrong. All five activities whose title begins "collect" and all six beginning "direct" score zero.2 That does not mean they never happen with AI. Each one does appear in the conversation data, and where it appears the assistant completed the task between 60% and 100% of the time.2 They score zero because each falls under the study's frequency threshold, not because the assistant failed at them. Reading your score sets out how that cut-off works.
Read the verb rather than the subject matter, though, because the "collect" family is not what its name suggests. Two of its five activities are about gathering information from people; the other three are collecting physical samples, products or fares, which a chat assistant cannot do at all. Information gathering as such runs the other way: gathering information from physical or electronic sources scores 0.65 and obtaining information about goods or services 0.78. Of the 21 activity titles containing the word "information", 15 score above zero.2
At the top, the highest single values in the dataset are explaining medical information to patients (0.83), explaining technical details of products (0.81) and providing information to guests, clients or customers (0.80).2 The strongest measured performance is in turning knowledge into an explanation for a particular person.
A second dataset, built differently, lands in the same area
Anthropic publishes its own count of what people produce with Claude, built from entirely different data. It measures something different from the Microsoft study: the share of Claude chat and Cowork conversations between 10 April and 10 June 2026, work and personal alike, that produced each kind of output. Claude Code sessions and first-party API traffic sit outside this particular measurement, which removes most coding output, though the report it comes from covers them elsewhere. The most common were explanations at 17%, documents and reports at 15%, and guidance at 11%.3
Two projects, different companies, different data, and the same centre of gravity: explanation and documents. The two are not directly comparable, and both classify chat logs with language models rather than observing work, so treat the overlap as a consistent picture rather than confirmation.
Creative work does not move as one block
The public story says artists and designers were hit first. The data splits. The three lowest non-zero values in the whole dataset are creating visual designs (0.20), creating artistic designs or performances (0.23) and creating decorative objects (0.25).2 Making images and objects sits at the bottom of the measured range, below every other activity that registers at all.
Creative writing runs the other way. Writing material for artistic or commercial purposes scores 0.71, which is eighteenth of all 332 activities, and developing news, entertainment or artistic content scores 0.36.2 A copywriter and an illustrator are in different places on this map, and the usual "AI came for the creatives" framing collapses the two.
One clarification, because the dataset invites a wrong reading here. Its separate "design" family is engineering design, covering databases, structures, industrial systems and electrical equipment rather than artwork, and four of its six activities fall below the frequency threshold. Its 0.14 average says nothing about creative work either way.2
Be careful about what the low visual figures do and do not mean. They measure how often people used a text-and-image assistant for those activities in 2024 and how far it got, in a dataset built around a chat interface. They are not a verdict on dedicated image-generation tools, which are a different product category, and they are certainly not a promise that creative work is safe. What they do show is that the confident claim "AI came for the creatives first" is not what this evidence supports.
How fast these numbers expire
Between February and June 2025 the research group METR ran a randomised controlled trial with 16 experienced open-source developers on 246 real tasks in their own repositories, published that July. Allowing AI tools made them 19% slower, with a confidence interval from 2% to 39%.4
The durable finding was not the slowdown, it was the misjudgement. Those developers predicted AI would make them 24% faster; even after being measurably slowed, they still believed it had sped them up by 20%. Outside experts had forecast speedups too: 39% from economists and 38% from machine-learning researchers.4 Everyone, including the people doing the work, got the direction wrong.
And then the finding itself expired. METR has since attached a warning to that study stating the results "are out of date" and that it believes they "no longer reflect the current impact of AI models on open-source developer productivity"; its continuation, as reported in February 2026, showed raw evidence of a speedup instead, on estimates whose intervals cross zero, unlike the original 19% slowdown.4 A headline result reversed within a year. Any page telling you what AI is good at, including this one, is a photograph rather than a rule.
Where these numbers are weakest
The measurement of how much of an activity AI handled is the noisiest part of the source study, and the classification pipeline agrees with human annotators at a level its authors call generally low, a limit quantified in reading your score.1 The data covers one consumer assistant, in the United States, over nine months of 2024, so anything that moved since is invisible to it, as is all the work done inside specialist tools. And an activity can score zero for reasons set out in what "not observed" actually means.
What this means for you
If your week is mostly explaining, advising, drafting and summarising, you are in the part of the map with the heaviest measured overlap. The useful move is to get good at directing and checking these tools, which is an inference from that pattern rather than a measured finding. If your week is mostly collecting things in person, watching over a process, or deciding what other people should do, the measured overlap is near zero, though that is a statement about one consumer assistant in 2024 and not about other kinds of automation.
Your own occupation page shows which of these activities actually make up your job, with the measured value beside each one. Start from the calculator.
- Kiran Tomlinson, Sonia Jaffe, Will Wang, Scott Counts and Siddharth Suri, Working with AI: Measuring the Applicability of Generative AI to Occupations, Microsoft Research, 2025. arXiv:2507.07935, version 6 of 22 December 2025. Source of the coverage, completion and scope measures, the nine-month United States collection window, and the Cohen's kappa range of 0.34 to 0.53 between the classification pipeline and human annotators.
- Microsoft Research, Working with AI published result files, CC BY 4.0. github.com/microsoft/working-with-ai. The verb-family table, the highest and lowest individual activity values, the zero counts, the information-title counts and the completion rates for the collect and direct families were computed directly from the published
iwa_metrics.csvon 22 August 2026, using the user-goal series. Grouping activities by leading verb is JobRiskAI's own editorial choice and is not part of the source dataset. The table lists every verb family with at least five activities, nineteen of them, covering 163 of the 332 activities; smaller families are omitted because an average over one or two rows is not worth reading, and the highest-scoring family of all, the four activities beginning "explain" at 0.80, is named in the text instead. - Anthropic Economic Index report Cadences, published 26 June 2026, covering conversations sampled between 10 April and 10 June 2026. anthropic.com/economic-index. The report states that 93% of those conversations produced an artefact, and that the most common were explanations at 17% of conversations, documents and reports at 15% and guidance at 11%. These shares cover Claude chat and Cowork conversations only, work and personal alike. Claude Code sessions and first-party API traffic fall outside this particular measurement, which removes most coding output, although the report discusses both elsewhere. The measure is not occupation-weighted, so it is not directly comparable to the Microsoft figures.
- Joel Becker, Nate Rush, Beth Barnes and David Rein, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR, published 10 July 2025 with fieldwork between February and June 2025, together with METR's later update, We are Changing our Developer Productivity Experiment Design, 24 February 2026. metr.org. Quoted from the 2025 study and its paper (arXiv:2507.09089): the 19% slowdown with a 95% confidence interval of 2% to 39%, which excludes zero; the 24% predicted and 20% perceived speedups; and the 39% and 38% expert forecasts. Quoted from the 2026 update: METR's own notice that the results are out of date and no longer reflect current model impact, and its continuation showing raw evidence of a speedup on estimates whose intervals cross zero.
Every figure above is quoted with the caveat its own source states. Last verified 2026-08-22. Scores and bands used on this site are documented on the methodology page.