+55,8%
GitHub / Microsoft Research · 2023 · Peng, Kalliamvakou, Cihon, Demirer

The group with access to GitHub Copilot completed the task 55.8% faster than the control group.
Controlled experiment (RCT) with recruited software developers asked to implement an HTTP server in JavaScript as quickly as possible. Completion time was measured for the Copilot group versus the control group.
Caveat: It is a single well-scoped lab task, not real work in a production codebase, and the authors work at GitHub/Microsoft, the tool's vendor.
How we apply it: The big gain shows up on well-scoped tasks. That is why we teach teams to slice work into small, bounded specs before handing it to agents: the lab speedup only transfers to production when the task looks like the lab's.
Read the study →+25,1%
Harvard Business School / BCG · 2023 · Dell'Acqua et al. (Organization Science)

Across 18 tasks within the AI capability frontier, consultants using GPT-4 completed 12.2% more tasks, 25.1% faster, with 32% higher quality on average.
Preregistered field experiment with 758 BCG consultants randomly assigned to no AI, GPT-4, or GPT-4 plus a prompt-engineering overview, on realistic consulting tasks (creative and analytical) with a prior performance baseline.
Caveat: On a complex task chosen to sit outside the AI frontier, subjects using AI were 19% less likely to produce correct solutions than those without it.
How we apply it: We train exactly that frontier: the judgment of what to delegate to AI and what not to. In our programs, people practice recognizing when a task sits within AI's capabilities and when it will confidently walk you to the wrong place.
Read the study →+14%
NBER · 2023 · Brynjolfsson, Li, Raymond

Access to a generative AI assistant increased issues resolved per hour in customer support by 14% on average, with +34% for novice and low-skilled agents.
Study of the staggered rollout of a generative-AI conversational assistant using data from 5,179 customer support agents. Productivity was measured as issues resolved per hour.
Caveat: The average effect hides strong heterogeneity: almost all the benefit accrues to novices, with minimal impact on experienced, highly skilled agents.
How we apply it: AI performs best when it carries the judgment of your best people. In transformations we codify the team's senior know-how — specs, context, guardrails — so the tool lifts the whole team instead of just assisting those who already knew.
Read the study →hasta 2x
McKinsey Digital · 2023 · Karaci Deniz, Gnanasambandam, Harrysson et al.

With generative AI, code documentation was completed in half the time, new code in nearly half the time, and refactoring in nearly two-thirds the time — up to twice as fast on common tasks.
McKinsey lab study with more than 40 of its own developers across the US and Asia, performing code generation, refactoring, and documentation tasks over several weeks, crossing over between a two-AI-tool group and a control group.
Caveat: On high-complexity tasks savings shrank below 10%, and developers with under a year of experience took 7-10% LONGER on some tasks with the tools; it is also internal consulting research, not peer-reviewed.
How we apply it: The return depends on the task and the starting level — that is why we measure Tech Fluency and AI Fluency before and after training, and prioritize the tasks where the return is real instead of promising 2x on everything.
Read the study →