Contents 8
For a year now, my team and I have been working with AI agents. We hand them whole tasks: build a feature, migrate a service, write tests. It goes fast, sometimes very fast.
To explain what that changes, to developers as much as to people who have never opened a code editor, I use a single image: an apple tree and a basket. It starts from a phrase you hear in every meeting: low-hanging fruit, the fruit hanging low, the one you pick without a ladder. This article unfolds the image, with its opportunities and its dangers, then puts it up against the numbers: my team’s, and a year of my own days.
In one minute
A project is a tree, each task is an apple. Agents shrink the tree. They don’t make the apples any lighter, or the basket any bigger.
- The appleA piece of work worth doing: a feature, some debt, a tool, a piece of content.
- The heightThe effort to reach it. With agents, it collapses.
- The weightThe risk, and the time to taste it: review, test, ship. It doesn't move.
- The basketWhat we can taste and carry. It doesn't grow on its own.
Opportunities
Dangers
The tree shrinks
You used to need a 10-meter ladder. Now a stepladder will do. But some apples still weigh 50 kilos.
Effort used to decide. We picked the apples at the bottom, the ones within arm’s reach, and the top of the tree waited: tests for the old code, the version upgrade, the security audit. “Someday, maybe.”
That’s the meeting-room low-hanging fruit. The phrase is in print as early as 1968, and becomes a management tic in the 1980s (Language Log). Orchards, for their part, changed before we did: dwarf trees 1.5 to 2 meters tall replaced the big apple trees, and growers now plant up to 2,000 per acre, against 20 before. A geneticist interviewed by Priceonomics draws the conclusion: “The phrase is irrelevant” (Priceonomics). With agents, our projects are going through the same thing.
With an agent, a task of several hours starts with one sentence. And it’s improving fast: since 2024, the length of tasks an agent completes on its own has roughly doubled every three months (METR, January 2026). Measured at a 50% success threshold, METR points out: you still have to check.
But height isn’t weight. Changing a button’s color and rewriting an application in Rust are now one sentence away. The first can be checked in thirty seconds. The second takes weeks of review, testing and rollout, and it can break everything. An apple’s weight is the risk, and the time it takes to taste it. Agents don’t change that.
It’s a real opportunity. A field study of an autonomous agent shows that people take on more complex, multi-step projects, and step outside their original line of work: they go from operators to supervisors (HBR, July 2026). The top of the tree finally comes within reach.
Sorting changes
When everything is within reach, “it’s easy” no longer justifies anything. Only one question remains: is it worth it?
Effort used to do the sorting for us. We ranked topics on two axes, effort and value, and dreamed of quick wins: little effort, lots of value. That cell was almost always empty, because the low, sweet apples had been picked long ago. “All the low-hanging fruit has been picked”: the line already appears in New Scientist in 2003, about new drugs (Language Log).
Sorting by effort has a known flaw. Under heavy load, people pick the easy task: 76% of participants in one experiment, against 64% when the load is light. It feels like progress, but performance drops in the long run: the emergency physicians studied who take on more difficult cases improve more (KC, Staats, Kouchaki and Gino, Management Science).
With agents, effort no longer breaks the tie: everything lands in the quick wins cell. What’s left is value, and the basket. And the basket fills up by weight, not by count.
In my team, the share of tests in added code has gone from 12.8% to 21.6% since we started working with agents. That’s an apple from the top that we’re finally starting to pick.
The bottleneck moves
Picking costs almost nothing now. Choosing, tasting, shipping and maintaining still cost the same.
Tasting, in software, is review: someone reads the code before it goes to production. Code arrives as a pull request (PR): a proposed change that a colleague reviews and then merges. Then it has to ship, and be maintained for years: monitoring, updates, security.
tasted shipped untasted
Before basket load 100% · 12 picked · 9 tasted
With agents basket load 300% · 38 picked · 10 tasted · 13 shipped untasted
This is the Jevons paradox, observed in 1865 with coal: when a resource gets more efficient, we use more of it, not less. The easier picking gets, the more we pick. For a CTO, the real risk is producing “more software than our organizations can safely operate, secure, and maintain” (Mike Grouchy, quoted by Amazing CTO).
I see it in the team I lead. Since everyone started working with agents, each person merges nearly half again as many PRs, and PRs four times bigger. Time to merge hasn’t budged.
Merged PRs per person per month+46%
Median PR size, in lines×4
Median time to mergestable
×2.2 more code to review every month: 61,000 lines before, 132,000 with agents.
11 → 26% of PRs exceed 500 lines: one in four, against one in nine before.
The basket is the same, and it gets twice as much code to review. Either we review twice as fast, or we taste less. That’s the question to ask as a team.
How far can the basket grow?
The basket can grow, but not by adding agents: by tooling up the tasting.
Adding agents fills the basket faster. And supervising several agents is tiring: in a BCG survey of 1,488 employees, 14% describe mental fatigue from supervising AI tools. Perceived productivity rises up to three tools used in parallel, and drops from four onward (HBR, March 2026; Fortune).
The other lever is to have machines taste some of the apples: automated tests, continuous integration that runs them on every PR, automated review. It’s the only one that grows the basket without wearing people out. Play with both sliders.
- Picked this week
- 12
- Verified by tooling
- 2
- To taste by hand
- 10
- Basket size
- 10
Stretched load 100%
The basket is full. At the slightest surprise, we taste faster, so worse.
One picker, five trees
The basket is also attention. Mine has broken into crumbs: on the days when I have the most time, I switch projects twenty times instead of once.
Context switching is nothing new. Meetings, Slack and “got two minutes?” have always chopped up our days. With agents, I’m the one doing the chopping: I start a session in one repository, another in a second one, and I go back and forth while they work.
For a year now, a collector has filed all my Claude Code and Codex sessions in a journal. I asked it two questions: how many projects I touch per day, and how long I stay on one before moving to the next. In lean, you draw this kind of path on the floor to see wasted motion: it’s a spaghetti diagram. Here’s mine, over two typical days with comparable meeting loads.
In real life, my projects don’t look alike: each has its own code, its own language, its own people. Switching projects isn’t moving from one apple tree to another, it’s putting down the pruning shears to pick up the hammer. So here, each project is a different workstation, with its own tool.
Three two-month windows, a year apart between the first and the last. For each window, the median of my working days, and in parentheses the range where half the days fall.
| Oct–Nov 2025 | Mar–Apr 2026 | Aug–Sep 2026 | |
|---|---|---|---|
| Projects per day | 1 (1–2) | 2 (1–2) | 4.5 (4–5) |
| Project switches per day | 0 (0–2) | 1 (0–3) | 14.5 (5–25) |
| Average block on one project | 110 min | 117 min | 27 min |
Until March, nothing moves: one or two projects a day, blocks of nearly two hours. I was already working with agents, but on one workstream at a time. In April, blocks get cut in half. The juggling comes next: ten switches a day in July, around fifteen in August and September.
A day packed with meetings leaves less room for agents: the comparison could have been biased. I redid it at equal meeting load, on the calmest days of each window. The gap holds, and it widens: one switch a day and 81-minute blocks in 2025, 21 switches and 13-minute blocks in 2026. It’s when I have time that I scatter the most, because that’s when I run agents in parallel.
These numbers don’t measure the quality of the work, only its shape. But a day made of 13-minute blocks no longer looks like the one I used to know.
The orchard rules
Choose by value, finally pick the top, weigh before you pick, and only carry what you can taste.
- Choose by value. Everything is within reach, so “it’s easy” no longer justifies anything. Ask instead: how much does it bring in, for whom, and who’s going to maintain it.
- Finally pick the top. Tests for the old code, version upgrades, documentation, accessibility, the security audit: what we kept putting off for lack of a ladder moves to the top of the list. Professional pickers already work this way: “You always need to start by picking apples from the top, never the lowest-hanging fruits first” (an apple picker from Washington State, to Priceonomics).
- Weigh before you pick. Before starting an agent, estimate the time it will take to check its work. A 50-kilo apple gets cut into apples you can taste one at a time.
- Tool up the tasting. Tests, continuous integration, automated review: that’s where the time saved goes, because that’s what really grows the basket.
- Four slots, not four sessions. One topic per kind: business, technical foundations, team, exploration. Enforcing the categories keeps a hard topic open, instead of giving in to the ones that feel good. But four open topics don’t mean four agents in parallel:
- alternate by half-day, leaving a topic at a stopping point: switching tasks costs more when you leave the previous one unfinished, that’s “attention residue” (Leroy, 2009);
- no more than three active agent sessions at a time;
- no fifth topic: to open one, finish or freeze another.
- Batch reviews, instead of reacting to each agent as soon as it’s done.
We used to ask: what can we reach? Now: what deserves picking, and how much can we carry? See you in six months to find out whether my curves come back up.
The data
The spaghetti is replayed from real days, and the numbers come from the same journal. If you want to check, it’s all here.
See the real data
The three typical days
For each window, among the window's days with the fewest meetings and at least three hours of activity, the one closest to those days' medians. One row per project, one bar per block, one tick per typed message.
Month by month
Medians of each month's working days. Darker: the months of the three windows being compared.
Block length, day by day
Share of each window's days by that day's average block length. The darker, the longer the blocks.
At equal meeting load
A day packed with meetings leaves less room for agents. On each window's lightest meeting days, the gap holds, and it widens.
| M1 · Oct–Nov 2025 | M6 · Mar–Apr 2026 | M12 · Aug–Sep 2026 | |
|---|---|---|---|
| Days kept | 11 of 33 | 14 of 31 | 10 of 28 |
| Projects per day | 2 | 2 | 5 |
| Switches per day | 1 | 1 | 21.5 |
| Average block | 81 min | 86 min | 13 min |
How it's measured
- Source: messages typed into Claude Code and Codex, on two Macs, collected by my agent journal. No meetings, no Slack, no email: only my exchanges with agents.
- Scope: working days with at least 6 messages, 9am to 7pm, Paris time; excluding time off from August 20 to September 4, 2026.
- Project: a git repository or a working folder. Names are replaced with letters, in order of appearance during the day.
- Switch: two consecutive messages sent to two different projects. Average block: the time between the day's first and last message, divided by the number of blocks. It exists even on days without any switch.
- Windows: two months each (M1 Oct–Nov 2025, 33 days; M6 Mar–Apr 2026, 31 days; M12 Aug–Sep 2026, 28 days). Medians, not means: a few extreme days would pull the mean.
- Limits: a descriptive study of a single person. Meetings are context switches too, and this log doesn't count them. In October 2025, almost every message comes from just one of the two Macs.
Sources
- A. Ranganathan, X. M. Ye, “AI Doesn’t Reduce Work, It Intensifies It”, Harvard Business Review, February 2026.
- J. Bedard, M. Kropp et al. (BCG), “When Using AI Leads to ‘Brain Fry’”, Harvard Business Review, March 2026; Fortune, March 10, 2026.
- METR, “Time Horizon 1.1”, January 29, 2026.
- “Research: How AI Agents Broaden the Scope of Knowledge Work”, Harvard Business Review, July 2026.
- D. KC, B. Staats, M. Kouchaki, F. Gino, “Task Selection and Workload: A Focus on Completing Easy Tasks Hurts Performance”, Management Science, 2020 (summary by Kellogg).
- S. Leroy, “Why is it so hard to do my work? The challenge of attention residue when switching between work tasks”, Organizational Behavior and Human Decision Processes, 2009.
- S. Schmidt, “Everyone Gets Jevons Paradox Wrong”, Amazing CTO, 2026.
- M. Liberman, “Low-hanging fruit: the history”, Language Log, August 3, 2022 (quoting the Oxford English Dictionary).
- Z. Crockett, “Should You Literally Pick the Low-Hanging Fruit?”, Priceonomics, February 5, 2016.
- “Low-hanging fruit”, Merriam-Webster.
- Team figures: pulled from GitHub on October 6, 2026, merged human PRs, excluding dependency files.