AI · · 9 min read

The apple tree and the basket.

With AI agents, the tree shrinks and almost every apple comes within reach. The basket stays the same size.

Translated from French with AI assistance. Read the original

Contents 8
  1. In one minute
  2. The tree shrinks
  3. Sorting changes
  4. The bottleneck moves
  5. How far can the basket grow?
  6. One picker, five trees
  7. The orchard rules
  8. The data

For a year now, my team and I have been working with AI agents. We hand them whole tasks: build a feature, migrate a service, write tests. It goes fast, sometimes very fast.

To explain what that changes, to developers as much as to people who have never opened a code editor, I use a single image: an apple tree and a basket. It starts from a phrase you hear in every meeting: low-hanging fruit, the fruit hanging low, the one you pick without a ladder. This article unfolds the image, with its opportunities and its dangers, then puts it up against the numbers: my team’s, and a year of my own days.

In one minute

A project is a tree, each task is an apple. Agents shrink the tree. They don’t make the apples any lighter, or the basket any bigger.

  • The appleA piece of work worth doing: a feature, some debt, a tool, a piece of content.
  • The heightThe effort to reach it. With agents, it collapses.
  • The weightThe risk, and the time to taste it: review, test, ship. It doesn't move.
  • The basketWhat we can taste and carry. It doesn't grow on its own.

The tree shrinks

You used to need a 10-meter ladder. Now a stepladder will do. But some apples still weigh 50 kilos.

Effort used to decide. We picked the apples at the bottom, the ones within arm’s reach, and the top of the tree waited: tests for the old code, the version upgrade, the security audit. “Someday, maybe.”

That’s the meeting-room low-hanging fruit. The phrase is in print as early as 1968, and becomes a management tic in the 1980s (Language Log). Orchards, for their part, changed before we did: dwarf trees 1.5 to 2 meters tall replaced the big apple trees, and growers now plant up to 2,000 per acre, against 20 before. A geneticist interviewed by Priceonomics draws the conclusion: “The phrase is irrelevant” (Priceonomics). With agents, our projects are going through the same thing.

With an agent, a task of several hours starts with one sentence. And it’s improving fast: since 2024, the length of tasks an agent completes on its own has roughly doubled every three months (METR, January 2026). Measured at a 50% success threshold, METR points out: you still have to check.

But height isn’t weight. Changing a button’s color and rewriting an application in Rust are now one sentence away. The first can be checked in thirty seconds. The second takes weeks of review, testing and rollout, and it can break everything. An apple’s weight is the risk, and the time it takes to taste it. Agents don’t change that.

2 m4 m6 m8 m10 mButton color: level 5 before, 1 with agents. 1 kg, before and after.Button color: level 5 before, 1 with agents. 1 kg, before and after.New field: level 7 before, 1 with agents. 2 kg, before and after.New field: level 7 before, 1 with agents. 2 kg, before and after.Small A/B test: level 10 before, 1 with agents. 2 kg, before and after.Small A/B test: level 10 before, 1 with agents. 2 kg, before and after.Up-to-date docs: level 14 before, 2 with agents. 1 kg, before and after.Up-to-date docs: level 14 before, 2 with agents. 1 kg, before and after.Tests for old code: level 19 before, 2 with agents. 6 kg, before and after.Tests for old code: level 19 before, 2 with agents. 6 kg, before and after.Accessibility: level 24 before, 3 with agents. 5 kg, before and after.Accessibility: level 24 before, 3 with agents. 5 kg, before and after.Site in three languages: level 28 before, 3 with agents. 10 kg, before and after.Site in three languages: level 28 before, 3 with agents. 10 kg, before and after.Version upgrade: level 33 before, 4 with agents. 15 kg, before and after.Version upgrade: level 33 before, 4 with agents. 15 kg, before and after.Security audit: level 38 before, 4 with agents. 8 kg, before and after.Security audit: level 38 before, 4 with agents. 8 kg, before and after.Internal tools: level 42 before, 5 with agents. 4 kg, before and after.Internal tools: level 42 before, 5 with agents. 4 kg, before and after.Design system overhaul: level 46 before, 5 with agents. 20 kg, before and after.Design system overhaul: level 46 before, 5 with agents. 20 kg, before and after.Rust rewrite: level 49 before, 5 with agents. 50 kg, before and after.Rust rewrite: level 49 before, 5 with agents. 50 kg, before and after.Rust rewrite · 50 kgBefore · a 10 m ladderWith agents · a stepladder will do
Twelve typical items for a product team. Before, height decided: we picked the bottom, the top waited. With agents, everything is within arm's reach. But the apples don't slim down: the Rust rewrite still weighs 50 kg, as much as the whole basket.

It’s a real opportunity. A field study of an autonomous agent shows that people take on more complex, multi-step projects, and step outside their original line of work: they go from operators to supervisors (HBR, July 2026). The top of the tree finally comes within reach.

Sorting changes

When everything is within reach, “it’s easy” no longer justifies anything. Only one question remains: is it worth it?

Effort used to do the sorting for us. We ranked topics on two axes, effort and value, and dreamed of quick wins: little effort, lots of value. That cell was almost always empty, because the low, sweet apples had been picked long ago. “All the low-hanging fruit has been picked”: the line already appears in New Scientist in 2003, about new drugs (Language Log).

Sorting by effort has a known flaw. Under heavy load, people pick the easy task: 76% of participants in one experiment, against 64% when the load is light. It feels like progress, but performance drops in the long run: the emergency physicians studied who take on more difficult cases improve more (KC, Staats, Kouchaki and Gino, Management Science).

With agents, effort no longer breaks the tie: everything lands in the quick wins cell. What’s left is value, and the basket. And the basket fills up by weight, not by count.

1234501020304050effort: height in the old tree →valueQuick winsempty cellButton color: effort 5, value 1, 1 kg11New field: effort 7, value 1, 2 kg12Small A/B test: effort 10, value 2, 2 kg10Up-to-date docs: effort 14, value 3, 1 kg7Tests for old code: effort 19, value 4, 6 kg4Accessibility: effort 24, value 4, 5 kg3Site in three languages: effort 28, value 5, 10 kg1Version upgrade: effort 33, value 4, 15 kg6Security audit: effort 38, value 4, 8 kg5Internal tools: effort 42, value 3, 4 kg8Design system overhaul: effort 46, value 5, 20 kg2Rust rewrite: effort 49, value 3, 50 kg9Ranked by value1Site in three languages10 kg2Design system overhaul20 kg3Accessibility5 kg4Tests for old code6 kg5Security audit8 kg6Version upgrade15 kg7Up-to-date docs1 kg8Internal tools4 kg9Rust rewrite50 kg10Small A/B test2 kg11Button color1 kg12New field2 kgthe basket is full: 49 kg out of 50
The same twelve items as in the orchard; each apple's size follows its weight. When effort no longer decides, choosing becomes the job: rank by value, then fill the basket by weight, in order and without skipping. Version upgrade (15 kg) doesn't fit: we stop there. Up-to-date docs (1 kg) would still fit, but we don't skip a more useful apple for a lighter one: that's on purpose.

In my team, the share of tests in added code has gone from 12.8% to 21.6% since we started working with agents. That’s an apple from the top that we’re finally starting to pick.

The bottleneck moves

Picking costs almost nothing now. Choosing, tasting, shipping and maintaining still cost the same.

Tasting, in software, is review: someone reads the code before it goes to production. Code arrives as a pull request (PR): a proposed change that a colleague reviews and then merges. Then it has to ship, and be maintained for years: monitoring, updates, security.

in prodBeforeone apple picked, one apple tastedpicktasteshipin prod: tasted + untastedWith agentssame basket, three times as many applesshipped without tasting

tasted shipped untasted

Before basket load 100% · 12 picked · 9 tasted

With agents basket load 300% · 38 picked · 10 tasted · 13 shipped untasted

Same basket, same pace: the person tastes one apple at a time, at the same speed on both lines. With agents, three times as many apples arrive. The surplus goes to prod untasted, or gets caught up in overtime.

This is the Jevons paradox, observed in 1865 with coal: when a resource gets more efficient, we use more of it, not less. The easier picking gets, the more we pick. For a CTO, the real risk is producing “more software than our organizations can safely operate, secure, and maintain” (Mike Grouchy, quoted by Amazing CTO).

I see it in the team I lead. Since everyone started working with agents, each person merges nearly half again as many PRs, and PRs four times bigger. Time to merge hasn’t budged.

Merged PRs per person per month+46%

Before (Jan 2021 – Jan 2025): 15.115.1BeforeWith agents (Oct 2025 – Sep 2026): 22.122.1With agents

Median PR size, in lines×4

Before (Jan 2021 – Jan 2025): 3838BeforeWith agents (Oct 2025 – Sep 2026): 161161With agents

Median time to mergestable

Before (Jan 2021 – Jan 2025): 24.5 h24.5 hBeforeWith agents (Oct 2025 – Sep 2026): 24.4 h24.4 hWith agents

×2.2 more code to review every month: 61,000 lines before, 132,000 with agents.

11 → 26% of PRs exceed 500 lines: one in four, against one in nine before.

The team I lead, 12,628 PRs pulled from GitHub on October 6, 2026: Jan 2021 – Jan 2025 versus Oct 2025 – Sep 2026. About 5.8 active people before, 4.8 with agents. Sizes exclude dependency files and generated code. A descriptive study: other things changed at the same time.

The basket is the same, and it gets twice as much code to review. Either we review twice as fast, or we taste less. That’s the question to ask as a team.

How far can the basket grow?

The basket can grow, but not by adding agents: by tooling up the tasting.

Adding agents fills the basket faster. And supervising several agents is tiring: in a BCG survey of 1,488 employees, 14% describe mental fatigue from supervising AI tools. Perceived productivity rises up to three tools used in parallel, and drops from four onward (HBR, March 2026; Fortune).

The other lever is to have machines taste some of the apples: automated tests, continuous integration that runs them on every PR, automated review. It’s the only one that grows the basket without wearing people out. Play with both sliders.

Picked this week
12
Verified by tooling
2
To taste by hand
10
Basket size
10

Stretched load 100%

verified by toolingtasted by hand

The basket is full. At the slightest surprise, we taste faster, so worse.

An illustrative model, not a measurement: an agent picks about 4 average-weight apples a week, a person tastes about 10, and beyond 3 parallel agents, each one costs them 2 basket slots (beyond 3 tools, BCG observes a drop in perceived productivity).

One picker, five trees

The basket is also attention. Mine has broken into crumbs: on the days when I have the most time, I switch projects twenty times instead of once.

Context switching is nothing new. Meetings, Slack and “got two minutes?” have always chopped up our days. With agents, I’m the one doing the chopping: I start a session in one repository, another in a second one, and I go back and forth while they work.

For a year now, a collector has filed all my Claude Code and Codex sessions in a journal. I asked it two questions: how many projects I touch per day, and how long I stay on one before moving to the next. In lean, you draw this kind of path on the floor to see wasted motion: it’s a spaghetti diagram. Here’s mine, over two typical days with comparable meeting loads.

In real life, my projects don’t look alike: each has its own code, its own language, its own people. Switching projects isn’t moving from one apple tree to another, it’s putting down the pruning shears to pick up the hammer. So here, each project is a different workstation, with its own tool.

A · orchardB · beehiveC · workbenchD · wheelbarrowE · vineyard
A year ago Monday, October 6, 2025
1 switch blocks of 152 min
A year later Friday, September 11, 2026
20 switches blocks of 13 min
Two real days with comparable meeting loads, replayed in 30 seconds; stretches with no messages are shortened, the clock skips the gaps. Each project is a different workstation, with its own tool (orchard, beehive, workbench, wheelbarrow, vineyard, assigned to the letters at random): switching projects means switching trades. Me today as a solid figure, me a year ago faded, with a dotted trail. The ring around the station I'm at takes the project's color, like the bar below.

Three two-month windows, a year apart between the first and the last. For each window, the median of my working days, and in parentheses the range where half the days fall.

Oct–Nov 2025 Mar–Apr 2026 Aug–Sep 2026
Projects per day 1 (1–2) 2 (1–2) 4.5 (4–5)
Project switches per day 0 (0–2) 1 (0–3) 14.5 (5–25)
Average block on one project 110 min 117 min 27 min

Until March, nothing moves: one or two projects a day, blocks of nearly two hours. I was already working with agents, but on one workstream at a time. In April, blocks get cut in half. The juggling comes next: ten switches a day in July, around fifteen in August and September.

A day packed with meetings leaves less room for agents: the comparison could have been biased. I redid it at equal meeting load, on the calmest days of each window. The gap holds, and it widens: one switch a day and 81-minute blocks in 2025, 21 switches and 13-minute blocks in 2026. It’s when I have time that I scatter the most, because that’s when I run agents in parallel.

These numbers don’t measure the quality of the work, only its shape. But a day made of 13-minute blocks no longer looks like the one I used to know.

The orchard rules

Choose by value, finally pick the top, weigh before you pick, and only carry what you can taste.

  1. Choose by value. Everything is within reach, so “it’s easy” no longer justifies anything. Ask instead: how much does it bring in, for whom, and who’s going to maintain it.
  2. Finally pick the top. Tests for the old code, version upgrades, documentation, accessibility, the security audit: what we kept putting off for lack of a ladder moves to the top of the list. Professional pickers already work this way: “You always need to start by picking apples from the top, never the lowest-hanging fruits first” (an apple picker from Washington State, to Priceonomics).
  3. Weigh before you pick. Before starting an agent, estimate the time it will take to check its work. A 50-kilo apple gets cut into apples you can taste one at a time.
  4. Tool up the tasting. Tests, continuous integration, automated review: that’s where the time saved goes, because that’s what really grows the basket.
  5. Four slots, not four sessions. One topic per kind: business, technical foundations, team, exploration. Enforcing the categories keeps a hard topic open, instead of giving in to the ones that feel good. But four open topics don’t mean four agents in parallel:
    • alternate by half-day, leaving a topic at a stopping point: switching tasks costs more when you leave the previous one unfinished, that’s “attention residue” (Leroy, 2009);
    • no more than three active agent sessions at a time;
    • no fifth topic: to open one, finish or freeze another.
  6. Batch reviews, instead of reacting to each agent as soon as it’s done.

We used to ask: what can we reach? Now: what deserves picking, and how much can we carry? See you in six months to find out whether my curves come back up.

The data

The spaghetti is replayed from real days, and the numbers come from the same journal. If you want to check, it’s all here.

See the real data

The three typical days

For each window, among the window's days with the fewest meetings and at least three hours of activity, the one closest to those days' medians. One row per project, one bar per block, one tick per typed message.

M1 · Monday, October 6, 2025: 2 projects, 1 switch, blocks of 152 min on average.
9am11am1pm3pm5pm7pmABA: 13:14 → 15:00, 9 messagesB: 16:09 → 18:17, 19 messages
M6 · Tuesday, April 7, 2026: 2 projects, 1 switch, blocks of 133 min on average.
9am11am1pm3pm5pm7pmABA: 10:13 → 12:22, 18 messagesB: 14:27 → 14:40, 3 messages
M12 · Friday, September 11, 2026: 5 projects, 20 switches, blocks of 13 min on average.
9am11am1pm3pm5pm7pmABCDEA: 13:37 → 13:45, 3 messagesB: 13:46 → 13:46, 1 messageA: 13:55 → 14:02, 2 messagesC: 14:13 → 14:13, 1 messageD: 14:13 → 15:08, 14 messagesA: 15:19 → 15:19, 1 messageD: 15:20 → 16:14, 8 messagesA: 16:19 → 16:19, 1 messageD: 16:27 → 16:27, 1 messageA: 16:29 → 16:37, 2 messagesD: 16:39 → 17:07, 6 messagesA: 17:15 → 17:20, 2 messagesD: 17:21 → 17:35, 6 messagesA: 17:37 → 17:41, 3 messagesE: 17:43 → 17:43, 1 messageA: 17:43 → 17:43, 1 messageD: 17:44 → 17:46, 2 messagesE: 17:50 → 17:50, 1 messageD: 17:50 → 17:50, 1 messageE: 17:54 → 17:55, 2 messagesA: 17:58 → 17:59, 2 messages

Month by month

Medians of each month's working days. Darker: the months of the three windows being compared.

Project switches per day (median)
Oct 2025: 1 (17 working days)1Oct2025Nov 2025: 0 (16 working days)0NovDec 2025: 0 (8 working days)0DecJan 2026: 1 (20 working days)1Jan2026Feb 2026: 2 (14 working days)2FebMar 2026: 1 (15 working days)1MarApr 2026: 1 (16 working days)1AprMay 2026: 1 (13 working days)1MayJun 2026: 4 (17 working days)4JunJul 2026: 10 (19 working days)10JulAug 2026: 25 (11 working days)25AugSep 2026: 10 (17 working days)10Sep
Average block on one project, in minutes (median)
Oct 2025: 125 (17 working days)125Oct2025Nov 2025: 103 (16 working days)103NovDec 2025: 154 (8 working days)154DecJan 2026: 136 (20 working days)136Jan2026Feb 2026: 118 (14 working days)118FebMar 2026: 153 (15 working days)153MarApr 2026: 59 (16 working days)59AprMay 2026: 58 (13 working days)58MayJun 2026: 45 (17 working days)45JunJul 2026: 39 (19 working days)39JulAug 2026: 14 (11 working days)14AugSep 2026: 44 (17 working days)44Sep

Block length, day by day

Share of each window's days by that day's average block length. The darker, the longer the blocks.

M124%21%49%M613%19%13%48%M1236%18%21%21%< 15 min15–30 min30–60 min1–2 h≥ 2 h

At equal meeting load

A day packed with meetings leaves less room for agents. On each window's lightest meeting days, the gap holds, and it widens.

M1 · Oct–Nov 2025M6 · Mar–Apr 2026M12 · Aug–Sep 2026
Days kept11 of 3314 of 3110 of 28
Projects per day225
Switches per day1121.5
Average block81 min86 min13 min

How it's measured

  • Source: messages typed into Claude Code and Codex, on two Macs, collected by my agent journal. No meetings, no Slack, no email: only my exchanges with agents.
  • Scope: working days with at least 6 messages, 9am to 7pm, Paris time; excluding time off from August 20 to September 4, 2026.
  • Project: a git repository or a working folder. Names are replaced with letters, in order of appearance during the day.
  • Switch: two consecutive messages sent to two different projects. Average block: the time between the day's first and last message, divided by the number of blocks. It exists even on days without any switch.
  • Windows: two months each (M1 Oct–Nov 2025, 33 days; M6 Mar–Apr 2026, 31 days; M12 Aug–Sep 2026, 28 days). Medians, not means: a few extreme days would pull the mean.
  • Limits: a descriptive study of a single person. Meetings are context switches too, and this log doesn't count them. In October 2025, almost every message comes from just one of the two Macs.

Sources