Writing · September 2026

Teaching Claude Code to pace itself against the weekly quota

A status-line script, one line of arithmetic, a check every two hours, and a hook that survives restarts. The paragraph of instructions turned out to matter more than the shell.


A Claude subscription gives you a weekly quota and a five-hour window. Claude Code works hard and doesn’t know either of them exists. For interactive use that’s fine. Once it runs a handful of agents in parallel for days at a time, it’s a problem in both directions: spend too fast and the week ends early, with work stranded mid-branch; spend too slowly and paid capacity goes unused.

I hit the first one last weekend. I had seven Opus agents working in parallel, and the weekly meter went from about 12% to 52% in roughly a day. At that rate the week would have run out Monday afternoon, sixteen hours before the reset. So I asked Claude to pace itself: aim to land at 100% right at the reset, not before and not well short of it.

It took four pieces. None of them is clever, and together they’ve worked better than I expected.

A line chart of weekly quota used, in percent, from Saturday noon to the Tuesday 9 AM reset. A red dotted line starts at 12% on Saturday and climbs at 1.7 points an hour, reaching 100% on Monday around 5 PM. A green dashed target line runs straight from 52% on Sunday midday to 100% at the reset. The blue stepped line of actual usage starts at 52% and tracks a few points above the green line, reaching 92% by Monday night.
Weekly quota used, from the status line. The red dotted line is the rate before pacing, extended forward. The green dashed line is the target: a straight line from where pacing started to 100% at the reset. The Saturday starting point is approximate; everything from Sunday midday on is logged.

1. Let Claude see the meter

Claude Code’s status line runs a command you configure and pipes it a JSON blob. On a subscription, that blob includes a rate_limits object: five_hour and seven_day, each with used_percentage and resets_at. The status line is the one place Claude Code publishes those numbers, so the script writes them to a file Claude can read, and appends a line to a history file whenever they change:

S5=$(echo "$input" | jq -r '.rate_limits.five_hour.used_percentage // empty')
S7=$(echo "$input" | jq -r '.rate_limits.seven_day.used_percentage // empty')
if [ -n "$S5$S7" ]; then
  echo "$input" | jq -c --arg at "$(date -u +%FT%TZ)" \
    '{at: $at, rate_limits: .rate_limits}' > ~/.claude/usage-latest.json
  KEY="$S5|$S7"
  if [ "$KEY" != "$(cat ~/.claude/usage-last-key 2>/dev/null)" ]; then
    echo "$KEY" > ~/.claude/usage-last-key
    echo "$input" | jq -c --arg at "$(date -u +%FT%TZ)" \
      '{at: $at, five_hour: .rate_limits.five_hour, seven_day: .rate_limits.seven_day}' \
      >> ~/.claude/usage-history.jsonl
  fi
fi

(The real script writes to a temp file and renames it, so a reader never sees half a file.) The rule that goes with it: read the file, never guess. Before this, “how much quota is left?” got an estimate. Now it gets a number with a timestamp.

2. One line of arithmetic

The target rate is the quota left, divided by the hours left:

target = (100 − weekly % used) ÷ hours until seven_day.resets_at

The actual rate is the change in weekly % over the last couple of hours, from the history file. Compare the two:

The 20% band matters. Usage arrives in lumps (one agent finishing a big test run can move the meter two points), and without a band the check flips back and forth every run.

3. Throttle concurrency, not quality

The obvious ways to save quota are a cheaper model, lower effort, or skipping the independent review. I ruled all three out. The independent review catches real defects most times it runs; tonight it sent back a change with a flaw that every test had passed.

So the only lever is how many lanes run at once, a lane being one branch’s author plus its reviewer. Two is the cap. The pacing check picks a number between zero and two; it never touches how each lane does its work. That split made the instruction easy to follow: Claude never has to judge whether a review is “worth it” this week.

4. A clock, and a way to survive restarts

A rule Claude applies only when it happens to think of it isn’t a rule. Claude Code’s CronCreate schedules a prompt into the session, so the check runs every two hours (17 */2 * * *, off the hour on purpose), with a prompt that spells out the steps above.

There are two catches. Cron jobs live only in the session, and recurring ones expire after seven days. So a SessionStart hook in ~/.claude/settings.json prints a short reminders file into every new session:

"SessionStart": [{ "hooks": [{ "type": "command",
  "command": "cat ~/.claude/session-start-reminders.md 2>/dev/null || true" }] }]

The first line of that file says to recreate the pacing check. When I restarted the container this weekend, the new session read it, recreated the cron and carried on. To change what gets recreated, I edit a text file.

The failure I didn’t expect: an idle lane

The first bug went the other way. Both lanes finished and deployed within a few minutes of each other, then sat empty for about twenty-five minutes while Claude processed something else. I noticed only because I looked. The two-hour check judges the pace; it doesn’t notice an empty lane in between.

The fix was a second rule: refill on completion, not on the clock. When a lane finishes, start the next queued item in the same turn, unless the pace says slow down. An idle lane while the pace allows work counts as a defect, not a rest.

What it looks like in practice

Most checks say nothing. When one does speak, it’s a line or two. Tonight’s, with 12.6 hours to the reset:

Weekly 91%, target 0.71 pts/h, actual 0.89 over the last two hours (about 25% over). Holding: the running lane finishes its checks, and its re-review and the next lane start after the reset.

So the week lands slightly early, and it lands on purpose. On the chart, the blue line runs a few points above the green one the whole way. That’s the 20% band doing its job: close enough that nothing strands, loose enough that the check isn’t twitchy.

What I’m still measuring

I also log tokens for each interval between meter changes, by type and model, from Claude Code’s own transcripts, and fit how many quota points a million tokens of each type costs. The raw ratios so far: about 100k output tokens and about 60M cache-read tokens per weekly point. Long-running agents re-read their whole context on every call, so cache reads dominate by volume. The fit itself isn’t stable yet. Most intervals move the meter by only a point, and the token types rise and fall together, so the regression can’t pull them apart. I’ll report it when it settles.

If you want to copy this

It’s about forty lines of shell and a paragraph of instructions. The part that took a weekend was finding out that the paragraph matters more than the shell.