Claude Code Harness Example: A Full-Stack Monorepo Setup
A working Claude Code harness from a startup monorepo: blocking hooks, skills that open and answer pull requests, a CLAUDE.md byte budget, and what broke.

I am a full-stack developer at a startup, and for the last sixteen weeks most of the code in the monorepo I work on has gone through Claude Code. The repository holds an API, a few frontends, a shared UI package and a docs site. What follows is the harness I built around the agent to make that safe: which parts held, which parts drifted, and the code patterns I would copy into the next repo.
The short version: every rule that lived only as prose in the model’s context eventually got ignored, and every rule that lived in a script did not. Hooks, permission rules, lint checks and CI are what held the line. The CLAUDE.md files mattered, but mostly as a pointer to the things that enforce.
This is the worked example. For what each primitive is and what it costs in context, read the Claude Code harness guide. For picking a workflow (plan mode, fan-out, worktrees), read the agentic workflows playbook. All code below is a generic reconstruction of the pattern, with placeholder paths, not the production files.
The shape of the harness
The harness has six layers, and each exists because the one above it proved insufficient at some point:
- Instruction files. A root
CLAUDE.md, one per app, and path-scoped rules. - Permission rules. Deny edits to generated files outright.
- PreToolUse hooks. Ask or deny before a risky shell command runs.
- PostToolUse and prompt hooks. Lint after every edit, route prompts to skills, check review threads after a push.
- Skills and one subagent. Repeatable GitHub workflows (open a PR, review a PR, answer a review) and a read-only code reviewer.
- Git hooks and CI. The same guards again, for the commands a human types and for anything the agent layer missed.
The figure below maps each guard to the stage of a change where it fires. The left lane runs inside the Claude Code session. The right lane runs in git and CI, which is the backstop.
| Stage | In the agent session | In git and CI |
|---|---|---|
| Session start | Stale build warning (SessionStart) | None |
| Prompt | Skill router (UserPromptSubmit) | None |
| Edit a file | Edit deny on generated files; Lint and format check (exit 2) | None |
| Create branch | Issue-number branch guard (deny) | Branch name check (pre-commit) |
| Commit | Commit guard (ask) | Lint, format, secret scan |
| Push | Push guard (ask); Unanswered review threads | Required checks start |
| Review | Read-only reviewer subagent; Review and answer skills | Code owners review |
| Merge | Merge guard (ask) | Five required checks green |
CLAUDE.md got a byte budget after it hit 31.7KB
The first failure was the root CLAUDE.md itself. It started as a helpful page and grew one “just remember this” bullet at a time. There was a line-count ceiling written into it, and the file stayed under that ceiling while it swelled to 31.7KB, because the bullets had grown past 250 characters each. The line count measured the wrong thing.
The cost is easy to forget. Anthropic’s memory documentation says CLAUDE.md files load into every session and recommends targeting under 200 lines per file, because “longer files consume more context and reduce adherence.” A paragraph added to the root file is a paragraph re-sent on every turn for the life of the repository.
So the budget became a script that runs in CI. It measures bytes, since bytes track tokens at roughly 4:1 without needing a tokenizer, and it discovers files rather than listing them, so a new app is covered the day it lands:
// scripts/check-claude-md-size.mjs (a generic version of the pattern)
import { readdirSync, statSync, existsSync } from 'node:fs';
import { join } from 'node:path';
const BUDGETS = { root: 14_000, app: 10_000, rule: 8_000 };
const failures = [];
function check(file, budget) {
if (!existsSync(file)) return;
const bytes = statSync(file).size;
if (bytes > budget) failures.push(`${file}: ${bytes} bytes (budget ${budget})`);
}
check('CLAUDE.md', BUDGETS.root);
for (const app of readdirSync('apps')) check(join('apps', app, 'CLAUDE.md'), BUDGETS.app);
for (const rule of readdirSync('.claude/rules')) check(join('.claude/rules', rule), BUDGETS.rule);
if (failures.length) {
console.error('Over budget. Relocate the prose to the docs site, never raise the budget.');
console.error(failures.join('\n'));
process.exit(1);
}
The policy line that matters is in the error message: relocate, never raise. Explanations move to the docs site, and the instruction file keeps a one-line pointer. The real version also budgets the worst case (root file plus the largest app file) and the combined size of rules that share a path glob, because rules with the same scope load together.
Every app file follows the same shape: a warning at the top when the framework moves faster than training data, one paragraph on what the app is, a fenced block of commands, and bold imperative rules that each end with a pointer to the explanation. A rule reads like this:
- **Never edit `src/routeTree.gen.ts`.** It is generated and Edit-denied.
Why: docs/explanation/routing.md
Path-scoped rules instead of one giant file
Conventions that only matter in one part of the tree live in .claude/rules/, with a paths list in the frontmatter. The memory documentation states that these conditional rules “only apply when Claude is working with files matching the specified patterns,” and trigger when Claude reads a matching file, not on every tool use.
---
paths:
- 'apps/api/**'
- 'packages/contract/**'
---
# API contract conventions
- Define the procedure in `packages/contract` first, then implement it in `apps/api`.
- Every input schema is bounded: strings get a max length, arrays get a max size.
- Regenerate the OpenAPI spec with `pnpm openapi:generate`. Never edit it by hand.
The repository ended up with about a dozen of these: API conventions, dependency management, infrastructure, React rendering discipline, UI component sourcing. The one I would write first in any repo is a rule that says a skill is not authority on a library’s API. Vendored best-practice skills go stale, so the rule tells the agent to check the installed package version before trusting a pattern from a skill.
Deny edits to generated files, and use the right rule name
Generated files (route trees, the OpenAPI spec, the lockfile, infrastructure lock files) are denied outright in the shared settings file:
{
"permissions": {
"deny": [
"Edit(apps/web/src/routeTree.gen.ts)",
"Edit(apps/api/openapi.json)",
"Edit(pnpm-lock.yaml)"
]
}
}
The gotcha that cost me an afternoon: the rule has to be Edit(...). The permissions documentation says Edit rules apply to all built-in tools that edit files, and that a path rule written for Write is accepted but never consulted. Regeneration still works, because the generator runs through Bash and writes the file as a subprocess, which path rules do not cover.
PreToolUse hooks: ask for side effects, deny for broken conventions
This is where the harness started to feel trustworthy. Three PreToolUse hooks run on every Bash call. The hooks are wired in .claude/settings.json:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/git-mutation-guard.sh"
},
{
"type": "command",
"command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/branch-name-guard.sh"
},
{
"type": "command",
"command": "node",
"args": ["${CLAUDE_PROJECT_DIR}/.claude/hooks/infra-guard.mjs"]
}
]
}
]
}
}
The Node hook uses exec form (command plus args). The hooks reference says that when args is set the hook is spawned directly with no shell, which makes the hook behave the same on a teammate’s Windows machine.
The ask guard
My local allowlist auto-approves git and gh commands, because approving every git status is miserable. That creates a hole: the agent could commit, push and open a pull request without a single prompt. The first hook closes it by forcing a prompt for exactly those mutations:
#!/usr/bin/env bash
# .claude/hooks/git-mutation-guard.sh
cmd=$(jq -r '.tool_input.command // empty' 2>/dev/null) || exit 0
[ -z "$cmd" ] && exit 0
# Regexes live in variables: macOS bash 3.2 cannot parse ;& inside an inline [[ =~ ]].
git_re='(^|[[:space:];&|])git[[:space:]][^;&|]*(commit|push)([[:space:]]|$|[;&|])'
gh_re='(^|[[:space:];&|])gh[[:space:]]+pr[[:space:]]+(create|merge)([[:space:]]|$|[;&|])'
if [[ "$cmd" =~ $git_re || "$cmd" =~ $gh_re ]]; then
jq -n '{hookSpecificOutput: {
hookEventName: "PreToolUse",
permissionDecision: "ask",
permissionDecisionReason: "Commits, pushes and PR changes always need a human yes."
}}'
fi
exit 0
The [^;&|]* stops a match from running across a command separator. The pattern is deliberately loose: a false positive costs one extra prompt, and a false negative is an unreviewed push.
This hook is also why the create-pr skill can stay model-invocable. Claude can decide on its own that a pull request is the next step, draft the title and body, and run gh pr create, and I still approve the actual mutation.
The deny guard
Branch names in the repo must start with the issue number (123-fix-login-redirect), so every branch traces back to an issue. The second hook returns deny instead of ask, because there is nothing to approve. The name is simply wrong:
# .claude/hooks/branch-name-guard.sh (the core of it)
search_cmd="${cmd%%<<*}" # ignore everything after a heredoc marker
branch=$(printf '%s' "$search_cmd" |
sed -nE 's/.*(checkout -b|switch -c|worktree add -b)[[:space:]]+([^[:space:]]+).*/\2/p')
[ -z "$branch" ] && exit 0
issue_prefix_re='^[0-9]+-'
[[ "$branch" =~ $issue_prefix_re ]] && exit 0
jq -n --arg b "$branch" '{hookSpecificOutput: {
hookEventName: "PreToolUse",
permissionDecision: "deny",
permissionDecisionReason: ("Branch \($b) must start with its issue number, e.g. 123-short-slug.")
}}'
The heredoc trim came from a real false positive. A commit message that mentioned git checkout -b in its body got denied, because the hook was reading the message as a command. Cutting the string at the first << fixed it.
The rule every hook follows
A broken hook must never break the session. If jq is missing, the JSON does not parse, or gh is not logged in, the hook exits 0 with no output and the normal permission flow takes over. The hook that manages infrastructure commands (asking before apply, destroy or state surgery) is written in Node rather than bash for the same reason. Its header says it best: a guard that quietly stops guarding is worse than no guard, because you think you are covered.
Lint after every edit, and exit 2 on failure
The highest-value hook per line of code is the one that lints the file Claude just edited:
#!/usr/bin/env bash
# .claude/hooks/lint-after-edit.sh, registered on PostToolUse with matcher "Edit|Write"
file=$(jq -r '.tool_input.file_path // empty' 2>/dev/null) || exit 0
case "$file" in *.ts|*.tsx|*.js|*.jsx|*.mjs|*.cjs) ;; *) exit 0 ;; esac
problems=$(pnpm exec oxlint "$file" 2>&1) || true
fmt=$(pnpm exec oxfmt --check "$file" 2>&1) || problems="$problems
$fmt
Fix formatting with: pnpm exec oxfmt $file"
if [ -n "$(printf '%s' "$problems" | grep -E 'error|warning|Fix formatting')" ]; then
printf '%s' "$problems" >&2
exit 2
fi
exit 0
On PostToolUse, the hooks reference notes that exit code 2 cannot undo the edit (the tool already ran), but stderr is shown to Claude. In practice the agent reads the lint output and fixes it in the same turn, before I ever see the diff. A clean file exits 0 silently, so the hook costs no context on the happy path. That silence is the design goal for every hook in the repo.
A regex skill router beat model judgement
The repository has about 27 skills. Fourteen are vendored best-practice packs for the frameworks in use, and about ten are home-grown workflows: open a PR, review a PR, answer a review, file an issue, add an endpoint, write docs, explain a change. The problem was that Claude often improvised a workflow instead of loading the skill that described it. “Open a PR for this” would produce a pull request with the wrong body format and no issue labels.
The fix is a UserPromptSubmit hook that matches the prompt against a table of regexes and, on a hit, adds a line of context naming the skill:
// .claude/hooks/skill-router.mjs (condensed)
const ROUTES = [
{
skill: 'create-pr',
pattern: /\b(open|create|raise|put up|submit)\b[^.?!]{0,40}\b(pr|pull request)s?\b/i,
},
{
skill: 'answer-review',
exclusive: true, // a pasted review mentions everything; this route silences the rest
pattern: [/\b(review|finding|comment)s?\b/i, /\b(fix|address|resolve|answer)\b/i],
},
];
const input = JSON.parse(await new Response(process.stdin).text());
const hits = ROUTES.filter((r) => [r.pattern].flat().every((re) => re.test(input.prompt)));
const kept = hits.some((h) => h.exclusive) ? hits.filter((h) => h.exclusive) : hits;
if (kept.length) {
console.log(
JSON.stringify({
hookSpecificOutput: {
hookEventName: 'UserPromptSubmit',
additionalContext:
'This repo has a skill for what was just asked. Use it rather than improvising:\n' +
kept.map((h) => `- /${h.skill}`).join('\n'),
},
}),
);
}
The patterns are loose on the verb and tight on the noun, and [^.?!]{0,40} keeps a match inside one sentence. The “exclusive” flag came from a specific mess: pasting a whole review and saying “please fix these” fired every route, because the quoted review mentioned pull requests, endpoints and docs.
The comment at the top of the real file explains why this is a regex and not a smarter model call: judgement is what already existed, and it is what missed. The hook also ships a --self-test flag that runs 27 prompt fixtures, including near-misses that must stay silent (“what does this pull request do” should route nowhere). A git hook runs those tests whenever a file under .claude/hooks/ is staged.
The GitHub loop: open, review, answer
The skills that paid for the whole harness are the three that cover a pull request’s life. They are ordinary SKILL.md files in .claude/skills/<name>/, and each one mostly encodes GitHub gotchas I would otherwise re-learn weekly.
Opening a PR. The create-pr skill copies labels from the linked issue and fixes the body shape: a plain-language summary, what changed, what did not change, the checks that were run, a one-sentence version, and Closes #N (or Refs #N when the PR is only one step of a larger issue).
labels=$(gh issue view 123 -R OWNER/REPO --json labels -q '.labels[].name' | paste -sd, -)
gh pr create -R OWNER/REPO \
--title "feat(api): add rate limit headers" \
--body-file pr-body.md --assignee "@me" --label "$labels"
The explicit -R OWNER/REPO is there because an SSH remote alias broke gh repository inference, and every skill now passes it.
Reviewing a PR. gh pr review cannot attach inline comments to lines, so the review-pr skill posts one submission through the REST reviews endpoint:
cat > review.json <<'JSON'
{
"event": "COMMENT",
"body": "Two blocking findings, one nit.",
"comments": [
{ "path": "apps/api/src/users.ts", "line": 42, "side": "RIGHT",
"body": "**issue (blocking):** this query runs once per row. Batch it." }
]
}
JSON
gh api repos/OWNER/REPO/pulls/57/reviews -X POST --input review.json
Two gotchas are written into the skill because both bit: a line outside the diff returns a 422 and fails the entire submission, and GitHub does not let pull request authors approve their own pull requests, so on a self-review the skill always submits COMMENT.
Answering a review. The answer-review skill lists unresolved threads with GraphQL, decides each one (fix it, defer it to a new issue, or push back), commits once, and pushes before replying. A reply that says “Fixed in a1b2c3d” is worse than no reply if nobody can fetch that commit yet.
gh api graphql -f query='
query($owner:String!, $repo:String!, $pr:Int!) {
repository(owner:$owner, name:$repo) {
pullRequest(number:$pr) {
reviewThreads(first:100) {
nodes { id isResolved comments(last:1) { nodes { author { login } path line body } } }
}
}
}
}' -F owner=OWNER -F repo=REPO -F pr=57 \
--jq '.data.repository.pullRequest.reviewThreads.nodes[] | select(.isResolved | not)'
Replies go one per thread, then the thread is resolved with the resolveReviewThread mutation. Threads where the answer is a pushback stay open for the reviewer to close.
The hook that closes the loop
Skills only help when they get used, and the agent would happily push a fix and move on while three review threads sat unanswered. So a PostToolUse hook on Bash runs after every git push, confirms the push landed, and asks GitHub whether the pull request still has threads whose last comment is not mine:
// .claude/hooks/review-threads-after-push.mjs (the decision at its core)
const unanswered = (threads, viewer) =>
threads.filter((t) => {
if (t.isResolved) return false;
const last = t.comments.nodes.at(-1);
return last !== undefined && last.author?.login !== viewer;
});
const open = unanswered(threads, viewer);
if (open.length) {
console.log(
JSON.stringify({
hookSpecificOutput: {
hookEventName: 'PostToolUse',
additionalContext: `This push landed on PR #${pr}, which has ${open.length} review threads waiting. Use /answer-review now.`,
},
}),
);
}
The design note in that file is one I now apply everywhere: whether a push answers a review is a fact on GitHub, so the hook asks GitHub instead of guessing from the conversation.
One subagent, read-only on purpose
There is exactly one home-grown subagent, a code reviewer, and its most important property is the tool list:
---
name: code-reviewer
description: Reviews the working diff against the house checklist. Use before opening a PR.
tools: Read, Grep, Glob, Bash
---
Review `git diff main...HEAD`. Check: env access goes through the env module, new endpoints
are contract-first, colors use design tokens, dependencies come from the pnpm catalog,
generated files are untouched. Never edit files. Put the suggested fix inside the finding.
No Edit and no Write means a review can never turn into an unrequested refactor. The subagents documentation also supports a maxTurns field that stops a subagent and returns its output marked as partial. Four design-tool subagents installed from a third-party skill use it, and their prompts plan around it (“by roughly the tenth turn, stop reading and write”). That is a useful pattern for any subagent that tends to research forever.
Vendored skills need a lockfile and a guard
Vendored skills are pinned in a skills-lock.json with the source repository and a content hash:
{
"vitest": {
"source": "some-org/skills",
"sourceType": "github",
"skillPath": "skills/vitest/SKILL.md",
"computedHash": "2ec08c85..."
}
}
They are added one at a time with the skills CLI in copy mode, and never bulk-updated. That rule has a story. One run of the update command replaced the committed skill folders with symlinks into a gitignored cache and deleted 55 committed files in one command. A second incident: a vendored SKILL.md indexed rule files that had never shipped, and the lockfile could not catch it, because the hash only covers SKILL.md. A check-skills script in CI now guards both failure modes.
Skills with real side effects are the exception to model invocation. A deploy skill carries disable-model-invocation: true, which the skills documentation describes as preventing Claude from loading the skill automatically, “for workflows you want to trigger manually with /name.” Opening a PR stays model-invocable because the ask hook gates it. A deploy has no such gate, so it only runs when I type it.
Git hooks and CI are the backstop
The pre-commit hook (managed with lefthook) mirrors the agent guards for humans, and runs the hook self-tests whenever a hook changes:
# lefthook.yml
pre-commit:
parallel: false
commands:
branch-name:
run: pnpm branch-name:check
lint:
glob: '*.{ts,tsx,js,mjs}'
run: pnpm exec oxlint {staged_files}
secrets:
run: gitleaks protect --staged --no-banner
hook-self-tests:
glob: '.claude/hooks/**'
run: node .claude/hooks/skill-router.mjs --self-test && node .claude/hooks/review-threads-after-push.mjs --self-test
pre-push:
jobs: [] # the verify gate belongs to CI, not the laptop
The empty pre-push is intentional. It used to run the full verify suite, which took long enough that people started skipping it. CI now runs a secret scan, the checks, a build, a typecheck and sharded tests on every pull request, and five of those are required before merge. The root CLAUDE.md tells the agent never to run the full suite locally. It verifies only what it touched and leaves the rest to CI.
The last structural guard is in Turborepo’s boundaries config. Packages carry tags, and each tag denies dependencies on the layers above it, so an import cycle cannot even be expressed:
{
"boundaries": {
"tags": {
"shared": { "dependencies": { "deny": ["app"] } },
"contract": { "dependencies": { "deny": ["data", "ui"] } },
"data": { "dependencies": { "deny": ["ui"] } }
}
}
}
An agent that reaches for a quick cross-layer import gets a failing check instead of a code review argument.
What I would do differently
Three things, in order of how much they cost:
- Start with the byte budget, not a line count. The line ceiling let the root file reach 31.7KB before anyone noticed. A budget checked in CI from day one would have kept the root file lean without any cleanup.
- Generate the hand-synced lists. Adding a compiled package still means updating three lists by hand, and nothing fails if one is missed. The root
CLAUDE.mdadmits this in writing, which is honest but is not a guard. Anything that must stay in sync should be generated or checked. - Write the self-tests with the hook. The router’s fixtures came after it misfired. Near-miss fixtures (prompts that must not route) are the ones that catch real regressions.
The economics are the reason this is worth the effort. When an agent writes most of the code, the bottleneck moves to review, a cost covered in who pays for AI slop in code. Each guard in this harness removes a class of comment a reviewer would otherwise have to write. If you are weighing whether to build this much around an off-the-shelf agent, the trade-off is laid out in build your own agent harness or buy Claude Code, and more ready-made skills are catalogued in the Claude Skills directory. For comparisons with other agents, see AI coding agents compared.
Claude Code harness FAQ
- What is a Claude Code harness?
- It is the configuration around the agent that shapes and constrains what it does: CLAUDE.md instruction files, path-scoped rules, permission rules, hooks, skills, subagents, and the git hooks and CI checks that back them up. The model writes the code, and the harness decides what it is allowed to do and what gets checked.
- Should a rule go in CLAUDE.md or in a hook?
- If it must hold every time, put it in a hook, a permission rule or CI. Anthropic's memory documentation says CLAUDE.md is treated as context, not enforced configuration, and points to a PreToolUse hook for anything that must be blocked regardless of what Claude decides.
- How do you stop Claude Code from pushing or opening a PR without approval?
- Register a PreToolUse hook on Bash that matches git commit, git push, gh pr create and gh pr merge, and returns permissionDecision set to ask. That forces a prompt even when a broader allow rule would auto-approve git commands.
- Why does my Write permission rule not block file edits?
- Claude Code checks file permissions against Edit and Read path rules only. A path rule written for Write is accepted but never consulted, so deny edits with Edit(path), which covers every built-in tool that edits files.
- How can Claude Code post inline pull request review comments?
- gh pr review cannot attach comments to lines, so post one submission to the REST endpoint repos/OWNER/REPO/pulls/N/reviews with a comments array of path, line, side and body. Every line must be inside the diff, or the whole submission fails with a 422.
Sources
- Anthropic (2026). How Claude remembers your project. Claude Code documentation. https://code.claude.com/docs/en/memory
- Anthropic (2026). Hooks reference. Claude Code documentation. https://code.claude.com/docs/en/hooks
- Anthropic (2026). Configure permissions. Claude Code documentation. https://code.claude.com/docs/en/permissions
- Anthropic (2026). Extend Claude with skills. Claude Code documentation. https://code.claude.com/docs/en/skills
- Anthropic (2026). Create custom subagents. Claude Code documentation. https://code.claude.com/docs/en/sub-agents
- GitHub (2026). REST API endpoints for pull request reviews. GitHub Docs. https://docs.github.com/en/rest/pulls/reviews
- GitHub (2026). Reviewing proposed changes in a pull request. GitHub Docs. https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/reviewing-changes-in-pull-requests/reviewing-proposed-changes-in-a-pull-request
- Hook, skill and CI counts: the author’s own repository (first-party), read on 2026-09-25.