Skip to content

Anatomy of a run

A single run, traced end to end. This walkthrough follows the code-review runbook reviewing a test repository full of intentional bugs — but the shape is the same for every runbook: collect → one structured LLM call → publish. Use it to see exactly what Junior gathers, what the model receives, and what comes back out.

The full review output (all 38 findings, the formatted markdown) lives on the next page: Review output in detail.

Terminal window
OPENAI_API_KEY=sk-... junior run \
--runbook local_review \
--harness pydantic \
--prompt-file ./prompts/security.md \
--prompt-file ./prompts/logic.md \
--prompt-file ./prompts/design.md \
--source branch --target-branch main \
--project-dir ../junior-test-repo \
--publish > /tmp/junior_review_output.md

--publish makes local_review render the result as Markdown (Phase 5); the redirect just saves it. Drop --publish and you get the raw ReviewOutput JSON on stdout instead — the output contract every runbook follows.

SettingValue
Collectorlocal (no platform API)
Harnesspydantic (single structured call)
Publisherlocal (renders Markdown to stdout with --publish)
Modelopenai:gpt-5.4-mini
Promptssecurity.md, logic.md, design.md
  • main: one commit, just hello.py with a basic greet().
  • feature/auth-system: 2 commits adding an auth + database + API layer.

The diff under review — feature/auth-system vs main:

FileStatusLines
api.pyadded+109
auth.pyadded+103
database.pyadded+120
hello.pymodified+53

4 files, +385 lines.

junior.collect.localcollect.core.collect.collect_base(): runs git diff, parses changed files, reads commit messages, attaches any extra context. The output is a ReviewContext:

{
"target_branch": "main",
"commit_messages": [
"feat: add authentication and database layer\n...",
"feat: add API endpoints and integrate auth with greetings\n..."
],
"full_diff": "<11,801 chars — unified diff of 4 files>",
"changed_files": [
{"path": "api.py", "status": "added", "diff": "<...>", "content": "<109 lines>"},
{"path": "auth.py", "status": "added", "diff": "<...>", "content": "<103 lines>"},
{"path": "database.py", "status": "added", "diff": "<...>", "content": "<120 lines>"},
{"path": "hello.py", "status": "modified", "diff": "<...>", "content": "<58 lines>"}
],
"extra_context": {}
}

The local collector has no platform API, so MR title/description/labels and source_branch stay empty — git is the only source. Each ChangedFile carries both its diff and full content.

junior.runbooks.code_review.render.build_user_message() turns the ReviewContext into the markdown user message the model sees (here ~12.3k chars):

## Merge Request:
**Branches:** → main
### Commits (2)
- feat: add authentication and database layer
- feat: add API endpoints and integrate auth with greetings
### Changed Files
- `api.py` (added)
- `auth.py` (added)
- `database.py` (added)
- `hello.py` (modified)
### Diff
```diff
<full unified diff — 385 added lines across 4 files>
`--context` / `--context-file` entries (none here) would appear in an
"Additional Context" section near the top.
## Phase 3 — Build the system prompt
**`junior.runbooks.code_review.instructions.build_review_prompt()`** assembles
one system prompt by concatenating these parts, blank-line separated
(`merge_prompts` in `prompt_loader.py`):
1. the runbook's `SYSTEM_PROMPT` role, plus the user's prompt bodies
(`--prompt` / `--prompt-file`) — here `security.md`, `logic.md`, `design.md`;
2. the shared `BASE_RULES`.
The runbook does not inline `AGENT.md` / `AGENTS.md` / `CLAUDE.md` — a harness that
wants project memory reads it itself from its working directory (`claudecode` →
`CLAUDE.md`, `codex` → `AGENTS.md`).
There is no per-prompt fan-out: every body lands in the **same** system prompt,
and the pydantic harness makes **one** structured call.

[ SYSTEM_PROMPT + security + logic + design ] + BASE_RULES │ one system prompt ✕ one user message (Phase 2) │ ▼ single structured LLM call

## Phase 4 — The LLM review
**`junior.harnesses.pydantic`** makes one structured call with
`output_type=ReviewOutput`. The model returns `summary`, `recommendation`,
and the `comments` list directly. Junior then wraps that `ReviewOutput` into a
`ReviewResult` — composing it with the measured token usage and any errors:
```json
{
"output": {
"summary": "The code quality is poor overall, with multiple critical security flaws...",
"recommendation": "request_changes",
"comments": [/* 38 findings */]
},
"usage": {
"input_tokens": 28174,
"output_tokens": 7224,
"total_tokens": 35398
},
"errors": []
}
SeverityCount
🔴 Critical5
🟠 High20
🟡 Medium13
Total38

The full findings tables are on the next page.

With --publish, junior.publish.local.post_review()publish.core.formatter.format_summary() renders the ReviewResult to Markdown on stdout (here redirected to /tmp/junior_review_output.md). Without --publish, local_review prints the raw ReviewOutput (summary + recommendation + comments) as JSON instead.

Every run also writes a secret-free JSON trace to <project_dir>/.junior/output/{timestamp}.json (on by default; disable with --no-record).

See the rendered output on the next page.

git diff ──▶ collect.local ─────────▶ ReviewContext
render.build_user_message() ▼ ──▶ user message (~12KB markdown)
instructions.build_review_prompt() ▼ ──▶ system prompt (merged, ~4KB)
pydantic harness call ▼ ──▶ ReviewOutput ─▶ ReviewResult (+tokens)
(--publish) format_summary() ▼ ──▶ Markdown
local.post_review() ▼ ──▶ stdout
·
(no --publish) raw ReviewOutput JSON ─────▶ stdout / -o file
.junior/output/{ts}.json trace ─▶ always written

Every runbook — GitLab, GitHub, Bitbucket, or your own — swaps the collector and publisher but keeps this exact spine.