Skip to main content

What this is

Claude Code writes a full JSONL transcript of every session to ~/.claude/projects/ — every prompt, every tool call, every failure, every token spent. This tool turns that exhaust into an evidence-backed report: what you actually use the agent for, how often it fails and recovers, and what it costs. The pipeline is deterministic Python — no LLM calls, no network access — except for the optional intent-classification step, which asks a model to label each session’s use case.
This analyzes your own local transcripts and writes reports to your own machine. Nothing here is a GAIA product feature end users interact with — it’s a contributor/maintainer tool for understanding how Claude Code is actually being used on this codebase.

Prerequisites

If you can already run gaia commands in this repo, you’re set — the harvest scripts need nothing beyond the base install (no [dev], [eval], or other extras). If not, follow the Dev Guide first (uv venv && uv pip install -e .). The optional intent-classification step needs a way to run an LLM subagent (e.g. a Claude Code session), not an API key.

Quick start

Never redirect output into the repository. These tables are built from real transcripts — an untracked tables.md at the repo root is one git add -A away from being published. Everything below defaults to ~/.gaia/cache/factory/ (created automatically), which lives outside any repo so it can’t get committed by accident. As a second safety net, this repo’s .gitignore also blocks these four filenames at its own root, in case you forget $FACTORY and redirect bare.
On Windows, run these in Git Bash (installed alongside Git for Windows), not PowerShell. This output has em dashes and other non-ASCII characters; PowerShell’s console encoding silently corrupts them on redirect (>), while Git Bash writes clean UTF-8 like macOS/Linux.
1

Extract — deterministic, no LLM, no network

Roughly linear in transcript count. Writes traces.jsonl, intents.jsonl, and stats.json into $FACTORY.
2

Tables — absolute counts and share-of-total for every figure

With use-case labels, once you’ve done Classifying intent below:
3

Context — per-request prompt size and local KV-cache memory

Re-reads the raw transcripts, because scan/report aggregate per session and that hides how large any single request got.
With use-case labels:
4

Savings — what a proposed fix would actually save

In tokens and dollars.
No transcripts yet? scan fails loudly and tells you which directory it searched instead of silently reporting zero — that’s expected on a machine that has never run Claude Code.

Flags

Output files

Classifying intent

Optional — the reports above already work without this; it only adds a “by use case” breakdown. scan produces intents.jsonl but assigns no use-case label. There’s no script for this step — hand it to an AI assistant (e.g. ask Claude Code to follow the steps below), or do it yourself:
  1. Split the intents into batches of ~70.
  2. For each batch, have a subagent assign one use-case tag per session, returning strict JSON. Use the same taxonomy across every batch (e.g. pr_lifecycle, code_review, doc_audit, feature_impl, ci_debug, security_fix, research), extended only where the corpus demands it.
  3. Write $FACTORY/labels.txt as <8-char-session-prefix> <use-case>, one per line — it carries session-id prefixes, so it stays in the cache directory, never in a repo.
  4. Re-run report --labels "$FACTORY/labels.txt" (and context the same way).
Classify from the first user message, not the auto-generated session title — the title summarizes what happened, which leaks the outcome into the label.

Things that will mislead you if you skip them

  • Subagents are not sessions. A delegated Task/Agent run gets its own transcript under <session-uuid>/subagents/. Say explicitly which scope a number covers (main session only, or including subagents).
  • There is no ground truth for task success. Nothing in a transcript says whether the goal was actually met — never present a tool-failure rate as a task-failure rate.
  • Friction signals are proxies, not proof. Corrections and interrupts are evidence of friction, not confirmation of it.
  • Cost is API-equivalent, not money spent, if the sessions ran on a subscription rather than metered API usage.
  • Duration is unusable past the median. A session left open overnight reports the whole night as active time.
  • The corpus grows while you analyze it. The session doing the analysis is itself being recorded. Timestamp the snapshot you’re reporting on.

Privacy

Transcripts contain absolute paths, branch names, repository content, and anything pasted into a prompt. The pipeline writes only to ~/.gaia/cache/factory/. If a report needs to be shared, put it somewhere private and scrub paths first — nothing derived from a corpus belongs in a public repository or a GitHub comment.

Driving this from Claude Code

If you want Claude Code itself to run this pipeline, review the output, and produce a written report, use the analyzing-claude-sessions skill (.claude/skills/analyzing-claude-sessions/SKILL.md) — it carries the same commands above plus guidance on what to actually report and how to verify the analysis before publishing it.

Roadmap

The current pipeline covers extraction and reporting only. A broader design — turning harvested sessions into eval scenarios and distilled skills for local models — is proposed but not implemented in docs/plans/claude-session-harvest.md. Nothing on this page depends on that plan landing.