The reference for the HTTP API, the Python SDK and the open package.
This page documents the HTTP API at api.treecode.ai, the Python SDK (treecode-client) that wraps it, and the open package (treecode-calibrate) that computes the same report on a laptop without uploading anything. The machine-readable OpenAPI 3.1 document is at https://api.treecode.ai/openapi.json.
Getting started
Four steps take a new account from nothing to a report. Each is one request; the whole sequence is also one short script in the SDK (see the SDK).
1. Sign up. Name, work email, company, password; no invitation, no card. The account opens on the Free plan with the sample project already loaded, so a report is there to read before anything is uploaded.
curl -s https://api.treecode.ai/v1/auth/signup \
-H 'content-type: application/json' \
-d '{"name": "Ada Byron", "email": "ada@larkspur.example", "company": "Larkspur Learning", "password": "a-long-password-1"}'
The response sets the application’s session cookie and returns the user, the organization and a session token. Or sign up in the browser at app.treecode.ai/signup.
2. Create an API key (from the API keys page of the application, or with the session token from step 1). The key is shown once; store it.
curl -s https://app.treecode.ai/api/v1/product/keys \
-H "cookie: tc_session=$TC_SESSION" -H 'content-type: application/json' \
-d '{"name": "ci"}'
# -> {"key": {"id": "...", "name": "ci", "prefix": "tc_live_Oh8C", ...}, "token": "tc_live_...", "note": "the key is shown once; store it now"}
export TREECODE_API_KEY=tc_live_...
3. Create a grader and upload a table. A grader is what you are measuring (a reward model, a judge, a program check or a step grader) together with the group size it is used with. The table is CSV or JSON lines, one row per candidate.
curl -s https://api.treecode.ai/v1/graders \
-H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/json' \
-d '{"name": "Math judge", "group_size": 16, "label_source": "answer key"}'
# -> {"grader": {"id": "53ccd10b-...", ...}}
curl -s "https://api.treecode.ai/v1/graders/53ccd10b-.../samples?checkpoint=step%20100" \
-H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: text/csv' \
--data-binary @samples.csv
# -> {"sample_set": {"id": "014e31b0-...", "prompts": 2000, "candidates": 32000, "kind": "outcome", ...}, "warnings": []}
4. Calibrate, then read the report. The calibration is queued and runs on the platform’s workers; an outcome table of 2,000 prompts by 16 candidates takes a few seconds.
curl -s https://api.treecode.ai/v1/graders/53ccd10b-.../calibrate \
-H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/json' -d '{}'
# -> 202 {"calibration": {"id": "2a0bc1ee-...", "status": "queued", ...}}
curl -s https://api.treecode.ai/v1/calibrations/2a0bc1ee-... -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"calibration": {"status": "succeeded", "verdict_level": "usable", "headline": "Within a prompt, this grader ranks a right answer above a wrong one 84% of the time (margin 1.1σ; pooled 1.2σ)."}, "report": {...}}
The same report is on the grader’s page in the application, with the figures, the what-to-do list and the question “did this change a decision?”, whose answer is logged.
The sample table
CSV with a header row, JSON lines (one object per line) or a JSON array of objects. One row per candidate.
| column | type | required | meaning |
|---|---|---|---|
prompt_id |
string | yes | the prompt (task, question) the candidate answers |
candidate_id |
string | yes | unique within the prompt |
score |
number | yes | the grader’s score; higher is better |
label |
1/0, true/false, “correct”/“wrong” | yes | whether the candidate was right, from a source other than the grader |
tokens |
integer | no | output tokens of the candidate (for the cost page) |
step |
integer ≥ 1 | step graders | the step index inside the trajectory |
node_id |
string | step graders | the node (prefix) the step hangs from; siblings share it |
parent_correct |
true/false | step graders | whether the node’s prefix was right |
checkpoint |
string | runs | the training checkpoint the samples came from |
class |
string | no | a free tag, shown but not used |
Kinds. The kind is detected from the columns and the scores: outcome (whole answers scored by a reward model or a judge), step (step and node_id present), binary (the scores take at most two values that agree with the labels: a program check, nothing to calibrate, only the Rollout Budget applies).
How much. A table needs at least 30 prompts with a right and a wrong candidate somewhere; about 200 prompts at the group size in use give usable bands. The bands come from a bootstrap over prompts, so more prompts tighten them; more candidates per prompt sharpen the model-free curve. At most 50 MB per upload.
Where the labels come from. From a source other than the grader: a held-out answer key, a program check on a subset, Monte Carlo completion where a checker exists for the final answer, consensus of a stronger model on a few hundred prompts, or a human pass. State the source in the grader’s label_source; it is printed on the report, and the report’s label-noise row shows what 5% wrong labels would do to the numbers.
A small table, as CSV:
prompt_id,candidate_id,score,label,tokens
p001,p001-1,6.41,1,812
p001,p001-2,5.02,0,990
p001,p001-3,4.77,0,1043
p002,p002-1,5.90,1,655
Base URLs and authentication
| where | base | credential |
|---|---|---|
| the public API | https://api.treecode.ai/v1 |
an API key: Authorization: Bearer tc_live_... |
| through the application | https://app.treecode.ai/api/v1/product |
the application’s session cookie (tc_session), or the same API key |
The two mounts serve the same handlers; every route accepts both credentials. The public host serves only /v1/*, /health and /openapi.json; anything else on api.treecode.ai is 404 {"error": "not found"}. Every response is JSON; times are ISO 8601 in UTC; ids are UUIDs. Routes below are relative to the base URL.
API keys. POST /keys returns a key once: tc_live_ followed by 40 URL-safe characters. The service stores the key’s SHA-256 and its first 12 characters (shown in lists as prefix). A revoked or unknown key gets 401 {"error": "this API key is not valid: it is unknown or has been revoked"}. A key acts for its organization; requests made with a key have no user.
Sessions. The web application signs in with the session cookie, which acts for the user’s first organization.
Routes that need no credential: POST /auth/signup, GET /share/{token}, GET /openapi.json, GET /health.
Routes: accounts and keys
POST /auth/signup
No credential. Creates the user, the organization (its slug from the company name, made unique), the owner membership and a session; the tc_session cookie is set on the response. Passwords need at least 10 characters.
curl -s https://api.treecode.ai/v1/auth/signup -H 'content-type: application/json' \
-d '{"name": "Ada Byron", "email": "ada@larkspur.example", "company": "Larkspur Learning", "password": "a-long-password-1"}'
201:
{"user": {"id": "90df40a0-…", "email": "ada@larkspur.example", "name": "Ada Byron", "role": "member"},
"org": {"id": "21f53e61-…", "name": "Larkspur Learning", "slug": "larkspur-learning", "plan": "free"},
"token": "so-XCVStu8L76qndH8Q2JIBXdhSd_Fwtj7yfp_w6DLg"}
token is the session token (the same value as the cookie), for clients that cannot keep cookies. Errors: 409 the email already has an account; 422 an invalid body; 429 more than 10 sign-ups per hour from one address. When sign-up is set to a waiting list the route answers 202 {"waitlist": true, "message": "..."}; when it is closed, 403.
GET /me
The caller: the user (null with an API key), the organization, the plan’s allowances and this month’s usage.
curl -s https://api.treecode.ai/v1/me -H "authorization: Bearer $TREECODE_API_KEY"
{"user": null,
"org": {"id": "21f53e61-…", "name": "Larkspur Learning", "slug": "larkspur-learning", "plan": "free"},
"plan": {"plan": "free", "label": "Free", "graders": 1, "prompts_per_month": 1000, "calibrations_per_month": 4, "runs": false, "keys": 2, "price": 0, "overage_per_1000_prompts": null},
"usage": {"period": {"start": "2026-10-01T00:00:00+00:00", "end": "2026-11-01T00:00:00+00:00"},
"prompts_ingested": 0, "calibrations": 0, "checkpoints": 0, "graders": 0, "keys": 1, "runs": 0},
"auth": "api_key"}
POST /keys
{"name": "ci"} (optional; default default). 201 with the key object and, once, the token; 402 when the plan’s key allowance is used up (revoked keys do not count).
curl -s https://api.treecode.ai/v1/keys -H "authorization: Bearer $TREECODE_API_KEY" \
-H 'content-type: application/json' -d '{"name": "ci"}'
{"key": {"id": "fc53ca96-…", "org_id": "21f53e61-…", "name": "ci", "prefix": "tc_live_Oh8C", "created_by": null,
"created_at": "2026-10-08T21:45:19.601258+00:00", "last_used_at": null, "revoked_at": null},
"token": "tc_live_Oh8CtU_gHtRphBBNzSv0MPNIfvbpLSbYVaGdDtZJ",
"note": "the key is shown once; store it now"}
GET /keys
{"keys": [key, ...]}, newest first, revoked ones included (revoked_at set). last_used_at is refreshed at most once a minute.
curl -s https://api.treecode.ai/v1/keys -H "authorization: Bearer $TREECODE_API_KEY"
DELETE /keys/{key_id}
Revokes the key (a second call is harmless). 200 {"key": {..., "revoked_at": "..."}, "revoked": true}; 404 for an unknown key.
curl -s -X DELETE https://api.treecode.ai/v1/keys/fc53ca96-... -H "authorization: Bearer $TREECODE_API_KEY"
Routes: graders and samples
A grader is what you are measuring: a reward model, a judge, a program check, or a step grader, together with the group size it is used with. Demo graders belong to the sample organization, appear in every organization’s list with "demo": true, can be read by everyone and changed by no one (writes answer 403).
The grader object:
{"id": "53ccd10b-…", "org_id": "21f53e61-…", "name": "Math judge", "kind": "outcome", "group_size": 8, "depth": null,
"label_source": "answer key", "policy": "demo-8b", "checkpoint_label": null, "notes": "a judge on word problems", "settings": {},
"demo": false, "share_token": null, "created_by": "90df40a0-…", "created_at": "2026-10-08T21:45:19.926447+00:00", "deleted_at": null,
"sample_sets_count": 1,
"latest_calibration": {"id": "2a0bc1ee-…", "status": "succeeded", "verdict_level": "usable",
"headline": "Within a prompt, this grader ranks a right answer above a wrong one 84% of the time (margin 1.1σ; pooled 1.2σ).",
"kind": "outcome", "finished_at": "2026-10-08T21:45:23.134136+00:00"},
"own": true, "share_url": null}
| field | meaning |
|---|---|
kind |
auto, outcome, step or binary; auto becomes the detected kind at the first upload, and later uploads must match |
group_size |
the group size (training) or samples per query (serving) the grader is used with; the default of calibrations |
depth |
step graders: the number of steps of the search (default of calibrations) |
label_source |
where the labels come from (an answer key, a unit test, human review); printed on the report |
policy |
the model the samples came from (shown on the report, not used in the computation) |
checkpoint_label |
a label for the samples’ checkpoint when the table has no checkpoint column |
notes, settings |
free text and a free JSON object |
latest_calibration |
the most recent calibration, whatever its status, or null |
own |
the grader belongs to the calling organization |
share_url |
the public link when sharing is on (https://app.treecode.ai/share/<token>) |
POST /graders
Only name is required (kind defaults to auto, group_size to 8, label_source to unknown). 201 {"grader": {...}}; 402 on the grader allowance.
curl -s https://api.treecode.ai/v1/graders -H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/json' \
-d '{"name": "Math judge", "kind": "auto", "group_size": 8, "label_source": "answer key", "policy": "demo-8b", "notes": ""}'
GET /graders
The organization’s graders, newest first, then the demo graders.
curl -s https://api.treecode.ai/v1/graders -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"graders": [{"...": "own graders"}, {"...": "then the demo graders, \"demo\": true"}], "counts": {"own": 1, "demo": 2, "allowance": 1}}
GET /graders/{grader_id}
The grader, its sample sets and calibrations (newest first; calibrations as summaries without the report), the full report of the latest succeeded calibration, and its training runs.
curl -s https://api.treecode.ai/v1/graders/53ccd10b-... -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"grader": {...}, "sample_sets": [...], "calibrations": [...], "latest_report": {...} | null, "runs": [...]}
PATCH /graders/{grader_id}
Any of name, group_size, depth, label_source, policy, checkpoint_label, notes, settings; only the fields given change. 200 {"grader": {...}}.
curl -s -X PATCH https://api.treecode.ai/v1/graders/53ccd10b-... -H "authorization: Bearer $TREECODE_API_KEY" \
-H 'content-type: application/json' -d '{"group_size": 16, "notes": "the judge prompt of October 8"}'
DELETE /graders/{grader_id}
Soft-deletes the grader (it leaves every list and answers 404 afterwards; its calibrations and decisions stay in the database) and removes its sample files from storage.
curl -s -X DELETE https://api.treecode.ai/v1/graders/53ccd10b-... -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"deleted": true, "grader_id": "53ccd10b-...", "sample_files_removed": 2}
POST /graders/{grader_id}/share and DELETE /graders/{grader_id}/share
Turns public sharing of the latest report on (the token is stable across calls) or off. The shared page shows the report, the grader’s name and settings and the organization’s name; it never shows the sample rows.
curl -s -X POST https://api.treecode.ai/v1/graders/53ccd10b-.../share -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"share_token": "1_q-rRwZ-UmIw0ljzb44p0NvQqtbbzQf", "url": "https://app.treecode.ai/share/1_q-rRwZ-UmIw0ljzb44p0NvQqtbbzQf"}
curl -s -X DELETE https://api.treecode.ai/v1/graders/53ccd10b-.../share -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"share_token": null, "url": null}
POST /graders/{grader_id}/samples?checkpoint=<label>&format=<csv|jsonl|json>
The request body is the table itself: text/csv, application/x-ndjson (JSON lines) or application/json (an array of row objects, or {"rows": [...]}). The format is detected from the content; ?format= overrides. Multipart forms are not accepted (415): send the file’s bytes with its content type.
?checkpoint= labels the whole table. Without it, a table whose checkpoint column holds one value takes that value; a table with several checkpoint values is refused (422), since a sample set is one checkpoint.
curl -s "https://api.treecode.ai/v1/graders/53ccd10b-.../samples?checkpoint=step%20100" \
-H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/x-ndjson' --data-binary @samples.jsonl
The table is validated by the engine before anything is stored. Problems answer 422:
{"error": "the sample table has problems",
"problems": ["missing required column(s): label", "1 row(s) with a label that is not 1/0, true/false or correct/wrong (rows 1); each was dropped"]}
A table whose kind differs from the grader’s (outcome rows on a step grader) is also 422, with a sentence saying which columns made the difference. Then the prompt allowance is checked (402), the number of distinct prompts is metered, and the normalized rows are stored. 201:
{"sample_set": {"id": "014e31b0-…", "org_id": "21f53e61-…", "grader_id": "53ccd10b-…", "training_run_id": null, "checkpoint": "step 100",
"prompts": 40, "candidates": 320, "correct": 152, "wrong": 168, "kind": "outcome", "sha256": "6b7c9c91…",
"summary": {"prompts": 40, "candidates": 320, "correct": 152, "wrong": 168, "kind": "outcome", "has_tokens": true,
"checkpoints": ["step 100"], "group_sizes": {"min": 8, "median": 8, "max": 8}},
"created_by": "90df40a0-…", "created_at": "2026-10-08T21:45:20.621559+00:00"},
"warnings": []}
warnings are the engine’s non-blocking notes (for example a step table without parent_correct).
Routes: calibrations and decisions
POST /graders/{grader_id}/calibrate
Queues the Grader Check. All fields are optional: sample_set_id defaults to the grader’s latest sample set (400 when there is none), group_size and depth to the grader’s, target_success to 0.9, n_boot (bootstrap resamples, 2 to 2000) to 200, or 500 when the sample set has fewer than 5,000 candidates. 402 on the calibration allowance (metered when queued).
curl -s https://api.treecode.ai/v1/graders/53ccd10b-.../calibrate -H "authorization: Bearer $TREECODE_API_KEY" \
-H 'content-type: application/json' -d '{"sample_set_id": null, "group_size": null, "target_success": 0.9, "n_boot": null, "seed": 0}'
202:
{"calibration": {"id": "2a0bc1ee-…", "org_id": "21f53e61-…", "grader_id": "53ccd10b-…", "sample_set_id": "014e31b0-…",
"params": {"sample_set_id": "014e31b0-…", "group_size": 8, "depth": null, "target_success": 0.9, "n_boot": 500, "seed": 0},
"status": "queued", "run_id": "caa7aaca-…", "error": null, "created_by": "90df40a0-…",
"created_at": "2026-10-08T21:45:21.265451+00:00", "started_at": null, "finished_at": null,
"verdict_level": null, "headline": null, "kind": null, "seconds": null}}
Poll GET /calibrations/{id} until status is succeeded or failed. An outcome table of 2,000 prompts by 16 candidates takes a few seconds; a step grader’s report about two minutes. status moves queued → running → succeeded | failed (error carries the reason).
GET /calibrations/{calibration_id}
The calibration summary and, once it has succeeded, the Report.
curl -s https://api.treecode.ai/v1/calibrations/2a0bc1ee-... -H "authorization: Bearer $TREECODE_API_KEY"
{"calibration": {"...": "the summary above, with", "status": "succeeded", "started_at": "...", "finished_at": "...",
"verdict_level": "usable", "headline": "Within a prompt, this grader ranks a right answer above a wrong one 84% of the time (margin 1.1σ; pooled 1.2σ).",
"kind": "outcome", "seconds": 0.183},
"report": {"...": "the Report, or null until it succeeds"}}
GET /calibrations/{calibration_id}/events?after=<id>&limit=<n>
The worker’s progress and log lines, in order (after: only events with a larger id; limit at most 2000). Levels: info, warning, error, progress (with progress in [0, 1]).
curl -s "https://api.treecode.ai/v1/calibrations/2a0bc1ee-.../events?after=0&limit=200" -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"calibration_id": "...", "status": "succeeded", "events": [{"id": 15, "level": "progress", "message": "loaded 320 candidates over 40 prompts (outcome grader) in 0.0s", "progress": 0.03, ...}, ...]}
POST /calibrations/{calibration_id}/decision
What was decided after reading the report. answer is one of changed_group_size, changed_mix, changed_n, stopped_run, changed_grader, no_change, not_yet; note is free text. A decision on a demo grader’s calibration is recorded in the caller’s own log.
curl -s https://api.treecode.ai/v1/calibrations/2a0bc1ee-.../decision -H "authorization: Bearer $TREECODE_API_KEY" \
-H 'content-type: application/json' -d '{"answer": "changed_group_size", "note": "16 -> 8 on the hard class"}'
# -> 201 {"decision": {"id": "31854194-…", "calibration_id": "2a0bc1ee-…", "answer": "changed_group_size", "note": "16 -> 8 on the hard class", "ts": "...", "grader_name": "Math judge", ...}}
GET /decisions
The organization’s decision log, newest first (at most 200), with the list of allowed answers.
curl -s https://api.treecode.ai/v1/decisions -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"decisions": [decision, ...], "answers": ["changed_group_size", "changed_mix", "changed_n", "stopped_run", "changed_grader", "no_change", "not_yet"]}
Routes: budgets
Both budgets are computed on demand from a finished calibration’s stored Report; nothing is metered or stored. They need status: succeeded (409 otherwise).
POST /budgets/rollouts
The Rollout Budget. group_size_now defaults to the report’s group size; floors is an even group size of at least 4 for every class or {"<20%": 6, ...} per class; cap defaults to 64; cost_per_rollout (dollars) adds the dollar lines. 409 for a step grader’s report (no coverage per prompt).
curl -s https://api.treecode.ai/v1/budgets/rollouts -H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/json' \
-d '{"calibration_id": "2a0bc1ee-...", "group_size_now": 8, "floors": 4, "cap": 64, "cost_per_rollout": 0.004}'
The summary part of the answer:
{"group_size_now": 8, "floors": {"<20%": 4, "20–80%": 4, ">80%": 4}, "cap": 64, "cost_per_rollout": 0.004,
"now": {"rollouts_per_step": 320, "signal_concave": 67.896755, "signal_linear": 192.041023, "waste_fraction": 0.274571, "wasted_groups_per_step": 10.98, "dollars_per_step": 1.28},
"saving_range": {"lo": 0.01875, "hi": 0.1, "sentence": "Rollouts not drawn at equal signal: 2% (concave) to 10% (linear, a bound)."},
"weights": {"<20%": 0.509554, "20–80%": 1.783439, ">80%": 0.764331},
"sentences": ["Today's mix draws 320 rollouts per step at group size 8 over 40 prompts; 27% of the groups are all right or all wrong and carry no learning signal.", "..."]}
The whole object: classes (per coverage class: prompts, coverage, the wasted-group share today, the signal per rollout, the group size under the concave rule, the linear bound and the max-signal allocation), now, allocations (equal_signal_concave, equal_signal_linear_bound, max_signal_at_budget, each with group sizes per class, rollouts per step, the saving or the gain, and the dollar lines when a cost was given), frontier (signal against rollouts at budget fractions 0.5 to 1.2), saving_range, weights and sentences. weights are the factors a dynamic-sampling loop multiplies each class’s sampling rate by (prompt-weighted mean 1).
POST /budgets/serving
The Cost per Solved Task. tokens_per_candidate and price_per_mtok (the output price in dollars per million tokens) are required; grader_cost_per_candidate, target_success (0.9), failure_cost, prompt_tokens and price_in_per_mtok are optional. 409 for a step or binary grader’s report (no predicted best-of-N curve).
curl -s https://api.treecode.ai/v1/budgets/serving -H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/json' \
-d '{"calibration_id": "2a0bc1ee-...", "tokens_per_candidate": 900, "price_per_mtok": 0.60, "grader_cost_per_candidate": 0.0002, "target_success": 0.9, "failure_cost": 2.0}'
The answer carries cost_per_candidate, the arrays n, success, cost_per_task and cost_per_solved (aligned by index), n_for_target, cost_at_target, cost_per_solved_at_target and success_band_at_target (null when no N up to 64 reaches the target), n_min_total_cost and expected_total_cost (with a failure cost), vendors (the same numbers at each listed vendor’s dated output price) and sentences, such as “Each candidate costs $0.00074: 900 tokens at $0.60 per million plus $0.00020 for the grader.” and “The smallest N whose predicted success reaches 50% is 2: $0.0015 per task, $0.0027 per solved task (success band 45% to 62% at that N).”
GET /vendors/prices
The listed, dated per-token prices of the serving vendors used by the cost page: a price table, not a measurement of the vendors.
curl -s https://api.treecode.ai/v1/vendors/prices -H "authorization: Bearer $TREECODE_API_KEY"
# -> {"vendors": [{"name": "Together", "model": "gpt-oss-120b", "price_in_per_mtok": 0.15, "price_out_per_mtok": 0.6, "date": "2026-10-06", "source": "..."}, "..."], "price_date": "2026-10-06", "unit": "dollars per million tokens"}
Routes: training runs and checkpoints
Team plan and above. A training run is a sequence of checkpoints of one grader, each a sample set labeled with the checkpoint; the checkpoint series (coverage, signal per rollout, margin with its band, mean score, the alarm) is computed from them. The series is labeled experimental: below eight checkpoints nothing is extrapolated, and beyond that a fitted trend is shown with the words “a trend, not a forecast”.
POST /runs
402 (allowance.kind: "runs") on the Free and Pro plans.
curl -s https://api.treecode.ai/v1/runs -H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/json' \
-d '{"grader_id": "53ccd10b-...", "name": "GRPO on word problems", "group_size": 8}'
# -> 201 {"run": {"id": "9d0eca93-…", "grader_id": "53ccd10b-…", "name": "GRPO on word problems", "group_size": 8, "has_series": false, "checkpoints_count": 0, "grader_name": "Math judge", "demo": false, ...}}
POST /runs/{run_id}/checkpoints?checkpoint=<label>
The same upload as POST /graders/{id}/samples (body, formats, validation, the prompt allowance), with the checkpoint label required (400 without it, unless the table carries exactly one checkpoint value). Upload checkpoints in training order: the series follows the upload order. The series is recomputed on each upload.
curl -s "https://api.treecode.ai/v1/runs/9d0eca93-.../checkpoints?checkpoint=step%2050" \
-H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: text/csv' --data-binary @step_050.csv
# -> 201 {"sample_set": {...}, "warnings": [], "run": {...}, "series": {...}}
GET /runs/{run_id}
The run, its checkpoints in upload order and the series.
curl -s https://api.treecode.ai/v1/runs/9d0eca93-... -H "authorization: Bearer $TREECODE_API_KEY"
{"run": {"...": "as above, with", "has_series": true, "checkpoints_count": 2},
"checkpoints": [{"...": "sample sets in upload order, each with its checkpoint label"}],
"series": {"version": 1, "group_size": 8, "n_boot": 200, "seed": 0, "experimental": true,
"checkpoints": ["step 0", "step 50"],
"series": [{"checkpoint": "step 0", "prompts": 40, "candidates": 320, "coverage_mean": 0.475, "signal_per_rollout": 0.553648, "waste_fraction": 0.320368,
"margin_within": 1.14607, "band": [0.919379, 1.395098], "accuracy_within": 0.844202, "mean_score": 5.7208}, "..."],
"classes": [{"name": "<20%", "lo": 0.0, "hi": 0.2, "prompts_at_first": 15, "coverage": [0.11, 0.12], "prompts": [15, 15]}, "..."],
"alarm": {"fired": false, "at": [], "at_index": [], "rule": "margin_within[t] < band_lo[t-1] and margin_within[t-1] < band_lo[t-2] while mean_score rose over both steps"},
"trend": null, "condition": "progressing",
"sentences": ["2 checkpoint(s) at group size 8; mean coverage went from 48% to 51% ...", "..."]}}
GET /runs
The organization’s runs, newest first, then the demo runs ("demo": true).
curl -s https://api.treecode.ai/v1/runs -H "authorization: Bearer $TREECODE_API_KEY"
Routes: usage, billing, sharing and meta
GET /usage
This month’s usage against the plan’s allowances, and the statement.
curl -s https://api.treecode.ai/v1/usage -H "authorization: Bearer $TREECODE_API_KEY"
{"plan": "free", "period": {"start": "2026-10-01T00:00:00+00:00", "end": "2026-11-01T00:00:00+00:00"},
"prompts_ingested": 40, "calibrations": 1, "checkpoints": 0, "graders": 1, "keys": 1, "runs": 0,
"allowances": {"plan": "free", "label": "Free", "graders": 1, "prompts_per_month": 1000, "calibrations_per_month": 4, "runs": false, "keys": 2, "price": 0, "overage_per_1000_prompts": null},
"statement": {"plan_price": 0, "overage_prompts": 0, "overage_price": 0.0, "total": 0.0, "currency": "USD", "note": null}}
On the Scale plan overage_prompts is the number above 500,000 and overage_price is $4 per started 1,000; plan_price and total are null on custom plans.
POST /billing/plan
{"plan": "team"} with one of free, pro, team, scale (400 for anything else; enterprise is set by Treecode). The prototype collects no payment: the plan switches at once.
curl -s https://api.treecode.ai/v1/billing/plan -H "authorization: Bearer $TREECODE_API_KEY" -H 'content-type: application/json' -d '{"plan": "team"}'
# -> 200 {"org": {..., "plan": "team"}, "plan": {allowances}, "note": "no payment is collected by this prototype"}
GET /share/{token}
No credential. The latest report of a grader whose sharing is on, or a demo grader: demo-reward-model-math (the reward model on math word problems, simulated run) and demo-step-grader-l (the step grader on program search, Milestone 1, real data). 404 for an unknown token or when sharing was turned off.
curl -s https://api.treecode.ai/v1/share/demo-reward-model-math
# -> {"grader": {"name": "...", "kind": "outcome", "group_size": 16, "demo": true, ...}, "report": {...}, "org_name": "Sample project", "calibration": {...}}
GET /health and GET /openapi.json
No credential. GET /health answers {"status": "ok", "database": "ok", "dispatcher": "ecs", "code_version": "…", "time": "…"} (503 when the database is unreachable). GET /openapi.json is the OpenAPI 3.1 document, which carries the plan table under x-plans.
curl -s https://api.treecode.ai/health
curl -s https://api.treecode.ai/openapi.json | python3 -m json.tool | head
The Report JSON
The Report (version 1) is the engine’s output, stored as is. Top-level keys: version, kind, input, verdict, margins, ranking, distributions, coverage, best_of_n, label_noise, step, method. Keys are snake_case, arrays are aligned by index, and null stands where a number does not exist (never NaN). Class names are "<5%", "5–20%", "20–50%", "50–80%", ">80%" (an en dash) and merged forms such as "<20%" when a class has fewer than 10 prompts.
| key | what it holds | where the application shows it |
|---|---|---|
input |
prompts, candidates, correct, wrong, group_size, depth, group_sizes: {min, median, max}, has_tokens, label_source, policy, checkpoint |
the input stamp at the bottom of the report |
verdict |
level (strong, usable, weak for outcome graders; above, near, below for step graders; binary), headline, sentences, what_to_do (at most 3), conventions |
the badge, the first sentence, the what-to-do list |
margins |
pooled and within_prompt: delta, sigma_w, sigma_c, margin, band; by_class: the margin per coverage class with its band and counts |
the margin tiles and the margin-by-class chart |
ranking |
accuracy_pooled, accuracy_within, band_within: P(a right answer scores above a wrong one), ties one half |
the ranking-accuracy tile |
distributions |
edges (41), correct (40 counts), wrong (40 counts), score_min, score_max, mean_correct, mean_wrong |
the two score histograms |
coverage |
the fitted beta (a, b), mean, a histogram, and classes (name, bounds, prompts, coverage) |
the coverage histogram; the Rollout Budget reads it |
best_of_n |
n, predicted, band_lo, band_hi (1 to 64); empirical_k, empirical, empirical_prompts, empirical_se (the model-free curve from the samples); n_support and n_support_by_class; n_star_reference; target_success, n_for_target; agreement_within_band_to_k; success_at_group_size |
the success-against-N chart with its band, the dots, the support shading and the target line; the Cost per Solved Task reads it |
label_noise |
epsilon, and the margin_within, accuracy_within and success_at_group_size after flipping that share of the labels that disagree most with the grader; verdict_changes |
the label-noise sentence |
step |
step graders only: b, depth, the gap and spreads, margin with margin_band, the threshold v0 and v0_over_sigma with the Gaussian value beside it, rho_star, multiplicity, the correlation of wrong-step scores under one node, task_spread, widths with predicted and bands, by_level, largest_b_clearing, verdict |
the threshold verdict and the success-against-width chart |
method |
n_boot, seed, engine_version, notes (one line per computation), seconds |
the “how this was computed” notes |
For kind == "step", coverage, best_of_n and label_noise are null, the margin bands are null and by_class is empty. For kind == "binary" only input and coverage are filled: a program check’s score is its label, so there is nothing to calibrate and the Rollout Budget applies. The engine’s README (calibrate/README.md, “Output schemas”) is the authority on every key.
The SDK
treecode-client wraps the routes above (pip install treecode-client; from the checkout for now: pip install ./sdk/python). Treecode(api_key) tries https://api.treecode.ai/v1 and falls back to https://app.treecode.ai/api/v1/product with the same key, so a script runs either way; base_url overrides both.
from treecode_client import Treecode
tc = Treecode(api_key="tc_live_...") # or set TREECODE_API_KEY
g = tc.graders.create("Math judge", scores="answers", group_size=16, label_source="answer key")
g.upload("samples.jsonl")
report = g.calibrate() # queues the calibration and waits for it
report.verdict # "usable"
report.success_at(16) # the predicted success of best-of-16
tc.budgets.rollouts(report, cost_per_rollout=0.004, floors={"<5%": 8})
tc.budgets.serving(report, tokens_per_candidate=900, price_per_mtok=0.60,
grader_cost_per_candidate=0.0002, target_success=0.9)
run = tc.runs.create(g, "GRPO run 12", group_size=16) # checkpoints of a training run (Team plan and above)
run.add_checkpoint("samples_step300.jsonl", checkpoint="step 300")
run.alarm # {"fired": ..., "at": [...], "rule": ...}
scores="answers" says the grader scores whole answers (a reward model or a judge); label_source is where the right/wrong labels come from, printed on the report. A refusal raises a TreecodeError with the service’s sentence: AllowanceError (402) when the plan’s allowance is used up, ValidationError (422) with one sentence per problem in an uploaded table, AuthError (401) for a missing or revoked key. The SDK’s README in the repository (sdk/python/README.md) lists every method; the client example examples/larkspur/ runs a fictional team’s scenario end to end and writes its figures and a case study.
The laptop package
The engine behind every page is the open package treecode-calibrate. It runs on a laptop against a CSV or JSON lines file, prints the same report, and sends nothing anywhere, which is the answer for proprietary data.
pip install treecode-calibrate # from the checkout for now: pip install ./calibrate
treecode-calibrate samples.csv --group-size 16
treecode-calibrate samples.csv --group-size 16 --json report.json --figures figs/ \
--budget --cost-per-rollout 0.004 --serving --tokens 900 --price 0.60
Options: --depth D (step graders), --target 0.9, --n-boot 500, --seed 0, --json report.json (the Report, with the budget and the serving result when asked for), --figures DIR (SVG figures), --budget --cost-per-rollout 0.004 --floors 4 --cap 64, --serving --tokens 900 --price 0.60 --grader-cost 0.0002 --failure-cost 2.0, --checkpoints (treat the checkpoint column as a training run and print the series), --label-source, --policy, --checkpoint (stamps printed on the report).
The same functions from Python:
from treecode_calibrate import read_table, calibrate, rollout_budget, serving_budget, checkpoint_series
table = read_table("samples.csv")
report = calibrate(table, group_size=16, target_success=0.9, n_boot=500, seed=0)
budget = rollout_budget(report, group_size_now=16, floors=4, cost_per_rollout=0.004)
serving = serving_budget(report, tokens_per_candidate=900, price_per_mtok=0.60, grader_cost_per_candidate=0.0002, target_success=0.9)
Every function returns the same JSON as the API. The package’s tests compare the Gaussian case with its exact answers, the best-of-N simulation with the closed form, the model-free curve with brute-force enumeration, the signal formula with direct computation of the advantages, and the waste table with the formula.
Allowances and errors
| plan | graders | prompts / month | calibrations / month | training runs | API keys | price |
|---|---|---|---|---|---|---|
| Free | 1 | 1,000 | 4 | no | 2 | $0 |
| Pro | 3 | 10,000 | unlimited | no | 5 | $99 / month |
| Team | 10 | 100,000 | unlimited | yes | 20 | $990 / month |
| Scale | unlimited | 500,000, then $4 per 1,000 | unlimited | yes | unlimited | $2,900 / month |
| Enterprise | unlimited | unlimited | unlimited | yes | unlimited | custom |
Prompts are counted when a table is stored (the number of distinct prompts in the upload, not the number of candidate rows), calibrations when they are queued, checkpoints when a run’s checkpoint is uploaded, all within the current calendar month in UTC. Graders (not deleted) and keys (not revoked) are counted as they stand. On the Scale plan prompts beyond 500,000 are never refused; they appear as overage on the statement. See the pricing page for what each plan is for.
Every error is {"error": "<a sentence>"} plus, where it helps, more fields:
| status | when | extra fields |
|---|---|---|
| 400 | a malformed request (an empty upload, a missing ?checkpoint=, an unknown format or plan) |
|
| 401 | no valid credential | |
| 402 | an allowance of the plan is exceeded | allowance, upgrade: "/billing" |
| 403 | a write to a demo grader or demo run; sign-up closed | |
| 404 | no such grader, calibration, run, key, share link (also for ids that are not UUIDs, and for other organizations’ objects) | |
| 409 | the state does not allow it: a calibration not finished or failed; a budget on a report without the needed curve | |
| 413 | an upload above 50 MB | |
| 415 | a multipart upload (send the file body itself) | |
| 422 | an invalid JSON body (details: the field errors) or an unusable sample table (problems: one sentence per problem) |
details or problems |
| 429 | more than 10 sign-ups per hour from one address |
A 402 looks like this:
{"error": "the Free plan allows 1 grader; 1 used and 1 requested",
"allowance": {"plan": "free", "kind": "graders", "limit": 1, "used": 1, "requested": 1, "remaining": 0,
"message": "the Free plan allows 1 grader; 1 used and 1 requested"},
"upgrade": "/billing"}
allowance.kind is one of graders, prompts_per_month, calibrations_per_month, keys, runs.
Data handling
Uploaded samples are stored encrypted in the platform’s private bucket, used only to compute your reports, never used to train anything, deletable by you from the grader page (or DELETE /graders/{id}), and gone from the platform’s backups within thirty days of that deletion. A share link exposes the report, never the sample rows. The laptop package computes everything locally and sends nothing. The full statement, with what is kept and for how long, is on the data handling page.
How this was computed
Every report lists its computations in one line each under “How this was computed” (the Report’s method.notes). This section gives the formulas behind them. Notation: for candidate of prompt , write for the negated score, so lower is better and the grader picks the lowest; “right” and “wrong” are your labels.
Margins and ranking accuracy. Pooled over all candidates: the gap between the average wrong and the average right candidate, the spreads and , and the margin . Best-of- and group-based training compare candidates of one prompt, so the report also gives the margin within prompts: a fit with a level for each prompt gives
over the prompts that hold both a right and a wrong candidate, divided by the spread of the wrong candidates around their own prompt’s wrong average. The ranking accuracy is the chance that a right candidate scores above a wrong one (the AUC, ties counted one half), pooled and within prompts.
Coverage. Coverage is the chance that one sampled answer to a prompt is right. A beta-binomial fit to the counts of right answers per prompt gives the coverage distribution and, for each prompt, a coverage estimate that is never exactly 0 or 1 (a prompt with 0 right of 16 is hard, not impossible). Prompts are grouped into classes at 5%, 20%, 50% and 80%.
The predicted best-of- curve. For each the engine simulates prompts: a coverage is drawn from the fitted distribution, of the candidates are right, every candidate gets a score from the measured distributions of its coverage class, and the prompt is solved when the best-scored candidate is right. The scores are simulated in units of each prompt’s own spread: a prompt’s level and its spread both cancel when its own candidates are compared, and mixing prompts whose spreads differ would make the grader look worse than it is inside any one prompt. Each prompt’s spread is its own within-prompt variance, shrunk toward its class’s median with four prior degrees of freedom, so a prompt with few candidates borrows its class’s typical spread.
The model-free curve. The same table gives the best-of- curve with no model at all, for every up to the number of candidates per prompt: with a prompt’s candidates sorted best-first, the best of a random of them is the top-ranked member of the subset, so
averaged over prompts (ties averaged over random orders). The report draws it over the prediction: where the two agree, the prediction has passed a check you can read; beyond it is trusted only as far as the data supports it. The prediction at relies on the best-scored of the wrong candidates, so with wrong candidates in a class it is supported up to ; the report shades the curve beyond the smallest such among the classes that still contribute failures. The value (where one right answer among would start to lose) is drawn as a reference line and never used as a decision: with several right answers among the curve is usually still rising there.
Bands, label noise and the verdict. Bands are the 5th and 95th percentiles of a bootstrap that resamples whole prompts and repeats every fit above in each resample. A label-noise row flips a stated share of the labels where they disagree most with the grader and shows how much the numbers move. The verdict strong needs a within-prompt ranking accuracy of at least 0.90 and the model-free curve inside the band at every up to the group size; usable needs at least 0.75 or disagreement only beyond half the group size; otherwise weak.
Graders that score each step. For search over steps (several candidates per step, a chain of steps), each step is measured on its own, because a step’s scores depend on how many steps remain; the tasks used are the ones whose length is the search depth. At each step, the right and wrong children of the nodes on the paths of the known solutions give the margin, and the verdict compares it with that step’s threshold: the best score that the wrong branches of a tree with candidates per step reach per step, computed from the step’s measured wrong-score distribution,
which is for normally distributed scores (shown beside it). The grader’s verdict is its weakest step’s. Above the threshold at every step, a beam of modest width keeps the right path (success climbs to its ceiling within the first few widths), and if longer tasks score like these, the width they need grows slowly with their length; below it at some step, a search loses the right path there, and more so on longer tasks. This is the freezing transition of a directed polymer on a tree (Derrida and Spohn, 1988), applied to search. The success against beam width comes from simulated trees whose steps are drawn, level by level, from the measured distributions. On the sample project’s real step grader (Treecode’s Milestone 1) this predicts 85% at width 1 and 98% at width 26; the Milestone 1 harness measured 89% and 98.7% on its test tasks.
Rollout Budget. A group of answers to a prompt at coverage is all right or all wrong with probability ; such a group teaches nothing (dynamic sampling discards these groups after paying for them). The learning signal per rollout is
the size of the advantages group-normalized training assigns, divided by the rollouts spent. How learning grows with the rollouts spent on one prompt is not settled, so the page shows two proxies: linear ( per prompt) and concave (, since a group’s estimate sharpens as ). Under the concave proxy the cheapest allocation at equal signal gives each class a group size proportional to , within your floor (never below 4) and a cap; under the linear proxy the optimum is a corner, shown only as a bound. The saving is the rollouts not drawn at equal signal.
Cost per Solved Task. With tokens per candidate at price , and the grader’s own cost per candidate, the cost per solved task is . It does not fall as grows, so the page reports the smallest that meets your target and, if you give a cost of failure , the that minimizes ; the same numbers are priced at the dated, listed prices of serving vendors.
Checkpoints. The same measurements per training checkpoint. The alarm fires when the within-prompt margin falls below the previous checkpoint’s band for two checkpoints in a row while the average score rises, the signature of a grader being gamed (the gold-against-proxy test of Gao, Schulman and Hilton, 2022). Nothing is extrapolated below eight checkpoints.
How well the prediction does on the sample run. On the simulated run of the sample project, the engine predicts best-of- up to from 16 candidates per prompt; measured on an independent table of the same simulated population with 64 candidates per prompt, the prediction lies between 0.5 and 4 points below the measurement at every , never above it. The simulated population departs from the model in documented ways (bounded scores, prompt shifts, unequal spreads, label noise), so this tests robustness to those departures; it is not evidence about real graders, because a scenario generated from the model cannot validate it. The research program page says what has been measured on real graders so far and what has not.
What is prior art. The threshold, the AUC, the gold-against-proxy test and the wasted-group arithmetic are known. What the product adds is their combination on your own grader with error bars: within-prompt margins, the model-free check on the predicted best-of- curve, the measured threshold for search over steps, and the allocation at equal signal. The engine is the open package treecode-calibrate; its README states every formula and output schema.