Pricing

Subscription for the many teams, usage for the heavy ones, a report fee for vendors.

Prototype prices, to be tested with the first ten customers. Every plan has every page; what grows with the price is how many graders and prompts a month, whether training runs with checkpoints are included, and the support.

The plans

What each plan is for.

planpriceallowancefor whom
Free$01 grader, 1,000 prompts a month, 4 calibrations a month (about one a week), 2 API keys, every page, the laptop packagea team trying it on one grader
Pro$99 a month3 graders, 10,000 prompts a month, unlimited calibrations, 5 API keys, the API and the SDKone engineer with a few graders
Team$990 a month10 graders, 100,000 prompts a month, unlimited calibrations, training runs with checkpoints, 20 API keys, email supporta team running RL on open weights
Scale$2,900 a monthunlimited graders, 500,000 prompts a month then $4 per thousand, unlimited keys, priority support, single sign-on when built, one grader report includeda lab team or a platform
Grader report$2,500 per reportan independent report on one grader, re-measured on each new batch for a year; included on Scaleenvironment and grader vendors
Enterprisecustomprivate deployment, or a share of measured rollout savings (15%) in place of the subscriptionlaboratories, insurers, certifiers

A read-only report link can be shared on every plan. The laptop package is free on every plan and needs no account.

Every allowance

What the service enforces, per calendar month.

plangradersprompts a monthcalibrations a monthtraining runsAPI keysprice
Free11,0004no2$0
Pro310,000unlimitedno5$99 / month
Team10100,000unlimitedyes20$990 / month
Scaleunlimited500,000, then $4 per 1,000unlimitedyesunlimited$2,900 / month
Enterpriseunlimitedunlimitedunlimitedyesunlimitedcustom

Graders (not deleted) and API keys (not revoked) are counted as they stand; prompts, calibrations and checkpoints are counted within the current calendar month in UTC. Exceeding an allowance answers 402 with the allowance object and a link to the billing page; on the Scale plan prompts beyond 500,000 are never refused and appear as overage on the statement. The allowance object is in every GET /me and GET /usage response, so a script can read it before uploading.

What is metered

Prompts, not candidate rows; calibrations; a run counts once.

Prompts. A table is counted when it is stored, by its number of distinct prompts, not by its candidate rows: the larger groups that make the bands tighter cost nothing extra. The same table uploaded twice is counted twice; a calibration of a stored table with other parameters is not a new upload.

Calibrations. A calibration is counted when it is queued, whatever its outcome. The Free plan's four a month is about one a week; the other plans do not count them.

Runs and checkpoints. A training run is one run whatever the number of checkpoints it holds; the checkpoint tables are counted as prompts when each is uploaded, and the checkpoint series is recomputed free of charge each time one is added. Runs are a Team feature and above.

Free once a calibration exists. The Rollout Budget, the Cost per Solved Task, the vendor price table, the checkpoint series and the shared report link are computed from the stored report and metered by nothing: you are never charged for looking at a decision.

Why month two

A calibration is a measurement of a grader on a model at a checkpoint.

The number belongs to the three together, and all three move. Every new batch of samples re-measures the grader, which is how drift is caught: the Checkpoints page's alarm fires when the margin against ground truth falls for two consecutive checkpoints while the grader's mean score rises. Every new run brings checkpoint tables; every new grader, a new report; every change of group size, prompt mix or N, a decision the report asks about and logs. The first month answers whether the grader can be trusted; the months after keep the answer current.

The grader report

An independent report on one grader, $2,500, included on Scale.

For environment and grader vendors, whose customers ask how good the grader is and have nothing independent to point at. The report is the same Grader Check run on a sample set whose label source and provenance are stated on its face, published under Treecode's name as a read-only page the vendor can attach to the grader, and re-measured on each new batch of samples the vendor sends for a year. Treecode does not consult for the vendor and does not tune the measurement; what the samples say is what the report says, including the band and the model-free curve beside the prediction.

Enterprise

Private deployment, or a share of the measured saving.

Laboratories, insurers and certifiers get custom terms: a private deployment of the service, or, in place of the subscription, a share (15%) of the rollout savings the Rollout Budget produces. How the saving is measured: the rollouts the customer's previous allocation would have drawn on the same prompts, minus the rollouts drawn under the new allocation at equal learning signal, priced at the customer's own rate per rollout. The headline shown is the saving under the customer's own floors; the linear bound is labeled a bound; and the first published number will be a real customer's before-and-after bill, not the arithmetic on example mixes that the research program page is careful to call arithmetic.

Payment

Not connected in this prototype.

Payment collection is not connected in this prototype: plans change without a card, from the Billing page of the application or with POST /billing/plan, and take effect at once. The Billing page shows this month's usage against the allowance and the statement (the plan price, the overage and the total); invoices go by email. Enterprise and the grader report are arranged by email.