Braintrust Pricing in 2026: What You Actually Pay for LLM Evals
Braintrust bills by evaluation scores (rows x scorers x runs), not seats. We decode the meter, find the ~199,000-score Pro crossover, work three 30-day bills, and show a one-lever fix that cut a bill 71%.

On this page
Quick answer (2026): Braintrust has three tiers. Starter is free, Pro is $249 a month, and Enterprise is custom. The sticker price is almost never your real bill. Braintrust meters four separate things, and for most teams one of them, the evaluation "scores" meter, is what the invoice is actually made of. Starter includes 10,000 scores a month and charges $2.50 per 1,000 after that. Pro includes 50,000 scores and charges $1.50 per 1,000 after that. That single meter is why two teams running what looks like "the same" eval suite can be billed $5 and $3,378 in the same month. Everything below is sourced from Braintrust's pricing page (2026).
The plans, decoded (2026)
Scroll to see more
| Plan | Monthly | Scores included | Scores overage | Processed data | Data overage | Retention |
|---|---|---|---|---|---|---|
| Starter | $0 | 10,000 | $2.50 / 1,000 | 1 GB | $4 / GB | 14 days |
| Pro | $249 | 50,000 | $1.50 / 1,000 | 5 GB | $3 / GB | 30 days, then $0.50 / GB / mo |
| Enterprise | Custom | Custom | Custom | Custom | Custom | Custom |
Two things the table hides. First, both tiers also hand you model credits ($10 on Starter, $249 on Pro) that only apply to Braintrust's built-in models and Topics. If you bring your own API key, those credits are moot, and on Pro the $249 sticker becomes a pure platform fee that you recoup only through the cheaper overage rates and the larger free blocks. Second, and this is the important one, both tiers include unlimited users, projects, datasets, playgrounds, and experiments. Your headcount does not move this bill. Usage does.
What Braintrust actually meters: "scores"
A score is one output from a scorer logged against one row of an eval. A scorer can be an LLM-as-a-judge call, a built-in autoeval, or your own code check. The meter is not per eval run, and it is not per seat. It is:
scores = rows x scorers x runs.
That multiplication is where the surprises live. Add a fourth scorer to a suite and every run costs a third more. Move the suite from "run it when I remember" to "run it on every pull request" and the run count becomes your CI traffic, not your team size. A 500-row suite with 6 scorers is 3,000 scores per run before you have measured anything twice. The seat count never enters the equation.
The number that decides your plan: about 199,000 scores a month
Ignore data for a second and price the scores meter on its own. On Starter you pay $2.50 per 1,000 scores above the free 10,000. On Pro you pay $249 flat plus $1.50 per 1,000 above the free 50,000. Set those two lines equal and they cross at roughly 199,000 scores a month.
Below about 199,000 scores a month, Starter plus overage is cheaper than Pro. Above it, Pro's lower per-score rate and larger free block win, and the $249 stops being a sticker and starts being a discount. Most teams never do this subtraction. They either default to Pro far too early or cling to Starter well past the point where the overage rate is quietly costing them more than the subscription would.
Three 30-day bills, same product
Scroll to see more
| Team | Rows | Scorers | Runs / mo | Scores / mo | Data | Cheapest plan | 30-day bill |
|---|---|---|---|---|---|---|---|
| A. Indie, on-demand | 200 | 3 | 20 | 12,000 | 0.4 GB | Starter | ~$5 |
| B. CI on every PR, scorer sprawl | 500 | 6 | 450 | 1,350,000 | 8 GB | Pro | ~$2,208 |
| C. Heavy prod logging | 1,000 | 2 | 30 | 60,000 | 60 GB | Starter | ~$361 |
Team A runs a small suite a few times a week. At 12,000 scores it is 2,000 over the free block, so $5.00, and under 1 GB of data means no data charge. Braintrust's free tier genuinely covers real solo eval work.
Team B is the trap. Six scorers, a 500-row regression suite, running in CI on every pull request plus nightly, is about 1,350,000 scores a month. On Starter that is (1,350,000 - 10,000) / 1,000 x $2.50 = $3,350 in scores plus $28 of data overage, or about $3,378. On Pro it is $249 plus (1,350,000 - 50,000) / 1,000 x $1.50 = $1,950 in scores plus $9 of data, about $2,208. The $249 sticker is a rounding error next to the meter. And the driver was scorer sprawl on the CI path, not the size of the team.
Team C flips the story. Only two scorers, so just 60,000 scores, but 60 GB of production traces logged. Here the bill is data, not scores. On Starter it is $125 of scores plus $236 of data, about $361. On Pro it is $264 of scores plus $165 of data, about $429. Because Team C sits below the ~199,000-score crossover, Starter is actually the cheaper plan even though it logs a lot. Pro is not automatically the "serious team" choice.
The one lever that cut Team B's bill 71%
You do not need all six scorers on every pull request. Split the suite by path:
- Nightly, full coverage: all 6 scorers on all 500 rows, 30 runs a month = 90,000 scores.
- Per pull request, smoke signal: 1 scorer on all 500 rows, ~420 runs = 210,000 scores.
Total drops from 1,350,000 to 300,000 scores a month. Team B's Pro bill falls from about $2,208 to about $633, a 71% cut, with the full six-scorer regression still running every single night. Same tests, same rows, cheaper signal on the noisy CI path. If you also sample rows on pull requests, say 100 of 500, it drops further. The cheapest fix here was never a plan change. It was running fewer scorers where the runs are frequent.
Which meter is your driver?
Before you pick a plan, tag your own driver with two quick numbers. Multiply rows x scorers x runs for your scores figure, and estimate GB logged per month for your data figure. Whichever is bigger tells you which lever to pull, and, as Team C showed, can flip which plan is cheaper. Scores are the headline for teams doing heavy evaluation; processed data is the quiet one for teams logging large production traces. Do not assume you know which camp you are in until you have done the multiplication.
The escape hatch: run the scorers yourself
The scorers are open source. Braintrust maintains autoevals under the MIT license, the same LLM-as-a-judge, heuristic, and statistical scorers the hosted product runs. Run them in your own CI and log pass or fail to a store you own, and the scores meter drops to zero. You pay compute you already have instead of $1.50 per 1,000. What you give up is the hosted experiment UI, regression diffing, and shared experiments, which is precisely what you are renting on Pro.
That is the honest line. Managed wins when the diffing UI and shared experiments are real product surface for your team. Roll-your-own wins when all you need is a red-or-green gate in CI. If you want a middle ground, hosted tracing without an eval-scores meter, an open-core tool like Langfuse prices tracing differently (see our Langfuse pricing teardown), and the same build-versus-buy crossover applies.
Math check
A score is rows times scorers times runs, and Braintrust charges by the score, not the seat. Count that product before you count anything else. Below about 199,000 scores a month, stay on Starter and pay overage. Above it, Pro pays for itself. And if the number is scaring you, the cheapest move is usually not a plan change. Math check: 1,350,000 scores becomes 300,000 with one nightly-versus-per-PR split, and a $2,208 bill becomes $633, same nightly coverage.
Written by
Diego AguirreFrequently asked questions
How much does Braintrust cost in 2026?
Braintrust has a free Starter plan, a Pro plan at $249 a month, and custom Enterprise pricing. Your real bill is usually driven by usage overage, not the sticker: Starter includes 10,000 evaluation scores a month and charges $2.50 per 1,000 after that, while Pro includes 50,000 scores and charges $1.50 per 1,000 after that, plus processed-data charges on both.
What is a "score" in Braintrust billing?
A score is one output from a scorer (an LLM-as-a-judge call, a built-in autoeval, or a custom code check) logged against one row of an eval. It is not billed per run or per seat. The number that grows is rows x scorers x runs, so adding scorers or running the suite on every pull request is what drives the meter.
Is Braintrust Pro always cheaper than Starter?
No. On the scores meter alone the two plans cross at roughly 199,000 scores a month. Below that, Starter plus overage is cheaper; above it, Pro's lower per-score rate and larger free block win. Teams with modest scores but heavy data logging can also be cheaper on Starter, so check your driver before upgrading.
What is included free on Braintrust Starter?
Starter is $0 a month and includes 10,000 scores, 1 GB of processed data, 14-day retention, and $10 of model credits, plus unlimited users, projects, datasets, playgrounds, and experiments. Overage is $2.50 per 1,000 scores and $4 per GB of processed data.
How do I lower a large Braintrust bill?
Attack rows x scorers x runs on your noisiest path. Run the full multi-scorer regression once a night and a single smoke scorer on pull requests, and optionally sample rows on PR runs. In a worked example this cut 1,350,000 scores a month to 300,000 and dropped a Pro bill from about $2,208 to about $633, a 71% cut, with full nightly coverage intact.
Can I use Braintrust's scorers without paying for the platform?
Yes. Braintrust maintains autoevals under the MIT license, the same LLM-as-a-judge, heuristic, and statistical scorers the hosted product runs. Run them in your own CI and log results to a store you own and the scores meter goes to zero. You lose the hosted experiment UI, regression diffing, and shared experiments, which is what the paid plan rents you.
Do extra team members increase Braintrust's cost?
No. Users are unlimited on both Starter and Pro. The bill is usage (evaluation scores and processed data), not seats, so adding teammates does not change the invoice by itself.
Related reading
Langfuse Pricing in 2026: What a Unit Really Costs
Langfuse pricing in 2026 starts free (Hobby: 50,000 units per month), then $29/month (Core), $199 (Pro), and $2,499 (Enterprise), with extra units at $8 per 100,000. The catch: a unit is every trace, observation, and score, so one multi-step agent request can burn 20 or more units. That is why a busy agent app pays far more than its request count suggests.
Build vs buy: when DIY beats SaaS at scale
Build-vs-buy isn't a values debate, it's a breakeven date. SaaS wins early because it converts a big upfront build into a small monthly fee. DIY wins once your usage-based SaaS bill exceeds the fully-loaded cost of owning the code, typically around month 18 for infrastructure-style tools. This piece gives you the formula, a worked example, and the three traps that make teams build too early.
OpenRouter pricing 2026: real 30-day bill vs direct providers
OpenRouter takes 5.5% on credit purchases with a $0.80 minimum, then passes provider rates through. Here is the real 30-day bill across Anthropic, OpenAI, and Mistral at four workload mixes, with the anti-patterns where the routing tax stops being worth it.

