benchmark.space Open workspace
A NEW SPACE FOR MODEL EXPERIMENTS

Know what works.
Make it better.

Evaluate a model. Understand its limits. Explore what a small model
can do for your task, with an agent beside you and a budget in view.

INTERACTIVE PREVIEW
Cost preview before a run
Or start with
Explore a setup here. Launch experiments in your workspace.

A clear starting point

A budget you choose

Evidence you can inspect

FIRST, ASK A BETTER QUESTION

Big questions.
Well-known benchmarks.

Start with the skill you care about.
Find a test that brings it into focus.

Benchmark selection is a starting point. Your workspace confirms model compatibility, settings, and cost.

FROM “WHAT IF” TO “LET’S FIND OUT”

Small model.
Specific ambition.

A model doesn’t need to know everything to become useful at your thing.

Bring the task. Your agent helps investigate failures and explore the next experiment, from better examples to fine-tuning and RL.

Your research partner: Prime Agent
Support request routing
EXAMPLE
“Can a smaller model choose the right support tool? Let’s work within $10.”
A focused plan

First, define a correct tool call. Then measure where the model gets it right and where it needs help.

1Establish the baselineEvaluate
2Investigate the failuresUnderstand
3Try a focused improvementExperiment
4Test on unseen examplesVerify
YOUR EXPERIMENT BUDGET$10.00
Set the scope.
See the spend.

Illustrative workflow. Budget and method depend on the task.

A LITTLE LESS SETUP. A LOT MORE DISCOVERY.

From a question
to a clear next step.

One place to move from a question
to a result you can actually use.

01

Point us at the possibility.

Choose a model and a benchmark, or describe your task. Start with what you want to learn.

02

Give your agent a direction.

Review the plan and budget. Follow the evidence, inspect the details, and shape the next experiment.

03

Take the useful part with you.

Understand what changed, what still fails, and what it cost. Put the result to work.

CURIOSITY, WITH RECEIPTS.

The result is only as good
as the evidence behind it.

Know the cost.

Keep estimates, committed spend, and actual cost in the same conversation.

Look beyond one score.

Inspect the failures and regressions. See what improved and what needs another look.

Keep moving forward.

A useful result can be a better model, a better choice, or a clearer next question.

THE NEXT EXPERIMENT IS YOURS.

A little curiosity.
A better model.

Open your workspace

Start with a question worth answering.

YOUR STARTING POINT

An evaluation,
taking shape.

Model
Benchmark
Next step

This is a setup preview. No experiment has been launched and nothing has been charged. Continue in the workspace to configure your run.

Continue to workspace