Baseline
PricingGet started free

Prompt optimization without the guesswork

Prompt optimization is how you get a meaningfully better prompt without spending a week tweaking wording and hoping. Baseline treats it as a search. It generates and tests prompt variations, scores each against your Rubric, and hands back the version that measurably wins, with the proof attached.

What it is, and why it matters

Most teams improve prompts by hand: change a sentence, run a few examples, decide it feels better, ship it. It's slow, it doesn't scale past a couple of prompts, and "feels better" is exactly the unmeasured judgment evaluation exists to replace.

Prompt optimization makes the improvement systematic. The system explores many candidate prompts, scores each against the same criteria, and keeps what actually performs. It turns prompt engineering from one person's guesswork into a measured search.

A better score is worth more when you can defend it. When a new prompt beats the rubric your team agreed on, you can ship it knowing the gain is real and show that number to anyone who asks.

See it in Baseline, step by step

  1. Point a run at a rubric and an agent

    Pick the Rubric that defines success, the agent Connection whose prompt you want to improve, and a rollout budget. The run freezes a set of input Instances up front, so every candidate is judged on identical ground.

  2. Baseline searches, scores, and keeps the winners

    The run proposes prompt variants and tests each one against the frozen inputs. Reflective mode reads the judge's written feedback and rewrites with intent; Simple mode samples rewrites and keeps the best scorers.

  3. Ship the lift, with the proof attached

    The run reports before and after scores against your rubric and lays the optimized prompt beside the seed for every Module. Copy it out when you're convinced.

    A completed Optimization Run in Baseline showing a score lift from 74% to 86% and the seed prompt beside the optimized version.

How Baseline does it

Optimization Runs

Point a run at the prompt you want improved, set a budget, and it explores candidates while your team does something else.

Two Modes

Simple Mode samples scored rewrites and keeps the best, suited to narrow tasks. Reflective Mode learns from the judge's written feedback, built for demanding rubrics.

Proven lift

Every run reports before and after scores on the same rubric you evaluate with, so the gain is measured before you ship it.

Prompts you can take with you

Winning prompts sit beside their seeds, per Module, with one-click copy. Your agent, your prompt, your call.

What you get

  • Recover the engineering weeks spent hand-tuning prompt wording.
  • Let domain experts drive prompt quality through the rubric they own.
  • Ship prompt changes with the before-and-after number attached.
  • Turn a failed evaluation directly into a better prompt.

Frequently asked questions

How is this different from a prompt playground?
A playground helps you try prompts one at a time and judge by eye. An Optimization Run tests many candidates for you and scores each against your rubric, so the winner is the one that measurably performs.
How do I know the new prompt is actually better?
Every run reports the before-and-after score on the same rubric you evaluate with, and shows the optimized prompt side by side with the seed, so you review exactly what changed and what it gained.
What is a Module?
A named prompt inside your agent that a run can improve on its own. An agent Connection declares its Modules; a run optimizes one at a time and reports each prompt separately.
Who can run an optimization?
Anyone who can read results. The expert who owns the Rubric kicks off a run and reviews the proven lift, with no prompt-engineering background required.