New — Agent Runtime 2.0

Build AI that ships itself.

PROMPT ZERO is the operating system for LLM products. Version prompts, run evals and deploy agents from one calm, blazing-fast workspace.

~/agents/support-triage

LIVE

>

prompt-zero run triage --model auto --evals strict

prompt-zero run triage --model auto --evals strict

✓ routed to sonnet-5

212 ms

$0.0004 / call

48/48 evals passed

INTENT

refund_request

SENTIMENT

−0.62 frustrated

PRIORITY

P1 · escalate

Plugs into the models and stack you already run

  • OpenAI

  • Anthropic

  • Gemini

  • Ollama

  • Perplexity

  • LangChain

  • Replicate

  • Vercel

  • Supabase

  • PyTorch

  • NVIDIA

  • Cloudflare

  • OpenAI

  • Anthropic

  • Gemini

  • Ollama

  • Perplexity

  • LangChain

  • Replicate

  • Vercel

  • Supabase

  • PyTorch

  • NVIDIA

  • Cloudflare

[ 01 ] Manifesto

Most AI features die in a notebook. PROMPT ZERO turns fragile prompts into versioned, tested, observable software so your team ships on Friday and sleeps on Saturday.

MK

Mara Kessler

Co-founder & CEO

[ 02 ] Workflow

Three commands from idea to production.

No glue code, no spreadsheets of test cases, no Friday-night hotfixes. PROMPT ZERO gives prompts the same lifecycle as the rest of your code.

01

01

Write

Draft prompts in a typed editor with variables, versions and inline diffs your whole team can review.

$

prompt-zero init support-bot

02

02

Evaluate

Score every change against golden datasets, LLM judges and custom checks before anything merges.

$

prompt-zero eval --suite golden

03

03

Deploy

Ship to the edge with canary rollouts, automatic model fallbacks and one-click rollback.

$

prompt-zero deploy --canary 10%

[ 03 ] Platform

Everything your LLM stack was missing.

One workspace replaces the prompt spreadsheet, the eval scripts, the tracing dashboard and the deploy checklist.

One workspace replaces the prompt spreadsheet, the eval scripts, the tracing dashboard and the deploy checklist.

  • sonnet-5

  • gpt-5

  • gemini-3-pro

  • llama-4-70b

  • mistral-large

  • fable-5-1

  • qwen-3

  • haiku-4.5

  • o-series

  • deepseek-v4

  • command-r+

  • phi-5

Model router

Send every request to the cheapest model that passes your evals. Automatic fallbacks when a provider hiccups.

accuracy

94%

grounding

88%

tone

97%

latency

72%

Evals on every commit

Golden sets, LLM judges and regex checks run in CI. Regressions never reach users.

Live traces

Every token, tool call and cost, streamed in real time with full replay.

12

system: You are a support agent for Acme.

system: You are a support agent for Acme.

13

-

tone: Be friendly and helpful.

tone: Be friendly and helpful.

14

+

tone: Be concise. Max 3 sentences. No emojis.

tone: Be concise. Max 3 sentences. No emojis.

15

+

policy: Escalate refunds over $500 to a human.

policy: Escalate refunds over $500 to a human.

16

output: json(intent, sentiment, priority)

output: json(intent, sentiment, priority)

Prompt diffs & reviews

Branch, diff and review prompts like code. Comment on a single token, approve, merge.

Guardrails built in

PII redaction, jailbreak detection and policy checks on every request, in under 5 ms.

fra-1

11 ms

iad-1

14 ms

sfo-2

9 ms

gru-1

23 ms

sin-1

18 ms

syd-1

21 ms

nrt-1

12 ms

lhr-1

8 ms

Global edge deploys

Canary rollouts to 34 regions with instant rollback. Your users never wait on a cold start.

[ 04 ] Integrations

Every model. One interface.

Swap providers with a config change, not a rewrite. PROMPT ZERO speaks every major model API, vector store and tool protocol out of the box.

40+ models across 12 providers

MCP tools and function calling

Bring your own keys or use ours

>_

>_

12

ms

Median router overhead

12

ms

Median router overhead

99.99

%

API uptime, last 12 months

4.2

B

Tokens routed every month

−38

%

Average LLM spend after 30 days

[ 05 ] Wall of love

Loved by teams who ship on Fridays.

  • “We replaced four internal tools and a 2,000-row eval spreadsheet in one sprint. Our prompt regressions went to zero.”

    PR

    Priya Raman

    Staff ML Engineer, Northwind

  • “The model router alone paid for the year. Same quality, a third of the bill.”

    JW

    Jonas Weber

    CTO, Kitelane

  • “Finally, prompt reviews that feel like code reviews. Product and engineering speak the same language now.”

    AL

    Ava Lindqvist

    Head of AI, Parcelo

  • “Canary deploys for prompts sound boring until they save your launch. Twice.”

    MO

    Marcus Oyelaran

    Founder, Relaywise

  • “Traces are absurdly good. I can replay a bad conversation token by token and fix it before lunch.”

    SM

    Sofia Marín

    AI Lead, Brightloop

  • “Setup took eleven minutes. Our first eval suite ran before the coffee was done.”

    DC

    Daniel Cho

    Engineer, Quillbase

  • “Guardrails caught a PII leak in staging that three humans missed. That was the day we signed.”

    HB

    Hannah Becker

    Security, Tessaro

  • “It is the rare dev tool that makes the whole team faster, not just the person who set it up.”

    LM

    Leo Martins

    VP Engineering, Orbitdesk

[ 06 ] Get started

Start at zero. Ship today.

Free for your first 1M tokens. No credit card, no sales call, no nonsense.

psst — the keys are draggable

⌘

⌘

K

>_

>_

Esc

AI

Create a free website with Framer, the website builder loved by startups, designers and agencies.