RRackor
LLM benchmarking

Which model is best for

Benchmark any OpenAI-compatible model on your own tasks. Rackor shows you where a smaller, cheaper model holds up, with full transcripts and scores to back it.

llama.cpp·vLLM·Ollama·OpenRouter·OpenAI
Run a1b2c3d4 · liverunning
llama-3.1-8b168/240
gpt-4o
94%
llama-3.1-8b
89%
phi-3-mini
61%
llama-3.1-8b passes at ~1/20th the cost of gpt-4o.

Author your own benchmarks

Questions, prompt, sampling, scorer: versioned and re-runnable.

Full transcripts & scoring

Every prompt, response, and score. See exactly why it passed or failed.

Head-to-head compare

Models side by side, question by question. Find where cheap holds up.

Any OpenAI-compatible endpoint

Point at a URL and key, local or hosted. Nothing to rewrite.

How it works

Three steps to a real answer.

No SDK, no data export. Connect, benchmark, decide.

01

Connect an endpoint

Add any OpenAI-compatible URL and key. Test it in one click.

02

Pick or author a benchmark

Run a published version, or write and version your own.

03

Run and compare

Benchmark one or many. Track pass rates and compare head-to-head.

Pricing

Priced to run more evals, not fewer.

Billed in credits: 1 completed run = 1 credit; an AI-generated question = 3.

Free

Start evaluating models, no card required.

$0
Start for free
  • 2 seats
  • 10 credits / month
  • 30-day transcript history
  • Community-visible benchmarks
  • Add a card for pay-as-you-go: $0.10/credit, $20/mo cap
Individual

A monthly credit pool for steady, solo evaluation.

$5/ mo
Start for free
  • 1 seat
  • 50 credits / month
  • Indefinite transcript history
  • Community-visible benchmarks
  • $0.05 / credit overage
Enterprise

Negotiated volume and terms.

Custom
Contact sales
  • Unlimited seats
  • Negotiated credit volume
  • Private benchmarks & runs
  • 1-year transcript history
  • Custom terms
FAQ

Questions, answered plainly.

Anything that speaks the OpenAI chat-completions API: llama.cpp, vLLM, Ollama, OpenRouter, OpenAI, and most hosted providers. You add a base URL and a key, and Rackor runs against it. No SDK or code changes.

Stop guessing which model to ship.

Start free, no card required. Connect an endpoint and run your first benchmark in minutes.

Start for free