System One Bench

Loading the latest results.

System One models read a situation and answer typed questions in one fast call: yes or no, pick one option, or score a level, each with probabilities. Every model gets the same cases and the same request body, scored against each dataset's labels.

System One Index

Mean accuracy across all 16 datasets. Higher is better.

Latency

Median time per call, in milliseconds. APIs with 8 calls in flight; Laya locally, one at a time. Lower is better.

Cost

USD per 1,000 calls, as billed by the API. Lower is better.

System One Index

The index is the mean accuracy across datasets, with every dataset weighted equally, on a 0 to 100 scale. A model gets an index only if it ran every dataset in the set, so the Arabic fine-tunes appear in the Arabic index alone.

All datasets

Accuracy against speed and cost

The shaded corner is the most attractive: more accurate, and faster or cheaper. The line joins the models that nothing else beats on both measures.

Index against latency

System One Index (all datasets) against median latency per call. API models measured through OpenRouter with 8 calls in flight; Laya on a local Apple M5 GPU, so its latency excludes any network.

Index against cost

System One Index (all datasets) against USD per 1,000 calls. Local models have no per-call price.

Calibration, speed and cost

Calibration error is the average gap between a model's stated confidence and how often it's right, in percentage points. A well-calibrated model's 80% answers are right about 80% of the time.

Calibration error

Top-label ECE over every question, all datasets. Lower is better.

Latency, median and 95th percentile

Milliseconds per call. The bar is the median; the tick marks the 95th percentile, shown on hover. Lower is better.

Cost to run System One Bench

USD for all 3,250 cases, as billed. Lower is better.

Input tokens per call

Mean billed input tokens. Each tokenizer counts differently, so this explains cost rather than ranks models.

By dataset

One chart per dataset. Bars are colored by who makes the model.

    Show every dataset as a table

    All runs

    Every model run with its overall numbers. Overall accuracy here counts every question, so the 5-question typed-decisions dataset weighs more than in the index.

    ModelIndexAccuracyCalibration error Median latencyCost per 1,000 callsDatasets

    Roadmap

    Models we plan to add. Vendors publish their own scores for several of these; none appear on this page until they've run on System One Bench like every other model.

    Method

    What's measured

    • Each case is a state plus one to five named questions, sent unchanged to every model.
    • Accuracy compares the most likely answer with the dataset's label. A failed call counts as wrong.
    • Calibration error is top-label ECE over 10 bins. Log loss, Brier score and more are in results.json.
    • Cost is the summed usage.cost the API reports.
    • API latency was measured with 8 requests in flight through OpenRouter; Laya ran locally on an Apple M5 GPU, one call at a time.
    • Nothing was tuned on test data. Some open-weight fine-tunes trained on related data; the repo README lists which.

    Run it yourself

    export OPENROUTER_API_KEY=<your key>
    python run_decisions.py --model openai/gpt-6-luna-decisions --name luna
    python run_decisions.py --model typesafe/jev-1.13 --name or_jev
    python score.py preds_or_jev.jsonl preds_luna.jsonl

    To race two models live on the same cases, run python server.py and open localhost:8061. The race needs API keys, so it runs on your machine rather than on this page.