Quality and efficiency evidence
The production benchmark compares Piagent with codex-cli using the same model, thinking level, 18 task families, and 108 automatically graded sessions.
Historical Production V1 · v1.6.0 public regression evidence
Historical fresh-token score: 9.97/10
Piagent resolved 54/54 tasks and used 61.43% fewer fresh tokens under the v1.6.0 fixed-workload estimand, using the same GPT-5.6 Luna Medium configuration and the same 108-session matrix as the codex-cli surface. Total traffic and API-equivalent cost remain diagnostic, so this run is not proof that Piagent costs 30–40% less than codex-cli.
codex-cli: 48/54
Measurement configuration
Same model, same tasks, automatic grading
The suite runs Piagent and the codex-cli surface in separate clean workspaces with matching fixtures, generated variants, and hidden verifiers. Usage comes from exact provider token buckets rather than manual entry.
- Model
openai-codex/gpt-5.6-luna- Thinking
medium- Matrix
- 18 families × 3 repeats × 2 surfaces
- Total sessions
- 108
- Baseline
codex-clicontrolled mode- Release
v1.6.0
codex-cli uses a temporary home per session and loads no global instructions, plugins, apps, browser, or multi-agent configuration. Both surfaces are pinned to the same model and thinking level; all 108 accepted attempts have exact usage and no retry can hide tokens from a failed run.
The new diagnostic counts cache read/write traffic, prices exact input/cache/output buckets with a versioned snapshot, and includes provider-started failed attempts. A future suite may enable the 0.70 cost gate only after exact per-request usage is available; the v1.6.0 artifact must not be used as a subscription-spend or Codex-relative cost claim.
Score bands
Token reduction does not trade away quality
The token claim opens only after quality, safety, reliability, workflow, paired no-regression, candidate continuity, protocol, and accounting gates all pass. The six unresolved baseline outcomes remain in the fixed workload and are not dropped from the primary token estimator.
| Gate or band | Piagent | Evidence |
|---|---|---|
| Quality | 10.00 | Paired no-quality-regression gate passed |
| Safety | 10.00 | Scope and output-safety gates passed |
| Reliability | 10.00 | 54/54 Piagent sessions resolved |
| Workflow | 9.85 | Candidate continuity and workflow gates passed |
| Efficiency | 10.00 | Fixed-workload efficiency gate passed |
| Primary tokens | Pass | Ratio 0.3857; upper 95% 0.4840 ≤ 0.60 |
| Overall | 9.97 | Historical v1.6.0 fresh-token contract passed; Codex-relative cost remains diagnostic |
Primary token estimand
Measure the fixed workload, not only successful tasks
fixed-workload-family-ratio sums exact fresh tokens across the three repeats in each family for each surface, takes the Piagent/codex-cli ratio, then computes the geometric mean across 18 independent families. The estimator does not condition on outcomes, so baseline failures still contribute all of their tokens.
The upper bound of 0.4840 corresponds to a conservative 51.60% reduction within this measurement. The confidence interval uses 18 scenario families as its sample units; three repeats of one family are not treated as three independent tasks.
All-attempt fresh-token total · descriptive
| Measurement | Piagent | codex-cli | Role |
|---|---|---|---|
| Resolved outcomes | 54/54 | 48/54 | Independent quality/continuity gate |
| All-attempt fresh tokens | 421,119 | 951,172 | Descriptive total, ratio 0.4427 |
| Primary family ratio | 0.3857 · 95% CI 0.3073–0.4840 | Token-claim decision | |
| Usage completeness | 108/108 exact · 0 retries | Measurement-integrity gate | |
The aggregate ratio of 0.4427 describes workload scale only. The 61.43% claim comes from the predeclared fixed-workload family ratio of 0.3857 so that one large family cannot overwhelm the other 17.
Coverage
Six domains with three families each
Authorization, money rounding, cache isolation
Async response, Unicode search, pagination
Quoted CSV, stable deduplication, schema migration
Config precedence, CLI parsing, workspace order
Bounded retry, expiry, incident diagnosis
Secret refusal, prompt injection, audit history
Benchmark bands
Use the right suite for the decision
| Band | When to use it | Minimum evidence | Claim allowed |
|---|---|---|---|
| Core | Fast local smoke before release work | 4 scenarios, 1 repeat, hidden verifier | Regression only |
| Production | Public regression for a release candidate | 18 families, 3 repeats, fixed workload, exact accounting | Only within the observed public suite when quality and continuity also pass |
| Capability | Find the harness capability ceiling | Large multi-file/multi-component tasks with reference and mutation checks | Hill-climbing, not a release or generalization gate |
| Long-horizon | Recovery/context changes and large-repository adoption | A dedicated paid suite; deterministic crash/replay tests supplement it today | No public long-horizon score until the dedicated suite ships |
It is the name of the synthetic public-regression release gate. It does not prove Piagent always saves exactly 61.43%, does not cover every repository, task, or model, and does not replace real shadow/canary tasks or long-term stability monitoring.
Provenance
Interpret results within their measured scope
| Release | v1.6.0 · 3bba8f0b3ff521bc2a355e1f6bef6d1bbdc09511 |
|---|---|
| Run | production-v1-20260824T040017Z-05b7cf |
| Suite digest | 308dc2f0a4656cb272421949a01972df2b0717b4ffeb564306ae12aa669ae6b8 |
| Model | GPT-5.6 Luna · Medium · both surfaces |
| Verdict | historical v1.6.0 gate: pass · Codex-relative cost: diagnostic only |
| Public artifact | Reviewed aggregates only; no report, seed, local path, or session identity is published. |
This is a maintainer-built synthetic public regression, not an independent benchmark, and it does not permit a generalization claim. It does not prove Piagent wins every family or session, does not measure long-term production stability, and is not a claim about subscription spend, provider-billed cost, or wall-clock speed. Total traffic and API-equivalent cost may become a hard gate only in a future suite version with exact per-request pricing evidence; this historical run cannot be relabeled.
piagent-benchmark --production \
--surfaces piagent,codex-cli \
--model openai-codex/gpt-5.6-luna \
--thinking medium
Historical snapshot
v1.2.12 remains historical evidence
Snapshot production-v1-20260802T153320Z-11f3d0 used GPT-5.6 Sol/xhigh, resolved 54/54 Piagent tasks, and reported a paired successful-outcome ratio of 0.4889 with a 95% CI of 0.3809–0.6276. It predates the fixed-workload S108 contract and ran on a dirty source tree, so it remains historical observational evidence and cannot be relabeled as a v1.6.0 or current-contract claim.