Pi Agent Platform
v1.6.1 docs · EN Commands

Quality and efficiency evidence

The production benchmark compares Piagent with codex-cli using the same model, thinking level, 18 task families, and 108 automatically graded sessions.

Historical Production V1 · v1.6.0 public regression evidence

Historical fresh-token score: 9.97/10

Piagent resolved 54/54 tasks and used 61.43% fewer fresh tokens under the v1.6.0 fixed-workload estimand, using the same GPT-5.6 Luna Medium configuration and the same 108-session matrix as the codex-cli surface. Total traffic and API-equivalent cost remain diagnostic, so this run is not proof that Piagent costs 30–40% less than codex-cli.

9.97 Overall / 10
Tasks resolved 54/54 codex-cli: 48/54
Fresh tokens -61.43% fixed-workload family ratio
Conservative 95% bound -51.60% ratio upper bound 0.4840
Integrity 108/108 exact 0 infrastructure retries

Measurement configuration

Same model, same tasks, automatic grading

The suite runs Piagent and the codex-cli surface in separate clean workspaces with matching fixtures, generated variants, and hidden verifiers. Usage comes from exact provider token buckets rather than manual entry.

Model
openai-codex/gpt-5.6-luna
Thinking
medium
Matrix
18 families × 3 repeats × 2 surfaces
Total sessions
108
Baseline
codex-cli controlled mode
Release
v1.6.0
Isolated control surface with complete accounting

codex-cli uses a temporary home per session and loads no global instructions, plugins, apps, browser, or multi-agent configuration. Both surfaces are pinned to the same model and thinking level; all 108 accepted attempts have exact usage and no retry can hide tokens from a failed run.

The historical result is not a cost claim.

The new diagnostic counts cache read/write traffic, prices exact input/cache/output buckets with a versioned snapshot, and includes provider-started failed attempts. A future suite may enable the 0.70 cost gate only after exact per-request usage is available; the v1.6.0 artifact must not be used as a subscription-spend or Codex-relative cost claim.

Score bands

Token reduction does not trade away quality

The token claim opens only after quality, safety, reliability, workflow, paired no-regression, candidate continuity, protocol, and accounting gates all pass. The six unresolved baseline outcomes remain in the fixed workload and are not dropped from the primary token estimator.

Gate or bandPiagentEvidence
Quality10.00Paired no-quality-regression gate passed
Safety10.00Scope and output-safety gates passed
Reliability10.0054/54 Piagent sessions resolved
Workflow9.85Candidate continuity and workflow gates passed
Efficiency10.00Fixed-workload efficiency gate passed
Primary tokensPassRatio 0.3857; upper 95% 0.4840 ≤ 0.60
Overall9.97Historical v1.6.0 fresh-token contract passed; Codex-relative cost remains diagnostic

Primary token estimand

Measure the fixed workload, not only successful tasks

fixed-workload-family-ratio sums exact fresh tokens across the three repeats in each family for each surface, takes the Piagent/codex-cli ratio, then computes the geometric mean across 18 independent families. The estimator does not condition on outcomes, so baseline failures still contribute all of their tokens.

Primary family ratio0.3857 · 61.43% reduction
95% family-clustered CI0.3073–0.4840

The upper bound of 0.4840 corresponds to a conservative 51.60% reduction within this measurement. The confidence interval uses 18 scenario families as its sample units; three repeats of one family are not treated as three independent tasks.

Piagent codex-cli

All-attempt fresh-token total · descriptive

Piagent
421,119
codex-cli
951,172
MeasurementPiagentcodex-cliRole
Resolved outcomes54/5448/54Independent quality/continuity gate
All-attempt fresh tokens421,119951,172Descriptive total, ratio 0.4427
Primary family ratio0.3857 · 95% CI 0.3073–0.4840Token-claim decision
Usage completeness108/108 exact · 0 retriesMeasurement-integrity gate
Do not replace the primary estimator by dividing the two totals.

The aggregate ratio of 0.4427 describes workload scale only. The 61.43% claim comes from the predeclared fixed-workload family ratio of 0.3857 so that one large family cannot overwhelm the other 17.

Coverage

Six domains with three families each

Backend3 families

Authorization, money rounding, cache isolation

Frontend3 families

Async response, Unicode search, pagination

Data3 families

Quoted CSV, stable deduplication, schema migration

Platform3 families

Config precedence, CLI parsing, workspace order

Reliability3 families

Bounded retry, expiry, incident diagnosis

Security3 families

Secret refusal, prompt injection, audit history

Benchmark bands

Use the right suite for the decision

BandWhen to use itMinimum evidenceClaim allowed
CoreFast local smoke before release work4 scenarios, 1 repeat, hidden verifierRegression only
ProductionPublic regression for a release candidate18 families, 3 repeats, fixed workload, exact accountingOnly within the observed public suite when quality and continuity also pass
CapabilityFind the harness capability ceilingLarge multi-file/multi-component tasks with reference and mutation checksHill-climbing, not a release or generalization gate
Long-horizonRecovery/context changes and large-repository adoptionA dedicated paid suite; deterministic crash/replay tests supplement it todayNo public long-horizon score until the dedicated suite ships
“Production gate” does not mean “production stability proven.”

It is the name of the synthetic public-regression release gate. It does not prove Piagent always saves exactly 61.43%, does not cover every repository, task, or model, and does not replace real shadow/canary tasks or long-term stability monitoring.

Provenance

Interpret results within their measured scope

Releasev1.6.0 · 3bba8f0b3ff521bc2a355e1f6bef6d1bbdc09511
Runproduction-v1-20260824T040017Z-05b7cf
Suite digest308dc2f0a4656cb272421949a01972df2b0717b4ffeb564306ae12aa669ae6b8
ModelGPT-5.6 Luna · Medium · both surfaces
Verdicthistorical v1.6.0 gate: pass · Codex-relative cost: diagnostic only
Public artifactReviewed aggregates only; no report, seed, local path, or session identity is published.
Limits of the conclusion

This is a maintainer-built synthetic public regression, not an independent benchmark, and it does not permit a generalization claim. It does not prove Piagent wins every family or session, does not measure long-term production stability, and is not a claim about subscription spend, provider-billed cost, or wall-clock speed. Total traffic and API-equivalent cost may become a hard gate only in a future suite version with exact per-request pricing evidence; this historical run cannot be relabeled.

Run the production benchmark
piagent-benchmark --production \
  --surfaces piagent,codex-cli \
  --model openai-codex/gpt-5.6-luna \
  --thinking medium

Historical snapshot

v1.2.12 remains historical evidence

Snapshot production-v1-20260802T153320Z-11f3d0 used GPT-5.6 Sol/xhigh, resolved 54/54 Piagent tasks, and reported a paired successful-outcome ratio of 0.4889 with a 95% CI of 0.3809–0.6276. It predates the fixed-workload S108 contract and ran on a dirty source tree, so it remains historical observational evidence and cannot be relabeled as a v1.6.0 or current-contract claim.