Skip to content
Junior Agent
← How Junior fits your workflow

Cost & capability, with the context.

Compare concrete examples without mistaking a token price or a benchmark score for the outcome of your deliverable.

WHAT THE TOKENS COST

Different rates. The same workload.

An illustrative API basket: 1 million uncached input tokens + 100,000 output tokens. Concrete models make the price comparison auditable; Junior’s value proposition is the delegation workflow.

Direct-provider list prices in US dollars, checked October 6, 2026. Rates per million tokens.
Example modelInput / 1MOutput / 1MExample basket
CommodityDeepSeek V4.1 Flash · off-peak$0.15$0.60$0.21
CommodityDeepSeek V4.1 Flash · peak$0.30$1.20$0.42
CommodityDeepSeek V4 Pro · off-peak$0.66$1.98$0.858
Mid-tierClaude Sonnet 5.5$2.00$10.00$3.00
FrontierClaude Opus 5.5$4.00$20.00$6.00

Checked October 6, 2026. Sources: commodity example pricing and frontier and mid-tier pricing. These are direct API rates, not Claude Code or Codex subscription allowances. Cache discounts, reasoning/output volume, retries, routing-provider fees, and manager review change the bill. The Pro example also has peak rates: $1.32 input / $3.96 output.

Spend less on repeated execution

At the listed rates, this basket costs $0.21 for the Flash off-peak example and $6.00 for the premium example—about 29× apart. That is a token-price comparison, not a promised reduction in the cost of completing your task.

Count the whole handoff

Total delivery cost = planning + worker execution + retries + checks + review. Delegation earns its place when the work is substantial enough and the finish line is clear enough to outweigh that overhead.

CAPABILITY IS TASK-SPECIFIC

Very capable. Uneven gaps.

A low-cost model can be competitive on repository fixes while trailing on a harder terminal test. Choose around the deliverable, then verify the actual work.

Published snapshot: DeepSeek V4.1 Flash vs Claude Opus 5.0. Opus 5.0 is an older frontier model, not the latest release. Scores are percentages; gaps are percentage points.
EvaluationCommodity exampleFrontier snapshotGap
DeepSWE v1.1Repository issue resolution74.2%74.0%Commodity +0.2 points
Terminal-Bench 2.1Agentic terminal tasks90.6%89.1%Commodity +1.5 points
Terminal-Bench 4.0A different, harder terminal evaluation31.2%51.8%Frontier +20.6 points

Source: publisher model card, checked October 6, 2026. Publisher-reported results, not a Junior evaluation or an independently controlled head-to-head test. The commodity results use maximum reasoning effort and benchmark-specific harnesses. Its DeepSWE score with Pi is 66.2%, versus 74.2% with mini-SWE: the harness matters. Scores from different benchmark versions must not be compared.

The current frontier keeps moving

In Anthropic’s newer FrontierCode 1.1 results, proprietary Sonnet 5.5 scores 52.1% at Xhigh and Opus 5.5 scores 54.4%—a 2.3-point gap. This is a separate evaluation, not an open-weight comparison. A narrow benchmark gap does not establish equal judgment across open-ended work.

Set up Junior →