The short answer
Default to gpt-6-luna and move work to claude-haiku-5-5 where Claude-style instruction following changes the output. Luna 6 is ~$0.03 / $0.15 per 1M in/out; Haiku 5.5 is ~$0.07 / $0.37 — about 2.5x on both input and output. Both take a 1M-token context window and both are one model= change apart on the same OpenAI-compatible key, so the decision is not integration, it is unit cost against output quality.
Side by side
| Model ID | Input / output per 1M | Context | Family position |
|---|---|---|---|
gpt-6-luna | ~$0.03 / $0.15 | 1M | Cheapest GPT-class chat row on Keyo |
claude-haiku-5-5 | ~$0.07 / $0.37 | 1M | Cheapest Claude-class row on Keyo |
Neither is the only option below them. On the same key, deepseek-v4-flash-0731 sits at ~$0.09 / $0.17 and claude-sonnet-5 at ~$0.44 / $2.21 — so Haiku 5.5 is also the cheaper way into the Claude family than Sonnet 5. Rates are indicative sell prices; confirm live on /pricing-list.
What the 2.5x buys
Run a month of realistic bulk traffic — 10M input and 2M output tokens:
gpt-6-luna 10M in × $0.0294 = $0.294
2M out × $0.1471 = $0.294 ≈ $0.59 / month
claude-haiku-5-5 10M in × $0.0735 = $0.735
2M out × $0.3677 = $0.735 ≈ $1.47 / month
That is about $0.88 more per month at this volume — the absolute gap stays small until the volume is large. At 10x that traffic (100M in / 20M out) the difference becomes roughly $8.80 per month, and at 100x roughly $88. The practical read: at low volume the 2.5x is noise and you should pick on output quality; at high volume the gap is real and you should split the traffic.
The split that usually works: keep classification, extraction, routing and bulk drafting on gpt-6-luna, and route the steps where you have seen Luna miss — long instruction chains, structured output that must follow a schema, or Claude-specific writing style — to claude-haiku-5-5.
Which to pick when
- Pick Luna 6 when unit cost decides: high-volume classification, triage, tagging, first-pass drafting, anything you run millions of times and sample instead of read.
- Pick Haiku 5.5 when instruction following decides: multi-step extraction from long documents, schema-bound JSON output, or bulk work where you have to trust every row rather than spot-check it.
- Neither if you have not measured yet — prototype on the fixed $0 catalog IDs first (see /free-models), then promote only the traffic that actually fails your evaluation.
Both rows and the rest of their families are listed on /claude-api-pricing (Claude side) and /model/gpt-6-luna (Luna side, with gpt-6-sol and gpt-6-astra above it). For the per-token math behind any single request, see the worked example on Claude API Pricing per Token.