Technical benchmark
Qwen3 support analysis benchmark on 60 synthetic messages
Three independent ReplyPilot runs measuring topic stability, risk output, generated assets, latency, token usage, and estimated inference cost.
Published 2026-07-24 · Model pricing and behavior may change
3
Independent runs
3/3
Topic agreement
12.0s
Average latency
$0.000655
Average inference
Direct result
All three runs returned the same seven topic names, volumes, risk levels, readiness score of 68, and “needs validation” decision. Generated FAQ, high-risk-example, and action-asset counts varied substantially, so those outputs still require review.
Stable topic result
| Topic | Volume | Risk |
|---|---|---|
| Shipping & tracking | 19 | yellow |
| Order edits & promos | 13 | yellow |
| Returns, refunds & replacements | 10 | yellow |
| Risk, legal & escalation | 8 | red |
| Product info & fit | 8 | green |
| Billing, invoices & subscriptions | 1 | yellow |
| Other support questions | 1 | yellow |
Run-by-run measurements
| Run | Latency | Tokens | Cost | Risks | FAQs | Assets |
|---|---|---|---|---|---|---|
| 1 | 11.7s | 5,540 | $0.000604 | 3 | 0 | 7 |
| 2 | 8.5s | 5,342 | $0.000586 | 3 | 3 | 10 |
| 3 | 15.8s | 6,085 | $0.000774 | 8 | 7 | 17 |
Method
- Model: @cf/qwen/qwen3-30b-a3b-fp8
- Provider: Cloudflare Workers AI
- Dataset: 60 synthetic English support messages
- Stages: topic clustering then report composition
- Deterministic safety scan before generation
- JSON mode with Qwen thinking disabled
Conclusions
- Topic grouping was stable across all three runs.
- Generated FAQ, risk-example, and action-asset counts were not stable enough to use without review.
- The average model inference estimate was below one tenth of one cent per 60-message analysis.
- The benchmark validates a technical workflow, not customer savings, production accuracy, or willingness to pay.
Limitations
- The dataset is synthetic and contains only 60 English-language support messages.
- The benchmark does not include human labels for topic purity or a complete red-risk recall score.
- Costs use Cloudflare model pricing available when the benchmark was published and may change.
- Latency reflects three requests from one deployment and is not a service-level guarantee.
- No real customer adoption, edit-rate, or time-saved evidence is included.