ReplyPilot

Technical benchmark

Qwen3 support analysis benchmark on 60 synthetic messages

Three independent ReplyPilot runs measuring topic stability, risk output, generated assets, latency, token usage, and estimated inference cost.

Published 2026-07-24 · Model pricing and behavior may change

3

Independent runs

3/3

Topic agreement

12.0s

Average latency

$0.000655

Average inference

Direct result

All three runs returned the same seven topic names, volumes, risk levels, readiness score of 68, and “needs validation” decision. Generated FAQ, high-risk-example, and action-asset counts varied substantially, so those outputs still require review.

Stable topic result

TopicVolumeRisk
Shipping & tracking19yellow
Order edits & promos13yellow
Returns, refunds & replacements10yellow
Risk, legal & escalation8red
Product info & fit8green
Billing, invoices & subscriptions1yellow
Other support questions1yellow

Run-by-run measurements

RunLatencyTokensCostRisksFAQsAssets
111.7s5,540$0.000604307
28.5s5,342$0.0005863310
315.8s6,085$0.0007748717

Method

  • Model: @cf/qwen/qwen3-30b-a3b-fp8
  • Provider: Cloudflare Workers AI
  • Dataset: 60 synthetic English support messages
  • Stages: topic clustering then report composition
  • Deterministic safety scan before generation
  • JSON mode with Qwen thinking disabled

Conclusions

  • Topic grouping was stable across all three runs.
  • Generated FAQ, risk-example, and action-asset counts were not stable enough to use without review.
  • The average model inference estimate was below one tenth of one cent per 60-message analysis.
  • The benchmark validates a technical workflow, not customer savings, production accuracy, or willingness to pay.

Limitations

  • The dataset is synthetic and contains only 60 English-language support messages.
  • The benchmark does not include human labels for topic purity or a complete red-risk recall score.
  • Costs use Cloudflare model pricing available when the benchmark was published and may change.
  • Latency reflects three requests from one deployment and is not a service-level guarantee.
  • No real customer adoption, edit-rate, or time-saved evidence is included.