ReplyPilot

Open evidence

AI support benchmarks

Versioned technical tests with downloadable inputs, machine-readable results, explicit limitations, and no claim that synthetic performance equals customer value.

Published 2026-07-24

Qwen3 support analysis benchmark on 60 synthetic messages

Three independent ReplyPilot runs measuring topic stability, risk output, generated assets, latency, token usage, and estimated inference cost.

View benchmark

60

Synthetic messages

3/3

Topic agreement

12.0s

Average latency