elevenlabs.providers.sgit.ai / experiments / Model A/B — latency, size and cost on one text
Model A/B — latency, size and cost on one text
Four models, one text, one voice, several runs each. The vendor's model table tells you the price and the character limit; it cannot tell you what your text sounds like or how long your connection takes to get it.
Your own full key, in your own browser, bounded only by your plan's monthly quota. That is pattern 0 with a ceiling: the narrow case where it is defensible — the key's owner, testing their own key, on their own device — and not a pattern to publish. No key ships in this page: the field below is empty until you fill it.
ElevenLabs offers no per-key spend limit for text to speech, so nothing here caps what a leaked key could cost you except the account's own quota. Use a key scoped in the dashboard to the endpoints this lab needs, and forget it when you are done. The version of this page with no key box at all is pattern three, and it does not exist yet.
1 · The text, and the models to race
2 · Results
| Model | Run | Latency | Bytes | Spoken length | Cost | Listen |
|---|
3 · Log
What this tests
The published differences between the models are price, character limit and a one-line character sketch vendor docs 5 Sep 2026 vendor docs 5 Sep 2026. The differences that decide a pipeline are:
- Latency under your own network, which is the number that determines whether a 16-scene render takes ten seconds or two minutes. Our only measurement is three samples of one model from one connection: 5.0–5.8 s verified 5 Sep 2026. That is not a distribution, and this lab exists to replace it with one.
- Spoken duration for identical text, which changes how much script fits a 60-second cut. Two models reading the same words at the same nominal speed do not produce the same length.
- Whether the expensive model is audibly better on this text. Half the price is half the price; if flash is indistinguishable on your workload, that is the finding.
Method
Same voice, same text, same speed, several runs per model — because one run per model times the weather, not the model. Runs are sequential, not parallel: parallel runs would measure the concurrency limit instead, which is a different lab with a different failure mode written, not run.
The cost estimate above the button is the whole spend, computed before anything is sent: characters × rate × runs × models.
What to write down
Latency mean and spread per model, spoken duration per model, and — the part no table can give you — whether you could tell them apart with your eyes closed. The Copy results as markdown button produces a table with a timestamp; paste it into the repository that cares about the answer.
Our own projection says v3 costs about what the incumbent provider costs and flash costs half projected. Whether the extra buys anything on a technical explainer is exactly the question this lab is for, and we have not answered it written, not run.