elevenlabs.providers.sgit.ai / experiments / Sound effects and music
Sound effects and music
Two endpoints that are easy to overuse. They are here with a cost meter and a house rule attached, because the interesting question is not can it — it obviously can — but should it, under narration dense with numbers.
Your own full key, in your own browser, bounded only by your plan's monthly quota. That is pattern 0 with a ceiling: the narrow case where it is defensible — the key's owner, testing their own key, on their own device — and not a pattern to publish. No key ships in this page: the field below is empty until you fill it.
ElevenLabs offers no per-key spend limit for text to speech, so nothing here caps what a leaked key could cost you except the account's own quota. Use a key scoped in the dashboard to the endpoints this lab needs, and forget it when you are done. The version of this page with no key box at all is pattern three, and it does not exist yet.
1 · Sound effects
A prompt, a duration and how literally to take you. Billed per minute, so the meter below is in seconds — a two-second sting is a fraction of a cent, and a two-second sting under a title slide is the only use of this endpoint that survived our own review.
2 · Music
Natural language in, a track out, three seconds to five minutes. Cleared for broad commercial use on paid plans — check the current terms yourself before publishing anything under a brand; this page is not legal advice and the terms are the vendor's to change.
3 · The house rule this lab exists to test
A two-second sting under the title slide, and nothing else. A bed under narration hurts intelligibility, and reels dense with numbers can least afford it. Mix a sting at −18 dB under the voice and fade it before the first word. If you disagree, this lab is where you produce the counter-example — generate both, listen on a phone speaker, and write down which one you would rather watch.
4 · Log
The numbers
Sound effects take a prompt, a duration from 0.1 to 30 seconds (omit it for automatic), a prompt-influence control where high means literal, and an optional seamless loop for ambience. $0.12 per minute — so a two-second sting is about $0.004 vendor docs 5 Sep 2026.
Music takes a natural-language prompt and a length from 3 seconds to 5 minutes. $0.15 per minute — a 30-second bed is about $0.075, which is the entire ElevenLabs spend behind this site's verified claims measured 5 Sep 2026. The vendor states that music from paid plans is cleared for broad commercial use; read the current terms yourself before publishing under a brand, because that is a licence question and it is theirs to change.
The rule these were evaluated against
A two-second sting under the title slide, and nothing else. A bed under narration hurts intelligibility, and every reel in the source estate is dense with numbers that a viewer has one chance to hear. If a sting is used, mix it about 18 dB under the voice and fade it before the first word — the render already knows when the first word starts, because the alignment says so.
That is a rule, not a finding. It was written from experience with a different medium and has never been A/B tested here written, not run. This lab is where somebody disproves it: generate a bed, mix it under a real narration line, listen on a phone speaker at half volume, and write down which version you would rather watch.
Where this fits the site's argument
Nowhere, and that is worth saying. These endpoints spend the same account quota as speech and carry the same credential story: a key scoped for narration does not need them, and a key that can reach them can spend the account's whole quota generating five-minute tracks vendor docs 5 Sep 2026. If you scope a key for a render pipeline, scope these out — and use the key-scope probe to check that you actually did.